A Method and System for Ultrasonic Localization Microscopic Imaging Based on Joint Resolution Sensing and Visual Mamba
By constructing a high-fidelity dataset and utilizing a hybrid hierarchical feature pyramid fusion network and a visual Mamba diffusion reconstruction module, the imaging problem of ULM technology under different microbubble concentration conditions was solved, achieving efficient and accurate vascular imaging and supporting clinical applications.
Patent Information
- Application Number
- CN202510818550.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Existing ULM technology suffers from severe signal confusion in high-concentration microbubble scenarios, and data acquisition is time-consuming and costly in low-concentration scenarios, making it difficult to meet the needs of clinical promotion. Furthermore, the number of simulation datasets is limited, resulting in insufficient training and generalization capabilities.
An ultrasound-guided microscopic imaging method based on joint resolution perception and visual Mamba was adopted. By acquiring modal angiography data and combining it with blood flow velocity field simulation data to generate microbubble motion sequences, a high-fidelity dataset was constructed. A hybrid hierarchical feature pyramid fusion network was used for probability density estimation, and a high-resolution vascular image was generated by combining it with the visual Mamba diffusion reconstruction module.
It achieves a balance between imaging speed and accuracy under different microbubble concentrations, improves the model's generalization ability and imaging quality, reduces data acquisition time and cost, and supports large-scale, high-efficiency imaging.
Smart Images

Figure CN120672891B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of computer vision and deep learning technology, specifically relating to an ultrasonic positioning microscopic imaging method and system based on joint resolution perception and visual Mamba. Background Technology
[0002] Ultrasound localization microscopy (ULM) is a super-resolution vascular imaging technique based on microbubble contrast agents. It combines the penetrating power of ultrasound with the localization properties of microbubbles to achieve micron-level reconstruction of deep microvascular networks. Microbubbles, as one of the most commonly used ultrasound contrast agents, can be used for non-invasive imaging of various organs such as the brain and kidneys, and have been widely applied in disease diagnosis and clinical research. However, there is a significant contradiction between data acquisition and image quality in existing ULM technologies: high-concentration microbubbles can accelerate vascular filling and shorten acquisition time, but microbubble overlap leads to signal confusion, severely affecting the accuracy of vascular distribution reconstruction; low-concentration microbubbles, while beneficial for single-bubble localization and high-quality imaging, require a longer time to accumulate sufficient localization points, are prone to introducing tissue motion and respiratory artifacts, and increase time and storage costs, thus limiting the clinical application of ULM technology.
[0003] To address the aforementioned issues, various algorithmic improvements have been proposed in existing technologies. Traditional methods primarily utilize temporal-spatial filtering or sparse reconstruction techniques, separating overlapping microbubbles through Fourier transform or sparse recovery theory to enhance localization capabilities and reduce the number of sampling frames in high-concentration scenarios. For example, filtering methods based on diverse spatiotemporal flow characteristics can improve microbubble detection accuracy while maintaining imaging speed; while sparse reconstruction methods locate overlapping microbubbles through sparse recovery algorithms, reducing dependence on the number of ultrasound images. Deep learning-based methods, with optimized neural network structures at their core, significantly improve localization accuracy and data processing efficiency under high-concentration microbubble conditions. Among these, the scheme combining generative adversarial networks (GANs) and context-aware models can achieve efficient localization under high microbubble concentrations; probabilistic estimation models can predict the state and uncertainties of dense microbubbles; and the refinement strategy using sub-pixel convolutional neural networks further enhances localization performance. Furthermore, for low-concentration scenarios, conditional generative adversarial networks (cGANs) have also been used for super-resolution image reconstruction, combined with Gaussian fitting and deconvolution algorithms to improve image sharpness.
[0004] Although the above technologies have achieved certain results in their respective application scenarios, they still face the following shortcomings: First, the number of existing publicly available simulation datasets is limited, and they mostly rely on random microbubble motion simulation, lacking real dynamic features, resulting in insufficient model training and generalization capabilities; Second, microbubble aggregation is still serious in high-concentration scenarios, with overlapping microbubbles obscuring individual features, resulting in blurred imaging and decreased detection accuracy; Third, data acquisition in low-concentration scenarios is time-consuming and costly, making it difficult to meet the needs of large-scale and efficient imaging.
[0005] Therefore, there is an urgent need for a ULM technology that can balance imaging speed and accuracy under different microbubble concentration conditions, while constructing a high-fidelity simulation dataset to support the efficient training and robust application of deep learning models. Summary of the Invention
[0006] To address the aforementioned technical issues, this invention provides an ultrasound localization microscopy imaging method and system based on joint resolution perception and visual Mamba, aiming to balance microbubble concentration, imaging quality, and acquisition efficiency, and to provide strong support for the large-scale promotion of ULM in clinical and scientific research.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] An ultrasound localization microscopy imaging method based on joint resolution perception and visual Mamba includes the following steps:
[0009] Step S1: Acquire several modal angiography data and preprocess them, combine them with blood flow velocity field simulation data to generate microbubble motion sequences, and construct a dataset;
[0010] Step S2: Input the dataset into the hybrid hierarchical feature pyramid fusion network, perform probability density estimation of microbubble locations based on the real-time multi-scale density field estimation algorithm, introduce a resolution perception algorithm to enhance microbubble features, output a set of microbubble positioning coordinates, perform physical analysis and association on the set of microbubble positioning coordinates, and obtain several short trajectory images.
[0011] Step S3: Perform dynamic feature decoupling and time-frequency analysis on the short trajectory image, extract velocity component features and spatiotemporal context information of vascular structure and encode them, construct a dual-branch processing network based on visual Mamba to obtain images that reflect the contributions of forward motion and backward motion respectively, and finally generate a high-resolution vascular image through image fusion.
[0012] Furthermore, step S1 includes: acquiring several modal angiography data as basic input maps; constructing a high-fidelity vascular wall geometric model using adaptive threshold segmentation and three-dimensional topology reconstruction technology; calculating the optimal motion path of intravascular fluid based on the improved Dijkstra path tracking algorithm and fusing blood flow velocity field simulation data to generate a probability distribution heatmap of microbubble motion, thereby obtaining the possible motion trajectories of microbubbles; constructing a dynamic prompt word library based on clinical prior knowledge; applying a diffusion model constrained by vascular trajectory to generate data based on the prompt words, so that the possible motion trajectories become microbubble motion sequences that conform to biomechanical characteristics, thus completing data preparation.
[0013] Furthermore, step S1 specifically includes:
[0014] Step S11: Acquire basic angiography image data, and perform image vessel enhancement, tissue denoising and vessel segmentation operations on each image to obtain the corresponding vascular structure image;
[0015] Step S12: Apply a path tracing algorithm to each vascular structure image to extract the vascular connectivity structure and obtain blood flow velocity field simulation data; simulate the optimal blood flow path curve based on the microbubble motion physical model, and generate a probability distribution heatmap of microbubble motion by combining the blood flow velocity field simulation data, thereby obtaining the target motion path of the microbubbles. ;
[0016] Step S13: Construct a conditional cue vector and use it as the control condition for generating the diffusion model simulation graph;
[0017] Step S14: Input the conditional cue vector and the initial noise map into the backbone network of the diffusion model simulation graph generation model. The backbone network extracts semantic feature maps sequentially through its multiple levels, starting from the i-th level. The semantic feature graph is denoted as ;
[0018] Step S15: Introduce a trajectory control unit to utilize the target motion path of the microbubbles. Trajectory condition information is derived, and the i-th layer of the backbone network is... semantic feature map Modulation is performed, and the output trajectory-aware feature map is denoted as... ;
[0019] Step S16: Target motion path of the microbubbles As precise structural control conditions, the ControlNet module is used to extract the i-th layer of the backbone network from these structural control conditions. Corresponding path geometry and spatial constraint features ;Will Trajectory-aware feature maps of corresponding levels Perform deep fusion to generate intermediate fused image features ;
[0020] Step S17: Fuse image features at each level After upsampling and decoding, the edge is generated. Microbubble simulation images that are moving and meet the conditional cue vectors ;
[0021] Step S18: Process the generated microbubble simulation image Authenticity testing and format standardization are performed to obtain the final simulation image dataset.
[0022] Furthermore, step S2 specifically includes:
[0023] Step S21: Normalize the obtained simulation image dataset to obtain the global central tendency analysis measure and spatial variability quantification value of the images, which serves as the normalized image set. ;
[0024] Step S22: Normalize the image After feature extraction by the detection encoding module, the data is then fed into the detection decoding module and the density estimation decoding module, respectively. The detection decoding module outputs the microbubble location detection results, and the density estimation decoding module outputs the predicted heatmap. The predicted heatmap It reflects the dense distribution characteristics of microbubbles in the prediction space of the image, which can be used for subsequent supervised learning;
[0025] Step S23: Process the input image The pixel distribution was statistically tested, and the Kolmogorov-Smirnov dynamic resolution test method was used to calculate the empirical cumulative distribution function of the current image region. Compared with the preset standard distribution Maximum deviation between :
[0026] (1)
[0027] In formula (1), This represents the minimum value among all upper bounds of the set, based on the stated statistic. The range of values determines the current image or image region into high-complexity, medium-complexity, and low-complexity regions. For different complexity regions, the corresponding feature connection paths between the encoding and decoding modules are dynamically activated.
[0028] Furthermore, in step S22, the detection encoding module includes multiple feature extraction unit structures, which include convolutional units of convolutional layers, batch normalization layers and activation functions cascaded in sequence, fusion units for feature fusion, and spatial compression downsampling units for spatial information compression and downsampling.
[0029] The detection decoding module receives features output by the detection encoding module and gradually integrates multi-scale feature information to generate a detection result. Each set of detection results includes the center coordinates of the detected microbubble targets. ),width ,high and the corresponding confidence score , This represents the number of targets detected in the current image.
[0030] The density estimation decoding module employs a layer-by-layer upsampling mechanism and utilizes skip connections between layers to fuse corresponding layer features from the detection encoding module, ultimately outputting a predicted heatmap. .
[0031] Furthermore, step 2 also includes:
[0032] Step S24: Use the mean squared error loss function to predict the heatmap. Compared with the reference density map Constraints are applied between the elements, and density loss is calculated. As shown in formula (2):
[0033] (2)
[0034] In formula (2) and These represent the observed results and the true values, respectively, with H and W representing the pixel width and length of the image, respectively.
[0035] The detection loss is calculated based on the detection results output by the detection decoding module. Construct the overall loss function As shown in formula (3):
[0036] (3)
[0037] In formula (3) and These are preset weight hyperparameters used to balance the contributions of the detection task and the density regression task in the model learning process;
[0038] The model is trained by minimizing the overall loss function.
[0039] Furthermore, step 3 includes:
[0040] Step S31: Receive the short trajectory image data generated in step S2. The analysis and decoupling were performed to extract shared vascular structural features. and feature maps representing the forward motion of microbubbles. and characteristic maps representing the backward motion of microbubbles ;
[0041] Step S32: Extract the shared vascular structure features Respectively related to forward motion characteristics and backward motion characteristics The forward and backward feature flows are obtained by combining them;
[0042] Step S33: Construct two parallel SSM processing backbones with the same structure but partially shared parameters to process the forward and backward feature flows respectively;
[0043] Step S34: After each SSM processing backbone, integrate a residual noise reduction unit. The signals of the forward and backward characteristic flows are purified;
[0044] Step S35: Decode the purified feature maps obtained from the two parallel streams respectively and finally fuse them to obtain the final high-resolution output image. .
[0045] On the other hand, the present invention provides an ultrasound localization microscopic imaging system based on joint resolution perception and visual Mamba, comprising:
[0046] The simulation data generation unit is used to acquire several modal angiography data and preprocess them, and combine them with blood flow velocity field simulation data to generate microbubble motion sequences and construct a dataset.
[0047] The resolution sensing unit is used to input the dataset into the hybrid hierarchical feature pyramid fusion network, perform probability density estimation of microbubble locations based on the real-time multi-scale density field estimation algorithm, introduce the resolution sensing algorithm to enhance microbubble features, and output a set of microbubble positioning coordinates.
[0048] The Visual Mamba diffusion and reconstruction module is used to perform physical analysis and association on the microbubble positioning coordinate set to obtain several short trajectory images. Dynamic feature decoupling and time-frequency analysis are performed on the short trajectory images to extract velocity component features and spatiotemporal context information of vascular structure and encode them. A dual-branch processing network is constructed based on the Visual Mamba to obtain images that reflect the forward motion contribution and the backward motion contribution respectively. Finally, a high-resolution vascular image is generated through image fusion.
[0049] Thirdly, the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned ultrasonic localization microscopic imaging method based on joint resolution perception and visual Mamba.
[0050] Fourthly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned ultrasound localization microscopy imaging method based on joint resolution perception and visual Mamba, and periodically perform multi-cue continuous learning and updating of the aforementioned collaborative model base model using newly captured data.
[0051] The beneficial effects of this invention are as follows:
[0052] 1. This invention designs a vascular morphology-guided simulation data module based on multi-resolution perception, which realizes the simulation of the movement trajectory of microbubbles in real vascular networks at three resolution levels: coarse, medium, and fine. The morphological cues are fused using ControlNet, which makes the simulation data have high fidelity and diversity, significantly enhancing the coverage and realism of the training set. This provides key data support for subsequent training of generative models for vascular imaging, thereby improving the model's generalization ability under different vascular structures and imaging conditions.
[0053] 2. This invention designs a multi-task convolutional neural network that jointly performs target detection, density estimation, and resolution adaptation. By introducing a Feature Pyramid Network (FPN), a self-attention module, and a dynamic skip connection strategy based on the Kolmogorov-Smirnov test, it achieves refined processing of microbubble signals and adaptive optimization of the network structure. This network can output high-precision pixel-level location, local density, and resolution quality assessment information of microbubbles within the same framework. These accurate outputs collectively constitute the prior knowledge crucial for subsequent hemodynamic analysis and high-quality vascular reconstruction, laying a solid foundation for accurate imaging by the visual Mamba diffusion reconstruction module and improving the overall system performance.
[0054] 3. This invention designs a visual Mamba diffusion reconstruction module, which realizes the rapid generation of high-resolution vascular images by using a small number of localization points (precise output from step S2) and an extremely short sampling time as input, through an iterative denoising process combined with vascular prior inference, and utilizing the iterative denoising characteristics of the diffusion model and the multi-directional scanning mechanism. This module significantly removes artifacts and improves the signal-to-noise ratio while preserving the details of microvascular branches, providing high-quality imaging results for non-invasive clinical microvascular detection. Due to its efficient reconstruction method, it reduces data acquisition time and subsequent storage and computing costs while ensuring image quality.
[0055] 4. This invention integrates a simulation data module, a multi-task convolutional neural network, and a visual Mamba diffusion reconstruction module to construct a complete ultrasound localization microscopy imaging solution from data source to high-quality imaging. It not only improves the imaging accuracy and speed in complex scenarios, but also effectively addresses the challenges of traditional ULM technology in data acquisition, processing efficiency, and imaging quality. This provides strong technical support for the early and accurate diagnosis of clinical diseases, the evaluation of treatment effects, and related medical research, demonstrating significant clinical application value and potential. Attached Figure Description
[0056] Figure 1 This is a schematic diagram of the ultrasonic localization microscopic imaging method based on joint resolution perception and visual Mamba of the present invention.
[0057] Figure 2 This is a schematic diagram of the dataset construction process of this invention;
[0058] Figure 3 This is a schematic diagram of the detection algorithm based on multi-resolution sensing and real-time multi-scale density field estimation of the present invention.
[0059] Figure 4 This is a schematic diagram of the Mamba super-resolution generation network based on physical constraint prior reasoning in this invention.
[0060] Figure 5 This is a structural diagram of the ultrasonic positioning microscopic imaging system based on joint resolution perception and visual Mamba of the present invention. Detailed Implementation
[0061] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0062] like Figure 1 As shown, the present invention provides an ultrasound localization microscopic imaging method based on joint resolution perception and visual Mamba, comprising the following steps:
[0063] Step S1: Acquire several modal angiography data as basic input images, and construct a high-fidelity vessel wall geometric model using adaptive threshold segmentation and 3D topology reconstruction techniques. Calculate the optimal fluid path within the vessel based on an improved Dijkstra path tracking algorithm and fuse blood flow velocity field simulation data to generate a probability distribution heatmap of microbubble motion, thereby generating possible microbubble trajectories. Adjust hyperparameters by constructing a dynamic prompting lexicon based on prior clinical knowledge. On this basis, a diffusion model constrained by vessel trajectory is applied to generate data according to the prompts, making it a microbubble motion sequence that conforms to biomechanical characteristics. The output includes the original vessel topology, microbubble motion trajectory parameters, etc., and is used for dataset preparation. Specifically, such as... Figure 2 As shown, it includes:
[0064] Step S11: Acquire basic vascular image dataset For each image Perform vascular enhancement, tissue denoising, and vascular segmentation operations to obtain the corresponding vascular structure image. The vascular regions are extracted using a segmentation network and represented as binary images, resulting in a set of structured vascular images. ;
[0065] Step S12: For each vascular structure image A path tracing algorithm was applied to extract the vascular connectivity structure and construct its centerline trajectory model to obtain simulated blood flow velocity field data, which included intravascular flow velocity information. On the other hand, based on a microbubble motion physical model, simulated blood flow path curves were obtained through spline interpolation. To more accurately predict microbubble behavior, the obtained initial path curves were combined with the simulated blood flow velocity field data for optimization calculations to determine and generate the optimal target motion path of the microbubbles within the blood vessel. Based on this, a probability distribution map of microbubble motion can be further analyzed and constructed according to the optimal motion path, possible dynamic changes in the flow field, and model uncertainties. This probability distribution map can statistically characterize the probability of microbubbles appearing in a specific region or the distribution range of their motion trajectories, providing more comprehensive information for subsequent analysis.
[0066] Step S13: Control the physiological semantic features of the generated image through text keywords, such as lesion location and blood flow status. Adjust the dynamic attributes of the simulated image using time control parameters, such as the degree of microbubble movement in the frame sequence. Construct a conditional cue vector as the control condition for the generative model.
[0067] Step S14: Input the cue vector and the initial noise map into the backbone network of the vessel trajectory control generation network. The backbone network extracts semantic feature maps sequentially through its multiple levels. The semantic feature map of the i-th level is denoted as... ;
[0068] Step S15: Introduce the trajectory control unit, which utilizes the data generated in step S12. Trajectory condition information is derived, and specific layers of the backbone network are analyzed. Feature map Modulation is performed to incorporate path guidance information, outputting a trajectory-aware feature map, denoted as... ;
[0069] Step S16: Design a ControlNet module to generate the optimal motion path in step S12. This serves as a precise structural control condition. The module extracts the relevant information from this structural control condition for specific layers of the backbone network. Corresponding path geometry and spatial constraint features ; and these features Trajectory-aware feature maps of the corresponding level generated in step S15 Deep fusion is performed to ultimately generate an intermediate fused image feature. The purpose of this fusion operation is to combine precise path geometry constraints with features that already contain preliminary path guidance;
[0070] Step S17: Merge the features at each level After subsequent upsampling and decoding processes, the final edge-to-edge image is generated. Simulated images of moving microbubbles that meet preset prompts. ;
[0071] Step S18: Process the generated microbubble simulation image Realism detection and format standardization are performed. This includes determining whether microbubbles are inside the vascular pathway and whether they drift or exhibit abnormal generation. Finally, the output image format, size, and frame number are standardized to obtain the final simulation image dataset.
[0072] Step S2: Design a hybrid hierarchical feature pyramid fusion network. First, introduce a feature enhancement module based on cross-dimensional attention to simultaneously enhance microbubble feature representations in both channel and spatial dimensions to suppress interference from complex backgrounds. Simultaneously, employ a deformation adaptive network to adaptively adjust the microbubble sampling positions, improving edge localization accuracy. These two designs enhance feature representation and improve localization performance. Furthermore, deploy a dynamic link resolution-aware strategy between the encoding and detection / decoding modules. This adaptively adjusts skip connections and dynamically adjusts feature fusion weights based on data resolution. Additionally, a real-time multi-scale density field estimation algorithm is simultaneously constructed to model the spatial distribution probability of microbubbles to handle overlap issues in high-density scenes. Simultaneously, statistically driven self-supervised training of microbubble positions is implemented using prior microbubble density distribution to address microbubble occlusion and mutual interference in high-density scenes, ultimately providing pixel-level precise positional constraints for the network. Specifically, as shown... Figure 3 As shown:
[0073] Step S21: Input image data and perform preprocessing. The acquired raw ultrasound image dataset is denoted as: ,in, Indicates the first Zhang is an original grayscale image with height and width dimensions of [value missing]. The images in the final dataset obtained in S18 are normalized to obtain global central tendency analysis measures and spatial variability quantification values, resulting in a normalized image set. ;
[0074] Step S22: Normalize the image obtained in S21 The input is fed into the detection encoding module for feature extraction. This encoding module contains multiple feature extraction unit structures, each consisting of three parts: a convolutional unit composed of a convolutional layer, a batch normalization layer, and an activation function; a fusion unit for feature fusion; and a spatial compression and downsampling unit for spatial information compression and downsampling.
[0075] The features output by the aforementioned encoding module are then fed into the detection decoding module. The detection decoding module, for example, employs a two-layer upsampling structure and a feature fusion unit to progressively integrate multi-scale feature information to generate the detection result. Each set of results includes the center coordinates of the detected microbubble targets. ),width ,high and confidence score , This represents the number of targets detected in the current image.
[0076] Simultaneously, the features output by the aforementioned encoding module are also input to the density estimation decoding module. This module employs a layer-by-layer upsampling mechanism and utilizes skip connections between layers to fuse features from corresponding layers of the encoding module. Each upsampling layer contains a density estimation unit consisting only of convolutional layers. The density estimation decoding module ultimately outputs a heatmap. This heatmap reflects the dense distribution characteristics of microbubbles in the predicted space of the image, which can be used for subsequent supervised learning.
[0077] Step S23: Process the input image The pixel distribution was statistically tested, and the Kolmogorov-Smirnov dynamic resolution test method was used to calculate the empirical cumulative distribution function of the current image region. Compared with the preset standard distribution Maximum deviation between :
[0078] (1)
[0079] In formula (1), This represents the minimum value among all upper bounds of the set. Based on the stated statistic... The range of values for this parameter divides the current image or image region into high-complexity, medium-complexity, and low-complexity regions. For regions of different complexity, corresponding feature connection paths between the encoding and decoding modules (including the detection decoding module and the density estimation decoding module) are dynamically activated: high-complexity regions prioritize and strengthen connections from deep features to the density estimation branch; medium-complexity regions enable cross-level skip fusion paths to balance global and local information; and low-complexity regions focus on connections from shallow features to the detection branch. This mechanism enables dynamic adjustment and adaptive reconstruction of the model's connection structure to optimize feature utilization efficiency in different scenarios.
[0080] Step S24: To achieve self-supervised training of microbubble locations, a supervision term based on the density heatmap is introduced. The mean squared error loss function is used to predict the heatmap. Compare with a real density map or a reference density map generated by methods such as kernel density estimation. Constraints are applied between the elements, and density loss is calculated. As shown in formula (2):
[0081] (2)
[0082] In formula (2) and H and W represent the observed result and the true value, respectively, and represent the pixel width and length of the image, respectively. The detection loss is calculated by combining the detection results output by the detection decoding module. (Including classification loss and localization loss), construct the overall loss function. As shown in formula (3):
[0083] (3)
[0084] In formula (3) and The preset weight hyperparameters are used to balance the contributions of the detection and density regression tasks in the model learning process. Model training is performed by minimizing the overall loss function, where the density loss term provides pixel-level constraints for the precise location learning of microbubbles, helping to handle microbubble occlusion and mutual interference problems in high-density scenes.
[0085] Step S3: Dynamic feature decoupling and time-frequency analysis are performed on the short trajectory images representing microbubble motion generated in Step S2 to separate and extract velocity components with different dynamic characteristics and corresponding spatiotemporal context information of the vascular structure. Subsequently, this separated structural information is combined with motion information in each direction and deeply encoded. Based on this, two parallel core processing branches are constructed: one branch specifically handles forward motion-related features, and the other handles backward motion-related features. Both branches utilize shared structural context information. Each processing branch uses a State-Space Model (SSM) unit specifically optimized for two-dimensional image features and integrating an eight-directional scanning mechanism as its backbone. It performs deep contextual relationship modeling and feature purification on its corresponding directional motion feature block sequences and integrates a residual denoising unit. Each branch independently decodes and generates high-resolution images reflecting forward and backward motion contributions respectively. Finally, these two directional contribution images are superimposed through image fusion to form a final high-resolution image that clearly displays both vascular structure and bidirectional blood flow information. The training of the entire model is guided by optimizing the combined loss function. Figure 4 As shown, it specifically includes:
[0086] Step S31: Receive the short trajectory image data generated in step S2 This segment of trajectory image data encodes the short-term motion information of the microbubbles. For the aforementioned... The application uses a trajectory analysis and motion direction separation module, which decouples and extracts key vascular structural features from trajectory information. and feature maps characterizing the forward motion of microbubbles. and characteristic maps representing the backward motion of microbubbles ;
[0087] Step S32: To construct two parallel processing branches, firstly, the shared structural features extracted in step S31 are... Respectively related to forward motion characteristics and backward motion characteristics Combine:
[0088] 1. Forward feature flow construction: and By concatenating along the channel dimension, combined features are obtained. Subsequently, a depthwise separable convolution module was used. right Preliminary enhancements and channel adjustments are performed to obtain the enhanced feature map of the forward flow. .
[0089] 2. Backward feature flow construction: and By concatenating along the channel dimension, combined features are obtained. Subsequently, a depthwise separable convolution module was used. right Preliminary enhancements and channel adjustments were performed to obtain the enhanced feature map of the backflow. ;in The number of characteristic channels for each processing stream.
[0090] Subsequently, respectively and Disassemble into a series of non-overlapping, fixed-size components. The feature block sequences are denoted as follows: and , which is used as a structured input for subsequent parallel processing.
[0091] Step S33: Construct two parallel SSM processing backbones with identical structures but partially shared parameters to process the forward and backward feature flows respectively:
[0092] 1. Forward flow processing: processing the feature block sequence Input to the first SSM backbone.
[0093] 2. Backflow processing: Processing the feature block sequence... Input to the second SSM backbone.
[0094] Each SSM backbone employs an eight-directional scanning mechanism optimized for two-dimensional image features. The feature block sequence is rearranged according to eight main scanning directions (horizontal forward, horizontal reverse, vertical forward, vertical reverse, main diagonal forward, main diagonal reverse, secondary diagonal forward, and secondary diagonal reverse), forming eight sets of one-dimensional sequences in different directions. Each set of directional sequences The inputs are independently fed into a series of stacked SSM cells. Each SSM cell follows the core state propagation equation. , Perform calculations, its parameter matrix Input features The function, and for different scanning directions It has an independent parameter set. The SiLU nonlinear activation function is applied to the output. The SSM output feature representations from each scanning direction are then concatenated along the channel dimension and deeply fused using a convolutional network module enhanced by an attention mechanism to obtain the enhanced feature representations of the forward flow. Enhanced feature representation of backflow ;
[0095] Step S34: After the SSM backbone of each parallel processing stream, integrate residual noise reduction units respectively. .
[0096] 1. Forward flow denoising: Denoising the feature representation after forward flow fusion. (Reorganized into feature map form) Applying a noise reduction unit, we obtain :
[0097] (4)
[0098] In formula (4) It is the feature representation after forward flow fusion, reorganized into feature map form. This indicates normalization along the channel dimension, with the aim of reducing internal covariate offset.
[0099] 2. Backflow Denoising: Feature Representation After Backflow Fusion (Reorganized into feature map form) Applying a noise reduction unit, we obtain The calculation method is the same as that of formula (4). This step aims to purify the signals of the two streams separately, retaining valuable structural and directional dynamic characteristics.
[0100] Step S35: Decode the purified feature maps obtained from the two parallel streams separately and then fuse them.
[0101] 1. Forward image decoding: Through depthwise separable convolutional modules and subsequent decoding network (Using sub-pixel convolutional upsampling and feature fusion), a high-resolution image reflecting the contribution of forward motion is reconstructed. .
[0102] 2. Backward image decoding: Through depthwise separable convolutional modules and subsequent decoding network Reconstruct a high-resolution image reflecting the contribution of backward motion. .
[0103] 3. Final image fusion: The resulting forward contribution images are then fused together. and backward contribution image The final high-resolution output image is obtained by fusing pixels-level images through addition operations. This image overlay operation combines flow information from both directions, providing a complete depiction of the vascular structure.
[0104] To optimize the parameters of the entire model (including all components of the two parallel branches), a composite loss function is designed. This function primarily operates on the final fused image. The loss function includes: pixel-level reconstruction loss. : Measure the final fused image Compared with real high-resolution reference images The difference between them is analyzed using L1 loss. (Perceptual loss) : Utilizing the feature extraction capabilities of pre-trained deep networks, comparison and Similarity in the feature space is used to enhance visual realism. Overall loss function. :
[0105] (5)
[0106] In formula (5) Let be the weight hyperparameters for each loss term. Minimize ... The model is trained end-to-end. This architecture is designed to enable the network to learn and reconstruct vascular details related to blood flow in a specific direction more precisely by processing features in different directions of motion separately, and finally to reveal the complete hemodynamic situation through fusion.
[0107] This invention supports an ultrasound-guided microscopy imaging system based on joint resolution perception and visual Mamba. Through multi-resolution simulated image generation and S2 design network, it achieves synergistic iteration of high-density microbubble localization and deep vascular super-resolution reconstruction, improving the spatial resolution and signal-to-noise ratio of imaging, enhancing real-time processing capabilities and system robustness, reducing hardware deployment costs, and facilitating the clinical application and rapid promotion of ultrasound-guided microscopy imaging technology.
[0108] like Figure 5 As shown, this invention provides an ultrasound localization microscopic imaging system based on joint resolution sensing and visual Mamba, comprising:
[0109] The simulation data generation module 51 is used to receive ultrasound or optical imaging sequences of the target vascular region, perform pixel-level segmentation and three-dimensional reconstruction of the sequences based on a multi-resolution perception strategy, and extract vascular structure models at three levels: coarse, medium, and fine. The module simulates the movement trajectory of microbubbles at each level, and inputs the trajectory information and vascular morphology cues together into a diffusion generation model guided by ControlNet morphology to generate multi-scale, high-fidelity microbubble simulation image sequences, providing a realistic and diverse ULM dataset for subsequent network training.
[0110] The resolution perception module 52 loads the multi-scale microbubble simulation image sequence and original ultrasound frames output by the simulation data generation module 51, and calls a quantized, deployed multi-task perceptual convolutional neural network model on a hardware-accelerated edge device to perform joint inference on the input data. This network utilizes a feature pyramid network (FPN) and a self-attention mechanism to achieve multi-scale feature fusion and complex background interference suppression, thereby accurately locating microbubble targets. Simultaneously, pixel-level probability loss is introduced to constrain high-density overlapping microbubble regions to improve density estimation accuracy in dense scenes. Furthermore, based on the Kolmogorov–Smirnov test, the cross-layer skip connection structure is dynamically adjusted, enabling the network to adaptively allocate shallow detection, cross-layer fusion, and deep density branches among regions of different image complexity, thereby optimizing inference efficiency while ensuring high detection accuracy. After the collaborative processing of these multiple tasks, the resolution perception module 52 outputs multi-resolution microbubble localization points and their confidence scores, providing accurate localization information for subsequent visual Mamba diffusion reconstruction.
[0111] The visual Mamba diffusion and reconstruction module 53 is used to input the multi-resolution localization points output by the resolution perception module 52 and the corresponding original ultrasound frames into the visual Mamba diffusion reconstruction network. The network first maps the localization points to several non-overlapping feature blocks through a feature splitting layer, and then embeds depthwise separable convolutions to enhance the perception of local details. In the diffusion backbone, Mamba blocks in the horizontal, vertical and diagonal directions are executed alternately to capture global structural information, and structural consistency is maintained during the iterative noise reduction process through conditional layer normalization and vascular prior inference. Finally, the features optimized by multi-layer Mamba diffusion and noise reduction are reconstructed by the decoder into a high signal-to-noise ratio and high resolution deep vascular imaging result.
[0112] Through the coordinated operation of the three modules mentioned above, the system of the present invention can generate large-scale and diverse simulation training data, realize real-time and efficient microbubble detection and high-precision vascular imaging reconstruction, and significantly improve the resolution and signal-to-noise ratio of localization microscopic imaging.
[0113] Thirdly, the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned ultrasonic localization microscopic imaging method based on joint resolution perception and visual Mamba.
[0114] Fourthly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned ultrasound localization microscopy imaging method based on joint resolution perception and visual Mamba, and periodically perform multi-cue continuous learning and updating of the aforementioned collaborative model base model using newly captured data.
[0115] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for ultrasound localization microscopy imaging based on joint resolution perception and visual Mamba, characterized in that, Includes the following steps: Step S1: Acquire several modal angiography data and preprocess them, combine them with blood flow velocity field simulation data to generate microbubble motion sequences, and construct a dataset; Step S2: Input the dataset into a hybrid hierarchical feature pyramid fusion network, perform probability density estimation of microbubble locations based on a real-time multi-scale density field estimation algorithm, introduce a resolution-aware algorithm to enhance microbubble features, output a set of microbubble location coordinates, and perform physical analysis and correlation on the set of microbubble location coordinates to obtain several short trajectory images; including: Step S21: Normalize the obtained simulation image dataset to obtain the global central tendency analysis measure and spatial variability quantification value of the images, which serves as the normalized image set. ; Step S22: Normalize the image After feature extraction by the detection encoding module, the data is then fed into the detection decoding module and the density estimation decoding module, respectively. The detection decoding module outputs the microbubble location detection results, and the density estimation decoding module outputs the predicted heatmap. The predicted heatmap It reflects the dense distribution characteristics of microbubbles in the prediction space of the image, which can be used for subsequent supervised learning; Step S23: Process the input image The pixel distribution was statistically tested, and the Kolmogorov-Smirnov dynamic resolution test method was used to calculate the empirical cumulative distribution function of the current image region. Compared with the preset standard distribution Maximum deviation between : (1) In formula (1), This represents the minimum value among all upper bounds of the set, based on the statistic. The range of values is used to divide the current image or image region into high-complexity, medium-complexity and low-complexity regions. For different complexity regions, the corresponding feature connection paths between the encoding module and the decoding module are dynamically activated. Step S3: Perform dynamic feature decoupling and time-frequency analysis on the short trajectory image, extract velocity component features and spatiotemporal context information of vascular structure and encode them, construct a dual-branch processing network based on visual Mamba to obtain images reflecting the contributions of forward motion and backward motion respectively, and finally generate a high-resolution vascular image through image fusion; including: Step S31: Receive the short trajectory image data generated in step S2. The analysis and decoupling were performed to extract shared vascular structural features. and feature maps representing the forward motion of microbubbles. and characteristic maps representing the backward motion of microbubbles ; Step S32: Extract the shared vascular structure features Respectively related to forward motion characteristics and backward motion characteristics The forward and backward feature flows are obtained by combining them; Step S33: Construct two parallel SSM processing backbones with the same structure but partially shared parameters to process the forward and backward feature flows respectively; Step S34: After each SSM processing backbone, integrate a residual noise reduction unit. The signals of the forward and backward characteristic flows are purified; Step S35: Decode the purified feature maps obtained from the two parallel streams respectively and finally fuse them to obtain the final high-resolution output image. .
2. The ultrasonic localization microscopic imaging method based on joint resolution perception and visual Mamba as described in claim 1, characterized in that, Step S1 includes: acquiring several modal angiography data as basic input maps; constructing a high-fidelity vascular wall geometric model using adaptive threshold segmentation and three-dimensional topology reconstruction technology; calculating the optimal motion path of intravascular fluid based on the improved Dijkstra path tracking algorithm and fusing blood flow velocity field simulation data to generate a probability distribution heatmap of microbubble motion, thereby obtaining the possible motion trajectories of microbubbles; constructing a dynamic prompt word library based on clinical prior knowledge; applying a diffusion model constrained by vascular trajectory to generate data based on the prompt words, so that the possible motion trajectories become microbubble motion sequences that conform to biomechanical characteristics, thus completing data preparation.
3. The ultrasonic localization microscopic imaging method based on joint resolution perception and visual Mamba as described in claim 2, characterized in that, Step S1 specifically includes: Step S11: Acquire basic angiography image data, and perform image vessel enhancement, tissue denoising and vessel segmentation operations on each image to obtain the corresponding vascular structure image; Step S12: Apply a path tracing algorithm to each vascular structure image to extract the vascular connectivity structure and obtain blood flow velocity field simulation data; simulate the optimal blood flow path curve based on the microbubble motion physical model, and generate a probability distribution heatmap of microbubble motion by combining the blood flow velocity field simulation data, thereby obtaining the target motion path of the microbubbles. ; Step S13: Construct a conditional cue vector and use it as the control condition for generating the diffusion model simulation graph; Step S14: Input the conditional cue vector and the initial noise map into the backbone network of the diffusion model simulation graph generation model. The backbone network extracts semantic feature maps sequentially through its multiple levels, starting from the i-th level. The semantic feature graph is denoted as ; Step S15: Introduce a trajectory control unit to utilize the target motion path of the microbubbles. Trajectory condition information is derived, and the i-th layer of the backbone network is... semantic feature map Modulation is performed, and the output trajectory-aware feature map is denoted as... ; Step S16: Target motion path of the microbubbles As precise structural control conditions, the ControlNet module is used to extract the i-th layer of the backbone network from these structural control conditions. Corresponding path geometry and spatial constraint features ;Will Trajectory-aware feature maps of corresponding levels Perform deep fusion to generate intermediate fused image features ; Step S17: Fuse image features at each level After upsampling and decoding, the edge is generated. Microbubble simulation images that are moving and meet the conditional cue vectors ; Step S18: Process the generated microbubble simulation image Authenticity testing and format standardization are performed to obtain the final simulation image dataset.
4. The ultrasonic localization microscopic imaging method based on joint resolution perception and visual Mamba as described in claim 1, characterized in that, In step S22, the detection encoding module includes multiple feature extraction unit structures, which include convolutional units of convolutional layers, batch normalization layers and activation functions cascaded in sequence, fusion units for feature fusion, and spatial compression downsampling units for spatial information compression and downsampling. The detection decoding module receives features output by the detection encoding module and gradually integrates multi-scale feature information to generate a detection result. Each set of detection results includes the center coordinates of the detected microbubble targets. ),width ,high and the corresponding confidence score , This represents the number of targets detected in the current image. The density estimation decoding module employs a layer-by-layer upsampling mechanism and utilizes skip connections between layers to fuse corresponding layer features from the detection encoding module, ultimately outputting a predicted heatmap. .
5. The ultrasound localization microscopy imaging method based on joint resolution perception and visual Mamba as described in claim 1, characterized in that, Step 2 also includes: Step S24: Use the mean squared error loss function to predict the heatmap. Compared with the reference density map Constraints are applied between the elements, and density loss is calculated. As shown in formula (2): (2) In formula (2) and These represent the observed results and the true values, respectively, with H and W representing the pixel width and length of the image, respectively. The detection loss is calculated based on the detection results output by the detection decoding module. Construct the overall loss function As shown in formula (3): (3) In formula (3) and These are preset weight hyperparameters used to balance the contributions of the detection task and the density regression task in the model learning process; The model is trained by minimizing the overall loss function.
6. An ultrasonic positioning microscopic imaging system based on joint resolution perception and visual Mamba, applied to the method described in any one of 1-5, characterized in that, include: The simulation data generation unit is used to acquire several modal angiography data and preprocess them, and combine them with blood flow velocity field simulation data to generate microbubble motion sequences and construct a dataset. The resolution sensing unit is used to input the dataset into the hybrid hierarchical feature pyramid fusion network, perform probability density estimation of microbubble locations based on the real-time multi-scale density field estimation algorithm, introduce the resolution sensing algorithm to enhance microbubble features, and output a set of microbubble positioning coordinates. The Visual Mamba diffusion and reconstruction module is used to perform physical analysis and association on the microbubble positioning coordinate set to obtain several short trajectory images. Dynamic feature decoupling and time-frequency analysis are performed on the short trajectory images to extract velocity component features and spatiotemporal context information of vascular structure and encode them. A dual-branch processing network is constructed based on the Visual Mamba to obtain images that reflect the forward motion contribution and the backward motion contribution respectively. Finally, a high-resolution vascular image is generated through image fusion.
7. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the ultrasound localization microscopy imaging method based on joint resolution perception and visual Mamba as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed by a processor, enable the processor to implement the ultrasound localization microscopy imaging method based on joint resolution perception and visual Mamba as described in any one of claims 1-5.
Citation Information
Patent Citations
Signal processing method and system for ultrasonic micro blood flow imaging
CN114533122A
Quantitative ultrasonic positioning microscopic imaging method based on attention learning mechanism
CN114557719A