Ultrasonic positioning microscopic imaging method and system based on joint resolution perception and vision Mangbar

By combining resolution perception with the ultrasound localization microscopy imaging method of Visual Mamba, the imaging problem of ULM technology in high and low concentration scenarios is solved, and efficient and accurate vascular imaging is achieved to support clinical applications.

CN120672891AActive Publication Date: 2025-09-19CAPITAL NORMAL UNIVERSITY

Patent Information

Application Number
CN202510818550.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-19
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The existing ULM technology suffers from severe signal confusion in high-concentration microbubble scenarios, and the data acquisition time is long and the cost is high in low-concentration scenarios, making it difficult to meet the needs of clinical promotion. In addition, the number of simulation data sets is limited, and the training and generalization capabilities are insufficient.

Method used

An ultrasound localization microscopy imaging method based on joint resolution perception and visual Mamba is adopted. The microbubble position probability density is estimated through a hybrid hierarchical feature pyramid fusion network. Combined with the visual Mamba diffusion reconstruction module, a high-fidelity simulation dataset is constructed. A multi-task convolutional neural network is designed for microbubble feature processing to achieve efficient and accurate vascular imaging.

Benefits of technology

By balancing imaging speed and accuracy under different microbubble concentration conditions, the model's generalization ability is improved, data acquisition time and cost are reduced, and high-quality microvascular images are generated to support clinical disease diagnosis and research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672891A_ABST
    Figure CN120672891A_ABST
Patent Text Reader

Abstract

The invention discloses an ultrasonic positioning microscopic imaging method and system based on joint resolution perception and vision Mangbar, and the method specifically comprises the steps: obtaining and preprocessing a plurality of modal angiography data, generating a microbubble motion sequence in combination with blood flow velocity field simulation data, and constructing a data set; inputting the data set into a mixed hierarchical feature pyramid fusion network, carrying out probability density estimation on a microbubble position based on a real-time multi-scale density field estimation algorithm, and introducing a resolution perception algorithm to enhance microbubble features to obtain a plurality of short-track images; performing dynamic feature decoupling and time-frequency analysis on the short-track image, extracting and encoding velocity component features and spatial-temporal context information of a vascular structure, and constructing a double-branch processing network based on visual Mangbar to obtain images respectively reflecting forward motion contribution and backward motion contribution, and finally generating a high-resolution blood vessel image through image fusion. The method has an important value for improving the efficiency and precision in microvascular imaging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and deep learning technology, and specifically relates to an ultrasonic positioning microscopy imaging method and system based on joint resolution perception and visual Mamba. Background Art

[0002] Ultrasound localization microscopy (ULM) is a super-resolution vascular imaging technology based on microbubble contrast agents. It combines the penetrating power of ultrasound with the localizable properties of microbubbles to achieve micron-scale reconstruction of deep microvascular networks. As one of the most commonly used ultrasound contrast agents, microbubbles can be used for noninvasive imaging of various organs, including the brain and kidneys, and have been widely used in disease diagnosis and clinical research. However, existing ULM technology suffers from a significant conflict between data acquisition and imaging quality. High microbubble concentrations can accelerate vascular filling and shorten acquisition time, but they can lead to signal confusion due to microbubble overlap, seriously compromising the accuracy of vascular distribution reconstruction. Low microbubble concentrations, while advantageous for single-bubble localization and high-quality imaging, require longer time to accumulate sufficient localization points, are prone to introducing tissue motion and respiratory artifacts, and increase time and storage costs, thus limiting the clinical application of ULM technology.

[0003] To address these issues, various algorithmic improvements have been proposed. Traditional methods primarily utilize spatiotemporal filtering or sparse reconstruction techniques, using Fourier transforms or sparse recovery theory to separate overlapping microbubbles, improving localization capabilities in high-concentration scenarios and reducing the number of sampling frames. For example, filtering methods based on diverse spatiotemporal flow characteristics can improve microbubble detection accuracy while maintaining imaging speed. Sparse reconstruction methods, on the other hand, use sparse recovery algorithms to locate overlapping microbubbles, reducing reliance on the number of ultrasound images. Deep learning-based approaches, centered on optimizing neural network architectures, have significantly improved localization accuracy and data processing efficiency in high-concentration microbubbles. Among them, a solution combining generative adversarial networks (GANs) with context-aware models achieves efficient localization at high microbubble concentrations. Methods utilizing probabilistic estimation models can predict the state and uncertainty of dense microbubbles. A sub-pixel convolutional neural network refinement strategy further enhances localization performance. Furthermore, for low-concentration scenarios, conditional generative adversarial networks (cGANs) have been used for super-resolution image reconstruction, combining Gaussian fitting and deconvolution algorithms to improve image clarity.

[0004] Although the above technologies have achieved certain results in their respective application scenarios, they still face the following shortcomings: First, the number of existing public simulation data sets is limited, and most of them rely on random microbubble motion simulation, lacking real dynamic characteristics, resulting in insufficient model training and generalization capabilities; second, the microbubble aggregation phenomenon in high-concentration scenarios is still serious, and overlapping microbubbles obscure individual features, resulting in blurred imaging and reduced detection accuracy; third, data acquisition in low-concentration scenarios is long and costly, making it difficult to meet the needs of large-scale and efficient imaging.

[0005] Therefore, there is an urgent need for a ULM technology that can balance imaging speed and accuracy under different microbubble concentration conditions, and at the same time build a high-fidelity simulation data set to support the efficient training and robust application of deep learning models. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides an ultrasound localization microscopy imaging method and system based on joint resolution perception and visual Mamba, aiming to balance microbubble concentration, imaging quality and acquisition efficiency, and provide strong support for the large-scale promotion of ULM in clinical and scientific research.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] An ultrasound positioning microscopy imaging method based on joint resolution perception and visual Mamba includes the following steps:

[0009] Step S1: Acquire and pre-process several modal angiography data, generate microbubble motion sequences based on blood flow velocity field simulation data, and construct a data set;

[0010] Step S2: Inputting the data set into a hybrid hierarchical feature pyramid fusion network, performing probability density estimation on the microbubble positions based on a real-time multi-scale density field estimation algorithm, introducing a resolution-aware algorithm to enhance microbubble features, outputting a microbubble location coordinate set, and performing physical analysis and association on the microbubble location coordinate set to obtain a number of short trajectory images;

[0011] Step S3: Dynamic feature decoupling and time-frequency analysis are performed on the short trajectory image to extract and encode velocity component features and spatiotemporal context information of vascular structure. A dual-branch processing network is constructed based on Visual Mamba to obtain images reflecting the forward motion contribution and the backward motion contribution respectively. A high-resolution vascular image is finally generated through image fusion.

[0012] Furthermore, step S1 includes: obtaining a number of modal angiography data as a basic input image, using adaptive threshold segmentation and three-dimensional topology reconstruction technology to construct a high-fidelity vascular wall geometric model, calculating the optimal movement path of the intravascular fluid based on the improved Dijkstra path tracking algorithm and fusing the blood flow velocity field simulation data, generating a probability distribution heat map of microbubble movement and then obtaining the possible movement trajectory of the microbubble, constructing a dynamic prompt word library based on clinical prior knowledge, applying a diffusion model constrained by vascular trajectory to generate data based on the prompt words, so that the possible movement trajectory becomes a microbubble movement sequence that conforms to the biomechanical characteristics, and completing data preparation.

[0013] Furthermore, step S1 specifically includes:

[0014] Step S11: collecting basic angiography image data, performing image vessel enhancement, tissue denoising and vessel segmentation operations on each image to obtain a corresponding vascular structure image;

[0015] Step S12: Apply the path tracking algorithm to each vascular structure image to extract the vascular connectivity structure and obtain blood flow velocity field simulation data; simulate the optimal blood flow path curve based on the microbubble motion physical model, and generate a probability distribution heat map of microbubble motion based on the blood flow velocity field simulation data to obtain the target motion path of the microbubble. ;

[0016] Step S13: constructing a conditional prompt vector and using it as a control condition for the diffusion model simulation graph generation model;

[0017] Step S14: Input the conditional prompt vector and the initial noise map into the backbone network of the diffusion model simulation map generation model, and the backbone network sequentially extracts semantic feature maps through its multiple layers, i-th layer The semantic feature map is recorded as ;

[0018] Step S15: Introduce a trajectory control unit and use the target motion path of the microbubble Derive trajectory condition information and perform the calculation on the i-th level of the backbone network Semantic feature map of Modulate and output the trajectory perception feature map, recorded as ;

[0019] Step S16: The target movement path of the microbubble As a precise structural control condition, the ControlNet module is used to extract the structure control condition and the backbone network layer i Corresponding path geometry and spatial constraint features ;Will Trajectory perception feature map of the corresponding level Perform deep fusion to generate intermediate fusion image features ;

[0020] Step S17: Fusing image features at each level After upsampling and decoding, the edge Simulated image of microbubbles that move and meet the conditional prompt vector ;

[0021] Step S18: Generate microbubble simulation image Perform authenticity detection and format standardization to obtain the final simulation image dataset.

[0022] Furthermore, step S2 specifically includes:

[0023] Step S21: normalize the obtained simulated image data set to obtain the global central tendency analysis measure and spatial variation degree quantitative value of the image as the normalized image set. ;

[0024] Step S22: normalize the image The input is input into the detection coding module for feature extraction and then input into the detection decoding module and density estimation decoding module respectively. The detection decoding module outputs the microbubble position detection result, and the density estimation decoding module outputs the predicted heat map. , the predicted heat map Reflects the predicted spatially dense distribution characteristics of microbubbles on the image for subsequent supervised learning;

[0025] Step S23: input image The pixel distribution is statistically tested, and the dynamic resolution test method based on Kolmogorov-Smirnov is used to calculate the empirical cumulative distribution function of the current image area. With the preset standard distribution The maximum deviation between :

[0026] (1)

[0027] In formula (1), Represents the minimum value of all upper bounds of the set, according to the statistic The value range of is used to divide the current image or image region into high complexity, medium complexity and low complexity regions. For regions of different complexity, the corresponding feature connection paths between the encoding module and the decoding module are dynamically activated.

[0028] Furthermore, in step S22, the detection coding module includes a plurality of feature extraction unit structures, wherein the feature extraction unit structure includes a convolutional layer, a batch normalization layer, and a convolution unit of an activation function cascaded in sequence, a fusion unit for feature fusion, and a spatial compression and downsampling unit for spatial information compression and downsampling;

[0029] The detection decoding module receives the features output by the detection encoding module and gradually integrates the multi-scale feature information to generate the detection results. , each set of detection results contains the center coordinates of the detected microbubble target ( ),width ,high and the corresponding confidence score , is the number of targets detected in the current image;

[0030] The density estimation decoding module adopts a layer-by-layer upsampling mechanism and uses jump connections between layers to fuse the corresponding layer features from the detection encoding module, and finally outputs the predicted heat map .

[0031] Furthermore, the step 2 further includes:

[0032] Step S24: Use the mean square error loss function to predict the heat map With reference density map Constraints are imposed between them to calculate the density loss , as shown in formula (2):

[0033] (2)

[0034] Among them, in formula (2) and Represent the observation result and the true value respectively, H and W represent the pixel width and length of the image respectively;

[0035] The detection loss calculated based on the detection results output by the detection decoding module , construct the overall loss function , as shown in formula (3):

[0036] (3)

[0037] Among them, in formula (3) and It is a preset weight hyperparameter used to balance the contribution of detection task and density regression task in the model learning process;

[0038] Model training is performed by minimizing the overall loss function.

[0039] Furthermore, the step 3 includes:

[0040] Step S31: Receive the short track image data generated by step S2 And perform analysis and decoupling to extract shared vascular structure features , and characteristic diagrams representing the forward motion of microbubbles and characteristic graphs characterizing the backward movement of microbubbles ;

[0041] Step S32: extract the shared vascular structure features Forward motion features and backward motion characteristics Combine to obtain forward and backward feature flows;

[0042] Step S33: construct two parallel SSM processing trunks with the same structure but partially shared parameters to process the forward and backward feature streams respectively;

[0043] Step S34: After each SSM processing trunk, integrate the residual noise reduction unit , purify the signals of the forward and backward feature streams;

[0044] Step S35: Decode the purified feature maps obtained by the two parallel streams separately and finally fuse them to obtain the final high-resolution output image. .

[0045] In another aspect, the present invention provides an ultrasound positioning microscopic imaging system based on joint resolution perception and visual Mamba, comprising:

[0046] A simulation data generation unit is used to obtain and pre-process a plurality of modal angiography data, generate a microbubble motion sequence in combination with the blood flow velocity field simulation data, and construct a data set;

[0047] a resolution perception unit, configured to input the data set into a hybrid hierarchical feature pyramid fusion network, perform probability density estimation of microbubble positions based on a real-time multi-scale density field estimation algorithm, introduce a resolution perception algorithm to enhance microbubble features, and output a set of microbubble location coordinates;

[0048] The Visual Mamba diffusion and reconstruction module unit is used to perform physical analysis and association on the microbubble positioning coordinate set to obtain a number of short-trajectory images, perform dynamic feature decoupling and time-frequency analysis on the short-trajectory images, extract velocity component features and spatiotemporal context information of vascular structure and encode them, construct a dual-branch processing network based on the Visual Mamba to obtain images reflecting the forward motion contribution and the backward motion contribution respectively, and finally generate a high-resolution vascular image through image fusion.

[0049] In a third aspect, the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned ultrasonic positioning microscopy imaging method based on joint resolution perception and visual mamba.

[0050] In a fourth aspect, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enables the processor to implement the aforementioned ultrasonic positioning microscopy method based on joint resolution perception and visual mamba, and regularly use newly captured data to continuously learn and update the above-mentioned collaborative model base model with multiple prompts.

[0051] The beneficial effects of the present invention are:

[0052] 1. This paper designs a vascular morphology-guided simulation data module based on multi-resolution perception, which simulates the movement trajectory of microbubbles in real vascular networks at three resolution levels: coarse, medium, and fine. It also performs morphological cue fusion through ControlNet, making the simulation data high-fidelity and diverse, significantly enhancing the coverage and authenticity of the training set, and providing key data support for subsequent vascular imaging generative model training, thereby improving the model's generalization ability under different vascular structures and imaging conditions.

[0053] 2. This paper designs a multi-task convolutional neural network that combines target detection, density estimation, and resolution adaptation. By introducing a feature pyramid network (FPN), a self-attention module, and a dynamic skip-connection strategy based on the Kolmogorov-Smirnov test, it achieves refined processing of microbubble signals and adaptive optimization of the network structure. This network outputs highly accurate microbubble pixel-level location, local density, and resolution quality assessment information within a single framework. These precise outputs collectively constitute crucial prior knowledge for subsequent hemodynamic analysis and high-quality vascular reconstruction, laying a solid foundation for the Visual Mamba diffusion reconstruction module to achieve precise imaging and improve overall system performance.

[0054] 3. The present invention designs a visual Mamba diffusion reconstruction module, which uses a small number of positioning points (accurate output from step S2) and an extremely short sampling time as input. Through an iterative denoising process combined with vascular prior reasoning, it utilizes the iterative denoising characteristics of the diffusion model and a multi-directional scanning mechanism to rapidly generate high-resolution vascular images. While preserving the details of microvascular branches, this module significantly removes artifacts and improves the signal-to-noise ratio, providing high-quality imaging results for clinical non-invasive microvascular detection. Due to its efficient reconstruction method, it reduces data acquisition time and subsequent storage and computing costs while ensuring image quality.

[0055] 4. This invention integrates a simulation data module, a multi-task convolutional neural network, and a visual Mamba diffusion reconstruction module to construct a complete ultrasound localization microscopy imaging solution from data source to high-quality imaging. This not only improves imaging accuracy and speed in complex scenarios, but also effectively addresses the challenges of traditional ULM technology in data acquisition, processing efficiency, and imaging quality. This provides strong technical support for early and accurate diagnosis of clinical diseases, evaluation of treatment effects, and related medical research, demonstrating significant clinical application value and potential. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 This is a schematic diagram of the ultrasonic positioning microscopy imaging method based on joint resolution perception and visual Mamba of the present invention;

[0057] Figure 2 A schematic diagram of the process of constructing the data set of the present invention;

[0058] Figure 3 This is a schematic diagram of the detection algorithm based on multi-resolution perception and real-time multi-scale density field estimation of the present invention;

[0059] Figure 4 This is a schematic diagram of the Mamba super-resolution generation network based on physical constraint prior reasoning of the present invention;

[0060] Figure 5 This is a structural diagram of the ultrasonic positioning microscopy system based on joint resolution perception and visual Mamba in the present invention. DETAILED DESCRIPTION

[0061] The present invention will be further described below with reference to the accompanying drawings and examples.

[0062] like Figure 1 As shown, the present invention provides an ultrasound positioning microscopy imaging method based on joint resolution perception and visual Mamba, comprising the following steps:

[0063] Step S1: Acquire several modal angiography data as the basic input image, and use adaptive threshold segmentation and three-dimensional topology reconstruction technology to construct a high-fidelity vascular wall geometric model. Based on the improved Dijkstra path tracking algorithm, the optimal movement path of the intravascular fluid is calculated and the blood flow velocity field simulation data is integrated to generate a probability distribution heat map of microbubble movement and then generate the possible movement trajectory of the microbubble. A dynamic prompt vocabulary is constructed based on clinical prior knowledge to adjust the hyperparameters. On this basis, the diffusion model of vascular trajectory constraints is applied to generate data according to the prompts, so that it becomes a microbubble motion sequence that conforms to the biomechanical characteristics. The output includes the original vascular topology, microbubble motion trajectory parameters, etc. The above output is used for data set preparation. Specifically, if Figure 2 As shown, including:

[0064] Step S11: Acquire basic vascular image dataset , for each image Perform vessel enhancement, tissue denoising and vessel segmentation operations to obtain the corresponding vessel structure image , where the vascular region is extracted through the segmentation network and represented in the form of a binary image to obtain a set of structured vascular images ;

[0065] Step S12: For each vascular structure image A path tracking algorithm is applied to extract the vascular connectivity structure and construct a centerline trajectory model to obtain blood flow velocity field simulation data. This simulation data contains the flow velocity information within the blood vessels. On the other hand, based on the physical model of microbubble motion, a simulated blood flow path curve is obtained through spline interpolation. In order to more accurately predict the behavior of microbubbles, the obtained initial path curve is combined with the blood flow velocity field simulation data to perform optimization calculations to determine and generate the target optimal motion path of the microbubbles within the blood vessels. Based on this, we can further analyze and construct a probability distribution map of microbubble motion based on the optimal motion path, possible flow field dynamics, and model uncertainty. This probability distribution map statistically characterizes the likelihood of microbubbles appearing in a specific area or the distribution range of their motion trajectories, providing more comprehensive information for subsequent analysis.

[0066] Step S13: The generated image's physiological semantic features, such as lesion location and blood flow status, are controlled using text keywords. Time control parameters are also used to adjust the dynamic properties of the simulated image, such as the degree of microbubble movement within the frame sequence. This constructs a conditional cue vector, which serves as the control condition for the generative model.

[0067] Step S14: Input the hint vector and the initial noise map into the backbone network of the vascular trajectory control generation network. The backbone network sequentially extracts semantic feature maps through its multiple layers. The semantic feature map of the i-th layer is recorded as ;

[0068] Step S15: introduce the trajectory control unit, which uses the Derive trajectory condition information and perform specific layer-level operations on the backbone network Feature map Modulate to incorporate path guidance information and output trajectory perception feature map, denoted as ;

[0069] Step S16: Design a ControlNet module to convert the most efficient motion path generated in step S12 into As a precise structural control condition. This module extracts the specific layer of the backbone network from the structural control condition. Corresponding path geometry and spatial constraint features ; and these features The trajectory perception feature map of the corresponding level generated in step S15 Perform deep fusion and finally generate an intermediate fusion image feature The purpose of this fusion operation is to combine the precise path geometry constraints with the features that already contain the preliminary path guidance;

[0070] Step S17: Fusion features of each level After subsequent upsampling and decoding, the final edge Simulated images of microbubbles that are moving and meet preset prompt conditions ;

[0071] Step S18: Generate microbubble simulation image Authenticity testing and format standardization are performed. This includes determining whether microbubbles are located within the vascular pathway, drift, or abnormal generation. Finally, the output image format, size, and frame number are standardized to obtain the final simulated image dataset.

[0072] Step S2: Design a hybrid hierarchical feature pyramid fusion network. First, introduce a feature enhancement module based on cross-dimensional attention to simultaneously enhance the microbubble feature representation in the channel and spatial dimensions to suppress the interference of complex backgrounds; at the same time, use a deformation adaptive network to adaptively adjust the microbubble sampling position to improve the positioning accuracy of its edges. The above two designs can strengthen feature representation and improve positioning effects. In addition, a dynamic link resolution perception strategy is deployed between the encoding module and the detection and decoding module to adaptively adjust the jump connection and dynamically adjust the feature fusion weights based on the perception of the data resolution. In addition, a real-time multi-scale density field estimation algorithm is constructed synchronously to model the spatial distribution probability of microbubbles to deal with overlapping problems in high-density scenes. At the same time, the microbubble density distribution prior is used to realize statistically driven self-supervised training of microbubble positions to deal with microbubble occlusion and mutual interference in high-density scenes, and finally provide pixel-level precise position constraints for the network. Specifically, such as Figure 3 As shown:

[0073] Step S21: input image data and perform preprocessing. The acquired original ultrasound image data set is recorded as: ,in, Indicates the The original grayscale image has a height and width of The final dataset image obtained by S18 is normalized to obtain the global central tendency analysis measure and spatial variation degree quantitative value of the image, and the normalized image set is obtained. ;

[0074] Step S22: normalize the image obtained in S21 The input is sent to the detection encoding module for feature extraction. This encoding module contains multiple feature extraction unit structures, each of which consists of three parts: a convolution unit consisting of a convolution layer, a batch normalization layer, and an activation function; a fusion unit for feature fusion; and a spatial compression and downsampling unit for spatial information compression and downsampling.

[0075] The features output by the encoding module will enter the detection decoding module. The detection decoding module gradually integrates multi-scale feature information to generate detection results by adopting, for example, a double-layer upsampling structure and a feature fusion unit. Each set of results contains the center coordinates of the detected microbubble target ( ),width ,high and confidence score , is the number of objects detected in the current image.

[0076] At the same time, the features output by the encoding module are also input to the density estimation decoding module. This module adopts a layer-by-layer upsampling mechanism and uses jump connections between layers to fuse the corresponding layer features from the encoding module. Each upsampling layer contains a density estimation unit consisting only of convolutional layers. The density estimation decoding module finally outputs a heat map ,The heat map reflects the predicted spatial dense distribution characteristics of microbubbles on the image, which is used for subsequent supervised learning;

[0077] Step S23: Input image The pixel distribution is statistically tested, and the dynamic resolution test method based on Kolmogorov-Smirnov is used to calculate the empirical cumulative distribution function of the current image area. With the preset standard distribution The maximum deviation between :

[0078] (1)

[0079] In formula (1), Represents the minimum value of all upper bounds of the set. According to the statistic The current image or image region is divided into high-complexity, medium-complexity, and low-complexity regions based on the value range of . For regions of different complexity, the corresponding feature connection paths between the encoding module and the decoding module (including the detection decoding module and the density estimation decoding module) are dynamically activated: high-complexity regions prioritize activating and strengthening the connection between deep features and the density estimation branch; medium-complexity regions enable cross-level jump fusion paths to balance global and local information; and low-complexity regions focus on connecting shallow features to the detection branch. This mechanism enables dynamic regulation and adaptive reconstruction of the model's connection structure to optimize feature utilization efficiency in different scenarios.

[0080] Step S24: In order to realize the self-supervised training of microbubble positions, a supervision term based on the density heat map is introduced. The predicted heat map is trained using the mean square error loss function. Compared with the real or reference density map generated by methods such as kernel density estimation Constraints are imposed between the two and density loss is calculated. , as shown in formula (2):

[0081] (2)

[0082] Among them, in formula (2) and Represent the observation result and the true value respectively, H and W represent the pixel width and length of the image respectively. At the same time, the detection loss calculated by combining the detection result output by the detection decoding module is (including classification loss and positioning loss), construct the overall loss function , as shown in formula (3):

[0083] (3)

[0084] Among them, in formula (3) and is a preset weight hyperparameter used to balance the contributions of the detection and density regression tasks during model learning. Model training is performed by minimizing this overall loss function, where the density loss term provides pixel-level constraints for precise microbubble location learning, helping to address microbubble occlusion and mutual interference in high-density scenarios.

[0085] Step S3: Perform dynamic feature decoupling and time-frequency analysis on the short trajectory image representing the microbubble movement generated by step S2 to separate and extract the velocity component features with different dynamic characteristics and the corresponding spatiotemporal context information of the vascular structure; then, combine these separated structural information with the motion information in each direction and perform deep encoding; on this basis, construct two parallel core processing branches, one branch specifically processes forward motion-related features, and the other branch specifically processes backward motion-related features. Both branches utilize shared structural context information. Each processing branch uses a state space model (SSM) unit that is specially optimized for two-dimensional image features and integrates an eight-directional scanning mechanism as the backbone. It performs deep context relationship modeling and feature purification on its corresponding directional motion feature block sequence, and integrates a residual denoising unit. Each branch independently decodes to generate high-resolution images that reflect the forward motion contribution and the backward motion contribution respectively. Finally, the two directional contribution images are superimposed through image fusion operations to form a final high-resolution image that can clearly display the vascular structure and bidirectional blood flow information at the same time. The training of the entire model is guided by optimizing the combined loss function. Figure 4 As shown, specifically including:

[0086] Step S31: Receive the short trajectory image data generated by step S2 , this segment of trajectory image data encodes the movement information of microbubbles in a short time. Apply trajectory analysis and motion direction separation module, which decouples and extracts main vascular structural features from trajectory information , and characteristic graphs characterizing the forward motion of microbubbles and characteristic graphs characterizing the backward movement of microbubbles ;

[0087] Step S32: To construct two parallel processing branches, firstly extract the shared structural features from step S31. Forward motion features and backward motion characteristics To combine:

[0088] 1. Forward feature flow construction: and Splicing is performed on the channel dimension to obtain combined features ; Then, through the depth-wise separable convolution module right Perform preliminary enhancement and channel adjustment to obtain the enhanced feature map of the forward flow .

[0089] 2. Backward feature flow construction: and Splicing is performed on the channel dimension to obtain combined features ; Then, through the depth-wise separable convolution module right Perform preliminary enhancement and channel adjustment to obtain the enhanced feature map of the backward flow ;in is the number of feature channels for each processing stream.

[0090] Afterwards, the and Split into a series of non-overlapping, fixed-size The feature block sequences are recorded as and , used as structured input for subsequent parallel processing.

[0091] Step S33: Construct two parallel SSM processing trunks with the same structure but partially shared parameters to process the forward and backward feature flows respectively:

[0092] 1. Forward flow processing: Sequence the feature blocks Input to the first SSM trunk.

[0093] 2. Backward flow processing: Sequence the feature blocks Input to the second SSM trunk.

[0094] Each SSM backbone uses an eight-directional scanning mechanism optimized for two-dimensional image features. The feature block sequence is rearranged according to the eight main scanning directions (horizontal forward, horizontal reverse, vertical forward, vertical reverse, main diagonal forward, main diagonal reverse, sub-diagonal forward, sub-diagonal reverse) to form eight groups of one-dimensional sequences in different directions. Each set of directional sequences are independently input into a series of stacked SSM units. Each SSM unit propagates the state according to the core equation , Calculate the parameter matrix is the input feature function, and for different scanning directions Has an independent set of parameters. The SiLU nonlinear activation function is applied to the output The SSM output feature representations of each scanning direction are then concatenated in the channel dimension and deeply fused through a convolutional network module enhanced by the attention mechanism to obtain the enhanced feature representations of the forward flow. and enhanced feature representation of backward flow ;

[0095] Step S34: After the SSM backbone of each parallel processing stream, integrate the residual noise reduction unit .

[0096] 1. Forward flow denoising: feature representation after forward flow fusion (reorganized into feature map form) Apply the noise reduction unit to obtain :

[0097] (4)

[0098] Among them, in formula (4) It is the feature representation after the forward flow fusion is reorganized into the form of feature map, Indicates normalization along the channel dimension to reduce internal covariate shift.

[0099] 2. Backward flow denoising: Feature representation after backward flow fusion (reorganized into feature map form) Apply the noise reduction unit to obtain , the calculation method is the same as formula (4). This step aims to purify the signals of the two streams separately and retain the valuable structural and directional dynamic features.

[0100] Step S35: Decode the purified feature maps obtained by the two parallel streams separately and finally fuse them:

[0101] 1. Forward image decoding: Through depth-wise separable convolution modules and the subsequent decoding network (Using sub-pixel convolution upsampling and feature fusion) to reconstruct a high-resolution image reflecting the contribution of forward motion .

[0102] 2. Backward image decoding: Through depth-wise separable convolution modules and the subsequent decoding network , reconstructing a high-resolution image reflecting the contribution of backward motion .

[0103] 3. Final image fusion: The forward contribution image and the backward contribution image The final high-resolution output image is obtained by fusion through pixel-level image addition operation. This image overlay operation combines flow information from both directions, providing a complete depiction of the vascular structure.

[0104] To optimize the parameters of the entire model (including all components of the two parallel branches), a composite loss function is designed , this function mainly acts on the final fused image The loss function includes: pixel-level reconstruction loss : Measure the final fused image with real high-resolution reference images The difference between them, L1 loss is used. Perceptual loss : Using the feature extraction capabilities of pre-trained deep networks, comparison and Similarity in feature space to enhance visual realism. Overall loss function :

[0105] (5)

[0106] Among them, in formula (5) is the weight hyperparameter of each loss. By minimizing The model is trained end-to-end. This architecture is designed to separate and process features of different motion directions, enabling the network to more finely learn and reconstruct vascular details related to blood flow in a specific direction, and ultimately fusion to present a complete hemodynamic picture.

[0107] This invention supports an ultrasound-localized microscopy system based on joint resolution perception and visual Mamba. Through multi-resolution simulation image generation and S2 design network, it realizes the coordinated iteration of high-density microbubble localization and deep blood vessel super-resolution reconstruction, improves the spatial resolution and signal-to-noise ratio of imaging, enhances real-time processing capabilities and system robustness, reduces hardware deployment costs, and facilitates the clinical application and rapid promotion of ultrasound-localized microscopy technology.

[0108] like Figure 5 As shown, the present invention provides an ultrasound positioning microscopic imaging system based on joint resolution perception and visual Mamba, comprising:

[0109] The simulation data generation module 51 is used to receive ultrasound or optical imaging sequences of the target vascular area, perform pixel-level segmentation and three-dimensional reconstruction of the sequences based on a multi-resolution perception strategy, and extract vascular structure models at three levels: coarse, medium, and fine. The module simulates the motion trajectory of microbubbles at each level and inputs the trajectory information together with vascular morphology cues into the diffusion generation model guided by the fusion ControlNet morphology to generate multi-scale, high-fidelity microbubble simulation image sequences, providing a realistic and diverse ULM dataset for subsequent network training.

[0110] The resolution perception module 52 loads the multi-scale microbubble simulation image sequences and original ultrasound frames output by the simulation data generation module 51 and invokes a quantized, multi-task perception convolutional neural network model deployed on the hardware-accelerated edge device to perform joint inference on the input data. This network utilizes a feature pyramid network (FPN) and a self-attention mechanism to achieve multi-scale feature fusion and suppress complex background interference, thereby accurately localizing microbubble targets. Furthermore, a pixel-level probabilistic loss is introduced to constrain regions of high-density overlapping microbubbles to improve density estimation accuracy in dense scenes. Furthermore, a cross-layer skip connection structure is dynamically adjusted based on Kolmogorov–Smirnov test statistics, enabling the network to adaptively allocate shallow detection, cross-layer fusion, and deep density branches across regions of varying image complexity, thereby optimizing inference efficiency while maintaining high detection accuracy. After collaborative processing of these multiple tasks, the detection module 52 outputs multi-resolution microbubble localization points and their confidence scores, providing precise localization information for subsequent visual Mamba diffusion reconstruction.

[0111] The visual Mamba diffusion and reconstruction module 53 is used to input the multi-resolution positioning points output by the multi-task detection module 52 and the corresponding original ultrasound frames into the visual Mamba diffusion reconstruction network. The network first maps the positioning points to several non-overlapping feature blocks through a feature splitting layer, and then embeds depth-separable convolution to enhance local detail perception; in the diffusion trunk, Mamba blocks are alternately executed in the horizontal, vertical and diagonal directions to capture global structural information, and structural consistency is maintained during the iterative denoising process through conditional layer normalization and vascular prior inference; finally, the features after multi-layer Mamba diffusion and denoising optimization are reconstructed by the decoder into high signal-to-noise ratio, high-resolution deep vascular imaging results.

[0112] Through the collaborative work of the above three modules, the system of the present invention can generate large-scale and diverse simulation training data, realize real-time and efficient microbubble detection and high-precision vascular imaging reconstruction, and significantly improve the resolution and signal-to-noise ratio of positioning microscopy imaging.

[0113] In a third aspect, the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned ultrasonic positioning microscopy imaging method based on joint resolution perception and visual mamba.

[0114] In a fourth aspect, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enables the processor to implement the aforementioned ultrasonic positioning microscopy method based on joint resolution perception and visual mamba, and regularly use newly captured data to continuously learn and update the above-mentioned collaborative model base model with multiple prompts.

[0115] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An ultrasound positioning microscopy imaging method based on joint resolution perception and visual Mamba, characterized in that: The steps include: Step S1: Acquire and pre-process several modal angiography data, generate microbubble motion sequences based on blood flow velocity field simulation data, and construct a data set; Step S2: Inputting the data set into a hybrid hierarchical feature pyramid fusion network, performing probability density estimation on the microbubble positions based on a real-time multi-scale density field estimation algorithm, introducing a resolution-aware algorithm to enhance microbubble features, outputting a microbubble location coordinate set, and performing physical analysis and association on the microbubble location coordinate set to obtain a number of short trajectory images; Step S3: Dynamic feature decoupling and time-frequency analysis are performed on the short trajectory image to extract and encode velocity component features and spatiotemporal context information of vascular structure. A dual-branch processing network is constructed based on Visual Mamba to obtain images reflecting the forward motion contribution and the backward motion contribution respectively. A high-resolution vascular image is finally generated through image fusion.

2. The ultrasound positioning microscopy imaging method based on joint resolution perception and visual Mamba according to claim 1 is characterized in that: The step S1 includes: obtaining a plurality of modal angiography data as a basic input image, constructing a high-fidelity vascular wall geometric model using adaptive threshold segmentation and three-dimensional topological reconstruction technology, calculating the optimal motion path of the intravascular fluid based on an improved Dijkstra path tracking algorithm and fusing the blood flow velocity field simulation data, generating a probability distribution heat map of microbubble motion and thereby obtaining possible motion trajectories of the microbubbles, constructing a dynamic prompt word library based on clinical prior knowledge, applying a diffusion model constrained by vascular trajectories to generate data based on prompt words, so that the possible motion trajectories become microbubble motion sequences that conform to biomechanical properties, and completing data preparation.

3. The ultrasound positioning microscopy imaging method based on joint resolution perception and visual Mamba according to claim 2 is characterized in that: Step S1 specifically includes: Step S11: collecting basic angiography image data, performing image vessel enhancement, tissue denoising and vessel segmentation operations on each image to obtain a corresponding vascular structure image; Step S12: Apply the path tracking algorithm to each vascular structure image to extract the vascular connectivity structure and obtain blood flow velocity field simulation data; simulate the optimal blood flow path curve based on the microbubble motion physical model, and generate a probability distribution heat map of microbubble motion based on the blood flow velocity field simulation data to obtain the target motion path of the microbubble. ; Step S13: constructing a conditional prompt vector and using it as a control condition for the diffusion model simulation graph generation model; Step S14: Input the conditional prompt vector and the initial noise map into the backbone network of the diffusion model simulation map generation model, and the backbone network sequentially extracts semantic feature maps through its multiple layers, i-th layer The semantic feature map is recorded as ; Step S15: Introduce a trajectory control unit and use the target motion path of the microbubble Derive trajectory condition information and perform the calculation on the i-th level of the backbone network Semantic feature map of Modulate and output the trajectory perception feature map, recorded as ; Step S16: The target movement path of the microbubble As a precise structural control condition, the ControlNet module is used to extract the structure control condition and the backbone network layer i Corresponding path geometry and spatial constraint features ;Will Trajectory perception feature map of the corresponding level Perform deep fusion to generate intermediate fusion image features ; Step S17: Fusing image features at each level After upsampling and decoding, the edge Simulated image of microbubbles that move and meet the conditional prompt vector ; Step S18: Generate microbubble simulation image Perform authenticity detection and format standardization to obtain the final simulation image dataset.

4. The ultrasound positioning microscopy imaging method based on joint resolution perception and visual Mamba according to claim 1, characterized in that: Step S2 specifically includes: Step S21: normalize the obtained simulated image data set to obtain the global central tendency analysis measure and spatial variation degree quantitative value of the image as the normalized image set. ; Step S22: normalize the image The input is input into the detection coding module for feature extraction and then input into the detection decoding module and density estimation decoding module respectively. The detection decoding module outputs the microbubble position detection result, and the density estimation decoding module outputs the predicted heat map. , the predicted heat map Reflects the predicted spatially dense distribution characteristics of microbubbles on the image for subsequent supervised learning; Step S23: input image The pixel distribution is statistically tested, and the dynamic resolution test method based on Kolmogorov-Smirnov is used to calculate the empirical cumulative distribution function of the current image area. With the preset standard distribution The maximum deviation between : (1) In formula (1), Represents the minimum value of all upper bounds of the set, according to the statistic The value range of is used to divide the current image or image region into high complexity, medium complexity and low complexity regions. For regions of different complexity, the corresponding feature connection paths between the encoding module and the decoding module are dynamically activated.

5. The ultrasound positioning microscopy imaging method based on joint resolution perception and visual Mamba according to claim 4 is characterized in that: In step S22, the detection coding module includes a plurality of feature extraction unit structures, wherein the feature extraction unit structure includes a convolutional layer, a batch normalization layer, and a convolution unit of an activation function cascaded in sequence, a fusion unit for feature fusion, and a spatial compression and downsampling unit for spatial information compression and downsampling; The detection decoding module receives the features output by the detection encoding module and gradually integrates the multi-scale feature information to generate the detection results. , each set of detection results contains the center coordinates of the detected microbubble target ( ),width ,high and the corresponding confidence score , is the number of targets detected in the current image; The density estimation decoding module adopts a layer-by-layer upsampling mechanism and uses jump connections between layers to fuse the corresponding layer features from the detection encoding module, and finally outputs the predicted heat map .

6. The ultrasound positioning microscopy imaging method based on joint resolution perception and visual Mamba according to claim 4 is characterized in that: The step 2 further comprises: Step S24: Use the mean square error loss function to predict the heat map With reference density map Constraints are imposed between the two and density loss is calculated. , as shown in formula (2): (2) Among them, in formula (2) and Represent the observation result and the true value respectively, H and W represent the pixel width and length of the image respectively; The detection loss calculated based on the detection results output by the detection decoding module , construct the overall loss function , as shown in formula (3): (3) Among them, in formula (3) and It is a preset weight hyperparameter used to balance the contribution of detection task and density regression task in the model learning process; Model training is performed by minimizing the overall loss function.

7. The ultrasound localization microscopy imaging method based on joint resolution perception and visual Mamba according to claim 1, characterized in that: The step 3 comprises: Step S31: Receive the short track image data generated by step S2 And perform analysis and decoupling to extract shared vascular structure features , and characteristic diagrams representing the forward motion of microbubbles and characteristic graphs characterizing the backward movement of microbubbles ; Step S32: extract the shared vascular structure features Forward motion features and backward motion characteristics Combine to obtain forward and backward feature flows; Step S33: construct two parallel SSM processing trunks with the same structure but partially shared parameters to process the forward and backward feature streams respectively; Step S34: After each SSM processing trunk, integrate the residual noise reduction unit , purify the signals of the forward and backward feature streams; Step S35: Decode the purified feature maps obtained by the two parallel streams separately and finally fuse them to obtain the final high-resolution output image. .

8. Ultrasound positioning microscopy imaging system based on joint resolution perception and visual Mamba, characterized by: include: A simulation data generation unit is used to obtain and pre-process a plurality of modal angiography data, generate a microbubble motion sequence in combination with the blood flow velocity field simulation data, and construct a data set; a resolution perception unit, configured to input the data set into a hybrid hierarchical feature pyramid fusion network, perform probability density estimation of microbubble positions based on a real-time multi-scale density field estimation algorithm, introduce a resolution perception algorithm to enhance microbubble features, and output a set of microbubble location coordinates; The Visual Mamba diffusion and reconstruction module unit is used to perform physical analysis and association on the microbubble positioning coordinate set to obtain a number of short-trajectory images, perform dynamic feature decoupling and time-frequency analysis on the short-trajectory images, extract velocity component features and spatiotemporal context information of vascular structure and encode them, construct a dual-branch processing network based on the Visual Mamba to obtain images reflecting the forward motion contribution and the backward motion contribution respectively, and finally generate a high-resolution vascular image through image fusion.

9. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs; Wherein, when one or more programs are executed by the one or more processors, the one or more processors implement the ultrasonic positioning microscopy imaging method based on joint resolution perception and visual Mamba as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that Executable instructions are stored thereon, which, when executed by a processor, enable the processor to implement the ultrasonic positioning microscopy imaging method based on joint resolution perception and visual Mamba as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Signal processing method and system for ultrasonic micro blood flow imaging

    CN114533122A

  • Quantitative ultrasonic positioning microscopic imaging method based on attention learning mechanism

    CN114557719A

  • Super-resolution contrast imaging method, blood vessel imaging method and ultrasonic imaging device

    CN116458923A

  • Microvascular blood flow ultrasonic imaging method and system

    CN117679073A

  • Super-resolution contrast imaging method and ultrasonic imaging device

    CN118356212A

Cited By

  • Craniopharyngeal tubuloma image segmentation method based on morphological perception hierarchical information bottleneck network

    CN121053154A

  • AI vision and multi-target tracking combined livestock number identification method and system

    CN121095983A

  • Livestock number recognition method and system combining AI vision and multi-target tracking

    CN121095983B

  • Brain tumor detection method based on attention mechanism and MRI (Magnetic Resonance Imaging) multi-modal fusion

    CN121353272A