Ultrasonic image processing method, electronic equipment and storage medium
Through the TransUNet deep learning model combined with global context and local detail features, the problems of anatomical edge blur and insufficient contrast in traditional ultrasound imaging are solved, and high-precision segmentation and visualization of ultrasound images are realized, supporting more accurate ultrasound-guided regional anesthesia.
Patent Information
- Application Number
- CN202510855866.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Traditional ultrasound imaging technology is difficult to accurately identify the target nerve and surrounding tissue structures under problems such as blurred anatomical edges and insufficient contrast, which affects the anesthesia and analgesic effects.
The TransUNet deep learning model is used to enhance and segment ultrasound images in real time. By combining global context information and local detail features, the target feature map is generated and the segmentation mask is constructed, which significantly improves the image visualization effect.
It significantly improves the visualization effect of ultrasound images, improves image segmentation accuracy and identification accuracy of key anatomical structures, and supports more accurate ultrasound-guided regional anesthesia operation.
Smart Images

Figure CN120374420A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of ultrasonic image processing, and particularly relates to an ultrasonic image processing method, an electronic device, and a storage medium. Background Art
[0002] The ultrasonic-guided regional anesthesia technology is an effective anesthesia and analgesia measure. This technology requires an anesthesiologist to accurately identify nerves and their adjacent tissue structures through regional anesthesia ultrasonic images, and inject local anesthetic around the nerves under ultrasonic guidance to achieve the purpose of regional anesthesia. However, in the practical application of traditional ultrasonic-guided regional anesthesia, due to the limitations of ultrasonic imaging by various factors, such as noise, occlusion by bone structures, etc., the obtained ultrasonic images often have problems such as blurred anatomical edges and insufficient contrast, resulting in difficulty for doctors to accurately identify the target nerves and surrounding tissue structures from the ultrasonic images with poor visual effects.
[0003] Currently, some studies have attempted to improve the quality of ultrasonic images through image processing techniques, such as using traditional filtering or enhancement algorithms for image optimization. However, these methods are often designed for specific ultrasonic imaging problems and are difficult to handle the complex and variable characteristics under different organs and pathological conditions. For example, in abdominal ultrasound, artifacts generated by adipose tissue and air cavities cannot be effectively processed. Moreover, although these methods can reduce noise, they are prone to losing small structures and detailed information during the denoising process, affecting the diagnostic accuracy. For example, in the detection of liver nodules, excessive smoothing may mask the true morphology of early small tumors and interfere with the accurate judgment of doctors.
[0004] Therefore, there is currently a technical problem of poor visualization effect in ultrasonic imaging. Summary of the Invention
[0005] The purpose of the present application is to provide an ultrasonic image processing method, an electronic device, and a storage medium to solve the above problems.
[0006] To achieve the above purpose, in a first aspect, the present application proposes an ultrasonic image processing method, and the method includes: Obtain an initial ultrasonic image, and segment the initial ultrasonic image into a plurality of sub-images; Extract the global context information and local detail features of the plurality of sub-images; Generate a target feature map based on the global context information and the local detail features; Generate a type label for each pixel in the initial ultrasonic image by performing a convolution operation on the target feature map, where the type label includes one or more of nerves, blood vessels, muscles, fascia, and bones; Construct a segmentation mask for the initial ultrasound image based on the type label of each pixel, and overlay the segmentation mask on the initial ultrasound image to generate a target ultrasound image.
[0007] In some embodiments, the extracting of the global context information and local detail features of the several sub-images includes: Map the several sub-images into vector representations of a preset dimension; Add the position encoding of the corresponding sub-image to the vector representation of each sub-image; Based on the position encoding corresponding to each sub-image, calculate the global context information of the several sub-images through a multi-head self-attention mechanism; Based on the global context information, perform a non-linear transformation on the vector representation of each sub-image through a feed-forward network to obtain the local detail features of the several sub-images.
[0008] In some embodiments, the generating of the target feature map based on the global context information and the local detail features includes: Fuse the global context information and the local detail features to generate a fused feature map; Input the fused feature map into an encoder network to generate multi-level initial feature maps; Through skip connections, transfer the low-level initial feature maps to the corresponding levels of the decoder network to obtain low-level features; Upsample the high-level initial feature maps to obtain high-level features; Fuse the low-level features and the high-level features to generate a target feature map.
[0009] In some embodiments, after constructing the segmentation mask for the initial ultrasound image based on the type label of each pixel, and overlaying the segmentation mask on the initial ultrasound image to generate a target ultrasound image, it further includes: Based on a preset regional anesthesia standard image set, determine a standard image subset corresponding to the target ultrasound image; Perform a fusion process on each image in the standard image subset to obtain a fused standard image; Match the fused standard image and the target ultrasound image to determine the scanning accuracy of the target ultrasound image.
[0010] In some embodiments, after matching the fused standard image and the target ultrasound image to determine the scanning accuracy of the target ultrasound image, it further includes: When the scanning accuracy is lower than a preset matching degree threshold, obtain a set of consecutive target ultrasound images corresponding to consecutive frames of the initial ultrasound image; Based on the type label distribution of each frame of image in the set of consecutive target ultrasound images, determine the key anatomical structures of each frame of image in the set of consecutive target ultrasound images; Based on the position parameters of each of the key anatomical structures, construct a virtual anatomical model; Based on the virtual anatomical model and the fused standard image, determine the target scanning parameters of the ultrasound probe; Obtain the actual physical parameters of the ultrasound probe, and output a scanning guidance path according to the actual physical parameters and the target scanning parameters.
[0011] In some embodiments, the position information includes center coordinates, volume information, direction vectors, and depth information. Before constructing a virtual anatomical model based on the position parameters of each of the key anatomical structures, it further includes: According to the boundary information of each of the key anatomical structures, calculate the center coordinates and volume information of each of the key anatomical structures; By respectively fitting the main axes of each of the key anatomical structures, determine the direction vectors of each of the key anatomical structures; According to the original gray values of the target ultrasound image, determine the depth information of each of the key anatomical structures.
[0012] In some embodiments, constructing a virtual anatomical model based on the position parameters of each of the key anatomical structures further includes: Integrate the center coordinates, volume information, direction vectors, and depth information of each of the key anatomical structures in each frame of image in the set of consecutive target ultrasound images to generate a set of spatial parameters for each of the key anatomical structures; Based on the set of spatial parameters, map each of the key anatomical structures into a three-dimensional space to form a spatial distribution representation of each of the key anatomical structures; Render the spatial distribution representation of each of the key anatomical structures to generate a virtual anatomical model.
[0013] In some embodiments, after constructing a segmentation mask of the initial ultrasound image based on the type label of each pixel and overlaying the segmentation mask on the initial ultrasound image to generate a target ultrasound image, it further includes: According to the type label distribution in the target ultrasound image, determine the target anatomical structure; According to the tissue characteristics of the target anatomical structure, determine the target pressure range of the ultrasound probe; Obtain the real-time pressure of the ultrasonic probe, and output a correction prompt when the real-time pressure exceeds the target pressure range.
[0014] In a second aspect, the present application provides an electronic device, including: One or more processors; A memory for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the ultrasonic image processing method as described above.
[0015] In a third aspect, the present application provides a storage medium storing executable instructions, which when executed by a processor cause the processor to execute the ultrasonic image processing method as described above.
[0016] Compared with the prior art, the beneficial effects of the present application include: In the first aspect, by combining global context information and local detail features, the anatomical relationship of the target nerve and surrounding tissues can be captured more accurately, thus significantly improving the accuracy of image segmentation. In the second aspect, the pixel-level annotation method can describe the distribution of different tissues more meticulously compared with the traditional region-based or edge-based processing methods. In the third aspect, in the process of generating the target ultrasonic image, by superimposing the segmentation mask on the initial ultrasonic image, not only the information of the initial ultrasonic image can be retained, but also the key anatomical structures (such as nerves and blood vessels) can be displayed through visual markers, significantly improving the visualization effect of ultrasonic imaging. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope of the present application.
[0018] Figure 1 It is a schematic flowchart of an ultrasonic image processing method in an embodiment; Figure 2 It is a schematic flowchart of extracting the global context information and local detail features of the several sub-images in an embodiment; Figure 3 It is a schematic flowchart of generating a target feature map based on the global context information and the local detail features in an embodiment; Figure 4 It is a schematic flowchart of an ultrasonic image processing method in another embodiment; Figure 5 It is a schematic flowchart of an ultrasonic image processing method in yet another embodiment; Figure 6 Schematic flowchart of constructing a virtual anatomical model based on the position parameters of each of the key anatomical structures in an embodiment; Figure 7 Schematic structural diagram of an electronic device involved in the ultrasonic image processing method in an embodiment of the present application. Detailed implementation manners
[0019] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0020] All terms (including technical and scientific terms) used in the present application have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0021] For example, terms such as "first" and "second" used in the present application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from another element.
[0022] For another example, terms such as "including" and "comprising" used in the present application indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0023] As described above, there have been some studies attempting to improve the quality of ultrasonic images through image processing techniques, such as using traditional filtering or enhancement algorithms for image optimization. However, these methods are often designed for specific ultrasonic imaging problems and are difficult to handle the complex and variable features under different organs and pathological conditions. For example, in abdominal ultrasound, artifacts generated by adipose tissue and air cavities cannot be effectively processed. Also, although these methods can reduce noise, they are prone to losing tiny structures and detailed information during the denoising process, affecting the diagnostic accuracy. For example, in liver nodule detection, excessive smoothing may obscure the true morphology of early tiny tumors and interfere with the accurate judgment of doctors. Therefore, there is currently a technical problem of poor visualization effect in ultrasonic imaging. For this reason, the present application proposes an ultrasonic image processing method, an electronic device, and a storage medium, which can effectively improve the visualization effect of ultrasonic images by performing real-time enhancement and segmentation on ultrasonic images through the TransUNet deep learning model.
[0024] As Figure 1 shown, an ultrasonic image processing method proposed in an embodiment of the present application includes the following steps: Step S10: Obtain an initial ultrasound image and segment the initial ultrasound image into several sub-images.
[0025] In this embodiment, the initial ultrasound image refers to the raw ultrasound image (usually a two-dimensional grayscale image) obtained by an ultrasound device, which is used as the input for a pre-trained regional anesthesia ultrasound image recognition model. The regional anesthesia ultrasound image recognition model is a deep learning model that combines the Transformer and U-Net architectures and is specifically used for real-time enhancement and segmentation of ultrasound images. Specifically, the regional anesthesia ultrasound image recognition model includes a Transformer-based encoder and a U-Net-based decoder. The Transformer-based encoder is stacked by multiple Transformer layers (also called encoder layers). Each Transformer layer contains two main sub-layers: the self-attention mechanism and the feed-forward network. The self-attention mechanism in the Transformer can help the model better understand the relationship between key anatomical structures such as nerves and blood vessels in the ultrasound image and the surrounding tissues, and capture long-range dependence features that are difficult to discover by traditional convolutional neural networks. The feed-forward network in the Transformer can extract local detail features of the image. The U-Net-based decoder is used to integrate global context information and local detail features. By skip connections (concatenating feature maps of different levels in the encoder with the corresponding feature maps in the decoder), local detail information and global context information are continuously integrated, so as to predict the class label corresponding to each pixel (such as nerves, blood vessels, etc.) and generate a high-resolution segmentation mask. The combination of the two enables the model to more precisely depict the boundaries and shapes of key anatomical structures on the basis of obtaining global information when processing ultrasound image enhancement and segmentation tasks.
[0026] Before step S10, a set of regional anesthesia ultrasound sample images manually segmented and labeled by a team of regional anesthesia experts can be obtained. Specifically, the experts will perform precise manual annotation on the ultrasound sample images according to the anatomical structures in the ultrasound sample images, such as nerves, blood vessels, muscles, etc., segment and label different tissue structures, and form labeled training data. Subsequently, these labeled ultrasound sample images are used to train the regional anesthesia ultrasound image recognition model. Through deep learning algorithms, the model can learn the characteristics and patterns of different tissue structures, so as to automatically segment and recognize ultrasound images in subsequent applications.
[0027] After obtaining the initial ultrasound image and before inputting it into the pre-trained regional anesthesia ultrasound image recognition model, some preprocessing can be performed on the initial ultrasound image, such as normalization, resolution adjustment, denoising, or contrast enhancement. Among them, normalization refers to scaling the pixels of the initial ultrasound image to a fixed range (such as 0 to 1). Resolution adjustment refers to adjusting the resolution of the initial ultrasound image to meet the input size requirements of the model.
[0028] Input the preprocessed initial ultrasound image into the pre-trained regional anesthesia ultrasound image recognition model. Among them, based on the received initial ultrasound image, in the encoder stage, the regional anesthesia ultrasound image recognition model cuts the initial ultrasound image into several sub-images. Each sub-image mainly contains all or partial features of one or several parts / organs. There may be overlapping parts between adjacent sub-images cut out, so that the information of one part / organ can be completely included in one or several sub-images.
[0029] Step S20: Extract the global context information and local detail features of the several sub-images.
[0030] In this embodiment, the regional anesthesia ultrasound model extracts the global context information and local detail features of the several sub-images through an encoder based on Transformer. Among them, the global context information refers to the overall features of the entire initial ultrasound image, and these features can reflect the mutual relationship and overall structure between different sub-images of the initial ultrasound image. In the ultrasound image, the global context information helps the regional anesthesia ultrasound image recognition model to subsequently understand the spatial relationship and interaction between key anatomical structures such as nerves and blood vessels and surrounding tissues. For example, through the global context information, the regional anesthesia ultrasound image recognition model can identify the relative positions of nerves and blood vessels, as well as their layout relationships with other tissues such as bones and muscles. This information is crucial for accurately segmenting and recognizing key anatomical structures in the ultrasound image because it provides the overall semantics and structural background of the image. The local detail features refer to the detailed features in each small sub-image of the initial ultrasound image, such as edges, textures, and gray-scale changes. In the ultrasound image, the local detail features help the regional anesthesia ultrasound image recognition model accurately identify and segment the boundaries and internal features of key anatomical structures such as nerves and blood vessels. For example, through the local detail features, the model can identify the fine branches of nerve fibers, as well as the wall thickness and internal blood flow conditions of blood vessels. These detail information are very important for accurately segmenting and recognizing key anatomical structures in the ultrasound image because they provide the fine structure and texture information of the image.
[0031] In this embodiment, the global context information and local detailed features complement each other and jointly provide rich image information for the regional anesthesia ultrasound image recognition model. The global context information helps the regional anesthesia ultrasound image recognition model understand the overall structure and layout of the image, while the local detailed features ensure that the regional anesthesia ultrasound image recognition model can perform fine analysis and recognition on each part of the image. This combination enables the regional anesthesia ultrasound image recognition model to achieve better performance in complex ultrasound image recognition and segmentation tasks.
[0032] In some embodiments, as Figure 2 shown, step S20 includes: Step S21, mapping the several sub-images into vector representations of a preset dimension.
[0033] Mapping each sub-image into a vector representation of a preset dimension is a process of compressing the pixel values of the sub-image into a fixed-length vector. For example, several sub-images (e.g., each sub-image is 16×16 pixels in size) are mapped into vectors of a preset dimension (e.g., 64 dimensions).
[0034] In some embodiments, for each sub-image, a small convolutional neural network (CNN) can be used as a projection head to map the two-dimensional image patch into a one-dimensional vector. This projection head can extract the basic features of the sub-image, such as edges, textures, etc., by learning the weights of different convolutional kernels, and combine these features into a vector representation of a fixed dimension through a fully connected layer.
[0035] In step S21, by mapping the sub-image from the high-dimensional pixel space to the low-dimensional vector space, the dimensionality of the data is reduced, and the complexity of subsequent calculations is lowered. At the same time, this process also performs preliminary feature extraction on the sub-image, converting pixel-level information into vector representations that can reflect the basic features of the sub-image. These vector representations contain important features such as the local texture and grayscale of the sub-image, providing the basic data for subsequent self-attention mechanisms and non-linear transformations.
[0036] Step S22, adding positional encoding corresponding to each sub-image to the vector representation of each sub-image.
[0037] In this embodiment, the positional encoding is a signal or vector used to represent the spatial position of the sub-image in the initial ultrasound image. Since the Transformer model inherently does not have the ability to perceive the order or spatial position of elements in the input sequence, it is necessary to explicitly add positional information to the vector representation of the sub-image. The positional encoding can be a fixed positional encoding based on sine and cosine functions, or a trainable positional encoding obtained through learning.
[0038] For example, for the i-th sub-image, a position encoding vector composed of two-dimensional sine and cosine functions can be generated, where the frequency changes proportionally with the position. This position encoding vector is consistent with the vector representation of the sub-image in dimension, and then the position encoding is combined with the vector representation of the sub-image through simple element-wise addition or concatenation operations. In this way, the vector representation of each sub-image not only contains its own feature information but also contains its spatial position information in the original image.
[0039] In some embodiments, the two-dimensional position coordinates of the sub-image can be first calculated according to the row number and column number of the sub-image in the initial ultrasound image, and then these coordinates are mapped to the position encoding vectors of the corresponding dimensions through sine and cosine functions. For example, for the position (row, col), it can be decomposed into two dimensions of row and column, and the sine and cosine values are calculated respectively, and then concatenated into a long vector as the position encoding.
[0040] In step S22, by adding the position encoding, additional dimensional information is added to the originally identical sub-image features, enabling the regional anesthesia ultrasound image recognition model to distinguish sub-images that were originally difficult to distinguish due to similar texture or gray-scale features, thereby improving the overall distinguishability of the features and the representation ability of the regional anesthesia ultrasound image recognition model.
[0041] Step S23, based on the position encoding corresponding to each block of sub-images, calculates the global context information of the several blocks of sub-images through the multi-head self-attention mechanism.
[0042] In this embodiment, the multi-head self-attention mechanism is one of the core components of the regional anesthesia ultrasound image recognition model and is used to calculate the correlation between different sub-images. In the extraction of the global context information, through the multi-head self-attention mechanism, the vector representation of each block of sub-images is interacted with the vector representations of all other blocks of sub-images, thereby obtaining the global information about the entire initial ultrasound image.
[0043] In some embodiments, based on the position encoding corresponding to each block of sub-images, a set of query vectors (Query), key vectors (Key), and value vectors (Value) corresponding to each block of sub-images are generated through the multi-head self-attention mechanism. By calculating the dot product similarity between the query vector and the key vector, the attention weights between different sub-images are obtained. Among them, the attention weights reflect the correlation between different sub-images, and the value vectors represent the feature content of the sub-images. The attention weights and the value vectors are weighted and summed to obtain the global context information of each block of sub-images.
[0044] Exemplarily, assume that there are N sub-images, and the vector representation of each sub-image is d-dimensional. The multi-head self-attention mechanism divides these vectors into H heads, and calculates the attention weights and weighted sums separately in each head. This is equivalent to extracting global information from different aspects in multiple different subspaces, and finally combining the information of all heads to form a richer and more comprehensive global context information vector.
[0045] In step S23, considering that some important anatomical structures may be distributed in different regions of the initial ultrasound image and there are complex long-distance dependencies between them. For example, a nerve may be associated with distant blood vessels or muscle tissues. The multi-head self-attention mechanism can capture these global dependencies without being restricted by distance, enabling the regional anesthesia ultrasound image recognition model to understand the structure and layout of the entire initial ultrasound image. By integrating the attention information from different sub-images, the regional anesthesia ultrasound image recognition model can generate a feature representation with a global perspective, which is of great significance for the detection of medium and large targets (such as brachial plexus nerves, etc.) and image segmentation tasks in the initial ultrasound image.
[0046] In step S24, based on the global context information, the vector representation of each piece of sub-image is non-linearly transformed through a feed-forward network to obtain the local detail features of the several pieces of sub-images.
[0047] In this embodiment, the feed-forward network consists of two fully connected layers and a non-linear activation function. In the first fully connected layer, the input vector representation (including the global context information processed by the self-attention mechanism) is expanded to a higher dimension, and then through the transformation of the non-linear activation function (such as ReLU or GELU), non-linear features are introduced. Then, in the second fully connected layer, the features are shrunk back to the original dimension. During the entire transformation process, the feed-forward network can learn more complex feature mappings, so as to better capture the local detail features in the sub-images.
[0048] Exemplarily, assume that the dimension of the input vector is d_model. In the first fully connected layer of the feed-forward network, it is expanded to a dimension of 4d_model. Then through the ReLU activation function, the features are non-linearly projected and folded in the high-dimensional space. Finally, through the second fully connected layer, the features are shrunk back to the d_model dimension to obtain the local detail features after non-linear transformation.
[0049] In step S24, due to the diverse shapes and complex textures of the anatomical structures in the ultrasound image, simple linear transformation may not be able to fully express these features. Through the non-linear transformation of the feed-forward network, the features can be more finely adjusted and combined, so as to extract richer local detail features.
[0050] Step S30: Generate a target feature map based on the global context information and the local detail features.
[0051] Separate global context information and local detail features can only reflect part of the features of the initial ultrasound image. Global information provides the overall structure and semantics, while local details focus on fine textures and edges. By integrating the global context information and local detail features into a target feature map through a U-Net-based decoder, it can not only contain the macroscopic anatomical structure layout but also retain the subtle tissue texture changes, which is crucial for accurately identifying and segmenting complex structures in ultrasound images.
[0052] In some embodiments, as Figure 3 shown, step S30 includes: Step S31: Fuse the global context information and the local detail features to generate a fused feature map.
[0053] In this embodiment, the global context information and the local detail features are features at two different levels, which respectively contain the overall structure and local details of the image. To generate a richer feature representation, the global context information and the local detail features are fused to generate a fused feature map.
[0054] In some embodiments, fusion can be performed through simple element-wise weighted summation or feature concatenation. For example, assuming the dimension of the global context information is d_g and the dimension of the local detail features is d_l, then a fused feature map can be generated through simple element-wise weighted summation, that is, fused feature map = global context information × w_g + local detail features × w_l, where w_g and w_l are learnable weights used to adjust the relative importance of the two features.
[0055] In some embodiments, more complex fusion strategies can also be used, such as using a small neural network to learn how to fuse these two features. For example, the global context information and the local detail features can be concatenated and then passed through several fully connected layers or convolutional layers to generate a fused feature map.
[0056] Step S32: Input the fused feature map into an encoder network to generate a multi-level initial feature map.
[0057] In this embodiment, the U-Net-based decoder applies an image segmentation network structure, including two parts: an encoder network and a decoder network. Among them, the encoder network is a multi-layer neural network composed of multiple convolutional layers and pooling layers, which is used to perform multi-level feature extraction and compression on the input fused feature map to generate initial feature maps at different levels.
[0058] In some embodiments, the fused feature map is subjected to multiple convolution and downsampling operations through an encoder network to gradually extract features at different levels. For example, the first convolution block may include two or three convolutional layers for extracting low-level local features, and then a pooling layer is used to halve the spatial resolution of the fused feature map. The second convolution block further extracts deeper features on this basis and again reduces the spatial resolution of the fused feature map through another pooling layer. This process is repeated multiple times until a multi-level initial feature map is generated.
[0059] In this way, the encoder network can generate a multi-level initial feature map, where the initial feature maps at different levels are inversely proportional in terms of spatial resolution and semantic level. For example, the low-level initial feature map may have a low semantic level but a high spatial resolution (containing rich local texture and edge information), while the high-level initial feature map may contain higher-level semantic information, such as the overall shape and distribution of anatomical structures, but a low spatial resolution.
[0060] In step S32, the multi-level initial feature map provides a rich feature basis for subsequent feature fusion and decoding processes. Among them, the low-level initial feature map can be used to restore the details of the image, while the high-level initial feature map can be used to guide the overall segmentation and recognition.
[0061] In step S33, through skip connections, the low-level initial feature map is passed to the corresponding layer of the decoder network to obtain low-level features.
[0062] In this embodiment, skip connection is a technique that directly passes the low-level initial feature map in the encoder network to the corresponding layer of the decoder network to retain and utilize low-level detailed features.
[0063] In step S34, the high-level initial feature map is upsampled to obtain high-level features.
[0064] The high-level initial feature map usually has a low spatial resolution but a high semantic level. To generate the final segmentation mask or target feature map, these high-level feature maps need to be upsampled to restore to a spatial resolution close to that of the initial ultrasound image.
[0065] In some embodiments, upsampling can be performed by methods such as bilinear interpolation, nearest neighbor interpolation, and transposed convolution. Among them, bilinear interpolation is an upsampling method based on interpolation, which increases the spatial resolution by inserting new pixel values in the high-level initial feature map, and these pixel values are calculated by linear interpolation. Nearest neighbor interpolation fills new spatial positions by copying the nearest pixel values. Transposed convolution is an upsampling method based on convolution, which performs upsampling by learning convolution kernels and has stronger expression power. For example, a transposed convolution layer can be used to upsample the high-level initial feature map while increasing the number of channels, thereby generating high-level features with higher spatial resolution and more channels.
[0066] In step S34, since the spatial resolution of the high-level initial feature map is low, it cannot be directly used to generate a high-resolution segmentation mask or target feature map. By upsampling, the spatial resolution of the high-level initial feature map can be restored to a level close to that of the initial ultrasound image, thereby providing a more refined feature representation for subsequent processing.
[0067] Step S35, fusing the low-level features and the high-level features to generate a target feature map.
[0068] The low-level features that mainly contain local detail information and the high-level features that mainly contain global semantic information are fused again, which is similar to the secondary purification and enhancement of features. It can explore more potential connections hidden in the global and local information, making the features richer and more discriminative. In some embodiments, fusion can be performed by element-wise weighted summation, feature concatenation, and multi-layer perceptron (MLP). For example, low-level features and high-level features can be element-wise weighted summation, where the weights can be learned by a small neural network. Alternatively, the two features can be concatenated and then fused through several convolutional layers or fully connected layers.
[0069] In the U-Net architecture of this embodiment, the upsampled high-level features and the low-level features transmitted through the skip connection can be concatenated and then processed by a small convolutional network. This small convolutional network can include several convolutional layers and nonlinear activation functions to learn how to fuse the two features.
[0070] Step S40, generating a type label for each pixel in the initial ultrasound image by performing a convolution operation on the target feature map.
[0071] In this embodiment, the type label refers to a label used to identify the tissue structure type to which each pixel belongs in an ultrasound image. These labels typically include nerves, blood vessels, muscles, fascia, and bones, etc. The type labels are obtained by segmenting and recognizing the ultrasound image, and each pixel is assigned to the corresponding tissue structure category, thereby achieving accurate annotation and segmentation of the ultrasound image.
[0072] It should be understood that during the training phase, the regional anesthesia ultrasound image recognition model learns the characteristics and patterns of different tissue structures (such as nerves, blood vessels, muscles, etc.) through a large number of manually annotated ultrasound sample images. These characteristics and patterns are encoded in the weights of the convolutional kernels, enabling the regional anesthesia ultrasound image recognition model to automatically extract and recognize these features during the convolution operation.
[0073] In some embodiments, when performing a convolution operation on the target feature map through a convolutional layer, multiple convolutional kernels (filters) contained in the convolutional layer slide on the target feature map to extract the feature vector of each pixel. After the last convolutional layer, the feature vector of each pixel will be passed to a classifier. The classifier is usually a fully connected layer or a convolutional layer, which is used to output the probability distribution of each pixel. The classifier marks the type with the highest probability for each pixel as the type label by calculating the probability distribution of each pixel belonging to various types (such as nerves, blood vessels, muscles, fascia, and bones). For example, if a pixel has the highest probability of belonging to a nerve, then the pixel will be marked as a nerve.
[0074] Step S50: Based on the type labels of each pixel, construct a segmentation mask for the initial ultrasound image, and superimpose the segmentation mask on the initial ultrasound image to generate a target ultrasound image.
[0075] In this embodiment, the segmentation mask is a color marker that matches the size of the initial ultrasound image and is used to mark specific regions or structures in the image. For example, nerves are red (specific RGB values), blood vessels are blue, etc.
[0076] In some embodiments, when fusing the segmentation mask with the initial ultrasound image, image fusion techniques such as opacity blending or intensity adjustment can be used. Opacity blending makes the mask cover the original image in a semi-transparent form by specifying the transparency of the mask, thereby simultaneously showing the texture details of the original image and the key regions marked by the mask.
[0077] Taking opacity blending as an example, when the transparency of the mask is set to 50%, the doctor can clearly see the grayscale information of the initial ultrasound image and at the same time observe the segmented results with color markings, effectively avoiding the complete occlusion of the original image by the mask information and achieving an organic visual combination of the two.
[0078] In step S50, by presenting the type tags at the pixel level in an intuitive image form, it is convenient for subsequent visualization and analysis, clearly identifying the distribution of various anatomical structures in the initial ultrasound image, and providing doctors with accurate lesion and key structure location information.
[0079] In some embodiments, according to the type tag distribution in the target ultrasound image, the target anatomical structure is determined. According to the tissue characteristics of the target anatomical structure, the target pressure range of the ultrasound probe is determined. For example, for softer nerve tissue, the pressure range may be smaller to avoid damage. The real-time pressure of the ultrasound probe is obtained, and a correction prompt is output when the real-time pressure exceeds the target pressure range. Here, the type tag distribution refers to the distribution of type tags of each pixel in the target ultrasound image, and these tags are used to identify different tissue structures, such as nerves, blood vessels, muscles, etc. The target anatomical structure refers to the anatomical structure that needs to be focused on in the target ultrasound image, usually determined according to the type tag distribution, such as nerves, blood vessels, etc. The tissue characteristics refer to the physical and physiological characteristics of the target anatomical structure, such as soft tissue, hard tissue, etc. The target pressure range refers to the pressure range that the ultrasound probe needs to reach when applying pressure to ensure that the characteristics of the target anatomical structure can be accurately captured. The real-time pressure refers to the pressure value measured by the ultrasound probe in real-time during actual operation. By determining the target pressure range, the pressure application process of the ultrasound probe can be optimized, improving the accuracy and safety of pressure application.
[0080] In the ultrasound image processing method proposed in the embodiments of the present application, on the one hand, by combining global context information and local detail features, the anatomical relationship between the target nerve and surrounding tissues can be captured more accurately, thus significantly improving the accuracy of image segmentation. On the other hand, compared with traditional region-based or edge-based processing methods, the pixel-level annotation method can describe the distribution of different tissues in more detail. On the third hand, during the generation of the target ultrasound image, by superimposing the segmentation mask on the initial ultrasound image, not only can the information of the initial ultrasound image be retained, but also the key anatomical structures (such as nerves and blood vessels) can be displayed through visual markers, significantly improving the visualization effect of ultrasound imaging.
[0081] In one embodiment, as Figure 4 shown, after step S50, it further includes: Step S60, based on a preset regional anesthesia standard image set, determine the standard image subset corresponding to the target ultrasound image.
[0082] In this embodiment, the preset regional anesthesia standard image set is a group of high-quality ultrasound images that have been marked and verified by experts. These images cover various common regional anesthesia scenarios and anatomical structures. Each regional anesthesia scenario or anatomical structure has a corresponding series of consecutive ultrasound image frames, and a dynamic standard image range is constructed to capture the anatomical structure changes at different time points, ensuring that the standard images can reflect the anatomical states at different time points. Each standard image has clear anatomical edges, good contrast, and accurate annotations, and can be used as an object for evaluation and comparison to evaluate the quality and accuracy of the target ultrasound image. By comparing with the standard image set, the scanning accuracy and quality of the target ultrasound image can be determined.
[0083] Through the feature matching algorithm, the key features of the target ultrasound image are compared with the images in the regional anesthesia standard image set. The key features can include features such as the edges, textures, and shapes of the images, and can also include key anatomical structures (such as nerves, blood vessels, muscles, etc.) within the images.
[0084] According to the comparison results (similarity calculation results), the standard images that are most similar to the target ultrasound image (higher than the preset similarity threshold or the preset number in the positive arrangement of similarity) are selected as the standard image subset. The number of images in the standard image subset is an integer greater than one, and the specific number depends on the application scenario and computing resources.
[0085] Step S70, perform fusion processing on each image in the standard image subset to obtain a fused standard image.
[0086] Adopt multi-scale deep learning (such as neural network models like U-net, etc.) to perform fusion processing on each image in the standard image subset. By fusing each image in the standard image subset, the anatomical structure information at different time points can be integrated to generate a fused standard image containing richer features. This includes not only static anatomical structures but also dynamic change information, such as the movement or deformation of nerves and blood vessels.
[0087] In step S70, the fused standard image can provide more comprehensive anatomical information, which helps to more accurately evaluate the scanning accuracy of the target ultrasound image. By including the features at different time points, the fused standard image can better reflect the dynamic changes of the anatomical structure, thereby improving the accuracy and reliability of the evaluation.
[0088] Step S80, match the fused standard image and the target ultrasound image to determine the scanning accuracy of the target ultrasound image.
[0089] Use the image matching algorithm to match the fused standard image with the target ultrasound image. The matching algorithm can be based on feature point matching, cross-correlation coefficient, or other similarity measurement methods.
[0090] According to the matching result (the calculation result of the matching algorithm), determine whether the scanning accuracy of the target ultrasonic image reaches a preset matching degree threshold. If it reaches the preset matching degree threshold, determine that the current scanning operation is qualified; if it is lower than the preset matching degree threshold, determine that the current scanning operation is unqualified.
[0091] In some embodiments, when the scanning accuracy is lower than the preset matching degree threshold, as Figure 5 shown, after the step S80, the following steps are further included: Step A10, obtain a set of consecutive target ultrasonic images corresponding to consecutive frames of the initial ultrasonic image.
[0092] In the actual operation of ultrasound-guided regional anesthesia, the ultrasound device continuously acquires ultrasonic images at a high frame rate. These continuously acquired image frames are continuous in the time series and can reflect the dynamic changes of the anatomical structure. In this embodiment, the consecutive frames of the initial ultrasonic image refer to the ultrasonic image frames that are continuous with the initial ultrasonic image in the time series. The set of consecutive target ultrasonic images refers to the image set composed of the initial ultrasonic image and its consecutive frames.
[0093] Step A20, based on the type label distribution of each frame of image in the set of consecutive target ultrasonic images, determine the key anatomical structures of each frame of image in the set of consecutive target ultrasonic images.
[0094] For each frame of image in the set of consecutive target ultrasonic images, analyze its type label distribution. The type labels include nerves, blood vessels, muscles, fascia, and bones, etc. By analyzing the distribution of these type labels, the positions and proportions of the key anatomical structures in each frame of image can be determined.
[0095] Step A30, construct a virtual anatomical model based on the position parameters of each of the key anatomical structures.
[0096] In this embodiment, the position parameters include center coordinates, volume information, direction vectors, and depth information. According to the boundary information of each of the key anatomical structures, calculate the center coordinates and volume information of each of the key anatomical structures; by respectively fitting the main axes of each of the key anatomical structures, determine the direction vectors of each of the key anatomical structures; according to the original gray value of the target ultrasonic image, determine the depth information of each of the key anatomical structures.
[0097] In some embodiments, as Figure 6 shown, the step A30 includes: Step A31: Integrate the central coordinates, volume information, direction vectors, and depth information of each key anatomical structure in each frame of the continuous target ultrasound image set to generate a spatial parameter set for each key anatomical structure.
[0098] Specifically, integrate the central coordinates, volume information, direction vectors, and depth information of each key anatomical structure in each frame of the continuous target ultrasound image set into a unified data structure to form a spatial parameter set for each key anatomical structure. The data structure can be a list, an array, or a dictionary, where each element corresponds to a key anatomical structure and contains all its position parameters.
[0099] By aggregating the position parameters scattered in different image frames, more comprehensive and accurate analysis can be performed for each key anatomical structure.
[0100] Step A32: Based on the spatial parameter set, map each key anatomical structure into a three-dimensional space to form a spatial distribution representation of each key anatomical structure.
[0101] Based on the spatial parameter set, map the key anatomical structures into a unified three-dimensional or pseudo-three-dimensional space coordinate system through normalized coordinate transformation or affine transformation. In the three-dimensional or pseudo-three-dimensional space, each key anatomical structure is represented as a point or a region, and these points or regions are used to construct the spatial distribution representation of the key anatomical structures, showing their positions and shapes in the three-dimensional space.
[0102] Step A33: Render the spatial distribution representation of each key anatomical structure to generate a virtual anatomical model.
[0103] In this embodiment, the virtual anatomical model refers to a three-dimensional or pseudo-three-dimensional model constructed based on the spatial parameter set of the key anatomical structures, which can intuitively display the spatial distribution and interrelationships of the anatomical structures. By using computer graphics techniques (such as OpenGL or DirectX), present the spatial distribution representation of each key anatomical structure in a visual manner, and specific colors can be assigned to each key anatomical structure to facilitate the distinction of different tissue types, thus forming a virtual anatomical model.
[0104] In some embodiments, the transparency and color contrast between the key anatomical structures in the virtual anatomical model can also be adjusted to form an optimized anatomical model, so as to more clearly display the hierarchical relationship of the key anatomical structures.
[0105] Step A40: Based on the virtual anatomical model and the fused standard image, determine the target scanning parameters of the ultrasound probe.
[0106] In this embodiment, the target scanning parameters refer to the parameters that the ultrasonic probe needs to achieve during the scanning of the fused standard image, including the position, angle, etc. of the ultrasonic probe.
[0107] Align and match the virtual anatomical model with the fused standard image to determine the sectional image in the virtual anatomical model that matches the fused standard image; based on the first position and the first angle of the sectional image in the virtual anatomical model, through the coordinate transformation algorithm, convert the first position of the matched sectional image into the second position of the ultrasonic probe; and through the direction vector transformation algorithm, convert the first angle of the matched sectional image into the second angle of the ultrasonic probe. Among them, the second position and the second angle are the target scanning parameters that the ultrasonic probe needs to achieve.
[0108] Step A50, obtain the actual physical parameters of the ultrasonic probe, and output a scanning guidance path according to the actual physical parameters and the target scanning parameters.
[0109] In this embodiment, the actual physical parameters refer to the physical parameters of the ultrasonic probe during actual operation, including the current position, angle, motion state, etc. The scanning guidance path is a path generated from the actual physical parameters to the target scanning parameters, and is used to guide the doctor to adjust the position and angle of the probe.
[0110] In some embodiments, the position, angle, motion state and other actual physical parameters of the ultrasonic probe can be obtained in real time through the sensors of the ultrasonic device. According to the actual physical parameters and the target scanning parameters, a scanning guidance path from the actual physical parameters to the target scanning parameters is generated. Through the scanning guidance path, the number of trial-and-error times of the doctor during the scanning process can be reduced, and the operation efficiency and accuracy can be improved.
[0111] In the ultrasonic image processing method proposed by the embodiments of the present application, in the first aspect, by continuously collecting high-frame-rate ultrasonic image frames and forming a continuous set of target ultrasonic images, the accurate capture of the dynamic changes of anatomical structures is ensured. In the second aspect, by analyzing the distribution of type tags (such as nerves, blood vessels, muscles, etc.) in each frame of the image, the positions and proportions of key anatomical structures in each frame of the image are determined, which not only improves the accuracy of positioning but also provides reliable basic data for subsequent steps. In the third aspect, based on the position parameters of these key anatomical structures (including central coordinates, volume information, direction vectors, and depth information), a virtual anatomical model is constructed. This model forms a more comprehensive and accurate set of spatial parameters by integrating the position parameters in multiple frames of images, and maps these key anatomical structures into three-dimensional space for rendering, intuitively showing their spatial distribution and interrelationships. In the fourth aspect, by aligning and matching the virtual anatomical model with the fused standard image, the target scanning parameters (including position and angle) of the ultrasonic probe are determined, thereby guiding the doctor to adjust the probe to the optimal operating position. In the fifth aspect, according to the actual physical parameters and target scanning parameters of the ultrasonic probe, a scanning guidance path is generated, reducing the number of trial-and-error operations of the doctor and significantly improving the efficiency and accuracy of the scanning process.
[0112] In one embodiment, a computer-readable storage medium is provided, on which executable instructions are stored. When the instructions are executed by a processor, the processor executes the steps in the above method embodiments.
[0113] In one embodiment, an electronic device is further provided, including one or more processors; a memory, and one or more programs are stored in the memory. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors execute the steps in the above method embodiments.
[0114] In one embodiment, as Figure 7 shown, it shows a schematic structural diagram of an electronic device for implementing the embodiments of the present application. The electronic device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0115] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as required. A removable medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 710 as required so that a computer program read therefrom is installed into the storage section 708 as required.
[0116] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product including a computer-readable medium carrying instructions. In such an embodiment, the instructions can be downloaded and installed from a network via the communication section 709 and / or installed from the removable medium 711. When the instructions are executed by a central processing unit (CPU) 701, the various method steps described in the present application are executed.
[0117] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0118] In addition, those skilled in the art can understand that although some of the embodiments herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present application and forms different embodiments. For example, any one of the embodiments claimed in the present application can be used in any combination. The information disclosed in this background art section is only intended to deepen the understanding of the overall background art of the present application, and should not be regarded as an admission or any form of suggestion that this information constitutes prior art already known to those skilled in the art.
Claims
1. An ultrasonic image processing method, characterized in that, The method includes: Obtaining an initial ultrasound image and segmenting the initial ultrasound image into a plurality of sub-images; Extracting the global context information and local detail features of the plurality of sub-images; Generating a target feature map based on the global context information and the local detail features; Generating a type label for each pixel in the initial ultrasound image by performing a convolution operation on the target feature map, where the type label includes one or more of nerve, blood vessel, muscle, fascia, and bone; Constructing a segmentation mask of the initial ultrasound image based on the type label of each pixel and superimposing the segmentation mask on the initial ultrasound image to generate a target ultrasound image.
2. The ultrasonic image processing method according to claim 1, characterized in that, The extracting the global context information and local detail features of the plurality of sub-images includes: Mapping the plurality of sub-images into vector representations of a preset dimension; Adding position encodings corresponding to the sub-images to the vector representations of each sub-image; Calculating the global context information of the plurality of sub-images through a multi-head self-attention mechanism based on the position encodings corresponding to each sub-image; Performing a non-linear transformation on the vector representations of each sub-image through a feed-forward network based on the global context information to obtain the local detail features of the plurality of sub-images.
3. The ultrasonic image processing method according to claim 1, characterized in that, The generating a target feature map based on the global context information and the local detail features includes: Fusing the global context information and the local detail features to generate a fused feature map; Inputting the fused feature map into an encoder network to generate multi-level initial feature maps; Transmitting the low-level initial feature maps to the corresponding levels of the decoder network through skip connections to obtain low-level features; Upsampling the high-level initial feature maps to obtain high-level features; Fusing the low-level features and the high-level features to generate a target feature map.
4. The ultrasonic image processing method according to claim 1, characterized in that, After constructing the segmentation mask of the initial ultrasound image based on the type label of each pixel and superimposing the segmentation mask on the initial ultrasound image to generate a target ultrasound image, it further includes: Determining a subset of standard images corresponding to the target ultrasound image based on a preset regional anesthesia standard image set; Performing a fusion process on each image in the subset of standard images to obtain a fused standard image; Matching the fused standard image and the target ultrasound image to determine the scanning accuracy of the target ultrasound image.
5. The ultrasonic image processing method according to claim 4, characterized in that, After matching the fused standard image and the target ultrasound image to determine the scanning accuracy of the target ultrasound image, it further includes: When the scanning accuracy is lower than a preset matching degree threshold, obtaining a set of consecutive target ultrasound images corresponding to consecutive frames of the initial ultrasound image; Determining the key anatomical structures of each frame of image in the set of consecutive target ultrasound images based on the type label distribution of each frame of image in the set of consecutive target ultrasound images; Constructing a virtual anatomical model based on the position parameters of each of the key anatomical structures; Determining the target scanning parameters of the ultrasound probe based on the virtual anatomical model and the fused standard image; Obtain the actual physical parameters of the ultrasonic probe, and output a scanning guidance path according to the actual physical parameters and the target scanning parameters.
6. The ultrasonic image processing method according to claim 5, characterized in that, The position information includes center coordinates, volume information, direction vectors, and depth information. Before constructing the virtual anatomical model based on the position parameters of each of the key anatomical structures, it further includes: Calculate the center coordinates and volume information of each of the key anatomical structures according to the boundary information of each of the key anatomical structures; Determine the direction vector of each of the key anatomical structures by respectively fitting the main axes of each of the key anatomical structures; Determine the depth information of each of the key anatomical structures according to the original gray value of the target ultrasonic image.
7. The ultrasonic image processing method according to claim 6, wherein The constructing of the virtual anatomical model based on the position parameters of each of the key anatomical structures further includes: Integrate the center coordinates, volume information, direction vectors, and depth information of each of the key anatomical structures in each frame of the continuous target ultrasonic image set to generate a set of spatial parameters for each of the key anatomical structures; Based on the set of spatial parameters, map each of the key anatomical structures into a three-dimensional space to form a spatial distribution representation of each of the key anatomical structures; Render the spatial distribution representation of each of the key anatomical structures to generate a virtual anatomical model.
8. The ultrasonic image processing method according to any one of claims 1 to 7, characterized in that After constructing the segmentation mask of the initial ultrasonic image based on the type label of each pixel and superimposing the segmentation mask on the initial ultrasonic image to generate a target ultrasonic image, it further includes: Determine the target anatomical structure according to the type label distribution in the target ultrasonic image; Determine the target pressure range of the ultrasonic probe according to the tissue characteristics of the target anatomical structure; Obtain the real-time pressure of the ultrasonic probe, and output a correction prompt when the real-time pressure exceeds the target pressure range.
9. An electronic device, characterized in that, Includes: One or more processors; A memory for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors execute the ultrasonic image processing method according to any one of claims 1 to 8.
10. A storage medium, characterized in that, The storage medium stores executable instructions, and when the instructions are executed by a processor, the processor executes the ultrasonic image processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method and system for quickly acquiring cardiac ultrasonic standard scanning section
CN111449684A
Multi-modal fusion imaging method and device and storage medium
CN113129342A
Ultrasonic image quantification method based on interactive fusion Transform
CN114863111A
Space guidance method for ultrasonic image and related device
CN116563191A
Positioning assistance system, method and equipment for transesophageal echocardiography and medium
CN117582251A