Method and Application for Implementing Ultrasound-Guided Assisted Popliteal Sciatic Nerve Block Based on U-Net

Through the improved U-Net model, combined with multi-scale hollow convolution and spatial attention mechanism, efficient identification and positioning of target tissues in popliteal ultrasound images is achieved, solving the problem of insufficient accuracy and safety of popliteal sciatic nerve block in the prior art, and improving the success rate and safety of popliteal sciatic nerve block.

CN119992189BActive Publication Date: 2025-07-25THE NAVAL MEDICAL UNIV OF PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510068897.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-07-25
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

Existing ultrasound image recognition methods are difficult to accurately identify sciatic nerves and muscles in the popliteal fossa area, resulting in insufficient accuracy and safety of sciatic nerve block in popliteal fossa.

Method used

The improved U-Net model is adopted, combining multi-scale hollow convolution, channel attention mechanism and spatial attention mechanism to build an encoder and decoder, and the feature fusion of popliteal ultrasound images is carried out through the jump connection structure to achieve automatic identification and positioning of the target tissue.

Benefits of technology

It improves the accuracy of ultrasonic image segmentation recognition of popliteal fossa, can accurately identify and locate target tissues such as biceps femoris, tibial nerve, and common peroneal nerve, which enhances the success rate and safety of sciatic nerve block of popliteal fossa.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992189B_ABST
    Figure CN119992189B_ABST
Patent Text Reader

Abstract

The present invention provides a method and application for realizing ultrasound-guided assisted popliteal sciatic nerve block based on U-Net, which relates to the technical field of medical image processing. The method includes: constructing a popliteal ultrasound data set and a U-Net model suitable for identifying target tissues in popliteal ultrasound images; the U-Net model includes an encoder, a decoder, and a skip connection structure connecting the aforementioned encoder and decoder; wherein, the encoder includes a plurality of encoder blocks; in each encoder block, an efficient cascaded multi-scale dilated convolution ECMAC and an ESAM spatial attention mechanism are sequentially executed; an encoder cross-region feature fusion structure is arranged in the encoder; the decoder includes a plurality of decoder blocks with an ESAM spatial attention mechanism; a decoder cross-region feature fusion structure is arranged in the decoder; collecting popliteal ultrasound images through an ultrasound probe device; using the trained U-Net model to perform segmentation and recognition on the collected popliteal ultrasound images to extract the image features of the target tissues in the aforementioned popliteal ultrasound images. The present invention can improve the accuracy of popliteal ultrasound image segmentation and recognition through the improved U-net model, and provides an analysis tool for realizing ultrasound-guided assisted popliteal sciatic nerve block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and particularly to a method for improving ultrasound-guided popliteal sciatic nerve block based on U-Net. Background Art

[0002] The popliteal fossa is a diamond-shaped area behind the knee joint, containing many important anatomical structures, such as the sciatic nerve, popliteal artery, popliteal vein, and popliteal lymph nodes. Ultrasonic examination of the popliteal fossa area can help doctors evaluate the status of these structures for relevant diagnosis or operations.

[0003] As an example, there are the popliteal artery and popliteal vein in the popliteal fossa area. Ultrasonic images can be used to evaluate the status of blood vessels, such as checking for thrombosis, arteriosclerosis, vascular stenosis, etc. When there may be a mass or lymph node enlargement in the popliteal fossa area, ultrasound can be used to evaluate the nature of these abnormal structures to assist in diagnosing diseases such as infections and tumors. Also, for example, soft tissues in the popliteal fossa, such as muscles and ligaments, can also be examined by ultrasound to assist in diagnosing injuries or lesions of the knee joint or surrounding tissues.

[0004] Popliteal sciatic nerve block (PSNB), as a common anesthetic technique, is widely used in lower limb surgeries. To improve the block accuracy, ultrasonic imaging technology is often used to guide and locate the nerve. However, due to limited ultrasonic image quality (such as noise, high reflection, etc.), accurately identifying the position of the sciatic nerve is challenging.

[0005] During the anesthetic process, popliteal fossa ultrasonic images can assist anesthesiologists in locating the sciatic nerve and its branches (tibial nerve and common peroneal nerve), providing precise guidance for popliteal sciatic nerve block. Nerve block under ultrasound guidance can help anesthesiologists accurately inject anesthetic drugs, thereby achieving effective local anesthesia.

[0006] In existing methods, manual analysis of ultrasonic images has low efficiency and strong subjectivity. Traditional nerve image recognition algorithms are difficult to accurately extract nerve contours in complex ultrasonic backgrounds.

[0007] Therefore, the present invention provides a method and application for improving ultrasound-guided popliteal sciatic nerve block based on U-Net to automatically identify and locate the sciatic nerve and muscles in the popliteal fossa area, thereby improving the success rate and safety of nerve block, which is an urgent technical problem to be solved currently. Summary of the Invention

[0008] The objective of the present invention is to overcome the deficiencies of the prior art and provide a method and application for improving ultrasound-guided popliteal sciatic nerve block based on U-Net. The present invention can perform image processing on popliteal ultrasound images using an improved U-net model, and then automatically identify and locate the sciatic nerve and muscles in the popliteal region to improve the success rate and safety of nerve block.

[0009] To solve the existing technical problems, the present invention provides the following technical solutions:

[0010] A method for realizing ultrasound-guided popliteal sciatic nerve block based on U-Net, including:

[0011] Construct a popliteal ultrasound dataset and a U-Net model suitable for identifying target tissues in popliteal ultrasound images; the popliteal ultrasound dataset includes multiple popliteal ultrasound images processed by data cropping, data marking, and data augmentation operations; the U-Net model includes an encoder, a decoder, and a skip connection structure connecting the aforementioned encoder and decoder; wherein,

[0012] The encoder includes multiple cascaded encoder blocks; in each encoder block, an efficient cascaded multi-scale dilated convolution ECMAC and a spatial attention mechanism ESAM are sequentially executed; wherein, the efficient cascaded multi-scale dilated convolution ECMAC is composed of a multi-scale dilated convolution and a channel attention mechanism ECA; the encoder can respectively capture the image features presented by the target tissues in the aforementioned popliteal ultrasound images at different scales; corresponding to the cascaded encoder blocks, shallow feature maps of different scales can be respectively obtained;

[0013] An encoder cross-region feature fusion structure is arranged in the encoder, and the encoder cross-region feature fusion structure can perform convolution processing on the shallow feature maps of different scales respectively obtained through the cascaded encoder blocks and perform feature fusion with the deep feature maps obtained in the decoder;

[0014] The decoder includes multiple decoder blocks with an ESAM spatial attention mechanism; each decoder block can correspondingly output a deep feature map;

[0015] A decoder cross-region feature fusion structure is arranged in the decoder, and the decoder cross-region feature fusion structure can perform transposed convolution processing on the deep feature maps output by each layer and perform feature fusion with the shallow feature maps in the encoder;

[0016] The skip connection structure can correspondingly splice the aforementioned encoder blocks and decoder blocks;

[0017] Collect popliteal ultrasound images through an ultrasound probe device;

[0018] Use the trained U-Net model to segment and recognize the collected popliteal fossa ultrasound images, so as to extract the image features of the target tissue in the aforementioned popliteal fossa ultrasound images.

[0019] Furthermore, the target tissue includes at least one of biceps femoris, tibial nerve, common peroneal nerve, artery and semimembranosus.

[0020] The U-Net model uses a preset proportion of popliteal fossa ultrasound images in the popliteal fossa ultrasound dataset as training samples and inputs them into the model for training.

[0021] After the training is completed, the remaining proportion of popliteal fossa ultrasound images in the aforementioned popliteal fossa ultrasound dataset are used as test samples and input into the model for testing, and the dice score is used as an evaluation index to measure the performance of the U-Net model in the ultrasound popliteal fossa image segmentation task.

[0022] Furthermore, when the multi-scale dilated convolution is executed, it specifically includes the steps of:

[0023] For the input feature map X ∈ R C×H×W , where C, H, and W are the number of channels, height, and width respectively, let the input feature map X ∈ R C×H×W successively pass through three convolutions with the same kernel size of 3×3 and dilation rates of 1, 2, and 5 respectively, and after each convolution, the BN layer and ReLU activation function are used for processing;

[0024] Use residual connection to add the feature maps output after the first convolution and the second convolution element-wise to supplement the details of the image features;

[0025] The feature map obtained by element-wise addition is then concatenated with the feature map after the third convolution in channels to form a new feature map X1 ∈ R 2C×H×W .

[0026] Furthermore, the execution of the ECA channel attention mechanism includes the steps of:

[0027] For the concatenated feature map X1 ∈ R 2C×H×W , through global average pooling, the average value of each channel is obtained:

[0028]

[0029] where z c represents the average value of all pixel points on the c-th channel, and X c,i,j refers to the pixel value of the feature map X at the c-th channel, height i, and width j;

[0030] For the average value z c ∈R C, a convolution operation with a one-dimensional convolution kernel of size k is used to obtain s = Conv1D(z, k), where k is a positive integer;

[0031] The result s obtained by the convolution operation for each channel is passed through the Sigmoid activation function to obtain the corresponding attention weight α c = σ(s); where σ(·) represents the Sigmoid activation function;

[0032] Apply the aforementioned attention weight α c to the aforementioned concatenated feature map X1 and re-weight each channel to obtain where, is the pixel value at channel c, height i, and width j in the re-weighted feature map X1.

[0033] Furthermore, the execution of the ESAM spatial attention mechanism specifically includes the steps:

[0034] For the input feature map P ∈ R 2C×H×W After passing through max-pooling and average-pooling respectively, and compressing along the channel dimension, the feature map P is obtained through max-pooling and 1×1 convolution M ∈ R 1×H×W , and the feature map obtained through average-pooling and 1×1 convolution is P A ∈ R 1×H×W ; At the same time, the input feature map P is used with depthwise separable convolution to obtain the feature map P D ∈ R 1×H×W ;

[0035] Concatenate the aforementioned feature maps P M ∈ R 1×H×W , P A ∈ R 1×H×W and P D ∈ R 1×H×W on the channel dimension to obtain the feature fusion map F = concat(P M , P A , P D );

[0036] Use a 7×7 convolution on the aforementioned feature fusion map F to obtain the feature fusion map Conv(F) ∈ R 1×H×W ; Pass the feature fusion map Conv(F) ∈ R 1×H×W through the Sigmoid activation function to obtain the attention map M = σ(Conv(F)), M ∈ R 1×H×W ;

[0037] Multiply the input feature map P element-wise with the aforementioned attention map M to obtain the weighted feature map P' = M ⊙ P; where ⊙ represents element-wise multiplication;

[0038] Apply a depthwise separable convolution to the weighted feature map P' to obtain the feature map P * ∈R C×H×W .

[0039] Furthermore, the encoder cross-region feature fusion structure can increase the number of channels while reducing the size of the aforementioned shallow feature maps through step-by-step 3×3 convolutions until the size and number of channels of the aforementioned shallow feature maps are the same as those of any deep feature map in the decoder, and then perform feature fusion between the aforementioned shallow feature map and the aforementioned deep feature map.

[0040] Furthermore, the decoder cross-region feature fusion structure can perform feature fusion between the aforementioned deep feature map and the aforementioned shallow feature map when the size and number of channels of the aforementioned deep feature map are made consistent with those of any shallow feature map through step-by-step upsampling operations.

[0041] An apparatus for realizing ultrasound-guided assisted popliteal sciatic nerve block based on U-Net, comprising:

[0042] A data and model construction unit for constructing a popliteal ultrasound data set and a U-Net model suitable for identifying target tissues in popliteal ultrasound images; the popliteal ultrasound data set includes multiple popliteal ultrasound images processed by data cropping, data labeling, and data augmentation operations; the U-Net model includes an encoder, a decoder, and a skip connection structure connecting the aforementioned encoder and decoder; wherein, the encoder includes multiple cascaded encoder blocks; in each encoder block, an efficient cascaded multi-scale dilated convolution ECMAC and an ESAM spatial attention mechanism are sequentially executed; wherein, the efficient cascaded multi-scale dilated convolution ECMAC is composed of a multi-scale dilated convolution and an ECA channel attention mechanism; the encoder can respectively capture the image features presented by the target tissues in the aforementioned popliteal ultrasound images at different sizes; corresponding to the cascaded encoder blocks, shallow feature maps of different sizes can be respectively obtained; an encoder cross-region feature fusion structure is arranged in the encoder, and the encoder cross-region feature fusion structure can perform feature fusion between the shallow feature maps of different sizes respectively obtained through the above cascaded encoder blocks and the deep feature maps obtained in the decoder after convolution processing; the decoder includes multiple decoder blocks with an ESAM spatial attention mechanism; each decoder block can output a corresponding deep feature map; a decoder cross-region feature fusion structure is arranged in the decoder, and the decoder cross-region feature fusion structure can perform feature fusion between the deep feature maps output by each layer and the shallow feature maps in the encoder after transposed convolution processing; the skip connection structure can splice the aforementioned encoder blocks and decoder blocks correspondingly;

[0043] An image acquisition unit for acquiring popliteal fossa ultrasound images through an ultrasound probe device;

[0044] An image recognition unit for segmenting and recognizing the acquired popliteal fossa ultrasound images using a trained U-Net model to extract the image features of the target tissue in the aforementioned popliteal fossa ultrasound images.

[0045] A system for realizing ultrasound-guided assisted popliteal fossa sciatic nerve block based on U-Net, comprising:

[0046] A network node for transmitting and receiving popliteal fossa ultrasound images;

[0047] An image segmentation module for identifying the target tissue in the aforementioned popliteal fossa ultrasound images and performing segmentation processing;

[0048] A system server, the system server being connected to the network node and the image segmentation module;

[0049] The system server is configured to: construct a popliteal fossa ultrasound dataset and a U-Net model suitable for identifying the target tissue in popliteal fossa ultrasound images; the popliteal fossa ultrasound dataset includes multiple popliteal fossa ultrasound images processed by data cropping, data labeling, and data augmentation operations; the U-Net model includes an encoder, a decoder, and a skip connection structure connecting the aforementioned encoder and decoder; wherein, the encoder includes multiple cascaded encoder blocks; in each encoder block, an efficient cascaded multi-scale dilated convolution ECMAC and an ESAM spatial attention mechanism are sequentially executed; wherein, the efficient cascaded multi-scale dilated convolution ECMAC is composed of a multi-scale dilated convolution and an ECA channel attention mechanism; the encoder can respectively capture the image features presented by the target tissue in the aforementioned popliteal fossa ultrasound images at different scales; corresponding to the cascaded encoder blocks, shallow feature maps of different scales can be respectively obtained; an encoder cross-region feature fusion structure is arranged in the encoder, and the encoder cross-region feature fusion structure can perform convolution processing on the shallow feature maps of different scales respectively obtained through the cascaded encoder blocks and perform feature fusion with the deep feature maps obtained in the decoder; the decoder includes multiple decoder blocks with an ESAM spatial attention mechanism; each decoder block can correspondingly output a deep feature map; a decoder cross-region feature fusion structure is arranged in the decoder, and the decoder cross-region feature fusion structure can perform transposed convolution processing on the deep feature maps output by each layer and perform feature fusion with the shallow feature maps in the encoder; the skip connection structure can correspondingly splice the aforementioned encoder blocks and decoder blocks; acquire popliteal fossa ultrasound images through an ultrasound probe device; use a trained U-Net model to segment and recognize the acquired popliteal fossa ultrasound images to extract the image features of the target tissue in the aforementioned popliteal fossa ultrasound images.

[0050] A computer-readable storage medium stores a computer program therein, and when the computer program is executed by a processor, the implementation steps of the method described in any one of the above are realized.

[0051] Based on the above advantages and positive effects, the advantages of the present invention are as follows: A U-Net model suitable for identifying target tissues in popliteal fossa ultrasound images is designed; the encoder of the U-Net model includes multiple cascaded encoder blocks; in each encoder block, an Efficient Cascaded Multi-Scale Dilated Convolution (ECMAC) and a Spatial Attention Mechanism (ESAM) are sequentially executed; wherein, the Efficient Cascaded Multi-Scale Dilated Convolution (ECMAC) is composed of a multi-scale dilated convolution and a Channel Attention Mechanism (ECA); the encoder can respectively capture the image features presented by the target tissues in the popliteal fossa ultrasound images at different scales; corresponding to the cascaded encoder blocks, shallow feature maps of different scales can be respectively obtained; an encoder cross-region feature fusion structure is arranged in the encoder, and the encoder cross-region feature fusion structure can perform convolution processing on the shallow feature maps of different scales respectively obtained through the cascaded encoder blocks and then fuse the features with the deep feature maps obtained in the decoder; the decoder includes multiple decoder blocks with the ESAM spatial attention mechanism; each decoder block can correspondingly output a deep feature map; a decoder cross-region feature fusion structure is arranged in the decoder, and the decoder cross-region feature fusion structure can perform transposed convolution processing on the deep feature maps output by each layer and then fuse the features with the shallow feature maps in the encoder; the skip connection structure can correspondingly splice the aforementioned encoder blocks and decoder blocks.

[0052] Furthermore, by using a multi-scale dilated convolution and an ECA channel attention mechanism in the encoder, the U-Net model can capture image features of different scales, compensate for the loss of feature information during the double-layer convolution operation of the U-Net, and reduce the number of parameters and the amount of calculation.

[0053] Furthermore, the ESAM spatial attention mechanism is used in both the encoder and the decoder, which enables the U-Net model to better focus on important spatial positions, can well suppress irrelevant information during the process of segmenting and identifying popliteal fossa ultrasound images, and thus improves the overall performance of the U-Net model.

[0054] Furthermore, an encoder cross-region feature fusion structure and a decoder cross-region feature fusion structure are respectively arranged in the encoder and the decoder. Therefore, global and local context information can be better captured, which helps to alleviate the problems of gradient disappearance and gradient explosion, makes the training more stable and efficient, thereby improving the network detection accuracy and enhancing the recognition and positioning of small target tissues.

[0055] Furthermore, the improved U-Net model has a high recognition accuracy for popliteal fossa ultrasound images and can accurately identify and locate the biceps femoris, sciatic nerve (tibial nerve and common peroneal nerve), and semimembranosus muscle. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a schematic flowchart provided by an embodiment of the present invention.

[0057] Figure 2 It is a schematic structural diagram of the U-Net model provided by an embodiment of the present invention.

[0058] Figure 3 It is a schematic structural diagram of the efficient cascaded multi-scale dilated convolution ECMAC provided by an embodiment of the present invention.

[0059] Figure 4 It is a schematic structural diagram of the ESAM spatial attention mechanism provided by an embodiment of the present invention.

[0060] Figure 5 It is a schematic diagram of the encoder cross-level feature fusion structure provided by an embodiment of the present invention.

[0061] Figure 6 It is a schematic structural diagram of the decoder cross-level feature fusion structure provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] The following further elaborates in detail on a method and application for ultrasound-guided assisted popliteal fossa sciatic nerve block based on U-Net improvement disclosed in the present invention in conjunction with the accompanying drawings and specific embodiments. It should be noted that the technical features described or combined in the following embodiments should not be considered in isolation. They can be combined with each other to achieve better technical effects. In the accompanying drawings of the following embodiments, the same reference numerals in each drawing represent the same features or components and can be applied to different embodiments. Therefore, once a certain item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0063] It should be noted that the structures, ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the invention. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the invention can produce and the purposes that can be achieved, should fall within the scope covered by the technical content disclosed in the invention. The scope of the preferred embodiments of the present invention includes additional implementations, in which the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order described or discussed. This should be understood by those skilled in the technical field of the embodiments of the present invention.

[0064] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be regarded as part of the authorized specification. In all the examples shown and discussed here, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0065] Embodiment

[0066] See Figure 1 As shown, it is a flowchart provided by the present invention. The implementation steps S100 of the method are as follows:

[0067] S101, construct a popliteal fossa ultrasound dataset and a U-Net model suitable for identifying target tissues in popliteal fossa ultrasound images.

[0068] The popliteal fossa ultrasound dataset includes multiple popliteal fossa ultrasound images processed by data cropping, data labeling, and data augmentation operations (wherein the data augmentation operations include occlusion, mirroring, rotation, etc.). The popliteal fossa ultrasound image refers to an image obtained through ultrasound examination in the popliteal fossa area (i.e., the posterior knee area).

[0069] When constructing the popliteal fossa ultrasound dataset, it is preferred that the popliteal fossa ultrasound images obtained at the popliteal fossa sciatic nerve are of multiple categories. By labeling the target tissues included in the popliteal fossa ultrasound images, and finally constructing them through data augmentation. The target tissues include but are not limited to at least one of the biceps femoris, tibial nerve, common peroneal nerve, artery, and semimembranosus muscle.

[0070] The popliteal fossa ultrasound image dataset in this embodiment is derived from the practice collation of clinical anesthesia in a certain hospital. 500 patients' popliteal fossa ultrasound images are collected and sorted to construct the popliteal fossa ultrasound image dataset, which contains approximately 3000 images taken during popliteal fossa ultrasound examinations.

[0071] During the collation, the original popliteal fossa ultrasound images are first renumbered to protect the privacy information of patients. Then, 1500 popliteal fossa ultrasound images containing the biceps femoris, tibial nerve, common peroneal nerve, artery, and semimembranosus muscle are selected. To ensure the invariance of the characteristics of the tissue area and reduce the computational amount, this embodiment preferably performs cropping and occlusion processing on some of the original popliteal fossa ultrasound images. For each image containing partial target tissues, three technicians in this field segment it according to the different anatomical structures of the target tissues and use the LabelMe tool for annotation.

[0072] Among them, it is worth noting that the U-Net model uses a preset proportion of popliteal ultrasound images in the popliteal ultrasound dataset as training samples for model training; after training is completed, the remaining proportion of popliteal ultrasound images in the aforementioned popliteal ultrasound dataset is used as test samples for model testing, and the dice score is used as an evaluation index to measure the performance of the U-Net model in the ultrasound popliteal image segmentation task.

[0073] In this embodiment, the U-Net model preferably uses 80% of the popliteal ultrasound images in the popliteal ultrasound dataset as training samples for model training; after training is completed, the remaining 20% of the popliteal ultrasound images in the aforementioned popliteal ultrasound dataset are used as test samples for model testing.

[0074] Combined with Figure 2 As shown, the U-Net model in this embodiment includes an encoder, a decoder, and a skip connection structure connecting the aforementioned encoder and decoder.

[0075] Specifically, the encoder includes multiple cascaded encoder blocks; in each encoder block, an efficient cascaded multi-scale atrous convolution ECMAC and an ESAM spatial attention mechanism are sequentially executed; among them, the efficient cascaded multi-scale atrous convolution ECMAC is composed of a multi-scale atrous convolution and a channel attention mechanism ECA.

[0076] Specifically combined with Figure 3 As shown, it is a schematic diagram of the structure of the efficient cascaded multi-scale atrous convolution ECMAC in the encoder provided by the embodiment of the present invention.

[0077] In the efficient cascaded multi-scale atrous convolution (Efficient cascaded multi-scale atrous convolutions, abbreviated as ECMAC), a multi-scale atrous convolution and an ECA channel attention mechanism are specifically used, which enables the U-Net model to capture image features of different sizes and compensates for the loss of feature information during the double-layer convolution operation of the U-Net.

[0078] When performing the multi-scale atrous convolution operation, combined with Figure 3 As shown, it specifically includes step S110:

[0079] S111, for the input feature map X ∈ R C×H×W , where C, H, and W are the channels, height, and width respectively, let the input feature map X ∈ R C×H×W sequentially pass through three convolutions with the same kernel size of 3×3 and dilation rates of 1, 2, and 5 respectively, and after each convolution, a BN layer and a ReLU activation function are used for processing.

[0080] Since multi-scale dilated convolutions may produce artifacts or discontinuous features when processing boundary information. Therefore, when performing step S112, residual connections are used to solve this problem.

[0081] S112. Use residual connections to add the feature maps output after the first convolution and the second convolution through element-wise addition to supplement the details of the image features.

[0082] S113. Concatenate the feature map obtained by element-wise addition with the feature map after the third convolution in the channel dimension to form a new feature map X1 ∈ R 2C×H×W 。

[0083] The above execution of multi-scale dilated convolution operations can improve the richness and diversity of the feature maps. However, in this operation, since some of the feature maps may contribute more to the relevant information of the recognition target, while some feature maps may contribute less to the relevant information of the recognition target.

[0084] For this reason, in this embodiment, it is preferably combined with the ECA channel attention mechanism. Preferably, by calculating the intensity weights of each channel, the important features (preferably the target tissue in this embodiment) in the feature map are enhanced, and the unimportant features (preferably other tissues in the popliteal ultrasound image except the aforementioned target tissue) are suppressed, thereby improving the expression ability of the feature map, enabling the model to better capture global and local information, improving the accuracy and robustness of segmentation, and finally outputting a feature map with a size of 2C×H×W.

[0085] Specifically, the execution of the ECA channel attention mechanism includes step S120:

[0086] S121. For the concatenated feature map X1 ∈ R 2C×H×W , through global average pooling (GAP), aggregate the spatial information of each channel into a scalar, that is, obtain the average value of each channel:

[0087]

[0088] where z c represents the average value of all pixel points on the c-th channel, and X c,i,j refers to the pixel value of the feature map X1 at the c-th channel, height i, and width j.

[0089] After obtaining the aforementioned average value z c , preferably, a one-dimensional convolution operation is used to capture the cross-channel interdependencies. That is:

[0090] S122. For the average value z of each channelc ∈R C , perform a convolution operation with a one-dimensional convolution kernel of size k to obtain s = Conv1D(z, k), where k is a positive integer.

[0091] Specifically, R C refers to a C-dimensional real vector space, and each vector element therein corresponds to the average value of one channel. Among them, R represents the set of real numbers, and R C represents the set of all vectors containing C real numbers. For example, when C = 3, R C represents a three-dimensional real space, that is, all vectors of the form (x1, x2, x3), where x1, x2, x3 are real numbers.

[0092] In this embodiment, the size k of the convolution kernel adaptively selects a smaller value according to the number of channels, and preferably k = 3 here.

[0093] S123, pass the result s obtained corresponding to each channel after the convolution operation through the Sigmoid activation function to obtain the corresponding attention weight α c = σ(s); where σ(·) represents the Sigmoid activation function.

[0094] S124, apply the aforementioned attention weight α c to the aforementioned concatenated feature map X1 and re-weight each channel to obtain where is the pixel value at channel c, height i, and width j in the re-weighted feature map X1.

[0095] As another preferred embodiment of this embodiment, for the input feature map P ∈ R 2C×H×W , where C is the number of channels, H is the height, and W is the width, use the ESAM spatial attention mechanism.

[0096] Specifically, as shown in Figure 4 , the execution of the ESAM spatial attention mechanism specifically includes step S130:

[0097] S131, respectively perform max pooling and average pooling on the input feature map P ∈ R 2C×H×W , and after compression along the channel dimension, obtain the feature map P M ∈ R 1×H×W through max pooling and 1×1 convolution, and obtain the feature map as P A ∈ R 1×H×W ; at the same time, obtain the feature map P D ∈ R 1×H×W after using depthwise separable convolution on the aforementioned input feature map P.

[0098] It should be noted that the depthwise separable convolution can reduce the computational complexity and the number of parameters. This is because the depthwise separable convolution can decompose the standard convolution into a depthwise convolution and a pointwise convolution.

[0099] In the standard convolution, the size of the convolution kernel is usually K×K×C, while the depthwise convolution is a convolution operation performed separately for each input channel, and the size of the convolution kernel is K×K×1.

[0100] Combined with Figure 4 As shown, the size of the input tensor is C×H×W. In the depthwise convolution, a convolution kernel with a size of k = 3 is used to perform a 3×3 convolution operation on each channel independently. For the convenience of subsequent image processing, the stride and padding are both set to 1 to ensure that the number of channels, height, and width of the output image remain unchanged.

[0101] Then, a pointwise convolution is performed. The pointwise convolution, as a 1×1 convolution operation, is used to linearly combine the channels at each position. The output tensor size of the depthwise convolution is C×H×W. The pointwise convolution uses a 1×1×C×C′ convolution kernel, where C′ is the number of output channels, that is, C in the figure. The pointwise convolution produces an output tensor of C′×H×W by performing a 1×1 convolution at each position. This operation linearly combines the features obtained from the depthwise convolution in the channel dimension to finally obtain the required output feature map.

[0102] S132, concatenate the aforementioned feature maps P M ∈R 1×H×W 、P A ∈R 1×H×W and P D ∈R 1×H×W in the channel dimension to obtain a feature fusion map F = concat(P M ,P A ,P D ). At this time, the obtained feature fusion map is a feature map with a size of 3×H×W.

[0103] In actual operation, to capture more subtle local features in the feature map P and thus extract more useful information, while using max pooling and average pooling, a depthwise separable convolution is used to obtain the feature map P D ∈R 1×H×W . After that, the feature maps P M ∈R 1×H×W 、P A ∈R 1×H×W and P D ∈R 1×H×WConcatenate on the channel dimension, and the resulting feature fusion map F = concat(P M , P A , P D ), thereby enhancing the richness of feature expression.

[0104] S133, apply a 7×7 convolution to the aforementioned feature fusion map F to obtain the feature fusion map Conv(F) ∈ R 1×H×W ; Pass the feature fusion map Conv(F) ∈ R 1×H×W through the Sigmoid activation function to obtain the attention map M = σ(Conv(F)), M ∈ R 1 ×H×W .

[0105] To further enhance the features at the spatial positions considered important in the attention map M, steps S134 and S135 are executed.

[0106] S134, perform an element-wise multiplication (element-wise product) of the input feature map P and the aforementioned attention map M to obtain the weighted feature map P' = M ⊙ P; where, ⊙ represents element-wise multiplication.

[0107] S135, apply a depthwise separable convolution to the weighted feature map P' to obtain the feature map P * ∈R C×H×W for subsequent image processing.

[0108] Among them, performing step S135 to apply a depthwise separable convolution to the weighted feature map P' can halve the dimension while extracting more advanced feature information, thereby optimizing the feature representation and improving the overall performance of the U-Net model.

[0109] It is also worth noting that in the encoder part of this embodiment, the encoder can respectively capture the image features presented by the target tissue in the popliteal fossa ultrasound image at different sizes; the corresponding cascaded encoder blocks can respectively obtain shallow feature maps of different sizes.

[0110] Specifically, as shown in Figure 2 , after image processing through four encoder blocks, shallow feature maps F1, F2, F3, and F4 with sizes of 64×512×512, 128×256×256, 256×128×128, and 512×64×64 are obtained in sequence.

[0111] In addition, it is also worth noting that an encoder cross-region feature fusion structure is further provided in the encoder. The encoder cross-region feature fusion structure can perform feature fusion on the shallow feature maps of different sizes respectively obtained through the cascaded encoder blocks after convolution processing with the deep feature maps obtained in the decoder.

[0112] Specifically, as shown in Figure 5 , the encoder cross-region feature fusion structure can increase the number of channels while reducing the size of the aforementioned shallow feature maps through step-by-step 3×3 convolutions until the size and number of channels of the aforementioned shallow feature maps are the same as those of any deep feature map in the decoder, and then perform feature fusion on the aforementioned shallow feature map and the aforementioned deep feature map.

[0113] The step-by-step 3x3 convolution can decompose a standard 3x3 convolution operation into two smaller convolution (e.g., 1x3 and 3x1 convolution) operations for execution, so as to effectively capture spatial relationships while reducing computational complexity and memory usage.

[0114] As one of the preferred embodiments of this embodiment, when the shallow feature maps are the feature maps F1, F2, F3, and F4 of the first four layers in the encoder structure, after each shallow feature map undergoes a step-by-step 3×3 convolution operation and the corresponding size is adjusted to 1024×32×32, it is then preferably added element-wise with the deep feature map 1024×32×32 in the decoder.

[0115] Preferably, after element-wise addition, the preferably obtained feature map

[0116] where F k (i,j) represents the value of the kth feature map at position (i,j).

[0117] The process of element-wise addition in the encoder cross-region feature fusion structure effectively ensures the gradual extraction and integration of information in the shallow feature maps, thereby retaining multi-scale context information.

[0118] In the decoder of this embodiment, it includes multiple decoder blocks with an ESAM spatial attention mechanism; each decoder block can correspondingly output a deep feature map.

[0119] It is worth noting that the ESAM spatial attention mechanism is used both in the downsampling part of the encoder and in the decoder to further enhance the feature representation ability of the model and emphasize the extraction of effective features.

[0120] This is because considering that the skip connections in the original U-Net model directly concatenate the features of the encoder and the decoder, which may contain a large amount of redundant or irrelevant information. To suppress this irrelevant information, the ESAM spatial attention mechanism is added, which can help the model focus on important spatial positions and enhance the expression of useful features.

[0121] In addition, a decoder cross-region feature fusion structure is also provided in the decoder. The decoder cross-region feature fusion structure can perform stepwise upsampling on the size and number of channels of the aforementioned deep feature map, and when the size and number of channels of the deep feature map are consistent with those of any shallow feature map, fuse the deep feature map with the shallow feature map.

[0122] Combined Figure 6 As shown, the decoder cross-region feature fusion structure fuses the deep feature maps with sizes of 1024×32×32, 512×64×64, 256×128×128, and 128×256×256 with the shallow feature map 64×512×512 by stepwise upsampling.

[0123] In actual operation, transposed convolution is used to change the size of the deep feature map from smaller to larger, and the number of channels gradually decreases until it becomes 64×512×512. Then, the upsampled feature map is further processed through a 3×3 convolutional layer and a ReLU activation function to extract more useful information, which is very beneficial for identifying small targets such as nerves.

[0124] The specific formula is as follows: x out = ReLU(Conv2d(Deconv(x))). Here, x is an input tensor with the shape of C×H×W, Deconv is the transposed convolution, its kernel size is 4×4, the stride is 2, and the padding is 1 to ensure that the image size remains unchanged after this convolution. Conv2d uses a 3×3 convolution and is followed by a ReLU activation function to obtain the output map x out .

[0125] It should be noted that in this embodiment, the decoder cross-region feature fusion structure uses the stepwise upsampling method and needs to repeat the above operation according to the tensor size of the input image. As an example, the deepest feature map needs to go through the above steps 4 times to obtain an image with a size of 64×512×512, while the second-layer feature map only needs 1 time.

[0126] It is also worth noting that in this embodiment, cross-level feature fusion is introduced in both the encoder and the decoder, so that global and local context information can be better captured, which helps to alleviate the problems of gradient disappearance and gradient explosion, making the training more stable and efficient, thereby improving the network detection accuracy and enhancing the recognition and localization of small targets.

[0127] In addition, the skip connection structure in this embodiment can correspondingly splice the aforementioned encoder block and decoder block.

[0128] So far, the configuration of the U-net model is completed.

[0129] In this embodiment, in order to verify the performance of the improved U-Net model on the custom popliteal ultrasound dataset and analyze the experimental results. Therefore, in actual operation, after the U-net model training is completed, steps S102 and S103 are preferably executed.

[0130] S102, collect popliteal ultrasound images through an ultrasound probe device.

[0131] S103, use the trained U-Net model to segment and recognize the collected popliteal ultrasound images to extract the image features of the target tissue in the aforementioned popliteal ultrasound images.

[0132] The experimental results show that for popliteal ultrasound images, the improved U-Net model can accurately identify the nerves (sciatic nerve, tibial nerve, and common peroneal nerve) and muscles (biceps femoris and semimembranosus) and arteries at the popliteal fossa. Among them, the mIOU reaches above 0.89, and the Dice coefficient reaches above 0.94, which indicates that the U-Net model has high reliability and very accurate recognition and localization.

[0133] In addition, it is also worth emphasizing that for the rotated popliteal ultrasound images and popliteal ultrasound images under special environments mentioned in this embodiment, the algorithm can also accurately segment the target positions.

[0134] Other technical features refer to the previous embodiments and will not be elaborated here.

[0135] In addition, the present invention also gives an embodiment, providing a device for realizing ultrasound-guided assisted popliteal sciatic nerve block based on U-Net, including:

[0136] A data and model construction unit for constructing a popliteal ultrasound dataset and a U-Net model suitable for identifying target tissues in popliteal ultrasound images; the popliteal ultrasound dataset includes multiple popliteal ultrasound images processed by data cropping, data labeling, and data augmentation operations; the U-Net model includes an encoder, a decoder, and a skip connection structure connecting the aforementioned encoder and decoder; wherein, the encoder includes multiple cascaded encoder blocks; in each encoder block, an Efficient Cascaded Multi-Scale Dilated Convolution (ECMAC) and a Spatial Attention Mechanism (ESAM) are sequentially executed; wherein, the Efficient Cascaded Multi-Scale Dilated Convolution (ECMAC) is composed of a multi-scale dilated convolution and a Channel Attention Mechanism (ECA); the encoder composed of the encoder blocks can respectively capture the image features presented by the target tissues in the aforementioned popliteal ultrasound images at different scales; corresponding to the cascaded encoder blocks, shallow feature maps of different scales can be respectively obtained; an encoder cross-region feature fusion structure is arranged in the encoder, and the encoder cross-region feature fusion structure can perform convolution processing on the shallow feature maps of different scales respectively obtained through the cascaded encoder blocks, and then perform feature fusion with the deep feature maps obtained in the decoder; the decoder includes multiple decoder blocks with an ESAM spatial attention mechanism; each decoder block can correspondingly output a deep feature map; a decoder cross-region feature fusion structure is arranged in the decoder, and the decoder cross-region feature fusion structure can perform transposed convolution processing on the deep feature maps output by each layer and then perform feature fusion with the shallow feature maps in the encoder; the skip connection structure can correspondingly splice the aforementioned encoder blocks and decoder blocks.

[0137] An image acquisition unit for acquiring popliteal ultrasound images through an ultrasound probe device.

[0138] An image recognition unit for using the trained U-Net model to perform segmentation and recognition on the acquired popliteal ultrasound images to extract the image features of the target tissues in the aforementioned popliteal ultrasound images.

[0139] For other technical features, refer to the previous embodiments and will not be elaborated here.

[0140] In addition, the present invention also gives an embodiment, providing a system for realizing ultrasound-guided assisted popliteal sciatic nerve block based on U-Net, including:

[0141] A network node for receiving and transmitting popliteal ultrasound images.

[0142] An image segmentation module for identifying target tissues in the aforementioned popliteal ultrasound images and performing segmentation processing.

[0143] A system server, and the system server is connected to the network node and the image segmentation module.

[0144] The described system server is configured to: construct a popliteal fossa ultrasound dataset and a U-Net model applicable to identifying target tissues in popliteal fossa ultrasound images; the popliteal fossa ultrasound dataset includes multiple popliteal fossa ultrasound images processed by data cropping, data labeling, and data augmentation operations; the U-Net model includes an encoder, a decoder, and a skip connection structure connecting the aforementioned encoder and decoder; wherein, the encoder includes multiple cascaded encoder blocks; in each encoder block, an Efficient Cascaded Multi-Scale Atrous Convolution (ECMAC) and a Spatial Attention Mechanism (ESAM) are sequentially executed; wherein, the Efficient Cascaded Multi-Scale Atrous Convolution (ECMAC) is composed of a multi-scale atrous convolution and a Channel Attention Mechanism (ECA); the encoder can respectively capture the image features presented by the target tissues in the aforementioned popliteal fossa ultrasound images at different scales; corresponding to the cascaded encoder blocks, shallow feature maps of different scales can be respectively obtained; an encoder cross-region feature fusion structure is arranged in the encoder, and the encoder cross-region feature fusion structure can perform convolution processing on the shallow feature maps of different scales respectively obtained through the cascaded encoder blocks, and then perform feature fusion with the deep feature maps obtained in the decoder; the decoder includes multiple decoder blocks with the ESAM spatial attention mechanism; each decoder block can correspondingly output a deep feature map; a decoder cross-region feature fusion structure is arranged in the decoder, and the decoder cross-region feature fusion structure can perform transposed convolution processing on the deep feature maps output by each layer and then perform feature fusion with the shallow feature maps in the encoder; the skip connection structure can correspondingly splice the aforementioned encoder blocks and decoder blocks; collect popliteal fossa ultrasound images through an ultrasound probe device; use the trained U-Net model to perform segmentation and recognition on the collected popliteal fossa ultrasound images to extract the image features of the target tissues in the aforementioned popliteal fossa ultrasound images.

[0145] For other technical features, refer to the previous embodiments and will not be elaborated here.

[0146] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a program is stored for the aforementioned system for improving ultrasound-guided assisted popliteal sciatic nerve block based on U-Net. When the program is executed by a processor, it can implement the steps of the method for improving ultrasound-guided assisted popliteal sciatic nerve block based on U-Net described in any one of the above.

[0147] For other technical features, refer to the previous embodiments and will not be elaborated here.

[0148] In the foregoing description, within the scope of the object of the present disclosure, the various components may be selectively and operably combined in any number. Additionally, terms such as "including", "comprising", and "having" shall be construed as inclusive or open-ended by default, rather than exclusive or closed, unless expressly limited to the contrary. All technical, scientific, or other terms shall have the meaning understood by those skilled in the art, unless limited to the contrary. Common terms found in dictionaries shall not be construed too idealistically or too unrealistically in the context of the relevant technical documents, unless the present disclosure expressly so limits them.

[0149] Although the exemplary aspects of the present disclosure have been described for purposes of illustration, those skilled in the art will recognize that the foregoing description is only of the preferred embodiments of the invention and is not any limitation on the scope of the invention. The scope of the preferred embodiments of the invention includes additional implementations where functions may be performed out of the order noted or discussed. Any changes or modifications made by those of ordinary skill in the art based on the foregoing disclosure are within the scope of the claims.

Claims

1. A method for realizing ultrasound-guided assisted popliteal sciatic nerve block based on U-Net, characterized in that, Including: Constructing a popliteal fossa ultrasound dataset and a U-Net model suitable for identifying target tissues in popliteal fossa ultrasound images; the popliteal fossa ultrasound dataset includes multiple popliteal fossa ultrasound images processed by data cropping, data labeling, and data augmentation operations; the U-Net model includes an encoder, a decoder, and a skip connection structure connecting the aforementioned encoder and decoder; wherein, The encoder includes multiple cascaded encoder blocks; in each encoder block, an efficient cascaded multi-scale dilated convolution ECMAC and a spatial attention mechanism ESAM are sequentially executed; wherein, the efficient cascaded multi-scale dilated convolution ECMAC is composed of a multi-scale dilated convolution and a channel attention mechanism ECA; the encoder can respectively capture the image features presented by the target tissues in the aforementioned popliteal fossa ultrasound images at different scales; corresponding to the cascaded encoder blocks, shallow feature maps of different scales can be respectively obtained; An encoder cross-region feature fusion structure is arranged in the encoder, and the encoder cross-region feature fusion structure can perform feature fusion on the shallow feature maps of different scales respectively obtained through the cascaded encoder blocks after convolution processing with the deep feature maps obtained in the decoder; The decoder includes multiple decoder blocks with an ESAM spatial attention mechanism; each decoder block can correspondingly output a deep feature map; A decoder cross-region feature fusion structure is arranged in the decoder, and the decoder cross-region feature fusion structure can perform feature fusion on the deep feature maps output by each layer after transposed convolution processing with the shallow feature maps in the encoder; The skip connection structure can correspondingly splice the aforementioned encoder block and decoder block; Collecting popliteal fossa ultrasound images through an ultrasound probe device; Using the trained U-Net model to perform segmentation and recognition on the collected popliteal fossa ultrasound images to extract the image features of the target tissues in the aforementioned popliteal fossa ultrasound images.

2. The method according to claim 1, wherein The target tissues include at least one of biceps femoris, tibial nerve, common peroneal nerve, artery, and semimembranosus; The U-Net model uses a preset proportion of popliteal fossa ultrasound images in the popliteal fossa ultrasound dataset as training samples and inputs them into the model for training; After the training is completed, the remaining proportion of popliteal fossa ultrasound images in the aforementioned popliteal fossa ultrasound dataset is used as test samples and input into the model for testing, and the dice score is used as an evaluation index to measure the performance of the U-Net model in the ultrasound popliteal fossa image segmentation task.

3. The method according to claim 1, characterized in that, When the multi-scale dilated convolution is executed, it specifically includes the steps of: For the input feature map X ∈ R C×H×W , where C, H, and W are the number of channels, height, and width respectively, let the input feature map X ∈ R C×H×W successively pass through three convolutions with the same kernel size of 3×3 and dilation rates of 1, 2, and 5 respectively, and after each convolution, use the BN layer and ReLU activation function for processing; Using a residual connection to add the feature maps output after the first convolution and the second convolution element by element through element-wise addition to supplement the details of the image features; The feature map obtained by adding elements level by level is then concatenated with the feature map after the third convolution in the channel dimension to form a new feature map X1 ∈ R 2C×H×W .

4. The method according to claim 3, wherein The execution of the ECA channel attention mechanism includes the steps of: For the stitched feature map X1 ∈ R 2C×H×W , through global average pooling, the average value of each channel is obtained: Among them, z c represents the average value of all pixel points on the c-th channel, and X c,i,j refers to the pixel value of the feature map X at the c-th channel, height i, and width j; The average value z for each channel c ∈R C , a convolution operation with a one-dimensional convolution kernel of size k is used to obtain s = Conv1D(z, k), where k is a positive integer; For each result s obtained after the convolution operation on each channel, pass it through the Sigmoid activation function to obtain the corresponding attention weight α c = σ(s); where σ(·) represents the Sigmoid activation function; apply the aforementioned attention weight α c to the aforementioned concatenated feature map X1 and re-weight each channel to obtain where is the pixel value at channel c, height i, and width j in the re-weighted feature map X1.

5. The method according to claim 1, wherein The execution of the ESAM spatial attention mechanism specifically includes the steps of: For the input feature map P ∈ R 2C×H×W After max pooling and average pooling respectively, and compression along the channel dimension, the feature map P is obtained through max pooling and 1×1 convolution M ∈ R 1×H×W , and the feature map obtained through average pooling and 1×1 convolution is P A ∈ R 1 ×H×W ; Meanwhile, after using depthwise separable convolution on the aforementioned input feature map P, the feature map P D ∈ R 1×H×W ; Concatenate the foregoing feature maps P M ∈R 1×H×W , P A ∈R 1×H×W and P D ∈R 1×H×W along the channel dimension to obtain a feature fusion map F = concat(P M , P A , P D ); Apply a 7×7 convolution to the aforementioned feature fusion graph F to obtain the feature fusion graph Conv(F) ∈ R 1×H×W ; The feature fusion graph Conv(F) ∈ R 1×H×W Through the Sigmoid activation function, the attention graph M = σ(Conv(F)) is obtained, where M ∈ R 1×H×W ; Multiplying the input feature map P element by element with the aforementioned attention map M to obtain a weighted feature map P′ = M ⊙ P; wherein, ⊙ represents element-wise multiplication; The weighted feature map P′ is passed through a depthwise separable convolution to obtain the feature map P * ∈R C×H×W .

6. The method according to claim 1, wherein The encoder cross-region feature fusion structure can reduce the size of the aforementioned shallow feature maps through step-by-step 3×3 convolutions while increasing the number of channels. When the size and number of channels of the aforementioned shallow feature maps are the same as those of any deep feature map in the decoder, the aforementioned shallow feature maps and the aforementioned deep feature maps are then subjected to feature fusion.

7. The method according to claim 1, wherein The decoder cross-region feature fusion structure can, through step-by-step upsampling operations, make the size and number of channels of the aforementioned deep feature maps consistent with those of any shallow feature map, and then perform feature fusion between the aforementioned deep feature maps and the aforementioned shallow feature maps.

8. An apparatus for ultrasound-guided assisted popliteal sciatic nerve block based on U-Net according to the method described in any one of claims 1-7, characterized in that Comprising: A data and model construction unit for constructing a popliteal ultrasound dataset and a U-Net model suitable for identifying target tissues in popliteal ultrasound images; the popliteal ultrasound dataset includes a plurality of popliteal ultrasound images processed through data cropping, data labeling, and data augmentation operations; the U-Net model includes an encoder, a decoder, and a skip connection structure connecting the aforementioned encoder and decoder; wherein, the encoder includes a plurality of cascaded encoder blocks; in each encoder block, an efficient cascaded multi-scale dilated convolution ECMAC and a spatial attention mechanism ESAM are sequentially executed; wherein, the efficient cascaded multi-scale dilated convolution ECMAC consists of a multi-scale dilated convolution and a channel attention mechanism ECA; the encoder can respectively capture the image features presented by the target tissues in the aforementioned popliteal ultrasound images at different sizes; corresponding to the cascaded encoder blocks, shallow feature maps of different sizes can be respectively obtained; an encoder cross-region feature fusion structure is provided in the encoder, and the encoder cross-region feature fusion structure can perform convolution processing on the shallow feature maps of different sizes respectively obtained through the above-mentioned cascaded encoder blocks and then perform feature fusion with the deep feature maps obtained in the decoder; the decoder includes a plurality of decoder blocks with an ESAM spatial attention mechanism; each decoder block can correspondingly output a deep feature map; a decoder cross-region feature fusion structure is provided in the decoder, and the decoder cross-region feature fusion structure can perform transposed convolution processing on the deep feature maps output by each layer and then perform feature fusion with the shallow feature maps in the encoder; the skip connection structure can correspondingly splice the aforementioned encoder blocks and decoder blocks; An image acquisition unit for acquiring popliteal ultrasound images through an ultrasound probe device; An image recognition unit for using the trained U-Net model to perform segmentation recognition on the acquired popliteal ultrasound images to extract the image features of the target tissues in the aforementioned popliteal ultrasound images.

9. A system for ultrasound-guided assisted popliteal sciatic nerve block based on U-Net according to the method of any one of claims 1-7, characterized in that Comprising: A network node for receiving and transmitting popliteal ultrasound images; An image segmentation module for identifying the target tissues in the aforementioned popliteal ultrasound images and performing segmentation processing; A system server, and the system server is connected to the network node and the image segmentation module; The described system server is configured to: construct a popliteal fossa ultrasound dataset and a U-Net model applicable to identifying target tissues in popliteal fossa ultrasound images; the popliteal fossa ultrasound dataset includes multiple popliteal fossa ultrasound images processed by data cropping, data labeling, and data augmentation operations; the U-Net model includes an encoder, a decoder, and a skip connection structure connecting the aforementioned encoder and decoder; among them, the encoder includes multiple cascaded encoder blocks; in each encoder block, an efficient cascaded multi-scale dilated convolution ECMAC and a spatial attention mechanism ESAM are sequentially executed; among them, the efficient cascaded multi-scale dilated convolution ECMAC is composed of a multi-scale dilated convolution and a channel attention mechanism ECA; the encoder can respectively capture the image features presented by the target tissues in the aforementioned popliteal fossa ultrasound images at different scales; corresponding to the cascaded encoder blocks, shallow feature maps of different scales can be respectively obtained; an encoder cross-region feature fusion structure is arranged in the encoder, and the encoder cross-region feature fusion structure can perform convolution processing on the shallow feature maps of different scales respectively obtained through the cascaded encoder blocks and fuse the features with the deep feature maps obtained in the decoder; the decoder includes multiple decoder blocks with an ESAM spatial attention mechanism; each decoder block can correspondingly output a deep feature map; a decoder cross-region feature fusion structure is arranged in the decoder, and the decoder cross-region feature fusion structure can perform transposed convolution processing on the deep feature maps output by each layer and fuse the features with the shallow feature maps in the encoder; the skip connection structure can correspondingly splice the aforementioned encoder blocks and decoder blocks; collect popliteal fossa ultrasound images through an ultrasound probe device; use the trained U-Net model to perform segmentation and recognition on the collected popliteal fossa ultrasound images to extract the image features of the target tissues in the aforementioned popliteal fossa ultrasound images.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Multi-feature cyclic convolution saliency target detection method based on attention mechanism

    CN110648334A

  • Coronary artery image segmentation method and system based on deep learning

    CN117495876A