Lumbar plexus ultrasound segmentation and recognition method based on U-Net and application thereof

By constructing a U-Net model and combining it with DCDR and AEH modules, precise segmentation and identification of the lumbar plexus nerves were achieved, solving the problem of difficult lumbar plexus nerve localization and improving the accuracy and efficiency of anesthesia procedures.

CN120612576BActive Publication Date: 2025-11-28THE NAVAL MEDICAL UNIV OF PLA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510606146.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-11-28
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

In existing technologies, accurate identification and localization of the lumbar plexus nerves are difficult in ultrasound-guided lumbar plexus nerve block procedures, especially due to the influence of patient position and the surgeon's skill, making it difficult for anesthesiologists to quickly and accurately locate the block position.

Method used

A U-Net-based ultrasound segmentation method for lumbar plexus nerves is constructed. By building a U-Net model, a multi-layer dynamic context-aware extended residual module DCDR and a multi-layer attention-enhanced hybrid module AEH are used for ultrasound image preprocessing and recognition, achieving accurate segmentation and recognition of the psoas major, quadratus lumborum, erector spinae, transverse processes, and lumbar plexus nerves.

Benefits of technology

It improves the accuracy of lumbar plexus nerve identification, reduces patient waiting time and pain, assists anesthesiologists in quickly and accurately locating the block area, and reduces the risk of operational errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612576B_ABST
    Figure CN120612576B_ABST
Patent Text Reader

Abstract

The application provides a lumbar plexus nerve ultrasound segmentation and recognition method based on U-Net and application, and relates to the technical field of medical auxiliary diagnosis. The method comprises the following steps: constructing a lumbar ultrasound image dataset; constructing a U-Net model capable of recognizing the psoas major muscle, the quadratus lumborum muscle, the erector spinae muscle, the transverse process and the lumbar plexus nerve, and training the U-Net model; the U-Net model comprises an encoder, a jump connection and a decoder; the encoder uses a multi-layer dynamic context perception expansion residual module (DCDR) to perform a downsampling operation; the decoder uses a multi-layer attention enhancement hybrid module (AEH) to perform an upsampling operation; the trained U-Net model is configured in an ultrasound device; an image recognition is performed on the collected ultrasound image, and a recognition result is obtained. The application can help clinicians to quickly and accurately interpret the cloverleaf view ultrasound image containing the lumbar plexus nerve tissue, so as to assist them in accurately positioning the block area and injecting local anesthetic.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical auxiliary diagnosis, and particularly relates to a lumbar plexus nerve ultrasound segmentation and recognition method based on U-Net. BACKGROUND

[0002] Lumbar plexus block is a technique commonly used for anesthesia, often used for analgesia in abdominal, hip, lower limb and other surgeries, which achieves anesthetic effect by injecting anesthetic near the nerve plexus of the waist. At present, the nerve block is often performed under the guidance of ultrasound by searching for the lumbar plexus nerve at the posterior one-third of the psoas major muscle.

[0003] The lumbar plexus block under the guidance of ultrasound has the following advantages: (1) safety: ultrasound guidance can provide real-time visual feedback, making anesthesia operation safer and reducing operation errors and complications; (2) accuracy: ultrasound can display the position, shape and surrounding tissue structure of the nerve in real time, helping doctors to accurately find the nerve and implement block, and avoiding damage to other tissues; (3) reducing pain: ultrasound can clearly show the position of the target nerve and the relationship between the surrounding tissues, so that doctors can accurately control the direction and depth of the needle, and can make patients obtain better pain control during the operation, thereby reducing intraoperative and postoperative discomfort.

[0004] Although the lumbar plexus block under the guidance of ultrasound has been widely used in clinical practice, the accurate recognition and positioning of the lumbar plexus nerve is still a difficulty faced by anesthesiologists. Due to different patient positions and different probe positions, angles and forces of the doctor during ultrasound operation, it is very difficult to find the lumbar plexus nerve.

[0005] In actual operation, the "Shamrock" method can be used to find the lumbar plexus nerve. Anesthesiologists can use the "Shamrock" method to locate the Shamrock view containing the lumbar plexus nerve. However, for anesthesiologists, the proficiency of ultrasound technique has a great influence on finding the lumbar plexus nerve in the Shamrock view at the posterior one-third of the psoas major muscle.

[0006] Therefore, it is of great clinical significance to automatically recognize the psoas major muscle, quadratus lumborum muscle, erector spinae muscle, transverse process and lumbar plexus nerve at the posterior one-third of the psoas major muscle in the Shamrock view on the ultrasound device, so as to ensure that anesthesiologists can accurately and quickly find the position of the lumbar plexus block.

[0007] Therefore, the present application provides a transversus abdominis muscle plane tissue structure recognition method based on U-Net and application to help anesthesiologists quickly and accurately obtain the position of the lumbar plexus block, which is a technical problem to be solved at present. SUMMARY

[0008] The purpose of the present application is to overcome the shortcomings of the prior art, provide a U-Net lumbar plexus ultrasound segmentation recognition method and application, the present application constructs a U-Net model suitable for identifying the lumbar plexus, to help clinicians quickly and accurately interpret the clover view ultrasound image containing the lumbar plexus tissue, to assist them in accurately positioning the block area and injecting local anesthetic.

[0009] To solve the existing technical problems, the present application provides the following technical solutions:

[0010] A lumbar plexus ultrasound segmentation method based on U-Net, specifically comprising:

[0011] The ultrasound image is preprocessed for the preset clover view, and a lumbar ultrasound image dataset is constructed; the clover view is an ultrasound image; the clover view includes psoas major muscle, quadratus lumborum muscle, erector spinae muscle, transverse process and lumbar plexus;

[0012] A U-Net model capable of identifying psoas major muscle, quadratus lumborum muscle, erector spinae muscle, transverse process and lumbar plexus is constructed and trained; the U-Net model includes an encoder, a skip connection and a decoder; wherein the encoder uses a multi-layer dynamic context perception expansion residual module DCDR to perform downsampling operation; the decoder uses a multi-layer attention enhancement hybrid module AEH to perform upsampling operation;

[0013] The trained U-Net model is configured in an ultrasound device, and the ultrasound device can collect the clover view;

[0014] The collected clover view is identified using the aforementioned U-Net model configured in the ultrasound device, and the identification result is obtained.

[0015] Further, the ultrasound image preprocessing includes data labeling and data enhancement; wherein the data labeling can use Labelme labeling software to draw a frame for the psoas major muscle, quadratus lumborum muscle, erector spinae muscle, transverse process and lumbar plexus in the aforementioned clover structure to complete the labeling;

[0016] The data enhancement includes processing the aforementioned clover view in the manner of occlusion, cropping, random rotation, random flipping and / or random mirroring;

[0017] 80% of the clover views in the lumbar ultrasound image dataset constitute a training set for training the aforementioned U-Net model; the remaining 20% of the clover views constitute a test set for verifying the aforementioned U-Net model.

[0018] Further, in the encoder, the DCDR module is configured with an adaptive multi-scale expansion branch AMD and a deep feature residual branch DFR which are executed in parallel;

[0019] In parallel execution, the aforementioned AMD branch is used to extract the context information in the clover view; the DFR branch extracts the local detail information in the clover view;

[0020] After fusing the output features of the AMD branch and the DFR branch, the expression of the image features is enhanced through the ECA attention mechanism.

[0021] Further, before executing the AMD branch, a multi-stage dilation rate predictor and a context-aware dilation rate predictor are used to respectively predict N different scale dilation rates, N being an integer greater than or equal to 2; wherein,

[0022] When using the multi-stage dilation rate predictor, sequentially perform the initial step S110 and the refinement step S120; wherein,

[0023] When performing the initial step S110, sequentially use a 1x1 convolution, a ReLU function and a 1x1 convolution to extract local information from the input feature map and generate an initial dilation rate prediction value;

[0024] When performing the S120 refinement step, it includes: S121, using the aforementioned initial dilation rate prediction value to perform dilated convolution on the input feature map, and then splicing with the initial feature map; S122, sequentially using a 1x1 convolution, a ReLU function and a 1x1 convolution to capture the context information in the feature map to correct the error of the aforementioned initial dilation rate prediction value, and output a corrected dilation rate prediction value;

[0025] When using the context-aware dilation rate predictor, it further includes step S130: S131, extracting the global information of the input feature map through adaptive average pooling AAP to compress the spatial dimension of the input feature map; S132, performing at least one 1x1 convolution on the compressed input feature map to obtain three different scale dilation rate prediction values;

[0026] The dilation rate prediction values of the multi-stage dilation rate predictor and the context-aware dilation rate predictor are respectively normalized using Softmax;

[0027] After normalization, through element-by-element addition and average fusion, three channel dilation rate weight values d1, d2 and d3 are obtained, the aforementioned dilation rate weight values d1, d2 and d3 are respectively amplified, and corresponding dilation rate prediction values D1, D2 and D3 are obtained.

[0028] Further, when executing the AMD branch, it includes step S140:

[0029] S141, for the feature Figure XThe aforementioned expansion rate prediction values D1, D2 and D3 are respectively applied to 3x3 expansion convolution to adjust the receptive field of the convolution operation, and different scale features are obtained after the aforementioned convolution operation Figure X 1. X2 and X3;

[0030] S142, the aforementioned features Figure X 1. X2 and X3 are respectively weighted according to corresponding weights d1, d2 and d3;

[0031] S143, the weighted features Figure X 1. X2 and X3, and the features Figure X Element-wise addition is performed;

[0032] S144, after element-wise addition, ECA attention mechanism is used on the element-wise added feature map, and new features Figure X ’.

[0033] Further, when performing the DFR branch, the step S150 is included:

[0034] S151, adjusting the channel number of the input feature map to obtain a feature map Out1;

[0035] S152, for the aforementioned feature map Out1, after respectively passing through a plurality of convolution operations of different sizes, splicing is performed, and convolution operation is performed on the spliced feature map, and the channel number of the feature map is adjusted to obtain a feature map Out2;

[0036] S153, for the aforementioned feature map Out2, after respectively passing through a plurality of convolution operations of different sizes, splicing is performed, and convolution operation is performed on the spliced feature map, and the channel number of the feature map is adjusted to obtain a feature map Out3;

[0037] S154, residual connection is performed on the aforementioned feature maps Out1, Out2 and Out3 with the same size; wherein learnable weight coefficients W1, W2 and W3 are respectively set for the aforementioned feature maps Out1, Out2 and Out3; the learnable weight coefficients W1, W2 and W3 can be respectively used to control the proportion weight of the aforementioned feature maps Out1, Out2 and Out3, and the values of the aforementioned learnable weight coefficients W1, W2 and W3 are optimized and adjusted in the U-Net model training process.

[0038] Further, the AEH module is configured with an enhanced residual multi-head attention mechanism module ER-MHA and a spatial attention mechanism module SA; wherein,

[0039] The number of layers of the AEH module is set according to the number of layers of the encoder; each layer of the AEH module corresponds to using a transpose convolution operation;

[0040] After using the transpose convolution operation on each layer of the decoder, a ER-MHA module is further used to transmit feature maps of different sizes layer by layer; and a SA module is executed for the output feature map after the up-sampling operation of at least one layer in the decoder, and the feature maps of different scales are spliced at the last layer of the decoder to obtain the final output feature map.

[0041] Further, the ER-MHA is configured to: after generating an output value using the multi-head attention mechanism module, a feedforward network FFN is added; at the same time, the input feature map is combined with the output of the aforementioned feedforward network through a residual connection; then, the output feature map is obtained by using the layer normalization operation and the preset learnable scaling factor.

[0042] A lumbar plexus ultrasound segmentation system based on a U-Net includes:

[0043] An ultrasound device is used to acquire a clover view;

[0044] A model configuration module is used to configure the trained U-Net model on the aforementioned ultrasound device;

[0045] A system server is connected to the ultrasound device and the model configuration module;

[0046] The system server is configured to: perform ultrasound image preprocessing on a preset clover view to construct a lumbar ultrasound image dataset; the clover view is an ultrasound image; the clover view includes the psoas major muscle, the quadratus lumborum muscle, the erector spinae muscle, the transverse process, and the lumbar plexus; a U-Net model capable of identifying the psoas major muscle, the quadratus lumborum muscle, the erector spinae muscle, the transverse process, and the lumbar plexus is constructed and trained; the U-Net model includes an encoder, a jump connection, and a decoder; wherein the encoder uses a multi-layer dynamic context-aware dilated residual module DCDR to perform a down-sampling operation; the decoder uses a multi-layer attention-enhanced hybrid module AEH to perform an up-sampling operation; the trained U-Net model is configured in the ultrasound device, which can acquire the clover view; the acquired clover view is subjected to image recognition using the aforementioned U-Net model configured in the ultrasound device, and an identification result is obtained.

[0047] A computer-readable storage medium has a computer program stored therein, and the computer program, when executed by a processor, implements the implementation steps of any of the above methods.

[0048] Based on the above advantages and positive effects, the advantages of the present application are:

[0049] A plurality of dynamic context-aware dilated residual modules DCDR are used in the encoder of the U-Net to capture the feature details of the ultrasound image to obtain multi-scale information.

[0050] Further, an entire attention-enhanced hybrid module AEH is used in the decoder to solve the problems of insufficient feature fusion, missing context information and loss of high-frequency details, which can effectively improve the accuracy of the U-Net model in segmenting the lumbar plexus. The AEH module consists of two parts. In the first part, an ER-MHA module is added after each upsampling operation. After adding the ER-MHA module, not only is the excessive attention to irrelevant information by the traditional splicing method avoided, but also the model can weight the features in a larger range, thereby improving the understanding of complex structures and helping the model better understand the global context. In the second part, a spatial attention mechanism module is preferably used for each layer of the decoder, and the feature maps of different scales are spliced in the last layer of the decoder to fully exploit the detail information of each layer, enhance the global perception ability of the model, and at the same time avoid potential interference during interlayer information transmission.

[0051] Further, the ER-MHA module is improved on the basis of the multi-head attention mechanism module, which can ensure that the model converges more quickly and stably, enhance the flexibility and task adaptability of the model, and thus enhance the capture of fuzzy areas.

[0052] Further, the U-Net model is applied to the ultrasound image dataset of the clover view to realize accurate segmentation and recognition of the clover view, which is helpful for the accurate positioning of the lumbar plexus by anesthetists during the anesthesia operation when applied to the ultrasound device, thereby greatly reducing the waiting time of patients and the pain of patients. BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1 The method flowchart provided for the embodiments of the application.

[0054] Figure 2 The structure diagram of the U-Net model provided for the embodiments of the application.

[0055] Figure 3 The structure diagram of the DCDR module in the encoder provided for the embodiments of the application.

[0056] Figure 4 The structure diagram of the multi-stage dilation rate predictor in the AMD branch provided for the embodiments of the application.

[0057] Figure 5 The structure diagram of the context-aware dilation rate predictor in the AMD branch provided for the embodiments of the application.

[0058] Figure 6 A structure diagram of the AMD branch provided for an embodiment of the present application is shown in FIG. 3.

[0059] Figure 7 A structure diagram of the AMD branch provided for an embodiment of the present application is shown in FIG. 3.

[0060] Figure 8 A structure diagram of the DFR branch provided for an embodiment of the present application is shown in FIG. 4.

[0061] Figure 9 A structure diagram of the AEH module provided for an embodiment of the present application is shown in FIG. 5.

[0062] Figure 10 A structure diagram of the ER-MHA module provided for an embodiment of the present application is shown in FIG. 6.

[0063] Figure 11 A structure diagram of the system provided for an embodiment of the present application is shown in FIG. 7.

[0064] Legend of reference signs:

[0065] System 200, ultrasound device 201, model configuration module 202, system server 203. DETAILED DESCRIPTION

[0066] The present application discloses a lumbar plexus nerve ultrasound segmentation and recognition method based on U-Net and application. It should be noted that the technical features or combinations of technical features described in the following embodiments should not be considered in isolation, and they can be combined with each other to achieve better technical effects. In the drawings of the following embodiments, the same reference signs in different drawings represent the same features or components, which can be applied to different embodiments. Therefore, once a feature is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0067] It should be noted that the structures, proportions, sizes, etc. shown in the drawings attached to the present specification are only used to cooperate with the content disclosed in the specification, so that those skilled in the art can understand and read, and are not used to limit the implementation conditions of the present application. Any modification of structure, change of proportion relationship or adjustment of size, which does not affect the effect and purpose of the present application, should be included in the scope of the disclosed technology. The scope of the preferred embodiments of the present application includes additional implementations, in which the functions can be performed in the order discussed or in reverse order, or in a substantially simultaneous manner according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0068] Techniques, methods, and apparatus known to those of ordinary skill in the relevant art can not be discussed in detail in this document, but the techniques, methods, and apparatus should be considered part of the disclosure in suitable circumstances. In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary, and not as a limitation. Thus, other examples of the exemplary embodiments can have different values.

[0069] Embodiments

[0070] Reference Figure 1 As shown, a method flowchart is provided for the present application. The implementation steps S100 of the method are as follows:

[0071] S101, pre-process the ultrasound image for the preset clover view, and construct a lumbar ultrasound image dataset.

[0072] The clover view is an ultrasound image. The clover view includes the psoas major muscle, the quadratus lumborum muscle, the erector spinae muscle, the transverse process, and the lumbar plexus nerve at the posterior one-third of the psoas major muscle. Among them, the psoas major muscle is in front, the erector spinae muscle is in back, and the quadratus lumborum muscle is located at the top of the transverse process.

[0073] In this embodiment, the clover view in the ultrasound image dataset is derived from a certain authoritative comprehensive three-level first-class hospital. A senior clinician of the hospital collected and arranged about 1000 clover views taken by ultrasound examination of about 150 patients.

[0074] The ultrasound image preprocessing includes data labeling and data enhancement. Among them, the data labeling can be completed by using the Labelme labeling software to draw a frame on the psoas major muscle, the quadratus lumborum muscle, the erector spinae muscle, the transverse process, and the lumbar plexus nerve in the aforementioned clover structure; the data enhancement includes processing the aforementioned clover view in the manner of occlusion, cropping, random rotation, random flipping, and / or random mirroring.

[0075] In addition, when pre-processing the ultrasound image, first, the original clover views are renumbered to protect patient privacy information.

[0076] Then, select the clover view with complete and clear boundaries of the psoas major muscle, the quadratus lumborum muscle, the erector spinae muscle, the transverse process, and the lumbar plexus nerve, and for each clover view, the psoas major muscle, the quadratus lumborum muscle, the erector spinae muscle, the transverse process, and the lumbar plexus nerve on the clover view are identified by a plurality of anesthesiologists with many years of clinical experience.

[0077] Finally, the professional personnel confirm that the image data marked in the aforementioned clover view is valid, generate a json file format, and convert it into a PASCAL VOC dataset format to provide the U-Net model in this embodiment for image segmentation identification.

[0078] S102, a U-Net model capable of identifying the psoas major muscle, the quadratus lumborum muscle, the erector spinae muscle, the transverse process and the lumbar plexus is constructed and trained.

[0079] In the present embodiment, 80% of the clover views in the lumbar ultrasound image dataset constitute a training set for training the aforementioned U-Net model; the remaining 20% of the clover views constitute a test set for verifying the aforementioned U-Net model.

[0080] Referring to Figure 2 Fig. 1 shows a structural schematic diagram of the U-Net model provided in the present embodiment.

[0081] The U-Net model comprises an encoder, a skip connection and a decoder; the encoder uses a multi-layer dynamic context-aware dilated residual module DCDR to perform a downsampling operation; the decoder uses a multi-layer attention-enhanced hybrid module AEH to perform an upsampling operation.

[0082] The English full name of the dynamic context-aware dilated residual module DCDR is Dynamic Context-Aware Dilated Residual; the English full name of the attention-enhanced hybrid module AEH is Attention-Enhanced Hybrid Module.

[0083] Referring to Figure 3 Fig. 2 shows a structural schematic diagram of the DCDR module in the encoder provided in the present embodiment. The DCDR module is configured with an adaptive multi-scale dilated branch AMD and a deep feature residual branch DFR which are executed in parallel. The English full name of the adaptive multi-scale dilated branch AMD is Adaptive Multi-Scale Dilated; the English full name of the deep feature residual branch DFR is Deep Feature Residual Branch.

[0084] When executed in parallel, the context information in the clover view is preferably extracted through the aforementioned AMD branch; the low-level features in the clover view, such as local detail information, are extracted through the DFR branch. In the present embodiment, after fusing the output features of the AMD branch and the DFR branch, the expression of image features is preferably enhanced through the ECA attention mechanism after 1x1 convolution, batch normalization (Batch Normalization, abbreviated as BN) and rectified linear unit (Rectified Linear Unit, abbreviated as ReLU).

[0085] As one of the preferred embodiments of the present embodiment, the AMD branch can not only process the context information in the image by dynamically adjusting the dilation rate, but also combine the context-aware features of different scales.

[0086] Therefore, before the AMD branch is executed, the multi-stage dilation rate predictor and the context-aware dilation rate predictor are used to respectively predict N different scales of dilation rates, N being an integer greater than or equal to 2.

[0087] The full English name of the multi-stage dilation rate predictor is MultiStage Dilation Predictor, and the full English name of the context-aware dilation rate predictor is Context-Aware Dilation Predictor.

[0088] In the present embodiment, the multi-stage dilation rate predictor and the context-aware dilation rate predictor can respectively predict three different scales of dilation rates. Moreover, the multi-stage dilation rate predictor and the context-aware dilation rate predictor can dynamically predict appropriate dilation rates for different scales according to the input feature map.

[0089] Referring to Figure 4 As shown in the figure, when the multi-stage dilation rate predictor is used, the initial step S110 and the refinement step S120 are sequentially performed to predict the dilation rate.

[0090] In the initial stage, when the initial step S110 is performed, 1×1 convolution, ReLU function and 1×1 convolution are sequentially used to extract local information from the input feature map and generate an initial dilation rate prediction value. The initial step S110 adopts a lightweight convolution operation, which can quickly extract key features and lay a foundation for subsequent fine prediction.

[0091] In the refinement stage, when the refinement step S120 is performed, the initial dilation rate prediction value is adjusted to improve the accuracy of the dilation rate prediction value.

[0092] Specifically, the refinement step S120 includes:

[0093] S121, after the input feature map is dilated convolved using the initial dilation rate prediction value, the initial feature map is spliced. This operation can enhance the diversity of image features and context information, and provide a context basis for accurate prediction of the dilation rate.

[0094] S122, 1×1 convolution, ReLU function and 1×1 convolution are sequentially used to capture the context information in the feature map to correct the error of the initial dilation rate prediction value, and output the corrected dilation rate prediction value to provide more accurate input for subsequent dynamic dilated convolution.

[0095] In the use of the context-aware dilation rate predictor, a step S130 is further included:

[0096] S131, global information of the input feature map is extracted by adaptive average pooling (AAP) to compress the spatial dimension of the input feature map. Referring to Figure 5 As shown, using adaptive average pooling (AAP) can compress the size of the feature map from C, H, W to C, 1, 1. This operation helps to capture the overall context features of the image.

[0097] S132, the compressed input feature map is subjected to at least one 1x1 convolution to obtain three dilation rate prediction values of different scales.

[0098] As one of the preferred embodiments of the present embodiment, it is preferred to perform two 1x1 convolutions. Among them, the first 1x1 convolution aims to map to a new feature space for further processing of global information. Then through the second 1x1 convolution, three dilation rate prediction values of different scales are generated corresponding to three channels.

[0099] By introducing the context-aware dilation rate predictor in the present embodiment, errors caused by insufficient local features can be effectively avoided, thereby improving the effect of subsequent convolution operations.

[0100] After obtaining the dilation rate prediction values of the multi-stage dilation rate predictor and the context-aware dilation rate predictor, the Figure 6 As shown, it is preferred to use Softmax to normalize the dilation rate prediction values of the multi-stage dilation rate predictor and the context-aware dilation rate predictor respectively, so as to obtain a standardized weight distribution to represent the importance of each dilation rate prediction value in different contexts.

[0101] After normalization, through element-by-element addition and average fusion, three-channel dilation rate weight values d1, d2 and d3 are obtained. The aforementioned dilation rate weight values d1, d2 and d3 are respectively amplified, and corresponding dilation rate prediction values D1, D2 and D3 are obtained.

[0102] When performing the AMD branch, in combination with Figure 7 As shown, step S140 is included:

[0103] S141, for the feature Figure X The aforementioned dilation rate prediction values D1, D2 and D3 are respectively applied to the 3x3 dilated convolution to adjust the receptive field of the convolution operation, and different scale features X1, X2 and X3 are obtained after the aforementioned convolution operation. Figure X 1, X2 and X3.

[0104] In this embodiment, the AMD branch can use three 3×3 dilated convolutional blocks to control the degree of dilation for different convolution operations by using the obtained dilation rate prediction values ​​D1, D2, and D3 respectively.

[0105] S142, the aforementioned features Figure X 1. X2 and X3 are weighted according to their corresponding weights d1, d2 and d3 respectively.

[0106] S143, weighted features Figure X 1. X2 and X3, with features sequentially processed by 1×1 and 3×3 convolutions. Figure X Perform element-by-element addition.

[0107] S144: After element-wise addition, apply the ECA attention mechanism to the resulting feature map and output new features. Figure X '.

[0108] The above operations can ensure that features at different scales are fused together, thereby improving the diversity of image features.

[0109] When executing the DFR branch, see [link / reference]. Figure 8 As shown, step S150 is included:

[0110] S151, adjust the number of channels in the input feature map to obtain feature map Out1.

[0111] As an example rather than a limitation, it is preferable to obtain feature map Out1 by adjusting the number of channels of the input feature map through a 1x1 convolution.

[0112] S152, for the aforementioned feature map Out1, after performing multiple convolution operations of different sizes, the feature maps are concatenated, and a convolution operation is performed on the concatenated feature maps. By adjusting the number of channels of the feature maps, feature map Out2 is obtained.

[0113] This is an example, not a limitation. For the aforementioned feature map Out1, after performing 1x1, 3x3, and 5x5 convolution operations respectively, it is concatenated to form a multi-scale feature representation. A 1x1 convolution operation is then performed on the concatenated feature map to reduce the number of channels, resulting in feature map Out2. Furthermore, the same operation is applied to feature map Out2, i.e., step S153 is executed to obtain feature map Out3.

[0114] S153, for the aforementioned feature map Out2, after performing multiple convolution operations of different sizes, the feature maps are concatenated, and a convolution operation is performed on the concatenated feature maps. By adjusting the number of channels of the feature maps, feature map Out3 is obtained.

[0115] S154, residual connection is performed on the aforementioned feature maps Out1, Out2 and Out3 with the same size. This embodiment preferably performs element-by-element addition on the feature maps Out1, Out2 and Out3 with the same C, H and W size through residual connection.

[0116] wherein the learnable weight coefficients W1, W2 and W3 are respectively set for the aforementioned feature maps Out1, Out2 and Out3. The learnable weight coefficients W1, W2 and W3 can be used to control the proportion of the aforementioned feature maps Out1, Out2 and Out3, respectively, and the values of the aforementioned learnable weight coefficients W1, W2 and W3 are optimized and adjusted during the training of the U-Net model.

[0117] The DFR branch combines the design ideas of multi-scale convolution, weighted summation and residual connection. When obtaining the feature maps Out2 and Out3, different size features are extracted through parallel use of 1x1, 3x3 and 5x5 convolution kernels to enhance the feature expression ability of the network for complex images. Moreover, the presence of the learnable weight coefficients W1, W2 and W3 in the DFR branch enables dynamic adjustment of the contribution of different convolution branches to the output.

[0118] Based on this, the advantage of setting the DFR branch in this embodiment is that the learnable weight coefficients W1, W2 and W3 can be set to control the contribution of each convolution path. These learnable weight coefficients are learnable parameters in this embodiment, which can be optimized during the training process. In this way, the DFR branch can process multi-scale features while maintaining the stability of model training to better adapt to specific task requirements.

[0119] In addition, it is also worth mentioning that since the residual connection can connect the feature information of low and high layers through cross-layer connection to solve the problems of gradient vanishing and gradient explosion, the use of residual connection to combine the weighted feature maps and the original input feature maps to obtain the final output feature maps can better transmit features between layers, thereby improving the training efficiency, model performance and generalization ability.

[0120] Therefore, the introduction of residual connection in the DFR branch can effectively alleviate the gradient vanishing problem in the deep layers of the model, ensuring the effective propagation of gradients and the stability of network training. When the number of input and output channels does not match, channel adjustment is performed through 1x1 convolution to ensure smooth execution of the residual connection.

[0121] In addition, the embodiment also considers that the boundary of the clover view to be segmented and recognized may be unclear, and therefore the decoder structure of the conventional U-Net model may have problems such as insufficient feature fusion, missing of context information, and loss of high-frequency details, and therefore the embodiment preferably uses the AEH module to improve the segmentation accuracy.

[0122] In combination Figure 9 As one of the preferred embodiments of the embodiment, a structural schematic diagram of the AEH module is provided for the embodiment.

[0123] The AEH module is configured with an enhanced residual multi-head attention mechanism module ER-MHA and a spatial attention mechanism module SA.

[0124] The English full name of the enhanced residual multi-head attention mechanism module ER-MHA is Enhanced Residual Multi-Head Attention module, and the English full name of the spatial attention mechanism module SA is Spatial Attention Module.

[0125] The number of layers of the AEH module is set according to the number of layers of the encoder; each layer of the AEH module uses a transposed convolution operation; in combination Figure 9 As shown in the figure, X 0 -X 4 are respectively the up-sampling transposed convolution operations performed by each layer in the encoder, that is, the conventional up-sampling operation.

[0126] After using the transposed convolution operation on each layer of the decoder, the ER-MHA module is further used to pass the feature maps of different sizes layer by layer, and the SA module is executed for the output feature map after the up-sampling operation of at least one layer of the decoder, and the feature maps of different scales are spliced in the last layer of the decoder to obtain the final output feature map.

[0127] As one of the preferred embodiments of the embodiment, the AEH module is composed of two parts. In the first part, the ER-MHA module is added after each up-sampling operation, which not only avoids the excessive attention of the traditional splicing method to irrelevant information, but also enables the model to weight the features in a larger range, thereby improving the understanding of complex structures and helping the model to better understand the global context. In the second part, the spatial attention mechanism module is preferably used for each layer of the decoder, and the feature maps of different scales are spliced in the last layer of the decoder to fully exploit the detailed information of each layer and enhance the global perception ability of the model while avoiding potential interference during inter-layer information transmission.

[0128] As another preferred embodiment of the present embodiment, on the basis of the multi-head attention mechanism, the present embodiment proposes an enhanced residual multi-head attention mechanism ER-MHA.

[0129] In the present embodiment, in combination with Figure 10 As shown in the figure, the ER-MHA is preferably configured to, after generating an output value using a multi-head attention mechanism module, add a feedforward network FFN; at the same time, combine the input feature map with the output of the aforementioned feedforward network through a residual connection; and then use a layer normalization operation and a preset learnable scaling factor to obtain an output feature map.

[0130] Among them, the feedforward network is added after the multi-head attention mechanism module generates an output value, which can enhance the nonlinear transformation ability of the ER-MHA module, so that it can capture more data when processing complex tasks. The multi-head attention mechanism module is the existing technology in the art, so it will not be described in more detail here.

[0131] At the same time, the input feature map and the output of the feedforward network are combined through a residual connection, which can ensure the effective propagation of the gradient.

[0132] And the layer normalization and learnable scaling factor are added in the ER-MHA module, which not only improves the stability of the training and ensures that the model can converge more quickly and stably during the training process, but also enhances the flexibility and task adaptability of the model, thereby enhancing the capture of fuzzy areas.

[0133] S103, configure the trained U-Net model in an ultrasound device, and the ultrasound device can collect a trident view.

[0134] This operation can deploy the trained U-Net model to the local ultrasound device, so that it can assist the anesthesiologist in the clinical process of the lumbar plexus nerve tissue.

[0135] S104, using the aforementioned U-Net model configured in the ultrasound device to perform image recognition on the collected trident view, and obtaining a recognition result.

[0136] The improved U-Net model can quickly identify and locate the lumbar plexus nerve when the anesthesiologist performs anesthesia operations, and can also perform real-time identification and detection on the trident view during the use of the ultrasound device.

[0137] The U-Net model described in the present embodiment can use the mIOU coefficient as an evaluation index for measuring the performance of the model in the segmentation task to verify the final model effect.

[0138] The mIOU (Mean Intersection over Union) is used as an indicator to evaluate the segmentation accuracy, defined as the ratio of the intersection to the union of the predicted and true regions. Specifically, the IOU is calculated for each class, and then averaged over all classes.

[0139] The mIOU measures the segmentation accuracy of the model on different abdominal tissue structures. A high mIOU indicates that the model performs well in identifying and segmenting these structures.

[0140] To verify the recognition effect of the U-Net model on the clover view, the experimental results show that for normal ultrasound images, the psoas major muscle, quadratus lumborum muscle, erector spinae muscle, transverse process, and lumbar plexus nerve in the clover view can be accurately identified, among which the mIoU of the psoas major muscle, quadratus lumborum muscle, erector spinae muscle, and transverse process is 0.845, 0.806, 0.837, and 0.876 respectively, achieving good results, and the mIoU of the lumbar plexus nerve is 0.629. The original U-Net only identifies the psoas major muscle and transverse process, and does not achieve recognition of the remaining parts.

[0141] It can be seen that the improved U-Net model has a higher mIoU value and more accurate recognition and positioning.

[0142] Other technical features are described in the previous embodiments and will not be repeated here.

[0143] In addition, an embodiment of the present application is provided, which provides a U-Net-based lumbar plexus nerve ultrasound segmentation and recognition system 200, comprising:

[0144] An ultrasound device 201 is used to collect a clover view.

[0145] A model configuration module 202 is used to configure the trained U-Net model on the aforementioned ultrasound device.

[0146] A system server 203 is connected to the ultrasound device and the model configuration module 202.

[0147] The system server 203 is configured to: pre-process an ultrasound image for a preset clover view, and construct a lumbar ultrasound image dataset; the clover view is an ultrasound image; the clover view includes psoas major muscle, quadratus lumborum muscle, erector spinae muscle, transverse process and lumbar plexus nerve; a U-Net model capable of identifying the psoas major muscle, the quadratus lumborum muscle, the erector spinae muscle, the transverse process and the lumbar plexus nerve is constructed and trained; the U-Net model includes an encoder, a skip connection and a decoder; the encoder uses a multi-layer dynamic context perception expansion residual module DCDR to perform a downsampling operation; the decoder uses a multi-layer attention enhancement hybrid module AEH to perform an upsampling operation; the trained U-Net model is configured in an ultrasound device capable of collecting the clover view; the collected clover view is subjected to image recognition using the aforementioned U-Net model configured in the ultrasound device, and an identification result is obtained.

[0148] Other technical features are described in the preceding embodiments and will not be repeated here.

[0149] In addition, the embodiment of the present application also provides a computer readable storage medium, which stores a program for use in the aforementioned U-Net-based lumbar plexus nerve ultrasound segmentation and identification system, and the program can implement the steps of any one of the aforementioned U-Net-based lumbar plexus nerve ultrasound segmentation and identification methods when executed by a processor.

[0150] Other technical features are described in the preceding embodiments and will not be repeated here.

[0151] In the above description, within the target protection scope of the present disclosure, each component can be selectively and operatively combined in any number. In addition, the terms such as "include", "comprise", and "have" should be interpreted as inclusive or open, rather than exclusive or closed, unless they are explicitly defined as the opposite meaning. All technical, scientific or other terms are in accordance with the meaning understood by those skilled in the art, unless they are defined as the opposite meaning. Common terms found in dictionaries should not be interpreted too ideally or too unrealistically in the context of relevant technical documents, unless the present disclosure explicitly limits them as such.

[0152] Although the example aspects of the present disclosure have been described for illustrative purposes, those skilled in the art should realize that the above description is only a description of preferred embodiments of the present application, and is not any limitation on the scope of the present application, the scope of the preferred embodiments of the present application includes additional implementations, in which the functions can be performed in an order different from that described or discussed. Any modification or modification made by those skilled in the art according to the above disclosure is within the protection scope of the claims.

Claims

1. A U-Net-based method for ultrasound segmentation of the lumbar plexus, characterized in that, Specifically, it includes: Ultrasound image preprocessing is performed on a preset clover view to construct a lumbar ultrasound image dataset; the clover view is an ultrasound image; the clover view includes the psoas major muscle, quadratus lumborum muscle, erector spinae muscle, transverse process and lumbar plexus nerve; A U-Net model capable of recognizing the psoas major, quadratus lumborum, erector spinae, transverse processes, and lumbar plexus nerves was constructed and trained. The U-Net model includes an encoder, skip connections, and a decoder. The encoder uses a multi-layer dynamic context-aware dilated residual module (DCDR) to perform downsampling. Within the encoder, the DCDR module is configured with a parallel-executed adaptive multi-scale dilated branch (AMD) and a depth feature residual branch (DFR). During parallel execution, the AMD branch extracts contextual information from the cloverleaf view, while the DFR branch extracts local detail information. After fusing the output features of the AMD and DFR branches, an ECA attention mechanism is used to enhance the representation of image features. The decoder uses a multi-layer attention-enhanced hybrid module (AEH) to perform upsampling. The AEH module is configured with enhancement... The decoder includes a residual multi-head attention mechanism module (ER-MHA) and a spatial attention mechanism module (SA). The number of layers in the AEH module corresponds to the number of layers in the encoder. Each layer of the AEH module uses a transposed convolution operation. After applying transposed convolution operations to each layer of the decoder, the ER-MHA module is used layer by layer to pass feature maps of different sizes. The SA module is executed on the output feature map after at least one upsampling operation in the decoder, and feature maps of different scales are concatenated in the last layer of the decoder to obtain the final output feature map. The ER-MHA is configured to: after generating output values ​​using the multi-head attention mechanism module, add a feedforward network (FFN); simultaneously, combine the input feature map with the output of the aforementioned feedforward network through residual connections; then, use layer normalization operations and a preset learnable scaling factor to obtain the output feature map. The trained U-Net model is configured in an ultrasound device capable of acquiring clover views; The acquired clover view was used to perform image recognition using the aforementioned U-Net model configured in the ultrasound device, and the recognition results were obtained.

2. The method according to claim 1, characterized in that, The ultrasound image preprocessing includes data annotation and data augmentation; wherein, the data annotation can be completed by drawing borders for the psoas major, quadratus lumborum, erector spinae, transverse processes and lumbar plexus in the aforementioned cloverleaf structure using Labelme annotation software; The data augmentation includes processing the aforementioned clover view using methods such as occlusion, cropping, random rotation, random flipping, and / or random mirroring; The 80% cloverleaf view of the lumbar ultrasound image dataset is used as the training set for training the aforementioned U-Net model; the remaining 20% ​​of the cloverleaf view is used as the test set for validating the aforementioned U-Net model.

3. The method according to claim 1, characterized in that, Before executing the AMD branch, a multi-stage inflation rate predictor and a context-aware inflation rate predictor are used to predict inflation rates at N different scales, where N is an integer greater than or equal to 2; where, When using the multi-stage expansion rate predictor, the initial step S110 and the refinement step S120 are executed sequentially; wherein, When performing the initial step S110, 1×1 convolution, ReLU function and 1×1 convolution are used sequentially to extract local information from the input feature map and generate the initial dilation rate prediction value. When performing the S120 refinement step, the steps include: S121, using the aforementioned initial dilation rate prediction value to perform dilation convolution on the input feature map, and then concatenating it with the initial feature map; S122, sequentially using 1×1 convolution, ReLU function and 1×1 convolution to capture the context information in the feature map, so as to correct the error of the aforementioned initial dilation rate prediction value, and outputting the corrected dilation rate prediction value. When using the context-aware dilation rate predictor, the method further includes step S130: S131, extracting global information from the input feature map through adaptive average pooling (AAP) and compressing the spatial dimension of the input feature map; S132, performing at least one 1×1 convolution on the compressed input feature map to obtain dilation rate prediction values ​​at three different scales. The inflation rate predictions of the multi-stage inflation rate predictor and the context-aware inflation rate predictor are normalized using Softmax. After normalization, the expansion rate weight values ​​d1, d2 and d3 of the three channels are obtained by adding and averaging each element. The expansion rate weight values ​​d1, d2 and d3 are amplified respectively, and the corresponding expansion rate prediction values ​​D1, D2 and D3 are obtained.

4. The method according to claim 3, characterized in that, When executing the AMD branch, step S140 is included: S141, For the feature map X, the aforementioned dilation rate prediction values ​​D1, D2 and D3 are applied to a 3×3 dilated convolution to adjust the receptive field of the convolution operation. After the aforementioned convolution operation, feature maps X1, X2 and X3 of different scales are obtained. S142, the aforementioned feature maps X1, X2 and X3 are weighted according to their corresponding weights d1, d2 and d3 respectively; S143, add the weighted feature maps X1, X2 and X3 to the feature map X that has been sequentially convolved with 1×1 and 3×3, element by element; S144: After element-wise addition, apply the ECA attention mechanism to the feature map after element-wise addition and output a new feature map X'.

5. The method according to claim 1, characterized in that, When executing the DFR branch, step S150 is included: S151, Adjust the number of channels in the input feature map to obtain feature map Out1; S152, for the aforementioned feature map Out1, after performing multiple convolution operations of different sizes, the feature map is concatenated, and a convolution operation is performed on the concatenated feature map to obtain feature map Out2 by adjusting the number of channels of the feature map; S153, for the aforementioned feature map Out2, after performing multiple convolution operations of different sizes, the feature map is stitched together, and a convolution operation is performed on the stitched feature map. The feature map Out3 is obtained by adjusting the number of channels of the feature map. S154, perform residual connections on the aforementioned feature maps Out1, Out2 and Out3 with the same size; wherein, learnable weight coefficients W1, W2 and W3 are set for the aforementioned feature maps Out1, Out2 and Out3 respectively; the learnable weight coefficients W1, W2 and W3 can be used to control the proportion weight of the aforementioned feature maps Out1, Out2 and Out3 respectively, and optimize and adjust the values ​​of the aforementioned learnable weight coefficients W1, W2 and W3 during the U-Net model training process.

6. A U-Net-based lumbar plexus ultrasound segmentation system according to any one of claims 1-5, characterized in that... include: Ultrasonic equipment used to acquire views of clover; The model configuration module is used to configure the trained U-Net model on the aforementioned ultrasound device; System server, which connects to the ultrasound equipment and the model configuration module; The system server is configured to: perform ultrasound image preprocessing on a preset cloverleaf view to construct a lumbar ultrasound image dataset; the cloverleaf view is an ultrasound image; the cloverleaf view includes the psoas major, quadratus lumborum, erector spinae, transverse processes, and lumbar plexus; construct and train a U-Net model capable of recognizing the psoas major, quadratus lumborum, erector spinae, transverse processes, and lumbar plexus; the U-Net model includes an encoder, skip connections, and a decoder; wherein, the encoder uses a multi-layer dynamic context-aware extended residual module DCDR to perform downsampling operations; in the encoder, the DCDR module is configured with a parallel-executed adaptive multi-scale expansion branch AMD and a depth feature residual branch DFR; during parallel execution, the aforementioned AMD branch is used to extract contextual information in the cloverleaf view; the DFR branch extracts local detail information in the cloverleaf view; after fusing the output features of the AMD branch and the DFR branch, the image feature expression is enhanced through an ECA attention mechanism; the decoder uses a multi-layer attention enhancement hybrid module AEH to perform upsampling operations; the AEH module is configured with... The system includes an Enhanced Residual Multi-Head Attention (ER-MHA) module and a Spatial Attention (SA) module. The number of layers in the AEH module corresponds to the number of layers in the encoder. Each layer of the AEH module uses a transposed convolution operation. After applying transposed convolution operations to each layer of the decoder, the system further includes using the ER-MHA module layer by layer to pass feature maps of different sizes. The SA module is executed on the output feature map after at least one layer of upsampling in the decoder, and the feature maps of different scales are concatenated in the last layer of the decoder to obtain the final output feature map. The ER-MHA is configured to: after generating output values ​​using the multi-head attention module, add a feedforward network (FFN); simultaneously, combine the input feature map with the output of the aforementioned feedforward network through residual connections; then, use layer normalization operations and a preset learnable scaling factor to obtain the output feature map; configure the trained U-Net model in an ultrasound device capable of acquiring cloverleaf views; perform image recognition on the acquired cloverleaf views using the aforementioned U-Net model configured in the ultrasound device, and obtain the recognition result.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-5.

Citation Information

Patent Citations

  • Method for realizing ultrasonic guidance assisted popliteal sciatic nerve block based on U-Net and application

    CN119992189A