Transverse abdominal muscle plane tissue structure recognition method and application based on RMSD-Net

By constructing the RMSD-Net model, the problem of inaccurate anatomical structure identification in ultrasound-guided transverse abdominal plane block was solved, the automatic segmentation and identification of the transverse abdominal plane tissue structure was achieved, and the positioning accuracy and safety of the block area were improved.

CN119206215BActive Publication Date: 2025-09-12THE NAVAL MEDICAL UNIV OF PLA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411238076.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2025-09-12
Estimated Expiration
2044-09-05

AI Technical Summary

Technical Problem

In the existing technology, ultrasound-guided transverse abdominal plane block has problems such as inaccurate anatomical structure identification, high operation dependence, and high risk of complications. Especially in the case of poor ultrasound imaging quality or anatomical variation, it is difficult to achieve accurate positioning of the block area.

Method used

The transverse abdominal muscle plane tissue structure recognition method based on RMSD-Net is adopted. By constructing the RMSD-Net model with a U-shaped network structure, combined with the multi-scale residual feature extraction module RMDC, skip connection structure and attention mechanism, the automatic segmentation and recognition of the transverse abdominal muscle plane tissue structure is achieved.

Benefits of technology

It improves the recognition accuracy of the transverse abdominal muscle plane tissue structure, assists clinicians in quickly and accurately locating the blocking area, reduces the risk of misoperation, and improves the success rate and safety of ultrasound-guided TAP block.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206215B_ABST
    Figure CN119206215B_ABST
Patent Text Reader

Abstract

The present invention provides a method and application for identifying the plane tissue structure of the transverse abdominal muscle based on RMSD-Net, and relates to the field of image processing technology. The method comprises the following steps: performing data preprocessing on an acquired ultrasound image with a plane structure of the transverse abdominal muscle to establish an abdominal ultrasound image dataset; wherein the ultrasound image with the plane structure of the transverse abdominal muscle includes key abdominal tissue structure information that can be observed during TAP block; constructing an RMSD-Net model with a U-shaped network structure; the RMSD-Net model comprises an encoder, a jump connection structure, and a decoder; a multi-scale residual feature extraction module RMDC module is added to each layer of the encoder; the jump connection structure uses a feature fusion module SFF module layer by layer; an attention mechanism GAM module is added to the last layer of the decoder; a segmentation task is performed on the aforementioned key abdominal tissue structure information, the aforementioned abdominal ultrasound image dataset is divided into a training set and a test set, and the aforementioned RMSD-Net model is used for training and testing respectively; the tested RMSD-Net model is used to segment and identify the key abdominal tissue structure information of the newly acquired ultrasound image with the plane structure of the transverse abdominal muscle. The present invention can help clinicians to quickly and accurately interpret ultrasound images with the transverse abdominal muscle plane structure, so as to assist them in accurately locating the blocking area and injecting local anesthetics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a transverse abdominal muscle plane tissue structure recognition method and application based on RMSD-Net. Background Art

[0002] Nerve blocks (NBs) are a technique that uses various physical and chemical methods to temporarily or permanently block nerve conduction at ganglia, roots, plexuses, trunks, and nerve endings. They are widely used in the clinical diagnosis and treatment of chronic pain disorders and in regional anesthesia. Traditional peripheral nerve block techniques typically rely on surface paresthesia, nerve stimulation, and physician experience to locate the nerves and blockade area, resulting in low efficiency and numerous postoperative complications.

[0003] In recent years, with the development and popularization of ultrasound technology, ultrasound imaging provides a visual field of view, and ultrasound-guided nerve block has become the mainstream method for anesthesiologists and pain physicians to perform nerve block, greatly improving accuracy and safety.

[0004] Among them, the ultrasound-guided transversus abdominis plane block (TAPB) is a widely used nerve block technique in clinical practice. It involves injecting local anesthetic into the neurofascial plane between the internal oblique and transverse abdominal muscles, blocking afferent nerves from T6 to L1. This provides effective analgesia for various abdominal surgeries, such as gynecological laparoscopy, laparoscopic colorectal surgery, pediatric hernia surgery, and cesarean sections. TAP blocks include four approaches: subcostal, anterior, posterior, and lateral. Different approaches result in different block ranges and corresponding clinical applications. The appropriate block approach can be selected based on clinical needs.

[0005] With the increasing clinical application of TAP block, in order to meet the requirements of comfortable medical care and precise diagnosis and treatment, ultrasound-guided TAP block has gradually become a gold standard for regional block anesthesia.

[0006] However, whether it is traditional TAP block or ultrasound-guided TAP block, incorrect or inaccurate positioning of the block area may lead to poor anesthesia effect, and there is a risk of accidentally entering the neurovascular space, resulting in complications such as local anesthetic poisoning, local hematoma, and nerve damage.

[0007] Due to the low spatial resolution and speckle noise of ultrasound imaging, subtle anatomical features cannot be clearly distinguished from the surrounding background by the naked eye. In particular, clinicians who have only received limited ultrasound image reading training face great challenges in interpreting ultrasound images.

[0008] At the same time, for deeper blocking targets, poor ultrasound imaging quality may result in the puncture needle not being clearly visualized. In this case, blind needle insertion may increase the incidence of complications.

[0009] In addition, the success of ultrasound-guided TAP block is highly dependent on the quality of ultrasound image acquisition and the skill of the operator, especially in cases of anatomical variations or scanning difficulties (such as obesity). The entire process requires the doctor to have strong image recognition capabilities.

[0010] Based on this, in view of the complexity of the abdominal anatomical structure, the variability of ultrasound images and operator dependence, the present invention provides a transverse abdominal muscle plane tissue structure recognition method and application based on RMSD-Net to solve the problem of the inability to accurately and automatically segment and recognize the transverse abdominal muscle plane tissue structure information, which is a technical problem that needs to be solved urgently. Summary of the Invention

[0011] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a method and application for identifying the plane tissue structure of the transverse abdominal muscle based on RMSD-Net. The present invention realizes the automatic segmentation and recognition of the plane tissue structure information of the transverse abdominal muscle by constructing an efficient deep learning model, thereby helping clinicians to quickly and accurately interpret ultrasound images of the TAP area, and assist them in accurately locating the block area and injecting local anesthetics.

[0012] In order to solve the existing technical problems, the present invention provides the following technical solutions:

[0013] A method for identifying the planar tissue structure of the transverse abdominal muscle based on RMSD-Net, comprising the following steps:

[0014] Data preprocessing is performed on the acquired ultrasound images with the transverse abdominal muscle plane structure to establish an abdominal ultrasound image dataset; wherein the ultrasound images with the transverse abdominal muscle plane structure include key abdominal tissue structure information that can be observed during TAP block; the key abdominal tissue structure information includes the external oblique muscle, internal oblique muscle, transverse abdominal muscle, rectus abdominis muscle, and anterior superior iliac spine;

[0015] Constructing an RMSD-Net model with a U-shaped network structure; the RMSD-Net model includes an encoder, a skip connection structure, and a decoder; the encoder uses the original double convolution structure in each layer, and adds a multi-scale residual feature extraction module RMDC module in each layer to extract features of the input ultrasound image with the transverse abdominal muscle plane structure layer by layer, thereby strengthening the important features of the input ultrasound image with the transverse abdominal muscle plane structure layer by layer, and suppressing secondary information at the same time; the skip connection structure uses a feature fusion module SFF module layer by layer to enhance feature expression and promote the integration of contextual information; an attention mechanism GAM module is added to the last layer of the decoder to combine channel attention and spatial attention to integrate local information and global information in the feature map, and dynamically adjust the weight to provide feature representation for subsequent feature weighting;

[0016] Perform segmentation tasks on the aforementioned key abdominal tissue structure information, divide the aforementioned abdominal ultrasound image dataset into a training set and a test set, and use the aforementioned RMSD-Net model for training and testing respectively;

[0017] The tested RMSD-Net model is used to segment and identify key abdominal tissue structure information in newly acquired ultrasound images with transverse abdominal muscle plane structure.

[0018] Furthermore, the ultrasound image with the transverse abdominal muscle plane structure is a frame image in an ultrasound imaging video that can observe the corresponding muscle tissue and is recorded during the transverse abdominal muscle plane scanning process;

[0019] The data preprocessing includes data labeling and data enhancement;

[0020] The data annotation can use Labelme annotation software to annotate the key abdominal tissue structure information in the aforementioned ultrasound image by drawing contour lines;

[0021] The data enhancement includes processing the ultrasound image having the transverse abdominal muscle plane structure by random rotation, random flipping and / or random mirroring;

[0022] 80% of the ultrasound images with the transverse abdominal muscle plane structure were extracted from the abdominal ultrasound image dataset to form a training set for training the aforementioned RMSD-Net model, and the remaining 20% ​​of the ultrasound images with the transverse abdominal muscle plane structure were composed of a test set for verifying the aforementioned RMSD-Net model.

[0023] Furthermore, in the encoder, the execution of the RMDC module includes step S110:

[0024] S111, using three dilated convolutional layers with different dilation rates to capture spatial information at different scales, thereby effectively extracting global and local information of the image; when the dilation rates are 1, 2, and 4, the corresponding feature representations of different scales are {X1, X2, X3}, thus obtaining X i =F i (X); where i = 1, 2 or 3, F i (X) represents the convolution operation with different expansion rates, and X represents the input feature map;

[0025] S112: After passing the output features of all dilated convolutional layers through the BN layer, they are concatenated along the channel dimension, and then integrated through the voting convolution layer voteConv layer. The weight map W of the fusion feature is obtained through the Sigmoid function σ. 11 , that is, W 11 =σ(F vote ([X1, X2, X3])); where F vote () indicates the voting convolution layer operation, and [X1, X2, X3] indicates the feature concatenation operation; the voteConv layer can strengthen key structural features and suppress secondary information by adjusting and filtering features extracted at different scales;

[0026] S113, use residual connection to combine the weighted feature map with the original input feature map X to obtain the final output feature map Y, that is, Y = X + X * W 11 .

[0027] Furthermore, the output feature map Y of each layer encoder corresponds to the kth layer and is marked as the output feature map E k In the jump connection structure, the SFF module converts the output feature map E of the encoder corresponding to each layer k After 3*3 convolution, the input feature map SFF corresponding to the SFF module k+1 After upsampling and element-by-element addition, features are extracted and output through the multi-scale depth-separable hybrid attention mechanism module MDSA module, where k is a positive integer. The overall structure of the MDSA module includes channel attention and spatial attention. The spatial attention is composed of a multi-scale depth-separable convolution module to dynamically allocate attention weights in the channel dimension and the spatial dimension.

[0028] Furthermore, the execution of the MDSA module includes step S120:

[0029] S121, obtain the input feature map, and obtain the spatial information of the input feature map after channel attention aggregation average pooling and maximum pooling operations;

[0030] S122, processing the aforementioned spatial information through a shared multi-layer perceptron (MLP) to generate a channel attention map; wherein the output feature map after the channel attention mechanism is obtained by element-wise multiplication with the input feature map;

[0031] S123, input the channel prior into the deep convolution module to generate a spatial attention map, receive the spatial attention feature map through a 1*1 convolution block and perform channel fusion;

[0032] S124, multiply the channel mixing result by the channel prior element by element to obtain the refined features as output.

[0033] Furthermore, the execution of the GAM module includes step S130:

[0034] S131, using a channel attention module to process the input; the channel attention module first adjusts the shape of the input tensor to H*W*C through a sizepermute operation, that is, moves the channel dimension to the end, so as to perform global feature calculation in the spatial dimension; then processes the adjusted tensor through a multi-layer perceptron (MLP) to generate global features; the MLP consists of two fully connected layers, using a reduction ratio r to reduce computational complexity, so that the first fully connected layer reduces the number of channels from C to C / r, and the second fully connected layer restores the number of channels from C / r to C; the processed tensor dimension is restored to the original shape C*H*W through a size reverse operation for further operation, and finally a channel attention weight is generated through a sigmoid activation function;

[0035] S132, the result of the processing of the aforementioned channel attention module is processed by the spatial attention module to obtain an output feature map; the execution of the spatial attention module includes: S1321, using two convolutional layers to perform spatial information fusion, S1322, using the same reduction ratio r as the channel attention module in the spatial attention module, S1323, using the Sigmoid activation function to weight the feature map of each position to improve the spatial expression ability of the feature map and reduce the influence of secondary information.

[0036] Furthermore, after the aforementioned tests, mIOU, mPA, Accuracy, and Dice coefficient were used as evaluation indicators to measure the performance of the RMSD-Net model in the segmentation task to verify the final model effect; among them, the mIOU is an indicator for evaluating segmentation accuracy; the mPA refers to the average pixel accuracy in each category; the Accuracy refers to the ratio of correctly classified pixels in the segmentation result to the total number of pixels; and the Dice coefficient is an indicator for evaluating the degree of overlap between the segmentation result and the true annotation.

[0037] A transverse abdominal muscle plane tissue structure recognition device based on RMSD-Net, comprising:

[0038] a data construction unit, configured to perform data preprocessing on an acquired ultrasound image having a transverse abdominal muscle plane structure to establish an abdominal ultrasound image dataset; wherein the ultrasound image having a transverse abdominal muscle plane structure includes key abdominal tissue structure information observable during TAP block; the key abdominal tissue structure information includes the external oblique muscle, internal oblique muscle, transverse abdominal muscle, rectus abdominis muscle, and anterior superior iliac spine;

[0039] A model construction unit is used to construct an RMSD-Net model with a U-shaped network structure; the RMSD-Net model includes an encoder, a skip connection structure, and a decoder; the encoder uses the original double convolution structure in each layer and adds a multi-scale residual feature extraction module RMDC module in each layer to extract features of the input ultrasound image with the transverse abdominal muscle plane structure layer by layer, thereby strengthening the important features of the input ultrasound image with the transverse abdominal muscle plane structure layer by layer and suppressing secondary information at the same time; the skip connection structure uses a feature fusion module SFF module layer by layer to enhance feature expression and promote the integration of contextual information; an attention mechanism GAM module is added to the last layer of the decoder to combine channel attention and spatial attention to integrate local information and global information in the feature map, and dynamically adjust the weight to provide feature representation for subsequent feature weighting;

[0040] A model training and testing unit, configured to perform segmentation tasks for the aforementioned key abdominal tissue structure information, divide the aforementioned abdominal ultrasound image dataset into a training set and a test set, and use the aforementioned RMSD-Net model for training and testing, respectively;

[0041] The image recognition unit is used to use the tested RMSD-Net model to segment and identify key abdominal tissue structure information of the newly acquired ultrasound image with the transverse abdominal muscle plane structure.

[0042] A transverse abdominal muscle plane tissue structure recognition system based on RMSD-Net, comprising:

[0043] A network node, used for transmitting and receiving an ultrasound image of a transverse abdominal muscle plane structure acquired by an ultrasound device;

[0044] A model configuration module is used to configure an RMSD-Net model for segmenting and identifying key abdominal tissue structure information in an ultrasound image having a transverse abdominal muscle planar tissue structure; the key abdominal tissue structure information includes the external abdominal oblique muscle, internal abdominal oblique muscle, transverse abdominal muscle, rectus abdominis muscle, and anterior superior iliac spine;

[0045] A system server, the system server connecting the network nodes and the model arrangement module;

[0046] The system server is configured to: perform data preprocessing on the acquired ultrasound images with the transverse abdominal muscle plane structure to establish an abdominal ultrasound image dataset; wherein, the ultrasound images with the transverse abdominal muscle plane structure include key abdominal tissue structure information that can be observed during TAP block; construct an RMSD-Net model with a U-shaped network structure; the RMSD-Net model includes an encoder, a jump connection structure and a decoder; the encoder uses the original double convolution structure in each layer and adds a multi-scale residual feature extraction module RMDC module in each layer to perform feature extraction on the input ultrasound image with the transverse abdominal muscle plane structure layer by layer, thereby strengthening the aforementioned input ultrasound image with the transverse abdominal muscle plane structure layer by layer. important features and suppress secondary information at the same time; the skip connection structure uses the feature fusion module SFF module layer by layer to enhance feature expression and promote the integration of context information; an attention mechanism GAM module is added to the last layer of the decoder to combine channel attention and spatial attention to integrate local information and global information in the feature map, and dynamically adjust the weights to provide feature representation for subsequent feature weighting; a segmentation task is performed on the aforementioned key abdominal tissue structure information, the aforementioned abdominal ultrasound image dataset is divided into a training set and a test set, and the aforementioned RMSD-Net model is used for training and testing respectively; the tested RMSD-Net model is used to perform segmentation and recognition of key abdominal tissue structure information on the newly acquired ultrasound image with the transverse abdominal muscle plane structure.

[0047] A computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the implementation steps of any of the above methods.

[0048] Based on the above advantages and positive effects, the present invention has the advantage of constructing a data set of abdominal ultrasound images suitable for processing nerve block technology.

[0049] Furthermore, an RMSD-Net model suitable for segmenting and identifying key abdominal tissue structure information in nerve block technology was constructed, and corresponding training and testing were completed.

[0050] Furthermore, an RMDC module is added to the encoder of the RMSD-Net model. The RMDC module can extract feature information of different scales by performing convolution operations on the input feature map under different receptive fields. This enables the model to capture image features at different levels and scales, which helps to improve the model's ability to understand complex scenes.

[0051] Furthermore, an SFF module is added to the skip connection structure of the RMSD-Net model. The SFF module extracts more fine-grained features through the multi-scale depth-separable hybrid attention mechanism module MDSA module and outputs them. The MDSA attention mechanism selects and weights the feature information of different channels by learning the importance weight of each channel in the feature map, which helps the model focus on the most relevant and most discriminative features and improves the expressiveness of the features. At the same time, the attention mechanism can reduce the influence of irrelevant or unimportant feature channels by performing weighted summation on the feature map, thereby reducing the computational complexity of the model.

[0052] Furthermore, a GAM module is added to the decoder of the RMSD-Net model. The GAM module models the correlation between channels and spaces in the feature map on a global scale, thereby better capturing the global dependencies of the feature map, improving the model's overall understanding and expression of input features, and at the same time reducing feature loss during upsampling. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A flowchart of a method provided by an embodiment of the present invention.

[0054] Figure 2 A schematic diagram of the structure of the RMSD-Net model provided in an embodiment of the present invention.

[0055] Figure 3 This is a schematic diagram of the structure of the RMDC module in the encoder provided by an embodiment of the present invention.

[0056] Figure 4 This is a structural diagram of the SFF module in the jump connection structure provided by an embodiment of the present invention.

[0057] Figure 5 This is a schematic structural diagram of a GAM module in a decoder provided by an embodiment of the present invention.

[0058] Figure 6 This is the visualization result of the ablation experiment provided by the embodiment of the present invention.

[0059] Figure 7 This is the visualization result of the comparative experiment provided by the embodiment of the present invention. DETAILED DESCRIPTION

[0060] The following is a further detailed description of a method for identifying the transverse abdominal muscle plane tissue structure based on RMSD-Net and its application disclosed in the present invention in conjunction with the accompanying drawings and specific embodiments. It should be noted that the technical features or combinations of technical features described in the following embodiments should not be considered isolated, and they can be combined with each other to achieve better technical effects. In the drawings of the following embodiments, the same reference numerals appearing in each drawing represent the same features or components, which can be applied to different embodiments. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0061] It should be noted that the structures, proportions, sizes, etc. illustrated in the drawings of this specification are only used to match the contents disclosed in the specification for people familiar with this technology to understand and read, and are not intended to limit the conditions under which the invention can be implemented. Any structural modification, change in proportional relationship, or adjustment of size should fall within the scope of the technical content disclosed in the invention without affecting the efficacy and purpose of the invention. The scope of the preferred embodiments of the present invention includes alternative implementations, in which the functions can be performed in a non-described or discussed order, including performing the functions in a substantially simultaneous manner or in a reverse order according to the functions involved, which should be understood by those skilled in the art of the art to which the embodiments of the present invention belong.

[0062] Technologies, methods, and apparatus known to persons of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and apparatus should be considered part of the specification. In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0063] Example

[0064] See also Figure 1 FIG. 1 is a flow chart of the present invention. The implementation step S100 of the method is as follows:

[0065] S101 , performing data preprocessing on the acquired ultrasound image with the transverse abdominal muscle plane structure to establish an abdominal ultrasound image dataset.

[0066] The ultrasound image with the transverse abdominal muscle plane structure is a frame image in the ultrasound image video recorded during the transverse abdominal muscle plane scanning process, which can clearly observe the corresponding muscle tissue.

[0067] Considering the current lack of publicly available data sets of relevant abdominal ultrasound images, this embodiment collects a large amount of clinical ultrasound data.

[0068] In actual operation, since the collected original data is in video format, in order to select frame images with clear image quality, clear structure, and different content as data for subsequent training, technical personnel in this field (such as professional anesthesiologists) mark the areas to be segmented on the aforementioned frame images. The ultrasound image dataset includes the corresponding key abdominal tissue structure information that can be clearly observed during TAP block.

[0069] It is also worth mentioning that the following targets to be segmented that are of clinical interest are identified and labeled as the transversus abdominis muscle (TA), the rectus abdominis muscle (RA), the external oblique muscle (EO), the internal oblique muscle (IO), and the anterior superior iliac spine (ASIS).

[0070] Corresponding to the aforementioned target to be segmented, it can be determined that the key abdominal tissue structure information includes the external oblique muscle, the internal oblique muscle, the transverse abdominal muscle, the rectus abdominis muscle, and the anterior superior iliac spine.

[0071] The data preprocessing includes data labeling and data enhancement.

[0072] The data annotation can use Labelme annotation software to annotate the key abdominal tissue structure information in the aforementioned ultrasound image by drawing contour lines.

[0073] The data enhancement includes processing the ultrasound image with the transverse abdominal muscle plane structure by random rotation, random flipping and / or random mirroring.

[0074] Specifically, the ultrasound images used in this embodiment are from an authoritative comprehensive tertiary-level Class A hospital. Senior clinical physicians of the hospital collected and compiled approximately 300 videos of the transverse abdominal muscle plane tissue structure of approximately 100 patients.

[0075] These transverse abdominal muscle plane tissue structure videos were captured frame by frame to obtain approximately 8,000 original transverse abdominal muscle plane structure images.

[0076] In order to protect the privacy of patients, these original transverse abdominal muscle plane structure images were renumbered, and 1,980 clear and identifiable images were selected.

[0077] Finally, professionals confirmed that the image data marked in the above pictures were valid, generated in the json file format, and converted into the Pascal voc dataset format for input into the program.

[0078] In order to ensure that the data in the Pascal voc dataset is of high quality, data enhancement methods such as random rotation, random flipping, and random mirroring are used to ensure the quality of the input data, thereby improving the robustness of the trained RMSD-Net model.

[0079] S102, constructing the RMSD-Net model with a U-shaped network structure.

[0080] See also Figure 2 , which is a structural diagram of the RMSD-Net model provided in this embodiment.

[0081] The RMSD-Net model includes an encoder, a skip connection structure, and a decoder.

[0082] The encoder uses the U-Net model in each layer. On the basis of the original double convolution structure of the encoder part, a multi-scale residual feature extraction module RMDC module is added in each layer to perform feature extraction on the input ultrasound image with the transverse abdominal muscle plane structure layer by layer, thereby strengthening the important features of the aforementioned input ultrasound image with the transverse abdominal muscle plane structure layer by layer, and at the same time suppressing secondary information.

[0083] The skip connection structure uses the feature fusion module SFF module layer by layer to enhance feature expression and promote the integration of context information.

[0084] In addition, an attention mechanism GAM module is added to the last layer of the decoder to combine channel attention and spatial attention to integrate local and global information in the feature map, thereby dynamically adjusting the weights and providing feature representation for subsequent feature weighting.

[0085] It is worth noting that in the deep layers of the encoder, two 3×3 convolutional layers are preferably used to extract local image detail features. This is because the spatial size of the feature map has been reduced in the deep layers of the encoder, which means that each pixel represents a larger area in the original image. At this scale, more contextual information can be captured. To this end, the RMDC module is preferably used to further enhance this context capture capability by applying attention at multiple spatial scales.

[0086] For details, see Figure 3 FIG. 4 is a schematic diagram of the structure of the RMDC module in the encoder provided in this embodiment.

[0087] Correspondingly, the execution of the RMDC module includes step S110:

[0088] S111 uses three dilated convolutional layers with different dilation rates to capture spatial information at different scales, thereby effectively extracting global and local information of the image.

[0089] It is worth emphasizing that the use of different dilation rates in the RMDC module in this embodiment is primarily to capture information at different scales, thereby better identifying and segmenting the target region, that is, determining key abdominal tissue structure information such as the external oblique muscles, internal oblique muscles, transverse abdominal muscles, and rectus abdominis muscles. Specifically, using different dilation rates can help the RMSD-Net model extract features in different receptive fields, thereby improving the expressive power and segmentation accuracy of the RMSD-Net model.

[0090] Therefore, in actual operation, it is preferred to select the optimal configuration of the expansion rate parameter to ensure the expressiveness and segmentation accuracy of the RMSD-Net model.

[0091] This embodiment provides two combinations of different expansion rates for detailed description:

[0092] Combination 1: Expansion rates are 1, 2, and 4.

[0093] When the expansion rate is 1, the receptive field is smaller, allowing for the capture of detailed information and local features. This helps identify and segment small or subtle structures, such as muscle fibers or small muscle groups. During the segmentation of EO, IO, TA, and RA, the identification of fascia is crucial. Setting the expansion rate to 1 allows for the initial capture of fascia features, laying the foundation for the subsequent use of attention mechanisms to capture muscle features.

[0094] A dilation rate of 2 creates a medium receptive field, balancing detailed information with global features. This allows for capturing larger structures and medium-scale features, facilitating the identification and segmentation of medium-sized muscle groups. Since the EO, IO, and TA muscles are clinically similar in size, a dilation rate of 2 is more effective.

[0095] When the dilation rate is 4, the receptive field is larger, capturing global information and large-scale features. This helps identify and segment large structures and overall contours, such as entire muscle groups or the boundaries between muscles and other tissues. The EO, IO, and TA muscles appear in the subcostal, anterior, posterior, and lateral approaches, but the RA only appears in the subcostal approach, and the ASIS only appears in the anterior approach. Therefore, after segmenting the EO, IO, and TA muscles, further identifying the RA and ASIS can help doctors better determine which approach to use. Each approach requires different anesthesia methods and anesthetic dosages.

[0096] In addition, there is a situation where the back and the outside are considered to be a complete scanning process by default, so these two sides can be processed together or separately.

[0097] Combination 2: Select experimental results with expansion rates of 2, 4, 8, or even larger expansion rates.

[0098] Among them, if the expansion rates of 2, 4, and 8 are used, the capture of detail information will be lost when the expansion rate is 1, resulting in insufficient small-scale feature extraction.

[0099] Moreover, when the dilation rate is set to 2, 4, or 8, the setting has a large jump in scale change, which may lead to information loss or discontinuity during feature extraction, especially for complex structures that need to span multiple scales.

[0100] Furthermore, at a larger expansion rate (such as 8, 16), the area covered by the convolution kernel is larger, which can easily lead to a hollow effect in the receptive field (i.e., the effective information area is reduced), resulting in some important features not being captured.

[0101] Finally, convolution with a larger dilation rate will increase the computational complexity of the network because the convolution kernel covers a larger area and involves more computation.

[0102] Therefore, the expansion rates of 1, 2, and 4 selected in this embodiment have significant advantages in capturing information of different scales, smooth transition, and computational efficiency.

[0103] Therefore, when the expansion rates are 1, 2, and 4 respectively, the corresponding features of different scales are represented as {X1, X2, X3}, thus obtaining X i =F i (X); where i = 1, 2, 3, F i (X) represents the convolution operation with different expansion rates, and X represents the input feature map.

[0104] S112: After passing the output features of all dilated convolutional layers through the BN layer, they are concatenated along the channel dimension, and then integrated through the voting convolution layer voteConv layer. The weight map W of the fusion feature is obtained through the Sigmoid function σ. 11 , that is, W 11 =σ(F vote ([X1, X2, X3])); where F vote () indicates a voting convolution operation, and [X1,X2,X3] indicates a feature concatenation operation.

[0105] It is worth mentioning that the output features of all dilated convolutional layers are spliced ​​along the channel dimension after passing through the BN layer, and combined through the voting convolutional layer, and then the Sigmoid activation function is used to play the role of spatial attention.

[0106] Furthermore, the voteConv layer adjusts and filters features extracted at different scales, enhancing key structural features while suppressing secondary information. This voting convolution operation is an effective feature selection and enhancement mechanism in practice, helping improve the model's understanding of complex image content.

[0107] S113, use residual connection to combine the weighted feature map with the original input feature map X to obtain the final output feature map Y, that is, Y = X + X * W 11 .

[0108] Among them, when obtaining the final output feature map Y, it is preferably activated by the ReLU function to combine the features extracted from different scales and contexts.

[0109] It is worth emphasizing that since the residual connection can connect the feature information of low and high layers through cross-layer connections to solve the problems of gradient disappearance and gradient explosion, the use of residual connections to combine the weighted feature map with the original input feature map X to obtain the final output feature map Y can better transfer features between layers, thereby improving training efficiency, model performance and generalization ability.

[0110] Therefore, the above formulas can be combined to determine that the RMDC module can be expressed as: Y = X + X * σ (F vote ([F1(X), F2(X), F3(X)]))

[0111] Among them, F1(X), F2(X) and F3(X) represent the convolution operations of the dilated convolution layer with dilation rates of 1, 2 and 4 respectively. vote () is the voting convolution layer operation, σ is the Sigmoid activation function, X is the input feature map, and Y is the final output feature map obtained after processing by the RMDC module. In this embodiment, the output feature map Y of each layer encoder corresponds to the kth layer marked as the output feature map E k .

[0112] Based on this, the RMDC module can extract feature information of different scales by performing convolution operations on the input feature map under different receptive fields, which enables the model to capture image features of different levels and scales, which helps to improve the model's ability to understand complex scenes.

[0113] To better transfer features between layers, preferably, in the skip connection structure, the SFF module dynamically adjusts the channel weights of the feature map based on the global information of the input feature map, thereby enhancing the network's focus on important features and suppressing responses to minor features. Therefore, the feature fusion module SFF module is used to replace the traditional U-Net skip connection structure, aiming to enhance feature expression and promote the integration of contextual information.

[0114] For details, see Figure 4 FIG. 1 is a schematic structural diagram of an SFF module in a jump connection structure provided in this embodiment.

[0115] In the skip connection structure, the SFF module converts the output feature map E of the encoder corresponding to each layer into k After 3*3 convolution, the input feature map SFF corresponding to the SFF module k+1 After upsampling and element-by-element addition, features are extracted and output through the multi-scale depth-separable hybrid attention mechanism module MDSA module, where k is a positive integer.

[0116] The overall structure of the MDSA module includes channel attention and spatial attention; the spatial attention is composed of a multi-scale depthwise separable convolution module (Multi-Scale Depthwise Separable Convolution Module) to dynamically allocate attention weights in the channel dimension and spatial dimension.

[0117] From this, it can be determined that the multi-scale depth-separable hybrid attention mechanism module MDSA module can extract more fine-grained features and output them.

[0118] As one of the preferred implementations of this embodiment, the execution of the MDSA module includes step S120:

[0119] S121, obtain the input feature map, and obtain the spatial information of the input feature map after channel attention aggregation average pooling and maximum pooling operations.

[0120] S122, the aforementioned spatial information is processed through a shared multi-layer perceptron (MLP) to generate a channel attention map; wherein, the output feature map after processing by the channel attention mechanism is obtained by element-wise multiplication with the input feature.

[0121] S123, input the channel prior into the deep convolution module to generate a spatial attention map, receive the spatial attention feature map through a 1*1 convolution block and perform channel fusion.

[0122] S124, multiply the channel mixing result by the channel prior element by element to obtain the refined features as output.

[0123] As another preferred implementation of this embodiment, specifically combined with Figure 4 , describes the attention mechanism of the MDSA module in detail. In this embodiment, the input feature map is F∈R C*H*W , where C is the number of channels, H and W are the height and width of the feature map respectively, and R represents a set of real numbers to indicate that all elements of the input feature map F are real numbers.

[0124] First, use the Channel Attention Module (CAM) to calculate a 1D channel attention map M c ∈R C*1*1 , then M c Multiply the input feature map F element by element to obtain the refined feature map F with channel attention c ∈R C*H*W . F c Represents the output feature map after the channel attention mechanism CA.

[0125] Secondly, the Spatial Attention Module (SAM) is used to process F c Generate 3D spatial attention map M s ∈R C*H*W .

[0126] Finally, by element-by-element s With F c Multiply to get the final output feature map

[0127] Based on this, the entire attention process can be summarized as: and

[0128] It is worth noting that the channel attention map is generated by the channel attention module, which explores the inter-channel relationships in the features to generate the channel attention map. Specifically, it uses average pooling and max pooling operations to aggregate spatial information from the feature map, which generates two independent spatial context descriptors.

[0129] The spatial context descriptors are then fed into a shared multi-layer perceptron (MLP), and the outputs of the shared MLP are combined via element-wise summation to obtain a channel attention map.

[0130] In order to reduce parameter overhead, it is preferred in this embodiment that the shared MLP contains only one hidden layer, and the size of the hidden layer activation is set to R C / r*1*1 , where r represents the dimensionality reduction ratio.

[0131] To this end, the calculation of channel attention is expressed as:

[0132] CA(F)=σ(MLP(F avg )+MLP(F max )).

[0133] The MLP() represents processing the input vector through one or more fully connected layers (which contain nonlinear activation functions) and outputting a vector of the same size; avg represents the feature map after average pooling, F max Represents the feature map after maximum pooling.

[0134] The spatial attention map is generated by extracting the spatial mapping relationship. Based on practical considerations, in this embodiment, it is considered that the spatial attention map of each channel should not be forced to remain consistent. Instead, it is considered more realistic to dynamically allocate attention weights across channels and spatial dimensions.

[0135] Correspondingly, Figure 4 We demonstrate the use of depthwise convolution to capture spatial relationships between features, ensuring that computational complexity is reduced while preserving inter-channel relationships.

[0136] In order to enhance the ability of convolution operation to capture spatial relationships, this embodiment preferably adopts a multi-scale structure. That is, at the end of the spatial attention module, a more refined attention map is generated by using 1×1 convolution to mix channels; where DwConv represents depthwise separable convolution, Branch i represents the i-th branch, i∈{1,2,3};

[0137] The calculation of the spatial attention is expressed as:

[0138]

[0139] Therefore, it is worth emphasizing that the MDSA attention mechanism in the SFF module selects and weights feature information from different channels by learning the importance weight of each channel in the feature map. This helps the model focus on the most relevant and discriminative features, improving the expressiveness of features. At the same time, this attention mechanism can reduce the influence of irrelevant or unimportant feature channels by performing a weighted sum on the feature map, thereby reducing the computational complexity of the model.

[0140] In the decoder structure, combined with Figure 2As shown in the figure, the output features obtained by downsampling the last layer of the encoder and the output features obtained by the SFF module in the corresponding layer are used as the input features of the first upsampling operation in the decoder structure. The feature maps provided by the input are concatenated along the specified dimension through element-wise concatenation, and the corresponding feature maps are obtained after upsampling. Each subsequent upsampling operation uses the output features obtained by the SFF module in the corresponding layer and the output features of the previous upsampling operation as input information. The feature maps provided by the input are concatenated along the specified dimension through element-wise concatenation, and the corresponding feature maps are obtained through upsampling. After the last upsampling operation, the feature maps obtained after upsampling of all layers in the decoder structure are combined. The attention mechanism GAM module integrates the local and global information in the feature maps, thereby dynamically adjusting the weights and providing a richer feature representation for subsequent feature weighting. Finally, a 1*1 convolution is performed to obtain the final output feature map.

[0141] Specifically, a GAM module is added to the decoder. The GAM module performs channel attention processing on the input based on the combination of channel attention and spatial attention, and then performs spatial attention processing on the result of channel attention.

[0142] Preferably, the execution of the GAM module includes step S130:

[0143] S131, process the input by the channel attention module.

[0144] In the GAM module, the channel attention module first reshapes the input tensor to H*W*C through a size permute operation, that is, moving the channel dimension to the end to facilitate global feature calculation in the spatial dimension. The reshaped tensor is then processed by a multilayer perceptron (MLP) to generate global features. The MLP consists of two fully connected layers (FC layers) that use a reduction ratio r to reduce computational complexity. The first FC layer reduces the number of channels from C to C / r, and the second FC layer restores the number of channels from C / r to C. The processed tensor is restored to its original shape C*H*W through a size reverse operation for further operations. Finally, the channel attention weights are generated through a sigmoid activation function.

[0145] S132, the results of the processing of the aforementioned channel attention module are processed by the spatial attention module to obtain an output feature map.

[0146] The execution of the spatial attention module includes:

[0147] S1321, uses two convolutional layers to perform spatial information fusion.

[0148] S1322, use the same reduction ratio r in the spatial attention module as in the channel attention module.

[0149] S1323, use the Sigmoid activation function to weight the feature map at each position to improve the spatial expression ability of the feature map and reduce the influence of secondary information.

[0150] As another preferred implementation of this embodiment, specifically combined with Figure 2 and Figure 5 As shown, for the input feature map X∈R C*H*W , where C is the number of channels, H and W are the height and width of the feature map respectively.

[0151] First, the four upsampled features in the decoder are concatenated and output to the GAM module, outputting F cat ∈R 5C*H*W , that is, F cat =[X1,X2,X3,X4,X5],F cat Indicates that all output feature maps are concatenated.

[0152] For the channel attention module, assuming the input feature map F1∈R C*H*W , where C is the number of channels, H and W are the height and width of the feature map respectively.

[0153] First, perform global average pooling and global maximum pooling operations to generate two feature vectors: F avg ∈R C*1*1 , F max ∈R C*1*1 .Right now:

[0154] and

[0155] Among them, F avg represents the feature map after global average pooling; F1c,i,j represents the feature value of position (i,j) on channel c; F max Represents the feature map after global maximum pooling.

[0156] The two pooled feature vectors are input into a shared multi-layer perceptron (MLP) to generate two new feature vectors F′ avg ,F′ max , to represent the feature map after global average pooling and the feature map after global maximum pooling in the multi-layer perceptron, namely:

[0157] F′ avg =σ(W1(W0(F avg ))) and F′ max =σ(W1(W0(F max ))).

[0158] Among them, W1 and W0 are shared weight matrices, and σ is the activation function.

[0159] The final channel attention map M is obtained by element-wise summing of the two feature vectors c , that is, M c =σ(F′ avg +F′ max ).

[0160] Finally, the channel attention map is multiplied element-wise with the original feature map, and F2 is the output, that is,

[0161] To focus on spatial information, the spatial attention module first uses two convolutional layers to fuse spatial information. Secondly, the channel attention module uses the same reduction ratio r as the proposed spatial attention module. Furthermore, since the max pooling operation negatively impacts information utilization, it is removed to further preserve the feature map. Finally, a sigmoid activation function is used to weight the feature map at each position, thereby improving the spatial expressiveness of the feature map and reducing the influence of irrelevant information.

[0162] Therefore, for the spatial attention module, assuming that the input feature map F2∈R C*H*W , where C is the number of channels, H and W are the height and width of the feature map respectively. First, global average pooling and global maximum pooling operations are performed to generate two feature vectors: F avg ∈R 1*H*W , F max ∈R 1*H*W .

[0163] in, and

[0164] The above two feature maps are spliced ​​in the channel dimension to generate a new feature map, F cat ∈R 2*H*W+ , that is, F cat =[F avg ; F max ].

[0165] The concatenated feature map undergoes two 7×7 convolution operations to generate the final spatial attention map M. s , that is: M s =σ(Conv(F cat)). Where Conv is a convolution operation, preferably using a 1×1 convolution kernel, and σ is an activation function.

[0166] Finally, the spatial attention map is element-wise multiplied with the input feature map. F3 is the output, which is expressed as:

[0167] It can be seen that through the processing sequence of the above step S130, the GAM module can model the correlation between channels and spaces in the feature map on a global scale, thereby better capturing the global dependencies of the feature map, and overall improving the model's understanding and expression capabilities of the input features. At the same time, it also reduces the loss of features during the upsampling process.

[0168] S103 , performing a segmentation task for the aforementioned key abdominal tissue structure information, dividing the aforementioned abdominal ultrasound image dataset into a training set and a test set, and using the aforementioned RMSD-Net model for training and testing, respectively.

[0169] In this embodiment, 80% of the ultrasound images with the transverse abdominal muscle plane structure are extracted from the abdominal ultrasound image dataset to form a training set for training the aforementioned RMSD-Net model, and the remaining 20% ​​of the ultrasound images with the transverse abdominal muscle plane structure are composed of a test set for verifying the aforementioned RMSD-Net model.

[0170] After testing, the RMSD-Net model can use mIOU, mPA, Accuracy and Dice coefficient as evaluation indicators to measure model performance in the segmentation task to verify the final model effect.

[0171] Among them, the mIOU is an indicator for evaluating segmentation accuracy; the mPA refers to the average pixel accuracy in each category; the Accuracy refers to the ratio of correctly classified pixels to the total number of pixels in the segmentation result; and the Dice coefficient is an indicator for evaluating the degree of overlap between the segmentation result and the true annotation.

[0172] Specifically, in the process of segmenting the abdominal tissue structure in the aforementioned ultrasound image, the evaluation indicators mIOU, mPA, Accuracy and Dice coefficient can be used to measure the performance of the segmentation algorithm.

[0173] The mIOU (Mean Intersection over Union) is a metric used to evaluate segmentation accuracy and is defined as the ratio of the intersection over union of the predicted and true regions. The IOU is calculated for each category and then averaged across all categories.

[0174] The mIOU measures the segmentation accuracy of the model on different abdominal tissue structures. A high mIOU indicates that the model performs well in identifying and segmenting these structures.

[0175] The mPA (Mean Pixel Accuracy) is the average pixel accuracy for each category. It calculates the ratio of correctly classified pixels in each category to the total number of pixels in that category, and then averages it across all categories.

[0176] The mPA can reflect the classification accuracy of each tissue structure (such as liver, kidney, etc.) and help evaluate the performance balance of the model on different tissues.

[0177] The accuracy (pixel accuracy) refers to the ratio of correctly classified pixels to the total number of pixels in the segmentation result.

[0178] The Accuracy measures the segmentation accuracy of all pixels in the entire image. However, since it is insensitive to category imbalance (ie, background pixels in a large area may result in a high Accuracy value), it is usually used together with other indicators in ultrasound image segmentation.

[0179] The Dice coefficient is a commonly used metric for evaluating the degree of overlap between segmentation results and true annotations. It is defined as twice the intersection of the predicted area and the true area divided by their sum.

[0180] Application in abdominal tissue structure: The Dice coefficient is particularly suitable for evaluating ultrasound image segmentation because it is highly sensitive to overlapping areas and can better reflect the segmentation performance of the model on small-scale tissues (such as small tumors or organ boundaries).

[0181] The above indicators are of great significance in the image segmentation of abdominal tissue structures. They can comprehensively evaluate the performance of the segmentation model from different angles and help optimize and improve the model.

[0182] S104, using the tested RMSD-Net model to segment and identify key abdominal tissue structure information of the newly acquired ultrasound image with the transverse abdominal muscle plane structure.

[0183] It is worth emphasizing here that the combination of TAP nerve block can realize the selection of four-side approach schemes from the subcostal margin, anterior side, posterior side, and lateral side. Therefore, after the RMSD-Net model performs segmentation and recognition of key abdominal tissue structure information on the transverse abdominal muscle plane tissue structure image, the implementation plan of TAP nerve block can be determined at the same time.

[0184] By way of example and not limitation, the EO, IO, and TA muscles will appear during the subcostal, anterior, posterior, and lateral approaches, but the RA will only appear during the subcostal approach, and the ASIS will only appear during the anterior approach.

[0185] This means that the RMSD-Net model segments and identifies key abdominal tissue structure information in the transverse abdominal muscle plane tissue structure image. If RA is identified, the muscle group belongs to the subcostal approach; if ASIS is identified, the muscle group belongs to the anterior approach.

[0186] Therefore, after segmenting and identifying the EO, IO, and TA muscles, the RA and ASIS can be further identified, which helps guide doctors to determine which muscle to approach. At the same time, combined with the anesthesia method and dosage of each approach, it lays the foundation for subsequent ultrasound guidance of each measurement approach.

[0187] See also Figure 6 The above is the visualization result of the ablation experiment provided by this embodiment. At the same time, Table 1 provides the numerical results of the evaluation indicators under different models in the ablation experiment.

[0188] Table 1 Numerical results of evaluation indicators under different models in ablation experiments

[0189]

[0190] Combine Figure 6 As shown in the figure, in the ablation experiment, the M, C, and G modules represent the RMDC module, the SFF module, and the GAM module, respectively. The baseline is the benchmark model U-Net, and the RMSD-Net is Baseline+RMDC+SFF+GAM.

[0191] Combine Figure 6 As shown in Table 1, the RMDC module, SFF module, and GAM module have a positive effect on the feature extraction and prediction capabilities of the model, and the numerical results of each evaluation index have been improved. Among them, the Dice coefficient reached 93.55%, 94.98%, and 95.45% respectively, and the IOU coefficient reached 88.12%, 90.57%, and 91.38% respectively. Although there is a great improvement in the numerical results, combined with Figure 6 It can be determined that the visual display effect of the test set is very poor, and the slightly blurred muscle images are not fully recognized or cannot be recognized at all (for example Figure 6 The visualization effects corresponding to the 2nd, 3rd, and 4th columns of image d).

[0192] Depend on Figure 6As can be seen from Table 1, the Dice score of Baseline+RMDC+GAM reached 94.51%, and the IOU reached 89.74%. The Dice score of Baseline+RMDC+SFF reached 95.01%, and the IOU reached 90.77%. The Dice score of Baseline+SFF+GAM reached 95.40%, and the IOU reached 91.29%. Therefore, it can be determined that the combination of RMDC module, SFF module, and GAM module added to the U-Net model is improved compared with the U-Net model adding modules separately. Among them, the combination of Baseline+SFF+GAM achieved the highest score. However, the visualization results show that although all muscles can be recognized, some noise will be predicted incorrectly (for example, Figure 6 The visualization effect corresponding to the penultimate column of image d in is shown in Figure 3), and the edge prediction of the fascia is also incomplete (e.g. Figure 6 The visualization effects corresponding to columns 5, 6, and 7 of image d).

[0193] Depend on Figure 6 As can be seen in Table 1, the baseline + RMDC + SFF + GAM (RMSD-Net) achieved the highest recognition results, with a Dice score of 95.70% and an IOU of 91.83%. The prediction results also largely resolved issues such as misidentification, missed identification, and noise. This is because the RMSD-Net model incorporates these three improvements, enabling it to extract and fuse features at different scales and in global context, improving the network's segmentation accuracy and robustness for complex backgrounds and low-contrast areas.

[0194] See also Figure 7 As shown, a comparative experiment is provided for this embodiment to demonstrate the visual comparison effect of the RMSD-Net model and different segmentation models.

[0195] The comparative experiment is a key step in evaluating the segmentation ability between the RMSD-Net network model and the current mainstream models, and can also verify the generalization performance of the model.

[0196] Tables 2 and 3 respectively show the evaluation values ​​given by the comparative experiment for the evaluation coefficient, as well as the specific numerical results of each part in the comparative experiment.

[0197] Table 2 shows the numerical results of the comparative experiment

[0198]

[0199] Table 3 Specific numerical results of each part in the comparative experiment

[0200]

[0201] Depend on Figure 7 As can be seen, segmentation models such as U-Net, U-Net++, and Attention U-Net perform poorly for transverse abdominal plane musculature segmentation, primarily due to missed recognition of the RA or ASIS. Even excluding these two areas and focusing solely on the EO, IO, and TA, Table 3 shows that these models perform relatively poorly compared to the RMSD-Net model, resulting in inferior predictions. While models such as CMU-Net and MACU-Net achieve similar recognition accuracy to RMSD-Net, they significantly increase the computational complexity of these models.

[0202] Based on this, it is emphasized again that the advantages of the RMSD-Net model provided in this embodiment are:

[0203] First, a multi-scale residual feature extraction (RMDC) module is added to each layer of the encoder, allowing convolution kernels with different dilation rates to capture information at different scales while suppressing less important information. The external oblique, internal oblique, transverse abdominal, and rectus abdominis muscles differ in size and morphology, and multi-scale dilated convolutions can adaptively process these varying scales.

[0204] Secondly, the skip connection structure uses the feature fusion module SFF module layer by layer to enhance feature expression and promote the integration of contextual information. The MDSA attention mechanism in the module combines channel attention and spatial attention, and dynamically allocates attention weights through a multi-scale depth-separable convolution module, while retaining channel priors. This combination can help the network better capture important features and improve the feature representation ability.

[0205] Third, an attention mechanism (GAM) module is added to the last layer of the decoder to provide feature representation for subsequent feature weighting. In ultrasound images, the background is complex and noisy. The GAM attention mechanism can help the model focus on the target area and effectively filter out background noise and interference, allowing the model to capture the details of different muscle boundaries more finely, which is very important for the identification of complex structures and interconnected muscle parts.

[0206] In addition, the RMSD-Net model preserves the original 3*3 convolution, which helps to extract local features in the image, which is crucial for distinguishing different muscle tissues (which have large differences).

[0207] Other technical features are described in the previous embodiments and will not be repeated here.

[0208] In addition, the present invention also provides an embodiment, which provides a transverse abdominal muscle plane tissue structure recognition device 200 based on RMSD-Net, comprising:

[0209] A data construction unit 201 is used to perform data preprocessing on the acquired ultrasound image having the transverse abdominal muscle plane structure to establish an abdominal ultrasound image dataset; wherein the ultrasound image having the transverse abdominal muscle plane structure includes key abdominal tissue structure information that can be observed during the TAP block process; the key abdominal tissue structure information includes the external oblique muscle, internal oblique muscle, transverse abdominal muscle, rectus abdominis muscle and anterior superior iliac spine.

[0210] The model construction unit 202 is used to construct an RMSD-Net model with a U-shaped network structure; the RMSD-Net model includes an encoder, a skip connection structure and a decoder; the encoder uses the original double convolution structure in each layer, and adds a multi-scale residual feature extraction module RMDC module in each layer to perform feature extraction on the input ultrasound image with the transverse abdominal muscle plane structure layer by layer, thereby strengthening the important features of the aforementioned input ultrasound image with the transverse abdominal muscle plane structure layer by layer, and suppressing secondary information at the same time; the skip connection structure uses the feature fusion module SFF module layer by layer to enhance feature expression and promote the integration of context information; the attention mechanism GAM module is added to the last layer of the decoder to combine channel attention and spatial attention to integrate local information and global information in the feature map, and dynamically adjust the weights to provide feature representation for subsequent feature weighting.

[0211] The model training and testing unit 203 is used to perform a segmentation task for the aforementioned key abdominal tissue structure information, divide the aforementioned abdominal ultrasound image dataset into a training set and a test set, and use the aforementioned RMSD-Net model for training and testing respectively.

[0212] The image recognition unit 204 is configured to segment and recognize key abdominal tissue structure information of a newly acquired ultrasound image having a transverse abdominal muscle plane structure using the tested RMSD-Net model.

[0213] For other technical features, please refer to the previous embodiments and will not be repeated here.

[0214] In addition, the present invention also provides an embodiment, which provides a transverse abdominal muscle plane tissue structure recognition system 200 based on RMSD-Net, including:

[0215] The network node 301 is configured to transmit and receive ultrasound images of the transverse abdominal muscle plane structure acquired by ultrasound equipment.

[0216] The model arrangement module 302 is used to configure an RMSD-Net model for segmenting and identifying key abdominal tissue structure information in an ultrasound image having a transverse abdominal muscle plane tissue structure; the key abdominal tissue structure information includes the external oblique muscle, internal oblique muscle, transverse abdominal muscle, rectus abdominis muscle and anterior superior iliac spine.

[0217] The system server 303 is used to connect the network node 301 and the model arrangement module 302 .

[0218] The system server 303 is configured to: perform data preprocessing on the acquired ultrasound image with the transverse abdominal muscle plane structure to establish an abdominal ultrasound image data set; wherein, the ultrasound image with the transverse abdominal muscle plane structure includes key abdominal tissue structure information that can be observed during the TAP block; construct an RMSD-Net model with a U-shaped network structure; the RMSD-Net model includes an encoder, a jump connection structure and a decoder; the encoder uses the original double convolution structure in each layer and adds a multi-scale residual feature extraction module RMDC module in each layer to perform feature extraction on the input ultrasound image with the transverse abdominal muscle plane structure layer by layer, thereby strengthening the aforementioned input ultrasound image with the transverse abdominal muscle plane structure layer by layer. The invention discloses a novel method for segmenting the key abdominal tissue structure information of the present invention, which can identify the important features in the image and suppress the secondary information at the same time; the skip connection structure uses the feature fusion module SFF module layer by layer to enhance the feature expression and promote the integration of context information; the attention mechanism GAM module is added to the last layer of the decoder to combine channel attention and spatial attention to integrate the local information and global information in the feature map, and dynamically adjust the weights to provide feature representation for subsequent feature weighting; the segmentation task is performed on the aforementioned key abdominal tissue structure information, the aforementioned abdominal ultrasound image dataset is divided into a training set and a test set, and the aforementioned RMSD-Net model is used for training and testing respectively; the tested RMSD-Net model is used to perform segmentation and recognition of the key abdominal tissue structure information on the newly acquired ultrasound image with the transverse abdominal muscle plane structure.

[0219] For other technical features, please refer to the previous embodiments and will not be repeated here.

[0220] In addition, an embodiment of the present invention also provides a computer-readable storage medium on which a program is stored for use in the aforementioned RMSD-Net-based transverse abdominal muscle plane tissue structure identification system. When the program is executed by the processor, it can implement any of the steps of the RMSD-Net-based transverse abdominal muscle plane tissue structure identification method described above.

[0221] For other technical features, please refer to the previous embodiments and will not be repeated here.

[0222] In the above description, the components may be selectively and operatively combined in any number within the scope of the intended protection of the present disclosure. In addition, terms such as "include," "encompass," and "have" should be interpreted as inclusive or open-ended rather than exclusive or closed by default, unless expressly defined to the contrary. All technical, technological, or other terms have the meanings understood by those skilled in the art, unless they are defined to the contrary. Common terms found in dictionaries should not be interpreted in an overly idealized or unrealistic manner in the context of the relevant technical documentation, unless expressly defined to that extent by the present disclosure.

[0223] Although example aspects of the present disclosure have been described for illustrative purposes, those skilled in the art will appreciate that the foregoing description is merely a description of preferred embodiments of the present invention and does not limit the scope of the present invention in any way. The scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order in which they appear or are discussed. Any changes or modifications made by those skilled in the art based on the foregoing disclosure are intended to fall within the scope of the claims.

Claims

1. A method for identifying the transverse abdominal muscle plane tissue structure based on RMSD-Net, characterized in that: Specifically include: Data preprocessing is performed on the acquired ultrasound images with the transverse abdominal muscle plane structure to establish an abdominal ultrasound image dataset; wherein the ultrasound images with the transverse abdominal muscle plane structure include key abdominal tissue structure information that can be observed during TAP block; the key abdominal tissue structure information includes the external oblique muscle, internal oblique muscle, transverse abdominal muscle, rectus abdominis muscle, and anterior superior iliac spine; Construct an RMSD-Net model with a U-shaped network structure; the RMSD-Net model includes an encoder, a jump connection structure and a decoder; the encoder uses the original double convolution structure in each layer, and adds a multi-scale residual feature extraction module RMDC module in each layer to extract features of the input ultrasound image with the transverse abdominal muscle plane structure layer by layer, thereby strengthening the important features of the aforementioned input ultrasound image with the transverse abdominal muscle plane structure layer by layer, and suppressing secondary information at the same time; the jump connection structure uses the feature fusion module SFF module layer by layer to enhance feature expression and promote the integration of contextual information; an attention mechanism GAM module is added to the last layer of the decoder to combine channel attention and spatial attention to integrate local information and global information in the feature map, and dynamically adjust the weights to provide feature representation for subsequent feature weighting; wherein, the execution of the GAM module includes steps S130: S131, using the channel attention module to process the input; the channel attention module first uses size permute The operation adjusts the shape of the input tensor to H*W*C, that is, moves the channel dimension to the end, so that global feature calculation can be performed in the spatial dimension; the adjusted tensor is then processed by a multi-layer perceptron MLP to generate global features; the MLP consists of two fully connected layers, and the reduction ratio r is used to reduce the computational complexity so that the first fully connected layer reduces the number of channels from C to C / r, and the second fully connected layer restores the number of channels from C / r to C; the size reverse operation is used to restore the processed tensor dimension to the original shape C*H*W for further operation, and finally the Sigmoid The activation function generates a channel attention weight; S132, the result of the processing of the aforementioned channel attention module is processed by the spatial attention module to obtain an output feature map; the execution of the spatial attention module includes: S1321, using two convolutional layers to perform spatial information fusion, S1322, using the same reduction ratio r as the channel attention module in the spatial attention module, S1323, using the Sigmoid activation function to weight the feature map of each position to improve the spatial expression ability of the feature map and reduce the influence of secondary information; Perform segmentation tasks on the aforementioned key abdominal tissue structure information, divide the aforementioned abdominal ultrasound image dataset into a training set and a test set, and use the aforementioned RMSD-Net model for training and testing respectively; The tested RMSD-Net model is used to segment and identify key abdominal tissue structure information in newly acquired ultrasound images with transverse abdominal muscle plane structure.

2. The method according to claim 1, characterized in that The ultrasound image with the transverse abdominal muscle plane structure is a frame image in the ultrasound image video that can observe the corresponding muscle tissue and is recorded during the transverse abdominal muscle plane scanning process; The data preprocessing includes data labeling and data enhancement; The data annotation can use Labelme annotation software to annotate the key abdominal tissue structure information in the aforementioned ultrasound image by drawing contour lines; The data enhancement includes processing the ultrasound image having the transverse abdominal muscle plane structure by random rotation, random flipping and / or random mirroring; 80% of the ultrasound images with the transverse abdominal muscle plane structure were extracted from the abdominal ultrasound image dataset to form a training set for training the aforementioned RMSD-Net model, and the remaining 20% ​​of the ultrasound images with the transverse abdominal muscle plane structure were extracted to form a test set for verifying the aforementioned RMSD-Net model.

3. The method according to claim 1, characterized in that In the encoder, the execution of the RMDC module includes step S110: S111, using three dilated convolutional layers with different dilation rates to capture spatial information at different scales, thereby effectively extracting global and local information of the image; when the dilation rates are 1, 2, and 4, the corresponding feature representations of different scales are {X1, X2, X3}, thus obtaining X i =F i (X); where i = 1, 2, or 3, F i (X) represents the convolution operation with different expansion rates, and X represents the input feature map; S112: After passing the output features of all dilated convolutional layers through the BN layer, they are concatenated along the channel dimension, and then integrated through the voting convolution layer voteConv layer. The weight map W of the fusion feature is obtained through the Sigmoid function σ. 11 , that is, W 11 =σ(F vote ([X1,X2,X3])); where F vote () indicates the voting convolution layer operation, and [X1, X2, X3] indicates the feature concatenation operation; the voteConv layer can strengthen key structural features and suppress secondary information by adjusting and filtering features extracted at different scales; S113, use residual connection to combine the weighted feature map with the original input feature map X to obtain the final output feature map Y, that is, Y=X+X*W 11 .

4. The method according to claim 3, characterized in that The output feature map Y of each layer encoder corresponds to the kth layer and is marked as the output feature map E k In the jump connection structure, the SFF module converts the output feature map E of the encoder corresponding to each layer k After 3*3 convolution, the input feature map SFF corresponding to the SFF module k+1 After upsampling and element-by-element addition, features are extracted and output through the multi-scale depth-separable hybrid attention mechanism module MDSA module, where k is a positive integer. The overall structure of the MDSA module includes channel attention and spatial attention. The spatial attention is composed of a multi-scale depth-separable convolution module to dynamically allocate attention weights in the channel dimension and the spatial dimension.

5. The method according to claim 4, characterized in that The execution of the MDSA module includes step S120: S121, obtain the input feature map, and obtain the spatial information of the input feature map after channel attention aggregation average pooling and maximum pooling operations; S122, processing the aforementioned spatial information through a shared multi-layer perceptron (MLP) to generate a channel attention map; wherein the output feature map after the channel attention mechanism is obtained by element-wise multiplication with the input feature map; S123, input the channel prior into the deep convolution module to generate a spatial attention map, receive the spatial attention feature map through a 1*1 convolution block and perform channel fusion; S124, multiply the channel mixing result by the channel prior element by element to obtain the refined features as output.

6. The method according to claim 1, characterized in that After the above tests, mIOU, mPA, Accuracy, and Dice coefficient were used as evaluation indicators to measure the performance of the RMSD-Net model in the segmentation task to verify the final model effect; Among them, the mIOU is an indicator for evaluating segmentation accuracy; The mPA refers to the average pixel accuracy of each category; Accuracy refers to the ratio of correctly classified pixels to the total number of pixels in the segmentation result; The Dice coefficient is an indicator for evaluating the degree of overlap between the segmentation result and the true annotation.

7. A device for identifying the planar tissue structure of the transverse abdominal muscle based on RMSD-Net according to any one of claims 1 to 6, characterized in that include: a data construction unit, configured to perform data preprocessing on an acquired ultrasound image having a transverse abdominal muscle plane structure to establish an abdominal ultrasound image dataset; wherein the ultrasound image having a transverse abdominal muscle plane structure includes key abdominal tissue structure information observable during TAP block; the key abdominal tissue structure information includes the external oblique muscle, internal oblique muscle, transverse abdominal muscle, rectus abdominis muscle, and anterior superior iliac spine; A model construction unit is used to construct an RMSD-Net model with a U-shaped network structure; the RMSD-Net model includes an encoder, a jump connection structure and a decoder; the encoder uses the original double convolution structure in each layer, and adds a multi-scale residual feature extraction module RMDC module in each layer to extract features of the input ultrasound image with the transverse abdominal muscle plane structure layer by layer, thereby strengthening the important features of the aforementioned input ultrasound image with the transverse abdominal muscle plane structure layer by layer, and suppressing secondary information at the same time; the jump connection structure uses the feature fusion module SFF module layer by layer to enhance feature expression and promote the integration of context information; an attention mechanism GAM module is added to the last layer of the decoder to combine channel attention and spatial attention to integrate local information and global information in the feature map, and dynamically adjust the weights to provide feature representation for subsequent feature weighting; wherein, the execution of the GAM module includes steps S130: S131, processing the input using the channel attention module; the channel attention module first uses size permute The operation adjusts the shape of the input tensor to H*W*C, that is, moves the channel dimension to the end, so that global feature calculation can be performed in the spatial dimension; the adjusted tensor is then processed by a multi-layer perceptron MLP to generate global features; the MLP consists of two fully connected layers, and the reduction ratio r is used to reduce the computational complexity so that the first fully connected layer reduces the number of channels from C to C / r, and the second fully connected layer restores the number of channels from C / r to C; the size reverse operation is used to restore the processed tensor dimension to the original shape C*H*W for further operation, and finally the Sigmoid The activation function generates a channel attention weight; S132, the result of the processing of the aforementioned channel attention module is processed by the spatial attention module to obtain an output feature map; the execution of the spatial attention module includes: S1321, using two convolutional layers to perform spatial information fusion, S1322, using the same reduction ratio r as the channel attention module in the spatial attention module, S1323, using the Sigmoid activation function to weight the feature map of each position to improve the spatial expression ability of the feature map and reduce the influence of secondary information; A model training and testing unit, configured to perform segmentation tasks for the aforementioned key abdominal tissue structure information, divide the aforementioned abdominal ultrasound image dataset into a training set and a test set, and use the aforementioned RMSD-Net model for training and testing, respectively; The image recognition unit is used to use the tested RMSD-Net model to segment and identify key abdominal tissue structure information of the newly acquired ultrasound image with the transverse abdominal muscle plane structure.

8. A transverse abdominal muscle plane tissue structure recognition system based on RMSD-Net according to any one of claims 1 to 6, characterized in that include: A network node, used for transmitting and receiving an ultrasound image of a transverse abdominal muscle plane structure acquired by an ultrasound device; A model configuration module is used to configure an RMSD-Net model for segmenting and identifying key abdominal tissue structure information in an ultrasound image having a transverse abdominal muscle planar tissue structure; the key abdominal tissue structure information includes the external abdominal oblique muscle, internal abdominal oblique muscle, transverse abdominal muscle, rectus abdominis muscle, and anterior superior iliac spine; A system server, the system server connecting the network nodes and the model arrangement module; The system server is configured to: perform data preprocessing on the acquired ultrasound images with the transverse abdominal muscle plane structure to establish an abdominal ultrasound image dataset; wherein, the ultrasound images with the transverse abdominal muscle plane structure include key abdominal tissue structure information that can be observed during TAP block; construct an RMSD-Net model with a U-shaped network structure; the RMSD-Net model includes an encoder, a jump connection structure and a decoder; the encoder uses the original double convolution structure in each layer and adds a multi-scale residual feature extraction module RMDC module in each layer to perform feature extraction on the input ultrasound image with the transverse abdominal muscle plane structure layer by layer, thereby extracting features layer by layer. Strengthen the important features of the aforementioned input ultrasound image with the transverse abdominal muscle plane structure, and at the same time suppress secondary information; the skip connection structure uses the feature fusion module SFF module layer by layer to enhance feature expression and promote the integration of context information; add an attention mechanism GAM module in the last layer of the decoder to combine channel attention and spatial attention to integrate local information and global information in the feature map, and dynamically adjust the weights to provide feature representation for subsequent feature weighting; wherein, the execution of the GAM module includes steps S130: S131, using the channel attention module to process the input; the channel attention module first uses size The permute operation adjusts the shape of the input tensor to H*W*C, that is, the channel dimension is moved to the end, so that global feature calculation can be performed in the spatial dimension; the adjusted tensor is then processed by a multi-layer perceptron MLP to generate global features; the MLP consists of two fully connected layers, and the reduction ratio r is used to reduce the computational complexity, so that the first fully connected layer reduces the number of channels from C to C / r, and the second fully connected layer restores the number of channels from C / r to C; the sizereverse operation is used to restore the processed tensor dimension to the original shape C*H*W for further operation, and finally the Sigmoid The activation function generates a channel attention weight; S132, the result of the processing of the aforementioned channel attention module is processed by the spatial attention module to obtain an output feature map; the execution of the spatial attention module includes: S1321, using two convolutional layers to perform spatial information fusion, S1322, using the same reduction ratio r as the channel attention module in the spatial attention module, S1323, using the Sigmoid activation function to perform weighted processing on the feature map of each position to improve the spatial expression ability of the feature map and reduce the influence of secondary information; performing a segmentation task for the aforementioned key abdominal tissue structure information, dividing the aforementioned abdominal ultrasound image dataset into a training set and a test set, and using the aforementioned RMSD-Net model for training and testing respectively; using the tested RMSD-Net model to perform segmentation and recognition of key abdominal tissue structure information on the newly acquired ultrasound image with the transverse abdominal muscle plane structure.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps of any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Semi-supervised lung image segmentation method based on improved U-Net

    CN118097128A