Medical Image Segmentation Method Based on Neural Network with Reverse Axial Attention Mechanism
Through the reverse axial attention mechanism neural network, combined with the packet residual backbone network and the cascaded hollow convolution network, the foreground objects are gradually erased, solving the problem of inaccurate segmentation of existing medical image segmentation methods under complex backgrounds and noises, and achieving efficient and accurate polyp region identification and segmentation.
Patent Information
- Application Number
- CN202510458879.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing medical image segmentation methods are not effective in complex backgrounds and noise situations, require manual intervention, and are prone to loss of detailed information, resulting in inaccurate segmentation, especially in gastrointestinal polyp images, which is difficult to accurately capture the boundary details of tissue areas.
Using a neural network based on the reverse axial attention mechanism, primary and advanced features are extracted through the grouping residual backbone network, combined with the cascaded hollow convolution network and the reverse attention mechanism network, the foreground objects are gradually erased, combined with the axial attention mechanism network, multi-scale feature information is extracted, global feature maps are generated, manual intervention is reduced, and segmentation accuracy is improved.
It significantly improves the accuracy and robustness of medical image segmentation, and can accurately identify and segment polyp areas under complex backgrounds and noise interference, reduces manual intervention, and improves processing speed and accuracy.
Smart Images

Figure CN119992106B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image segmentation, and particularly relates to a medical image segmentation method based on a reverse axial attention mechanism neural network. Background Art
[0002] Gastrointestinal polyps are a common digestive tract disease, usually caused by chronic inflammation, and manifested as raised lesions formed by mucosal hyperplasia and hypertrophy. Although most patients have no obvious symptoms, they are often detected during endoscopic examination or imaging examination. Early detection and accurate segmentation of gastrointestinal polyps are crucial for timely treatment and improving the diagnostic efficiency.
[0003] Gastrointestinal polyp images are taken by an endoscope, and are affected by factors such as the endoscopic shooting angle, illumination, and shadow, which have a certain impact on the image resolution and clarity. The commonly used method for identifying gastrointestinal polyps is to use image segmentation technology to separate the polyps in the image from the surrounding background. Then, feature extraction technology can be used to extract features in the image, such as color, texture, and shape, etc. Next, a classifier can be trained using machine learning algorithms to match the extracted features with labels (polyp / non-polyp), so as to identify the gastrointestinal polyps in the image.
[0004] Although existing technologies and methods have made significant progress in the field of medical image segmentation, there are still some deficiencies. The designs of some advanced segmentation models are too complex, and the segmentation effect is poor in the case of complex backgrounds and noises, requiring manual intervention and adjustment, which affects the accuracy of disease condition judgment. In addition, many existing methods lose important detail information during the downsampling process, resulting in a decrease in segmentation accuracy. Moreover, most methods are limited to attention in the spatial or channel dimensions, and due to the existence of factors such as image shadows, the boundary details of tissue regions cannot be accurately captured, resulting in inaccurate segmentation of tissue regions and affecting the accurate judgment of gastrointestinal polyps. Summary of the Invention
[0005] The present invention provides a medical image segmentation method based on a reverse axial attention mechanism neural network to solve the problems of subjective differences and difficulty in disease condition judgment caused by complex image segmentation and the need for manual intervention and adjustment, as well as inaccurate tissue region segmentation caused by the loss of detail information during the sampling process of existing methods and the inability to accurately capture the boundary details of tissue regions.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A medical image segmentation method based on a reverse axial attention mechanism neural network, comprising:
[0008] According to the medical image, through a grouped residual backbone network, primary features and high-level features are obtained in sequence, where the area of the image region corresponding to the high-level features is larger than that of the primary features;
[0009] Aggregate the high-level features, and through a cascaded dilated convolutional network, extract multi-scale feature information and combine to obtain a global feature map;
[0010] According to the global feature map, through a reverse attention mechanism network, erase the foreground objects to obtain a tissue region, and combine with an axial attention mechanism network to obtain a key tissue region.
[0011] The medical image segmentation method based on the reverse axial attention mechanism neural network in the present invention further includes the following additional technical features:
[0012] According to the medical image, through a grouped residual backbone network, primary features and high-level features are obtained in sequence, specifically:
[0013] The grouped residual backbone network includes a first group of residual backbone networks and a second group of residual backbone networks with multiple layers of residual backbone networks,
[0014] According to the medical image, through the first group of residual backbone networks, the primary features are obtained;
[0015] According to the primary features, through the second group of residual backbone networks, the high-level features are obtained;
[0016] Where the number of layers of the residual backbone network in the second group of residual backbone networks is greater than that of the first group of residual backbone networks.
[0017] Aggregate the high-level features, and through a cascaded dilated convolutional network, extract multi-scale feature information and combine to obtain a global feature map, specifically:
[0018] According to the high-level features, through a multi-layer cascaded dilated convolutional network, extract feature information of different regional areas in sequence to obtain the multi-scale feature information;
[0019] Among them, the regional area for the multi-layer cascaded dilated convolutional network to extract feature information gradually increases.
[0020] The regional area for the multi-layer cascaded dilated convolutional network to extract feature information gradually increases, specifically:
[0021] The multi-layer cascaded dilated convolutional network has different dilation rates,
[0022] The dilation rates corresponding to the multi-layer cascaded dilated convolutional network gradually increase, so that the regional area for the corresponding cascaded dilated convolutional network to extract feature information gradually increases.
[0023] Extract multi-scale feature information and combine it to obtain a global feature map, specifically:
[0024] Perform a 1*1 convolution operation on the multi-scale feature information to obtain the global feature map;
[0025] The convolution kernel of the cascaded dilated convolution network has a size of 3*3.
[0026] According to the global feature map, use the reverse attention mechanism network to erase foreground objects to obtain the tissue region, specifically:
[0027] According to the global feature map, use the reverse attention mechanism network to start from any point in the foreground of the global feature map and gradually erase the foreground objects;
[0028] According to the boundary information of the tissue region, stop the foreground object erasing operation to retain the tissue region.
[0029] Obtain the tissue region and combine it with the axial attention mechanism network to obtain the key tissue region, specifically:
[0030] Set different adjustment weights according to multiple tissue regions, and the adjustment weights are determined according to the positions of the tissue regions in the medical image;
[0031] Combine with the axial attention mechanism network, calculate the horizontal axis and the vertical axis to obtain the key tissue region.
[0032] The medical image segmentation method based on the reverse axial attention mechanism neural network further includes:
[0033] Perform denoising processing on the medical image,
[0034] Delete the data with the blue channel value greater than 175 in the image;
[0035] For the data with the blue channel value less than or equal to 175 in the image, perform weighted averaging according to the red channel value, green channel value, and blue channel value to obtain a grayscale image.
[0036] The medical image segmentation method based on the reverse axial attention mechanism neural network further includes:
[0037] Perform quality assessment according to the key tissue region;
[0038] Perform post-processing operations according to the quality assessment results;
[0039] Among them, the post-processing operation at least includes Gaussian filtering denoising.
[0040] The present invention also provides an electronic device, including:
[0041] One or more central processing units,
[0042] One or more memories, in which a computer program is stored,
[0043] The central processing unit is used to execute the computer program to implement the medical image segmentation method based on the reverse axial attention mechanism neural network.
[0044] Due to the adoption of the above technical solution, the beneficial effects obtained by the present invention are as follows:
[0045] 1. In the present invention, according to the medical image, primary features and high-level features are sequentially obtained through a grouped residual backbone network, wherein the image area corresponding to the high-level features is larger than that of the primary features. Through the design of the grouped residual backbone network, the present invention significantly improves the ability of multi-scale feature extraction.
[0046] Specifically, the present invention can efficiently extract primary and high-level features, thereby significantly improving the overall feature representation ability. Especially in capturing rich feature information at different scales, the design of the cascaded dilated convolutional network plays an important role.
[0047] The present invention can capture rich feature information at different scales, effectively retain the detail information of the image, and improve the segmentation accuracy. This makes the present invention perform better in dealing with complex backgrounds and noise interference, significantly superior to the existing multi-scale feature extraction methods. For example, in the gastrointestinal polyp image segmentation task, the present invention can more accurately identify and segment the polyp area, providing a more refined and reliable segmentation result.
[0048] 2. In the present invention, the high-level features are aggregated, and multi-scale feature information is extracted through a cascaded dilated convolutional network to combine and obtain a global feature map. Through the design of the cascaded dilated convolutional network, the present invention significantly improves the ability to capture multi-scale feature information.
[0049] The present invention retains more detail information during the downsampling process through the cascaded dilated convolutional network, solves the problem that traditional methods are prone to losing detail information during the downsampling process, significantly improves the effect of multi-scale feature extraction, and provides higher accuracy and reliability for medical image segmentation.
[0050] In a preferred embodiment of the present invention, the area of the region for extracting feature information by the multi-level cascaded dilated convolutional network gradually increases. Specifically, the multi-level cascaded dilated convolutional network has different dilation rates, and the dilation rates corresponding to the multi-level cascaded dilated convolutional network gradually increase, so that the area of the region for extracting feature information by the corresponding cascaded dilated convolutional network gradually increases. The cascaded dilated convolutional network expands the receptive field without increasing parameters through different dilation rates, thereby extracting features at multiple scales and ensuring that the fine structures in the image are retained.
[0051] 3. In the present invention, according to the global feature map, the foreground object is erased through the reverse attention mechanism network to obtain the tissue region, and combined with the axial attention mechanism network, the key tissue region is obtained. By introducing the combination of the reverse axial attention mechanism and the axial attention mechanism, the present invention significantly enhances the feature extraction ability and segmentation accuracy.
[0052] Specifically, the reverse axial attention mechanism network is used to erase the foreground object, gradually excavate and identify the tissue region, while the axial attention mechanism network is used to maintain the global connection and effectively calculate, so as to capture more fine features. This combination method can more comprehensively describe the boundary and internal structure of the target object, making the model perform better when dealing with complex backgrounds and noise interference. By introducing the reverse axial attention mechanism, the present invention not only optimizes the extraction of local features, but also effectively combines global features, further improving the segmentation accuracy.
[0053] Especially in the medical image segmentation task, it can accurately identify the polyp region in the complex gastrointestinal polyp image and accurately segment its boundary and internal structure. By erasing the foreground object through the reverse attention mechanism and gradually refining the segmentation result, and combining the axial attention mechanism to maintain the consistency of global features, the present invention significantly improves the segmentation accuracy under complex backgrounds and noise interference.
[0054] In summary, through the design of the reverse axial attention mechanism, the present invention not only enables the model to more comprehensively describe the features of the target object, significantly improving the segmentation accuracy, but also greatly improves the reliability and accuracy in practical applications.
[0055] 4. The present invention adopts a neural network based on the reverse axial attention mechanism, enabling the computer to automatically summarize the criteria for segmenting images during multiple training and learning processes, so as to efficiently and accurately segment the gastroscope image object without manual intervention. Specifically, through the automated training and inference process of the neural network model, the present invention realizes highly automated image segmentation, significantly reducing the labor cost and greatly shortening the time for segmenting images, and greatly improving the work efficiency. And it can automatically learn and adapt to the complex patterns and structures in different images, reducing the dependence on manual intervention.
[0056] In addition, the highly automated feature of the present invention not only improves the processing speed but also excels in terms of accuracy. The medical image segmentation method based on the reverse axial attention mechanism neural network significantly outperforms traditional methods under complex backgrounds and noise interference. In practical applications, it can not only quickly generate high-quality segmentation results but also ensure the consistency and reliability of the results.
[0057] In summary, the present invention achieves high automation through neural network technology, significantly reducing the need for manual intervention, lowering labor costs, and greatly improving the efficiency and accuracy of image segmentation. This improvement is not only applicable to the gastrointestinal polyp image segmentation task but can also be widely applied to other medical image analysis fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0059] Figure 1 is a schematic flow chart of the medical image segmentation method based on the reverse axial attention mechanism neural network according to an embodiment of the present invention;
[0060] Figure 2 is a schematic structural diagram of the model based on the reverse axial attention mechanism neural network according to an embodiment of the present invention;
[0061] Figure 3 is a schematic structural diagram of the reverse axial attention mechanism network according to an embodiment of the present invention;
[0062] Figure 4 is a schematic structural diagram of the cascaded dilated convolutional network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] In order to more clearly illustrate the overall concept of the present invention, the following will be described in detail by way of examples with reference to the accompanying drawings of the specification.
[0064] In the following description, many specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0065] As Figure 1 shown, the medical image segmentation method based on the reverse axial attention mechanism neural network includes:
[0066] S100: According to the medical image, successively obtain the primary features and the high-level features through the grouped residual backbone network, where the area of the image region corresponding to the high-level features is larger than that of the primary features.
[0067] The main purpose of this step is to extract the primary features and the high-level features from the medical image through the grouped residual backbone network. These features will be used in subsequent processing to generate high-quality segmentation results. The primary features contain more detailed information, while the high-level features provide more abstract semantic information, which helps to identify and segment the target object.
[0068] Among them, the grouped residual backbone network is a convolutional neural network architecture improved from ResNet. By introducing grouped convolution and residual connection, it enhances the feature extraction ability. Res2Net can capture multi-scale features within each residual block, thereby improving the model's expressive ability and generalization performance.
[0069] The primary features refer to the low-level features extracted by the shallow network, usually containing more detailed information such as edges and textures. These features help to accurately depict the boundary of the target object.
[0070] The high-level features refer to the high-level features extracted by the deep network, usually containing more abstract semantic information such as shapes and positions. These features help to identify and classify the target object and provide global context information.
[0071] In this step, the medical image is input. The medical image to be segmented (such as a gastrointestinal endoscope image) is used as the input.
[0072] Extract features through the grouped residual backbone network. First, perform primary feature extraction, and use the shallow structure in the grouped residual backbone network to extract the primary features. These features usually correspond to a smaller receptive field and can capture the detailed information in the image.
[0073] Then, perform high-level feature extraction, and use the deep structure to extract the high-level features. The high-level features correspond to a larger receptive field and can capture the context information in a larger range, providing a more abstract semantic representation.
[0074] In this step, through the grouped residual backbone network, the primary and high-level features can be efficiently extracted, significantly improving the overall feature representation ability. By combining the primary and high-level features, the segmentation accuracy under complex backgrounds and noise interference is improved, making the segmentation results more accurate and reliable.
[0075] In this step, primary and high-level features are extracted from medical images through a grouped residual backbone network, aiming to provide rich feature representations for subsequent image segmentation tasks. The primary features focus on detailed information, while the high-level features provide global context information. The combination of the two significantly improves the segmentation accuracy and robustness. This design not only enhances the quality of the segmentation results but also increases the reliability of the model in practical applications.
[0076] S200: Aggregate the high-level features, extract multi-scale feature information through a cascaded dilated convolutional network, and combine them to obtain a global feature map.
[0077] The main purpose of this step is to aggregate high-level features through a cascaded dilated convolutional network, extract multi-scale feature information, and combine them to obtain a global feature map. This step aims to enhance the model's ability to capture features at different levels and scales, thereby improving the segmentation accuracy and robustness.
[0078] Among them, the cascaded dilated convolutional network refers to a special convolutional neural network structure that expands the receptive field without increasing parameters by cascading multiple convolutional layers with different dilation rates, thereby extracting features at multiple scales.
[0079] Multi-scale feature information refers to feature information extracted from different scales (such as local details, medium-range structures, global context). This information helps to describe the target object and its background more comprehensively.
[0080] The global feature map refers to a feature map composed of features at multiple scales, which can simultaneously contain local details and global context information, providing rich feature representations for subsequent segmentation tasks.
[0081] In this step, high-level features are aggregated. Among them, high-level features are features extracted from deep networks, which usually contain more abstract semantic information and receptive fields of different sizes. Multiple high-level features are fused to obtain a rich feature representation with both detailed information and context information.
[0082] Extract multi-scale feature information through a cascaded dilated convolutional network. Use the cascaded dilated convolutional network to process the aggregated features. Dilated convolution expands the receptive field without increasing parameters through different dilation rates, thereby extracting features at multiple scales. Specifically, cascade multiple convolutional layers with different dilation rates to gradually capture multi-scale feature information from local to global.
[0083] Finally, combine to obtain a global feature map. Combine the features extracted by each level of dilated convolutional layer to generate a global feature map. This global feature map fuses information at different scales and can describe the target object and its background more comprehensively.
[0084] Through the cascaded dilated convolutional network, the present invention can extract rich feature information at different scales, significantly enhancing the feature representation ability. During the multi-scale feature extraction process, more detailed information is retained, avoiding the decrease in segmentation accuracy caused by information loss. By combining multi-scale feature information, the segmentation accuracy under complex backgrounds and noise interference is improved, making the segmentation results more accurate and reliable.
[0085] Generally speaking, this step aggregates high-level features through the cascaded dilated convolutional network, extracts multi-scale feature information, and combines them to obtain a global feature map. The aim is to enhance the model's ability to capture features at different levels and scales, thereby improving the segmentation accuracy and robustness. In this way, the present invention can not only extract rich feature information at multiple scales but also effectively retain the detailed information of the image, providing strong support for subsequent high-quality segmentation. This design significantly improves the performance of the model in practical applications, especially under complex backgrounds and noise interference, and can still provide high-precision segmentation results.
[0086] S300: According to the global feature map, through the reverse attention mechanism network, erase the foreground object to obtain the tissue region, and combine it with the axial attention mechanism network to obtain the key tissue region.
[0087] The main purpose of this step is to erase the foreground object according to the global feature map through the reverse attention mechanism network to obtain the tissue region, and further refine it by combining the axial attention mechanism network to obtain the key tissue region. This step aims to improve the accuracy and precision of segmentation, ensuring that the target object (such as gastrointestinal polyps) can be accurately identified and segmented.
[0088] Among them, the reverse attention mechanism refers to a special attention mechanism used to gradually erase the foreground object (i.e., the background or non-target region) to highlight the tissue region of interest. It starts from any point in the foreground and gradually erases the parts irrelevant to the target, finally retaining the tissue region.
[0089] The tissue region refers to the preliminary segmentation result obtained after being processed by the reverse attention mechanism, which contains the main part of the target object but may still contain some noise or irrelevant information.
[0090] The axial attention mechanism refers to an attention mechanism that extracts global dependencies and local representations by calculating the horizontal and vertical axes, thereby capturing more fine-grained features. It can effectively calculate while maintaining global connections, further optimizing the feature extraction process.
[0091] The key tissue region is the final segmentation result obtained after being processed by the axial attention mechanism, which more accurately depicts the boundary and internal structure of the target object.
[0092] First, input the global feature map. Use the global feature map generated from the previous steps as the input. The global feature map fuses information at different scales and contains rich details and context information.
[0093] Erase the foreground objects through the reverse attention mechanism network. The reverse attention mechanism is used to gradually erase the foreground objects (i.e., the background or non-target regions) to highlight the tissue regions of interest. Specifically, the reverse attention mechanism starts from the foreground regions in the global feature map and gradually erases the parts irrelevant to the target until clear boundaries of the tissue regions are retained.
[0094] Obtain the tissue regions. After being processed by the reverse attention mechanism, preliminary tissue regions are obtained. These regions contain the main parts of the target objects but may still contain some noise or irrelevant information.
[0095] Combine with the axial attention mechanism network to obtain the key tissue regions. The axial attention mechanism is used to further optimize the tissue regions. By calculating the horizontal and vertical axes, it extracts global dependencies and local representations to capture more fine-grained features. Specifically, the axial attention mechanism further refines the tissue regions, enhances the attention to key features, and finally obtains more accurate key tissue regions.
[0096] In this step, the reverse attention mechanism can gradually erase the foreground objects, making the tissue regions clearer and reducing the influence of background noise. By combining the reverse attention mechanism and the axial attention mechanism, the present invention significantly improves the segmentation accuracy under complex backgrounds and noise interferences, ensuring that the segmentation results are more accurate and reliable. The axial attention mechanism can capture more fine-grained features while maintaining global connections, further improving the quality of the segmentation results.
[0097] Generally speaking, in this step, the reverse attention mechanism network is used to erase the foreground objects to obtain the tissue regions, and then combined with the axial attention mechanism network for further refinement to obtain the key tissue regions. The aim is to improve the segmentation accuracy and precision to ensure that the target objects can be accurately recognized and segmented. In this way, the present invention can not only effectively remove interferences under complex backgrounds but also capture more fine-grained features, thus providing high-quality segmentation results.
[0098] As a preferred embodiment of the present invention, according to the medical images, through the grouped residual backbone network, primary features and high-level features are obtained in sequence, specifically:
[0099] The grouped residual backbone network includes a first group of residual backbone networks and a second group of residual backbone networks with multiple layers of residual backbone networks.
[0100] Based on the medical image, the primary features are obtained through the first group of residual backbone networks;
[0101] Based on the primary features, the high-level features are obtained through the second group of residual backbone networks;
[0102] where the number of residual backbone network layers in the second group of residual backbone networks is greater than that in the first group of residual backbone networks.
[0103] The main purpose of this implementation is to sequentially extract primary features and high-level features from medical images through grouped residual backbone networks. The primary features capture low-level information in the image, such as details like edges and textures; while the high-level features contain more high-level semantic information and context information, such as global information like the shape of objects and combined textures. This method can effectively extract multi-scale features, improve segmentation accuracy, and through feature extraction at different levels, can better handle complex backgrounds and noise interference.
[0104] Input the medical image into the grouped residual backbone network. The medical image is a preprocessed endoscopic image of gastrointestinal polyps. These images have high resolution and clarity, but are also affected by factors such as shooting angle, lighting, and shadows.
[0105] The primary features are obtained through the first group of residual backbone networks. The first group of residual backbone networks contains multiple layers of residual backbone networks, and each residual backbone network is composed of multiple residual blocks, mainly used to extract primary features. The primary features mainly focus on local details, such as low-level information like edges and textures. These features usually represent the basic structure and detailed parts of the image. Specifically, when implementing, the medical image is input into the first group of residual backbone networks, and through a series of convolutional operations and residual connections, the primary features are gradually extracted.
[0106] The high-level features are obtained through the second group of residual backbone networks. The second group of residual backbone networks also contains multiple layers of residual backbone networks, and each residual backbone network is composed of multiple residual blocks, but has more layers, used to further extract high-level features. The high-level features contain more high-level semantic information and context information, such as global information like the shape of objects and combined textures. These features can better describe the overall structure and context environment of the target object. Specifically, when implementing, the primary features are used as the input and fed into the second group of residual backbone networks, and through deeper convolutional operations and residual connections, the high-level features are gradually extracted.
[0107] Among them, the number of residual backbone network layers in the second group of residual backbone networks is greater than that in the first group of residual backbone networks. This design is to ensure that more context information and global features can be captured when extracting high-level features, thereby improving the accuracy and robustness of segmentation.
[0108] In one embodiment, the first group of residual backbone networks extracts primary features. The first group of residual backbone networks contains two layers of residual backbone networks. Among them, the stem layer uses three 3×3 convolutional kernels (stride = 2, padding = 1) for downsampling while keeping the output feature resolution unchanged. The second layer consists of 3 residual blocks. The number of input channels changes from 64 to 64×4 = 256, and the image resolution remains unchanged. This stage mainly extracts primary features, focusing on low-level information such as edges and textures in the image.
[0109] The second group of residual backbone networks extracts high-level features. The second group of residual backbone networks contains three layers of residual backbone networks. Among them, the third layer of residual backbone networks consists of 4 residual blocks. The first residual block is a 3×3 convolution with a stride of 2, and the stride of the residual connection is also 2. The fourth layer consists of 6 residual blocks. The first residual block is a 3×3 convolution with a stride of 2, and the stride of the residual connection is also 2. The fifth layer consists of 3 residual blocks. The first residual block is a 3×3 convolution with a stride of 2, and the stride of the residual connection is also 2. This stage mainly extracts high-level features, focusing on high-level semantic information and context information such as shapes and combined textures in the image.
[0110] The extracted primary features and high-level features are aggregated through a parallel decoder, aiming to minimize the training parameters as much as possible while ensuring the receptive field and aggregating high-level features. This design not only improves the feature extraction ability but also simplifies the model structure and reduces the computational complexity.
[0111] In this embodiment, through the design of grouped residual backbone networks, multi-scale features can be effectively extracted, including primary features and high-level features. Primary features capture detailed information in the image, while high-level features contain more context information, which helps to improve the segmentation accuracy. By increasing the number of layers of the second group of residual backbone networks, more context information and global features can be captured when extracting high-level features, thereby enhancing the robustness of the model to complex backgrounds and noise interference. The skip connections in the residual blocks can effectively alleviate the gradient vanishing problem in deep networks, enabling the network to be trained deeper, thus improving the expression ability and generalization ability of the model.
[0112] As a preferred embodiment of the present invention, to aggregate the high-level features, multi-scale feature information is extracted through a cascaded dilated convolutional network, and a global feature map is combined as follows:
[0113] According to the high-level features, feature information of different regional areas is sequentially extracted through a multi-level cascaded dilated convolutional network to obtain the multi-scale feature information;
[0114] Among them, the regional areas for the multi-level cascaded dilated convolutional network to extract feature information gradually increase.
[0115] The main purpose of this embodiment is to aggregate high-level features through a cascaded dilated convolutional network, extract multi-scale feature information, and combine them to obtain a global feature map. This method can capture multi-level information in the image, including local details and global context, thereby improving the accuracy and robustness of segmentation.
[0116] First, high-level features are input into the cascaded dilated convolutional network. High-level features are features extracted from a deep network, containing more high-level semantic information and context information, such as global information like the shape of an object and combined texture.
[0117] Multi-scale feature information is extracted through a multi-level cascaded dilated convolutional network. A cascaded dilated convolutional network refers to a special convolutional neural network structure that expands the receptive field without increasing parameters by cascading multiple convolutional layers with different dilation rates, thereby extracting features at multiple scales. Dilated convolution gradually expands the receptive field through different dilation rates, enabling the capture of a larger range of information while maintaining the resolution.
[0118] Specifically, when implementing, the primary features and high-level features are used as inputs and fed into the multi-level cascaded dilated convolutional network to extract feature information of different regional areas layer by layer. In the first layer, a smaller dilation rate (e.g., r = 1) is used to extract local detail information. As the number of layers increases, the dilation rate gradually increases (e.g., r = 2, 4, 8) to extract context information over a larger range.
[0119] Specifically, the specific design of the cascaded dilated convolutional network is as follows:
[0120] In the first layer, the dilation rate is 1 (r = 1), and a 3*3 convolutional kernel is used to extract local detail information;
[0121] In the second layer, the dilation rate is 2 (r = 2), and a 3*3 convolutional kernel is used to further expand the receptive field and extract feature information over a slightly larger range;
[0122] In the third layer, the dilation rate is 4 (r = 4), and a 3*3 convolutional kernel is used to continue expanding the receptive field and extract feature information over a larger range;
[0123] In the fourth layer, the dilation rate is 8 (r = 8), and a 3*3 convolutional kernel is used to maximize the receptive field and extract global context information;
[0124] In the last layer, a 1×1 convolutional kernel is used to reduce the feature dimension and simplify the model structure.
[0125] As an example under this embodiment, the regional area for the multi-level cascaded dilated convolutional network to extract feature information gradually increases, specifically:
[0126] The multi-level cascaded dilated convolutional networks have different dilation rates,
[0127] The dilation rate of the multi-level cascaded dilated convolutional network gradually increases, so that the area of the region for extracting feature information by the corresponding cascaded dilated convolutional network gradually increases.
[0128] In this embodiment, in order to gradually increase the area of the region of feature information, the dilation rate of each dilated convolutional layer gradually increases, so that the receptive field of each layer also increases accordingly, thereby being able to extract feature information of different scales. For example, the dilation rate of the first layer is 1, the dilation rate of the second layer is 2, the dilation rate of the third layer is 4, the dilation rate of the fourth layer is 8, and so on, finally forming a feature representation containing multi-scale information.
[0129] As another embodiment under this implementation manner, to extract multi-scale feature information and combine it to obtain a global feature map, specifically:
[0130] Perform a 1*1 convolution operation on the multi-scale feature information to obtain the global feature map;
[0131] The size of the convolutional kernel of the cascaded dilated convolutional network is 3*3.
[0132] In this embodiment, a 1*1 convolution operation is performed to obtain a global feature map. After the extraction of multi-scale feature information is completed, these features are fused. The specific method is to perform dimensionality reduction processing on the multi-scale features through a 1×1 convolutional kernel to generate a global feature map. The main purpose of this step is to reduce the feature dimension, simplify the model structure, and avoid overfitting problems.
[0133] In addition, the size of the convolutional kernel of the cascaded dilated convolutional network is 3*3. Except for the last layer using a 1×1 convolutional kernel, the remaining layers all use 3×3 convolutional kernels. The 3×3 convolutional kernel can effectively extract rich multi-scale feature information while maintaining high resolution.
[0134] Specifically, as Figure 4 shown, it is a schematic diagram of the structure of the cascaded dilated convolutional network. Among them, Input refers to the input image.
[0135] dilated Conv r=1 refers to the dilated convolutional layer, and the dilation rate (rate) is 1.
[0136] dilated Conv r=2 refers to the dilated convolutional layer, and the dilation rate (rate) is 2.
[0137] dilated Conv r=4 refers to the dilated convolutional layer, and the dilation rate (rate) is 4.
[0138] dilated Conv r=8 refers to the dilated convolutional layer, and the dilation rate (rate) is 8.
[0139] The feature map refers to the feature map, which is the feature map extracted through multiple dilated convolutional layers.
[0140] Conv 1*1 refers to the 1*1 convolutional layer, which is used to fuse multi-scale features and generate the final output.
[0141] In this embodiment, in a specific medical image segmentation task, this multi-scale feature extraction method significantly improves the performance of the model. For example, in the gastrointestinal polyp image segmentation task, the primary features can capture details such as the edges and textures of the polyps, while the high-level features can identify the overall shape and location of the polyps. Through the multi-scale feature extraction of the cascaded dilated convolutional network, the model can not only accurately depict the boundaries of the polyps, but also maintain a high segmentation accuracy under complex backgrounds and noise interference.
[0142] In addition, the last layer uses a 1×1 convolutional kernel, which effectively reduces the feature dimension, simplifies the model structure, and avoids the overfitting problem. This makes the model more stable and efficient in practical applications.
[0143] As a preferred embodiment of the present invention, according to the global feature map, through the reverse attention mechanism network, the foreground object is erased to obtain the tissue region, specifically:
[0144] According to the global feature map, through the reverse attention mechanism network, starting from any point in the foreground of the global feature map, the foreground object is gradually erased;
[0145] According to the boundary information of the tissue region, the foreground object erasing operation is stopped to retain the tissue region.
[0146] The main purpose of this embodiment is to gradually erase the foreground object from the global feature map through the reverse attention mechanism network, so as to obtain a clear tissue region. This method can effectively remove background noise, highlight the target object (such as gastrointestinal polyps), and improve the segmentation accuracy and robustness.
[0147] Among them, the tissue region refers to the preliminary segmentation result obtained after being processed by the reverse attention mechanism, which contains the main part of the target object, but may still contain some noise or irrelevant information.
[0148] The boundary information refers to the edge information of the tissue region, which is used to judge when to stop the erasing operation to ensure the retention of the complete tissue region.
[0149] In this embodiment, the global feature map is input into the reverse attention mechanism network. The global feature map is a feature representation extracted and fused by a multi-level cascaded dilated convolutional network, which contains multi-level information from local details to global context.
[0150] The foreground object is gradually erased through the reverse attention mechanism network. The reverse attention mechanism starts from the foreground region in the global feature map and gradually erases the parts irrelevant to the target until the clear boundary of the tissue region is retained.
[0151] In specific implementation, any point in the global feature map is selected as the starting point to gradually erase the foreground object. This process is similar to "digging out" the target area from the image and gradually removing the background noise.
[0152] Finally, the erasing operation is stopped according to the boundary information of the tissue region. During the erasing process, the boundary information of the tissue region is used to judge when to stop the erasing operation. When reaching the boundary of the target area, further erasing is stopped to ensure the integrity of the tissue region. The key to this step is to accurately identify and retain the boundary information of the tissue region to avoid losing the information of the target object due to excessive erasing.
[0153] In a specific embodiment, the specific design of the reverse attention mechanism network is as follows:
[0154] Starting point selection: Start from any point in the global feature map and select a foreground point as the starting point;
[0155] Gradual erasing: Starting from the starting point, gradually erase the foreground object. This process is similar to "digging out" the target area from the image and gradually removing the background noise. The specific method is to calculate the importance weight of each pixel point through the reverse attention mechanism network, and gradually reduce the importance weight of the background area until these areas are completely ignored;
[0156] Boundary detection: During the erasing process, the boundary information of the tissue region is used to judge when to stop the erasing operation. When reaching the boundary of the target area, further erasing is stopped to ensure the integrity of the tissue region.
[0157] After the above processing, a clear tissue region is obtained, which contains the main part of the target object and removes most of the background noise.
[0158] In this specific medical image segmentation task, this reverse attention mechanism significantly improves the performance of the model. For example, in the gastrointestinal polyp image segmentation task, the reverse attention mechanism can gradually remove the background noise and highlight the polyp region, making the segmentation result more accurate.
[0159] The reverse attention mechanism can effectively remove the background noise, making the polyp region clearer and reducing the influence of the background noise on the segmentation result. By gradually erasing the background noise, the segmentation accuracy under complex background and noise interference is significantly improved, ensuring that the segmentation result is more accurate and reliable. Using the boundary information of the polyp region to judge when to stop the erasing operation ensures the integrity and accuracy of the polyp region.
[0160] As a preferred embodiment of the present invention, an organizational region is obtained, and in combination with an axial attention mechanism network, a key organizational region is obtained, specifically as follows:
[0161] According to the multiple organizational regions, different adjustment weights are set, and the adjustment weights are determined according to the positions of the organizational regions in the medical image;
[0162] In combination with the axial attention mechanism network, the horizontal axis and the vertical axis are calculated to obtain the key organizational region.
[0163] The main purpose of this embodiment is to combine the axial attention mechanism network to further highlight the key organizational region in the medical image by calculating the weights of the horizontal axis and the vertical axis. This method can not only improve the attention to specific tissues, but also effectively enhance the accuracy and robustness of the segmentation results, especially in the case of dealing with complex backgrounds or the presence of noise.
[0164] In this embodiment, different adjustment weights are first set.
[0165] The adjustment weights are determined according to the positions of the multiple organizational regions. First, it is necessary to analyze the multiple identified organizational regions and assign different adjustment weights according to their positions in the medical image.
[0166] These weights reflect the importance of each organizational region relative to the entire image and its key degree in the diagnosis process.
[0167] The principle of weight setting is that generally, regions closer to the center or with higher clinical significance will be given higher weights; while edge regions or non-critical parts will be given lower weights.
[0168] It can be understood that for each initially identified organizational region, its relative position in the image needs to be evaluated. For example, regions near the center of the image may be considered more important because they are closer to the observer's line of sight focus; on the contrary, edge regions may be considered less important.
[0169] Suppose there are three initially identified organizational regions A, B, and C, where A is located in the center of the image, B is near the edge, and C is between the two, then the following weights can be assigned to these three regions: A = 0.8, B = 0.2, C = 0.5.
[0170] Secondly, the horizontal axis and the vertical axis are calculated in combination with the axial attention mechanism network. The axial attention mechanism refers to a special attention mechanism dedicated to calculating the distribution of feature importance in the horizontal and vertical directions of the image.
[0171] On the horizontal axis, calculate the importance of each pixel point along the horizontal direction; similarly, on the vertical axis, calculate the importance of each pixel point along the vertical direction.
[0172] Specifically, first apply the axial attention mechanism to calculate the weight matrices of the horizontal axis and the vertical axis respectively. Then, combine these weight matrices with the original tissue region map to enhance or weaken the saliency of certain regions. Finally, by integrating the information of the horizontal axis and the vertical axis, a more accurate map of key tissue regions is obtained.
[0173] Specifically, input the tissue region with weights into the axial attention mechanism network to calculate the attention weight matrices on the horizontal axis and the vertical axis respectively.
[0174] On the horizontal axis, if a certain region becomes less important from left to right, the weight of this region on the horizontal axis may show a decreasing trend; vice versa.
[0175] Similarly, on the vertical axis, the corresponding weight matrices are also calculated. These weight matrices are then used to adjust the original tissue region map to enhance the saliency of those regions considered to be more important.
[0176] It should be noted that in this embodiment, a threshold can be set, and only when the comprehensive weight of a certain region exceeds this threshold, it is regarded as a key tissue region. At the same time, according to the requirements of different application scenarios, dynamically adjust the weight allocation method and threshold standard to adapt to different types of medical images and diagnostic needs.
[0177] Specifically, obtain the key tissue regions. By comprehensively considering the weight matrices on the horizontal axis and the vertical axis, we can obtain a new image representation in which the key tissue regions are strengthened. For example, in the above example, region A, which was originally in the center and had a high weight, will become more prominent and become the final key focus.
[0178] In this embodiment, through the axial attention mechanism, it is possible to more accurately locate and highlight the key tissue regions in medical images, which helps to improve the diagnostic efficiency and accuracy of doctors. In the face of complex backgrounds or noise interference, the axial attention mechanism can enhance the stability and robustness of the model by emphasizing key information and suppressing irrelevant details.
[0179] In practical applications, this axial attention mechanism is particularly suitable for the analysis of high-resolution medical images. For example, in the task of gastrointestinal polyp detection, it can effectively help doctors quickly lock in potential lesion regions and reduce the possibility of missed diagnoses. In addition, since this method can highlight key information while maintaining the overall structural integrity, it is also very suitable for the fields of teaching and research, providing strong support for the development of medical imaging.
[0180] In a specific embodiment, the specific structure of the model based on the reverse axial attention mechanism neural network is as Figure 2 shown. The model consists of multiple 3×3 convolutional kernels, 1×1 convolutional kernels, residual blocks, 3×3 max pooling layers, reverse attention mechanism, axial attention mechanism, and cascaded dilated convolutional blocks. The activation function used in the convolutional layer is the ReLU function, and the softmax layer is used at the end of the model.
[0181] As Figure 3 shown, it is an example of the reverse axial attention mechanism network and the feature map after passing through the reverse axial attention mechanism network. In the reverse axial attention mechanism network, the global feature can only capture the approximate position of the tissue, and the reverse attention is used to erase the foreground object, gradually mining and identifying the tissue area. In addition, the axial attention is used to maintain the global connection and calculate effectively, enabling it to capture more fine features.
[0182] Among them, f1, f2, f3, f4, f5 represent feature maps at different levels obtained through the Grouped Residual Backbone Network. From f1 to f5, the level of the feature map gradually deepens, and the corresponding image area gradually increases.
[0183] CDCM refers to the cascaded dilated convolutional module, which is used to capture rich feature information at different scales while retaining more detailed information.
[0184] A-RA refers to the axial reverse attention mechanism, which is used to erase the foreground object, gradually mining and identifying the tissue area.
[0185] Axial-attention refers to the axial attention mechanism module, which is used to perform attention calculations on the height axis and width axis respectively. Height axis refers to performing attention calculation on the height axis. Width axis refers to performing attention calculation on the width axis.
[0186] PD refers to the partial decoder, which is used to aggregate the primary features and high-level features to generate the final segmentation result.
[0187] Si represents different stages of the global feature map. Sg represents the global feature map, which is obtained by aggregating the primary features and high-level features.
[0188] Down-sampling refers to the down-sampling process, which is used to reduce the resolution of the feature map so as to extract higher-level features.
[0189] The Sigmoid function (S) refers to the Sigmoid activation function, which is used to convert the feature map into a probability distribution.
[0190] Feature flow refers to the feature flow, which represents the transmission path of features in the neural network.
[0191] Map flow refers to the mapping flow, which represents the mapping relationship of feature maps between different modules.
[0192] Prediction refers to the final prediction result, that is, the medical image segmentation result after processing.
[0193] f1' is the feature map after preliminary processing.
[0194] Generally speaking, in this embodiment, the processing flow based on the reverse axial attention mechanism neural network model is as follows: Input image: Input the original medical image.
[0195] Feature extraction: Sequentially obtain primary features (f1, f2, f3, f4, f5) through the grouped residual backbone network, where the image area corresponding to the high-level features is larger than that of the primary features.
[0196] Multi-scale feature extraction: Capture rich feature information at different scales through the cascaded dilated convolution module (CDCM) and retain more detailed information.
[0197] Axial reverse attention mechanism: Erase foreground objects through the axial reverse attention mechanism (A-RA), and gradually mine and identify tissue regions.
[0198] Aggregate features: Aggregate primary features and high-level features through the partial decoder (PD) to generate a global feature map (Sg).
[0199] Final prediction: Convert the feature map into a probability distribution through the sigmoid function to generate the final segmentation result.
[0200] Among them, the processing flow of the neural network structure with the reverse axial attention mechanism is as follows:
[0201] Input feature map (f1'): Input the feature map after preliminary processing.
[0202] Axial attention mechanism module processing (Axial-attention): Through the axial attention mechanism module, perform attention calculations on the height axis and width axis respectively to capture more fine features.
[0203] Obtain the global feature map (Si): Different stages of the global feature map obtained through the axial attention mechanism.
[0204] Use the Sigmoid function (S): Use the Sigmoid function to convert the feature map into a probability distribution.
[0205] Generate final prediction: Generate the final segmentation result.
[0206] As a preferred embodiment of the present invention, the medical image segmentation method based on the reverse axial attention mechanism neural network also includes:
[0207] According to the medical image, denoising is performed.
[0208] Delete the data with blue channel value greater than 175 in the image;
[0209] For the data with blue channel value less than or equal to 175 in the image, weighted average is performed according to the red channel value, green channel value, and blue channel value to obtain a grayscale image.
[0210] The main purpose of this embodiment is to preprocess the input medical image, especially to remove noise. By deleting the data with blue channel value greater than 175 and performing weighted averaging on the data with blue channel value less than or equal to 175 to generate a grayscale image, noise interference can be effectively reduced and the accuracy and robustness of subsequent segmentation tasks can be improved.
[0211] The blue channel value (B value) refers to a component in the RGB color model, representing the blue intensity of each pixel in the image. The larger the B value, the more blue the pixel is.
[0212] The red channel value (R value) refers to a component in the RGB color model that represents the red intensity of each pixel in the image. The larger the R value, the more red the pixel is.
[0213] The green channel value (G value) refers to a component in the RGB color model that represents the green intensity of each pixel in the image. The larger the G value, the more green the pixel is.
[0214] Grayscale images are images where each pixel is represented by a single brightness value. Grayscale images are often used to simplify image processing tasks because they only contain brightness information and no color information.
[0215] In this embodiment, first, a medical image is input. The input is an unprocessed medical image, such as an endoscopic image of a gastrointestinal polyp. These images may be affected by factors such as shooting angle and lighting conditions, and contain varying degrees of noise and artifacts, such as local overexposure caused by the endoscope light source, or bright spots caused by reflections.
[0216] Secondly, perform denoising processing. Delete the data with a blue channel value greater than 175. For each pixel, if its blue channel value (B value) is greater than 175, it is considered that the pixel may be noise or an outlier, and it is deleted from the image. For example, at a specific pixel, if R = 120, G = 130, B = 180, then this pixel will be deleted. This step helps to remove noise in high-brightness areas, especially those bright spots caused by uneven illumination or reflection.
[0217] Perform weighted averaging on the data with a blue channel value less than or equal to 175. For the pixels with a blue channel value less than or equal to 175, perform weighted averaging based on the red channel value (R value), green channel value (G value), and blue channel value (B value) to generate a grayscale image.
[0218] The weighted averaging formula can be adjusted according to specific situations. A common practice is to use the simple average (R + G + B) / 3 or a more complex weighted formula 0.299*R + 0.587*G + 0.114*B.
[0219] Finally, generate a grayscale image. Combine the processed results above into a new grayscale image, which serves as the basic input for subsequent feature extraction and segmentation tasks. The grayscale image can better highlight the structural information in the image, while reducing the interference brought by color information, making the subsequent processing more efficient and accurate.
[0220] In this embodiment, by deleting the pixels with too high blue channel values and performing weighted averaging, the noise and artifacts in the image can be effectively reduced, and the overall quality of the image can be improved. The generated grayscale image can better highlight the structural information in the image, providing a clearer input for subsequent feature extraction and segmentation tasks. A high-quality preprocessed image can significantly improve the performance of the segmentation algorithm, reduce the possibility of mis-segmentation, and improve the final diagnostic accuracy.
[0221] It should be noted that before inputting the medical image, since there are restrictions on the image size for subsequent image reading. Therefore, correct the image to eliminate the influence of factors such as image stretching and rotation on the image, and change the image to an appropriate size.
[0222] In addition, after obtaining the grayscale image, for the computer to read the image, perform image digitization, that is, convert it into three-dimensional numerical values in the RGB color space. Due to computer characteristics and human regulations, the range of the three RGB components of the image is from 0 to 255. For the convenience of subsequent process calculations, while reading the image, normalize the three RGB component values of each pixel in the image. The normalization method is to divide the component value of each pixel in the image by 127.5, and then subtract 1. That is, the normalization range of the image is [-1, 1].
[0223] The core step in image preprocessing is denoising, which can effectively suppress noise interference and enhance the features of the area to be segmented. The present invention does not limit the preprocessing steps such as image correction and image digitization.
[0224] As a preferred embodiment of the present invention, the medical image segmentation method based on the reverse axial attention mechanism neural network further includes:
[0225] Performing quality assessment according to the key tissue area;
[0226] Performing post-processing operations according to the quality assessment results;
[0227] Among them, the post-processing operation at least includes Gaussian filtering denoising.
[0228] The main purpose of this embodiment is to perform quality assessment on the key tissue area obtained by the reverse axial attention mechanism neural network, and perform post-processing operations according to the assessment results, such as Gaussian filtering denoising. This process aims to further improve the quality of the segmentation results, reduce noise interference, ensure that the finally output image has high clarity and accuracy, and thus improve the reliability and efficiency of medical diagnosis.
[0229] Among them, quality assessment refers to evaluating the overall quality and applicability of an image by calculating and analyzing a series of predefined quality indicators during the image processing process.
[0230] A Gaussian filter refers to a linear filter based on the Gaussian function, which is used to smooth the image and remove noise while trying to retain the edge information of the image.
[0231] First, input the key tissue area. The key tissue area map obtained from the reverse axial attention mechanism neural network is used as the input. These areas are the key parts after preliminary segmentation and highlighting, containing the main target objects (such as polyps) and their surrounding background information.
[0232] Secondly, perform quality assessment. Define the quality assessment criteria to clarify the criteria for evaluating the image quality. Common evaluation indicators include signal-to-noise ratio (SNR), contrast, edge sharpness, etc.
[0233] Among them, the signal-to-noise ratio (SNR) is used to measure the ratio of the signal intensity to the noise intensity in the image, and the higher the value, the better the image quality.
[0234] Contrast is used to reflect the brightness difference between the target object and the background. The higher the contrast, the easier it is to identify the target object.
[0235] Edge sharpness is used to evaluate whether the boundary of the target object is clear. The clearer the edge, the higher the image quality.
[0236] Calculate quality assessment metrics. For each key tissue region, calculate the above quality assessment metrics. For example, calculate the signal-to-noise ratio (SNR). Use the SNR formula SNR = 10 * log10(signal power / noise power) to calculate the SNR for each key tissue region. A lower SNR value indicates more noise and further processing is required.
[0237] Evaluate contrast. By means of histogram analysis, calculate the luminance difference between the target object and the background. Regions with low contrast may affect the recognition effect of the target object.
[0238] Check edge sharpness. Use the Canny edge detection algorithm to evaluate the sharpness of the boundary of the target object. Blurred edges may lead to inaccurate segmentation.
[0239] Secondly, perform post-processing operations according to the quality assessment results. If the SNR of a key tissue region is low (e.g., below a certain threshold), that is, the quality assessment result shows more noise, a Gaussian filter can be applied for denoising. Among them, the Gaussian filter refers to a commonly used linear smoothing filter that removes high-frequency noise in the image by weighted averaging while trying to retain the edge information of the image.
[0240] Select an appropriate Gaussian kernel size (such as 5x5 or 7x7) and set the standard deviation (σ) parameter. Generally, the larger the σ value, the stronger the filtering effect. Apply the Gaussian filter to the key tissue region to generate a denoised image. Check the quality of the denoised image and repeat the above steps if necessary to adjust the parameters to obtain the best effect.
[0241] Of course, other possible post-processing operations can also be adopted, and the present invention does not limit this. For example, morphological operations, such as opening and closing operations, are used to remove small noise points or fill holes. Use the opening operation (erode first and then dilate) to remove isolated small noise points; use the closing operation (dilate first and then erode) to fill holes to ensure the integrity of the target object. Among them, morphological operations refer to an image processing method based on set theory, which is often used for shape analysis and image enhancement. It mainly includes operations such as erosion, dilation, opening, and closing.
[0242] Enhance contrast. Through histogram equalization or adaptive contrast enhancement techniques, further improve the visual effect of the image to make the target object more clearly visible. For example, at a specific pixel point, if the original gray value range is narrow (such as 100 - 150), it can be extended to the entire gray range (0 - 255) through histogram equalization, thereby improving the visual effect of the image. Among them, contrast refers to the luminance difference between the brightest region and the darkest region in the image. High contrast helps to distinguish the target object and the background more clearly.
[0243] In this embodiment, through quality assessment and corresponding post-processing operations, especially Gaussian filtering denoising, the quality and accuracy of the image can be further improved based on the segmentation result. This method not only reduces noise interference and improves the overall clarity of the image, but also provides a better basis for subsequent advanced image analysis tasks.
[0244] The present invention also provides an electronic device, including:
[0245] One or more central processing units,
[0246] One or more memories, in which computer programs are stored,
[0247] The central processing unit is used to execute the computer program to implement the medical image segmentation method based on the reverse axial attention mechanism neural network.
[0248] Therefore, this electronic device can achieve any effect of the medical image segmentation method based on the reverse axial attention mechanism neural network, which will not be elaborated here.
[0249] What is not described in the present invention can be realized by adopting or referring to the existing technologies.
[0250] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments.
[0251] The above are only the embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A medical image segmentation method based on a reverse axial attention mechanism neural network, characterized in that: include: According to the medical image, primary features and advanced features are obtained in sequence through a grouped residual backbone network, wherein the grouped residual backbone network includes a first group of residual backbone networks and a second group of residual backbone networks having multiple layers of residual backbone networks, the number of residual backbone network layers in the second group of residual backbone networks is greater than that in the first group of residual backbone networks, the primary features are obtained through the first group of residual backbone networks, and the advanced features are obtained through the second group of residual backbone networks; the image region corresponding to the advanced features is larger than the primary features; Aggregate the high-level features, extract multi-scale feature information through cascaded dilated convolutional networks, and combine them to obtain a global feature map; According to the global feature map, the foreground object is erased through the reverse attention mechanism network to obtain the tissue area. Different adjustment weights are set according to the multiple tissue areas. The adjustment weights are determined according to the positions of the tissue areas in the medical image. Combined with the axial attention mechanism network, the horizontal axis and the vertical axis are calculated to obtain the key tissue area.
2. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 1 is characterized in that: Aggregate the high-level features, extract multi-scale feature information through cascaded dilated convolutional networks, and combine them to obtain a global feature map, specifically: According to the high-level features, feature information of different areas is extracted in sequence through a multi-layer cascaded atrous convolutional network to obtain the multi-scale feature information; Among them, the area where the multi-layer cascaded atrous convolutional network extracts feature information gradually increases.
3. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 2 is characterized in that: The area of feature information extracted by multi-layer cascaded dilated convolutional networks gradually increases, specifically: The multi-layer cascaded hole convolutional network has different hole rates. The hole rate corresponding to the multi-layer cascaded atrous convolutional network gradually increases, so that the area where the corresponding cascaded atrous convolutional network extracts feature information gradually increases.
4. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 2, characterized in that: Extract multi-scale feature information and combine them to obtain the global feature map, specifically: Performing a 1*1 convolution operation on the multi-scale feature information to obtain the global feature map; The size of the convolution kernel of the cascaded dilated convolutional network is 3*3.
5. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 1, characterized in that: According to the global feature map, the foreground object is erased through the reverse attention mechanism network to obtain the tissue area, specifically: According to the global feature map, starting from any point in the foreground of the global feature map, gradually erase the foreground object through a reverse attention mechanism network; According to the boundary information of the tissue area, the foreground object erasing operation is stopped to reserve the tissue area.
6. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 1, characterized in that: Also includes: According to the medical image, denoising is performed. Delete the data with blue channel value greater than 175 in the image; For the data with blue channel value less than or equal to 175 in the image, weighted average is performed according to the red channel value, green channel value, and blue channel value to obtain a grayscale image.
7. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 1, characterized in that: Also includes: Conduct quality assessments based on key organizational areas; Perform post-processing operations based on the quality assessment results; Wherein, the post-processing operation at least includes Gaussian filtering and denoising.
8. An electronic device, characterized in that: include: One or more CPUs, one or more memories having a computer program stored therein, The central processing unit is used to execute the computer program to implement the medical image segmentation method based on the reverse axial attention mechanism neural network as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Four-axis fusion method based on CNN and Transform
CN116188928A
Medical image segmentation method based on full convolutional neural network
CN119206237A