Medical image segmentation method based on reverse axial attention mechanism neural network

Through the medical image segmentation method based on the reverse axial attention mechanism neural network, the problem of poor segmentation effect in the existing technology under complex background and noise is solved, and high-precision and automated image segmentation are achieved.

CN119992106AActive Publication Date: 2025-05-13INSPUR GENERSOFT CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510458879.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing medical image segmentation technology has poor segmentation effect in complex backgrounds and noise situations, requiring manual intervention and adjustment, and it is easy to lose detailed information during the downsampling process, resulting in a decrease in segmentation accuracy.

Method used

The medical image segmentation method based on the reverse axial attention mechanism neural network is adopted to extract primary and advanced features through grouping residual backbone networks, the cascaded hollow convolution network extracts multi-scale features, the reverse attention mechanism erases the foreground objects, and the axial attention mechanism optimizes the organizational area.

Benefits of technology

It significantly improves the ability of multi-scale feature extraction, maintains the detailed information of the image, improves segmentation accuracy, reduces the need for manual intervention, and realizes highly automated image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992106A_ABST
    Figure CN119992106A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image segmentation method based on a reverse axial attention mechanism neural network, and belongs to the technical field of image processing, and the method comprises the steps: sequentially obtaining primary features and advanced features through a grouping residual backbone network according to a medical image, and enabling the area of an image region corresponding to the advanced features to be larger than that of the primary features; high-level features are aggregated, multi-scale feature information is extracted through a cascaded cavity convolutional network, and a global feature map is obtained through combination; and according to the global feature map, through a reverse attention mechanism network, a foreground object is erased to obtain a tissue area, and in combination with an axial attention mechanism network, a key tissue area is obtained. According to the method, multi-scale features are efficiently extracted through the grouping residual error backbone network and the cascade cavity convolution network, local and global features are optimized in combination with a reverse axial attention mechanism, and the medical image segmentation precision and robustness are remarkably improved. In addition, manual intervention is reduced, and the processing speed and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image segmentation, and in particular relates to a medical image segmentation method based on a reverse axial attention mechanism neural network. Background Art

[0002] Gastrointestinal polyps are a common digestive tract disease, usually caused by chronic inflammation, and manifest as raised lesions formed by mucosal hyperplasia and hypertrophy. Although most patients have no obvious symptoms, they are often found during endoscopy or imaging examinations. Early detection and accurate segmentation of gastrointestinal polyps are crucial for timely treatment and improved diagnostic efficiency.

[0003] Gastrointestinal polyp images are taken by endoscopes, which are affected by factors such as the endoscope's shooting angle, lighting, and shadows, which have a certain impact on image resolution and clarity. The commonly used method for identifying gastrointestinal polyps is to use image segmentation technology to separate the polyps in the image from the surrounding background. Then, feature extraction technology can be used to extract features in the image, such as color, texture, and shape. Next, a machine learning algorithm can be used to train a classifier to match the extracted features with labels (polyp / non-polyp) to identify gastrointestinal polyps in the image.

[0004] Although existing technologies and methods have made significant progress in the field of medical image segmentation, there are still some shortcomings. Some advanced segmentation models are too complex in design, and the segmentation effect is poor under complex background and noise conditions, requiring manual intervention and adjustment, which affects the accuracy of disease diagnosis. In addition, many existing methods lose important detail information during the downsampling process, resulting in reduced segmentation accuracy. Moreover, most methods are limited to attention in the spatial or channel dimensions. Due to the presence of factors such as image shadows, they cannot accurately capture the boundary details of tissue regions, resulting in inaccurate tissue region segmentation, affecting the accurate judgment of gastrointestinal polyps. Summary of the invention

[0005] The present invention provides a medical image segmentation method based on a reverse axial attention mechanism neural network to solve the problem of subjective differences and difficulty in disease judgment caused by the complexity of image segmentation, the need for manual intervention and adjustment, as well as the problem of inaccurate tissue region segmentation caused by the loss of detail information during the sampling process of the existing method and the inability to accurately capture the boundary details of the tissue region.

[0006] The technical solution adopted by the present invention is: Medical image segmentation method based on reverse axial attention mechanism neural network, including: According to the medical image, primary features and advanced features are obtained in sequence through a grouped residual backbone network, wherein the image region corresponding to the advanced features is larger than the primary features; Aggregate the high-level features, extract multi-scale feature information through cascaded dilated convolutional networks, and combine them to obtain a global feature map; According to the global feature map, the foreground object is erased through the reverse attention mechanism network to obtain the tissue area, and the key tissue area is obtained by combining with the axial attention mechanism network.

[0007] The medical image segmentation method based on the reverse axial attention mechanism neural network described in the present invention also includes the following additional technical features: According to the medical image, the primary features and advanced features are obtained in turn through the grouped residual backbone network, specifically: The grouped residual backbone network includes a first group of residual backbone networks and a second group of residual backbone networks having multiple layers of residual backbone networks, According to the medical image, obtaining the primary features through the first group of residual backbone networks; According to the primary features, the high-level features are obtained through the second group of residual backbone networks; The number of residual backbone network layers in the second group of residual backbone networks is greater than that in the first group of residual backbone networks.

[0008] Aggregate the high-level features, extract multi-scale feature information through cascaded dilated convolutional networks, and combine them to obtain a global feature map, specifically: According to the high-level features, feature information of different areas is extracted in sequence through a multi-layer cascaded atrous convolutional network to obtain the multi-scale feature information; Among them, the area where the multi-layer cascaded atrous convolutional network extracts feature information gradually increases.

[0009] The area of ​​feature information extracted by multi-layer cascaded dilated convolutional networks gradually increases, specifically: The multi-layer cascaded hole convolutional network has different hole rates. The hole rate corresponding to the multi-layer cascaded atrous convolutional network gradually increases, so that the area where the corresponding cascaded atrous convolutional network extracts feature information gradually increases.

[0010] Extract multi-scale feature information and combine them to obtain the global feature map, specifically: Performing a 1*1 convolution operation on the multi-scale feature information to obtain the global feature map; The size of the convolution kernel of the cascaded dilated convolutional network is 3*3.

[0011] According to the global feature map, the foreground object is erased through the reverse attention mechanism network to obtain the tissue area, specifically: According to the global feature map, starting from any point in the foreground of the global feature map, gradually erase the foreground object through a reverse attention mechanism network; According to the boundary information of the tissue area, the foreground object erasing operation is stopped to reserve the tissue area.

[0012] Get the organization area, combine it with the axial attention mechanism network, and get the key organization area, specifically: According to the plurality of tissue regions, different adjustment weights are set, wherein the adjustment weights are determined according to the positions of the tissue regions in the medical image; Combined with the axial attention mechanism network, the horizontal and vertical axes are calculated to obtain the focused organizational area.

[0013] The medical image segmentation method based on the reverse axial attention mechanism neural network also includes: According to the medical image, denoising is performed. Delete the data with blue channel value greater than 175 in the image; For the data with blue channel value less than or equal to 175 in the image, weighted average is performed according to the red channel value, green channel value, and blue channel value to obtain a grayscale image.

[0014] The medical image segmentation method based on the reverse axial attention mechanism neural network also includes: Conduct quality assessments based on key organizational areas; Perform post-processing operations based on the quality assessment results; Wherein, the post-processing operation at least includes Gaussian filtering and denoising.

[0015] The present invention also provides an electronic device, comprising: One or more CPUs, one or more memories having a computer program stored therein, The central processing unit is used to execute the computer program to implement the medical image segmentation method based on the reverse axial attention mechanism neural network.

[0016] Due to the adoption of the above technical solution, the beneficial effects achieved by the present invention are as follows: 1. In the present invention, primary features and advanced features are obtained in sequence through a grouped residual backbone network based on medical images, wherein the image area corresponding to the advanced features is larger than the primary features. The present invention significantly improves the ability of multi-scale feature extraction through the design of a grouped residual backbone network.

[0017] Specifically, the present invention can efficiently extract primary and advanced features, thereby significantly improving the overall feature representation capability. In particular, the design of the cascaded dilated convolutional network plays an important role in capturing rich feature information at different scales.

[0018] The present invention can capture rich feature information at different scales, effectively maintain the detailed information of the image, and improve the segmentation accuracy. This makes the present invention perform better when dealing with complex backgrounds and noise interference, and is significantly better than existing multi-scale feature extraction methods. For example, in the gastrointestinal polyp image segmentation task, the present invention can more accurately identify and segment the polyp area, providing more refined and reliable segmentation results.

[0019] 2. In the present invention, the high-level features are aggregated, and multi-scale feature information is extracted through a cascaded dilated convolutional network, and a global feature map is obtained by combining them. The present invention significantly improves the ability to capture multi-scale feature information through the design of a cascaded dilated convolutional network.

[0020] The present invention retains more detail information during the downsampling process by cascading atrous convolutional networks, solves the problem that traditional methods easily lose detail information during the downsampling process, significantly improves the effect of multi-scale feature extraction, and provides higher accuracy and reliability for medical image segmentation.

[0021] In a preferred embodiment of the present invention, the area of ​​the feature information extracted by the multi-layer cascaded dilated convolutional network is gradually increased, specifically: the multi-layer cascaded dilated convolutional network has different dilation rates, and the dilation rates corresponding to the multi-layer cascaded dilated convolutional network are gradually increased, so that the area of ​​the feature information extracted by the corresponding cascaded dilated convolutional network is gradually increased. The cascaded dilated convolutional network expands the receptive field without increasing parameters through different dilation rates, thereby extracting features at multiple scales, ensuring that the subtle structure in the image is preserved.

[0022] 3. In the present invention, according to the global feature map, the foreground object is erased through the reverse attention mechanism network to obtain the tissue area, and the key tissue area is obtained by combining the axial attention mechanism network. The present invention significantly enhances the feature extraction capability and segmentation accuracy by introducing the combination of the reverse axial attention mechanism and the axial attention mechanism.

[0023] Specifically, the reverse axial attention mechanism network is used to erase foreground objects, gradually dig out and identify tissue areas, while the axial attention mechanism network is used to maintain global connections and effectively calculate to capture more subtle features. This combination can more comprehensively describe the boundaries and internal structures of the target objects, making the model perform better when dealing with complex backgrounds and noise interference. By introducing the reverse axial attention mechanism, the present invention not only optimizes the extraction of local features, but also effectively combines global features, further improving the segmentation accuracy.

[0024] Especially in the task of medical image segmentation, it can accurately identify the polyp area in complex gastrointestinal polyp images and accurately segment its boundaries and internal structures. By erasing the foreground object through the reverse attention mechanism, gradually refining the segmentation results, and combining the axial attention mechanism to maintain the consistency of global features, the present invention significantly improves the segmentation accuracy under complex background and noise interference.

[0025] In summary, the reverse axial attention mechanism design in the present invention not only enables the model to more comprehensively describe the characteristics of the target object and significantly improves the segmentation accuracy, but also greatly improves the reliability and accuracy in practical applications.

[0026] 4. The present invention adopts a neural network based on the reverse axial attention mechanism, which enables the computer to automatically summarize the criteria for segmenting images in multiple training and learning, so that the gastrointestinal endoscope image objects can be efficiently and accurately segmented without human intervention. Specifically, through the automated training and reasoning process of the neural network model, the present invention realizes highly automated image segmentation, significantly reduces labor costs and greatly shortens the time for segmenting images, greatly improving work efficiency. It can also automatically learn and adapt to complex patterns and structures in different images, reducing dependence on human intervention.

[0027] In addition, the highly automated nature of the present invention not only improves processing speed, but also performs well in terms of accuracy. The medical image segmentation method based on the reverse axial attention mechanism neural network performs significantly better than traditional methods under complex backgrounds and noise interference. In practical applications, it can not only quickly generate high-quality segmentation results, but also ensure the consistency and reliability of the results.

[0028] In summary, the present invention achieves a high degree of automation through neural network technology, significantly reduces the need for manual intervention, reduces labor costs, and greatly improves the efficiency and accuracy of image segmentation. This improvement is not only applicable to gastrointestinal polyp image segmentation tasks, but can also be widely used in other medical image analysis fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 A schematic diagram of a process of a medical image segmentation method based on a reverse axial attention mechanism neural network according to an embodiment of the present invention; Figure 2 It is a structural schematic diagram of the model based on the reverse axial attention mechanism neural network in one embodiment of the present invention; Figure 3 A schematic diagram of the structure of the reverse axial attention mechanism network according to one embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of the cascaded atrous convolutional network according to one embodiment of the present invention. DETAILED DESCRIPTION

[0030] In order to more clearly illustrate the overall concept of the present invention, a detailed description is given below in an exemplary manner in conjunction with the accompanying drawings.

[0031] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.

[0032] like Figure 1 As shown, the medical image segmentation method based on the reverse axial attention mechanism neural network includes: S100: According to the medical image, primary features and high-level features are obtained in sequence through a grouped residual backbone network, wherein the image region corresponding to the high-level feature is larger in area than the primary feature.

[0033] The main purpose of this step is to extract primary and high-level features from medical images through a grouped residual backbone network. These features will be used to generate high-quality segmentation results in subsequent processing. Primary features contain more detailed information, while high-level features provide more abstract semantic information, which helps to identify and segment the target object.

[0034] Among them, the grouped residual backbone network is a convolutional neural network architecture improved from ResNet, which enhances feature extraction capabilities by introducing grouped convolution and residual connections. Res2Net can capture multi-scale features in each residual block, thereby improving the expressiveness and generalization performance of the model.

[0035] Primary features refer to low-level features extracted by shallow networks, which usually contain more detailed information, such as edges, textures, etc. These features help to accurately depict the boundaries of target objects.

[0036] High-level features refer to high-level features extracted by deep networks, which usually contain more abstract semantic information, such as shape, position, etc. These features help to identify and classify target objects and provide global context information.

[0037] In this step, a medical image is input. A medical image to be segmented (such as a gastrointestinal endoscopy image) is taken as input.

[0038] Feature extraction through the grouped residual backbone network. First, primary feature extraction is performed, using the shallow structure in the grouped residual backbone network to extract primary features. These features usually correspond to a smaller receptive field and can capture the details in the image.

[0039] Then, high-level feature extraction is performed to extract high-level features using deep structures. High-level features correspond to larger receptive fields, can capture contextual information in a larger range, and provide more abstract semantic representations.

[0040] In this step, the residual backbone network is grouped to efficiently extract primary and advanced features, significantly improving the overall feature representation capability. Combining primary and advanced features improves the segmentation accuracy under complex backgrounds and noise interference, making the segmentation results more accurate and reliable.

[0041] This step extracts primary and advanced features from medical images through a grouped residual backbone network, aiming to provide rich feature representations for subsequent image segmentation tasks. Primary features focus on detail information, while advanced features provide global context information. The combination of the two significantly improves segmentation accuracy and robustness. This design not only improves the quality of segmentation results, but also enhances the reliability of the model in practical applications.

[0042] S200: Aggregate the high-level features, extract multi-scale feature information through a cascaded atrous convolutional network, and combine them to obtain a global feature map.

[0043] The main purpose of this step is to aggregate high-level features through cascaded dilated convolutional networks, extract multi-scale feature information, and combine them to obtain a global feature map. This step aims to enhance the model's ability to capture features at different levels and scales, thereby improving segmentation accuracy and robustness.

[0044] Among them, the cascaded atrous convolutional network refers to a special convolutional neural network structure, which expands the receptive field without increasing parameters by cascading multiple convolutional layers with different atrous rates, thereby extracting features at multiple scales.

[0045] Multi-scale feature information refers to feature information extracted from different scales (such as local details, medium-range structures, and global context). This information helps to more comprehensively describe the target object and its background.

[0046] A global feature map refers to a feature map composed of features at multiple scales, which can contain both local details and global context information, providing rich feature representation for subsequent segmentation tasks.

[0047] In this step, high-level features are aggregated. High-level features are features extracted from deep networks, usually containing more abstract semantic information and receptive fields of different sizes. Multiple high-level features are fused to obtain a rich feature representation that contains both detailed information and contextual information.

[0048] Multi-scale feature information is extracted through cascaded dilated convolutional networks. The aggregated features are processed using cascaded dilated convolutional networks. Dilated convolutions use different dilation rates to expand the receptive field without increasing parameters, thereby extracting features at multiple scales. In specific implementation, multiple convolutional layers with different dilation rates are cascaded to gradually capture multi-scale feature information from local to global.

[0049] Finally, the global feature map is obtained by combining the features extracted by each level of dilated convolutional layer to generate a global feature map. This global feature map integrates information of different scales and can more comprehensively describe the target object and its background.

[0050] Through the cascaded atrous convolutional network, the present invention can extract rich feature information at different scales, significantly improving the feature representation capability. In the multi-scale feature extraction process, more detailed information is retained, avoiding the decrease in segmentation accuracy due to information loss. Combined with multi-scale feature information, the segmentation accuracy under complex background and noise interference is improved, making the segmentation result more accurate and reliable.

[0051] In summary, this step aggregates high-level features through cascaded dilated convolutional networks, extracts multi-scale feature information, and combines them to obtain a global feature map. It aims to enhance the model's ability to capture features of different levels and scales, thereby improving segmentation accuracy and robustness. In this way, the present invention can not only extract rich feature information at multiple scales, but also effectively maintain the detailed information of the image, providing strong support for subsequent high-quality segmentation. This design significantly improves the performance of the model in practical applications, especially under complex backgrounds and noise interference, and can still provide high-precision segmentation results.

[0052] S300: According to the global feature map, the foreground object is erased through the reverse attention mechanism network to obtain the tissue area, and the key tissue area is obtained by combining the axial attention mechanism network.

[0053] The main purpose of this step is to erase the foreground object through the reverse attention mechanism network according to the global feature map to obtain the tissue area, and further refine it in combination with the axial attention mechanism network to obtain the key tissue area. This step aims to improve the precision and accuracy of segmentation and ensure that the target object (such as gastrointestinal polyps) can be accurately identified and segmented.

[0054] Among them, the reverse attention mechanism refers to a special attention mechanism that is used to gradually erase the foreground object (i.e., the background or non-target area) to highlight the tissue area of ​​interest. It starts from any point in the foreground and gradually erases the parts that are not related to the target, and finally retains the tissue area.

[0055] The tissue region refers to the preliminary segmentation result obtained after the inverse attention mechanism, which contains the main part of the target object but may still contain some noise or irrelevant information.

[0056] Axial attention refers to an attention mechanism that extracts global dependencies and local representations by calculating horizontal and vertical axes to capture more subtle features. It can effectively calculate while maintaining global connections, further optimizing the feature extraction process.

[0057] The key tissue area is the final segmentation result obtained after processing by the axial attention mechanism, which more accurately depicts the boundary and internal structure of the target object.

[0058] First, input the global feature map. The global feature map generated from the previous steps is used as input. The global feature map integrates information of different scales and contains rich details and context information.

[0059] Erasing foreground objects through the inverse attention network. The inverse attention mechanism is used to gradually erase foreground objects (i.e., background or non-target areas) to highlight the tissue area of ​​interest. In specific implementation, the inverse attention mechanism starts from the foreground area in the global feature map and gradually erases parts that are not related to the target until a clear tissue area boundary is retained.

[0060] Get the organization region. After the reverse attention mechanism is processed, the preliminary organization region is obtained. These regions contain the main part of the target object, but may still contain some noise or irrelevant information.

[0061] Combined with the axial attention mechanism network, the key tissue area is obtained. The axial attention mechanism is used to further optimize the tissue area, extracting global dependencies and local representations by calculating the horizontal and vertical axes, thereby capturing more subtle features. In specific implementation, the axial attention mechanism further refines the tissue area, enhances the focus on key features, and finally obtains a more accurate key tissue area.

[0062] In this step, the reverse attention mechanism can gradually erase the foreground object, making the tissue area clearer and more specific, and reducing the impact of background noise. By combining the reverse attention mechanism and the axial attention mechanism, the present invention significantly improves the segmentation accuracy under complex background and noise interference, ensuring that the segmentation results are more accurate and reliable. The axial attention mechanism can capture more subtle features while maintaining global connections, further improving the quality of the segmentation results.

[0063] In summary, this step erases the foreground object through the reverse attention mechanism network to obtain the tissue area, and further refines it in combination with the axial attention mechanism network to obtain the key tissue area. It aims to improve the precision and accuracy of segmentation and ensure that the target object can be accurately identified and segmented. In this way, the present invention can not only effectively remove interference in a complex background, but also capture more subtle features, thereby providing high-quality segmentation results.

[0064] As a preferred implementation of the present invention, primary features and advanced features are obtained in sequence through a grouped residual backbone network according to medical images, specifically: The grouped residual backbone network includes a first group of residual backbone networks and a second group of residual backbone networks having multiple layers of residual backbone networks, According to the medical image, obtaining the primary features through the first group of residual backbone networks; According to the primary features, the high-level features are obtained through the second group of residual backbone networks; The number of residual backbone network layers in the second group of residual backbone networks is greater than that in the first group of residual backbone networks.

[0065] The main purpose of this implementation is to extract primary features and high-level features from medical images in sequence through a grouped residual backbone network. Primary features capture low-level information in the image, such as details such as edges and textures, while high-level features contain more high-level semantic information and contextual information, such as global information such as the shape of an object and combined textures. This method can effectively extract multi-scale features, improve segmentation accuracy, and better handle complex backgrounds and noise interference through feature extraction at different levels.

[0066] Medical images are input to the grouped residual backbone network. The medical images are pre-processed endoscopic images of gastrointestinal polyps. These images have high resolution and clarity, but are also affected by factors such as shooting angle, lighting, and shadows.

[0067] Primary features are obtained through the first group of residual backbone networks. The first group of residual backbone networks contains multiple layers of residual backbone networks, each of which consists of multiple residual blocks and is mainly used to extract primary features. Primary features focus on local details, such as low-level information such as edges and textures. These features usually represent the basic structure and details of the image. In specific implementation, the medical image is input into the first group of residual backbone networks, and primary features are gradually extracted through a series of convolution operations and residual connections.

[0068] The high-level features are obtained through the second set of residual backbone networks. The second set of residual backbone networks also contains multiple layers of residual backbone networks. Each residual backbone network consists of multiple residual blocks, but with more layers, which are used to further extract high-level features. High-level features contain more high-level semantic information and contextual information, such as global information such as the shape of the object and combined texture. These features can better describe the overall structure and context of the target object. In specific implementation, the primary features are used as input and passed into the second set of residual backbone networks. Through deeper convolution operations and residual connections, high-level features are gradually extracted.

[0069] Among them, the number of residual backbone network layers in the second group of residual backbone networks is larger than that in the first group of residual backbone networks. This design is to ensure that more contextual information and global features can be captured when extracting high-level features, thereby improving the accuracy and robustness of segmentation.

[0070] In one embodiment, the first set of residual backbone networks extracts primary features. The first set of residual backbone networks contains two layers of residual backbone networks. The stem layer uses three 3×3 convolution kernels (with a step size of 2 and padding=1) for downsampling to keep the output feature resolution unchanged. The second layer consists of three residual blocks, and the number of input channels changes from 64 to 64×4=256, and the image resolution remains unchanged. This stage mainly extracts primary features, focusing on low-level information such as edges and textures in the image.

[0071] The second group of residual backbone networks extracts high-level features. The second group of residual backbone networks contains three layers of residual backbone networks. Among them, the third layer of residual backbone network consists of 4 residual blocks, the first residual block is a 3×3 convolution with a step size of 2, and the step size of the residual connection is also 2. The fourth layer consists of 6 residual blocks, the first residual block is a 3×3 convolution with a step size of 2, and the step size of the residual connection is also 2. The fifth layer consists of 3 residual blocks, the first residual block is a 3×3 convolution with a step size of 2, and the step size of the residual connection is also 2. This stage mainly extracts high-level features, focusing on high-level semantic information and contextual information such as shape and combined texture in the image.

[0072] The extracted primary features and high-level features are aggregated through parallel decoders, with the goal of minimizing training parameters and aggregating high-level features while ensuring the receptive field. This design not only improves the ability of feature extraction, but also simplifies the model structure and reduces computational complexity.

[0073] This embodiment can effectively extract multi-scale features, including primary features and high-level features, through the design of a grouped residual backbone network. Primary features capture detailed information in the image, while high-level features contain more contextual information, which helps to improve the accuracy of segmentation. By increasing the number of layers of the second group of residual backbone networks, more contextual information and global features can be captured when extracting high-level features, thereby enhancing the robustness of the model to complex backgrounds and noise interference. The jump connection in the residual block can effectively alleviate the gradient vanishing problem in the deep network, allowing the network to train deeper layers, thereby improving the expressiveness and generalization capabilities of the model.

[0074] As a preferred implementation of the present invention, the high-level features are aggregated, and multi-scale feature information is extracted through a cascaded dilated convolutional network, and a global feature map is obtained by combining them, specifically: According to the high-level features, feature information of different areas is extracted in sequence through a multi-layer cascaded atrous convolutional network to obtain the multi-scale feature information; Among them, the area where the multi-layer cascaded atrous convolutional network extracts feature information gradually increases.

[0075] The main purpose of this implementation is to aggregate high-level features through cascaded dilated convolutional networks, extract multi-scale feature information, and combine them to obtain a global feature map. This method can capture multi-level information in the image, including local details and global context, thereby improving the accuracy and robustness of segmentation.

[0076] First, high-level features are input into the cascaded atrous convolutional network. High-level features are features extracted from deep networks, containing more high-level semantic information and contextual information, such as the shape of the object, combined texture and other global information.

[0077] Multi-scale feature information is extracted through multi-layer cascaded atrous convolutional networks. Cascaded atrous convolutional networks refer to a special convolutional neural network structure that expands the receptive field without increasing parameters by cascading multiple convolutional layers with different atrous rates, thereby extracting features at multiple scales. Atrous convolution gradually expands the receptive field through different atrous rates, and can capture a wider range of information while maintaining resolution.

[0078] In the specific implementation, the primary features and advanced features are used as input and passed into the multi-layer cascaded dilated convolutional network to extract feature information of different areas layer by layer. The first layer uses a smaller dilation rate (such as r=1) to extract local detail information. As the number of layers increases, the dilation rate gradually increases (such as r=2, 4, 8) to extract a wider range of context information.

[0079] Specifically, the specific design of the cascaded hole convolutional network is as follows: In the first layer, the void rate is 1 (r=1), and a 3*3 convolution kernel is used to extract local detail information; In the second layer, the dilation rate is 2 (r=2), and a 3*3 convolution kernel is used to further expand the receptive field and extract feature information in a slightly larger range; The third layer has a void rate of 4 (r=4) and uses a 3*3 convolution kernel to further expand the receptive field and extract feature information in a wider range; The fourth layer has a void rate of 8 (r=8) and uses a 3*3 convolution kernel to maximize the receptive field and extract global context information; The last layer uses a 1×1 convolution kernel to reduce the feature dimension and simplify the model structure.

[0080] As an example of this implementation, the area of ​​the multi-layer cascaded hole convolution network to extract feature information is gradually increased, specifically: The multi-layer cascaded hole convolutional network has different hole rates. The hole rate corresponding to the multi-layer cascaded atrous convolutional network gradually increases, so that the area where the corresponding cascaded atrous convolutional network extracts feature information gradually increases.

[0081] In this embodiment, in order to gradually increase the area of ​​the feature information, the hole rate of each level of hole convolution layer is gradually increased, so that the receptive field of each layer is also increased, so that feature information of different scales can be extracted. For example, the hole rate of the first layer is 1, the hole rate of the second layer is 2, the hole rate of the third layer is 4, the hole rate of the fourth layer is 8, and so on, finally forming a feature representation containing multi-scale information.

[0082] As another example of this implementation, multi-scale feature information is extracted and combined to obtain a global feature map, specifically: Performing a 1*1 convolution operation on the multi-scale feature information to obtain the global feature map; The size of the convolution kernel of the cascaded dilated convolutional network is 3*3.

[0083] In this embodiment, a 1*1 convolution operation is performed to obtain a global feature map. After the multi-scale feature information is extracted, these features are fused. The specific method is to perform dimensionality reduction processing on the multi-scale features through a 1×1 convolution kernel to generate a global feature map. The main purpose of this step is to reduce the feature dimension, simplify the model structure, and avoid overfitting problems.

[0084] In addition, the convolution kernel size of the cascaded dilated convolutional network is 3*3. Except for the last layer which uses a 1×1 convolution kernel, the remaining layers all use a 3×3 convolution kernel. The 3×3 convolution kernel can effectively extract rich multi-scale feature information while maintaining high resolution.

[0085] Specifically, Figure 4 As shown in Figure 1, it is a schematic diagram of the structure of the cascaded hole convolutional network. Here, Input refers to the input image.

[0086] dilated Conv r=1 refers to the dilated convolution layer with a dilation rate of 1.

[0087] dilated Conv r=2 refers to the dilated convolution layer with a dilation rate of 2.

[0088] dilated Conv r=4 refers to the dilated convolution layer with a dilation rate of 4.

[0089] dilated Conv r=8 refers to the dilated convolution layer with a dilation rate of 8.

[0090] Feature map refers to a feature map, which is a feature map extracted by multiple hole convolutional layers.

[0091] Conv 1*1 refers to a 1*1 convolutional layer, which is used to fuse multi-scale features and generate the final output.

[0092] In this implementation, this multi-scale feature extraction method significantly improves the performance of the model in specific medical image segmentation tasks. For example, in the gastrointestinal polyp image segmentation task, the primary features can capture detailed information such as the edges and textures of the polyps, while the high-level features can identify the overall shape and position of the polyps. Through the multi-scale feature extraction of the cascaded dilated convolutional network, the model can not only accurately depict the boundaries of the polyps, but also maintain a high segmentation accuracy under complex background and noise interference.

[0093] In addition, the last layer uses a 1×1 convolution kernel, which effectively reduces the feature dimension, simplifies the model structure, and avoids the overfitting problem. This makes the model more stable and efficient in practical applications.

[0094] As a preferred implementation of the present invention, according to the global feature map, the foreground object is erased through the reverse attention mechanism network to obtain the tissue area, specifically: According to the global feature map, starting from any point in the foreground of the global feature map, gradually erase the foreground object through a reverse attention mechanism network; According to the boundary information of the tissue area, the foreground object erasing operation is stopped to reserve the tissue area.

[0095] The main purpose of this implementation is to gradually erase the foreground object from the global feature map through the reverse attention mechanism network, so as to obtain a clear tissue area. This method can effectively remove background noise, highlight the target object (such as gastrointestinal polyps), and improve segmentation accuracy and robustness.

[0096] Among them, the tissue area refers to the preliminary segmentation result obtained after processing by the reverse attention mechanism, which contains the main part of the target object but may still contain some noise or irrelevant information.

[0097] Boundary information refers to edge information of a tissue region, and is used to determine when to stop the erasing operation to ensure that a complete tissue region is preserved.

[0098] In this embodiment, a global feature map is input into the reverse attention mechanism network. The global feature map is a feature representation extracted and fused by a multi-layer cascaded dilated convolutional network, and contains multi-level information from local details to global context.

[0099] The foreground object is gradually erased through the reverse attention mechanism network. The reverse attention mechanism starts from the foreground area in the global feature map and gradually erases the parts that are not related to the target until a clear tissue area boundary is retained.

[0100] In the specific implementation, we select any point in the global feature map as the starting point and gradually erase the foreground object. This process is similar to "digging out" the target area from the image and gradually removing the background noise.

[0101] Finally, the erasing operation is stopped according to the boundary information of the tissue region. During the erasing process, the boundary information of the tissue region is used to determine when to stop the erasing operation. When the boundary of the target region is reached, further erasing is stopped to ensure that the complete tissue region is preserved. The key to this step is to accurately identify and preserve the boundary information of the tissue region to avoid excessive erasing that causes loss of target object information.

[0102] In a specific embodiment, the specific design of the reverse attention mechanism network is as follows: Starting point selection, starting from any point in the global feature map, select a foreground point as the starting point; Erasing gradually, starting from the starting point, gradually erases the foreground object. This process is similar to "digging out" the target area from the image, gradually removing the background noise. The specific method is to use the reverse attention mechanism network to calculate the importance weight of each pixel point, and gradually reduce the importance weight of the background area until these areas are completely ignored; Boundary detection,During the erasing process, the boundary information of the tissue region is used to determine when to stop the erasing operation.,When the boundary of the target region is reached, further erasing is stopped, ensuring that the complete tissue region is preserved.

[0103] After the above processing, a clear tissue area is obtained, which includes the main part of the target object and removes most of the background noise.

[0104] In this embodiment, in a specific medical image segmentation task, this reverse attention mechanism significantly improves the performance of the model. For example, in a gastrointestinal polyp image segmentation task, the reverse attention mechanism can gradually remove background noise, highlight the polyp area, and make the segmentation result more accurate.

[0105] The reverse attention mechanism can effectively remove background noise, making the polyp area clearer and reducing the impact of background noise on the segmentation results. By gradually erasing the background noise, the segmentation accuracy under complex background and noise interference is significantly improved, ensuring that the segmentation results are more accurate and reliable. The boundary information of the polyp area is used to determine when to stop the erasing operation, ensuring the integrity and accuracy of the polyp area.

[0106] As a preferred embodiment of the present invention, the tissue area is obtained, and the key tissue area is obtained by combining the axial attention mechanism network, specifically: According to the plurality of tissue regions, different adjustment weights are set, wherein the adjustment weights are determined according to the positions of the tissue regions in the medical image; Combined with the axial attention mechanism network, the horizontal and vertical axes are calculated to obtain the focused organizational area.

[0107] The main purpose of this implementation is to combine the axial attention mechanism network to further highlight the key tissue areas in the medical image by calculating the weights of the horizontal and vertical axes. This method can not only improve the attention to specific tissues, but also effectively enhance the accuracy and robustness of the segmentation results, especially when dealing with complex backgrounds or in the presence of noise.

[0108] In this embodiment, different adjustment weights are set first.

[0109] Determining adjustment weights according to the locations of multiple tissue regions: First, it is necessary to analyze multiple identified tissue regions and assign different adjustment weights according to their locations in the medical image.

[0110] These weights reflect the importance of each tissue region relative to the entire image and how critical it is in the diagnostic process.

[0111] The principle of weight setting is that, generally speaking, areas closer to the center or with higher clinical significance will be given higher weights, while marginal areas or non-critical parts will be given lower weights.

[0112] It is understandable that for each initially identified tissue region, its relative position in the image needs to be evaluated. For example, regions near the center of the image may be considered more important because they are closer to the observer's visual focus; conversely, edge regions may be considered less important.

[0113] Assume that there are three initially identified tissue regions A, B, and C, where A is located in the center of the image, B is close to the edge, and C is between the two. The following weights can be assigned to these three regions: A = 0.8, B = 0.2, C = 0.5.

[0114] Secondly, the horizontal and vertical axes are calculated by combining the axial attention mechanism network. The axial attention mechanism refers to a special attention mechanism that is specifically used to calculate the feature importance distribution in the horizontal and vertical directions of the image.

[0115] On the horizontal axis, the importance of each pixel along the horizontal direction is calculated; similarly, on the vertical axis, the importance of each pixel along the vertical direction is calculated.

[0116] Specifically, the axial attention mechanism is first applied to calculate the weight matrices of the horizontal and vertical axes respectively. Then, these weight matrices are combined with the original organizational region map to enhance or weaken the saliency of certain regions. Finally, a more accurate focused organizational region map is obtained by integrating the information of the horizontal and vertical axes.

[0117] Specifically, the tissue regions with weights are input into the axial attention mechanism network, and the attention weight matrices on the horizontal and vertical axes are calculated respectively.

[0118] On the horizontal axis, if an area becomes less important from left to right, the weight of the area on the horizontal axis may show a decreasing trend, and vice versa.

[0119] Similarly, on the vertical axis, corresponding weight matrices are calculated, which are then used to adjust the original tissue region map to enhance the prominence of those regions that are considered more important.

[0120] It should be noted that in this embodiment, a threshold can be set, and only when the comprehensive weight of a certain area exceeds this threshold, it is regarded as a key tissue area. At the same time, according to the needs of different application scenarios, the weight allocation method and threshold standard are dynamically adjusted to adapt to different medical image types and diagnostic needs.

[0121] Specifically, the key tissue area is obtained. Taking into account the weight matrices on the horizontal and vertical axes, we can obtain a new image representation in which the key tissue area is strengthened. For example, in the above example, area A, which is originally located in the center and has a higher weight, will become more prominent and become the final focus.

[0122] This implementation, through the axial attention mechanism, can more accurately locate and highlight key tissue areas in medical images, helping to improve doctors' diagnostic efficiency and accuracy. In the face of complex backgrounds or noise interference, the axial attention mechanism can enhance the stability and robustness of the model by emphasizing key information and suppressing irrelevant details.

[0123] In practical applications, this axial attention mechanism is particularly suitable for the analysis of high-resolution medical images. For example, in the gastrointestinal polyp detection task, it can effectively help doctors quickly lock potential lesion areas and reduce the possibility of missed diagnosis. In addition, since this method can highlight key information while maintaining the integrity of the overall structure, it is also very suitable for use in teaching and research fields, providing strong support for the development of medical imaging.

[0124] In a specific embodiment, the specific structure of the model based on the reverse axial attention mechanism neural network is as follows Figure 2 As shown in the figure, the model consists of multiple 3×3 convolution kernels, 1×1 convolution kernels, residual blocks, 3×3 max pooling layers, reverse attention mechanisms, axial attention mechanisms, and cascaded hole convolution blocks. The activation functions used in the convolution layers are all ReLU functions, and the model finally uses a softmax layer.

[0125] like Figure 3 The figure shows an example of a reverse axial attention network and a feature map after the reverse axial attention network. In the reverse axial attention network, global features can only capture the approximate location of the tissue. Reverse attention is used to erase foreground objects and gradually mine and identify tissue areas. In addition, axial attention is used to maintain global connections and effectively calculate, so that it can capture more subtle features.

[0126] Among them, f1, f2, f3, f4, and f5 represent feature maps of different levels obtained through the Grouped Residual Backbone Network. From f1 to f5, the level of the feature map gradually deepens, and the corresponding image area gradually increases.

[0127] CDCM refers to cascaded dilated convolutional module, which is used to capture rich feature information at different scales while retaining more detail information.

[0128] A-RA refers to the axial reverse attention mechanism, which is used to erase foreground objects and gradually mine and identify tissue regions.

[0129] Axial-attention refers to the axial attention mechanism module, which is used to perform attention calculations on the height axis and width axis respectively. Height axis refers to performing attention calculations on the height axis. Width axis refers to performing attention calculations on the width axis.

[0130] PD refers to partial decoder, which is used to aggregate primary features and high-level features to generate the final segmentation result.

[0131] Si represents different stages of the global feature map. Sg represents the global feature map, which is obtained by aggregating primary features and high-level features.

[0132] Down-sampling refers to the downsampling process, which is used to reduce the resolution of feature maps to extract higher-level features.

[0133] Sigmoid function (S) refers to the Sigmoid activation function, which is used to convert the feature map into a probability distribution.

[0134] Feature flow refers to feature flow, which represents the transmission path of features in the neural network.

[0135] Map flow refers to mapping flow, which indicates the mapping relationship between feature maps in different modules.

[0136] Prediction refers to the final prediction result, that is, the processed medical image segmentation result.

[0137] f1' is the feature map after preliminary processing.

[0138] In summary, in this embodiment, the processing flow of the neural network model based on the reverse axial attention mechanism is as follows: Input image: Input the original medical image.

[0139] Feature extraction: The primary features (f1, f2, f3, f4, f5) are obtained in sequence through the grouped residual backbone network, where the image area corresponding to the high-level features is larger than the primary features.

[0140] Multi-scale feature extraction: Capture rich feature information at different scales and retain more detailed information through the cascaded atrous convolution module (CDCM).

[0141] Axial Reverse Attention Mechanism: The foreground objects are erased through the Axial Reverse Attention Mechanism (A-RA) to gradually mine and identify tissue regions.

[0142] Aggregate features: The low-level features and high-level features are aggregated through the partial decoder (PD) to generate a global feature map (Sg).

[0143] Final prediction: The feature map is converted into a probability distribution through the sigmoid function to generate the final segmentation result.

[0144] Among them, the reverse axial attention mechanism structure neural network processing flow is as follows: Input feature map (f1'): Input feature map after preliminary processing.

[0145] Axial-attention: Through the axial attention module, attention calculations are performed on the height axis and width axis respectively to capture more subtle features.

[0146] Obtaining the global feature map (Si): Different stages of the global feature map obtained through the axial attention mechanism.

[0147] Use Sigmoid function (S): Use the Sigmoid function to convert the feature map into a probability distribution.

[0148] Generate final prediction: Generate the final segmentation result.

[0149] As a preferred embodiment of the present invention, the medical image segmentation method based on the reverse axial attention mechanism neural network also includes: According to the medical image, denoising is performed. Delete the data with blue channel value greater than 175 in the image; For the data with blue channel value less than or equal to 175 in the image, weighted average is performed according to the red channel value, green channel value, and blue channel value to obtain a grayscale image.

[0150] The main purpose of this embodiment is to preprocess the input medical image, especially to remove noise. By deleting the data with blue channel value greater than 175 and performing weighted averaging on the data with blue channel value less than or equal to 175 to generate a grayscale image, noise interference can be effectively reduced and the accuracy and robustness of subsequent segmentation tasks can be improved.

[0151] The blue channel value (B value) refers to a component in the RGB color model, representing the blue intensity of each pixel in the image. The larger the B value, the more blue the pixel is.

[0152] The red channel value (R value) refers to a component in the RGB color model that represents the red intensity of each pixel in the image. The larger the R value, the more red the pixel is.

[0153] The green channel value (G value) refers to a component in the RGB color model that represents the green intensity of each pixel in the image. The larger the G value, the more green the pixel is.

[0154] Grayscale images are images where each pixel is represented by a single brightness value. Grayscale images are often used to simplify image processing tasks because they only contain brightness information and no color information.

[0155] In this embodiment, first, a medical image is input. The input is an unprocessed medical image, such as an endoscopic image of a gastrointestinal polyp. These images may be affected by factors such as shooting angle and lighting conditions, and contain varying degrees of noise and artifacts, such as local overexposure caused by the endoscope light source, or bright spots caused by reflections.

[0156] Next, perform denoising. Delete data with a blue channel value greater than 175. For each pixel, if its blue channel value (B value) is greater than 175, it is considered that the pixel may be noise or an outlier and is deleted from the image. For example, at a specific pixel, if R=120, G=130, B=180, the pixel will be deleted. This step helps remove noise in high-brightness areas, especially those bright spots caused by uneven lighting or reflections.

[0157] The data with blue channel value less than or equal to 175 are weighted averaged. For pixels with blue channel value less than or equal to 175, the red channel value (R value), green channel value (G value) and blue channel value (B value) are weighted averaged to generate a grayscale image.

[0158] The weighted average formula can be adjusted according to the specific situation. A common practice is to use a simple average (R + G + B) / 3 or a more complex weighted formula 0.299*R + 0.587*G + 0.114*B.

[0159] Finally, a grayscale image is generated. The above processed results are combined into a new grayscale image as the basic input for subsequent feature extraction and segmentation tasks. The grayscale image can better highlight the structural information in the image while reducing the interference caused by color information, making subsequent processing more efficient and accurate.

[0160] This implementation can effectively reduce noise and artifacts in the image and improve the overall quality of the image by deleting pixels with too high blue channel values ​​and performing weighted averaging. The generated grayscale image can better highlight the structural information in the image and provide clearer input for subsequent feature extraction and segmentation tasks. High-quality preprocessed images can significantly improve the performance of the segmentation algorithm, reduce the possibility of mis-segmentation, and improve the final diagnostic accuracy.

[0161] It should be noted that before inputting medical images, the image size is limited by the subsequent image reading. Therefore, the image is corrected to eliminate the influence of factors such as image expansion and rotation on the image and change the image to a suitable size.

[0162] In addition, after obtaining the grayscale image, in order to read the image through the computer, the image is digitized, that is, converted into a three-dimensional value in the RGB color space. Due to computer characteristics and human regulations, the range of the three RGB components of the image is 0 to 255. In order to facilitate the calculation of subsequent processes, the three RGB component values ​​of each pixel of the image are normalized while reading the image. The normalization method is to divide the component value of each pixel of the image by 127.5 and then subtract 1. That is, the normalized range of the image is [-1,1].

[0163] The core step in image preprocessing is denoising, which can effectively suppress noise interference and enhance the features of the area to be segmented. The present invention does not limit the preprocessing steps such as image correction and image digitization.

[0164] As a preferred embodiment of the present invention, the medical image segmentation method based on the reverse axial attention mechanism neural network also includes: Conduct quality assessments based on key organizational areas; Perform post-processing operations based on the quality assessment results; Wherein, the post-processing operation at least includes Gaussian filtering and denoising.

[0165] The main purpose of this implementation is to evaluate the quality of the key tissue areas obtained by the reverse axial attention mechanism neural network, and perform post-processing operations such as Gaussian filtering denoising based on the evaluation results. This process aims to further improve the quality of the segmentation results, reduce noise interference, and ensure that the final output image has high clarity and accuracy, thereby improving the reliability and efficiency of medical diagnosis.

[0166] Among them, quality assessment refers to the evaluation of the overall quality and applicability of the image by calculating and analyzing a series of predefined quality indicators during the image processing process.

[0167] A Gaussian filter is a linear filter based on a Gaussian function, which is used to smooth an image and remove noise while preserving the edge information of the image as much as possible.

[0168] First, the key tissue regions are input. The key tissue region maps obtained from the reverse axial attention mechanism neural network are used as input. These regions are the key parts after preliminary segmentation and highlighting, containing the main target objects (such as polyps) and their surrounding background information.

[0169] Secondly, conduct quality assessment. Define quality assessment criteria and clarify the criteria used to assess image quality. Common evaluation indicators include signal-to-noise ratio (SNR), contrast, edge clarity, etc.

[0170] Among them, the signal-to-noise ratio (SNR) is used to measure the ratio of signal intensity to noise intensity in an image. The higher the value, the better the image quality.

[0171] Contrast is used to reflect the brightness difference between the target object and the background. The higher the contrast, the easier it is to identify the target object.

[0172] Edge clarity is used to evaluate whether the boundary of the target object is clear. The clearer the edge, the higher the image quality.

[0173] Calculate quality assessment indicators. For each key tissue area, calculate the above quality assessment indicators. For example, calculate the signal-to-noise ratio (SNR). Use the signal-to-noise ratio formula SNR = 10 * log10(signal power / noise power) to calculate the signal-to-noise ratio of each key tissue area. Lower SNR values ​​indicate that there is more noise and further processing is required.

[0174] Evaluate contrast. Calculate the brightness difference between the target object and the background through histogram analysis. Areas with low contrast may affect the recognition effect of the target object.

[0175] Check edge clarity. Use the Canny edge detection algorithm to assess the clarity of the target object's boundaries. Blurred edges may lead to inaccurate segmentation.

[0176] Secondly, post-processing operations are performed based on the quality assessment results. If the signal-to-noise ratio of a key tissue area is low (such as below a certain threshold), that is, the quality assessment results show that there is a lot of noise, a Gaussian filter can be applied for denoising. Among them, the Gaussian filter refers to a commonly used linear smoothing filter that removes high-frequency noise in the image by weighted averaging while retaining the edge information of the image as much as possible.

[0177] Select an appropriate Gaussian kernel size (e.g., 5x5 or 7x7) and set the standard deviation (σ) parameter. Generally, the larger the σ value, the stronger the filtering effect. Apply the Gaussian filter to the focal tissue region to generate a denoised image. Check the quality of the denoised image and repeat the above steps if necessary to adjust the parameters for the best results.

[0178] Of course, other possible post-processing operations can also be used, and the present invention does not limit this. For example, morphological operations, such as opening and closing operations, are used to remove small noise points or fill holes. Use opening operations (corrosion first and then expansion) to remove isolated small noise points; use closing operations (expansion first and then corrosion) to fill holes to ensure the integrity of the target object. Among them, morphological operations refer to an image processing method based on set theory, which is often used for shape analysis and image enhancement. It mainly includes operations such as corrosion, expansion, opening operations and closing operations.

[0179] Contrast enhancement, through histogram equalization or adaptive contrast enhancement technology, further improves the visual effect of the image and makes the target object more visible. For example, at a specific pixel, the original grayscale value range is narrow (such as 100-150), which can be expanded to the entire grayscale range (0-255) through histogram equalization, thereby improving the visual effect of the image. Among them, contrast refers to the brightness difference between the brightest and darkest areas in the image. High contrast helps to distinguish the target object from the background more clearly.

[0180] This implementation can further improve the quality and accuracy of the image based on the segmentation result through quality assessment and corresponding post-processing operations, especially Gaussian filtering denoising. This method not only reduces noise interference and improves the overall clarity of the image, but also provides a better foundation for subsequent advanced image analysis tasks.

[0181] The present invention also provides an electronic device, comprising: One or more CPUs, one or more memories having a computer program stored therein, The central processing unit is used to execute the computer program to implement the medical image segmentation method based on the reverse axial attention mechanism neural network.

[0182] Therefore, this electronic device can achieve any effect of the medical image segmentation method based on the reverse axial attention mechanism neural network, which will not be elaborated here.

[0183] Anything not described in the present invention can be achieved by adopting or drawing on existing technologies.

[0184] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

[0185] The above description is only an embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of the claims of the present invention.

Claims

1. A medical image segmentation method based on a reverse axial attention mechanism neural network, characterized in that: include: According to the medical image, primary features and advanced features are obtained in sequence through a grouped residual backbone network, wherein the image region corresponding to the advanced features is larger than the primary features; Aggregate the high-level features, extract multi-scale feature information through cascaded dilated convolutional networks, and combine them to obtain a global feature map; According to the global feature map, the foreground object is erased through the reverse attention mechanism network to obtain the tissue area, and the key tissue area is obtained by combining with the axial attention mechanism network.

2. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 1 is characterized in that: According to the medical image, the primary features and advanced features are obtained in turn through the grouped residual backbone network, specifically: The grouped residual backbone network includes a first group of residual backbone networks and a second group of residual backbone networks having multiple layers of residual backbone networks, According to the medical image, obtaining the primary features through the first group of residual backbone networks; According to the primary features, the high-level features are obtained through the second group of residual backbone networks; The number of residual backbone network layers in the second group of residual backbone networks is greater than that in the first group of residual backbone networks.

3. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 1, characterized in that: Aggregate the high-level features, extract multi-scale feature information through cascaded dilated convolutional networks, and combine them to obtain a global feature map, specifically: According to the high-level features, feature information of different areas is extracted in sequence through a multi-layer cascaded atrous convolutional network to obtain the multi-scale feature information; Among them, the area where the multi-layer cascaded atrous convolutional network extracts feature information gradually increases.

4. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 3 is characterized in that: The area of ​​feature information extracted by multi-layer cascaded dilated convolutional networks gradually increases, specifically: The multi-layer cascaded hole convolutional network has different hole rates. The hole rate corresponding to the multi-layer cascaded atrous convolutional network gradually increases, so that the area where the corresponding cascaded atrous convolutional network extracts feature information gradually increases.

5. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 3 is characterized in that: Extract multi-scale feature information and combine them to obtain the global feature map, specifically: Performing a 1*1 convolution operation on the multi-scale feature information to obtain the global feature map; The size of the convolution kernel of the cascaded dilated convolutional network is 3*3.

6. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 1, characterized in that: According to the global feature map, the foreground object is erased through the reverse attention mechanism network to obtain the tissue area, specifically: According to the global feature map, starting from any point in the foreground of the global feature map, gradually erase the foreground object through a reverse attention mechanism network; According to the boundary information of the tissue area, the foreground object erasing operation is stopped to reserve the tissue area.

7. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 1, characterized in that: Get the organization area, combine it with the axial attention mechanism network, and get the key organization area, specifically: According to the plurality of tissue regions, different adjustment weights are set, wherein the adjustment weights are determined according to the positions of the tissue regions in the medical image; Combined with the axial attention mechanism network, the horizontal and vertical axes are calculated to obtain the focused organizational area.

8. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 1, characterized in that: Also includes: According to the medical image, denoising is performed. Delete the data with blue channel value greater than 175 in the image; For the data with blue channel value less than or equal to 175 in the image, weighted average is performed according to the red channel value, green channel value, and blue channel value to obtain a grayscale image.

9. The medical image segmentation method based on the reverse axial attention mechanism neural network according to claim 1, characterized in that: Also includes: Conduct quality assessments based on key organizational areas; Perform post-processing operations based on the quality assessment results; Wherein, the post-processing operation at least includes Gaussian filtering and denoising.

10. An electronic device, characterized in that: include: One or more CPUs, one or more memories having a computer program stored therein, The central processing unit is used to execute the computer program to implement the medical image segmentation method based on the reverse axial attention mechanism neural network as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Polyp segmentation method based on lightweight network model and reverse attention module

    CN114627137A

  • Four-axis fusion method based on CNN and Transform

    CN116188928A

  • Salient target detection method based on adaptive feature fusion

    CN117115601A

  • Polyp image segmentation method and system

    CN119048538A

  • Medical image segmentation method based on full convolutional neural network

    CN119206237A