Polyp segmentation method and device based on state awareness attention network
By applying a state-perceptual attention network method in polyp detection, using channel state transfer and spatial state inference attention modules, the problem of difficult to identify similarly lesions and small polyps lesions in the prior art is solved, and more efficient and accurate polyps detection is achieved.
Patent Information
- Application Number
- CN202510165541.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
At this stage, automated polyp detection methods based on deep learning are difficult to accurately identify lesion polyps with similar shapes to normal tissues, and it is difficult to capture small polyps, resulting in an increased risk of missed detection, limiting its large-scale application in clinical practice.
The polyp segmentation method based on the state-perceptual attention network is adopted to model the long-term dependence between the lesion area and the normal area through the channel state transfer attention module to distinguish the lesion polyps from normal tissue; at the same time, the spatial state reasoning attention module is used to focus on the effective information of the target and capture small polyps.
It improves the efficiency and performance of polyp detection, can accurately identify polyps, including small polyps, reduces the risk of missed detection, and solves the limitations of the application of the existing technology in clinical practice.
Smart Images

Figure CN120107286A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing technology, and in particular to a polyp segmentation method and device based on a state-aware attention network. Background Art
[0002] Polyps are abnormal growths that form on the mucous membranes of human organs, such as the intestines, nasal cavity, and stomach, and usually appear as small lumps. Polyps often cause a variety of health problems, including bleeding, blockage or pain, and can even develop into cancer. These polyp tissues can be benign, but they can also be a precursor to malignant tumors. Therefore, it is important to detect polyps in gastrointestinal organs and treat them promptly to prevent the condition from worsening.
[0003] Automated detection methods based on deep learning have performed well in polyp detection, but are still greatly limited in clinical practice due to the following issues: First, diseased polyps are very similar to normal tissues in shape, color, and texture (e.g. Figure 1 The lesion area shown in the box is very similar to the normal area pointed by the arrow), which makes it difficult for doctors to accurately identify them during visual inspection and increases the risk of missed detection; secondly, small polyps (such as Figure 2 Polyps (circled), such as adenomatous and hyperplastic polyps, can form on the mucous membranes of the digestive tract, respiratory tract, or other organs and are often difficult to detect during endoscopy or imaging. Summary of the invention
[0004] In view of this, the purpose of an embodiment of the present invention is to provide a polyp segmentation method and device based on a state-aware attention network, which can accurately identify polyp lesions and capture small polyp lesions, improve the efficiency and performance of polyp detection, and solve the problem that the current automated polyp detection method based on deep learning cannot be applied on a large scale in clinical practice.
[0005] In a first aspect, an embodiment of the present invention provides a polyp segmentation method based on a state-aware attention network, comprising:
[0006] Acquire image data of the target area and extract a first feature map;
[0007] Input the first feature map into a state-aware attention network to obtain a second feature map with state-aware attention, wherein the second feature map includes a channel feature map and a spatial feature map; wherein the state-aware attention network includes a channel state transfer attention module and a spatial state reasoning attention module, wherein the channel state transfer attention module is used to model the long-term dependency between the lesion area and the normal area to distinguish between lesion polyps and normal tissues, and output the channel feature map; and the spatial state reasoning attention module is used to focus on target effective information to capture small polyp lesions, and output the spatial feature map;
[0008] The polyp region is segmented in the target region based on the second feature map.
[0009] Optionally, acquiring image data of the target area and extracting the first feature map includes:
[0010] Image data of a target area is acquired, and a first feature map of a preset size is extracted from the image data based on a residual network.
[0011] Optionally, the channel state transfer attention module includes a channel head perception block and a channel state transfer block, the channel head perception block is used to extract a channel feature map from the first feature map and output a third feature map;
[0012] The channel state transfer block removes redundant information in the channel feature map through a selective spatial state model mechanism to capture effective detail features in the channel and output a channel feature map to capture effective detail features in the channel and output a channel feature map.
[0013] Optionally, extracting a channel feature map from the first feature map and outputting a third feature map comprises:
[0014] Feed the first feature map to the linear layer, and generate the first query, the first key, and the first value after channel grouping;
[0015] Then reshape the first key and first value into shape;
[0016] Finally, the first query, the first key, and the first value are used as the input of the multi-head channel attention to model the sequence labeling, and then input into the feedforward neural network to obtain the third feature map with channel feature map.
[0017] Optionally, the redundant information in the channel feature map is removed by a selective spatial state model mechanism to capture effective detail features in the channel, and the output channel feature map includes:
[0018] Redundant information in the channel feature map is removed based on the following method:
[0019]
[0020] Among them, I represents the third feature map with channel feature map, A, B, C, D and is a learnable matrix, is the input holding step vector, n represents the step size, G and L are implemented by the linear layer, H is the hidden state of I, J and K are the learnable matrices obtained from the linear layer, and Y1 is the channel feature map generated after removing redundant information from I.
[0021] Optionally, the spatial state reasoning attention module includes a spatial head perception block and a spatial state reasoning block, and the spatial head perception block is used to obtain valid information from the spatial feature map and output a fourth feature map;
[0022] The spatial state inference block is used to capture fine-grained features from the fourth feature map and output a spatial feature map.
[0023] Optionally, acquiring valid information from the spatial feature map and outputting a fourth feature map includes:
[0024] The first feature map of the input is flattened in the spatial dimension and then fed into the linear layer, where the channels are grouped to produce the second query, the second key, and the second value;
[0025] Reshape the second query into The second key and the second value are reshaped into tensors with the same dimension after being processed by the convolution layer and the layer normalization;
[0026] The reshaped second query is multiplied by the reshaped second key, and the feature map S is generated after processing with the softmax function. The feature map S is multiplied by the reshaped second value, and a fourth feature map carrying valid information obtained from the small lesion area is output. The fourth feature map contains valid information obtained from the small lesion area.
[0027] Optionally, capturing fine-grained features from the fourth feature map and outputting a spatial feature map comprises:
[0028] Capturing fine-grained features from the effective information is based on the following methods:
[0029]
[0030] Among them, I represents the fourth feature map with effective information of the input, A, B, C, D and is a learnable matrix, is the input holding step vector, n represents the step size, G and L are implemented by the linear layer, H is the hidden state of I, J and K are the learnable matrices obtained from the linear layer, and Y2 is the spatial feature map formed from the fine-grained features captured in I.
[0031] Optionally, segmenting the polyp region in the target region based on the second feature map includes:
[0032] Upsampling the second feature map by bilinear interpolation, and then processing it with a skip connection block to generate a fifth feature map;
[0033] After upsampling the fifth feature map using a convolutional layer, a mask of the polyp area is predicted;
[0034] The polyp region is segmented in the target region based on the mask of the polyp region.
[0035] In a second aspect, an embodiment of the present invention provides a polyp segmentation device based on a state-aware attention network, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the computer program running in the following modules of the device:
[0036] A feature extraction module, used to obtain image data of a target area and extract a first feature map;
[0037] A state-aware attention generation module is used to input the first feature map into a state-aware attention network to obtain a second feature map with state-aware attention, wherein the second feature map includes a channel feature map and a spatial feature map; wherein the state-aware attention network includes a channel state transfer attention module and a spatial state reasoning attention module, wherein the channel state transfer attention module is used to model the long-term dependency between the lesion area and the normal area to distinguish between lesion polyps and normal tissues, and output the channel feature map; wherein the spatial state reasoning attention module is used to focus on target effective information to capture small polyp lesions, and output the spatial feature map;
[0038] A polyp segmentation module is used to segment the polyp area in the target area based on the second feature map.
[0039] The embodiments of the present invention include the following beneficial effects: the embodiments of the present invention model the long-term dependency between the diseased area and the normal area based on the channel state transfer attention module to distinguish between diseased polyps and normal tissues; and focus on the target effective information based on the spatial state reasoning attention module to capture small polyp lesions, thereby solving the current restriction problem of large-scale application of automated polyp detection methods based on deep learning in clinical practice. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0041] Figure 1 To detect polyps containing similar lesion areas and normal areas;
[0042] Figure 2 To provide pictures of small polyps that are difficult to find;
[0043] Figure 3is a schematic diagram of the architecture of a spin-adjacent attention network proposed in some embodiments of the present invention;
[0044] Figure 4 is a flow chart of a polyp segmentation method based on a state-aware attention network provided by an embodiment of the present invention;
[0045] Figure 5 is a schematic diagram of the structure of a channel state attention transfer module in some embodiments of the present invention;
[0046] Figure 6 is a schematic diagram of the structure of a spatial state reasoning attention module in some embodiments of the present invention;
[0047] Figure 7 is a schematic diagram of a polyp segmentation device based on a state-aware attention network according to some embodiments of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0049] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0051] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0052] The block diagrams shown in the accompanying drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware charging modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0053] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0054] Polyps are abnormal growths that form on the mucous membrane surfaces of human organs, such as the intestines, nasal cavity, and stomach, and usually appear as small lumps. Polyps always lead to various health problems, including bleeding, blockage or pain, and can even evolve into cancer. These polyp tissues can be benign, but they can also be a precursor to malignant tumors. Therefore, it is necessary to detect polyps in gastrointestinal organs and treat them in time to prevent the condition from worsening.
[0055] In clinical diagnosis, polyps are usually detected through endoscopic examinations by endoscopists, such as gastroscopy or colonoscopy. However, long-term overload work can cause visual fatigue in endoscopists during endoscopic observation. In addition, endoscopists often overlook similar lesions and small polyps, resulting in errors and omissions in diagnostic results. In addition, manual inspections are performed by experienced doctors through visual observation, which is insufficient for real-time medical diagnosis. Therefore, it is necessary to design an automatic tool to assist endoscopists in polyp detection to improve diagnostic efficiency and performance.
[0056] Deep learning-based methods have shown good performance in polyp detection. However, they cannot be applied in large-scale clinical practice due to the following problems. First, diseased polyps are very similar to normal tissues in shape, color, and texture (e.g. Figure 1 The lesion area shown in the box is very similar to the normal area pointed by the arrow), which makes it difficult for doctors to accurately identify them during visual inspection and increases the risk of missing them. Figure 2 Circled), such as adenomatous and hyperplastic polyps, can form on the mucosa of the digestive tract, respiratory tract, or other organs and are often difficult to detect during endoscopy or imaging. These issues make automated polyp detection using deep neural networks still a challenging task.
[0057] In view of the above-mentioned defects of the existing automatic polyp detection method based on deep neural network, the embodiments of the present invention propose a state-aware attention network, which includes a multi-layer cascaded channel state transfer attention module and a spatial state reasoning attention module. The channel state transfer attention module is used to model the long-term dependency between the lesion area and the normal area to distinguish the lesion polyp from the normal tissue, and the spatial state reasoning attention module is used to focus on the target effective information to capture small polyp lesions. In some embodiments of the present invention, the network structure of the state-aware attention network is as follows: Figure 3 shown.
[0058] The present invention also proposes a polyp segmentation method based on a state-aware attention network based on the perception attention network. Figure 4 As shown, the method comprises the following steps:
[0059] S410, acquiring image data of a target area and extracting a first feature map.
[0060] S420, input the first feature map into a state-aware attention network to obtain a second feature map with state-aware attention, wherein the second feature map includes a channel feature map and a spatial feature map; wherein the state-aware attention network includes a channel state transfer attention module and a spatial state reasoning attention module, wherein the channel state transfer attention module is used to model the long-term dependency between the lesion area and the normal area to distinguish between lesion polyps and normal tissues, and outputs a channel feature map; wherein the spatial state reasoning attention module is used to focus on target effective information to capture small polyp lesions, and outputs a spatial feature map. The channel state transfer attention module and the spatial state reasoning attention module are cascaded using skip connections, and output a channel feature map and a spatial feature map, respectively.
[0061] In clinical diagnosis, polyp lesion examination relies on experienced gastroenterologists with rich pathological knowledge. However, gastroenterologists often overlook some polyp lesions that are similar to normal tissues, which further deteriorates the patient's condition. For example, the polyp area is morphologically similar to the surrounding normal area, such as Figure 1 The lesion area indicated by the box is very similar to the normal area indicated by the arrow.
[0062] To solve this problem, the present invention models the long-term dependency between the lesion area and the normal area through the channel state transfer attention (CSTA) module to distinguish the polyp area from similar normal problems.
[0063] In some embodiments of the present invention, formally, the channel state transfer attention module can be expressed as:
[0064] CSTA(X)=CST[CHA(X)];
[0065] Wherein, X represents the first feature map, CHA and CST represent the channel head perception block and the channel state transfer block respectively, CSTA represents the channel state transfer attention module, and CSTA(X) represents the channel feature map. The channel head perception block is used to extract the channel feature map from the input first feature map X to obtain a third feature map with a channel feature map; the channel state transfer block removes redundant information in the channel feature map of the third feature map through a selective spatial state model mechanism to capture effective detail features in the channel, and obtains a channel feature map generated after removing redundant information.
[0066] The network structure of the channel state transfer attention module is as follows Figure 5 shown.
[0067] The channel head perception block can improve the computation and memory efficiency of the standard multi-head attention mechanism, which is formally defined as:
[0068] CHA(X)=FFNs{MHCA[L(X)]};
[0069] Among them, L, MHCA and FFNs represent linear layers, multi-head channel attention and feedforward neural networks respectively, L(X) represents the linear transformation of the first feature map X of the input, and CHA(X) represents the third feature map.
[0070] The specific implementation is:
[0071] First, the first feature map X∈R B×C×W×H Feed to the linear layer, and generate the first query Q, the first key K and the first value V after channel grouping, where B, C, W, H and S are the batch, channel, weight, height and group of the first feature map respectively;
[0072] Then reshape the first key K and the first value V into shape;
[0073] Finally, the first query Q, the first key K and the first value V are used as the input of the multi-head channel attention to model the sequence labeling, and then input into the feedforward neural network to obtain the feature map with channel feature map.
[0074] The channel state transfer block removes redundant information based on the selective spatial state model mechanism and retains relevant valid information during the learning process. Therefore, the channel state transfer block can further capture the effective detail features in the channel, which can be formally expressed as:
[0075]
[0076] Among them, I represents the third feature map with channel feature map, A, B, C, D and is a learnable matrix, is the input holding step vector, n represents the step size, G and L are implemented by the linear layer, H is the hidden state of I, J and K are the learnable matrices obtained from the linear layer, and Y is the channel feature map generated after removing redundant information from I.
[0077] The detailed process is as follows:
[0078] First: From Initialize the learnable matrix The input third feature map I is multiplied with the parameter matrices G and L respectively to obtain batch B and channel C, where G and L are obtained by the linear layer and D is initialized to a matrix of all 1s.
[0079] Second: After the third feature map I is multiplied by the matrices J and K, the learnable matrix is obtained after the Softplus function is processed. The Softplus function is expressed as:
[0080] Softplus(X)=In(1+e X );
[0081] Similarly, the learnable matrices J and K are obtained from the linear layers.
[0082] Third: Learnable Matrix The exponential result after multiplication with A, and the matrix The product of B and I is fed into the pscan function at the same time to solve the hidden state H of the spatial feature map I.
[0083] Fourth: Sum the multiplication results of C and H and the multiplication results of D and I to obtain the channel feature map Y1.
[0084] Another challenge in clinical polyp examination is that the lesions are very small, which can be easily overlooked by gastroenterologists. Figure 2 In this case, the attention mechanism should be guided by a specific transformation mechanism to focus on small target areas. Inspired by these observations, the present invention uses a spatial state reasoning attention module to capture small polyp lesions by focusing on the effective information in the feature map, such as Figure 6 shown.
[0085] The spatial state reasoning attention module can be formally expressed as:
[0086] SSIA(X)=STI[SHA(X)];
[0087] Among them, SSLA stands for the spatial state reasoning attention module, SSLA(X) represents the spatial feature map output by the spatial state reasoning attention module acting on the first feature map X; SHA represents the spatial head perception block of the spatial state reasoning attention module, which is used to obtain effective information from small lesions. STI represents the spatial state reasoning block of the spatial state reasoning attention module, which is used to capture fine-grained features.
[0088] Spatial Header Aware Block:
[0089] The traditional multi-head self-attention (CMSA) mechanism can be described as:
[0090]
[0091] Among them, CMSA(Q,K,V) represents the traditional multi-head attention mechanism acting on Q, K, V, where Q, K and V represent the query, key and value sequence graphs generated from linear or convolutional layers, respectively. However, the computational complexity of CMSA is too expensive for polyp detection in clinical diagnosis. In order to achieve fast inference efficiency and high performance, the empty head perception block is designed to capture details from the first feature map of the input. The formal definition is:
[0092] SHA(X)=MGRA{LCG[SF(X)]};
[0093] Among them, SF(X) represents the spatial flattening of the first feature map X, SF, LCG, MGRA represent spatial flattening, linear layer and channel grouping, and multiple groups of reshape attention, respectively, and SHA(X) represents the fourth feature map. Specifically:
[0094] The first feature map X of the input is flattened in the spatial dimension and then fed into the linear layer. After the channels are grouped, the second query Q, the second key K and the second value V are generated; where X∈R B×C×W×H ,
[0095] Reshape the second query into The second key and the second value are reshaped into tensors with the same dimension after being processed by the convolution layer and the layer normalization;
[0096] The reshaped second query is multiplied by the reshaped second key, and the feature map S is generated after processing with the softmax function. The feature map S is multiplied by the reshaped second value, and a fourth feature map carrying valid information obtained from the small lesion area is output. The fourth feature map contains valid information obtained from the small lesion area.
[0097] The present invention is based on the spatial state inference (STI) block, which uses a selective spatial state mechanism to select relevant information in the feature extraction process to capture fine-grained features from the spatial feature map.
[0098] The STI block can be represented as:
[0099]
[0100] Where I represents the fourth feature map of the input that carries the effective information obtained from the small lesion area, A, B, C, D and is a learnable matrix, is the input holding step vector, n represents the step size, G and L are implemented by the linear layer, H is the hidden state of I, J and K are the learnable matrices obtained from the linear layer, and Y is the spatial feature map formed from the fine-grained features captured in I.
[0101] The detailed process is as follows:
[0102] First: the learnable matrix A is initialized to [1,...,n], while B and C are obtained by multiplying the original input feature map I by vectors G and L respectively, where G and L are implemented by linear layers, and D is initialized to a full 1 vector.
[0103] Second, by multiplying the input feature map I with the vectors J and K and then processing it with the Softplus function, a learnable matrix ▽ is generated, and the learnable matrices J and K are obtained in the same way.
[0104] Third, the result of multiplying the learnable matrix ▽ and A and then processed by the exponential function, as well as the result of multiplying the matrix ▽ with B and I, are input into the pscan function at the same time to solve the hidden state H of the input feature map I.
[0105] Finally, the multiplication result of C and H is added to the multiplication result of matrices D and I to obtain the final spatial feature map.
[0106] S430: Segment a polyp region in the target region based on the second feature map.
[0107] The present invention generates a segmentation mask for polyp region prediction based on the second feature map of the target region, and segments the polyp region from the target region based on the segmentation mask.
[0108] The segmentation mask module aims to generate the final prediction of the polyp region from the image, and it consists of a simple skip block that can be described as:
[0109] First, a second feature map is obtained, where the second feature map includes a channel feature map and a spatial feature map;
[0110] Upsampling the second feature map by bilinear interpolation, and then processing it with a skip connection block to generate a fifth feature map;
[0111] Secondly, the fifth feature map obtained by upsampling is processed by the skip connection block, and the skip connection block processing is expressed as:
[0112] skip(Z)=Z+C(Z);
[0113] Among them, Z represents the second feature map after upsampling, C represents the convolution layer, and Skip(Z) represents the fifth feature map generated after performing skip connection processing on the second feature map after upsampling.
[0114] Finally, the fifth feature map will be further upsampled and processed by the convolutional layer to predict the final mask of the polyp region.
[0115] One embodiment of the present invention is Figure 4 The method described in S410-S430 (abbreviated as MambaFormer) was tested against other detection methods on two public polyp datasets (Kvasir dataset and ClinicDB dataset). The Kvasir dataset contains 1,000 images with corresponding polyp masks, and ClinicDB consists of 612 frames of images from 29 different subjects with pixel-level data labels. In the test, Kvasir and CVCClnicDB datasets were compared, where the Kvasir dataset was divided into a training set and a test set, with 880 and 120 images respectively, while the CVCClnicDB dataset had 550 and 62 images for training and testing.
[0116] Table 1: Quantitative results on the KVASIR dataset
[0117]
[0118]
[0119] Table 2: Quantitative results on the CVC-CLINICDB dataset
[0120] method years mDice(%) MIoU(%) Recovery rate (%) Early stage (%) U-Net 2015 87.2 80.4 86.8 91.7 U-Net++ 2018 88.1 81.9 91 88.5 Attention U-Net 2018 89 82.7 88.7 90.9 Swin-Unet 2022 90.6 84.9 91.8 90.7 SegFormer 2021 91.1 86 94.2 91.1 HarDNet-MSEG 2021 91.8 86.4 91.2 94.5 MCTrans 2021 92.3 - - - TransUnet 2021 92.3 86.9 94.2 91.7 DoubleU-Net 2020 92.4 86.1 84.6 95.9 FANet 2022 93.6 89.4 93.4 94 DS-TransUNet 2022 94.2 89.4 95 93.7 MambaFormer - 97.2 94.7 96.9 97.6
[0121] As can be seen from Tables 1 and 2, MambaFormer has the best polyp detection performance.
[0122] Figure 7 is a schematic diagram of a polyp segmentation device based on a state-aware attention network according to some embodiments of the present invention. Figure 7As shown, the polyp segmentation device 700 based on the state-aware attention network includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to run in the modules of the following devices:
[0123] A feature extraction module 710 is used to obtain image data of a target area and extract a first feature map;
[0124] A state-aware attention generation module 720 is used to input the first feature map into a state-aware attention network to obtain a second feature map with state-aware attention, wherein the state-aware attention network includes a multi-layer cascaded channel state transfer attention module and a spatial state reasoning attention module, the channel state transfer attention module is used to model the long-term dependency between the lesion and the normal area to distinguish between the lesion polyp and the normal tissue, and the spatial state reasoning attention module is used to focus on the target effective information to capture the small polyp lesion;
[0125] The polyp segmentation module 730 is configured to segment the polyp region in the target region based on the second feature map.
[0126] In summary, the polyp segmentation method and device based on the state-aware attention network provided by the embodiments of the present invention distinguish diseased polyps from normal tissues by modeling the long-term dependency between the diseased area and the normal area based on the channel state transfer attention module; and focusing on the target effective information based on the spatial state reasoning attention module to capture small polyp lesions, thereby solving the restriction problem of large-scale application of automated polyp detection methods based on deep learning in clinical practice at this stage.
[0127] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the charging modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0128] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional charging modules / units in the systems and devices may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0129] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0130] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0131] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0132] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0133] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0134] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0135] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A polyp segmentation method based on state-aware attention network, characterized in that: include: Acquire image data of the target area and extract a first feature map; Input the first feature map into a state-aware attention network to obtain a second feature map with state-aware attention, wherein the second feature map includes a channel feature map and a spatial feature map; wherein the state-aware attention network includes a channel state transfer attention module and a spatial state reasoning attention module, wherein the channel state transfer attention module is used to model the long-term dependency between the lesion area and the normal area to distinguish between lesion polyps and normal tissues, and output the channel feature map; and the spatial state reasoning attention module is used to focus on target effective information to capture small polyp lesions, and output the spatial feature map; The polyp region is segmented in the target region based on the second feature map.
2. The method according to claim 1, characterized in that: The acquiring image data of the target area and extracting the first feature map comprises: Image data of a target area is acquired, and a first feature map of a preset size is extracted from the image data based on a residual network.
3. The method according to claim 1, characterized in that: The channel state transfer attention module includes a channel head perception block and a channel state transfer block, wherein the channel head perception block is used to extract a channel feature map from the first feature map and output a third feature map; The channel state transfer block removes redundant information in the channel feature map through a selective spatial state model mechanism to capture effective detail features in the channel and output a channel feature map to capture effective detail features in the channel and output a channel feature map.
4. The method according to claim 3, characterized in that: The extracting a channel feature map from the first feature map and outputting a third feature map comprises: Feed the first feature map to the linear layer, and generate the first query, the first key, and the first value after channel grouping; Then reshape the first key and first value into shape; Finally, the first query, the first key, and the first value are used as the input of the multi-head channel attention to model the sequence labeling, and then input into the feedforward neural network to obtain the third feature map with channel feature map.
5. The method according to claim 3, characterized in that: The redundant information in the channel feature map is removed by the selective spatial state model mechanism to capture effective detail features in the channel, and the output channel feature map includes: Redundant information in the channel feature map is removed based on the following method: Among them, I represents the third feature map with channel feature map, A, B, C, D and is a learnable matrix, is the input holding step vector, n represents the step size, G and L are implemented by the linear layer, H is the hidden state of I, J and K are the learnable matrices obtained from the linear layer, and Y1 is the channel feature map generated after removing redundant information from I.
6. The method according to claim 1, characterized in that: The spatial state reasoning attention module includes a spatial head perception block and a spatial state reasoning block, wherein the spatial head perception block is used to obtain valid information from the spatial feature map and output a fourth feature map; The spatial state inference block is used to capture fine-grained features from the fourth feature map and output a spatial feature map.
7. The method according to claim 6, characterized in that: The obtaining effective information from the spatial feature map and outputting the fourth feature map comprises: The first feature map of the input is flattened in the spatial dimension and then fed into the linear layer, where the channels are grouped to produce the second query, the second key, and the second value; Reshape the second query into The second key and the second value are reshaped into tensors with the same dimension after being processed by the convolution layer and the layer normalization; The reshaped second query is multiplied by the reshaped second key, and the feature map S is generated after processing with the softmax function. The feature map S is multiplied by the reshaped second value, and a fourth feature map carrying valid information obtained from the small lesion area is output. The fourth feature map contains valid information obtained from the small lesion area.
8. The method according to claim 6, characterized in that: The capturing of fine-grained features from the fourth feature map and the outputting of the spatial feature map include: Capturing fine-grained features from the effective information is based on the following methods: Among them, I represents the fourth feature map with effective information of the input, A, B, C, D and is a learnable matrix, is the input holding step vector, n represents the step size, G and L are implemented by the linear layer, H is the hidden state of I, J and K are the learnable matrices obtained from the linear layer, and Y2 is the spatial feature map formed from the fine-grained features captured in I.
9. The method according to claim 1, characterized in that: The segmenting of the polyp region in the target region based on the second feature map comprises: Upsampling the second feature map by bilinear interpolation, and then processing it with a skip connection block to generate a fifth feature map; After upsampling the fifth feature map using a convolutional layer, a mask of the polyp area is predicted; The polyp region is segmented in the target region based on the mask of the polyp region.
10. A polyp segmentation device based on a state-aware attention network, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to run in the following modules of the device: A feature extraction module, used to obtain image data of a target area and extract a first feature map; A state-aware attention generation module is used to input the first feature map into a state-aware attention network to obtain a second feature map with state-aware attention, wherein the second feature map includes a channel feature map and a spatial feature map; wherein the state-aware attention network includes a channel state transfer attention module and a spatial state reasoning attention module, wherein the channel state transfer attention module is used to model the long-term dependency between the lesion area and the normal area to distinguish between lesion polyps and normal tissues, and output the channel feature map; wherein the spatial state reasoning attention module is used to focus on target effective information to capture small polyp lesions, and output the spatial feature map; A polyp segmentation module is used to segment the polyp area in the target area based on the second feature map.