Skin lesion image segmentation method, system, device and storage medium
By extracting the multi-level feature context information of skin lesion images and performing gating fusion decoder fusion, combining with the shape guide flow module, the lesion boundaries are gradually explored, and the problems of different sizes of lesion areas and irregular shapes and blurred boundaries in skin lesion image segmentation are solved, achieving more accurate segmentation results.
Patent Information
- Application Number
- CN202210751558.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-06-29
AI Technical Summary
The prior art is difficult to effectively deal with different sizes and irregular shapes of the lesion area in the segmentation of skin lesions, resulting in inaccurate target positioning and ambiguity between the lesion area and the background.
A skin lesion image segmentation method is proposed, which generates an initial guide feature map by extracting the context information of multi-level features and using a gated fusion decoder. Then, the lesion boundary is gradually tapped through the shape guide flow module and the shape of the lesion area is highlighted, and the segmentation result is finally obtained through gated convolutional fusion.
It effectively solves the problem of inaccurate target positioning caused by different sizes and irregular shapes of the lesion area, as well as the fuzzy problem between the lesion area and the background, and improves the accuracy of the segmentation results and the network training effect.
Smart Images

Figure CN114972324B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, system, device and storage medium for skin lesion image segmentation. Background Art
[0002] Automatic medical image segmentation plays a vital role in computer-aided diagnosis systems. It helps doctors develop treatment plans and monitor disease progression for patients with skin diseases and cancers, and greatly helps dermatologists improve the accuracy of analysis. In order to screen and diagnose skin diseases, dermatologists visually inspect and analyze skin lesions and surrounding tissues through dermoscopic images. However, manual inspection of dermoscopic images requires doctors with professional skills, and it is difficult to design a general method that can be applied to different tasks. Therefore, in order to detect and analyze skin diseases more effectively and accurately, a large number of studies are devoted to applying CAD technology to skin disease diagnosis. Due to the challenging problems of dermoscopic image datasets, such as the diversity of lesions, irregular shapes, and blurred boundaries, traditional algorithms have poor segmentation performance. In contrast, deep learning models can adaptively learn different features of dermoscopic images, and their segmentation performance is better than traditional algorithms. Therefore, deep learning has a dominant position in medical image segmentation fields such as skin lesion segmentation.
[0003] Many researchers have proposed classic encoder-decoder network structures for medical image segmentation tasks and achieved competitive results. In this network structure, the encoder is usually used to extract image features, while the decoder is usually used to restore the extracted features to the original image size and output the final prediction results.
[0004] Although existing methods have improved the performance of skin lesion segmentation, issues such as the color, texture, shape, size, appearance, or a combination of these features that affect segmentation performance still plague researchers, especially for the large diversity and irregular shapes of lesion areas, and the ambiguity between lesion areas and backgrounds. Summary of the invention
[0005] The present application aims to at least solve the technical problems existing in the prior art. To this end, the present application proposes a skin lesion image segmentation method, system and electronic device, which can well solve the problem of inaccurate target positioning caused by different sizes and irregularities of lesion areas in skin lesion images and the ambiguity between the lesion area and the background.
[0006] According to a first aspect of the present application, a skin lesion image segmentation method is provided, the skin lesion image segmentation method comprising:
[0007] Acquire images of skin lesions;
[0008] Extracting shallow features and multi-level features of the skin lesion image, and extracting context information corresponding to each level of features;
[0009] Inputting the context information of the multi-level features into a preset gated fusion decoder, so that the gated fusion decoder adaptively selects complementary information from the context information of the multi-level features, and fuses the context information of the multi-level features through an addition and gating mechanism to obtain an initial guided feature map output by the gated fusion decoder;
[0010] Constructing a shape-guided flow module, inputting the initial guided feature map and the context information of the multi-level features into the shape-guided flow module, so that the shape-guided flow module gradually mines the lesion boundary in the feature map and highlights the shape of the lesion area, and obtaining a final guided feature map output by the shape-guided flow module;
[0011] The final guided feature map and the shallow features are gated convolutionally fused to obtain a segmentation result of the skin lesion image.
[0012] According to the embodiments of the present application, at least the following technical effects are achieved:
[0013] This method first extracts the context information of the multi-level features of the skin lesion image; then the gated fusion decoder adaptively selects complementary information from the context information of the multi-level features, and fuses the context information of the multi-level features through the addition and gating mechanism to obtain the initial guided feature map. The gated fusion decoder can effectively fuse multi-scale detail information and semantic information to solve the problem of inaccurate target positioning due to the different sizes and irregularities of the lesion area; then the shape guided flow module gradually mines the lesion boundary in the feature map and highlights the shape of the lesion area to solve the boundary fuzziness problem; finally, the final guided feature map output by the shape guided flow module and the shallow features of the skin lesion image are gated and convoluted to obtain the final output, which can reduce noise interference and enhance the network training effect. Combined with experiments, it is proved that this method can well solve the problem of inaccurate target positioning due to the different sizes and irregularities of the lesion areas in the skin lesion image and the ambiguity between the lesion area and the background.
[0014] According to some embodiments of the present application, extracting context information corresponding to each level of features includes:
[0015] Construct a context feature extraction module, the context feature extraction module includes a receptive field module and a residual multi-core pooling module connected in sequence, the receptive field module includes a jump connection layer and a multi-branch convolution layer, wherein each branch convolution layer includes a 1x1 convolution layer, a 1xn convolution layer, an nx1 convolution layer and a hole convolution layer with a specific hole rate connected in sequence, and the output result of the jump connection layer and the output result of the multi-branch convolution layer are spliced to obtain the output result of the receptive field module; the residual multi-core pooling module includes four pools of different sizes A convolution layer, four 1×1 convolution layers and a bilinear interpolation layer, wherein the four pooling layers of different sizes are used to extract four feature maps of different sizes from the feature map input to the residual multi-core pooling module, respectively, the four 1×1 convolution layers are used to reduce the number of channels of the four feature maps of different sizes, the bilinear interpolation layer is used to upsample the feature maps output by the four 1×1 convolution layers, and the feature map input to the residual multi-core pooling module is spliced with the feature map output by the bilinear interpolation layer to obtain the output result of the residual multi-core pooling module;
[0016] Each level feature of the multi-level features is input into the context feature extraction module to obtain context information of each level feature in the context feature extraction module.
[0017] According to some embodiments of the present application, the expression of the gated fusion decoder includes:
[0018]
[0019] Among them, σ represents the sigmoid activation function, Indicates pixel-level multiplication in the channel, fi i represents the i-th level feature in the multi-level features, S g represents the initial guided feature map.
[0020] According to some embodiments of the present application, the multi-level features are N level features, wherein the first level feature is the lowest level feature among the N level features, and the level features from the first level feature to the Nth level feature are increased in sequence; the shape guidance flow module includes N channel reverse attention modules and N gated convolution fusion modules; the shape guidance flow module gradually mines the lesion boundary in the feature map and highlights the shape of the lesion area, and obtains the final guidance feature map output by the shape guidance flow module, including:
[0021] Using the initial guide feature map as the first guide feature map, and inputting the first guide feature map and the first hierarchical feature into the first channel reverse attention module, so that the first channel reverse attention module removes the foreground object by using reverse attention, and outputs the first attention feature map;
[0022] Inputting the first attention feature map and the first guide feature map into the first gated convolution fusion module to obtain a second guide feature map output by the first gated convolution fusion module;
[0023] Inputting the second guide feature map and the second hierarchical features into the second channel reverse attention module, so that the second channel reverse attention module removes the foreground object by using reverse attention and outputs a second attention feature map;
[0024] Inputting the second attention feature map and the second guide feature map into the second gated convolution fusion module to obtain a third guide feature map output by the second gated convolution fusion module;
[0025] And so on, until the Nth guide feature map and the Nth hierarchical feature are input into the Nth channel reverse attention module, so that the Nth channel reverse attention module removes the foreground object by using reverse attention and outputs the Nth attention feature map;
[0026] Inputting the Nth attention feature map and the Nth guide feature map into the Nth gated convolution fusion module to obtain a final guide feature map output by the Nth gated convolution fusion module;
[0027] Among them, the expression of the gated convolution fusion module includes:
[0028]
[0029] G i =((F′ i ⊙α i )+F′ i ) T ω i
[0030] Among them, α i represents the weight value, σ represents the sigmoid function, Conv 1×1 represents 1×1 convolution, F′ i represents the i-th attention feature map, S′ i represents the i-th guided feature map, i is an integer from 1 to N, ω i Represents the weight coefficient, G iRepresents the guided feature map output by the gated convolutional fusion module.
[0031] According to some embodiments of the present application, the channel reverse attention module includes a reverse attention module and a multi-frequency channel attention module connected in sequence, the reverse attention module uses reverse attention to eliminate foreground objects, and the multi-frequency channel attention module is used to process each channel of the input feature map according to discrete cosine transform to obtain different frequency components.
[0032] According to some embodiments of the present application, after the i-th guided feature map output by the i-th gated convolution fusion module, the method further includes:
[0033] The i-th guided feature map is deeply supervised.
[0034] According to some embodiments of the present application, performing gated convolution fusion on the final guided feature map and the shallow features to obtain a segmentation result of the skin lesion image includes:
[0035] Constructing a gated convolution fusion module;
[0036] The final guided feature map and the shallow features are input into the gated convolution fusion module to obtain the segmentation result of the skin lesion image output by the gated convolution fusion module.
[0037] According to a second aspect of the present application, a skin lesion image segmentation system is provided, wherein the skin lesion image segmentation system comprises:
[0038] An image acquisition unit, used for acquiring skin lesion images;
[0039] A feature extraction unit, used to extract shallow features and multi-level features of the skin lesion image, and extract context information corresponding to each level of features;
[0040] A multi-scale feature fusion unit, used for inputting the context information of the multi-level features into a preset gated fusion decoder, so that the gated fusion decoder adaptively selects complementary information from the context information of the multi-level features, and fuses the context information of the multi-level features through an addition and gating mechanism to obtain an initial guided feature map output by the gated fusion decoder;
[0041] A feature enhancement unit is used to construct a shape guidance flow module, and input the initial guidance feature map and the context information of the multi-level features into the shape guidance flow module, so that the shape guidance flow module gradually mines the lesion boundary in the feature map and highlights the shape of the lesion area, and obtains the final guidance feature map output by the shape guidance flow module;
[0042] The segmentation result acquisition unit is used to perform gated convolution fusion on the final guided feature map and the shallow features to obtain the segmentation result of the skin lesion image.
[0043] Since the skin lesion image segmentation system adopts all the technical solutions of the skin lesion image segmentation method of the above embodiment, it at least has all the beneficial effects brought by the technical solutions of the above embodiment.
[0044] According to an electronic device of an embodiment of the third aspect of the present application, the device comprises: at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute a skin lesion image segmentation method as described above.
[0045] Since the electronic device can implement all the technical solutions of the skin lesion image segmentation method of the above embodiment, it at least has all the beneficial effects brought by the technical solutions of the above embodiment.
[0046] According to the computer-readable storage medium of the fourth aspect embodiment of the present application, the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute a skin lesion image segmentation method as described above.
[0047] Since the computer-readable storage medium adopts all the technical solutions of the skin lesion image segmentation method of the above embodiment, it at least has all the beneficial effects brought by the technical solutions of the above embodiment.
[0048] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent from the description, or may be understood by practicing the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0050] Figure 1 is a schematic diagram of the structure of a GFA-Net network provided in an embodiment of the present application;
[0051] Figure 2 It is a flowchart of a skin lesion image segmentation method provided in one embodiment of the present application;
[0052] Figure 3 yes Figure 2 Schematic diagram of the process of step S200;
[0053] Figure 4is a structural diagram of a receptive field module provided in an embodiment of the present application;
[0054] Figure 5 It is a structural diagram of a residual multi-core pooling module provided in an embodiment of the present application;
[0055] Figure 6 yes Figure 2 Schematic diagram of the process of step S400;
[0056] Figure 7 It is a schematic diagram of the structure of a gated convolution fusion module provided in an embodiment of the present application;
[0057] Figure 8 is a schematic diagram of the structure of a gated fusion decoder provided in an embodiment of the present application;
[0058] Fig. 9 yes Figure 2 Schematic diagram of the process of step S500;
[0059] Fig.10 It is a structural schematic diagram of a skin lesion image segmentation system provided by an embodiment of the present application;
[0060] Fig.11 This is an embodiment of the present application that provides visualization of segmentation results of different networks on the ISIC 2018 skin lesion dataset. DETAILED DESCRIPTION
[0061] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.
[0062] In the description of this application, if there is a description of 1, 2, etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0063] Automatic medical image segmentation plays a vital role in computer-aided diagnosis systems. It helps doctors develop treatment plans and monitor disease progression for patients with skin diseases and cancers, and greatly helps dermatologists improve the accuracy of analysis. In order to screen and diagnose skin diseases, dermatologists visually inspect and analyze skin lesions and surrounding tissues through dermoscopic images. However, manual inspection of dermoscopic images requires doctors with professional skills, and it is difficult to design a general method that can be applied to different tasks. Therefore, in order to detect and analyze skin diseases more effectively and accurately, a large number of studies are devoted to applying CAD technology to skin disease diagnosis. Due to the challenging problems of dermoscopic image datasets, such as the diversity of lesions, irregular shapes, and blurred boundaries, traditional algorithms have poor segmentation performance. In contrast, deep learning models can adaptively learn different features of dermoscopic images, and their segmentation performance is better than traditional algorithms. Therefore, deep learning has a dominant position in medical image segmentation fields such as skin lesion segmentation.
[0064] Many researchers have proposed the classic encoder-decoder network structure and applied it to the medical image segmentation task, and achieved competitive results. In this network structure, the encoder is usually used to extract image features, while the decoder is usually used to restore the extracted features to the original image size and output the final prediction results. Although the existing methods have improved the performance of skin lesion segmentation, problems such as the color, texture, shape, size, appearance, or a combination of these features that affect the segmentation performance still plague researchers, especially for the large diversity and irregular shapes of lesion areas, and the ambiguity between lesion areas and backgrounds.
[0065] In order to solve the above technical defects, refer to Figure 1 and Figure 2 An embodiment of the present application provides a skin lesion image segmentation method, the method comprising the following steps S100 to S500:
[0066] Step S100: Acquire a skin lesion image.
[0067] Step S200: extracting shallow features and multi-level features of the skin lesion image, and extracting context information corresponding to each level of features.
[0068] Step S300, input the context information of the multi-level features into a preset gated fusion decoder, so that the gated fusion decoder adaptively selects complementary information from the context information of the multi-level features, and fuses the context information of the multi-level features through addition and gating mechanisms to obtain an initial guided feature map output by the gated fusion decoder.
[0069] Step S400, construct a shape guidance flow module, input the initial guidance feature map and the context information of the multi-level features into the shape guidance flow module, so that the shape guidance flow module gradually mines the lesion boundary in the feature map and highlights the shape of the lesion area, and obtains the final guidance feature map output by the shape guidance flow module.
[0070] Step S500: Perform gated convolution fusion on the final guided feature map and the shallow features to obtain a segmentation result of the skin lesion image.
[0071] This method first extracts the context information of the multi-level features of the skin lesion image; then the gated fusion decoder adaptively selects complementary information from the context information of the multi-level features, and fuses the context information of the multi-level features through the addition and gating mechanism to obtain the initial guided feature map. The gated fusion decoder can effectively fuse multi-scale detail information and semantic information to solve the problem of inaccurate target positioning due to the different sizes and irregularities of the lesion area; then the shape guided flow module gradually mines the lesion boundary in the feature map and highlights the shape of the lesion area to solve the boundary fuzziness problem; finally, the final guided feature map output by the shape guided flow module and the shallow features of the skin lesion image are gated and convoluted to obtain the final output, which can reduce noise interference and enhance the network training effect. Combined with experiments, it is proved that this method can well solve the problem of inaccurate target positioning due to the different sizes and irregularities of the lesion areas in the skin lesion image and the ambiguity between the lesion area and the background.
[0072] Reference Figure 3 In step S200 of some embodiments, extracting shallow features and multi-level features of the skin lesion image includes the following steps S201 and S202:
[0073] Step S201, obtain the HarDNet-68 network.
[0074] Step S202: input the skin lesion image into the HarDNet-68 network, so that the HarDNet-68 network extracts shallow features and multi-level features of the skin lesion image.
[0075] The shallow features in step S202 are as follows Figure 1 F 2 As shown in Figure 1 F 3 ,F 4 ,F 5 shown.
[0076] In some embodiments, extracting context information corresponding to each level of features in step S202 includes the following steps S2021 and S2022:
[0077] Step S2021, construct a context feature extraction module (CEM). The context feature extraction module is mainly used to solve the problem of large diversity and irregular shapes of lesion areas in skin lesion images, and extracts context information of three different levels of features, low, medium and high, and features of objects of different sizes.
[0078] The context feature extraction module includes a receptive field block (RFB) and a residual multi-kernel pooling (RMP) module connected in sequence. Figure 4 As shown in the figure, the receptive field module includes a skip connection layer and a multi-branch convolution layer, wherein each branch convolution layer includes a 1x1 convolution layer (used to reduce the number of channels of the feature map), a 1xn convolution layer, an nx1 convolution layer, and a dilated convolution layer with a specific dilation rate, which are connected in sequence. The output result of the skip connection layer and the output result of the multi-branch convolution layer are concatenated to obtain the output result of the receptive field module.
[0079] like Figure 5 As shown in Figure 2, the residual multi-core pooling module uses multiple effective fields of view to identify skin lesions of different sizes. The residual multi-core pooling module includes four pooling layers of different sizes ( Figure 5 In the figure, 2×2, 3×3, 5×5 and 6×6 are used), four 1×1 convolutional layers and a bilinear interpolation layer. Four pooling layers of different sizes are used to extract four feature maps of different sizes from the feature map input to the residual multi-core pooling module, respectively, to extract global context information. In order to reduce the amount of calculation and the dimension of the weight, four 1×1 convolutional layers are used to reduce the number of channels of the four feature maps of different sizes. The bilinear interpolation layer is used to upsample the feature maps output by the four 1×1 convolutional layers, and the feature map input to the residual multi-core pooling module is concatenated with the feature map output by the bilinear interpolation layer to obtain the output result of the residual multi-core pooling module.
[0080] Step S2022: Input each level of features in the multi-level features into the context feature extraction module to obtain context information of each level of features in the context feature extraction module.
[0081] In step S300 of some embodiments, a gated fusion decoder is constructed, and the fusion of context information of multi-level features is achieved through the constructed gated fusion decoder. Specifically:
[0082] Reference Figure 8 The gated fusion decoder (GFD) of this embodiment uses a gating mechanism to measure low, medium and high (such as Figure 1 F3 ,F 4 ,F 5 ) The context information in the three different levels of features and selectively fuse the feature information at different levels. Specifically: The three different levels of features, low, medium and high, are first passed through the context feature extraction module, and then the context information of the three different levels of features output by the context feature extraction module (such as Figure 1 f 3 f 4 ,f 5 ) is used as the input of the gated fusion decoder. For the i-th level feature of the input, the gated fusion decoder can adaptively select complementary information from the rest of the input features, and then fuse these features through addition and gating mechanisms:
[0083]
[0084] Among them, σ represents the sigmoid activation function, Indicates pixel-level multiplication in the channel, f i Represents the i-th level feature in the multi-level feature. The gated fusion decoder can effectively fuse the feature information of multiple levels, avoiding the introduction of too much redundant information, thereby retaining rich details and semantic information.
[0085] Reference Figure 6 In step S400 of some embodiments, a shape guidance stream module (Shape Guidance Stream) is constructed, and the shape guidance stream module includes N channel reverse attention modules (CRA) and N gated convolutional fusion modules (GCF). The shape guidance stream module gradually mines the lesion boundary in the feature map and highlights the shape of the lesion area, and outputs the final guidance feature map, including the following steps S401 to S406:
[0086] Step S401: Use the initial guided feature map as the first guided feature map, and input the first guided feature map and the first hierarchical feature into the first channel reverse attention module, so that the first channel reverse attention module removes the foreground object by using reverse attention and outputs the first attention feature map.
[0087] Step S402: input the first attention feature map and the first guide feature map into the first gated convolution fusion module to obtain the second guide feature map output by the first gated convolution fusion module.
[0088] Step S403: input the second guide feature map and the second hierarchical features into the second channel reverse attention module, so that the second channel reverse attention module removes the foreground object by using reverse attention and outputs the second attention feature map.
[0089] Step S404: input the second attention feature map and the second guide feature map into the second gated convolution fusion module to obtain the third guide feature map output by the second gated convolution fusion module.
[0090] Step S405, and so on, until the Nth guide feature map and the Nth hierarchical feature are input into the Nth channel reverse attention module, so that the Nth channel reverse attention module removes the foreground object by using reverse attention and outputs the Nth attention feature map.
[0091] Step S406: input the Nth attention feature map and the Nth guide feature map into the Nth gated convolution fusion module to obtain the final guide feature map output by the Nth gated convolution fusion module.
[0092] Among them, the expression of the gated convolution fusion module includes:
[0093]
[0094] G i =((F′ i ⊙α i )+F′ i ) T ω i
[0095] Among them, α i represents the weight value, σ represents the sigmoid function, Conv 1×1 represents 1×1 convolution, F′ i represents the i-th attention feature map, S′ i represents the i-th guided feature map, i is an integer from 1 to N, ω i Represents the weight coefficient, G i represents the guided feature map output by the gated convolutional fusion module, Represents feature map concatenation.
[0096] like Figure 1 As shown, the shape guided flow module includes three channel reverse attention modules (CRA) and three gated convolution fusion modules (GCF). In step S401 and step S402, the first guided feature map (S g ) and the first attention feature map (F 5) is input into the first channel reverse attention module. Different from the existing model that fuses all hierarchical features together, the aforementioned gated fusion decoder adaptively learns and fuses the low, medium and high level features of the encoder, and then outputs the initial guided feature map S g ,Since the gated fusion decoder can only roughly capture the approximate location of the skin lesion area and has no structural details, ,this embodiment removes the foreground objects by using reverse attention, ,so the skin lesion area and detail information can be gradually mined out the skin lesion area ,and detail information more accurately.
[0097] Reference Figure 1 In some embodiments, the channel reverse attention module is composed of a reverse attention module and a multi-frequency channel attention module. The reverse attention module uses reverse attention to remove foreground objects. In order to highlight the importance of each feature channel and better capture rich information, frequency channel attention is applied after the reverse attention. The multi-frequency channel attention module uses discrete cosine transform to process each channel of the input feature to obtain different frequency components, thereby solving the problem of information loss and lack of diversity of input features due to global average pooling.
[0098] It should be noted that the same applies to steps S403 and S404, and steps S405 and S406, which will not be described in detail here.
[0099] Reference Figure 1 In some embodiments, after the i-th guided feature map is output by the i-th gated convolution fusion module, deep supervision is applied to the i-th guided feature map. Specifically:
[0100] First, the three features S′ in deep supervision 4 ,S′ 5 ,S′ g Get the reverse attention weights respectively Then compare it with the three characteristics of low, medium and high 3 ,F 4 ,F 5 Multiply pixel by pixel, and finally get F′ through frequency attention i :
[0101]
[0102]
[0103] Among them, σ represents the sigmoid activation function, Indicates pixel-level multiplication in the channel, FCA i represents the i-th frequency attention.
[0104] Figure 7It is a gated convolution fusion module (GCF). In some embodiments, the attention feature map F′ output by the channel reverse attention module is firstly i S′ with the corresponding feature size obtained through the previous gated convolution fusion module and bilinear interpolation (upsampling or downsampling) i Splicing; then passing through the normalized 1x1 convolution Conv1x1 and sigmoid function σ, a weight value (attention map) is obtained. Then, through the residual connection, α i With input feature F′ i After doing pixel-level dot multiplication, we add them together and finally perform a convolution on the channel to perform weighted ω. i Then output feature map G i :
[0105]
[0106] G i =((F′ i ⊙α i )+F′ i ) T ω i
[0107] in, represents feature map concatenation, ⊙ represents pixel-level dot product, and i represents the i-th level feature.
[0108] Reference Fig. 9 In some embodiments, the final guided feature map and the shallow features are gated convolutionally fused in step S500 to obtain the segmentation result of the skin lesion image, including the following steps S501 and S502:
[0109] Step S501: construct a gated convolution fusion module.
[0110] Step S502: input the final guided feature map and the shallow features into the gated convolution fusion module to obtain the segmentation result of the skin lesion image output by the gated convolution fusion module.
[0111] like Figure 1 As shown, step S502 fuses the final guided feature map S′ through the gated convolution fusion module 3 and shallow features F′ 2 , the segmentation results of skin lesion images are obtained to reduce noise interference and enhance the network training effect.
[0112] Reference Fig.10In one embodiment of the present application, a skin lesion image segmentation device is provided. The skin lesion image segmentation device includes an image acquisition unit 1100, a feature extraction unit 1200, a multi-scale feature fusion unit 1300, a feature enhancement unit 1400 and a segmentation result acquisition unit 1500, wherein:
[0113] The image acquisition unit 1100 is used to acquire skin lesion images.
[0114] The feature extraction unit 1200 is used to extract shallow features and multi-level features of the skin lesion image, and extract context information corresponding to each level of features.
[0115] The multi-scale feature fusion unit 1300 is used to input the context information of the multi-level features into a preset gated fusion decoder so that the gated fusion decoder adaptively selects complementary information from the context information of the multi-level features, and fuses the context information of the multi-level features through addition and gating mechanisms to obtain the initial guided feature map output by the gated fusion decoder.
[0116] The feature enhancement unit 1400 is used to construct a shape-guided flow module, and input the initial guided feature map and the context information of the multi-level features into the shape-guided flow module, so that the shape-guided flow module gradually mines the lesion boundaries in the feature map and highlights the shape of the lesion area, and obtains the final guided feature map output by the shape-guided flow module.
[0117] The segmentation result acquisition unit 1500 is used to perform gated convolution fusion on the final guided feature map and the shallow features to obtain the segmentation result of the skin lesion image.
[0118] It should be noted that the present device embodiment and the above method embodiment are based on the same inventive concept, and therefore the relevant contents of the above method embodiment are also applicable to the present device embodiment and will not be repeated here.
[0119] The system first extracts the context information of the multi-level features of the skin lesion image; then the gated fusion decoder adaptively selects complementary information from the context information of the multi-level features, and fuses the context information of the multi-level features through the addition and gating mechanism to obtain the initial guided feature map. The gated fusion decoder can effectively fuse multi-scale detail information and semantic information to solve the problem of inaccurate target positioning due to the different sizes and irregularities of the lesion area; then the shape guided flow module gradually mines the lesion boundary in the feature map and highlights the shape of the lesion area to solve the boundary fuzziness problem; finally, the final guided feature map output by the shape guided flow module and the shallow features of the skin lesion image are gated and convoluted to obtain the final output, which can reduce noise interference and enhance the network training effect. Combined with experiments, it is proved that the system can well solve the problem of inaccurate target positioning due to the different sizes and irregularities of the lesion area in the skin lesion image and the ambiguity between the lesion area and the background.
[0120] Reference Figure 1 An embodiment of the present application provides a method for skin lesion image segmentation. The method constructs a GFA-Net network (Gated Fusion Attention Network for Skin Lesion Segmentation), inputs the skin lesion image to be segmented into the GFA-Net network, and obtains the segmentation result output by the GFA-Net network.
[0121] like Figure 1 As shown, the GFA-Net network constructed in the embodiment of the method includes: a feature extraction network module (FeatureExtraction Stream), a gated fusion decoder (GFD), a shape guidance stream module (Shape Guidance Stream) and a gated convolution fusion module (GCF).
[0122] The feature extraction network module can effectively extract detail information and semantic information. The shape-guided flow module uses the channel reverse attention (CRA) module, gated convolution and deep supervision to gradually mine boundary-related information and highlight the shape of the lesion area. The gated fusion decoder selectively fuses multiple levels of features and outputs the prediction results as the initial guidance map of the shape-guided flow module. The gated convolution fusion module (GCF) fuses shallow features and the prediction map of the shape-guided flow module (the final guidance feature map output by the shape-guided flow module) to improve the performance of the network.
[0123] like Figure 1As shown in Figure 2, the feature extraction network module uses HarDNet-68 as the basic network structure. The HarDNet-68 network is mainly used to extract the shallow features (F 2 ) and multi-level features (F 3 ,F 4 ,F 5 ), and then add a context feature extraction module to the HarDNet-68 network.
[0124] The context feature extraction module is used to solve the problem of large diversity and irregular shapes of lesion areas in skin lesion images. The main function of the context feature extraction module is to extract context information of three different levels of features, low, medium and high, and features of objects of different sizes. The context feature extraction module includes: receptive field module and residual multi-core pooling module.
[0125] 1) The receptive field module is a multi-branch convolutional layer composed of convolution kernels of different sizes and convolutions with different dilation rates. First, a 1x1 convolutional layer is used in each branch convolutional layer to reduce the number of channels of the feature map, and a set of 1xn and nx1 convolutional layers are added to replace the nxn convolutional layer. Then, dilation convolutions with specific dilation rates and jump connections are used, and finally the feature maps of each branch are spliced together.
[0126] 2) The residual multi-core pooling module uses multiple effective fields of view to identify skin lesion targets of different sizes. Specifically, the residual multi-core pooling module uses four pooling layers of different sizes to extract global context information, and then outputs four feature maps of different sizes. In order to reduce the amount of computation and the dimension of the weight, 1×1 convolution is used to reduce the number of channels of the four output feature maps. Then, according to the size of the input feature map, the low-dimensional feature map is upsampled by bilinear interpolation. Finally, the input feature map is concatenated with the upsampled feature map.
[0127] The gated fusion decoder uses a gating mechanism to weigh the contextual information in the low, medium and high level features and selectively fuses the feature information at different levels. For the i-th level feature of the input, the gated fusion decoder can adaptively select complementary information from the rest of the input features, and then fuse these features together through addition and gating mechanisms.
[0128] The shape-guided flow module includes a channel reverse attention module, a gated convolution fusion module, and deep supervision. Based on the rough lesion location, the shape-guided flow module gradually mines the boundary information of the lesion through the channel reverse attention module and the gated convolution fusion module, and uses a set of repeated channel reverse attention and gated convolution fusion modules to solve the boundary blur problem. The channel reverse attention module consists of a reverse attention module and a multi-frequency channel attention module. Reverse attention is used to remove foreground objects, thereby gradually mining the skin lesion area and detail information more accurately. In order to highlight the importance of each feature channel and better capture rich information, frequency channel attention is applied after reverse attention, which uses discrete cosine transform to process each channel of the input feature to obtain different frequency components, thereby solving the problem of information loss and lack of diversity of input features due to global average pooling.
[0129] Finally, the shallow features of the feature extraction module and the final output feature map of the shape-guided flow module are fused through the gated convolution fusion module to obtain the image segmentation result, so as to reduce noise interference and enhance the network training effect.
[0130] The following is a set of experimental results:
[0131] The performance of GFA-Net and ten excellent networks were quantitatively studied on the ISIC 2018 dataset. Among them, FCN, UNet, UNet++, AttUNet and DeepLabv3+ are widely used in medical image segmentation. CANet, Ms RED, CENet and CPFNet are all networks designed specifically for skin lesion segmentation. PraNet uses deep supervision to implement the strategy from coarse segmentation to fine segmentation. All evaluation indicators shown in Table 1 are the average and standard deviation of five-fold cross validation. As can be seen from the table, GFA-Net obtains the highest value in all five evaluation indicators. DeepLabv3+ using ResNet 50 as the backbone network performs better than FCN, UNet, UNet++ and AttUNet. Among the networks using the skin lesion dataset, CPFNet performs the best, while CANet has the worst segmentation performance. Compared with CANet, DeepLabv3+, PraNet and CPFNet, GFA-Net improves the main evaluation indicator JI by 2.11%, 1.74%, 1.54% and 0.61% respectively. These experimental results demonstrate that the proposed GFA-Net effectively segments the lesion area of the skin and outperforms other state-of-the-art methods.
[0132] Qualitative Research: Analysis Fig.11It can be seen that the segmentation images in the first and fifth rows show that GFA-Net can handle the problem of low contrast between the lesion area and its surrounding tissues well. The segmentation images in the third, sixth and seventh rows show that GFA-Net can handle the problem of large diversity and irregular shapes of lesion areas and blurred boundaries well. The segmentation images in the second and fourth rows show that GFA-Net can also identify the problem of small lesion areas well in the presence of interference noise.
[0133]
[0134] Table 1
[0135] Experiments were conducted on four public skin lesion segmentation datasets, namely ISIC 2016, ISIC 2017, ISIC 2018 and PH2. The experimental results show that the GFA-Net proposed in this embodiment is superior to most existing methods and can effectively handle problems such as irregular shapes, noise interference and blurred boundaries.
[0136] This method has at least the following beneficial effects:
[0137] (1) A gated fusion attention network (GFA-Net) is constructed to accurately segment skin lesion areas. It uses a gated fusion decoder to fuse the contextual information of multi-scale features and generate an initial guidance map for the shape-guided flow. Compared with most existing methods, it can effectively handle problems such as irregular shapes, noise interference, and blurred boundaries.
[0138] (2) In order to fully exploit boundary information and remove interference information, the channel reverse attention module is used in the shape-guided flow module to repeatedly refine the boundaries of skin lesions, and the gated convolution fusion module is used to retain important information, and the shallow features in the feature extraction flow and the final output of the shape-guided flow are fused to improve segmentation performance and promote network convergence. Due to this repeated calibration mechanism in the shape-guided flow, GFA-Net can more accurately identify skin lesion areas.
[0139] An embodiment of the present application provides an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor.
[0140] The processor and the memory may be connected via a bus or other means.
[0141] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0142] The non-transient software program and instructions required to implement the skin lesion image segmentation method of the above embodiment are stored in the memory. When executed by the processor, the skin lesion image segmentation method of the above embodiment is executed, for example, the above described Figure 1 Steps S100 to S500 of the method, Figure 3 Steps S201 and S202 in Figure 6 Steps S401 and S406 in Fig. 9 Steps S501 and S502 in .
[0143] The terminal embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0144] In addition, an embodiment of the present application provides a computer-readable storage medium, which stores computer-executable instructions. The computer-executable instructions are executed by a processor or a controller, for example, by a processor in the above terminal embodiment, so that the above processor can execute the skin lesion image segmentation method in the above embodiment, for example, execute the above described Figure 1 Steps S100 to S500 of the method, Figure 3 Steps S201 and S202 in Figure 6 Steps S401 and S406 in Fig. 9 Steps S501 and S502 in .
[0145] It will be appreciated by those skilled in the art that all or some of the steps and systems in the disclosed method above may be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that may be used to store desired information and may be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0146] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0147] The embodiments of the present application are described in detail above in conjunction with the accompanying drawings, but the present application is not limited to the above embodiments. Various changes can be made within the knowledge scope of ordinary technicians in the relevant technical field without departing from the purpose of the present application.
Claims
1. A skin lesion image segmentation method, characterized in that: The skin lesion image segmentation method comprises: Acquire images of skin lesions; Extracting shallow features and multi-level features of the skin lesion image, and extracting context information corresponding to each level of features; extracting context information corresponding to each level of features includes: Construct a context feature extraction module, the context feature extraction module includes a receptive field module and a residual multi-core pooling module connected in sequence, the receptive field module includes a jump connection layer and a multi-branch convolution layer, wherein each branch convolution layer includes a 1×1 convolution layer, a 1×n convolution layer, an n×1 convolution layer and a hole convolution layer with a specific hole rate connected in sequence, and the output result of the jump connection layer and the output result of the multi-branch convolution layer are spliced to obtain the output result of the receptive field module; the residual multi-core pooling module includes four pools of different sizes A convolution layer, four 1×1 convolution layers and a bilinear interpolation layer, wherein the four pooling layers of different sizes are used to extract four feature maps of different sizes from the feature map input to the residual multi-core pooling module, respectively, the four 1×1 convolution layers are used to reduce the number of channels of the four feature maps of different sizes, the bilinear interpolation layer is used to upsample the feature maps output by the four 1×1 convolution layers, and the feature map input to the residual multi-core pooling module is spliced with the feature map output by the bilinear interpolation layer to obtain the output result of the residual multi-core pooling module; Input each level feature of the multi-level features into the context feature extraction module to obtain context information of each level feature in the context feature extraction module; The context information of the multi-level features is input into a preset gated fusion decoder, so that the gated fusion decoder adaptively selects complementary information from the context information of the multi-level features, and fuses the context information of the multi-level features through addition and gating mechanisms to obtain an initial guided feature map output by the gated fusion decoder; the expression of the gated fusion decoder includes: Among them, σ represents the sigmoid activation function, Indicates pixel-level multiplication in the channel, f i represents the i-th level feature in the multi-level features, S g represents the initial guided feature map; A shape-guided flow module is constructed, and the context information of the initial guided feature map and the multi-level features is input into the shape-guided flow module, so that the shape-guided flow module gradually mines the lesion boundary in the feature map and highlights the shape of the lesion area, and obtains the final guided feature map output by the shape-guided flow module; the multi-level features are N level features, wherein the first level feature is the lowest level feature among the N level features, and the level features from the first level feature to the Nth level feature are increased in sequence; the shape-guided flow module includes N channel reverse attention modules and N gated convolution fusion modules; the shape-guided flow module gradually mines the lesion boundary in the feature map and highlights the shape of the lesion area, and obtains the final guided feature map output by the shape-guided flow module, including: Using the initial guide feature map as the first guide feature map, and inputting the first guide feature map and the first hierarchical feature into the first channel reverse attention module, so that the first channel reverse attention module removes the foreground object by using reverse attention, and outputs the first attention feature map; Inputting the first attention feature map and the first guide feature map into the first gated convolution fusion module to obtain a second guide feature map output by the first gated convolution fusion module; Inputting the second guide feature map and the second hierarchical features into the second channel reverse attention module, so that the second channel reverse attention module removes the foreground object by using reverse attention and outputs a second attention feature map; Inputting the second attention feature map and the second guide feature map into the second gated convolution fusion module to obtain a third guide feature map output by the second gated convolution fusion module; And so on, until the Nth guide feature map and the Nth hierarchical feature are input into the Nth channel reverse attention module, so that the Nth channel reverse attention module removes the foreground object by using reverse attention and outputs the Nth attention feature map; Inputting the Nth attention feature map and the Nth guide feature map into the Nth gated convolution fusion module to obtain a final guide feature map output by the Nth gated convolution fusion module; Among them, the expression of the gated convolution fusion module includes: G i =((F′ i ☉a i )+F′ i ) T oh i Among them, α i represents the weight value, σ represents the sigmoid function, Conv 1×1 represents 1×1 convolution, F′ i represents the i-th attention feature map, S i ′ represents the i-th guided feature map, i is an integer from 1 to N, ω i Represents the weight coefficient, G i Represents the guided feature map output by the gated convolutional fusion module; The final guided feature map and the shallow features are gated convolutionally fused to obtain a segmentation result of the skin lesion image.
2. The skin lesion image segmentation method according to claim 1, characterized in that: The channel reverse attention module includes a reverse attention module and a multi-frequency channel attention module connected in sequence, wherein the reverse attention module uses reverse attention to remove foreground objects, and the multi-frequency channel attention module is used to process each channel of the input feature map according to discrete cosine transform to obtain different frequency components.
3. The skin lesion image segmentation method according to claim 2, characterized in that: After the i-th guided feature map output by the i-th gated convolution fusion module, it also includes: The i-th guided feature map is deeply supervised.
4. The skin lesion image segmentation method according to claim 1, characterized in that: The step of performing gated convolution fusion on the final guided feature map and the shallow features to obtain a segmentation result of the skin lesion image includes: Constructing a gated convolution fusion module; The final guided feature map and the shallow features are input into the gated convolution fusion module to obtain the segmentation result of the skin lesion image output by the gated convolution fusion module.
5. A skin lesion image segmentation system, characterized in that: The skin lesion image segmentation system comprises: An image acquisition unit, used for acquiring skin lesion images; A feature extraction unit is used to extract shallow features and multi-level features of the skin lesion image, and extract context information corresponding to each level of features; the extraction of context information corresponding to each level of features includes: Construct a context feature extraction module, the context feature extraction module includes a receptive field module and a residual multi-core pooling module connected in sequence, the receptive field module includes a jump connection layer and a multi-branch convolution layer, wherein each branch convolution layer includes a 1×1 convolution layer, a 1×n convolution layer, an n×1 convolution layer and a hole convolution layer with a specific hole rate connected in sequence, and the output result of the jump connection layer and the output result of the multi-branch convolution layer are spliced to obtain the output result of the receptive field module; the residual multi-core pooling module includes four pools of different sizes A convolution layer, four 1×1 convolution layers and a bilinear interpolation layer, wherein the four pooling layers of different sizes are used to extract four feature maps of different sizes from the feature map input to the residual multi-core pooling module, respectively, the four 1×1 convolution layers are used to reduce the number of channels of the four feature maps of different sizes, the bilinear interpolation layer is used to upsample the feature maps output by the four 1×1 convolution layers, and the feature map input to the residual multi-core pooling module is spliced with the feature map output by the bilinear interpolation layer to obtain the output result of the residual multi-core pooling module; Input each level feature of the multi-level features into the context feature extraction module to obtain context information of each level feature in the context feature extraction module; The multi-scale feature fusion unit is used to input the context information of the multi-level features into a preset gated fusion decoder, so that the gated fusion decoder adaptively selects complementary information from the context information of the multi-level features, and fuses the context information of the multi-level features through addition and gating mechanisms to obtain an initial guided feature map output by the gated fusion decoder; the expression of the gated fusion decoder includes: Among them, σ represents the sigmoid activation function, Indicates pixel-level multiplication in the channel, f i represents the i-th level feature in the multi-level features, S g represents the initial guided feature map; The feature enhancement unit is used to construct a shape-guided flow module, input the context information of the initial guided feature map and the multi-level features into the shape-guided flow module, so that the shape-guided flow module gradually mines the lesion boundary in the feature map and highlights the shape of the lesion area, and obtains the final guided feature map output by the shape-guided flow module; the multi-level features are N hierarchical features, wherein the first hierarchical feature is the lowest hierarchical feature among the N hierarchical features, and the hierarchical features from the first hierarchical feature to the Nth hierarchical feature are increased in sequence; the shape-guided flow module includes N channel reverse attention modules and N gated convolution fusion modules; the shape-guided flow module gradually mines the lesion boundary in the feature map and highlights the shape of the lesion area, and obtains the final guided feature map output by the shape-guided flow module, including: Using the initial guide feature map as the first guide feature map, and inputting the first guide feature map and the first hierarchical feature into the first channel reverse attention module, so that the first channel reverse attention module removes the foreground object by using reverse attention, and outputs the first attention feature map; Inputting the first attention feature map and the first guide feature map into the first gated convolution fusion module to obtain a second guide feature map output by the first gated convolution fusion module; Inputting the second guide feature map and the second hierarchical features into the second channel reverse attention module, so that the second channel reverse attention module removes the foreground object by using reverse attention and outputs a second attention feature map; Inputting the second attention feature map and the second guide feature map into the second gated convolution fusion module to obtain a third guide feature map output by the second gated convolution fusion module; And so on, until the Nth guide feature map and the Nth hierarchical feature are input into the Nth channel reverse attention module, so that the Nth channel reverse attention module removes the foreground object by using reverse attention and outputs the Nth attention feature map; Inputting the Nth attention feature map and the Nth guide feature map into the Nth gated convolution fusion module to obtain a final guide feature map output by the Nth gated convolution fusion module; Among them, the expression of the gated convolution fusion module includes: G i =((F′ i ☉a i )+F′ i ) T oh i Among them, α i represents the weight value, σ represents the sigmoid function, Conv 1×1 represents 1×1 convolution, F′ i represents the i-th attention feature map, S i ′ represents the i-th guided feature map, i is an integer from 1 to N, ω i Represents the weight coefficient, G i Represents the guided feature map output by the gated convolutional fusion module; The segmentation result acquisition unit is used to perform gated convolution fusion on the final guided feature map and the shallow features to obtain the segmentation result of the skin lesion image.
6. An electronic device, characterized in that: include: at least one control processor and a memory for communicatively coupling with the at least one control processor; The memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute the skin lesion image segmentation method as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute a skin lesion image segmentation method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Medical image segmentation method based on deep learning
CN111145170A
Medical image segmentation method based on T-shaped attention structure
CN111612790A