Pigmented Skin Lesion Image Segmentation Method Based on Reverse Channel Filling CNN and Level Set

By using a joint method of filling CNN and level sets with reverse channels in pigmented lesions image segmentation, the problem of low segmentation accuracy in the prior art is solved, and higher segmentation accuracy and accuracy are achieved.

CN113989288BActive Publication Date: 2025-06-17GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111076764.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-14
Publication Date
2025-06-17
Estimated Expiration
2041-09-14

AI Technical Summary

Technical Problem

Pigmentogenic lesions image segmentation faces challenges, including low contrast, blurred boundaries, changeable color and human interference, resulting in less accuracy of existing unsupervised methods.

Method used

The segmentation method based on reverse channel filling CNN and horizontal set is adopted to learn the spatial position of the target through the backward propagation of BCF-CNN, and the segmentation result is driven through horizontal set evolution to improve the segmentation accuracy.

Benefits of technology

Improved the accuracy and accuracy of pigmented lesions image segmentation, especially when dealing with lesions with blurred boundaries and changing color, achieving higher IoU values.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113989288B_ABST
    Figure CN113989288B_ABST
Patent Text Reader

Abstract

The present invention provides a method for segmenting pigmented skin lesion images based on reverse-channel filling CNN and level set. This method is based on reverse-channel filling CNN and a joint backpropagation learning algorithm. It obtains the approximate spatial position of the target based on the attention mechanism, enhances the spatial position features of the target through reverse-channel filling, and improves the accuracy of the target position output by CNN. Further, the feature energy learned by CNN is input into the level set segmentation model to drive the level set evolution, and a backpropagation learning algorithm combining CNN and level set is established to further improve the segmentation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and more specifically, to a method for segmenting pigmented skin lesion images based on reverse channel filling CNN and level set. Background Art

[0002] Common forms of pigmented skin lesions include melanoma and melanocytic nevus. Among them, melanoma is one of the most common fatal forms of skin cancer, and more than 80% of melanomas originate from the malignant transformation of melanocytic nevi. Pigmented skin lesion image segmentation can provide information such as the location, shape, and size of the skin lesions, and is one of the key technologies for auxiliary diagnosis. The main difficulties faced by pigmented skin lesion image segmentation include: the low contrast between the skin lesion area and the surrounding skin; the blurred, irregular, and non-fixed shape of the skin lesion boundary; the influence of artificial or other features of the skin itself, such as hair, blood vessels, etc.; the variable color of the skin lesion area; and the possibility of being fragmented by scars, etc.

[0003] Pigmented skin lesion segmentation methods can generally be divided into two categories: unsupervised methods and supervised methods. Unsupervised methods mainly distinguish skin lesions from others by designing and extracting features of skin lesion images. Such as threshold method, active contour model, histogram statistics method, and region growing method, etc. The disadvantages of these methods are: if the features of the skin lesions themselves are not consistent (such as diverse colors, diverse textures, and unclear edges, etc.), or there are black frames, hair, and marking lines around the skin lesions, the accuracy of segmentation using these methods is relatively low. The main reason is that the features designed by unsupervised methods are for those "conventional" skin lesions, and for some "unconventional" skin lesions, either features specifically designed for them need to be supplemented, or some additional preprocessing needs to be added. Supervised methods have the potential to uniformly process conventional and unconventional skin lesions because they can automatically learn the features of these skin lesions and their surroundings, such as methods based on support vector machines, machine learning, etc. Among them, the skin lesion segmentation method based on deep learning has received particular attention. Summary of the Invention

[0004] The present invention provides a method for segmenting pigmented skin lesion images based on reverse channel filling CNN and level set with relatively high segmentation accuracy.

[0005] In order to achieve the above technical effects, the technical solution of the present invention is as follows:

[0006] A method for segmenting pigmented skin lesion images based on reverse channel filling CNN and level set, comprising the following steps:

[0007] S1: Through the first backward propagation of BCF-CNN, learn the spatial position of the target;

[0008] S2: Relax the coordinates of the learned spatial position to obtain a position prior;

[0009] S3: Inject the location prior in reverse into the original input, and then perform the second backpropagation: continue to learn the driving energy of the level set evolution and the initial level set to drive the level set evolution to obtain the final segmentation result.

[0010] Furthermore, the specific process of step S1 is as follows:

[0011] Perform a Concat operation on the original image and the idle channel, and input it into the CNN to learn the target location. After training is completed, the rough location of the target can be output. Among them, the idle channel is a channel C0 of all zeros with the same size as the original image. The introduction of the idle channel is to prepare for automatically adding the location prior.

[0012] Furthermore, the specific process of step S2 is as follows:

[0013] Relax the learned spatial target location into Extreme points. Extreme points include 4 coordinate points: left-most, right-most, top and bottom pixels, which are used to estimate the spatial location of the target; fill the Extreme points into the idle channel as the location prior, that is, the automatic addition of the location prior is completed.

[0014] Furthermore, BCF-CNN includes two parts: an encoder and a decoder. Among them, the encoder is similar to the encoder of Unet and contains 5 groups of convolutional layers and 4 max pooling layers. The decoder integrates the Attention Gate module to output the target location more accurately.

[0015] Furthermore, the decoder contains two branches: an external energy branch and an internal energy branch; both branches contain 5 groups of convolutional layers and 4 upsampling layers. Among them, the external energy branch contains two sub-branches in the last layer, which respectively output the external energy and the target location; the internal energy branch outputs the curvature energy coefficient.

[0016] Furthermore, in the first backpropagation, the input consists of the original image and 1 idle channel; in the second time, the input consists of the original image and the idle channel filled with the location prior.

[0017] Furthermore, the first backpropagation mainly learns the rough location of the target, converts the location probability into the spatial coordinates of the target, and obtains the location prior to fill it back into the idle channel in reverse;

[0018] The second time is to further finely tune the target position and learn the energy for the level set evolution: First, feature maps with a spatial resolution close to the input are extracted through separate encoding-decoding branches; then, a convolutional network is used to automatically learn the external energy, the initial level set, and the curvature energy coefficient.

[0019] Furthermore, the first backpropagation is for the BCF-CNN to learn the rough position of the target, which is only related to the BCF-CNN. The loss function L1 is calculated based on the error between the position estimation output of the BCF-CNN decoder and the ground truth (GT) of the target position, and its definition is as follows:

[0020]

[0021] where, φ t is the position estimation value output in the t-th training, and φ gt is the value of the level set function corresponding to the training sample GT.

[0022] Furthermore, the second backpropagation occurs after the first backpropagation. At this time, the initial level set φ′ t of the BCF-CNN after reverse padding, the external energy F e and the curvature energy coefficient m need to be input into the level set evolution equation for K times of level set evolution first, and then the loss function value is calculated according to the level set segmentation result φ″ t and the error is backpropagated to the BCF-CNN to achieve the coordination of level set evolution and BCF-CNN learning. Its loss function L2 is defined as follows:

[0023] L2 = L φ + L e

[0024]

[0025]

[0026]

[0027] where, L φ is the class-balanced binary cross-entropy loss, L e is the external energy angular loss, M gt is the GT segmentation result, H(·) is the Heaviside function, w n , w p are the proportions of the number of background and target pixels in a sample to the total amount respectively, DT(·) is the distance transform function, and φ gt represents the GT boundary.

[0028] Furthermore, the driving energy of level set evolution includes external energy and two internal energies, namely curvature and length constraints. The level set evolution formula is defined as follows:

[0029]

[0030]

[0031] Among them, the first term on the right side is the external energy term, and F e is the external energy; the second term is the curvature energy term, m is the curvature energy coefficient, and κ is the curvature of φ; the third term is the length constraint term, div(·) is the divergence, p(·) is the double-well potential function, and λ and μ are weight coefficients. Both F e and m are obtained through BCF-CNN learning.

[0032] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:

[0033] Based on the reverse channel filling CNN and the joint backpropagation learning algorithm, the present invention obtains the approximate spatial position of the target based on the attention mechanism, and enhances the spatial position features of the target through reverse channel filling, improving the accuracy of the CNN output target position; further, the feature energy learned by the CNN is input into the level set segmentation model to drive the level set evolution, and a joint CNN and level set backpropagation learning algorithm is established to further improve the segmentation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a schematic flow chart of the method of the present invention;

[0035] Figure 2 is a schematic diagram of the reverse channel filling process;

[0036] Figure 3 is the BCF-CNN network model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0038] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the size of the actual product;

[0039] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0040] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0041] As Figure 1-2As shown in the figure, a method for segmenting pigmented skin lesion images based on reverse channel filling CNN and level set includes the following steps:

[0042] S1: Through the first backward propagation of BCF-CNN, learn the spatial position of the target;

[0043] S2: Relax the learned spatial position coordinates to obtain a position prior;

[0044] S3: Inject the position prior back into the original input, and then perform the second backward propagation: continue to learn the driving energy of the level set evolution and the initial level set to drive the level set evolution to obtain the final segmentation result.

[0045] The specific process of step S1 is as follows:

[0046] Perform a Concat operation on the original image and the idle channel, and input it into the CNN to learn the target position. After the training is completed, the rough position of the target can be output. Among them, the idle channel is a zero-filled channel C0 with the same size as the original image. The introduction of the idle channel is to prepare for automatically adding a position prior.

[0047] The specific process of step S2 is as follows:

[0048] Relax the learned spatial target position into Extreme points. Extreme points include 4 coordinate points: left-most, right-most, top and bottom pixels, which are used to estimate the spatial position of the target; fill the Extreme points into the idle channel as the position prior, that is, the automatic addition of the position prior is completed.

[0049] As Figure 3 shown, BCF-CNN includes an encoder and a decoder. The encoder is similar to the encoder of Unet and contains 5 groups of convolutional layers and 4 max pooling layers. The decoder integrates an Attention Gate module to output the target position more accurately.

[0050] The decoder contains two branches: an external energy branch and an internal energy branch; both branches contain 5 groups of convolutional layers and 4 upsampling layers. The external energy branch contains two sub-branches in the last layer, which output the external energy and the target position respectively; the internal energy branch outputs the curvature energy coefficient.

[0051] In the first backward propagation, the input consists of the original image and 1 idle channel; in the second time, the input consists of the original image and the idle channel filled with the position prior.

[0052] The first backpropagation mainly learns the rough position of the target, converts the position probability into the spatial coordinates of the target, and fills the position prior back into the idle channels;

[0053] The second time is to further fine-tune the target position and learn the energy driving the level set evolution: First, feature maps with a spatial resolution close to the input are extracted through independent encoding-decoding branches respectively; then, a convolutional network is used to automatically learn the external energy, the initial level set, and the curvature energy coefficient.

[0054] The first backpropagation is that BCF-CNN learns the rough position of the target, which is only related to BCF-CNN. The loss function L1 is calculated according to the error between the position estimation output of the BCF-CNN decoder and the ground truth (GT) of the target position, and its definition is as follows:

[0055]

[0056] where, φ t is the position estimation value output at the t-th training, and φ gt is the value of the level set function corresponding to the training sample GT.

[0057] The second backpropagation occurs after the first backpropagation. At this time, the initial level set φ′ t of the BCF-CNN after reverse filling, the external energy F e and the curvature energy coefficient m need to be input into the level set evolution equation for K times of level set evolution first, and then the loss function value is calculated according to the level set segmentation result φ″ t and the error is backpropagated to BCF-CNN to realize the cooperation between level set evolution and BCF-CNN learning. Its loss function L2 is defined as follows:

[0058] L2 = L φ + L e

[0059]

[0060]

[0061]

[0062] where, L φ is the class-balanced binary cross-entropy loss, L e is the external energy angle loss, M gt is the GT segmentation result, H(·) is the Heaviside function, w n , w pare the proportions of the background and target pixel numbers in a sample to the total amount respectively, DT(·) is the distance transformation function, and φ gt represents the GT boundary.

[0063] The driving energy of the level set evolution includes external energy and two internal energies, i.e., curvature and length constraints. The level set evolution formula is defined as follows:

[0064]

[0065]

[0066] Among them, the first term on the right side is the external energy term, and F e is the external energy; the second term is the curvature energy term, m is the curvature energy coefficient, and κ is the curvature of φ; the third term is the length constraint term, div(.) is the divergence, p(·) is the double-well potential function, and λ and μ are the weight coefficients. Both F e and m are obtained by learning with BCF-CNN.

[0067] This method establishes a unified deep learning-based level set image segmentation model. The deep learning-based level set image segmentation model changes the way of empirically selecting feature energy in the traditional level set segmentation model to the way of adaptively selecting effective feature energy by learning from a large amount of empirical data; the fused CNN-LS image segmentation model not only retains the original advantages of the level set image segmentation but also adds the learning characteristics of the neural network, and integrates the CNN and the level set image segmentation model into a whole. This model endows the level set image segmentation model with learning ability and improves its ability to integrate prior features;

[0068] This method is based on the reverse channel filling CNN and the joint backpropagation learning algorithm. Based on the attention mechanism, the approximate spatial position of the target is obtained, and the spatial position features of the target are enhanced by reverse channel filling, improving the accuracy of the CNN output target position; further, the feature energy learned by the CNN is input into the level set segmentation model to drive the level set evolution, and a joint CNN and level set backpropagation learning algorithm is established to further improve the segmentation accuracy. The IoU of this method on the test set of ISIC 2017 is 0.769, which is comparable to other SOTA methods.

[0069] The same or similar reference numerals correspond to the same or similar components;

[0070] The descriptions of the positional relationships in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0071] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A method for segmenting pigmented skin lesion images based on reverse channel filling CNN and level set, characterized in that, It includes the following steps: S1: Through the first backward propagation of BCF-CNN, learn the spatial position of the target; S2: Relax the learned spatial position to obtain a position prior; S3: Inject the position prior back into the original input, and then perform the second backward propagation: continue to learn the driving energy of the level set evolution and the initial level set to drive the level set evolution to obtain the final segmentation result; BCF-CNN includes an encoder and a decoder. The encoder is similar to the encoder of Unet and contains 5 groups of convolutional layers and 4 max pooling layers. The decoder integrates an Attention Gate module to more accurately output the target position; The decoder contains two branches: an external energy branch and an internal energy branch; Both branches contain 5 groups of convolutional layers and 4 upsampling layers. The external energy branch contains two sub-branches in the last layer, which respectively output the external energy and the target position; the internal energy branch outputs the curvature energy coefficient.

2. The method for segmenting pigmented skin lesion images based on reverse channel filling CNN and level set according to claim 1, characterized in that, The specific process of step S1 is: Perform a Concat operation on the original image and the idle channel, and input it into the CNN for learning the target position. After training is completed, the rough position of the target can be output. Among them, the idle channel is a channel C0 of all zeros with the same size as the original image. The introduction of the idle channel is to prepare for automatically adding a position prior.

3. The method for segmenting pigmented skin lesion images based on reverse channel filling CNN and level set according to claim 2, characterized in that, The specific process of step S2 is: Relax the learned spatial target position into Extreme points. Extreme points include 4 coordinate points: left-most, right-most, top and bottom pixels, which are used to estimate the spatial position of the target; fill the Extreme points into the idle channels as the position prior, that is, the automatic addition of the position prior is completed.

4. The method for segmenting pigmented skin lesion images based on reverse channel filling CNN and level set according to claim 3, characterized in that, In the first backward propagation, the input consists of the original image and 1 idle channel; in the second time, the input consists of the original image and the idle channel filled with the position prior.

5. The method for segmenting pigmented skin lesion images based on reverse channel filling CNN and level set according to claim 4, characterized in that, The first backward propagation mainly learns the rough position of the target, converts the position probability into the spatial coordinates of the target, and fills the position prior back into the idle channel; The second time is to further fine-tune the target position and learn the energy to drive the level set evolution: First, extract feature maps with a spatial resolution close to the input through independent encoding-decoding branches respectively; then, use a convolutional network to automatically learn the external energy, the initial level set and the curvature energy coefficient.

6. The method for segmenting pigmented skin lesion images based on reverse channel filling CNN and level set according to claim 5, characterized in that, The first backward propagation is for BCF-CNN to learn the rough position of the target, which is only related to BCF-CNN. The loss function L1 is calculated according to the error between the position estimation output of the BCF-CNN decoder and the true value of the target position, and its definition is as follows: Among them, φ t is the position estimation value output by the t-th training, and φ gt is the level set function value corresponding to the training sample GT.

7. The method for segmenting pigmented skin lesion images based on reverse channel filling CNN and level set according to claim 6, characterized in that, The second backpropagation occurs after the first backpropagation. At this time, the initial level set φ of the BCF-CNN after reverse filling needs to be t ′, the external energy F e and the curvature energy coefficient m are input into the level set evolution equation for K times of level set evolution. Then, according to the level set segmentation result φ t ″, the loss function value is calculated, and the error is backpropagated to the BCF-CNN to achieve the coordination of level set evolution and BCF-CNN learning. The loss function L2 is defined as follows: L2 = L φ + L e Among them, L φ is the class-balanced binary cross-entropy loss, and L e is the external energy angle loss. M gt is the GT segmentation result, H(·) is the Heaviside function, and w n , w p are the proportions of the number of background and target pixels in a sample to the total amount respectively. DT(·) is the distance transform function, and φ gt represents the GT boundary.

8. The method for segmenting pigmented skin lesion images based on reverse channel filling CNN and level set according to claim 7, characterized in that, The driving energy of the level set evolution includes external energy and two internal energies, namely curvature and length constraints. The level set evolution formula is defined as follows: Among them, the first term on the right is the external energy term, F e is the external energy; the second term mκ|▽φ| is the curvature energy term, m is the curvature energy coefficient, κ is the curvature of φ; the third term is the length constraint term, div(·) is the divergence, p(·) is the double-well potential function, λ and μ are weight coefficients, F e and m are both obtained through BCF-CNN learning.

Citation Information

Patent Citations

  • Improved ceramic material member sequence image segmentation method of fully convolutional neural network

    CN106920243A

  • Layer segmentation method and system for retina layer and effusion area based on deep learning

    CN111583291A