Network model training method and skin lesion determination device

By constructing a lesion segmentation and classification network and integrating it into a mutual guide network model, the problem of doctors' manual judgment of low efficiency and poor accuracy is solved, and automated efficient and accurate lesion classification is achieved.

CN113812923BActive Publication Date: 2025-08-22SUZHOU CHUANGYING MEDICAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110981999.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-25
Publication Date
2025-08-22
Estimated Expiration
2041-08-25

AI Technical Summary

Technical Problem

In the prior art, doctors have low efficiency and poor accuracy when manually determining skin lesions.

Method used

The lesion segmentation network and the lesion classification network are constructed, fused into a mutual guide network model, and the skin lesion classification is automatically determined through sample skin image set training.

Benefits of technology

It improves the efficiency and accuracy of skin lesions determination and realizes automated lesion classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113812923B_ABST
    Figure CN113812923B_ABST
Patent Text Reader

Abstract

The present application discloses a network model training method and a skin lesion determination device, relating to the field of image processing technology. The method comprises: constructing a lesion segmentation network, the lesion segmentation network being used to segment lesion areas in a skin image; constructing a lesion classification network, the lesion classification network being used to determine the classification of skin lesions in the skin image based on the skin lesion areas; fusing the lesion segmentation network and the lesion classification network to obtain a mutually guided network model; training the mutually guided network model using a sample skin image set, and using the trained mutually guided network model to determine the classification of skin lesions in the image. This method solves the problem of low efficiency and poor accuracy in manual judgment by doctors in the prior art, and achieves the effect of automatically determining the lesion classification through the mutually guided network model, thereby improving the efficiency and accuracy of skin judgment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a network model training method and a skin lesion determination device, and belongs to the technical field of image processing. Background Art

[0002] With changing living environments, more and more patients are suffering from skin diseases. In existing technologies, doctors collect skin images using methods such as dermatoscopes and then manually make judgments based on the collected skin images. Obviously, this existing solution is affected by the doctor's level of experience, and the doctor's judgment results may be incorrect. Furthermore, manual judgment is inefficient. Summary of the Invention

[0003] The purpose of the present invention is to provide a network model training method and a skin lesion determination device to solve the problems existing in the prior art.

[0004] In order to achieve the above object, the present invention provides the following technical solutions:

[0005] According to a first aspect, an embodiment of the present invention provides a network model training method, the method comprising:

[0006] constructing a lesion segmentation network, wherein the lesion segmentation network is used to segment the skin lesion area in the image;

[0007] constructing a lesion classification network, wherein the lesion classification network is used to determine a skin lesion classification in an image based on the skin lesion area;

[0008] fusing the lesion segmentation network and the lesion classification network to obtain a mutually guided network model;

[0009] The mutually guided network model is trained using a sample skin image set, and the trained mutually guided network model is used to determine the classification of skin lesions in the skin images.

[0010] Optionally, the lesion segmentation network is a U-Net network, and the lesion segmentation network includes a convolution layer, a maximum pooling layer, a deconvolution layer and a ReLU nonlinear activation function.

[0011] Optionally, the encoding of the lesion segmentation network consists of ResNet50 pre-trained on the ImageNet dataset.

[0012] Optionally, the first convolutional layer in the ResNet50 is replaced by one 7*7 convolutional layer with two 3*3 convolutional layers, and the two 3*3 convolutional layers are used to maintain the size of the image input.

[0013] Optionally, the lesion classification network is composed of an Xception network pre-trained on the ImageNet dataset.

[0014] Optionally, the lesion classification network is composed of an improved Xception network, and the improved Xception network does not include the last pooling layer.

[0015] Optionally, the last two layers of separable convolutions in the improved Xception network are replaced with separable dilated convolutions.

[0016] Optionally, the fusing of the lesion segmentation network and the lesion classification network includes:

[0017] using the lesion segmentation network output as the input of the lesion classification network;

[0018] The pseudo labels output by the lesion classification network are input into the lesion segmentation network for training the lesion segmentation network.

[0019] Optionally, the sample skin images include a first sample image set and a second sample image set, the images in the first sample image set are provided with pixel-level segmentation annotations, and the images in the second sample image set are provided with image set annotations, and training the mutual guidance network model using the sample skin image set includes:

[0020] Training the lesion segmentation network using the first sample image set;

[0021] The mutual guidance network model is trained using the second sample image set.

[0022] In a second aspect, a skin lesion determination device is provided, the device comprising a memory and a processor, the memory storing at least one program instruction, the processor loading and executing the at least one program instruction to implement the following method:

[0023] Acquire images;

[0024] The image is input into a trained mutual guidance network model, and the trained mutual guidance network model outputs a classification of the skin lesions in the image; the mutual guidance network model is trained by the method described in the first aspect.

[0025] By constructing a lesion segmentation network for segmenting skin lesion areas in an image, and a lesion classification network for determining the classification of skin lesions in the image based on the skin lesion areas, and fusing the lesion segmentation network and the lesion classification network to form a mutually guided network model, the mutually guided network model is trained using a sample skin image set, and the trained mutually guided network model is used to determine the classification of skin lesions in the skin image. This solves the problem of low efficiency and poor accuracy of manual judgment by doctors in the prior art, and achieves the effect of automatically determining lesion classification through the mutually guided network model, thereby improving the efficiency and accuracy of skin judgment.

[0026] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 A flowchart of a network model training method provided by one embodiment of the present invention;

[0028] Figure 2 A schematic diagram of a possible network structure of a lesion segmentation network provided by one embodiment of the present invention;

[0029] Figure 3 A schematic diagram of a possible network structure of a lesion classification network provided by one embodiment of the present invention;

[0030] Figure 4 A schematic diagram of a possible network structure of a mutual guidance network model provided by one embodiment of the present invention;

[0031] Figure 5 A flowchart of a method for determining a skin disease provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0032] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0033] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0034] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0035] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0036] Please refer to Figure 1 , which shows a method flow chart of a network model training method provided by an embodiment of the present application, such as Figure 1 As shown, the method includes:

[0037] Step 101: constructing a lesion segmentation network, wherein the lesion segmentation network is used to segment skin lesion areas in a skin image;

[0038] In this embodiment, the lesion segmentation network is a U-Net network, such as Figure 2 As shown, the lesion segmentation network includes convolutional layers, max pooling layers, deconvolution layers, and the ReLU nonlinear activation function. In actual implementation, the encoding of the lesion segmentation network is composed of ResNet50 pre-trained on the ImageNet dataset, while the decoding remains unchanged. Using ResNet50 for encoding increases the number of layers in the encoding stage, improving the network's feature extraction capabilities and, consequently, segmentation accuracy. Furthermore, the residual structure of ResNet50 effectively prevents vanishing gradients and model degradation.

[0039] In addition, in this embodiment, in order to make the network structure of ResNet50 match the encoder structure of U-Net, in this embodiment, the last average pooling layer and the fully connected layer are removed to maintain the full convolutional structure of the network; and, in order to enable the network to retain more detailed information, the first 7×7 convolutional layer is replaced by two 3×3 convolutional layers that can maintain the image input size, while ensuring the size of the receptive field as much as possible.

[0040] The lesion segmentation network is optimized by minimizing the Dice coefficient loss, which is calculated as follows:

[0041]

[0042] X represents the result predicted by the lesion segmentation network, and Y represents the segmentation gold standard.

[0043] Optionally, the input of the lesion segmentation network is the skin image to be processed, and the output is the segmented lesion area.

[0044] The skin image described in this embodiment can be an image collected by a dermatoscope or an image obtained by taking a photo, and the specific collection method is not limited.

[0045] Step 102: constructing a lesion classification network, wherein the lesion classification network is used to determine the skin lesion classification in the skin image according to the skin lesion area;

[0046] The lesion classification network is composed of an Xception network pre-trained on the ImageNet dataset. In addition, in order to improve the classification performance of the lesion classification network, this application uses an improved Xception network. Specifically, the improved Xception network does not include the last pooling layer, thereby expanding the resolution of the feature map to avoid the loss of small lesion information during the progressive downsampling process; in addition, the last two layers of separable convolutions of the improved Xception network are replaced by separable dilated convolutions, and the dilation rate of the separable dilated convolutions is a preset value. The use of separable dilated convolutions can compensate for the reduction in receptive field caused by downsampling. After performing global average pooling (GAP), the generated features are input to a new fully connected (FC) layer, followed by a softmax activation function.

[0047] The weight of the 4th channel of the input (i.e., Mask) is initialized by averaging the weights of the other 3 channels (i.e., RGB image), and Mask-CN is optimized by minimizing the cross entropy loss, which is calculated as follows

[0048]

[0049] Where y is the label and p is the predicted output.

[0050] In actual implementation, the input of the lesion classification network is the lesion area output by the lesion segmentation network and the skin image to be processed. Since the skin lesion area usually only occupies a small part of the skin image, and most of the image is healthy skin tissue, the presence of artifacts such as hair, blood vessels and bubbles may interfere with the classification results of the lesion. The lesion mask (lesion area) obtained by the lesion segmentation network in this application can help the lesion classification network focus on the lesion area and remove interference on the skin image, thereby achieving accurate skin lesion classification. Figure 3 As shown in the figure, each training image and its corresponding lesion Mask are concatenated as the input of Mask-CN (lesion classification network), and Mask-CN output determines the obtained classification and pseudo label.

[0051] Mask-CN uses the class weights of the output layer to weight the feature map generated by the last convolutional layer, and then sums the weighted feature maps on all channels to generate a class activation map (CAM). The process is as follows Figure 3 As shown in the figure, after obtaining the CAM, the CAM is refined by the conditional random field (CRF) to obtain pseudo segmentation labels and used to train the lesion segmentation network to achieve weakly supervised semantic segmentation and obtain more accurate lesion segmentation.

[0052] Step 103, fusing the lesion segmentation network and the lesion classification network to obtain a mutual guidance network model;

[0053] This step includes:

[0054] First, the output of the lesion segmentation network is used as the input of the lesion classification network;

[0055] Second, the pseudo labels output by the lesion classification network are input into the lesion segmentation network for training the lesion segmentation network.

[0056] Optional, please refer to Figure 4 , which shows a possible network structure diagram of the fused mutual guidance network model. Figure 4 As shown in the figure, the image to be processed is segmented by the lesion segmentation network to obtain the lesion area (predicted Mask). The lesion area and the skin image are input into Mask-CN together. After obtaining the CAM, the CAM is refined by the conditional random field (CRF) to obtain the pseudo-label and the predicted classification. The pseudo-label is used to train the lesion segmentation network.

[0057] Step 104: training the mutual guidance network model using a sample skin image set. The trained mutual guidance network model is used to determine the classification of skin lesions in the image.

[0058] The sample skin images include a first sample image set and a second sample image set. The images in the first sample image set are provided with pixel-level segmentation annotations, and the images in the second sample image set are provided with image set annotations.

[0059] Specifically, the first sample image set contains 2,000 images for training and 600 images for testing. The training images include 374 melanoma cases, 254 seborrheic keratosis cases, and 1,372 benign nevus cases. The testing images include 117 melanoma cases, 97 seborrheic keratosis cases, and 393 benign nevus cases. Each image in this dataset has precise pixel-level segmentation annotations.

[0060] The second sample image set contains 10,015 images, including 1,113 cases of melanoma, 6,705 cases of benign moles, 514 cases of basal cell carcinoma, 327 cases of actinic keratosis, 1,099 cases of seborrheic keratosis, 115 cases of dermatofibroma, and 142 cases of vascular lesions. Each image in this dataset has only image-level annotations.

[0061] In actual implementation, in order to speed up the network reading speed and computational efficiency, the image sizes of the two sample image sets were all adjusted to 450×600 while maintaining the original aspect ratio of the images.

[0062] Accordingly, this step includes:

[0063] First, training the lesion segmentation network using the first sample image set;

[0064] In actual implementation, in order to increase the number of samples and thus improve the accuracy of network training, this step may also include:

[0065] (1) Perform data enhancement on the first sample image set.

[0066] The data enhancement method includes: at least one of random upside down flipping, random left-right flipping and random rotation.

[0067] (2) Train the lesion segmentation network using the enhanced first sample image set.

[0068] In actual implementation, the Adam optimizer can be used to update the network parameters. The training cycle is 120 rounds, the initial learning rate is set to 1×e-4, the batch size is set to 4, and the learning rate decays to 0.1 times the previous one every 30 rounds of training.

[0069] Second, the mutual guidance network model is trained using the second sample image set.

[0070] During training, the input images were randomly cropped to 224×224, and data augmentation was performed using random horizontal and vertical flipping. Stochastic gradient descent (SGD) was used to update network parameters. The lesion segmentation network and the lesion classification network were updated alternately and iteratively. Training was performed for 150 epochs, with an initial learning rate of 1e-4. This rate decayed to 0.1 times the previous value every 30 epochs, and the training batch size was set to 4. A test set from the first sample image set was used to monitor the performance of each network, and training was terminated if the network became overfitted.

[0071] Among them, the lesion classification output by the mutual guidance network model can be any one of melanoma, seborrheic keratosis, benign nevus, dermatofibroma, basal cell carcinoma, actinic keratosis, and vascular lesions. Of course, in actual implementation, it can also be other skin diseases, and this embodiment does not limit this.

[0072] In summary, by constructing a lesion segmentation network for segmenting skin lesion areas in skin images; constructing a lesion classification network for determining the classification of skin lesions in skin images based on the skin lesion areas; fusing the lesion segmentation network and the lesion classification network to obtain a mutually guided network model; and training the mutually guided network model using a sample skin image set, the trained mutually guided network model is used to determine the classification of skin lesions in skin images. This solves the problem of low efficiency and poor accuracy in manual judgment by doctors in the prior art, achieving the effect of automatically determining lesion classifications through the mutually guided network model, thereby improving the efficiency and accuracy of skin judgment.

[0073] Please refer to Figure 5 , which shows a flow chart of a method for determining skin diseases provided by an embodiment of the present application, such as Figure 5 As shown, the method includes:

[0074] Step 501, obtaining a skin image;

[0075] Step 502: Input the skin image into the trained mutual guidance network model, and output the classification of skin lesions in the skin image through the trained mutual guidance network model.

[0076] The mutual guidance network model is obtained by training using the method described in the above embodiment.

[0077] Specifically, the skin image is input into the trained U-Net segmentation network to obtain a segmentation mask; the skin image and the segmentation mask are spliced ​​together and input into the trained Mask-CN classification network for analysis; the Mask-CN classification network identifies the category of the skin image to be tested and assigns a classification label.

[0078] In summary, by acquiring a skin image, inputting the skin image into a trained mutual guidance network model, and outputting the classification of skin lesions in the skin image through the trained mutual guidance network model, the problem of low efficiency and poor accuracy of manual judgment by doctors in the prior art is solved, and the lesion classification can be automatically determined by the mutual guidance network model, thereby improving the efficiency and accuracy of skin judgment.

[0079] The present application also provides a skin lesion determination device, which includes a memory and a processor. The memory stores at least one program instruction, and the processor loads and executes the above-mentioned skin disease determination method by loading and executing the at least one program instruction.

[0080] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0081] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.

Claims

1. A network model training method, characterized in that: The method comprises: constructing a lesion segmentation network, wherein the lesion segmentation network is used to segment the skin lesion area in the image; constructing a lesion classification network, wherein the lesion classification network is used to determine a skin lesion classification in an image based on the skin lesion area; fusing the lesion segmentation network and the lesion classification network to obtain a mutually guided network model; The fusion of the lesion segmentation network and the lesion classification network includes: using the lesion segmentation network output as the input of the lesion classification network; Inputting the pseudo labels output by the lesion classification network into the lesion segmentation network for training the lesion segmentation network; The mutually guided network model is trained by a sample skin image set, and the trained mutually guided network model is used to determine the classification of skin lesions in an image, wherein the sample skin image set includes a first sample image set and a second sample image set, wherein the images in the first sample image set are provided with pixel-level segmentation annotations, and the images in the second sample image set are provided with image set annotations, and the mutually guided network model is trained by the sample skin image set, comprising: Training the lesion segmentation network using the first sample image set; The mutual guidance network model is trained using the second sample image set.

2. The method according to claim 1, characterized in that The lesion segmentation network is a U-Net network, which includes a convolutional layer, a maximum pooling layer, a deconvolution layer and a ReLU nonlinear activation function.

3. The method according to claim 2, characterized in that The encoding of the lesion segmentation network consists of ResNet50 pre-trained on the ImageNet dataset.

4. The method according to claim 3, characterized in that The first convolutional layer in the ResNet50 is replaced by one 7*7 convolutional layer with two 3*3 convolutional layers, and the two 3*3 convolutional layers are used to maintain the size of the image input.

5. The method according to claim 1, wherein The lesion classification network consists of an Xception network pre-trained on the ImageNet dataset.

6. The method according to claim 5, characterized in that The lesion classification network is composed of an improved Xception network, and the improved Xception network does not include the last pooling layer.

7. The method according to claim 6, characterized in that The last two layers of separable convolution in the improved Xception network are replaced by separable dilated convolution.

8. A skin lesion determination device, characterized in that: The device includes a memory and a processor, wherein the memory stores at least one program instruction, and the processor implements the following method by loading and executing the at least one program instruction: Get skin image; The skin image is input into a trained mutual guidance network model, and the classification of skin lesions in the skin image is output by the trained mutual guidance network model; the mutual guidance network model is trained by any one of the methods described in claims 1 to 7.

Citation Information

Patent Citations

  • Skin image processing method based on deep learning

    CN111951235A

  • Pathological image processing method and device, storage medium and processor

    CN112017162A

  • AMD lesion OCT image classification segmentation method and system based on bidirectional guide network

    CN113160226A