Neural Network-Based Dermoscopic Image Classification Method and System

By introducing coordinate attention and multi-head self-attention mechanisms into the neural network dermatoscope image classification method, the problems of large amount of parameters and low accuracy in the prior art are solved, and efficient dermatoscope image classification in resource-limited environments are achieved.

CN119339172BActive Publication Date: 2025-05-27ZHONGYU ZHICHUANG (SUZHOU) PHARMACEUTICAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411886012.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-27
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

In the prior art, dermatoscope image classification has the problem of large amount of algorithm parameters and low classification accuracy, which is difficult to popularize in medical institutions with limited resources.

Method used

A dermatoscope image classification method based on neural network is adopted to obtain pre-trained image classification model, and introduce coordinate attention processing module and multi-head self-attention feature extraction submodule into the model to reduce the amount of model parameters and improve classification accuracy.

Benefits of technology

It significantly reduces the model size, is suitable for deployment on embedded devices, and achieves a classification accuracy of 95.54% on the HAM10000 dataset, improving the classification performance and efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339172B_ABST
    Figure CN119339172B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image classification, and particularly to a dermoscopic image classification method and system based on a neural network, including obtaining spatially enhanced features from dermoscopic images; inputting the spatially enhanced features into a deep convolutional unit for local feature extraction of a single channel to obtain first local features of each channel, and performing linear combination between channels to obtain a first local feature map; inputting the first local feature map into the deep convolutional unit for global feature extraction of a single channel to obtain first global features of each channel; and performing linear combination between channels to obtain a first global feature map; inputting the spatially enhanced features, the first local feature map and the first global feature map into a feature fusion module for feature fusion to obtain a fused feature map; and classifying the dermoscopic images according to the fused feature map. The dermoscopic image classification method and system of the present invention significantly reduce the number of model parameters, and at the same time, the classification accuracy is improved to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image classification, and in particular to a dermoscopic image classification method and system based on a neural network. Background Art

[0002] Traditional means for screening skin diseases and skin cancers mainly include dermoscopy and pathological tissue examination; a dermoscope, also known as an epidermal transillumination microscope, can provide clear images of fine structures such as the living epidermis that are difficult to observe with the naked eye. Compared with the accuracy rate of about 70% for naked-eye diagnosis, the pathological coincidence rate of skin lesions diagnosed by dermoscopy is as high as 95%. By analyzing dermoscopic images, the medical community can detect early signs of skin cancer, thereby improving the cure rate of patients. However, manually analyzing these images is not only time-consuming and laborious, but also requires doctors to have rich clinical experience, which is prone to misdiagnosis.

[0003] In recent years, with the rapid development of deep learning technology, image classification methods based on neural networks have gradually become the mainstream means in dermoscopic image analysis. These technologies use deep learning models such as convolutional neural networks (CNNs) to extract features and classify dermoscopic images, promising to improve the accuracy and efficiency of classification. However, some deficiencies in the existing technologies in applications are still significant.

[0004] First of all, many dermoscopic image classification methods based on neural networks have the problem of a large number of algorithm parameters. Deep learning models usually consist of multiple layers, and the number of parameters in each layer often reaches millions or even tens of millions. This not only results in the need for a large amount of computing resources in the training and inference processes of the model, but also limits its application to environments with high-performance hardware and is difficult to popularize in medical institutions with relatively limited resources.

[0005] Secondly, although some advanced network structures have been proven to be able to extract rich features, in the actual application of dermoscopic image classification, the classification accuracy is still not high; this is because the lesion areas in dermoscopic images may be affected by problems such as uneven illumination, occlusion by hair or foreign objects, and unclear skin texture, resulting in the model being unable to effectively capture key features during feature extraction. In addition, the high similarity between diseased skin and benign moles also affects the accuracy of the algorithm. Existing classification models have limitations and relatively few reliable datasets. Summary of the Invention

[0006] Therefore, the technical problem to be solved by the present invention is to overcome the problems in the prior art that the dermoscopic image classification has a large number of algorithm parameters and low classification accuracy, and to provide a dermoscopic image classification method and system based on a neural network, which significantly reduces the number of model parameters and at the same time improves the classification accuracy to a certain extent, so as to better play a role in an environment with low computing resources.

[0007] In a first aspect, to solve the above technical problems, the present invention provides a dermoscopic image classification method based on a neural network, including,

[0008] Obtain a pre-trained image classification model; wherein, the image classification model includes a coordinate attention processing module and a plurality of feature extraction modules connected in sequence along the forward propagation direction, and each of the feature extraction modules includes a multi-head self-attention feature extraction sub-module; each of the multi-head self-attention feature extraction sub-modules includes a local feature extraction block, a global feature extraction block, and a feature fusion module; both the local feature extraction block and the global feature extraction block include a depth convolution unit and a pointwise convolution unit;

[0009] Input the dermoscopic image to be classified into the coordinate attention processing module to obtain a spatially enhanced feature;

[0010] Input the spatially enhanced feature into the depth convolution unit of the local feature extraction block of the first feature extraction module for single-channel local feature extraction to obtain first local features for each channel;

[0011] Input the first local features for each channel into the pointwise convolution unit of the local feature extraction block of the first feature extraction module for inter-channel linear combination to obtain a first local feature map;

[0012] Input the first local feature map into the depth convolution unit of the global feature extraction block of the first feature extraction module for single-channel global feature extraction to obtain first global features for each channel;

[0013] Input the first global features for each channel into the pointwise convolution unit of the global feature extraction block of the first feature extraction module for inter-channel linear combination to obtain a first global feature map;

[0014] Input the spatially enhanced feature, the first local feature map, and the first global feature map into the feature fusion module for feature fusion to obtain a fused feature map;

[0015] Use the fused feature map as the input of the next feature extraction module; classify the dermoscopic image according to the fused feature map output by the last feature extraction module.

[0016] In an embodiment of the present invention, inputting the dermoscopic image to be classified into the coordinate attention processing module to obtain a spatially enhanced feature includes,

[0017] Perform preliminary processing on the dermoscopic image to obtain preliminary features;

[0018] Perform an average pooling operation on the preliminary features in the horizontal direction to extract X-direction features;

[0019] Perform an average pooling operation on the preliminary features in the vertical direction to extract Y-direction features;

[0020] Perform concatenation on the X-direction features and the Y-direction features to obtain concatenated features;

[0021] Perform segmentation on the concatenated features to obtain two parts of segmented features;

[0022] Perform convolution operations on the two parts of segmented features respectively to obtain X-direction weight information and Y-direction weight information; wherein, the weight information represents the importance of the features at the corresponding positions;

[0023] Multiply the X-direction weight information by the preliminary features to obtain X-direction weighted features; multiply the Y-direction weight information by the preliminary features to obtain Y-direction weighted features; take the sum of the X-direction weighted features and the Y-direction weighted features as the weighted features;

[0024] Introduce a residual block, and obtain the spatial enhancement features according to the weighted features and the preliminary features; wherein, the spatial enhancement features are calculated in the following manner:

[0025] ;

[0026] wherein, F is the spatial enhancement feature; F1 is the weighted feature; F2 is the preliminary feature; is the weight parameter.

[0027] In an embodiment of the present invention, the weight parameter is obtained through learning, and it includes,

[0028] Initialize the weight parameter ; the weight parameter is defined as a vector having the same size as the preliminary features, and each element corresponds to a feature channel;

[0029] Add a network model in the residual block, and obtain the weight parameter by training the network model by adjusting hyperparameters;

[0030] wherein, training the network model includes,

[0031] Define a hyperparameter space, and set candidate values for each hyperparameter to be optimized;

[0032] Create a model function, and the model function is used to receive a set of the hyperparameters and return model performance metrics;

[0033] Train and validate all combinations of hyperparameters, and record the performance metrics for each set of hyperparameters;

[0034] Select the optimal set of hyperparameters based on the performance on the validation set;

[0035] Obtain the weight parameters according to the optimal set of hyperparameters.

[0036] In an embodiment of the present invention, the dermoscopic image is preliminarily processed to obtain preliminary features, including

[0037] Perform a channel dimension increase operation on the dermoscopic image, and then perform normalization and activation after the channel dimension increase operation to obtain channel dimension increase features;

[0038] Perform a diffusion separable convolution operation on the channel dimension increase features to obtain convolution features;

[0039] Perform normalization and activation on the convolution features to obtain the preliminary features;

[0040] After obtaining the spatial enhancement features, it further includes performing a channel dimension reduction operation on the spatial enhancement features, and inputting the result of the channel dimension reduction operation into the feature extraction module.

[0041] In an embodiment of the present invention, the depth convolution unit of the local feature extraction block includes a 5*5 diffusion separable convolution layer and a 7*7 dilated diffusion separable convolution layer; and along the forward propagation direction, the 5*5 diffusion separable convolution layer and the 7*7 dilated diffusion separable convolution layer are arranged in sequence; wherein, the dilation rate of the 7*7 dilated diffusion separable convolution layer is configured as 3.

[0042] In an embodiment of the present invention, a multi-head self-attention mechanism is embedded in the global feature extraction block, and the global feature extraction block further has a point convolution block, and the point convolution block is used to simulate the interaction operation between multiple heads in the multi-head self-attention mechanism.

[0043] In an embodiment of the present invention, the process of performing global feature extraction based on the multi-head self-attention mechanism includes

[0044] Perform dimension reduction on the input feature map based on the diffusion separable convolution operation;

[0045] Map the dimension-reduced feature map to query Q, key K, and value V respectively through linear transformation;

[0046] Obtain attention scores based on query Q and key K;

[0047] Apply the Softmax function to normalize the attention scores to obtain attention weights;

[0048] Multiply the attention weight by the value V to obtain a weighted value;

[0049] Perform normalization on the weighted value, and perform global feature extraction according to the result of the normalization;

[0050] Among them, the self-attention of the multi-head self-attention mechanism is calculated according to the following formula:

[0051] ;

[0052] Among them, Q represents the query of the multi-head self-attention mechanism; K represents the key of the multi-head self-attention mechanism; V represents the value of the multi-head self-attention mechanism; d k represents the dimension of K; K T represents the transpose of K; Conv() represents a diffusion separable convolution operation; Softmax() represents the Softmax function; IN() represents performing normalization.

[0053] In an embodiment of the present invention, the training process of the image classification model includes,

[0054] Obtain an image sample set;

[0055] Construct an image classification loss function based on the output of the last feature extraction module;

[0056] Use the image samples in the image sample set to iteratively train the coordinate attention processing module and multiple feature extraction modules until the value of the image classification loss function is minimized, and obtain a trained image classification model;

[0057] Among them, the image classification loss function is:

[0058] ;

[0059] i represents the sample serial number; n represents the total number of samples; y i represents the true value probability of the i-th sample; represents the probability that the i-th sample is evaluated as a positive sample; ln represents the natural logarithm function.

[0060] In an embodiment of the present invention, the image classification model further includes a coordinate attention downsampling module; among them, the output of the coordinate attention processing module is connected to the coordinate attention downsampling module; the output of the feature extraction module is connected to the coordinate attention downsampling module; the coordinate attention downsampling module is used to perform a downsampling operation on the input features connected.

[0061] Second aspect, to solve the above technical problems, the present invention provides a dermoscopic image classification system based on a neural network for performing the dermoscopic image classification method based on a neural network.

[0062] The above technical solutions of the present invention have the following beneficial effects compared with the prior art:

[0063] For the dermoscopic image classification method and system based on a neural network of the present invention, both the local feature extraction block and the global feature extraction block decompose a complete convolution into a depth convolution and a pointwise convolution. First, the depth convolution is used to extract the features of each channel by performing convolution operations only within a single channel, and then the pointwise convolution is used to perform a linear combination of the features of each channel between channels; the convolution operation within a single channel significantly reduces the amount of computation and the number of parameters compared with the inter-channel convolution operation of a general convolution network; in addition, although a significant reduction in the number of parameters will theoretically weaken the accuracy of image classification, in this solution, a coordinate attention mechanism is added to the residual structure of the network to increase the sensitivity to position information and enhance the information in specific regions, making the model more focused on key features and enhancing the network's ability to fuse high-dimensional features, ultimately achieving the purpose of improving the network classification performance; and a multi-head self-attention mechanism is introduced to process local and global feature information, optimizing the computational complexity of the algorithm while increasing the feature interaction between multiple heads, capturing the global dependencies in the feature map, and enhancing the network's ability to extract global information, thereby indirectly improving the efficiency and accuracy of model classification.

[0064] In summary, for the dermoscopic image classification method and system based on a neural network of the present invention, the model size is significantly reduced to 0.83 MB, and at the same time, a classification accuracy of 95.54% is achieved for the HAM10000 dataset, making it more suitable for deployment on embedded devices for clinical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to specific embodiments of the present invention in conjunction with the accompanying drawings, where

[0066] Figure 1 is a flowchart of the dermoscopic image classification method based on a neural network in a preferred embodiment of the present invention;

[0067] Figure 2 is an overall framework diagram of the image classification model in a preferred embodiment of the present invention;

[0068] Figure 3 is a structural block diagram of the MV2 module with a coordinate attention mechanism added in a preferred embodiment of the present invention;

[0069] Figure 4 is Figure 3The structural block diagram of coordinate attention in the MV2 module with coordinate attention mechanism added as shown;

[0070] Figure 5 The structural block diagram of the multi-head self-attention feature extraction sub-module in the preferred embodiment of the present invention;

[0071] Figure 6 The flow chart of obtaining spatially enhanced features in the preferred embodiment of the present invention;

[0072] Figure 7 The accuracy graph of the test set;

[0073] Figure 8 The loss graph of the test set;

[0074] Figure 9 The schematic diagram of the confusion matrix;

[0075] Explanation of the reference numerals in the specification drawings: 100 - Coordinate attention processing module;

[0076] 200 - Feature extraction module; 201 - 5*5 dilated separable convolution layer; 202 - 7*7 dilated separable convolution layer; 210 - Local feature extraction block; 220 - Global feature extraction block; 230 - Feature fusion module;

[0077] 300 - Coordinate attention downsampling module. Detailed implementation manners

[0078] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited are not intended to limit the present invention.

[0079] Embodiment 1

[0080] The embodiment of the present invention discloses a dermoscopic image classification method based on a neural network, including,

[0081] Obtaining a pre-trained image classification model; this step requires pre-training to obtain an image classification model, and the image classification model includes a coordinate attention processing module 100, a plurality of feature extraction modules 200 connected in sequence along the forward propagation direction, and a coordinate attention downsampling module 300;

[0082] Specifically, referring to Figure 2As shown, the image classification model includes a 3*3 convolution module, a coordinate attention processing module 100, a coordinate attention downsampling module 300, a coordinate attention processing module 100, a coordinate attention downsampling module 300, a feature extraction module 200, a coordinate attention downsampling module 300, a feature extraction module 200, a coordinate attention downsampling module 300, a feature extraction module 200, a 1*1 convolution module, and a global pooling layer, which are connected in sequence along the forward propagation direction. The dermoscopic images to be classified are processed through the above modules in sequence, and finally the classification results are output through the global pooling layer.

[0083] Among them, the coordinate attention processing module 100 is configured as an MV2 module with a coordinate attention mechanism, which is a variant of MobileNetV2 and incorporates the Coordinate Attention mechanism; the feature extraction module 200 is configured as an Efficient-MobileViT Block with a multi-head self-attention mechanism; the coordinate attention downsampling module 300 is configured as an MV2 module with a coordinate attention mechanism.

[0084] Refer to Figure 3 As shown, the MV2 module with a coordinate attention mechanism performs the following operations in sequence along the forward propagation direction:

[0085] The first 1×1 convolution layer is used to adjust the number of channels;

[0086] Bn: Normalization;

[0087] ReLU6: Activation function, used to introduce non-linearity;

[0088] MV2 block: MobileNetV2 block, which contains a depthwise separable convolution. The depthwise separable convolution can reduce the number of parameters when extracting features;

[0089] Coordinate Attention (CA): Coordinate attention mechanism, used to enhance the spatial information in the feature map;

[0090] Add: Residual connection, adding the input feature map to the processed feature map;

[0091] The second 1×1 convolution layer: Used to adjust the number of channels.

[0092] In summary, the processing flow of the MV2 module with the coordinate attention mechanism includes: the input feature map is adjusted in the number of channels through the first 1×1 convolutional layer, and the ReLU activation function introduces non-linearity; it is processed through the MV2 block, including depthwise separable convolution and the ReLU activation function; it is processed through the coordinate attention mechanism (Coordinate Attention); then it is processed through the second 1×1 convolutional layer and the ReLU activation function; finally, it is added to the input feature map through a residual connection (Add) to output the processed feature map.

[0093] Refer to Figure 4 As shown, the coordinate attention is carried out in sequence along the forward propagation direction:

[0094] X Avg Pool: Perform an average pooling operation on the input feature map in the horizontal direction;

[0095] Y Avg Pool: Perform an average pooling operation on the input feature map in the vertical direction;

[0096] Concat: Concatenate the pooling results in the horizontal and vertical directions;

[0097] Split: Split the feature map into two parts;

[0098] Alpha: Perform a weighted processing on the split feature map;

[0099] Add: Add the weighted processed feature map to the original input feature map.

[0100] In summary, the processing flow of the coordinate attention mechanism includes: the input feature map undergoes average pooling operations in the horizontal and vertical directions to extract features in the horizontal and vertical directions; the features in the horizontal and vertical directions are concatenated to retain more spatial information; it is processed through a residual block, including multiple convolutional layers and activation functions, for feature extraction and enhancement; the feature map is split into two parts for weighted processing; finally, it is added to the input feature map through the residual connection Add to output the processed feature map.

[0101] Refer to Figure 5 As shown, each of the multi-head self-attention feature extraction sub-modules includes a local feature extraction block 210, a global feature extraction block 220, and a feature fusion module 230; the local feature extraction block 210 includes multiple convolutional units and a coordinate attention mechanism for extracting local features; among them, the convolutional units include depth convolutional units and pointwise convolutional units; the global feature extraction block 220 includes an efficient multi-head self-attention mechanism EMSA and convolutional units for extracting global features; among them, the convolutional units include depth convolutional units and pointwise convolutional units; the feature fusion module 230 includes a concatenation operation, a convolutional layer, and a coordinate attention mechanism for fusing local features and global features.

[0102] The dermoscopic image classification method according to the embodiment of the present invention performs automatic classification through the image classification model obtained by the above training, referring to Figure 1 as shown, the classification method specifically includes the following steps:

[0103] S100. Input the dermoscopic image to be classified into the coordinate attention processing module to obtain spatially enhanced features;

[0104] S200. Input the spatially enhanced features into the depth convolution unit of the local feature extraction block of the first feature extraction module to perform local feature extraction for a single channel, and obtain first local features for each channel;

[0105] S300. Input the first local features for each channel into the pointwise convolution unit of the local feature extraction block of the first feature extraction module to perform linear combination between channels, and obtain a first local feature map;

[0106] S400. Input the first local feature map into the depth convolution unit of the global feature extraction block of the first feature extraction module to perform global feature extraction for a single channel, and obtain first global features for each channel;

[0107] S500. Input the first global features for each channel into the pointwise convolution unit of the global feature extraction block of the first feature extraction module to perform linear combination between channels, and obtain a first global feature map;

[0108] S600. Input the spatially enhanced features, the first local feature map, and the first global feature map into the feature fusion module to perform feature fusion, and obtain a fused feature map;

[0109] S700. Use the fused feature map as the input of the next feature extraction module; classify the dermoscopic image according to the fused feature map output by the last feature extraction module.

[0110] In the dermoscopic image classification method based on a neural network according to the present invention, both the local feature extraction block and the global feature extraction block decompose a complete convolution into a depth convolution and a pointwise convolution. First, the depth convolution is used to perform convolution operations only within a single channel to extract the features of each channel, and then the pointwise convolution is used to perform a linear combination of the features of each channel across channels; the convolution operation within a single channel significantly reduces the computational amount and the number of parameters compared to the cross-channel convolution operation of a general convolution network; in addition, the significant reduction in the number of parameters will theoretically weaken the accuracy of image classification, but in this solution, an improved coordinate attention mechanism is added to the residual structure of the network, adding position information sensitivity and enhancing the information in specific regions, improving the adaptability to various illuminations, colors, and texture transformations, making the model more focused on key features, enhancing the network's ability to fuse high-dimensional features, and ultimately achieving the purpose of improving the network classification performance; and a multi-head self-attention mechanism is introduced to process local and global feature information, optimizing the computational complexity of the algorithm while increasing the feature interaction between multiple heads, capturing the global dependencies in the feature map, and enhancing the network's ability to extract global information, thereby indirectly improving the efficiency and accuracy of model classification.

[0111] Specifically, in an embodiment of the present invention, the dermoscopic image to be classified is input into the coordinate attention processing module to obtain a spatially enhanced feature, as shown in Figure 6 and it includes the following steps:

[0112] Perform preliminary processing on the dermoscopic image to obtain preliminary features; perform average pooling operations in the horizontal direction on the preliminary features to extract X-direction features; perform average pooling operations in the vertical direction on the preliminary features to extract Y-direction features; perform concatenation on the X-direction features and the Y-direction features to obtain concatenated features; perform segmentation on the concatenated features to obtain two parts of segmented features; perform convolution operations on the two parts of segmented features respectively to obtain X-direction weight information and Y-direction weight information; where the weight information represents the importance of the features at the corresponding positions; multiply the X-direction weight information by the preliminary features to obtain X-direction weighted features; multiply the Y-direction weight information by the preliminary features to obtain Y-direction weighted features; take the sum of the X-direction weighted features and the Y-direction weighted features as the weighted features; introduce a residual block, and obtain the spatially enhanced feature according to the weighted features and the preliminary features; where the spatially enhanced feature is calculated as follows:

[0113] ;

[0114] where F is the spatially enhanced feature; F1 is the weighted feature; F2 is the preliminary feature; is the weight parameter.

[0115] In a specific application scenario, first, features are extracted from the X and Y directions respectively. Through splicing and splitting operations, useful information in different directions can be captured simultaneously, thereby enhancing the diversity of features. This enables the model to provide a more comprehensive feature description when faced with complex skin lesions. Second, by calculating the weight information in the X and Y directions and multiplying it with the preliminary features to obtain weighted features, this process dynamically adjusts through the weight parameters learned ; this feature enhancement mechanism can adaptively emphasize important features and suppress irrelevant features, thereby improving the feature representation ability and its contribution to classification decisions. Third, by introducing residual blocks, the problem of gradient disappearance in deep networks can be effectively alleviated, promoting the transmission of information flow, and ensuring the full utilization of features in the network. This enables the model to not only have faster training convergence but also maintain high performance in deep networks. Finally, through effective feature extraction and weighting mechanisms, the number of parameters to be processed is reduced. Compared with traditional deep learning models, it can significantly reduce the computational cost while maintaining a high classification accuracy.

[0116] Furthermore, different from setting weight parameters based on conventional expert experience , the weight parameters in the embodiments of the present invention are obtained through learning, and it includes

[0117] initializing the weight parameters ; the weight parameters are defined as vectors of the same size as the preliminary features, and each element corresponds to a feature channel; a network model is added in the residual block, and the weight parameters are obtained by training the network model by adjusting hyperparameters;

[0118] wherein, training the network model includes defining a hyperparameter space, setting candidate values for each hyperparameter to be optimized; creating a model function, which is used to receive a set of the hyperparameters and return model performance metrics; training and validating all hyperparameter combinations, recording the performance metrics under each set of hyperparameters; selecting the optimal set of hyperparameters according to the performance on the validation set; obtaining the weight parameters according to the optimal set of hyperparameters.

[0119] Above, in the embodiments of the present invention, first, configure the weight parameters as vectors of the same size as the preliminary features, and each element corresponds to a feature channel; different weights can be assigned to different channels, enabling the model to more finely adjust the importance of features. Second, through the learning of adaptive weight parameters and the effective optimization of hyperparameters, not only the accuracy and efficiency of dermoscopic image classification are improved, but also the parameter tuning burden is reduced, and the robustness and applicability of the model are enhanced.

[0120] Further, perform preliminary processing on the dermoscopic image to obtain preliminary features, including performing a channel dimension increase operation on the dermoscopic image, and after the channel dimension increase operation, performing normalization and activation to obtain channel dimension increase features; performing a diffusion separable convolution operation on the channel dimension increase features to obtain convolution features; performing normalization and activation on the convolution features to obtain the preliminary features; after obtaining the spatial enhancement features, further include performing a channel dimension reduction operation on the spatial enhancement features, and inputting the result of the channel dimension reduction operation into the feature extraction module.

[0121] Based on the above specific embodiments, performing a channel dimension increase operation on the input dermoscopic image increases the number of channels of the feature map, enabling the model to capture more local information and enhancing the feature representation ability; performing normalization and activation immediately after the channel dimension increase can accelerate the training process and make the training more stable; normalization helps to alleviate the problem of gradient vanishing or explosion, while the activation function introduces non-linearity and improves the expression ability of the model; the diffusion separable convolution decomposes the standard convolution into two steps of depth convolution and pointwise convolution, significantly reducing the amount of computation and the number of parameters; in addition, although the amount of computation is reduced, the diffusion separable convolution can still effectively capture local features, ensuring the quality of feature extraction; finally, performing a channel dimension reduction operation reduces the number of channels of the feature map, thereby further reducing the size and number of parameters of the model.

[0122] In the publicly disclosed MobileViT block, the local feature extraction block (Local Represention) uses 3*3 convolution and 1*1 convolution for local representation, and the receptive field of the 3*3 convolution is too small; however, if large kernel convolution is used, although the receptive field is increased, the amount of computation and the number of parameters are also increased significantly. To solve the above contradiction, refer to Figure 5As shown in the figure, the depth convolution unit of the local feature extraction block in the solution of the embodiment of the present invention includes a 5*5 diffusion separable convolution layer 201 and a 7*7 dilated diffusion separable convolution layer 202; and along the forward propagation direction, the 5*5 diffusion separable convolution layer 201 and the 7*7 dilated diffusion separable convolution layer 202 are arranged in sequence; among them, the dilation rate of the 7*7 dilated diffusion separable convolution layer 202 is configured to be 3. The combination of a 5*5 diffusion separable convolution layer 201 and a 7*7 dilated diffusion separable convolution layer 202 with a dilation rate of 3 achieves the effect of a 21*21 convolution. Specifically, it can be understood as follows: decomposing the large kernel 21*21 convolution, and the decomposition principle is: decomposing the N*N large kernel convolution into a (2d - 1)*(2d - 1) diffusion separable convolution and a (N / d)*(N / d) dilated diffusion separable convolution, where d is the dilation rate; based on this decomposition principle, the 21*21 convolution is decomposed into a 5*5 diffusion separable convolution and a 7*7 dilated diffusion separable convolution with a dilation rate of 3; where N is taken as 21; d is taken as 3.

[0123] For a feature map with a size of H*W, the parameter quantity and floating-point operation quantity of a conventional convolution are as follows:

[0124] P = K*K*C i *C o ;

[0125] F = K*K*C i *C o *H*W;

[0126] P is the parameter quantity; F is the floating-point operation quantity; K is the size of the convolution kernel; C i is the number of input channels of the feature map; C o is the number of output channels of the feature map; H is the height of the feature map; W is the width of the feature map.

[0127] The parameter quantity and floating-point operation quantity after decomposing the large kernel convolution are as follows:

[0128] ;

[0129] ;

[0130] Among them, P is the parameter quantity; F is the floating-point operation quantity; K is the size of the convolution kernel; d is the dilation rate; C i is the number of input channels of the feature map; C o is the number of output channels of the feature map; H is the height of the feature map; W is the width of the feature map.

[0131] It can be seen that when the dilation rate d increases, the increase in the denominator will cause the floating-point operation amount and the number of parameters to decrease correspondingly. Therefore, using the decomposed large kernel convolution can not only increase the receptive field of the Local Represention part, but also reduce the resource occupation, compensating for the increase in the number of parameters and computational amount caused by the introduction of the coordinate attention mechanism.

[0132] Furthermore, a multi-head self-attention mechanism is embedded in the global feature extraction block, and the global feature extraction block also has a point convolution block, which is used to simulate the interaction operations between multiple heads in the multi-head self-attention mechanism. By using the point convolution block to simulate the interaction operations between multiple heads in the multi-head self-attention mechanism, the number of parameters of the model is further reduced.

[0133] The process of performing global feature extraction based on the multi-head self-attention mechanism includes

[0134] Performing dimensionality reduction on the input feature map based on the diffusion separable convolution operation; linearly transforming the dimension-reduced feature map to map it to the query Q, key K, and value V respectively; obtaining attention scores based on the query Q and key K; applying the Softmax function to normalize the attention scores to obtain attention weights; multiplying the attention weights by the value V to obtain the weighted value; performing implementation normalization on the weighted value, and performing global feature extraction according to the result of the implementation normalization; where, the self-attention of the multi-head self-attention mechanism is calculated according to the following formula:

[0135] ;

[0136] where, Q represents the query of the multi-head self-attention mechanism; K represents the key of the multi-head self-attention mechanism; V represents the value of the multi-head self-attention mechanism; d k represents the dimension of K; K T represents the transpose of K; Conv() represents the diffusion separable convolution operation; Softmax() represents the Softmax function; IN() represents the implementation normalization.

[0137] Based on the above specific embodiments, the solution of the embodiment of the present invention reduces the computational complexity and the number of parameters by optimizing the standard multi-head self-attention mechanism, while maintaining or improving the performance of the model.

[0138] Furthermore, the training process of the image classification model includes obtaining an image sample set; constructing an image classification loss function based on the output of the last feature extraction module; using the image samples in the image sample set to iteratively train the coordinate attention processing module and multiple feature extraction modules until the value of the image classification loss function is the smallest, and obtaining a trained image classification model; where, the image classification loss function is:

[0139] ;

[0140] i represents the sample serial number; n represents the total number of samples; y i represents the true value probability of the i-th sample; represents the probability that the i-th sample is evaluated as a positive sample; ln represents the natural logarithm function.

[0141] To verify the effectiveness of the image classification model provided by this application, the coordinate attention processing module and the feature extraction module in the model are used in the embodiments of this application to replace the preprocessing module and the feature extraction module in various deep learning networks, and experiments are carried out on several visual benchmark datasets. The following is a specific description of the experiments:

[0142] The dataset used for training and testing is the HAM dataset, with a quantity of 10,015; a resolution of 600×450, and it contains dermoscopic images of 7 different categories of pixels, namely: melanocytic nevi (NV), basal cell carcinoma (BCC), melanoma (MEL), benign keratosis (BKL), dermatofibroma (DF), vascular skin lesion (VASC), and actinic keratosis (AKIEC). Among them, except for basal cell carcinoma (BCC) and melanoma (MEL) which are malignant tumors, the rest of the skin-related lesions are benign. The diagnosis of all melanomas is confirmed by histopathological evaluation of biopsies, and the diagnosis of melanocytic nevi is confirmed by histopathological examination (24%), expert consensus (54%), or other diagnostic methods.

[0143] The experiments and the software and hardware are shown in Table 1 below. When training, the dataset is divided into a training set, a validation set, and a test set according to the ratio of 8:1.5:0.5, and there is no intersection between the training set and the validation set. The experiment sets hyperparameters to train the network, with a batch size of 128, the model loss function is set to the cross-entropy loss function, the learning rate is 0.0002, the optimizer is selected as AdamW, and the number of training epochs is 150.

[0144] Table 1

[0145] Hardware Software CPU: 8 × Xeon E5-2686 v4 Matplotlib 3.3.4 Memory: 30 GB Visual Studio code 1.78.0 GPU: NVIDIA RTX A4000 Pytorch 1.12.1 Video memory: 16 GB CUDA 11.3 Operating system: Ubuntu 18.04 Python3.9

[0146] Comprehensive evaluation is carried out from four dimensions: precision, recall, specificity, and F1 value. The specific evaluation results are shown in Table 2 below:

[0147] Table 2

[0148] Clinical diagnosis Precision Recall Specificity F1 AKIEC 98.35 97.35 99.73 97.84 BCC 96.00 98.20 99.31 97.08 BKL 93.98 88.37 99.06 91.08 DF 99.09 99.90 99.85 99.49 MEL 93.18 90.02 98.89 91.57 NV 88.85 95.12 97.98 91.88 VASC 99.90 100.00 99.98 99.95

[0149] Among them, the units of precision, recall, specificity, and F1 are all %.

[0150] The image classification model proposed in the solution of the present invention achieved an overall accuracy of 95.32, an average precision of 95.35%, an average recall of 95.35%, an average specificity of 99.22%, and an average F1 value of 95.35% for the classification task of the HAM dataset. As can be seen from Table 2, the classification effect of the VASC class is the best, exceeding other classes in all indicators. As Figure 7 shown is the accuracy graph of the test set of the image classification model, as Figure 8 shown is the loss graph of the test set of the image classification model. The number of training rounds is 150. It can be seen that as the number of training rounds increases, the accuracy steadily improves and finally stabilizes, reaching a maximum of 95.54%. The loss value gradually decreases and stabilizes near 0.13. Experiments prove that the model of the implementation solution of the present invention can effectively classify and predict the HAM10000 dataset.

[0151] The confusion matrix of the image classification model of the present invention for each skin condition is as Figure 9 shown. According to Figure 9 it can be clearly seen that the classification effect of VASC is the best among all types. Except for a very small confusion with the NV class, there are almost no other errors, and the recall rate reaches 100%. The classification effect of the DF class is also good, with an accuracy as high as 99.09%. The recognition effects of VASC and DF classes are the best, with F1 values as high as 99.95% and 99.49%. These two clinical diagnoses are relatively obvious under dermoscopy. DF is a dermatofibroma, a benign tumor with obvious signs of fibrous hyperplasia, which is relatively special among other types. And VASC is a vascular skin disease, with obvious blood vessels shown in the dermoscopy image, which is also relatively special compared with other types of lesions and is easy to identify. It can be seen from the confusion matrix that the confusion degree between NV and MEL is relatively high. NV is a benign nevus cell, and MEL is a malignant cell developed from a dysplastic nevus cell. Therefore, there is a smooth transition process between NV and MEL, and they may develop from the same origin, which makes it challenging to identify them. The implementation solution of the present invention makes the clinical manifestations of each type increase as much as possible in various evaluation indicators, and the specificity reaches more than 95%, proving that the method implemented by the present invention effectively excludes confusing targets and is more robust.

[0152] Ablation experiment:

[0153] To explore the effects of the coordinate attention mechanism, the multi-head self-attention mechanism, and the decomposed large kernel convolution on network improvement, ablation experiments were set up. The impacts of different methods on the model are shown in Table 3 below:

[0154] After adding the CA attention mechanism to the image classification model, the number of parameters increased by approximately 0.01 mb, and the floating-point computational volume increased by 0.8%. However, the model accuracy increased by 2.39%. After adding the decomposed large kernel convolution to the model, the number of parameters decreased by 0.1 mb, the floating-point operation volume decreased by approximately 8.24%, and the performance was slightly improved. Then, the efficient multi-head attention mechanism was added to make the model's number of parameters reach 0.82 mb. The floating-point operation rate decreased by approximately 15.71% compared to the original model, and the accuracy increased by approximately 1.48% compared to the original model. Using the decomposed large kernel convolution and the efficient multi-head attention mechanism better balanced the model inference speed brought by the CA coordinate attention. Therefore, the implementation scheme of the present invention combines the decomposed large kernel convolution, the efficient multi-head attention mechanism, and the CA coordinate attention mechanism to improve the model. The network precision increased by 3.51% compared to the original network, the network recall rate increased by 3.37% compared to the original network, the final accuracy increased by 3.42%, the computational complexity decreased by approximately 14.89%, and the number of parameters decreased by approximately 8.79%. The method implemented in the present invention has a good classification effect on the HAM10000 dataset.

[0155] Table 3

[0156] Method Final accuracy / % Precision / % Recall / % Number of parameters / MB <![CDATA[Floating-point operation amount / 10 6 > MobileViT 92.12 92.11 92.20 0.91 546.13 MobileViT + CA 94.51 94.50 94.54 0.92 550.60 MobileViT + EMSA 93.45 93.66 93.46 0.91 505.31 MobileViT + LKA 92.35 92.37 92.35 0.81 501.09 MobileViT + EMSA + LKA 93.58 93.58 93.61 0.82 460.28 The solution of the present invention 95.54 95.62 95.57 0.83 464.76

[0157] Example Two

[0158] In a second aspect, to solve the above technical problems, the present invention provides a dermoscopic image classification system based on a neural network for performing the dermoscopic image classification method based on a neural network, which can effectively implement the image classification of dermoscopic images. The technical effects that can be achieved are as described in the above embodiments and will not be elaborated here.

[0159] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0160] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0161] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0162] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0163] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to exhaustively list all implementations here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A dermatoscopic image classification method based on a neural network, characterized in that: include, Acquire a pre-trained image classification model; wherein the image classification model includes a coordinate attention processing module and a plurality of feature extraction modules sequentially connected along the forward propagation direction, each of the feature extraction modules includes a multi-head self-attention feature extraction submodule; each of the multi-head self-attention feature extraction submodule includes a local feature extraction block, a global feature extraction block and a feature fusion module; the local feature extraction block and the global feature extraction block both include a depth convolution unit and a point-by-point convolution unit; Inputting the dermoscopic image to be classified into the coordinate attention processing module to obtain spatial enhancement features; Inputting the spatial enhancement feature into the deep convolution unit of the local feature extraction block of the first feature extraction module to extract the local features of a single channel, and obtaining the first local features of each channel; Inputting the first local features of each channel into the point-by-point convolution unit of the local feature extraction block of the first feature extraction module to perform linear combination between channels to obtain a first local feature map; Inputting the first local feature map into the deep convolution unit of the global feature extraction block of the first feature extraction module to perform global feature extraction of a single channel to obtain a first global feature of each channel; Inputting the first global features of each channel into the point-by-point convolution unit of the global feature extraction block of the first feature extraction module to perform linear combination between channels to obtain a first global feature map; Inputting the spatial enhancement feature, the first local feature map and the first global feature map into the feature fusion module for feature fusion to obtain a fused feature map; Using the fused feature map as the input of the next feature extraction module; classifying the dermoscopic image according to the fused feature map output by the last feature extraction module; The dermoscopic image to be classified is input into the coordinate attention processing module to obtain spatial enhancement features, including: Performing preliminary processing on the dermoscopic image to obtain preliminary features; Performing an average pooling operation in the horizontal direction on the preliminary features to extract X-direction features; Performing an average pooling operation in the vertical direction on the preliminary features to extract features in the Y direction; Performing stitching on the X-direction feature and the Y-direction feature to obtain a stitching feature; Segmenting the splicing feature to obtain two parts of segmentation features; Performing convolution operations on the two parts of the segmentation features respectively to obtain X-direction weight information and Y-direction weight information; wherein the weight information represents the importance of the corresponding position feature; Multiplying the X-direction weight information by the preliminary feature to obtain an X-direction weighted feature; multiplying the Y-direction weight information by the preliminary feature to obtain a Y-direction weighted feature; taking the sum of the X-direction weighted feature and the Y-direction weighted feature as a weighted feature; A residual block is introduced, and the spatial enhancement feature is obtained according to the weighted feature and the preliminary feature; wherein the spatial enhancement feature is obtained by calculating in the following manner: ; Among them, F is the spatial enhancement feature; F1 is the weighted feature; F2 is the preliminary feature; is the weight parameter.

2. The neural network-based dermoscopic image classification method according to claim 1, characterized in that: The weight parameter Acquired through learning, which includes, Initialize the weight parameters ; The weight parameter is defined as a vector of the same size as the preliminary feature, each element corresponding to a feature channel; Add a network model to the residual block, and train the network model by adjusting hyperparameters to obtain the weight parameters; Wherein, training the network model includes: Define the hyperparameter space and set candidate values ​​for each hyperparameter that needs to be optimized; Creating a model function, the model function is used to receive a set of the hyperparameters and return a model performance indicator; Train and validate all hyperparameter combinations, and record the performance indicators under each set of hyperparameters; Select the best set of hyperparameters based on the performance on the validation set; The weight parameters are obtained according to the optimal set of hyperparameters.

3. The neural network-based dermoscopic image classification method according to claim 1, characterized in that: The dermoscopic image is preliminarily processed to obtain preliminary features, including: Performing a channel dimension-upgrading operation on the dermoscopic image, and then normalizing and activating the channel dimension-upgrading operation to obtain a channel dimension-upgrading feature; Performing a diffusion separable convolution operation on the channel dimension-raising feature to obtain a convolution feature; Normalizing and activating the convolutional features to obtain the preliminary features; After obtaining the spatial enhancement feature, the method further includes performing a channel dimensionality reduction operation on the spatial enhancement feature, and inputting the result of the channel dimensionality reduction operation into the feature extraction module.

4. The neural network-based dermoscopic image classification method according to claim 1, characterized in that: The depth convolution unit of the local feature extraction block includes a 5*5 diffusion separable convolution layer and a 7*7 dilated diffusion separable convolution layer; and along the forward propagation direction, the 5*5 diffusion separable convolution layer and the 7*7 dilated diffusion separable convolution layer are arranged in sequence; wherein the dilation rate of the 7*7 dilated diffusion separable convolution layer is configured to be 3.

5. The neural network-based dermoscopic image classification method according to claim 1, characterized in that: A multi-head self-attention mechanism is embedded in the global feature extraction block, and the global feature extraction block also has a point convolution block, which is used to simulate the interactive operation between multiple heads in the multi-head self-attention mechanism.

6. The neural network-based dermoscopic image classification method according to claim 5, characterized in that: The process of performing global feature extraction based on the multi-head self-attention mechanism includes: Reduce the dimension of the input feature map based on the diffusion separable convolution operation; Map the dimensionally reduced feature map to query Q, key K, and value V respectively through linear transformation; Get an attention score based on the query Q and key K; Apply the Softmax function to normalize the attention score and obtain the attention weight; Multiply the attention weight by the value V to obtain a weighted value; Normalizing the weighted values, and performing global feature extraction according to the normalized results; The self-attention of the multi-head self-attention mechanism is calculated according to the following formula: ; Where Q represents the query of the multi-head self-attention mechanism; K represents the key of the multi-head self-attention mechanism; V represents the value of the multi-head self-attention mechanism; d k represents the dimension of K; K T Represents the transpose of K; Conv() represents diffuse separable convolution; Softmax() represents the Softmax function; IN() represents the implementation of normalization.

7. The neural network-based dermoscopic image classification method according to claim 1, characterized in that: The training process of the image classification model includes: Get a sample set of images; Constructing an image classification loss function based on the output of the last feature extraction module; Iteratively training the coordinate attention processing module and the plurality of feature extraction modules using the image samples in the image sample set until the value of the image classification loss function is minimized, thereby obtaining a trained image classification model; Among them, the image classification loss function is: ; i represents the sample number; n represents the total number of samples; y i Represents the true value probability of the i-th sample; represents the probability that the i-th sample is evaluated as a positive sample; ln represents the natural logarithm function.

8. The neural network-based dermoscopic image classification method according to claim 1, characterized in that: The image classification model also includes a coordinate attention downsampling module; wherein the output of the coordinate attention processing module is connected to the coordinate attention downsampling module; the output of the feature extraction module is connected to the coordinate attention downsampling module; the coordinate attention downsampling module is used to perform downsampling operations on the connected input features.

9. A dermoscopic image classification system based on a neural network, characterized in that: Used to perform the neural network-based dermoscopic image classification method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Medical image classification method and classification device thereof

    CN112700434A

  • Anti-noise interference dermatoscope image cancer focus segmentation method and system, and medium

    CN118037744A