A Multi-classification System for Children's Skin Diseases Based on Improved ResNet

Through the improved ResNet network model, combined with multiple convolutional structures, the problem of insufficient accuracy of image classification of children's skin disease is solved, efficient automatic diagnosis assistance is achieved, and diagnostic efficiency and accuracy are improved.

CN119380077BActive Publication Date: 2025-07-18ZHEJIANG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411400145.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-09
Publication Date
2025-07-18
Estimated Expiration
2044-10-09

AI Technical Summary

Technical Problem

Existing skin disease image classification techniques have problems with insufficient accuracy when facing the diversity and complexity of skin diseases in children, especially the similarity and image quality effects among common skin diseases in children lead to diagnosis difficulties.

Method used

The improved ResNet network model is adopted, combining residual module, mesoscale expansion convolution, wide-scale expansion convolution and multi-scale fusion cavity convolution structure, and automatically classifies skin disease images through feature extraction and prediction units, and uses feature fusion and standardization techniques to improve identification accuracy.

Benefits of technology

It improves the accuracy of identification and classification of skin diseases in children, provides doctors with auxiliary diagnostic suggestions, shortens diagnosis time, improves the diagnostic capabilities of primary medical institutions, and ensures that more children receive timely and accurate medical services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380077B_ABST
    Figure CN119380077B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-classification system for children's skin diseases based on an improved ResNet, which includes a computer memory and a classification model stored in the computer memory; the classification model includes a feature extraction unit and a prediction unit; the feature extraction unit includes a residual module structure, a medium-scale dilated convolution structure, a wide-scale dilated convolution structure, and a multi-scale fusion dilated convolution structure; the image sequentially passes through a convolutional layer, a max pooling layer, a convolutional layer, and a batch normalization layer to extract features, enters the above four structures used in parallel, and outputs their respective feature maps; the prediction unit fuses the feature maps output by each structure and inputs them into a batch normalization layer to standardize the fused features; then, the standardized feature maps are downsampled through an average pooling layer; finally, the features are transformed into a classification probability distribution of skin diseases through a softmax layer. By using the present invention, the recognition and classification accuracy of common children's skin diseases can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical artificial intelligence, and particularly relates to a multi-classification system for children's skin diseases based on improved ResNet. Background Art

[0002] Common skin diseases in children (such as atopic dermatitis, urticaria, hemangioma, diaper dermatitis, etc.) seriously affect the quality of life of children, and easily lead to limitations in the daily activities of child patients, sleep disorders, poor performance at school, slow growth and development, social barriers, anxiety, depression, etc. And their family members will have problems such as lack of sleep, fatigue, anxiety, depression, and impact on work and life due to taking care of the children and constantly seeking medical treatment. Both the children and their families are severely troubled by this, endangering physical and mental health, and even threatening life, causing a heavy economic and social burden.

[0003] Dermatology is a morphological science with strong intuitiveness. Its diagnosis and treatment monitoring largely depend on the morphology of various skin diseases. Therefore, the traditional process of diagnosing skin diseases is mainly that dermatologists judge diseases based on the patient's medical history, clinical manifestations combined with image information such as the morphology, size, and color of skin lesions. However, this method has certain subjectivity and errors. Therefore, the research in the field of skin disease image classification aims to use computer vision and image processing technologies to automatically classify and diagnose skin disease images, thereby improving the diagnosis efficiency and accuracy of skin diseases.

[0004] For example, the Chinese patent document with the publication number CN117333710A discloses a method for classifying skin disease images based on deep learning; the Chinese patent document with the publication number CN111444960A discloses a skin disease image classification system based on multi-modal data input.

[0005] Although some progress has been made in the research in the field of skin disease image classification, there are still many challenges. For example, factors such as the diversity and complexity of skin disease images, the similarity between different skin diseases, and the impact of image quality all bring difficulties to the classification and diagnosis of skin disease images. Summary of the Invention

[0006] The present invention provides a multi-classification system for children's skin diseases based on improved ResNet, which can realize the automatic classification of children's skin disease images and improve the recognition and classification accuracy of common children's skin diseases.

[0007] A multi-classification system for children's skin diseases based on improved ResNet includes a computer memory, a computer processor, and a computer program stored in the computer memory and executable on the computer processor. A trained classification model is stored in the computer memory;

[0008] The classification model described above uses an improved ResNet network model, which includes a feature extraction unit and a prediction unit;

[0009] The feature extraction unit includes a residual module structure, a medium-scale dilated convolution structure, a wide-scale dilated convolution structure, and a multi-scale fusion atrous convolution structure; After passing through a convolutional layer, a max pooling layer, a convolutional layer, and a batch normalization layer in sequence, features are extracted from the image and enter the above four structures. The four structures are used in parallel and output their respective feature maps;

[0010] In the prediction unit described above, the feature maps output by each structure are fused and then input into a batch normalization layer. The fused features are standardized through normalization techniques; Then, the normalized feature maps are downsampled through an average pooling layer, and the feature maps are compressed into a single vector, reducing the spatial dimension of the feature maps while retaining the most important feature information; Finally, the features are transformed into the classification probability distribution of skin diseases through a softmax layer.

[0011] When the computer processor executes the computer program, the following steps are implemented: Input the images of common skin diseases in children to be classified into the trained classification model to obtain the skin disease classification result.

[0012] Furthermore, the training process of the classification model is as follows:

[0013] Collect the images of the affected areas of children with atopic dermatitis, urticaria, hemangioma, diaper dermatitis, and the skin images of normal children, and preprocess the image data;

[0014] Randomly select 70% of the data as the training set, randomly select 10% of the data as the validation set, and the remaining 20% of the data as the test set; Among them, the training set is used for model training; The validation set is used for model selection and hyperparameter tuning; The test set is used for final evaluation of the model performance.

[0015] The preprocessing of the image data includes the following steps:

[0016] Crop the skin images to reduce irrelevant areas; Image standardization: Normalize the images through the mean and standard deviation to ensure that all images have the same brightness and contrast level; Image enhancement: Use histogram equalization and color space conversion techniques to enhance the image details; Size standardization: Resize all input images to a unified size.

[0017] The residual module structure described above includes three residual modules, where two residual modules are connected in series and used in parallel with another residual module;

[0018] Each residual module contains two parallel branches. One branch contains a 3×3 convolutional layer, a batch normalization layer, a 3×3 convolutional layer, and a batch normalization layer connected in sequence, and the other branch is a 1×1 convolutional layer. The features obtained from the two branches are added together and then passed through a batch normalization layer to obtain the output feature map.

[0019] The described medium-scale dilated convolutional structure contains a 3×3 convolutional layer, a batch normalization layer, a 3×3 dilated convolutional layer, and a batch normalization layer connected in sequence. Among them, the dilation rate of the 3×3 dilated convolutional layer is 2, and the medium-scale dilated convolutional structure is used to simulate the receptive field of a 7×7 convolutional kernel.

[0020] The described wide-scale dilated convolutional structure contains a 3×3 convolutional layer, a batch normalization layer, a 3×3 dilated convolutional layer, a batch normalization layer, a 3×3 dilated convolutional layer, and a batch normalization layer connected in sequence. Among them, the dilation rate of the 3×3 dilated convolutional layer is 2, and the wide-scale dilated convolutional structure is used to simulate the receptive field of a 15×15 convolutional kernel.

[0021] The described multi-scale fusion atrous convolutional structure contains two parallel branches. One branch contains three atrous convolutional layers connected in sequence, and each atrous convolutional layer uses a different dilation rate to expand the receptive field. After the input feature map undergoes three atrous convolutional operations in sequence, the intermediate feature maps corresponding to the first two convolutions and the feature map of the last convolution are concatenated. The other branch is a 1×1 convolutional layer. The outputs of the two branches are added together and then passed through a 1×1 convolutional layer to obtain the output feature map.

[0022] Compared with the prior art, the present invention has the following beneficial effects:

[0023] 1. The classification model in the present invention adopts an improved Resnet network model, and uses three different methods combined with residual modules, medium-scale dilated convolutions, wide-scale dilated convolutions, and multi-scale fusion atrous convolutions to simulate various convolutional kernel sizes for feature extraction. This means that the network can decide how to trade off and compensate between different methods of simulating convolutional kernels, and can enable the model to effectively learn feature information at different scales.

[0024] 2. The present invention improves the accuracy of the recognition and classification of skin diseases by learning features from a large number of skin disease images, provides auxiliary diagnosis suggestions for doctors, shortens the diagnosis time, helps primary medical institutions improve their diagnostic capabilities, and enables more children to receive timely and accurate medical services. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a flowchart of the implementation of a multi-classification system for children's skin diseases based on an improved ResNet according to an embodiment of the present invention;

[0026] Figure 2Schematic diagram of the overall structure of the classification model in the embodiments of the present invention;

[0027] Figure 3 Schematic diagram of the residual module structure in the embodiments of the present invention;

[0028] Figure 4 Schematic diagram of the medium-scale dilated convolution structure in the embodiments of the present invention;

[0029] Figure 5 Schematic diagram of the wide-scale dilated convolution structure in the embodiments of the present invention;

[0030] Figure 6 Schematic diagram of the multi-scale fusion atrous convolution structure in the embodiments of the present invention. Detailed implementation manners

[0031] The present invention will be further described in detail below with reference to the drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention and do not impose any limitations on it.

[0032] A classification system for common skin diseases in children includes a computer memory, a computer processor, and a computer program stored in the computer memory and executable on the computer processor. The trained classification model is stored in the computer memory.

[0033] As Figure 1 shown, the implementation process of the entire system is as follows:

[0034] 1. Data collection

[0035] Collect the images of the affected areas of children with atopic dermatitis, urticaria, hemangioma, diaper dermatitis, and the skin images of normal children, and preprocess the image data. All images should be labeled by professional dermatologists to ensure the accuracy and consistency of the data.

[0036] 2. Image preprocessing

[0037] The preprocessing of the pictures includes the following steps: Crop the skin pictures to reduce irrelevant areas, enabling the model to focus on key areas and improving the efficiency and performance of the model; Image normalization: Normalize the images through the mean and standard deviation to ensure that all images have the same brightness and contrast levels; Image enhancement: Use techniques such as histogram equalization and color space conversion to enhance the details of the images; Size normalization: Resize all input images to a unified size for easy model processing.

[0038] 3. Data grouping

[0039] Randomly select 70% of the data as the training set, 10% of the data as the validation set, and the remaining 20% of the data as the test set. Among them, the training set is used for model training; the validation set is used for model selection and hyperparameter tuning; the test set is used for final evaluation of the model performance.

[0040] 4. Model construction

[0041] In the present invention, the classification model adopts an improved ResNet network model, which includes a feature extraction unit and a prediction unit.

[0042] The feature extraction unit includes a residual module structure, a medium-scale dilated convolution structure, a wide-scale dilated convolution structure, and a multi-scale fusion dilated convolution structure; after the image sequentially passes through a convolutional layer, a max pooling layer, a convolutional layer, and a batch normalization layer, features are extracted from the image and enter the above four structures. The four structures are used in parallel and output their respective feature maps;

[0043] In the prediction unit, the feature maps output by each structure are fused and then input into a batch normalization layer, and the fused features are standardized through normalization technology; then, the standardized feature maps are downsampled through an average pooling layer, and the feature maps are compressed into a single vector, reducing the spatial dimension of the feature maps while retaining the most important feature information; finally, the features are transformed into the classification probability distribution of skin diseases through a softmax layer.

[0044] As Figure 2 shown, in this embodiment, the input size of the model is a skin disease picture of [3, 224, 224], indicating that both the height and width of the image are 224 pixels, and the number of channels is 3; the feature maps pass through a convolutional layer, a max pooling layer, and a convolutional layer respectively. The convolutional layer extracts features from the image and generates a new feature map. The two convolutional layers capture the patterns and features of the image at different scales. Through the max pooling layer, the size of the feature map is reduced while retaining important features. Max pooling can reduce the spatial dimension of the feature map, thereby reducing the computational cost.

[0045] The residual module is one of the basic building blocks for constructing a residual network, aiming to improve the training and performance of the network through residual connections. The design of the residual module enables the network to learn and represent complex features more effectively.

[0046] As Figure 3 shown, the residual module structure includes three residual modules, where two residual modules are connected in series and used in parallel with another residual module. Each residual module contains two parallel branches. One branch includes a 3×3 convolutional layer, a batch normalization layer, a 3×3 convolutional layer, and a batch normalization layer connected in sequence, and the other branch is a 1×1 convolutional layer; the features obtained from the two branches are added together and then passed through a batch normalization layer to obtain the output feature map.

[0047] In this embodiment, one and two residual modules are adopted. The specific structure of each residual module is as follows: a feature map with an input scale size of [c, w, h] is used, where c represents the number of channels, and w and h respectively represent the width and height of the feature map; a 1×1 convolutional layer is used to change the number of channels of the feature map to match the number of channels required by the subsequent convolutional layer; a 3×3 convolutional layer is responsible for feature extraction; a non-linear activation function is added to the two convolutional layers to introduce non-linear features; a batch normalization layer is used to standardize the input features, which helps to accelerate the training process and prevent gradient vanishing or explosion; an addition operation adds the newly learned features and the input features, enabling the training of deeper networks.

[0048] As Figure 4 shown, the medium-scale dilated convolutional structure includes a 3×3 convolutional layer, a batch normalization layer, a 3×3 dilated convolution, and a batch normalization layer connected in sequence; among them, the dilation rate of the 3×3 dilated convolution is 2. As Figure 5 shown, the wide-scale dilated convolutional structure includes a 3×3 convolutional layer, a batch normalization layer, a 3×3 dilated convolutional layer, a batch normalization layer, a 3×3 dilated convolutional layer, and a batch normalization layer connected in sequence; among them, the dilation rate of the 3×3 dilated convolutional layer is 2.

[0049] In this embodiment, the medium-scale dilated convolution and the wide-scale dilated convolution adopt dilated convolution to reduce the parameters required by a larger convolutional kernel, enabling the network to understand higher-level features. The specific structure of this module is as follows: a feature map with an input scale size of [c, w, h] is used, where c represents the number of channels, and w and h respectively represent the width and height of the feature map; a 3×3 convolutional layer is used to learn features, and the dilated convolutional layer can expand the receptive field of the model, which helps to capture long-range dependencies. The medium-scale dilated convolution simulates the receptive field of a 7×7 convolutional kernel, while the wide-scale dilated convolution simulates the receptive field of a 15×15 convolutional kernel; a batch normalization layer is used to standardize the input features, which helps to accelerate the training process and prevent gradient vanishing or explosion.

[0050] As Figure 6 shown, the multi-scale fusion atrous convolutional structure includes two parallel branches. One branch includes three atrous convolutions connected in sequence. After the input feature map undergoes three atrous convolution operations in sequence, the intermediate feature maps corresponding to the first two convolutions and the feature map of the last convolution are feature concatenated; the other branch is a 1×1 convolution; the outputs of the two branches are added by features and then passed through a 1×1 convolution to obtain the output feature map.

[0051] In this embodiment, each dilated convolutional layer uses a different dilation rate to expand the receptive field. The input feature map has a size of [c, w, h], where c represents the number of channels, and w and h represent the width and height of the feature map respectively. After three convolutional operations, the intermediate feature maps are concatenated to obtain a feature map with a size of [3c, w, h], and 1×1 convolution is used for channel compression. Finally, the compressed feature map and the input feature map are combined through a residual connection. The output feature map aggregates the receptive fields of these three different scales at the same time, enabling the model to effectively learn the feature information at different scales.

[0052] As Figure 2 shown, in the network architecture of the present invention, all the above-mentioned module structures are used in parallel, enabling the network to use the behavior it deems optimal at each step. By simulating various convolutional kernel sizes in three different ways, the different modules output feature maps with the same size of [c, w, h]. The feature maps of different modules are concatenated in the channel c dimension through the splicing method, preserving the integrity of the features of the feature maps output by different modules and making full use of the feature information obtained by different modules. The network can decide how to trade off and compensate between different ways of simulating convolutional kernels. Having multiple convolutional kernel sizes means that it can find the general area of the target and also correctly find the edges.

[0053] After fusing the feature maps output by each module, the input is sent to the batch normalization layer, and the fused features are standardized through the normalization technique to keep the data distribution stable. Then, the feature map is downsampled through the average pooling layer to compress the feature map into a single vector, reducing the spatial dimension of the feature map while retaining the most important feature information. Finally, the features are transformed into an output in the form of a probability distribution through the softmax layer, converting the scores of each category into probability values. The output of the softmax layer can be regarded as the probability of each category.

[0054] 5. Model Training and Classification

[0055] The training set is fed into the classification model. Through model training, the model parameters are updated by backpropagation. At the same time, the model is evaluated on the validation set, and the model parameters are optimized according to the evaluation results. After several rounds of training, the trained model is tested on the test set to evaluate the final performance of the classification model.

[0056] 6. Model Evaluation

[0057] After the model training is completed, the classification effect is evaluated on an independent test set, and indicators such as accuracy, recall rate, F1 value, and confusion matrix are used to evaluate the performance of the model on the test set.

[0058] The multi-classification system for common childhood skin diseases based on the improved ResNet provides an efficient and accurate method for the auxiliary diagnosis of common childhood skin diseases. Through carefully designed data collection, model training, and optimization strategies, a powerful auxiliary diagnosis tool can be constructed to help doctors better identify and manage childhood skin diseases.

[0059] The above-described embodiments have detailed the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, supplements, and equivalent replacements made within the scope of the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multi-classification system for children's skin diseases based on improved ResNet, characterized in that, It includes a computer memory, a computer processor, and a computer program stored in the computer memory and executable on the computer processor. A trained classification model is stored in the computer memory; The classification model uses an improved ResNet network model, which includes a feature extraction unit and a prediction unit; The feature extraction unit includes a residual module structure, a medium-scale dilated convolution structure, a wide-scale dilated convolution structure, and a multi-scale fusion atrous convolution structure; After the image sequentially passes through a convolutional layer, a max pooling layer, a convolutional layer, and a batch normalization layer, features are extracted from the image and enter the above four structures. The four structures are used in parallel and output their respective feature maps; The residual module structure includes three residual modules, where two residual modules are connected in series and used in parallel with another residual module; Each residual module includes two parallel branches. One branch includes a 3×3 convolutional layer, a batch normalization layer, a 3×3 convolutional layer, and a batch normalization layer connected in sequence, and the other branch is a 1×1 convolutional layer; The features obtained from the two branches are added and then passed through a batch normalization layer to obtain the output feature map; The medium-scale dilated convolution structure includes a 3×3 convolutional layer, a batch normalization layer, a 3×3 dilated convolution, and a batch normalization layer connected in sequence; Among them, the dilation rate of the 3×3 dilated convolution is 2, and the medium-scale dilated convolution structure is used to simulate the receptive field of a 7×7 convolutional kernel; The wide-scale dilated convolution structure includes a 3×3 convolutional layer, a batch normalization layer, a 3×3 dilated convolutional layer, a batch normalization layer, a 3×3 dilated convolutional layer, and a batch normalization layer connected in sequence; Among them, the dilation rate of the 3×3 dilated convolutional layer is 2, and the wide-scale dilated convolution structure is used to simulate the receptive field of a 15×15 convolutional kernel; The multi-scale fusion atrous convolution structure includes two parallel branches. One branch includes three atrous convolutions connected in sequence. Each atrous convolution uses a different dilation rate to expand the receptive field. After the input feature map is sequentially subjected to three atrous convolution operations, the intermediate feature maps corresponding to the first two convolutions and the feature map of the last convolution are feature concatenated; The other branch is a 1×1 convolutional layer; The outputs of the two branches are added and then passed through a 1×1 convolutional layer to obtain the output feature map; In the prediction unit, the feature maps output by each structure are fused and then input into a batch normalization layer, and the fused features are standardized through normalization technology; Then, the standardized feature map is downsampled through an average pooling layer to compress the feature map into a single vector, reducing the spatial dimension of the feature map while retaining the most important feature information; Finally, the features are converted into a classification probability distribution of skin diseases through a softmax layer; When the computer processor executes the computer program, the following steps are implemented: Input the image of common skin diseases in children to be classified into the trained classification model to obtain the skin disease classification result.

2. The multi-classification system for children's skin diseases based on the improved ResNet according to claim 1, characterized in that, The training process of the classification model is as follows: Collect the images of the affected areas of children with atopic dermatitis, urticaria, hemangioma, diaper dermatitis, and the skin images of normal children, and preprocess the image data; Randomly select 70% of the data as the training set, randomly select 10% of the data as the validation set, and the remaining 20% of the data as the test set; among them, the training set is used for model training; the validation set is used for model selection and hyperparameter tuning; the test set is used for final evaluation of model performance.

3. The multi-classification system for children's skin diseases based on the improved ResNet according to claim 2, characterized in that, Preprocess the image data, including the following steps: Crop the skin images to reduce irrelevant regions; Image normalization: Normalize the images by mean and standard deviation to ensure that all images have the same brightness and contrast levels; Image enhancement: Use histogram equalization and color space conversion techniques to enhance image details; Size normalization: Resize all input images to a unified size.

Citation Information

Patent Citations

  • Skin disease image classification system based on multi-modal data input

    CN111444960A

  • Deep learning-based skin disease image classification method

    CN117333710A

  • Skin disease identification model and auxiliary diagnosis platform based on VGG-16 fusion residual network

    CN114693976A

  • Breast cancer pathological image classification device and method based on deep learning

    CN116524226A

  • Ground object target extraction method and device based on deep learning

    CN116883679A