Skin image-based psoriasis traditional Chinese medicine syndrome classification method, equipment and medium
Through deep learning technology, using skin recognition models and psoriasis syndrome classification models to automatically identify and classify psoriatic skin images, solving the problems of strong subjectivity and low efficiency in traditional diagnostic methods, and achieving high accuracy of traditional Chinese medicine syndrome classification.
Patent Information
- Application Number
- CN202510161259.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-03
AI Technical Summary
Traditional psoriasis diagnosis relies on doctors' observation and manual annotation, which has problems such as strong subjectivity, long-term and difficult to meet large-scale clinical needs. At the same time, the lesions of psoriasis are complex in shape and diverse in distribution, and traditional diagnostic methods have limited performance in small-target lesions detection and precise classification.
By obtaining skin images, inputting skin recognition model to obtain target lesions information, determining target lesions images, and entering them into the psoriasis syndrome classification model to obtain classification results for traditional Chinese medicine syndromes. This method includes skin recognition model and psoriasis syndrome classification model, and uses deep learning technologies such as convolutional neural networks to extract and classify features.
It realizes automatic identification and segmentation of skin images, improves the accuracy of the classification of Chinese medicine syndromes in psoriasis, reduces the errors that rely on physician experience, reduces the risk of misdiagnosis and missed diagnosis, and promotes the popularization of standardized Chinese medicine syndromes in psoriasis.
Smart Images

Figure CN120088554A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method, device, and medium for classifying traditional Chinese medicine syndromes of psoriasis based on skin images. Background Art
[0002] With the development of deep learning technology, significant progress has been made in automated diagnosis technology in medical image analysis. In the field of dermatology, the classification of traditional Chinese medicine syndromes of psoriasis is an important basis for diagnosis and treatment. Traditional methods mainly rely on doctors' observation and manual annotation of the skin lesion area, but this method is limited by diagnostic experience, time-consuming, and highly subjective, making it difficult to meet the large-scale clinical needs.
[0003] At the same time, the morphology of the psoriasis skin lesion area is complex and the distribution is diverse, and the traditional diagnosis method is relatively limited in the detection and precise classification of small target skin lesions. Due to clinical privacy, different from general big data models that use thousands of pictures for pre-training, there are few clinical RGB images of psoriasis target skin lesions. In order to better utilize clinical images and improve the accuracy of the model, the background such as the environment and clothing should be removed. However, currently, the well-trained large models that can be applied to identify the human body target the whole person, and have low specificity in identifying local human bodies such as the trunk and limbs, and cannot maximize the use of clinical images to extract features. Summary of the Invention
[0004] According to one aspect of the present invention, there is provided a method, device, and medium for classifying traditional Chinese medicine syndromes of psoriasis based on skin images, which can effectively exclude the interference of the environment and clothing, and has a high accuracy in classifying traditional Chinese medicine syndromes and a short determination time.
[0005] To solve the above technical problems, the first aspect of the present invention discloses a method for classifying traditional Chinese medicine syndromes of psoriasis based on skin images, including:
[0006] Obtain a skin image, input the skin image into a skin recognition model to obtain target skin lesion information;
[0007] Determine a target skin lesion image based on the target skin lesion information;
[0008] Input the target skin lesion image into a psoriasis syndrome classification model to obtain the classification result of the traditional Chinese medicine syndrome corresponding to the skin image.
[0009] In some embodiments, the psoriasis syndrome classification model includes an input end, a convolutional layer, a global average pooling layer, and a fully connected layer; the input end receives the target skin lesion image; the convolutional layer includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer, and multi-channel feature processing is performed through the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer to obtain a multi-channel two-dimensional feature map;
[0010] The global average pooling layer performs global average pooling operation on the two-dimensional feature map of each channel to generate a feature vector; the fully connected layer obtains the scores corresponding to each traditional Chinese medicine syndrome classification of the target skin lesion image through the feature vector, and outputs the classification result of the traditional Chinese medicine syndrome according to the scores.
[0011] In some embodiments, the psoriasis syndrome classification model is trained by a second training image set and a cross-entropy loss function, and the weights of the psoriasis syndrome classification model are updated by a backpropagation algorithm.
[0012] In some embodiments, the first convolutional layer is a 7×7 convolutional layer, and the stride of the first convolutional layer is 2; the first convolutional layer is used for performing preliminary feature extraction on the target skin lesion image to obtain input features;
[0013] The second convolutional layer includes two residual blocks, each residual block contains two 3×3 convolutional layers, and the number of feature maps is 64; through the skip connection mechanism of the residual block, the input features are added to the output of the convolutional layer; the second convolutional layer is used for extracting the edge and texture information of the target skin lesion image.
[0014] The third convolutional layer includes two residual blocks, each residual block contains two 3×3 convolutional layers, and the number of feature map channels is 128; the third convolutional layer is used for performing deep processing on the feature map.
[0015] The fourth convolutional layer includes two residual blocks, and the number of feature map channels is 264; the fourth convolutional layer is used for extracting the deep semantic features of the target skin lesion image to obtain a multi-channel two-dimensional feature map.
[0016] In some embodiments, the skin recognition model includes an input end, a feature extraction network, a feature fusion network, and an output end. The skin image enters through the input end, feature maps are extracted through the feature extraction network, feature fusion is performed through the feature fusion network, and target skin lesion information is output by the output end; the target skin lesion information includes skin tissue information and bounding box information of the prediction target. The skin tissue information indicates whether there is skin tissue in each grid unit of the feature map, and the bounding box information of the prediction target includes the center position of the bounding box and the parameters of the bounding box.
[0017] In some embodiments, the skin recognition model extracts feature maps through the feature extraction network, including:
[0018] Performing slicing operation on the skin image to divide it into a plurality of sub-images;
[0019] Concatenate the sub - graphs in the channel dimension to form a multi - channel feature map; extract local features of the skin image through a convolution operation to obtain an original feature map;
[0020] Perform three maximum pooling operations on the original feature map to obtain a pooled feature map;
[0021] Concatenate the original feature map and the pooled feature map to obtain global features.
[0022] In some embodiments, the skin recognition model performs feature fusion through the feature fusion network, including:
[0023] The feature fusion network includes a feature pyramid network and a path aggregation network. The feature pyramid network gradually improves the resolution of deep - layer features through up - sampling operations and concatenates the deep - layer features and shallow - layer features;
[0024] The path aggregation network gradually reduces the resolution of shallow - layer features through down - sampling operations and concatenates the shallow - layer features and corresponding deep - layer features.
[0025] In some embodiments, the skin recognition model is trained through a labeled first training image set and a multi - task loss function, and the multi - task loss function is used for model optimization during each training cycle; the multi - task loss function includes a bounding box regression loss function, an object confidence loss function, and a classification loss function. The bounding box regression loss function is the CIoU loss function, and the object confidence loss function and the classification loss function adopt binary cross - entropy loss functions.
[0026] According to the second aspect of the present invention, a computer device is disclosed, including: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of a method for classifying traditional Chinese medicine syndromes of psoriasis based on skin images as described in any one of the above.
[0027] According to the third aspect of the present invention, a computer storage medium is disclosed, on which a computer program is stored. When the computer program is executed by a processor, the steps of a method for classifying traditional Chinese medicine syndromes of psoriasis based on skin images as described in any one of the above are implemented.
[0028] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0029] The present invention provides a traditional Chinese medicine syndrome classification method, device, and medium for psoriasis based on skin images, which automatically recognize and segment skin images, and determine the classification results of skin images through a psoriasis syndrome classification model. The present invention does not rely on the experience of traditional Chinese medicine masters, reduces the error of relying on the subjective judgment of physicians, improves the accuracy of psoriasis classification, and reduces the risk of misdiagnosis and missed diagnosis. In addition, through the skin recognition model and the psoriasis syndrome classification model, the classification process is standardized and normalized, so that the syndrome determination results of different physicians and different medical institutions are consistent, which helps to promote the popularization of standardized syndromes of traditional Chinese medicine for psoriasis. Description of the Drawings
[0030] Figure 1 It is a schematic flow chart of a traditional Chinese medicine syndrome classification method for psoriasis based on skin images provided by the present invention;
[0031] Figure 2 It is an example diagram of obtaining the coordinate information of the original image using yolov5 in the present invention.
[0032] Figure 3 It is an example diagram of constructing a second training set using segment anything and opencv in the present invention.
[0033] Figure 4 It is an architecture diagram of the convolutional neural network used in the psoriasis syndrome classification model of the present invention.
[0034] Figure 5 It is an example diagram of the image input into the psoriasis syndrome classification model and the heat map generated using the Grad-CAM technique in the present invention.
[0035] Figure 6 It is the receiver operating characteristic curve and the area under the curve of the internal validation set of the psoriasis syndrome classification model of the present invention.
[0036] Figure 7 It is the receiver operating characteristic curve and the area under the curve of the external validation set of the psoriasis syndrome classification model of the present invention. Detailed Embodiments
[0037] For better understanding and implementation, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] As used in the embodiments of the present invention, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or modules need not be limited to those steps or modules clearly listed, but may include other steps or modules not clearly listed or inherent to these processes, methods, products, or devices.
[0039] An embodiment of the present invention discloses a method for classifying traditional Chinese medicine syndromes of psoriasis based on skin images, which can effectively exclude the interference of the environment and clothing, and has a high accuracy rate for classifying traditional Chinese medicine syndromes and a short determination time.
[0040] As Figure 1 shown, the present invention provides a method for classifying psoriasis syndromes based on skin images, including:
[0041] Step S1, obtain a skin image, input the skin image into a skin recognition model, and obtain target skin lesion information. The skin image is generally a digital RGB two-dimensional image, and common formats include JPEG, PNG, BMP, etc. The image resolution and size can be adjusted according to application requirements. The skin image includes the skin of the trunk, limbs, or other local human body parts. The obtained skin image, as Figure 2 shown, directly obtains skin tissue coordinates and target confidence information through the skin recognition model.
[0042] The skin recognition model needs to be trained before application, and is trained through a labeled first training image set and a multi-task loss function. The first training image set includes multiple training images manually labeled for the human skin of the RGB image. The yolov5 pre-trained model is trained based on the first training image set, and when the training meets the requirements, the skin recognition model is obtained.
[0043] The skin recognition model includes an input end, a feature extraction network, a feature fusion network, and an output end. The skin image enters through the input end, feature maps are extracted through the feature extraction network, feature fusion is performed through the feature fusion network, and target skin lesion information is output by the output end; the target skin lesion information includes the skin tissue information and the bounding box information of the prediction target. The skin tissue information indicates whether there is skin tissue in each grid unit of the feature map, and the bounding box information of the prediction target includes the center position of the bounding box and the parameters of the bounding box.
[0044] The input end is responsible for receiving the input skin images. Generally, the standard input size of skin images is 608x608 pixels. Before training or application, the skin images are preprocessed, including scaling and normalization, to meet the input requirements of the model. YOLOv5 uses the Mosaic data augmentation technique to improve the training speed and generalization ability of the model by stitching multiple images, ensuring its effectiveness in different environments.
[0045] The Backbone part, that is, the feature extraction network, is the core part of this skin recognition model, which is used to extract the basic features of skin tissues from the input skin images. In this application, YOLOv5m uses CSPDarknet53 as the baseline network and combines the Focus structure to efficiently extract local and global features. The skin recognition model extracts feature maps through the feature extraction network, including:
[0046] Slice the skin image to divide it into several sub-images;
[0047] Stitch the sub-images in the channel dimension to form a multi-channel feature map; extract the local features of the skin image through convolution operations to obtain the original feature map;
[0048] Perform three max-pooling operations on the original feature map to obtain the pooled feature map;
[0049] Stitch the original feature map and the pooled feature map to obtain the global feature.
[0050] Specifically, during the feature extraction process of the feature extraction network of YOLOv5m, the input skin image first undergoes the slicing and stitching operations of the Focus module, divides the image into 4 sub-images, and stitches them in the channel dimension to generate a multi-channel feature map with a higher number of channels. At the same time, the spatial resolution is reduced by half, for example, from 640×640 to 320×320, providing an efficient input for subsequent feature extraction.
[0051] The multi-channel feature map enters the Conv module, and local features are extracted through convolution operations. At the same time, batch normalization and activation functions are combined to make the feature expression more semantic. After the initial convolution processing, the multi-channel feature map enters the first C3 module. The C3 module further extracts deep features through residual connections and multiple Bottleneck operations, and at the same time fuses shallow and deep information, so as to enhance semantic expression while maintaining detailed features. After 4 Conv + C3, the number of channels and semantic information of the feature map gradually increase, and the spatial resolution is gradually compressed.
[0052] The original feature map enters the SPPF module (Spatial Pyramid Pooling - Fast). The SPPF module first performs a 5×5 max pooling on the input feature map to capture features in a larger range; the second pooling is a 5×5 max pooling that pools the result of the first pooling again to extract information in an even larger range; the third pooling is also a 5×5 max pooling that pools the result of the second pooling to obtain a pooled feature map, thereby further capturing more global information. Through the stacked pooling operations, feature representations with different receptive field ranges are generated.
[0053] The original feature map and the pooled feature map after three times of pooling are concatenated in the channel dimension and fused through a 1×1 convolutional layer to generate the final output feature. This structure significantly enhances the feature map's perception ability for multi-scale targets while keeping the spatial resolution unchanged (for example, both the input and output feature maps are 40×40). During the feature extraction process of YOLOv5m, through the progressive processing of multiple Conv and C3 combination modules, the features gradually evolve from shallow detailed information to deep semantic information, and finally, global feature integration is completed in the SPPF module, providing high-quality multi-scale feature representations for subsequent object detection.
[0054] The Neck part is located between the Backbone and the Head, that is, the feature fusion network, which is responsible for enhancing the diversity and robustness of features. In YOLOv5m, the Feature Pyramid Network (FPN) and the Path Aggregation Network (PAN) modules are adopted to integrate features of different scales. The feature fusion network includes the Feature Pyramid Network and the Path Aggregation Network. The Feature Pyramid Network gradually improves the resolution of deep features through upsampling operations and concatenates the deep features with shallow features; the Path Aggregation Network gradually reduces the resolution of shallow features through downsampling operations and concatenates the shallow features with the corresponding deep features.
[0055] Specifically, feature fusion is performed from deep to shallow through the FPN (Feature Pyramid Network). The FPN uses upsampling operations to raise the resolution of deep features with low resolution but rich semantic information (such as P5) to the resolution of shallow features (such as P4, P3), and concatenates the deep features with the shallow features, thereby endowing the shallow features with more semantic information and making the model better at detecting small targets.
[0056] Next, information is passed backward from the shallow layer to the deep layer through PAN (Path Aggregation Network). PAN uses downsampling operations to gradually transfer shallow features with high resolution but lacking semantic information into deep features, and splices the detailed information of shallow features with deep features, thereby enhancing the model's perception ability of large objects. Through this two-way information flow and fusion, the PAN-FPN structure ensures that the model can comprehensively capture the feature information of small and large objects, thereby improving the detection effect of different-scale features in skin tissue.
[0057] The Head part is the output end, responsible for generating object detection results. The multi-scale feature maps processed by the Neck, that is, the feature fusion network, are respectively input into the Head for the skin tissue detection task. The output end includes a classification branch and a regression branch. The features processed by the Neck are input into the Head. The classification branch determines whether there is skin tissue in each grid cell of the multi-scale feature map, while the regression branch outputs the bounding box of the predicted object, including the center position (x, y) of the bounding box and the width and height (w, h) of the bounding box. In the Head, the classification branch and the regression branch jointly generate a prediction result of the complete target skin lesion information. The target skin lesion information includes the skin tissue information and the bounding box information of the predicted object. The skin tissue information indicates whether there is skin tissue in each grid cell of the feature map. The bounding box information of the predicted object includes the center position of the bounding box and the parameters of the bounding box, or the upper-left corner coordinates and the lower-right corner coordinates of the bounding box. Specifically, the dimension formula of the prediction result of the target skin lesion information is:
[0058] [B, A, H, W, (5 + C)]
[0059] Among them, B is the batch size, A is the number of anchor boxes for each grid cell; H and W are the height and width of the feature map; 5 + C contains 5 regression parameters and the number of classes, 4 regression parameters (x, y, w, h), 1 object confidence, and the classification probabilities of C classes.
[0060] In each training cycle, the model is optimized through a multi-task loss function. The multi-task loss function includes a bounding box regression loss function, an object confidence loss function, and a classification loss function. The bounding box regression loss function is the CIoU loss function, and the object confidence loss function and the classification loss function use the binary cross-entropy loss function.
[0061] Among them, the bounding box regression loss function is used to optimize the position and size of the predicted skin bounding box in the skin recognition model, so that the predicted skin bounding box is closer to the real skin position. Among them, the bounding box regression loss function uses CIoU (Complete Intersection over Union) loss, and the formula is:
[0062]
[0063] Among them, gt is the ground-truth, and b and b gt represent the center points of the true bounding box and the predicted bounding box, ρ represents the Euclidean distance between the true bounding box and the predicted bounding box, c represents the distance of the diagonal of the closure region between the true bounding box and the predicted bounding box, v is used to measure the consistency of the relative ratio between the true bounding box and the predicted bounding box, and α is the weight coefficient. w and h are the width and height of the predicted bounding box, w gt 、h gt are the width and height of the true bounding box. IoU is the ratio of the intersection to the union of the true bounding box and the predicted bounding box,
[0064] and the object confidence loss function measures the gap between the confidence that each predicted box contains the object and the true value. In this application, binary cross-entropy loss (BCEloss) is used to calculate the loss of object confidence. The formula of BCELoss is:
[0065]
[0066] Among them, y i is the true label, taking values of 0 or 1; p(yi) is the probability that the i-th sample predicted by the model is a positive class; N is the number of groups of objects predicted by the model.
[0067] Furthermore, in this application, the classification loss function measures the gap between the predicted class distribution and the true class. Binary cross-entropy loss with label smoothing (BCE loss) is used to calculate the classification loss.
[0068] The finally used multi-task loss function weights and sums up the losses of three parts: bounding boxes, object confidence, and classification results through hyperparameters. During the training process, the training dataset is input into the yolov5 pre-trained model, and iterative training is performed through the backpropagation algorithm. Among them, the Backbone and the feature extraction network extract the features of the skin tissue, the Neck, that is, the feature fusion network integrates information of different scales, and the Head, that is, the output end generates the detection results, that is, the target skin lesion information. The number of training rounds is 100. After each training cycle, the network weights are optimized based on the total loss, and the comprehensive performance index (fitness) of the model on the validation set is calculated. The model with the best performance is saved and applied to the actual skin detection scenario. It is set that the comprehensive performance index Fitness consists of mAP@0.5 (Mean Average Precision at IoU = 0.5) and mAP@0.5:0.95 (Mean Average Precision at IoU = 0.5 to 0.95). mAP@0.5 is the average precision of all classes when the IoU (Intersection over Union) threshold is 0.5, which is obtained by calculating the average precision of different classes and taking the mean (precision = correct prediction / total samples). mAP@0.5:0.95 refers to the mean of all average precisions when IoU ranges from 0.5 to 0.95 (step size is 0.05), and it takes the average value of mAP at multiple IoU thresholds. It is set that the weight of mAP@0.5 is 0.1 and the weight of mAP@0.5:0.95 is 0.9.
[0069] Step S2: Determine the target skin lesion image based on the target skin lesion information;
[0070] The obtained target skin lesion information is used as the input of the prompt encoder and is input as a prompt into the Prompt Encoder of the SegmentAnything (SAM) model to represent the initial position of the target skin lesion area. The bounding box or coordinate information of the target skin lesion information is used as a guide. SAM first extracts features from the input image to generate multi-scale feature maps, fuses the prompt information, that is, the target skin lesion information, with the feature maps, and generates a preliminary segmentation mask. Based on the adaptive segmentation ability of deep learning, SAM gradually optimizes the mask to accurately cover the skin lesion area and remove the background. SAM outputs a segmented binary mask, indicating the pixel positions of the target skin lesion area. Among them, the pixel points with a mask value of 1 belong to the target skin lesion area, and the pixel points with a mask value of 0 belong to the background. Using the mask, the skin lesion area in the original image is extracted to generate an accurate target skin lesion image. Finally, as Figure 3As shown, the black background in the image is removed and cropped through OpenCV image processing technology, including grayscale conversion, threshold segmentation, contour detection and bounding rectangle calculation.
[0071] Step S3: input the target skin lesion image into a psoriasis syndrome classification model to obtain a classification result of the TCM syndrome corresponding to the skin image.
[0072] The psoriasis syndrome classification model includes an input end, a feature extraction network, a global average pooling layer, and a fully connected layer; the input end receives the target skin lesion image; the feature extraction network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer, and multi-channel feature extraction is performed through the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer to obtain a multi-channel two-dimensional feature map;
[0073] The global average pooling layer performs a global average pooling operation on the two-dimensional feature map of each channel to generate a feature vector; the fully connected layer obtains the score of each TCM syndrome classification corresponding to the target skin lesion image through the feature vector, and outputs the TCM syndrome classification result according to the score.
[0074] The neural network architecture of the psoriasis syndrome classification model is as follows Figure 4 As shown, training is required before use, and the training process can be as follows: based on step S2, a target skin lesion image without interference background is obtained, and each target skin lesion is labeled with a TCM syndrome type based on manual labeling, for example, the first target skin lesion image is labeled as blood heat syndrome, the second target skin lesion image is labeled as blood stasis syndrome, and the Nth target skin lesion image is labeled as blood dryness syndrome, to generate a second training image set.
[0075] In this application, 1786 typical skin lesions (i.e., the syndrome differentiation manifestations of skin lesions are consistent with the main symptoms of blood heat syndrome, blood stasis syndrome, and blood dryness syndrome in the "Guidelines for the Diagnosis and Treatment of Psoriasis in China (2023 Edition)") constitute the second training image set. Use resnet18 for transfer learning, learn the characteristics of characteristic skin lesions, and save the trained mature model for application in future clinical Chinese medicine syndrome judgment. The pre-trained model of resnet18 is trained based on the second training image set. Stop training when the training termination conditions are met to obtain a psoriasis syndrome classification model. The training termination conditions include reaching the highest point in accuracy or reaching a threshold in the number of training times, etc., to determine the model parameters with the best verification performance during the training process.
[0076] During annotation, the TCM syndrome differentiation of psoriasis in the "Chinese Guidelines for Psoriasis Diagnosis and Treatment (2023 Edition)" is referred to in the manual annotation stage. ① Blood-heat syndrome: Main symptoms: The skin lesions are bright red, and new rashes are constantly increasing or rapidly expanding. Secondary symptoms: Irritability, yellow urine, red or crimson tongue, string-taut, slippery or rapid pulse. ② Blood stasis syndrome: Main symptoms: The skin lesions are dark red; the skin lesions are thickened, infiltrated, and do not heal for a long time. Secondary symptoms: Dry and rough skin, sallow complexion or purplish lips and nails; dark menstrual blood in women, or accompanied by blood clots; purple or dark tongue with stasis points or ecchymoses; string-taut, fine or slow pulse. ③ Blood dryness syndrome: Main symptoms: The skin lesions are light red, and the scales are dry. Secondary symptoms: Dry mouth and throat; pale tongue, little or thin white tongue coating; fine or thready and rapid pulse.
[0077] During the training process, the second training dataset is input into the convolutional layer of ResNet-18, and image features are extracted through the residual blocks in its multiple stages. These pre-trained weights can help the model effectively extract low-level to high-level features. Especially when the number of training images in the second training set is limited, using the pre-trained model can accelerate training and improve the model performance. Then, the extracted features are transformed into a 256-dimensional vector of a fixed size through the global average pooling layer, and the classification results are output through the fully connected layer. During the training process, the cross-entropy loss function is used to calculate the difference between the model prediction and the true label, and the weights of the model are updated through the backpropagation algorithm. The formula for the cross-entropy loss function is:
[0078]
[0079] where N is the number of samples, C is the number of classes, and y n,c is the true label of the nth sample in class c. For multi-classification problems, one-hot encoding is usually used. y n,c = 1 indicates that sample n belongs to class c, otherwise y n,c = 0, is the probability that the psoriasis syndrome classification model predicts the nth sample as class c.
[0080] During the training stage, a total of 50 iterations are performed. After each training cycle, the model performance is evaluated on the validation set, and the parameters are adjusted and optimized according to the validation set accuracy. At the same time, the model with the best accuracy in the training set is saved. The accuracy is the ratio of the number of correctly predicted samples to the total number of samples. This model can accurately classify the syndrome types of skin lesion images and provide effective support for the TCM syndrome classification and diagnosis of psoriasis.
[0081] The psoriasis syndrome classification model obtained through training includes an input end, a convolutional layer, a global average pooling layer, and a fully connected layer; the input end receives the target skin lesion image; the convolutional layer includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer, and multi-channel feature processing is performed through the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer to obtain a multi-channel two-dimensional feature map;
[0082] The global average pooling layer performs global average pooling operations on the two-dimensional feature maps of each channel to generate feature vectors; the fully connected layer obtains the scores corresponding to each traditional Chinese medicine syndrome classification of the target skin lesion image through the feature vectors, and outputs the classification results of traditional Chinese medicine syndromes according to the scores.
[0083] Specifically, the input end receives the target skin lesion image, and the standard input size is 341,512 pixels. Before input, the target skin lesion image will undergo preprocessing, including scaling and normalization, to ensure that it meets the input requirements of the model, helps the model converge faster and improves the training or application effect.
[0084] After the target skin lesion image is input, it undergoes feature processing through the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer of the convolutional layer. The convolutional layer is the core of ResNet-18 and consists of multiple residual blocks. Each residual block contains two 3x3 convolutional layers and a short connection, allowing the input to be directly added to the output after convolutional processing, thereby realizing residual learning. ResNet-18 stacks the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer to gradually extract the low-level features of the input image into high-level features, enhancing the learning ability of the network in the deep structure.
[0085] This part extracts features layer by layer through multiple convolutional layers, ReLU activation functions, and pooling layers, providing rich feature representations for subsequent processing. The first convolutional layer includes a 7x7 convolutional layer with a stride of 2, followed by a max pooling layer for preliminary feature extraction and reducing the dimension of the feature map. The 7x7 convolutional operation can capture large-scale features in the image, while max pooling further reduces the spatial dimension, laying a foundation for subsequent deep feature extraction.
[0086] The second convolutional layer contains 2 residual blocks, and each residual block is composed of 2 3x3 convolutional layers. At this stage, the number of channels of the feature map increases to 64. The skip connection of each residual block allows the input features to be directly added to the output of the convolutional layer, which can effectively avoid the problem of gradient disappearance, improve the training efficiency and performance of the network. The features extracted at this stage are mainly edge and texture information, providing support for subsequent complex feature learning.
[0087] The third convolutional layer also contains 2 residual blocks, and the number of channels of the feature map increases to 128. Similar to the second convolutional layer, the skip connection ensures the effective transmission of information. During the learning process of the third convolutional layer, the convolutional layer can learn more advanced features, such as more complex shapes and local patterns, which are very important for distinguishing different types of skin lesion images. The feature maps at this stage provide richer information to help the model make more accurate judgments during classification.
[0088] The fourth convolutional layer consists of 2 residual blocks, and the number of channels is further increased to 256 to extract more abstract and complex features, capable of capturing higher-level semantic information in the target skin lesion images. Through the gradually increasing number of channels and the maintained skip connections, ResNet-18 can construct a deep feature representation, thereby improving the accuracy of the classification task.
[0089] After the feature processing by the convolutional layer, the two-dimensional feature maps of each channel are globally average pooled by the global average pooling layer to generate feature vectors; the two-dimensional feature maps of each channel are compressed into a scalar by calculating the average value of each channel, and the resulting 256-dimensional feature vectors will be used as inputs and enter the fully connected layer. Finally, the feature vectors are mapped to 3-dimensional vectors by the fully connected layer, representing the scores of 3 categories, namely the blood-heat syndrome, blood stasis syndrome, and blood dryness syndrome. The Softmax function is used to convert the scores of each category into a probability, where the probability represents the likelihood that the target skin lesion image belongs to a certain category, and the largest probability value corresponds to the most likely category.
[0090] The model with the highest validation set accuracy is saved as the trained psoriasis syndrome classification model. The parameters of this model are as follows: the total number of weights in the model is 11,173,248, among which the number of weights in the convolutional layer is 10,994,880, and the number of weights in the fully connected layer is 1,536. The average value of the convolutional weights is -0.001304021380008936; the standard deviation of the convolutional weights is 0.01996901072561741; the average value of the fully connected weights is 0.0015171129489317536; the standard deviation of the fully connected weights is 0.02553003090303893; the average value of all weights is -0.001109592616558075; the standard deviation of all weights is 0.0243691597138749; the minimum weight is -0.8451254367828369; the maximum weight is 2.337225914001465.
[0091] In the model evaluation stage, the complete process from input images to feature extraction and classification can be divided into the following steps. Taking the verification of a single picture as an example: First, load the input images of the second training set and convert them to the RGB format. Through preprocessing, ensure that the images meet the requirements of the model input. The preprocessing includes resizing the pictures to (341, 512), converting them into tensors, and normalizing. The preprocessed image tensors are input into the above-mentioned well-trained psoriasis syndrome classification model. Image features are extracted through the convolutional layer and classified through the fully connected layer. The output results are converted into a probability distribution through the softmax function, indicating the confidence of the model in each category. To understand the decision-making process of the model, calculate the softmax values of each category and use the Grad-CAM technique to generate a heatmap. Generate an activation map by calculating the gradient of the target category (here it is the category with the highest probability) with respect to the output of the last convolutional layer, showing the image regions that the model focuses on when making predictions, and overlay the generated activation map on the original image to form a heatmap. Taking Figure 5 as an example, the softmax values output by the psoriasis syndrome classification model are blood stasis syndrome = 0.0005, blood dryness syndrome = 0.0001, and blood heat syndrome = 0.9994. With blood heat syndrome having the highest confidence, it is predicted that the image belongs to the "blood heat syndrome". As Figure 5 shown, it can be seen that the heatmap shows that the psoriasis syndrome classification model focuses on the lesioned areas in the skin tissue, which is consistent with the areas of concern in clinical decision-making for syndrome differentiation.
[0092] In the overall model verification, the results show that the overall accuracy rate is approximately 87.66%. Among them, the precision rate of blood stasis syndrome is 83%, the recall rate is 80%, and the ROC-AUC (Receiver Operating Characteristic - Area Under the Curve) is 0.94; the precision rate of blood stasis syndrome is 89%, the recall rate is 92%, and the ROC-AUC is 0.94; the precision rate of blood heat syndrome is 83%, the recall rate is 80%, and the AUC is 0.97. The ROC details can be seen in Figure 5 .
[0093] Furthermore, 277 target lesioned pictures of the external validation set data are used for further evaluation of the model. The results show that the overall accuracy rate of the model is approximately 74.64%. Among them, the precision rate of blood stasis syndrome is 69%, the recall rate is 57%, and the ROC-AUC (Receiver Operating Characteristic - Area Under the Curve) is 0.79; the precision rate of blood stasis syndrome is 80%, the recall rate is 85%, and the ROC-AUC is 0.87; the precision rate of blood heat syndrome is 64%, the recall rate is 78%, and the AUC is 0.94. The ROC details can be seen in Figure 7 .
[0094] The traditional classification of TCM syndromes for psoriasis vulgaris has traditionally relied on doctors' experience, and syndrome differentiation and treatment are carried out by combining the four diagnostic methods of "inspection, auscultation and olfaction, interrogation, and palpation". It has become an industry consensus that the three major syndromes are blood-heat syndrome, blood-stasis syndrome, and blood-dryness syndrome. However, the traditional method has limitations such as strong subjectivity, low efficiency, lack of quantitative criteria, and limited knowledge inheritance, making it difficult to meet the modern medical needs for efficient and accurate diagnosis. The deep learning method based on target skin lesion images combines the theory of TCM syndrome classification with modern computer vision technology, providing a brand-new tool for TCM diagnosis. By automatically analyzing target skin lesion images and classifying syndromes, it effectively makes up for the deficiencies of the traditional method and demonstrates significant advantages. First, the deep learning method has a high degree of objectivity and consistency. Traditional diagnosis relies on doctors' experience and is easily affected by subjective factors. The deep learning model is trained with a large number of labeled target skin lesion images, can automatically learn syndrome-related features, and output consistent diagnostic results. For example, the model can accurately distinguish blood-heat syndrome, blood-stasis syndrome, and blood-dryness syndrome, and output the confidence of the quantitative diagnostic results through probability, providing an objective reference basis for doctors. Second, the deep learning method significantly improves the diagnostic efficiency. The traditional diagnostic process is time-consuming and difficult to handle a large number of patients. The deep learning system based on target skin lesion images can complete image analysis and syndrome classification within seconds, which is especially suitable for quickly screening a large number of patients in clinical and research needs, greatly reducing the work burden of doctors and researchers and reducing the training cost of researchers.
[0095] A method for classifying TCM syndromes of psoriasis based on skin images provided by this application only needs to take skin images of relevant positions by patients or doctors to directly obtain the corresponding TCM syndrome classification results, that is, the TCM syndrome labels of the skin. The present invention does not rely on the experience of TCM masters, reduces the error of relying on doctors' subjective judgment, improves the accuracy of classification, and reduces the risk of misdiagnosis and missed diagnosis. In addition, through the training of the deep learning model, this system has realized the standardization and regularization of the diagnostic process, making the syndrome determination results of different doctors and different medical institutions consistent, which helps to promote the popularization of standardized TCM syndromes for psoriasis.
[0096] Based on the same inventive concept, this application also provides a computer device, including: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of a method for classifying TCM syndromes of psoriasis based on skin images as above.
[0097] This application embodiment also provides a non-transitory machine-readable storage medium, on which an executable program is stored. When the executable program runs on the processor, it causes the processor to execute the method provided in the above embodiment.
[0098] An embodiment of the present invention discloses a computer-readable storage medium that stores a computer program for electronic data exchange, wherein the computer program causes a computer to execute the described method.
[0099] An embodiment of the present invention discloses a computer program product that includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the described method.
[0100] The above-described embodiments are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.
[0101] Through the above specific descriptions of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, a magnetic disk memory, a tape memory, or any other computer-readable medium capable of carrying or storing data.
[0102] Finally, it should be noted that: The embodiments disclosed in the present invention are only the preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than limiting it; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: They can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for classifying TCM syndromes of psoriasis based on skin images, characterized in that: include: Acquire a skin image, input the skin image into a skin recognition model, and obtain target skin lesion information; Determining a target skin lesion image according to the target skin lesion information; The target skin lesion image is input into a psoriasis syndrome classification model to obtain a classification result of the TCM syndrome corresponding to the skin image.
2. The method for classifying TCM syndromes of psoriasis based on skin images according to claim 1, characterized in that: The psoriasis syndrome classification model includes an input end, a convolution layer, a global average pooling layer, and a fully connected layer; the input end receives the target skin lesion image; the convolution layer includes a first convolution layer, a second convolution layer, a third convolution layer, and a fourth convolution layer, and multi-channel feature processing is performed through the first convolution layer, the second convolution layer, the third convolution layer, and the fourth convolution layer to obtain a multi-channel two-dimensional feature map; The global average pooling layer performs a global average pooling operation on the two-dimensional feature map of each channel to generate a feature vector; the fully connected layer obtains the score of each TCM syndrome classification corresponding to the target skin lesion image through the feature vector, and outputs the TCM syndrome classification result according to the score.
3. The method for classifying TCM syndromes of psoriasis based on skin images according to claim 2, characterized in that: The psoriasis syndrome classification model is trained by a second training image set and a cross entropy loss function, and the weights of the psoriasis syndrome classification model are updated by a back propagation algorithm.
4. The method for classifying TCM syndromes of psoriasis based on skin images according to claim 2, characterized in that: The first convolution layer is a 7×7 convolution layer, and the stride of the first convolution layer is 2; the first convolution layer is used to perform preliminary feature extraction on the target skin lesion image to obtain input features; The second convolutional layer includes two residual blocks, each of which includes two 3×3 convolutional layers, and the number of feature maps is 64; the input features are added to the output of the convolutional layer through the skip connection mechanism of the residual block; the second convolutional layer is used to extract the edge and texture information of the target skin lesion image; The third convolutional layer includes two residual blocks, each of which includes two 3×3 convolutional layers, and the number of feature map channels is 128; the third convolutional layer is used to perform deep processing on the feature map; The fourth convolutional layer includes two residual blocks, and the number of feature map channels is 264; the fourth convolutional layer is used to extract deep semantic features of the target skin lesion image to obtain a multi-channel two-dimensional feature map.
5. The method for classifying TCM syndromes of psoriasis based on skin images according to any one of claims 1 to 4, characterized in that: The skin recognition model includes an input end, a feature extraction network, a feature fusion network and an output end. The skin image enters through the input end, a feature map is extracted through the feature extraction network, feature fusion is performed through the feature fusion network, and target skin lesion information is output by the output end; the target skin lesion information includes skin tissue information and bounding box information of the predicted target, the skin tissue information indicates whether there is skin tissue in each grid unit in the feature map, and the bounding box information of the predicted target includes the center position of the bounding box and the parameters of the bounding box.
6. The method for classifying TCM syndromes of psoriasis based on skin images according to claim 5, characterized in that: The skin recognition model extracts feature maps through a feature extraction network, including: Slicing the skin image into a plurality of sub-images; splicing the sub-images in the channel dimension to form a multi-channel feature map; extracting local features of the skin image through a convolution operation to obtain an original feature map; Performing three maximum pooling operations on the original feature map to obtain a pooled feature map; The original feature map and the pooled feature map are concatenated to obtain global features.
7. The method for classifying TCM syndromes of psoriasis based on skin images according to claim 5, characterized in that: The skin recognition model performs feature fusion through the feature fusion network, including: The feature fusion network includes a feature pyramid network and a path aggregation network. The feature pyramid network uses an upsampling operation to gradually improve the resolution of deep features and splice the deep features with shallow features. The path aggregation network uses a downsampling operation to gradually reduce the resolution of shallow features and splice the shallow features with corresponding deep features.
8. The method for classifying TCM syndromes of psoriasis based on skin images according to claim 4, characterized in that: The skin recognition model is trained by annotating a first training image set and a multi-task loss function, and the multi-task loss function is used to optimize the model during each training cycle; the multi-task loss function includes a bounding box regression loss function, a target confidence loss function and a classification loss function, the bounding box regression loss function is a CIoU loss function, and the target confidence loss function and the classification loss function use a binary cross loss function.
9. A computer device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the method for classifying TCM syndromes of psoriasis based on skin images as claimed in any one of claims 1 to 8.
10. A computer storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method for classifying TCM syndromes of psoriasis based on skin images as claimed in any one of claims 1 to 8 are implemented.