A skin disease image classification method based on deep learning

By employing reflectivity-illuminance decomposition and an improved PuzzleMix mosaicking technique, combined with topological threshold constraints, the problem of illumination and acquisition domain differences in skin disease image classification was solved, resulting in a more stable skin disease image classification model.

CN122135098APending Publication Date: 2026-06-02YANGTZE UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YANGTZE UNIVERSITY
Filing Date
2026-02-27
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies for classifying skin disease images suffer from insufficient classification stability and generalization ability due to differences in illumination and acquisition domain, especially in long-tail categories and small sample scenarios.

Method used

By employing reflectivity-illumination intrinsic decomposition, lesion segmentation, improved PuzzleMix mosaicking, and topological threshold constraints, we generate mixed samples and train a classification model by replacing reflectivity components in skin disease images and performing domain adversarial training, thus ensuring the morphological consistency of lesion regions and the stability of the acquisition domain.

Benefits of technology

It improves the generalization performance and classification stability of the skin disease image classification model under cross acquisition domain conditions, reduces the impact of illumination differences and acquisition domain bias, and enhances the model's adaptability in long-tail and small-sample scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135098A_ABST
    Figure CN122135098A_ABST
Patent Text Reader

Abstract

This invention discloses a deep learning-based method for classifying skin disease images, comprising the following steps: acquiring a training set of skin disease images containing category labels and acquisition domain labels, and setting a topological threshold; performing reflectance-illuminance intrinsic decomposition on the main image and paired images to obtain the main reflectance component, the main illumination component, and the paired reflectance component; performing lesion segmentation on the main image to obtain a lesion mask and extracting the main topological signature; generating an improved PuzzleMix mosaic mask within the lesion mask based on the category evidence map and the domain evidence map; replacing the main reflectance component to generate a mixed reflectance component and synthesizing a mixed image; masking the mosaic pieces based on the topological threshold and updating the mosaic mask and the mixed image; generating mixed labels according to the area ratio of the main mosaic pieces and combining them with domain adversarial training to obtain a skin disease image classification model. This invention improves the cross-acquisition domain classification generalization ability and output stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent medical image analysis technology, and in particular to a deep learning-based method for classifying skin disease images. Background Technology

[0002] Intelligent classification of skin disease images belongs to the field of medical image analysis. Common acquisition methods in clinical scenarios include mobile phone shooting and dermoscopy imaging. Images contain changes in illumination, reflection, skin color differences, noise and occlusion. The boundaries of lesion areas are irregular and have significant morphological differences. Current practices mostly use convolutional neural networks to complete feature extraction and category discrimination, and combine them with supervised learning to train classification models using labeled data to achieve automatic identification and assisted diagnosis of different skin disease categories.

[0003] To alleviate the problems of high annotation costs and insufficient samples, existing technologies typically introduce data augmentation and sample mixing strategies, such as color perturbation, geometric transformation, random erasure, MixUp, CutMix, and PuzzleMix. At the same time, they combine domain adaptation methods to align differences in the acquisition domain. For example, they use domain discriminators and gradient inversion mechanisms to constrain feature distributions to be similar. Some schemes further introduce lesion segmentation to locate lesion regions, or use brightness normalization and white balance correction to improve input consistency, thereby reducing the impact of differences in acquisition equipment and environment on classification.

[0004] Existing technologies still have significant shortcomings: general enhancement and hybridization strategies often perform replacements on the basis of the entire image or a fixed grid, which can easily introduce non-lesion background into the hybrid region, leading to a mismatch between the semantics of the hybrid image and the label weights and generating label noise; when the mosaic replacement lacks constraints on the lesion morphology, it may destroy the boundary structure of the lesion and the continuity of the internal texture, forming samples that do not conform to the laws of medical imaging, thereby inducing the model to learn shortcut features and reducing the generalization ability across acquisition domains; domain adversarial training may suppress class-related discriminative features when there are no significant constraints, and acquisition domain differences will still remain in the evidence region and cause domain bias, resulting in insufficient classification stability under conditions of acquisition domain changes, light reflection changes, and long-tailed categories.

[0005] Therefore, how to provide a deep learning-based method for classifying skin disease images is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a deep learning-based method for classifying skin disease images. This invention combines reflectance-illuminance intrinsic decomposition, lesion segmentation, improved PuzzleMix mosaic mixing, topological threshold constraints, and domain adversarial training to construct mixed samples of skin disease images and train a classification model that addresses differences in acquisition domains. It has the advantages of strong cross-acquisition domain generalization ability, high classification stability, and good adaptability to long-tail and small-sample scenarios.

[0007] A skin disease image classification method based on deep learning according to an embodiment of the present invention includes the following steps: Obtain a training set of skin disease images, which includes category labels and acquisition domain labels, and set a topological threshold; Perform reflectance-illumination intrinsic decomposition on the subject image and the paired image to obtain the subject reflectance component, the subject illumination component and the paired reflectance component; The lesion mask is obtained by performing lesion segmentation on the main image, and the main topological signature is extracted within the lesion mask; The classifier calculates the category evidence map for the main image, the domain evidence map is calculated by the domain discriminator, the puzzle piece set is divided within the lesion mask, the category evidence score is calculated according to the category evidence map, the domain evidence score is calculated according to the domain evidence map, and the improved PuzzleMix puzzle mask is generated. The mixed reflectance component is obtained by replacing the corresponding puzzle pieces of the main reflectance component with the puzzle mask according to the matching reflectance component, and then synthesizing the mixed image with the main illumination component. Extract the hybrid topological signature of the hybrid image within the lesion mask. When the difference between the hybrid topological signature and the main topological signature exceeds the topological threshold, mask out the jigsaw puzzle pieces that cause the threshold to be exceeded and update the jigsaw puzzle mask and the hybrid image. Hybrid labels are generated based on the area ratio of the main puzzle pieces in the puzzle mask. A classifier is trained using hybrid images and hybrid labels. A domain discriminator is trained based on the domain evidence map and the acquired domain labels. Domain adversarial measures are applied to the classifier to obtain a skin disease image classification model.

[0008] Optionally, acquiring the training set of skin disease images specifically includes: Collect images of skin diseases, assign category labels and acquisition domain labels to each image, and build a training set of skin disease images containing skin disease images, category labels, and acquisition domain labels; In the dermatology image training set, each dermatology image is set as the main image, and a matching image is determined for the main image. The matching image is selected from the dermatology image training set and is different from the dermatology image corresponding to the main image. The paired images are determined according to the pairing priority, which is as follows: the same category label and different acquisition domain label, the same acquisition domain label and different category label, and different acquisition domain label and different category label. Set a topological threshold, generate lesion masks for each image in the skin disease image training set and extract the main body topological signature, calculate the difference between any two main body topological signatures, the difference between the main body topological signatures is obtained by summing the absolute differences of each dimension of the main body topological signature, sort the main body topological signature differences by value and take the middle main body topological signature difference as the topological threshold.

[0009] Optionally, obtaining the main reflectivity component, the main illumination component, and the paired reflectivity component specifically includes: Perform a logarithmic transformation on the pixel values ​​at pixel coordinates of the main image to obtain a logarithmic image; The logarithmic image is subjected to edge smoothing with a smoothing parameter to obtain a logarithmic illumination map. The logarithmic illumination map is then subjected to an exponential transformation to obtain the subject illumination component. The subject image is divided pixel by pixel at pixel coordinates by the subject illumination component to obtain a ratio map. The ratio map is then subjected to logarithmic and anti-logarithmic transformations to obtain the subject reflectivity component. The reconstructed image is obtained by multiplying the subject reflectance component and the subject illumination component pixel by pixel. The pixel difference between the reconstructed image and the subject image at pixel coordinates is calculated. Pixel coordinates with pixel differences exceeding the error threshold are subjected to local edge-preserving smoothing correction in the logarithmic illumination map, and the subject illumination component and the subject reflectance component are updated. The paired images are subjected to the same logarithmic transformation, edge smoothing, exponential transformation, pixel-by-pixel division, reconstruction consistency verification, and local correction processing as the main images to obtain the paired reflectance components.

[0010] Optionally, the step of performing lesion segmentation on the main image to obtain a lesion mask, and extracting the main topological signature within the lesion mask specifically includes: Pixel-level lesion segmentation is performed on the main image to obtain a lesion probability map, which records the lesion probability at pixel coordinates. The lesion probability map is compared with the segmentation threshold. The pixel coordinates of the lesion probability are marked as lesion pixels, and the lesion pixels constitute the lesion mask. The connected regions of the lesion mask are marked according to the four-adjacent connectivity rule and the area of ​​the connected regions is calculated. The connected region with the largest area constitutes the lesion mask. The lesion mask is filled with holes to obtain a filled mask. The difference between the filled mask and the lesion mask constitutes the hole region. The hole region is marked with connected domains according to the four-adjacent connectivity rule to obtain the hole count. Extract boundary pixels and calculate boundary perimeter from lesion mask, calculate mask area from lesion mask, calculate minimum bounding rectangle area from lesion mask, and the boundary area ratio is the ratio of minimum bounding rectangle area to mask area; Skeletonization is performed on the lesion mask to obtain a skeleton map. Skeleton pixels with a degree greater than 2 in the skeleton map are counted as skeleton bifurcation points, and the skeleton bifurcation point count is counted. Skeleton pixels with a degree equal to 1 in the skeleton map are counted as skeleton endpoints. The main topological signature is formed by arranging the hole count, boundary perimeter, boundary area ratio, skeleton bifurcation point count, and skeleton endpoint count in a fixed dimensional order.

[0011] Optionally, the generation of the improved PuzzleMix puzzle mask specifically includes: The classifier consists of a convolutional feature extraction layer, a global average pooling layer, and a fully connected classification layer. The convolutional feature extraction layer calculates a convolutional feature map on the main image. The global average pooling layer calculates the average of the convolutional feature map according to spatial coordinates to obtain a classification feature vector. The fully connected classification layer maps the classification feature vector to obtain a class score vector. The class score vector and the class label are used to calculate the class loss. The gradient of the category loss with respect to the convolutional feature map is calculated. The gradient is averaged over spatial coordinates to obtain the channel weights. The channel weights are then weighted and summed with the convolutional feature map over the channel dimension and normalized to obtain the category evidence map. The domain discriminator includes a gradient inversion layer and a fully connected domain classification layer. The gradient inversion layer connects the convolutional feature map and scales the gradient by taking the negative value during backpropagation. The fully connected domain classification layer obtains the domain feature vector by global average pooling of the convolutional feature map and maps it to obtain the domain score vector. The domain score vector and the collected domain label are used to calculate the domain loss. The gradient of the domain loss with respect to the convolutional feature map is calculated. The average of the gradient is calculated using spatial coordinates to obtain the domain channel weights. The domain channel weights are then weighted and summed with the convolutional feature map according to the channel dimension and normalized to obtain the domain evidence map. Within the lesion mask, a set of jigsaw puzzle blocks is divided according to the block side length parameter. Each jigsaw puzzle block set contains a jigsaw puzzle block identifier and a set of jigsaw puzzle block pixel coordinates. The class evidence score is obtained by averaging the pixel coordinates of the puzzle pieces on the class evidence map, and the domain evidence score is obtained by averaging the pixel coordinates of the puzzle pieces on the domain evidence map. Given a suppression weight parameter, the overall score of the puzzle piece is obtained by subtracting the product of the suppression weight parameter and the domain evidence score from the category evidence score; Set a replacement ratio parameter. The number of replacement puzzle pieces is obtained by multiplying the puzzle piece count by the replacement ratio parameter and rounding down. Sort the puzzle pieces in descending order of their overall scores. Select the puzzle pieces corresponding to the number of replacement puzzle pieces and mark them as replacement areas within the lesion mask to form the improved PuzzleMix puzzle mask.

[0012] Optionally, the generation of the mixed image specifically includes: The set of pixel coordinates of the replacement region is determined based on the puzzle mask of the improved PuzzleMix. The set of pixel coordinates of the replacement region is the set of pixel coordinates marked as the replacement region in the puzzle mask of the improved PuzzleMix. A mixed reflectance component is generated from the main reflectance component. The pixel values ​​of the mixed reflectance component within the set of pixel coordinates of the replacement region are assigned the pixel values ​​of the paired reflectance component at the same pixel coordinates. The pixel values ​​of the mixed reflectance component outside the set of pixel coordinates of the replacement region are assigned the pixel values ​​of the main reflectance component at the same pixel coordinates, thus obtaining the mixed reflectance component. The mixed image is obtained by multiplying the mixed reflectivity component and the main illumination component pixel by pixel at the same pixel coordinates.

[0013] Optionally, the updated mosaic mask and blended image specifically include: Within the set of pixel coordinates defined by the lesion mask, the hole count, boundary perimeter, boundary area ratio, skeleton bifurcation point count, and skeleton endpoint count are calculated for the hybrid image, and arranged in a fixed dimensional order according to the main topological signature to form a hybrid topological signature; Calculate the difference between the hybrid topology signature and the main topology signature. The difference is obtained by summing the absolute differences between the hybrid topology signature and the main topology signature in each dimension. When the difference between the hybrid topological signature and the main topological signature is greater than the topological threshold, the puzzle pieces marked as replacement regions in the improved PuzzleMix puzzle mask are selected, and a masking evaluation is performed on the puzzle pieces. The masking evaluation includes assigning the pixel values ​​in the puzzle piece pixel coordinate set to the pixel values ​​of the main reflectance component at the same pixel coordinates to generate candidate hybrid reflectance components. The candidate hybrid reflectance components are multiplied pixel by pixel with the main illumination component to generate candidate hybrid images. The candidate hybrid topological signature is extracted in the pixel coordinate set defined by the lesion mask, and the difference between the candidate hybrid topological signature and the main topological signature is calculated. The puzzle piece with the smallest difference between the candidate hybrid topological signature and the main topological signature is selected as the masked puzzle piece. The set of pixel coordinates of the masked puzzle piece is deleted from the replacement region of the puzzle mask of the improved PuzzleMix, and the hybrid reflectivity component and the hybrid image are updated. Extract the hybrid topological signature from the updated hybrid image and calculate the difference between the hybrid topological signature and the main topological signature. When the difference between the hybrid topological signature and the main topological signature is not greater than the topological threshold, complete the mosaic mask update and hybrid image update.

[0014] Optionally, the generation of the skin disease image classification model specifically includes: The pixel count of the lesion mask is counted within the lesion mask, and the pixel count of the replacement area is counted within the puzzle mask of the improved PuzzleMix. The area ratio of the main puzzle piece is obtained by dividing the difference between the pixel count of the lesion mask and the pixel count of the replacement area by the pixel count of the lesion mask. The area ratio of the paired puzzle piece is obtained by dividing the pixel count of the replacement area by the pixel count of the lesion mask. The main image category label is encoded as a main label vector, the paired image category label is encoded as a paired label vector, and the mixed label is obtained by adding the result of multiplying the area ratio of the main jigsaw puzzle piece by the main label vector element by element and the result of multiplying the area ratio of the paired jigsaw puzzle piece by the paired label vector element by element. The classifier calculates the class probability vector for the mixed image. The class probability vector and the mixed label are used to calculate the classification loss. The classification loss is used to update the classifier parameters. The domain probability vector is calculated by the domain discriminator on the subject image. The domain probability vector and the domain label acquired from the subject image are used to calculate the domain loss. The domain evidence map is averaged within the lesion mask to obtain the domain weight. The domain loss is multiplied by the domain weight to obtain the weighted domain loss. The domain discriminator parameters are updated based on the weighted domain loss. The weighted domain loss is then passed through the gradient inversion layer and multiplied by the domain adversarial coefficient before being used to update the classifier parameters.

[0015] The beneficial effects of this invention are: This invention decouples illumination factors from reflectance factors through reflectance-illuminance intrinsic decomposition. During the mosaic replacement stage, replacement is performed only on the reflectance component, and reconstruction is completed using the main illumination component. This ensures that the imaging brightness distribution of the mixed image remains consistent with the main image, thereby reducing the interference of illumination differences introduced by different acquisition conditions on classification features. At the same time, this invention generates an improved PuzzleMix mosaic mask within the lesion mask, uses the category evidence map to select replacement regions with greater category discriminative power, and uses the domain evidence map to suppress acquisition domain-related regions from entering the replacement region. This makes the mixed sample more concentrated on lesion discriminative evidence and weakens acquisition domain bias.

[0016] This invention further extracts the main topological signature and the mixed topological signature within the lesion mask, and uses a topological threshold to constrain the puzzle piece masking and puzzle mask updating, so that the mixed sample remains consistent with the main sample at the lesion morphology and structure level, reducing morphological distortion and label noise caused by puzzle replacement. Based on the joint optimization of supervised training and domain adversarial training of mixed labels, the classifier reduces the sensitivity of the acquisition domain while maintaining the class discrimination ability, thereby improving the generalization performance and output stability of the skin disease image classification model under cross acquisition domain conditions. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a deep learning-based skin disease image classification method proposed in this invention; Figure 2This is a schematic diagram illustrating the generation of an improved PuzzleMix mosaic mask for a deep learning-based skin disease image classification method proposed in this invention. Figure 3 This is a schematic diagram of the mosaic block masking under the topological threshold constraint of a deep learning-based skin disease image classification method proposed in this invention. Detailed Implementation

[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0019] refer to Figure 1-3 A deep learning-based method for classifying skin disease images includes the following steps: Obtain a training set of skin disease images, which includes category labels and acquisition domain labels, and set a topological threshold; Perform reflectance-illumination intrinsic decomposition on the subject image and the paired image to obtain the subject reflectance component, the subject illumination component and the paired reflectance component; The lesion mask is obtained by performing lesion segmentation on the main image, and the main topological signature is extracted within the lesion mask; The classifier calculates the category evidence map for the main image, the domain evidence map is calculated by the domain discriminator, the puzzle piece set is divided within the lesion mask, the category evidence score is calculated according to the category evidence map, the domain evidence score is calculated according to the domain evidence map, and the improved PuzzleMix puzzle mask is generated. The mixed reflectance component is obtained by replacing the corresponding puzzle pieces of the main reflectance component with the puzzle mask according to the matching reflectance component, and then synthesizing the mixed image with the main illumination component. Extract the hybrid topological signature of the hybrid image within the lesion mask. When the difference between the hybrid topological signature and the main topological signature exceeds the topological threshold, mask out the jigsaw puzzle pieces that cause the threshold to be exceeded and update the jigsaw puzzle mask and the hybrid image. Hybrid labels are generated based on the area ratio of the main puzzle pieces in the puzzle mask. A classifier is trained using hybrid images and hybrid labels. A domain discriminator is trained based on the domain evidence map and the acquired domain labels. Domain adversarial measures are applied to the classifier to obtain a skin disease image classification model.

[0020] In this embodiment, obtaining the skin disease image training set specifically includes: Collect skin disease images, assign a unique image identifier to each skin disease image, record a category label and a collection domain label for each image identifier, and combine the image identifier, skin disease image, category label, and collection domain label to form a skin disease image training set. The value range of the category label is a preset category set, and the value range of the collection domain label is a preset collection domain set. In the dermatology image training set, a subject image setting is performed for each image identifier. The subject image setting includes the subject image identifier and the subject image dermatology image. For each subject image identifier, a matching image identifier is selected. The matching image identifier is different from the subject image identifier. The matching image setting includes the matching image identifier and the matching image dermatology image. Image identifiers for paired images are determined based on pairing priority. The first-level filtering condition is that the category label is the same as the category label of the main image and the acquisition field label is different from the acquisition field label of the main image. When the count of the first-level filtering result is 0, the second-level filtering condition is used. The second-level filtering condition is that the acquisition field label is the same as the acquisition field label of the main image and the category label is different from the category label of the main image. When the count of the second-level filtering result is 0, the third-level filtering condition is used. The third-level filtering condition is that the acquisition field label is different from the acquisition field label of the main image and the category label is different from the category label of the main image. When the count of the filtering result is greater than 0, image identifiers for paired images are determined according to preset sorting rules. A topological threshold is set. The topological threshold setting includes the preparation of topological threshold calculation data and the statistical value of topological threshold. The preparation of topological threshold calculation data includes generating lesion masks for each image in the skin disease image training set and extracting the main body topological signature. The main body topological signature is a vector arranged in a fixed dimension order. The statistical value of topological threshold includes calculating the difference between the main body topological signatures of any two images and collecting the difference set. The difference between the main body topological signatures is obtained by summing the absolute values ​​of the numerical differences between the main body topological signatures in the same dimension. The difference set is sorted in ascending order of numerical value, and the difference whose sequence number is equal to half of the total number of differences and rounded up is taken as the topological threshold.

[0021] In this embodiment, obtaining the main reflectivity component, the main illumination component, and the paired reflectivity component specifically includes: A logarithmic transformation is performed on the pixel value at each pixel coordinate of the main image. The logarithmic transformation uses the pixel value plus a bias as the logarithmic independent variable, and the bias is a positive constant, to obtain the logarithmic image. The logarithmic image is subjected to edge smoothing, and the smoothing intensity is controlled by the smoothing parameter. Edge smoothing performs weighted smoothing on the values ​​at pixel coordinates in the spatial neighborhood and keeps the gradient abrupt locations unsmoothed to obtain a logarithmic illumination map. The logarithmic illumination map is then subjected to an exponential transformation to obtain the main illumination component. The ratio map is obtained by dividing the subject image pixel by pixel at each pixel coordinate by the subject illumination component. The ratio map is then logarithmically transformed at each pixel coordinate to obtain the logarithmic reflectance map. Finally, the logarithmic reflectance map is anti-logarithmically transformed at each pixel coordinate to obtain the subject reflectance component. The main reflectance component and the main illumination component are multiplied pixel by pixel at each pixel coordinate to obtain the reconstructed image. The pixel difference between the reconstructed image and the main image at each pixel coordinate is calculated. Pixel coordinates with pixel differences exceeding the error threshold are subjected to local edge preservation smoothing correction within the logarithmic illumination map. The local edge preservation smoothing correction uses smoothing parameters and limits the correction neighborhood radius. After updating the logarithmic illumination map, an exponential transformation is performed to update the main illumination component. Then, the main reflectance component is updated by pixel-by-pixel division. The paired images are subjected to logarithmic transformation, edge-preserving smoothing, exponential transformation, pixel-by-pixel division, reconstruction consistency verification, and local correction processing with the same bias, smoothing parameters, error threshold, and corrected neighborhood radius as the main image, to obtain the paired reflectance components.

[0022] In this embodiment, the step of performing lesion segmentation on the main image to obtain a lesion mask, and extracting the main topological signature within the lesion mask specifically includes: Pixel-level lesion segmentation is performed on the main image. The pixel-level lesion segmentation outputs the lesion probability at each pixel coordinate. The lesion probabilities at all pixel coordinates are combined to form a lesion probability map. The lesion probability map is compared with the segmentation threshold, which is a scalar threshold parameter. When the lesion probability at a pixel coordinate is greater than or equal to the segmentation threshold, the pixel coordinate is marked as a lesion pixel. All lesion pixels form a lesion mask, which is a binary map. The lesion mask is marked with connected components according to the four-adjacent connectivity rule. The four-adjacent connectivity rule uses horizontal and vertical adjacency as connectivity conditions. The number of lesion pixels in each connected component is counted to obtain the area of ​​the connected component. The connected component with the largest area is retained as the lesion mask, and the connected components with an area smaller than the largest area are removed from the lesion mask. A hole filling is performed on the lesion mask. The hole filling process sets the non-lesion pixels surrounded by lesion pixels inside the lesion mask as lesion pixels to obtain the filling mask. The filling mask is subtracted from the lesion mask pixel by pixel to obtain the hole region. The hole region is marked according to the four-adjacent connectivity rule and the number of connected regions is counted to obtain the hole count. Extract boundary pixels from the lesion mask. Boundary pixels are the set of pixel coordinates in the lesion pixels that have four adjacent non-lesion pixels. The perimeter of the boundary is the number of boundary pixels multiplied by the pixel spacing coefficient, which is a preset constant. The mask area is obtained by counting the number of lesion pixels in the lesion mask. The minimum bounding rectangle area is calculated for the lesion mask. The minimum bounding rectangle area is determined by the minimum and maximum values ​​of the lesion pixels in the horizontal coordinate direction and the minimum and maximum values ​​in the vertical coordinate direction. The boundary area ratio is the ratio of the minimum bounding rectangle area to the mask area. Skeletonization is performed on the lesion mask. Skeletonization refines the lesion mask into a skeleton map with a width of one pixel. The degree of the skeleton pixel in the skeleton map is the number of skeleton pixels within the four adjacent ranges of the skeleton pixel. Skeleton pixels with a degree greater than 2 are counted as skeleton bifurcation points and the skeleton bifurcation point count is counted. Skeleton pixels with a degree equal to 1 are counted as skeleton endpoints and the skeleton endpoint count is counted. The main topological signature is formed by arranging the hole count, boundary perimeter, boundary area ratio, skeleton bifurcation point count, and skeleton endpoint count in a fixed dimensional order. The fixed dimensional order is set during the method initialization phase and remains unchanged during training.

[0023] In this embodiment, the generation of the improved PuzzleMix puzzle mask specifically includes: The classifier consists of a convolutional feature extraction layer, a global average pooling layer, and a fully connected classification layer. The convolutional feature extraction layer calculates a convolutional feature map on the main image. The convolutional feature map is a three-dimensional tensor containing channel dimension and spatial coordinate dimension. The global average pooling layer calculates the average value of the convolutional feature map channel by channel in the spatial coordinate dimension to form a classification feature vector. The fully connected classification layer performs a linear mapping on the classification feature vector to obtain a class score vector. The class loss is calculated from the class score vector and the class label. The gradient of the category loss with respect to the convolutional feature map is calculated. The gradient has the same dimension as the convolutional feature map. The gradient is averaged channel by channel in the spatial coordinate dimension to obtain the channel weights. The channel weights are multiplied with the convolutional feature map in the channel dimension and summed in the channel dimension to obtain the two-dimensional response map. The two-dimensional response map is linearly normalized according to the minimum and maximum values ​​to obtain the category evidence map. The domain discriminator includes a gradient inversion layer and a fully connected domain classification layer. The gradient inversion layer passes the convolutional feature map during the forward propagation stage and multiplies the gradient from the domain discriminator by a negative scaling factor during the backward propagation stage. The fully connected domain classification layer calculates the mean of the convolutional feature map channel by channel in the spatial coordinate dimension to form a domain feature vector and performs a linear mapping to obtain a domain score vector. The domain loss is calculated by the domain score vector and the collected domain labels. The gradient of the domain loss with respect to the convolutional feature map is calculated. The gradient is averaged channel by channel in the spatial coordinate dimension to obtain the domain channel weights. The domain channel weights are multiplied with the convolutional feature map along the channel dimension and summed along the channel dimension to obtain the domain response map. The domain response map is linearly normalized according to the minimum and maximum values ​​to obtain the domain evidence map. Within the lesion mask, a set of jigsaw puzzle blocks is divided according to the block side length parameter, where the block side length parameter is a positive integer. The lesion mask forms a grid block according to the block side length parameter. The intersection of the grid block and the lesion mask yields the set of pixel coordinates of the jigsaw puzzle block. The jigsaw puzzle block identifier is obtained by concatenating the grid row number and the grid column number. For each set of pixel coordinates of a puzzle piece, extract the category evidence value on the category evidence map and calculate the average to obtain the category evidence score. For each set of pixel coordinates of a puzzle piece, extract the domain evidence value on the domain evidence map and calculate the average to obtain the domain evidence score. The category evidence score and the domain evidence score are scalars. Let the suppression weight parameter be positive. The overall score of the puzzle piece is obtained by subtracting the product of the suppression weight parameter and the domain evidence score from the category evidence score. Let the replacement ratio parameter be a ratio value between 0 and 1. The number of replacement puzzle pieces is obtained by multiplying the puzzle piece count by the replacement ratio parameter and rounding down. The puzzle pieces are sorted from largest to smallest according to their overall puzzle piece score. The puzzle pieces that are ranked first and whose number is equal to the number of replacement puzzle pieces are marked as replacement areas in the lesion mask. The set of pixel coordinates of all replacement areas constitutes the puzzle mask of the improved PuzzleMix. The improved PuzzleMix calculates category evidence scores and domain evidence scores by puzzle pieces within the lesion mask. This suppresses the weighting of domain evidence scores into the overall puzzle piece score and selects replacement regions to form the puzzle mask according to the replacement ratio parameter. This makes the mixed image dominated by lesion regions with high category evidence and low domain evidence, thereby reducing shortcut feature learning caused by acquisition domain bias and reducing label noise. This improves the generalization ability and stability of the skin disease image classification model in cross-acquisition domain scenarios.

[0024] In this embodiment, the generation of the hybrid image specifically includes: The set of pixel coordinates of the replacement region is determined based on the puzzle mask of the improved PuzzleMix. The set of pixel coordinates of the replacement region consists of all pixel coordinates that satisfy the replacement mark of the puzzle mask. The replacement mark is a binary mark and takes the value of 0 or 1 at the pixel coordinate. A mixed reflectance component is generated from the main reflectance component, and the initial pixel value of the mixed reflectance component at each pixel coordinate is set to the pixel value of the main reflectance component at the same pixel coordinate. For each pixel coordinate, determine whether the pixel coordinate belongs to the set of pixel coordinates of the replacement region. If the pixel coordinate belongs to the set of pixel coordinates of the replacement region, set the pixel value of the mixed reflectance component at the pixel coordinate to the pixel value of the paired reflectance component at the same pixel coordinate. If the pixel coordinate does not belong to the set of pixel coordinates of the replacement region, keep the pixel value of the mixed reflectance component at the pixel coordinate equal to the pixel value of the main reflectance component at the same pixel coordinate, and obtain the mixed reflectance component. The mixed reflectivity component and the main illumination component are multiplied pixel by pixel at each pixel coordinate to obtain the mixed image. The product result at each pixel coordinate is written into the pixel value of the mixed image at the pixel coordinate.

[0025] In this embodiment, updating the mosaic mask and the blended image specifically includes: Within the set of pixel coordinates defined by the lesion mask, the hole count, boundary perimeter, boundary area ratio, skeleton bifurcation point count, and skeleton endpoint count are calculated for the hybrid image. The hole count is obtained by performing hole filling on the lesion mask to obtain the filling mask, and the hole region is obtained by the difference between the filling mask and the lesion mask. The number of connected components is then counted according to the four-adjacent connectivity rule. The boundary perimeter is obtained by extracting the boundary pixels of the lesion mask and counting the number of boundary pixels multiplied by the pixel spacing coefficient. The boundary area ratio is obtained by calculating the ratio of the area of ​​the minimum bounding rectangle of the lesion mask to the area of ​​the lesion mask. The skeleton bifurcation point count and skeleton endpoint count are obtained by performing skeletonization on the lesion mask to obtain the skeleton graph and counting the four-adjacent degree of the skeleton pixels. They are then arranged in the fixed dimensional order of the main topological signature to form a hybrid topological signature. Calculate the difference between the hybrid topology signature and the main topology signature. The difference is calculated by taking the absolute value of the difference between the hybrid topology signature component and the main topology signature component in each dimension and summing it over all dimensions. The difference is then compared with the topology threshold. When the difference between the hybrid topological signature and the main topological signature is greater than the topological threshold, a masking evaluation is performed on each puzzle piece corresponding to the replacement region in the improved PuzzleMix puzzle mask. The masking evaluation assigns the pixel value in the puzzle piece pixel coordinate set to the pixel value of the main reflectivity component at the same pixel coordinate in the hybrid reflectivity component to form a candidate hybrid reflectivity component. The candidate hybrid reflectivity component and the main illumination component are multiplied pixel by pixel at the same pixel coordinate to form a candidate hybrid image. The candidate hybrid image is extracted from the pixel coordinate set defined by the lesion mask according to the aforementioned hole count, boundary perimeter, boundary area ratio, skeleton bifurcation point count, and skeleton endpoint count, and the difference between the candidate hybrid topological signature and the main topological signature is calculated. The puzzle piece with the smallest difference between the candidate hybrid topological signature and the main topological signature is selected as the masked puzzle piece. The set of pixel coordinates of the masked puzzle piece is deleted from the replacement area of ​​the puzzle mask of the improved PuzzleMix and the set of pixel coordinates of the masked puzzle piece is retained as the set of pixel coordinates of the non-replacement area. The hybrid reflectivity component is updated and the hybrid image is updated by multiplying the updated hybrid reflectivity component with the main illumination component pixel by pixel. Extract the hybrid topological signature from the updated hybrid image and calculate the difference between the hybrid topological signature and the main topological signature. When the difference between the hybrid topological signature and the main topological signature is not greater than the topological threshold, end the masking evaluation and output the jigsaw mask and the hybrid image.

[0026] In this embodiment, the generation of the skin disease image classification model specifically includes: The pixel count of the lesion mask is obtained by counting the pixel coordinates of the binary mask with a value of 1 within the lesion mask. The pixel count of the replacement region is obtained by counting the pixel coordinates of the replacement marker with a value of 1 within the puzzle mask of the improved PuzzleMix. The area ratio of the main puzzle piece is obtained by subtracting the pixel count of the replacement region from the pixel count of the lesion mask and then dividing the difference by the pixel count of the lesion mask. The area ratio of the paired puzzle piece is obtained by dividing the pixel count of the replacement region by the pixel count of the lesion mask. Map the category labels of the main image to a main label vector. The length of the main label vector is equal to the number of values ​​of the category label. The main label vector is assigned a value of 1 at the corresponding position of the category label and a value of 0 at the other positions. Map the category labels of the paired images to a paired label vector. The paired label vector is assigned a value of 1 at the corresponding position of the category label and a value of 0 at the other positions. The first scaling vector is obtained by scaling the main label vector element by element according to the area ratio of the main puzzle piece, and the second scaling vector is obtained by scaling the paired label vector element by element according to the area ratio of the paired puzzle piece. The first scaling vector and the second scaling vector are added element by element to obtain the mixed label. The class probability vector is obtained by performing forward computation on the mixed image in the classifier. The class probability vector is obtained by normalizing the class score vector. The classification loss is calculated by the class probability vector and the mixed label. The classification loss is used to update the classifier parameters in the backpropagation. The domain probability vector is obtained by performing forward computation on the subject image in the domain discriminator. The domain probability vector is obtained by normalizing the domain score vector. The domain loss is calculated by combining the domain probability vector with the domain label acquired from the subject image. The domain evidence map is averaged on the domain evidence values ​​within the lesion mask to obtain the domain weight. The domain loss is multiplied by the domain weight to obtain the weighted domain loss. The weighted domain loss is used to update the domain discriminator parameters during backpropagation. The weighted domain loss is passed through the gradient inversion layer and multiplied by the domain adversarial coefficient before participating in the classifier parameter update. The gradient inversion layer negatively scales the gradient during the backpropagation stage.

[0027] Example 1: To verify the feasibility of this invention in practice, it was applied to an image-assisted classification scenario for common dermatological diseases. In actual business, image sources include mobile phone photography and dermoscopic imaging. Differences in acquisition domains mainly manifest in light intensity, reflection position, skin color distribution, imaging resolution, and white balance shift, resulting in significant appearance differences for the same category in different acquisition domains. At the same time, the training data exhibits a long-tail distribution, with sufficient images for head categories and fewer samples for tail categories. Traditional training easily learns shortcut features related to the acquisition domain, leading to increased misclassification rate and unstable output during cross-acquisition domain testing. In this embodiment, a training set of dermatological images is constructed with 8 dermatological disease categories. The training set contains 3 acquisition domain labels, totaling 12,360 images, of which approximately 2,600 are for the head category and approximately 480 are for the tail category. The training set, validation set, and test set are divided into 70% / 10% / 20% according to category and acquisition domain, and an additional cross-acquisition domain test set is constructed. The cross-acquisition domain test set only retains images that were not used as the main acquisition domain during the training phase to test the generalization ability.

[0028] In the application process, a training set of skin disease images is first acquired, and category labels and acquisition domain labels are assigned to each image. Simultaneously, lesion masks are generated for each image based on the training set, and the subject's topological signature is extracted. The difference in the subject's topological signature is calculated, and the difference at the center of the sorted values ​​is used as a topological threshold to constrain the consistency of lesion morphology in mixed samples. Subsequently, for each subject image, paired images are selected according to pairing priority, and reflectance-illuminance intrinsic decomposition is performed to obtain the subject reflectance component, the subject illumination component, and the paired reflectance component. Lesion segmentation is then performed on the subject image to obtain the lesion mask, and... The topological signature of the subject is extracted within the lesion mask. Then, the category evidence map and the domain evidence map are calculated on the subject image. The process is as follows: The classifier consists of a convolutional feature extraction layer, a global average pooling layer, and a fully connected classification layer. The gradient of the convolutional feature map is calculated based on the category loss to obtain the channel weights. The channel weights are weighted and summed with the convolutional feature map and normalized to obtain the category evidence map. The domain discriminator consists of a gradient inversion layer and a fully connected domain classification layer. The gradient of the convolutional feature map is calculated based on the domain loss to obtain the domain channel weights. The domain channel weights are weighted and summed with the convolutional feature map and normalized to obtain the domain evidence map. Within the lesion mask, a set of jigsaw puzzle pieces is divided according to the block side length parameter. For each piece, the mean value is calculated on both the category evidence map and the domain evidence map to obtain the category evidence score and the domain evidence score. The overall score of the jigsaw puzzle piece is obtained by subtracting the product of the suppression weight parameter and the domain evidence score from the category evidence score. Then, the number of replacement jigsaw puzzle pieces is determined according to the replacement ratio parameter, and the replacement regions are selected according to the overall score to generate an improved PuzzleMix jigsaw puzzle mask. Then, according to the jigsaw puzzle mask, the corresponding jigsaw puzzle pieces on the main reflectivity component are replaced with paired reflectivity components to obtain the mixed reflectivity component. This mixed reflectivity component is then multiplied pixel by pixel with the main illumination component to synthesize a mixed image. After the mixed image is completed, the mixed topological signature is extracted within the lesion mask and the difference is calculated with the main topological signature. When the difference exceeds the topological threshold, the jigsaw puzzle pieces in the replacement region are subjected to masking evaluation. The corresponding regions are replaced block by block with the main reflectance component to generate candidate mixed images and calculate the candidate difference. The puzzle block with the smallest candidate difference is selected as the masked puzzle block and the puzzle mask and mixed image are updated until the difference is no greater than the topological threshold. Finally, the area ratio of the main puzzle block is calculated according to the puzzle mask to generate mixed labels. The classifier is trained using the mixed image and mixed labels. At the same time, the domain discriminator is trained based on the domain evidence map and the acquired domain labels. The domain adversarial signal is applied to the classifier parameter update through the gradient inversion layer and the domain adversarial coefficient to obtain the skin disease image classification model. In this embodiment, the block side length parameter is 32, the replacement ratio parameter is 0.40, the suppression weight parameter is 0.70, the domain adversarial coefficient is 0.20, and the topological threshold is obtained by statistics from the training set and remains unchanged throughout the training process.

[0029] To verify the improvement effect of this invention compared to existing solutions, four comparative schemes were set up and evaluated under the same training set partition and the same training round: Scheme A is basic training without using jigsaw puzzle blending and domain adversarial training; Scheme B uses general jigsaw puzzle blending but does not introduce domain evidence graph suppression and topology threshold masking; Scheme C introduces domain adversarial training but does not introduce domain evidence graphs to participate in jigsaw puzzle mask generation and does not introduce topology threshold masking; Scheme D is the scheme of this invention. Evaluation metrics include overall accuracy (Acc), macro-average F1, cross-acquisition domain accuracy (Acc-CD), cross-acquisition domain macro-average F1-CD, and output calibration error (ECE). The statistical results are shown in the table below: Table 1. Cross-domain classification comparison results

[0030] Based on the cross-acquisition domain test set, the accuracy of Scheme A decreased significantly when the acquisition domain changed, with some tail categories having an F1 score below 0.55, demonstrating its sensitivity to acquisition domain bias and long-tail issues. Although Scheme B improved the overall performance, the improvement across acquisition domains was limited, and misjudgments occurred in images with strong reflections and unclear boundaries due to inconsistent mixed sample morphology. Scheme C continued to improve the cross-acquisition domain performance after introducing domain adversarial analysis, but it was still observed that some mosaic replacement areas were concentrated in domain-sensitive textures, resulting in limited improvement in cross-domain F1 scores. Scheme D improves the overall Acc by 0.023, while improving the cross-acquisition domain Acc-CD by 0.042, the Macro-F1-CD by 0.066, and reducing the ECE by 0.018, indicating that it maintains the classification and discrimination ability while providing more stable output. Statistics on the tail category show that the average F1 of the tail category improved from 0.681 in Scheme C to 0.742, and the average F1 of the tail category under the cross-acquisition domain condition improved from 0.608 to 0.705. This reflects that the improved PuzzleMix, by selecting replacement regions with category evidence scores within the lesion mask, suppressing domain-related regions with domain evidence scores, and using topological thresholds to shield inconsistent puzzle pieces, can reduce shortcut feature learning caused by differences in acquisition domains and reduce mixed label noise, thereby improving the classification stability in cross-acquisition domain generalization and long-tail scenarios.

[0031] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A deep learning-based method for classifying skin disease images, characterized in that, Includes the following steps: Obtain a training set of skin disease images, which includes category labels and acquisition domain labels, and set a topological threshold; Perform reflectance-illumination intrinsic decomposition on the subject image and the paired image to obtain the subject reflectance component, the subject illumination component and the paired reflectance component; The lesion mask is obtained by performing lesion segmentation on the main image, and the main topological signature is extracted within the lesion mask; The classifier calculates the category evidence map for the main image, the domain evidence map is calculated by the domain discriminator, the puzzle piece set is divided within the lesion mask, the category evidence score is calculated according to the category evidence map, the domain evidence score is calculated according to the domain evidence map, and the improved PuzzleMix puzzle mask is generated. The mixed reflectance component is obtained by replacing the corresponding puzzle pieces of the main reflectance component with the puzzle mask according to the matching reflectance component, and then synthesizing the mixed image with the main illumination component. Extract the hybrid topological signature of the hybrid image within the lesion mask. When the difference between the hybrid topological signature and the main topological signature exceeds the topological threshold, mask out the jigsaw puzzle pieces that cause the threshold to be exceeded and update the jigsaw puzzle mask and the hybrid image. Hybrid labels are generated based on the area ratio of the main puzzle pieces in the puzzle mask. A classifier is trained using hybrid images and hybrid labels. A domain discriminator is trained based on the domain evidence map and the acquired domain labels. Domain adversarial measures are applied to the classifier to obtain a skin disease image classification model.

2. The skin disease image classification method based on deep learning according to claim 1, characterized in that, The acquisition of the skin disease image training set specifically includes: Collect images of skin diseases, assign category labels and acquisition domain labels to each image, and build a training set of skin disease images containing skin disease images, category labels, and acquisition domain labels; In the dermatology image training set, each dermatology image is set as the main image, and a matching image is determined for the main image. The matching image is selected from the dermatology image training set and is different from the dermatology image corresponding to the main image. The paired images are determined according to the pairing priority, which is as follows: the same category label and different acquisition domain label, the same acquisition domain label and different category label, and different acquisition domain label and different category label. Set a topological threshold, generate lesion masks for each image in the skin disease image training set and extract the main body topological signature, calculate the difference between any two main body topological signatures, the difference between the main body topological signatures is obtained by summing the absolute differences of each dimension of the main body topological signature, sort the main body topological signature differences by value and take the middle main body topological signature difference as the topological threshold.

3. The skin disease image classification method based on deep learning according to claim 1, characterized in that, The specific methods for obtaining the main reflectivity component, the main illumination component, and the paired reflectivity component include: Perform a logarithmic transformation on the pixel values ​​at pixel coordinates of the main image to obtain a logarithmic image; The logarithmic image is subjected to edge smoothing with a smoothing parameter to obtain a logarithmic illumination map. The logarithmic illumination map is then subjected to an exponential transformation to obtain the subject illumination component. The subject image is divided pixel by pixel at pixel coordinates by the subject illumination component to obtain a ratio map. The ratio map is then subjected to logarithmic and anti-logarithmic transformations to obtain the subject reflectivity component. The reconstructed image is obtained by multiplying the subject reflectance component and the subject illumination component pixel by pixel. The pixel difference between the reconstructed image and the subject image at pixel coordinates is calculated. Pixel coordinates with pixel differences exceeding the error threshold are subjected to local edge-preserving smoothing correction in the logarithmic illumination map, and the subject illumination component and the subject reflectance component are updated. The paired images are subjected to the same logarithmic transformation, edge smoothing, exponential transformation, pixel-by-pixel division, reconstruction consistency verification, and local correction processes as the main images to obtain the paired reflectance components.

4. The skin disease image classification method based on deep learning according to claim 1, characterized in that, The step of performing lesion segmentation on the main image to obtain a lesion mask, and extracting the main topological signature within the lesion mask specifically includes: Pixel-level lesion segmentation is performed on the main image to obtain a lesion probability map, which records the lesion probability at pixel coordinates. The lesion probability map is compared with the segmentation threshold. The pixel coordinates of the lesion probability are marked as lesion pixels, and the lesion pixels constitute the lesion mask. The connected regions of the lesion mask are marked according to the four-adjacent connectivity rule and the area of ​​the connected regions is calculated. The connected region with the largest area constitutes the lesion mask. The lesion mask is filled with holes to obtain a filling mask. The difference between the filling mask and the lesion mask constitutes the hole region. The hole region is marked with connected components according to the four-adjacent connectivity rule to obtain the hole count. Extract boundary pixels and calculate boundary perimeter from lesion mask, calculate mask area from lesion mask, calculate minimum bounding rectangle area from lesion mask, and the boundary area ratio is the ratio of minimum bounding rectangle area to mask area; Skeletonization is performed on the lesion mask to obtain a skeleton map. Skeleton pixels with a degree greater than 2 in the skeleton map are counted as skeleton bifurcation points, and the skeleton bifurcation point count is counted. Skeleton pixels with a degree equal to 1 in the skeleton map are counted as skeleton endpoints. The main topological signature is formed by arranging the hole count, boundary perimeter, boundary area ratio, skeleton bifurcation point count, and skeleton endpoint count in a fixed dimensional order.

5. The skin disease image classification method based on deep learning according to claim 1, characterized in that, The generation of the improved PuzzleMix puzzle mask specifically includes: The classifier consists of a convolutional feature extraction layer, a global average pooling layer, and a fully connected classification layer. The convolutional feature extraction layer calculates a convolutional feature map on the main image. The global average pooling layer calculates the average of the convolutional feature map according to spatial coordinates to obtain a classification feature vector. The fully connected classification layer maps the classification feature vector to obtain a class score vector. The class score vector and the class label are used to calculate the class loss. The gradient of the category loss with respect to the convolutional feature map is calculated. The gradient is averaged over spatial coordinates to obtain the channel weights. The channel weights are then weighted and summed with the convolutional feature map over the channel dimension and normalized to obtain the category evidence map. The domain discriminator includes a gradient inversion layer and a fully connected domain classification layer. The gradient inversion layer connects the convolutional feature map and scales the gradient by taking the negative value during backpropagation. The fully connected domain classification layer obtains the domain feature vector by global average pooling of the convolutional feature map and maps it to obtain the domain score vector. The domain score vector and the collected domain label are used to calculate the domain loss. The gradient of the domain loss with respect to the convolutional feature map is calculated. The average of the gradient is calculated using spatial coordinates to obtain the domain channel weights. The domain channel weights are then weighted and summed with the convolutional feature map according to the channel dimension and normalized to obtain the domain evidence map. Within the lesion mask, a set of jigsaw puzzle blocks is divided according to the block side length parameter. Each jigsaw puzzle block set contains a jigsaw puzzle block identifier and a set of jigsaw puzzle block pixel coordinates. The class evidence score is obtained by averaging the pixel coordinates of the puzzle pieces on the class evidence map, and the domain evidence score is obtained by averaging the pixel coordinates of the puzzle pieces on the domain evidence map. Given a suppression weight parameter, the overall score of the puzzle piece is obtained by subtracting the product of the suppression weight parameter and the domain evidence score from the category evidence score; Set a replacement ratio parameter. The number of replacement puzzle pieces is obtained by multiplying the puzzle piece count by the replacement ratio parameter and rounding down. Sort the puzzle pieces in descending order of their overall scores. Select the puzzle pieces corresponding to the number of replacement puzzle pieces and mark them as replacement areas within the lesion mask to form the improved PuzzleMix puzzle mask.

6. The skin disease image classification method based on deep learning according to claim 1, characterized in that, The generation of the hybrid image specifically includes: The set of pixel coordinates of the replacement region is determined based on the puzzle mask of the improved PuzzleMix. The set of pixel coordinates of the replacement region is the set of pixel coordinates marked as the replacement region in the puzzle mask of the improved PuzzleMix. A mixed reflectance component is generated from the main reflectance component. The pixel values ​​of the mixed reflectance component within the set of pixel coordinates of the replacement region are assigned the pixel values ​​of the paired reflectance component at the same pixel coordinates. The pixel values ​​of the mixed reflectance component outside the set of pixel coordinates of the replacement region are assigned the pixel values ​​of the main reflectance component at the same pixel coordinates, thus obtaining the mixed reflectance component. The mixed image is obtained by multiplying the mixed reflectance component and the main illumination component pixel by pixel at the same pixel coordinates.

7. The skin disease image classification method based on deep learning according to claim 1, characterized in that, The updated mosaic mask and blended image specifically include: Within the set of pixel coordinates defined by the lesion mask, the hole count, boundary perimeter, boundary area ratio, skeleton bifurcation point count, and skeleton endpoint count are calculated for the hybrid image, and arranged in a fixed dimensional order according to the main topological signature to form a hybrid topological signature; Calculate the difference between the hybrid topology signature and the main topology signature. The difference is obtained by summing the absolute differences between the hybrid topology signature and the main topology signature in each dimension. When the difference between the hybrid topological signature and the main topological signature is greater than the topological threshold, the puzzle pieces marked as replacement regions in the improved PuzzleMix puzzle mask are selected, and a masking evaluation is performed on the puzzle pieces. The masking evaluation includes assigning the pixel values ​​in the puzzle piece pixel coordinate set to the pixel values ​​of the main reflectance component at the same pixel coordinates to generate candidate hybrid reflectance components. The candidate hybrid reflectance components are multiplied pixel by pixel with the main illumination component to generate candidate hybrid images. The candidate hybrid topological signature is extracted in the pixel coordinate set defined by the lesion mask, and the difference between the candidate hybrid topological signature and the main topological signature is calculated. The puzzle piece with the smallest difference between the candidate hybrid topological signature and the main topological signature is selected as the masked puzzle piece. The set of pixel coordinates of the masked puzzle piece is deleted from the replacement region of the puzzle mask of the improved PuzzleMix, and the hybrid reflectivity component and the hybrid image are updated. Extract the hybrid topological signature from the updated hybrid image and calculate the difference between the hybrid topological signature and the main topological signature. When the difference between the hybrid topological signature and the main topological signature is not greater than the topological threshold, complete the mosaic mask update and hybrid image update.

8. The skin disease image classification method based on deep learning according to claim 1, characterized in that, The generation of the skin disease image classification model specifically includes: The pixel count of the lesion mask is counted within the lesion mask, and the pixel count of the replacement area is counted within the puzzle mask of the improved PuzzleMix. The area ratio of the main puzzle piece is obtained by dividing the difference between the pixel count of the lesion mask and the pixel count of the replacement area by the pixel count of the lesion mask. The area ratio of the paired puzzle piece is obtained by dividing the pixel count of the replacement area by the pixel count of the lesion mask. The main image category label is encoded as a main label vector, the paired image category label is encoded as a paired label vector, and the mixed label is obtained by adding the result of multiplying the area ratio of the main jigsaw puzzle piece by the main label vector element by element and the result of multiplying the area ratio of the paired jigsaw puzzle piece by the paired label vector element by element. The classifier calculates the class probability vector for the mixed image. The class probability vector and the mixed label are used to calculate the classification loss. The classification loss is used to update the classifier parameters. The domain probability vector is calculated by the domain discriminator on the subject image. The domain probability vector and the domain label acquired from the subject image are used to calculate the domain loss. The domain evidence map is averaged within the lesion mask to obtain the domain weight. The domain loss is multiplied by the domain weight to obtain the weighted domain loss. The domain discriminator parameters are updated based on the weighted domain loss. The weighted domain loss is then passed through the gradient inversion layer and multiplied by the domain adversarial coefficient before being used to update the classifier parameters.