Skin image data processing and classifying method based on deep learning
By constructing multimodal skin image data processing and classification methods of multi-branch skin lesions classification module and long-tail distribution processing module, the problems of time-consuming, subjective impact and poor long-tail data processing performance in the prior art are solved, and high-accurate BCC diagnosis and histopathological subtype prediction are achieved.
Patent Information
- Application Number
- CN202510107832.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
AI Technical Summary
The existing diagnostic methods for skin lesions rely on doctors’ experience and visual examinations, which have problems with time-consuming and subjective influence. The single-modal-based machine learning method has poor performance when processing long-tail data distribution, making it difficult to effectively distinguish the histopathological subtypes of BCC.
Using multimodal skin image data processing and classification method based on deep learning, we use multi-branch skin lesion classification module and long-tail distribution processing module to integrate clinical and dermatoscope image data, and use fine-tuned multimodal neural networks and large pre-trained models to solve the challenge of long-tail distribution.
Improves the accuracy of accurate diagnosis and histopathological subtype prediction of BCC, solves the inherent challenges of long-tail distribution, and provides clinicians with tools for early detection and personalized management, improving patient prognosis.
Smart Images

Figure CN120047401A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing, and particularly to a method for processing and classifying skin image data based on deep learning. Background Art
[0002] Basal Cell Carcinoma (BCC) is the most common skin malignancy globally. Early detection and appropriate intervention are crucial for avoiding severe local destruction and improving prognosis. Existing methods for diagnosing skin lesions mainly rely on the experience and visual inspection of dermatologists, which are not only time-consuming but also vulnerable to subjective factors. In recent years, although some machine learning-based methods have been proposed for assisting diagnosis, such as applications using artificial intelligence (AI) technologies like Convolutional Neural Networks (CNN) for BCC diagnosis, these methods often only consider single-modal information and perform poorly when dealing with long-tail data distributions, such as Figure 11 shown, under the challenge. Therefore, there is an urgent need to propose a multi-modal method that can effectively distinguish various BCC histopathological subtypes. Summary of the Invention
[0003] To solve the above problems in the prior art, embodiments of the present invention provide a method for processing and classifying skin image data based on deep learning in view of the deficiencies in the prior art.
[0004] The first aspect of the present invention provides a method for processing and classifying skin image data based on deep learning, including:
[0005] Step S1: Obtain skin image data;
[0006] Step S2: Preprocess the skin image data to obtain first skin image data;
[0007] Step S3: Input the first skin image data into a skin image data classification processing model for classification processing;
[0008] Step S4: Obtain skin image classification data according to the output of the skin image data classification processing model;
[0009] Step S5: Output the skin image classification data.
[0010] Preferably, the skin image data classification processing model includes a multi-branch skin lesion classification module and a long-tail distribution processing module.
[0011] Further preferably, the steps of constructing the skin image data classification processing model include:
[0012] Steps for constructing a dataset, including data collection, data annotation, and dataset division; among them, the dataset includes a four-class set and a six-class set;
[0013] Steps for the model architecture, including adjusting attention weights through an information complementary block to obtain a multi-branch skin lesion classification module, and adjusting parameters based on a pre-trained model of CL IP to obtain a long-tail distribution processing module;
[0014] Steps for model training and evaluation, including the training process and the performance evaluation process.
[0015] Further preferably, the multi-branch skin lesion classification module includes: a dermoscopy branch, a clinical branch, a fusion branch, and an output layer.
[0016] Further preferably, the fusion branch includes a first information complementary block, a convolutional layer, and a second information complementary block, where,
[0017] The first information complementary block receives the feature data D1 from the dermoscopy branch and the feature data C1 from the clinical branch to process and obtain the first fusion feature data;
[0018] The convolutional layer performs convolutional processing on the first fusion feature data to obtain convolutional feature data;
[0019] The second information complementary block receives the convolutional feature data, the feature data D2 from the dermoscopy branch, and the feature data C2 from the clinical branch to perform classification processing and obtain skin lesion image classification data.
[0020] Further preferably, the skin lesion image classification data is skin fibroma image data, melanocytic nevus image data, seborrheic keratosis image data, or basal cell carcinoma image data.
[0021] Preferably, inputting the first skin image data into the skin image data classification processing model for classification processing specifically includes: using the first skin image data as the input of the multi-branch skin lesion classification module for lesion recognition and classification processing;
[0022] According to the output of the multi-branch skin lesion classification module, obtain skin lesion image classification data;
[0023] When it is determined that the skin lesion image classification data includes basal cell carcinoma image data:
[0024] Use the basal cell carcinoma image data as the input of the long-tail distribution processing module for long-tail recognition and classification processing;
[0025] According to the output of the long-tail distribution processing module, obtain long-tail distribution image data;
[0026] Based on the long-tailed distribution image data and skin lesion image classification data, obtain the classification result data to be output;
[0027] When it is determined that the skin lesion image classification data does not include basal cell carcinoma image data:
[0028] Obtain the classification result data to be output according to the skin lesion image classification data.
[0029] Further preferably, the long-tailed distribution image data includes one or more of nodular basal cell carcinoma, superficial basal cell carcinoma, a mixed form of subtypes, micronodular basal cell carcinoma, basal squamous cell carcinoma, and invasive basal cell carcinoma.
[0030] Preferably, the skin classification data includes skin image data and classification result description data, and outputting the skin image classification data specifically includes displaying the skin image classification data on a display device according to a preset output method.
[0031] A second aspect of the present invention provides an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the method for processing and classifying skin image data based on deep learning provided in the first aspect above.
[0032] A third aspect of the present invention provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method for processing and classifying skin image data based on deep learning provided in the first aspect above is implemented.
[0033] The exact beneficial effects of the method for processing and classifying skin image data based on deep learning provided by the present invention are as follows:
[0034] 1. A skin image data classification processing model is constructed by using a fine-tuned multi-modal neural network, effectively integrating clinical and dermoscopic images to improve the accurate diagnosis of BCC and the prediction of histopathological subtypes;
[0035] 2. By using a large pre-trained model and complex fusion techniques, not only the inherent challenges of the long-tailed distribution are solved, but also a powerful tool is provided for clinicians for the early detection and personalized management of BCC, thereby improving the prognosis of patients.
[0036] The method for processing and classifying skin image data based on deep learning provided by the present invention analyzes and judges the acquired skin image data through a pre-trained skin image data classification processing model, and directly outputs the skin image classification data. It has a fast speed for classifying skin image data, can accurately distinguish BCC from common benign skin tumors, and is a fast image data processing and classification method for predicting the histopathological subtypes of BCC. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flowchart of a method for constructing a skin image data classification processing model according to the present invention;
[0038] Figure 2 It is an architecture diagram of a multi-branch skin lesion classification module according to the present invention;
[0039] Figure 3 It is an architecture diagram of a long-tail distribution processing module according to the present invention;
[0040] Figure 4 It is a result diagram for evaluating the skin image data classification processing model of the present invention using sensitivity and specificity;
[0041] Figure 5 It is an ROC curve diagram of the skin image data classification processing model according to the present invention;
[0042] Figure 6 It is a confusion matrix diagram of the skin image data classification processing model according to the present invention;
[0043] Figure 7 It is a significance diagram of the skin image data classification processing model according to the present invention;
[0044] Figure 8 It is a t-SNE diagram of two classification tasks of the skin image data classification processing model according to the present invention;
[0045] Figure 9 It is a flowchart of a method for processing and classifying skin image data based on deep learning according to the present invention;
[0046] Figure 10 It is a structural block diagram of a computer device according to the present invention;
[0047] Figure 11 It is a long-tail distribution diagram of a six-classification data set. Detailed implementation manners
[0048] The following further elaborates on the present application in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are merely for explaining the relevant invention and not for limiting the invention. Additionally, it should be noted that for ease of description, only parts related to the relevant invention are shown in the accompanying drawings.
[0049] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The following will elaborate on the present application in detail with reference to the accompanying drawings and embodiments.
[0050] To more clearly describe a method for processing and classifying skin image data based on deep learning provided by the present invention, here, first, a method for training a skin image data classification processing model of the present invention will be introduced.
[0051] Figure 1 As shown in the flowchart of the method for training a skin image data classification processing model according to the present invention, the method for obtaining the skin image data classification processing model of the present invention includes a data set construction step (S10), a model architecture step (S20), and a model training and evaluation step (S30). Each step will be described in detail below with reference to the accompanying drawings.
[0052] S10: Data set construction step, including data collection, data annotation, and data set division.
[0053] Specifically, when constructing and training the skin image data classification processing model of the present invention, first, a data set needs to be constructed. Further constructing the data set includes:
[0054] Data collection: In the present invention, the data set includes clinical skin image data and dermoscopic image data marked as BCC, MN, SK, and DF in a specific skin image database during a specific time period. Among them, BCC, MN, SK, and DF correspond to Basal Cell Carcinoma (BCC), Melanocytic Nevus (MN), Seborrheic Keratosis (SK), and Dermatofibroma (DF), respectively. In the present invention, the specific time period is from June 2018 to January 2024; the specific skin image database is the skin image database of Peking Union Medical College Hospital (PUMCH). These images were captured by experienced technicians using a digital dermoscope system (MoleMax HD 1.0 dermoscope, Digital Image Systems, Vienna, Austria), and have high image standardization characteristics, which are particularly suitable for training the skin image data classification processing model provided by the present invention.
[0055] Data annotation: In the data annotation stage of the present invention, manual annotation is adopted. Specifically, two dermatologists with at least 5 years of clinical experience are invited to independently screen and annotate the data according to the patient's medical history, clinical manifestations, and dermoscopic features, including confirming that the classification of clinical skin image data and dermoscopic image data is one or more of BCC, MN, SK, and DF. After confirmation, add the labels BCC, MN, SK, and / or DF to the clinical skin image data and dermoscopic image data. In one solution of the present invention, if the clinical skin image data or dermoscopic image data shows multiple types of BCC, MN, SK, and DF, then add the determined multiple labels to the clinical skin image data or dermoscopic image data, rather than adding one label. In another solution of the present invention, if the clinical skin image data or dermoscopic image data shows multiple types of BCC, MN, SK, and DF, then after specific analysis by the annotator, add one of the determined labels to the clinical skin image data or dermoscopic image data.
[0056] In the embodiments of the present invention, for the clinical skin image data or dermoscopic image data corresponding to BCC cases, the corresponding histopathological subtypes are also annotated, including adding the labels NOD, SUP, MIX, MIC, BSQ, and INF, which respectively correspond to various types of BCC, including: Nodular BCC (NOD), Superficial BCC (SUP), Mixed forms Of BCC Subtypes (MIX), Micronodular BCC (MIC), Basosquamous Carcinoma (BSQ), and Infiltrative BCC (INF). To ensure the accuracy of the annotation of clinical skin image data and dermoscopic image data, the annotation results of all clinical skin image data or dermoscopic image data in the dataset are confirmed by doctors.
[0057] Dataset division: In order to enable the present invention to perform BCC detection classification and BCC tissue case subtype prediction on skin image data, a four-class set and a six-class set required by the invention need to be constructed. To complete the training of the skin image data classification processing model in the embodiments of the present invention, the dataset is also divided into a training set and a test set. In one training solution of the embodiments of the present invention, the number of samples in the dataset and the specific division are shown in Table 1 below:
[0058] Table 1: Division of the number of samples in the four-class set and six-class set in the dataset
[0059]
[0060] S20: Model architecture steps, including adjusting attention weights through an information complementary block and parameter adjustment based on a CLIP-based pre-trained model.
[0061] Specifically, in the embodiments of the present invention, after the construction of the dataset is completed, a skin image data classification processing model of the present invention is constructed by separately constructing a multi-branch skin lesion classification module and a long-tail distribution processing module. Next, in combination with Figure 2 and Figure 3 the constructed multi-branch skin lesion classification module and long-tail distribution processing module are introduced respectively.
[0062] [Multi-branch skin lesion classification module]
[0063] Adjust the attention weights through the information complementary block. Specifically, the ICB module receives the features from two modal encoders, processes them through a series of convolutional layers, batch normalization, and ReLU activation functions to learn the attention weights of the fused features, and obtains the multi-branch skin lesion classification module that constitutes the skin image data classification processing model of the present invention. The multi-branch skin lesion classification module of the present invention includes multiple stages and branches for processing different types of data. In one embodiment of the present invention, it is used to process dermoscopic image data and clinical skin image data.
[0064] In the embodiments of the present invention, as Figure 2 shown, the architecture of the multi-branch skin lesion classification module includes two branches: the dermoscopic branch, the clinical branch, the fusion branch, and the output layer. The dermoscopic branch and the clinical branch are respectively used to process dermoscopic image data and clinical skin image data. Both the dermoscopic branch and the clinical branch go through four stages (stage1, stage2, stage3, and stage4), and then are merged through the fusion branch. Finally, the outputs of the dermoscopic branch and the clinical branch are sent to the convolutional layer (conv) of the fusion branch for processing and output to the information complementary block (ICB) as a classifier, and output to four categories through the ICB: BCC, MN, SK, or DF. The fusion branch includes a first information complementary block, a convolutional layer, and a second information complementary block. Among them, the first information complementary block receives the feature data D1 from the dermoscopic branch and the feature data C1 from the clinical branch to process and obtain the first fused feature data; the convolutional layer performs convolutional processing on the first fused feature data to obtain convolutional feature data; the second information complementary block receives the convolutional feature data and the feature data D2 from the dermoscopic branch and the feature data C2 from the clinical branch to perform classification processing to obtain skin lesion image classification data.
[0065] AsFigure 2 As shown, the convolution (conv) therein is a common deep learning operation, which is used to extract features from images or other data. The convolutional layer is located after the fusion layer. It takes the fused data as input and outputs a higher-level feature representation, which can be further used for tasks such as classification and regression. Specifically, the convolutional layer scans the input data by using a set of learnable filters (or called convolutional kernels) to generate a new feature map. These filters can capture local patterns and structures in the input data, such as edges and textures. By stacking multiple convolutional layers, the model can learn increasingly complex feature representations, thus achieving in-depth understanding of the data.
[0066] As Figure 2 shown, the second ICB therein is used as the final classifier. It takes the data of the dermoscopy branch and the clinical branch and their fusion results as input and outputs the final classification result. Specifically, the first ICB extracts and processes the features of the data of the dermoscopy branch and the clinical branch respectively, and then inputs the results of these two branches into the convolutional layer of the fusion branch for fusion processing, such as using integration methods such as weighted average and voting for fusion to obtain the first fused feature data. The fused first fused feature data is sent to the convolutional layer for processing to generate a new feature representation, denoted as convolutional feature data. Finally, this convolutional feature data is sent to the second ICB as the classifier for classification to obtain the final skin lesion image classification data: one of BCC, MN, SK or DF. As Figure 2 shown, the information supplement block includes the following steps: First, the input data C passes through the convolutional layer (conv), batch normalization layer (batchnorm) and residual connection (relu). Then, the data passes through the pooling operation, linear layer (linear) and ReLU activation function. Next, the data is sent to the Sigmoid function for processing. Finally, the data passes through the scale operation and outputs.
[0067] The advantage of the multi-branch skin lesion classification module architecture design provided by the present invention is that it can make full use of the information of dermoscopy image data and clinical skin image data, and improve the accuracy and robustness of the model through multiple stages and the processing of the dermoscopy branch and the clinical branch. At the same time, the use of ICB also enables the model to better adapt to different types of data sources and task requirements, and has good generalization ability and adaptability.
[0068] [Long-tailed distribution processing module]
[0069] Embodiments of the present invention are based on the Contrastive Language-Image Pre-training (CLIP) model, and fine-tune the CLIP model, selectively optimizing a specified proportion of the model parameters. At the same time, the Logit-Adjusted Loss (LA loss) is used to adjust the Softmax output probability to solve the long-tail data distribution problem. When selecting the base model in the present invention, one or more CLIP models can be selected. In one embodiment of the present invention, a CLIP model is selected for the architecture skin image data classification processing model. In the embodiments of the present invention, for each weight matrix W∈R^(d_1×d_2) in the CLIP model, a specified proportion α of the parameters is selectively optimized, and the remaining parameters remain unchanged, so that only the parameters α, d1, and d2 are optimized. In the embodiments of the present invention, this selective optimization is achieved by applying a sparse 0-1 mask M∈{0,1)^(d1×d2), where d1 and d2 represent the dimensions of the weight matrix. This mask determines which parameters are optimized, so that only part of the parameters (specifically, the proportion is αd1d2) are updated. This method helps to reduce the demand for computing resources while maintaining the performance of the model. In a specific example of the embodiments of the present invention, α is preferably 20%.
[0070] In an embodiment provided by the present invention, the Information Complement Module (ICB) is used to receive the features from two modal encoders, and is processed through a series of convolutional layers, batch normalization, and ReLU activation functions to learn the attention weights of the fused features. Adopting this processing method helps the model to better understand and integrate information from different sources, and improve the accuracy of prediction or classification. It should be understood that by using the Information Complement Module to receive the features from two modal encoders, it indicates that the skin image data classification processing model of the present invention includes the Information Complement Module, and further includes a series of convolutional layers, batch normalization layers, and ReLU activation functions.
[0071] In an embodiment provided by the present invention, in order to compensate for class imbalance, before applying the softmax function, by merging the class priors (i.e., the proportion of each class in the dataset) into the Logits, the following LA loss is used for optimization:
[0072]
[0073] where y is the true class label, y^k is the predicted Logits for class k, πk is the prior probability of class k, τ is a scaling factor used to adjust the influence of the prior, and 1(·) is an indicator function that returns 1 when the internal condition is true and 0 otherwise.
[0074] In one embodiment provided by the present invention, as Figure 3 shown, the architecture of the long-tail distribution processing module constructed by the present invention includes: a text encoder: receiving BCC and outputting a text encoding result; an input embedding: performing embedding processing on the input data to obtain an embedding vector; an image encoder: including multiple layers, each layer consisting of an Add&Norm operation and a Feed Forward operation. The output of each layer will go through a Multi-Head Attention operation and then be passed to the next layer; an attention block: located on the far right of the image encoder, used to calculate the attention weights between different layers; an attention matrix: showing the attention weight distribution between different layers, and each row represents the influence degree of the output of one layer on all other layers.
[0075] In one embodiment provided by the present invention, as Figure 3 shown in the following figure of, the skin image data classification processing model of the present invention includes a text encoder, a dermoscopy branch, and a clinical branch, which summarize information into the image encoder through a fusion branch. The image encoder contains frozen parameters and adjustable parameters, which are connected to multiple transformation matrices, and finally, the loss is calculated through LALoss.
[0076] The above introduced the model architecture steps of the present invention. It can be seen that the skin image data classification processing model of the present invention includes an information complementary module and at least one fine-tuning module of a pre-trained model based on CLIP.
[0077] S30: Model training and evaluation steps, including a training process and a performance evaluation process.
[0078] Specifically, after the model architecture steps, the model is trained and evaluated.
[0079] In one embodiment provided by the present invention, during the training process, the four-class set and six-class set in the constructed dataset are used to train the model, and the LA loss is applied to improve the model's recognition ability for minority classes.
[0080] In one embodiment provided by the present invention, during the performance evaluation process, indicators such as sensitivity, specificity, accuracy, ROC curve, and confusion matrix are used to evaluate the performance of the skin image data classification processing model of the present invention.
[0081] Sensitivity and specificity evaluation, as Figure 4 shown, which shows the overall classification results of the sensitivity and specificity evaluation of the skin image data classification processing model of the present invention using the samples in the four-class set and six-class set in the dataset in Table 1. The corresponding classification result data is shown in Table 2 below, by Figure 4As can be seen from the results in Table 2, for the four-class set, the overall accuracy of the model reaches 0.97, the sensitivity is 0.97, and the specificity is 0.99. For the six-class set, the overall accuracy of the model is 0.90, the sensitivity is 0.88, and the specificity is 0.91. From the results, it can be known that the skin image data classification processing model of the present invention can accurately classify skin image data.
[0082] Table 2: Classification result data for sensitivity and specificity evaluation
[0083]
[0084] ROC curve evaluation, as Figure 5 shown, which shows the ROC curve display of evaluating the skin image data classification processing model of the present invention using the samples in the four-class set and six-class set in the dataset in Table 1. From Figure 5 the curve results in, it can be seen that the skin image data classification processing model of the present invention has high accuracy for both the four-class set and the six-class set.
[0085] Confusion matrix evaluation, as Figure 6 shown, which shows the confusion matrix result display of evaluating the skin image data classification processing model of the present invention using the samples in the four-class set and six-class set in the dataset in Table 1 respectively. Figure 6 The results in show that the four-class confusion matrix indicates the high accuracy of the skin image data classification processing model of the present invention for the BCC class and the MN class, and the six-class confusion matrix indicates the high accuracy of the skin image data classification processing model of the present invention for the MIX category.
[0086] Significance evaluation, as Figure 7 shown, which shows the significance result display of evaluating the skin image data classification processing model of the present invention using the samples in the four-class set and six-class set in the dataset in Table 1 respectively. Figure 7 The significance heatmap results in show that the skin image data classification processing model of the present invention can accurately focus on the lesion areas in clinical skin image data and dermoscopic image data, and distinguish different lesion categories based on the lesion areas.
[0087] t-SEN evaluation, as Figure 8 shown, which shows the T-SEN result display of evaluating the skin image data classification processing model of the present invention using the samples in the four-class set and six-class set in the dataset in Table 1 respectively. Figure 7 The 8t-SNE feature distribution results in show that the skin image data classification processing model of the present invention can significantly distinguish different category features.
[0088] In summary, the above evaluation results of sensitivity and specificity, ROC curve, confusion matrix, significance, and t-SEN indicate that the skin image data classification and processing model of the present invention exhibits excellent performance in both four-class and six-class classification tasks.
[0089] The construction of the skin image data classification and processing model of the present invention has been described in detail above. Next, in combination with the attached Figure 9 A method for processing and classifying skin image data based on deep learning provided by the present invention will be described in detail.
[0090] [Method for Processing and Classifying Skin Image Data Based on Deep Learning]
[0091] As Figure 9 shown, a method for processing and classifying skin image data based on deep learning provided by the present invention includes the following steps:
[0092] Step S1: Obtain skin image data.
[0093] Step S2: Preprocess the skin image data to obtain first skin image data.
[0094] Step S3: Input the first skin image data into the skin image data classification and processing model for classification processing.
[0095] Specifically, in an embodiment provided by the present invention, implementing this step specifically includes:
[0096] Step S301: Use the first skin image data as the input of the multi-branch skin lesion classification module for lesion recognition and classification processing.
[0097] Step S302: Obtain skin lesion image classification data according to the output of the multi-branch skin lesion classification module;
[0098] Step S303: Determine whether the skin lesion image classification data includes basal cell carcinoma image data. If the judgment result is "yes", execute steps S304 to S306. If the judgment result is "no", then execute step S307.
[0099] Step S304: Use the basal cell carcinoma image data as the input of the long-tail distribution processing module for long-tail recognition and classification processing;
[0100] Step S305: Obtain long-tail distribution image data according to the output of the long-tail distribution processing module;
[0101] Step S306: Obtain the classification result data to be output according to the long-tail distribution image data and the skin lesion image classification data;
[0102] Step S307: Obtain the classification result data to be output according to the skin lesion image classification data.
[0103] Step S4: Obtain skin image classification data according to the output of the skin image data classification processing model.
[0104] Step S5: Output the skin image classification data.
[0105] Through the above steps S1 to S5, the skin image data can be classified into BCC, MN, SK or DF for skin lesions. If the obtained skin lesion classification result is BCC, the BCC skin image data can be further processed for long-tail distribution classification to further obtain the classification of BCC histopathological subtypes NOD, SUP, MIX, MIC, BSQ or INF.
[0106] The embodiment of the present invention provides a method for processing and classifying skin image data based on deep learning. By constructing a skin image data classification processing model based on deep learning and using this model to classify the obtained skin image data, it is classified into BCC, MN, SK or DF for skin lesions. If it is of the BCC type, it can be further processed for long-tail distribution classification to confirm the classification of BCC histopathological subtypes NOD, SUP, MIX, MIC, BSQ or INF. This method has high classification accuracy and solves the technical defect of inability to perform long-tail classification in the prior art. Its exact beneficial effects are as follows:
[0107] 1. Construct a skin image data classification processing model by fine-tuning a multi-modal neural network, effectively integrating clinical and dermoscopic images to improve the accurate diagnosis of BCC and the prediction of histopathological subtypes.
[0108] 2. Utilize a large pre-trained model and complex fusion techniques to not only solve the inherent challenges of long-tail distribution but also provide a powerful tool for clinicians for early detection and personalized management of BCC, thereby improving patient prognosis.
[0109] The method for processing and classifying skin image data based on deep learning provided by the present invention analyzes and judges the obtained skin image data through a pre-trained skin image data classification processing model, and directly outputs the skin image classification data. It has a fast speed for classifying skin image data, can accurately distinguish BCC from common benign skin tumors, and is a fast image data processing and classification method for predicting the histopathological subtypes of BCC.
[0110] In some possible implementations, a method for processing and classifying skin image data based on deep learning provided by the present invention can be implemented by a processing component calling computer-readable instructions stored in a memory. In an example, the processing component includes, but is not limited to, a single processor, or discrete components, or a combination of a processor and discrete components. The processor may include a controller in an electronic device having an instruction execution function, and the processor may be implemented in any suitable manner. For example, it is implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. Inside the processor, executable instructions can be executed through hardware circuits such as logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers.
[0111] The first and second in the following text are only for distinction and have no other meaning.
[0112] Each step in the above method for processing and classifying skin image data based on deep learning of the present invention can be implemented in whole or in part by software, hardware, and their combination. The above steps can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above steps.
[0113] To solve the above technical problems, an embodiment of the present application also provides a computer device. For details, please refer to Figure 10 , Figure 10 which is the basic structural block diagram of the computer device in this embodiment.
[0114] The computer device 10 includes a memory 101, a processor 102, and a network interface 103 that are communicatively connected to each other via a system bus. It should be noted that only the computer device 10 with components connected to the memory 101, the processor 102, and the network interface 103 is shown in the figure. However, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0115] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touchpad, or a voice control device, etc.
[0116] The memory 101 includes at least one type of readable storage medium. The readable storage medium includes flash memory, hard disks, multimedia cards, card-type memories (such as SD or D interface display memories, etc.), random access memories (RAMs), static random access memories (SRAMs), read-only memories (ROMs), electrically erasable programmable read-only memories (EEPROMs), programmable read-only memories (PROMs), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the memory 101 can be an internal storage unit of the computer device 10, such as the hard disk or memory of the computer device 10. In other embodiments, the memory 101 can also be an external storage device of the computer device 10, such as a plug-in hard disk equipped on the computer device 10, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 101 can also include both the internal storage unit and the external storage device of the computer device 10. In this embodiment, the memory 101 is generally used to store the operating system installed on the computer device 10 and various application software, such as program codes for controlling electronic files, etc. In addition, the memory 101 can also be used to temporarily store various data that have been output or will be output.
[0117] In some embodiments, the processor 102 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 102 is generally used to control the overall operation of the computer device 10. In this embodiment, the processor 102 is used to run the program code stored in the memory 101 or process data, such as running the program code for controlling electronic files.
[0118] The network interface 103 may include a wireless network interface or a wired network interface. The network interface 103 is generally used to establish a communication connection between the computer device 10 and other electronic devices.
[0119] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing an interface display program, and the interface display program can be executed by at least one processor to enable the at least one processor to execute the steps of the control method of the electronic file as described above.
[0120] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods of the various embodiments of the present application.
[0121] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The accompanying drawings show the preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements for some of the technical features. Any equivalent structure made by using the specification and drawings of the present application, directly or indirectly applied to other related technical fields, is similarly within the scope of the patent protection of the present application.
[0122] In the above specific embodiments, the purpose, technical solution and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A skin image data processing and classification method based on deep learning, characterized in that: The method comprises: Step S1: Acquire skin image data; Step S2: preprocessing the skin image data to obtain first skin image data; Step S3: inputting the first skin image data into a skin image data classification processing model for classification processing; Step S4: obtaining skin image classification data according to the output of the skin image data classification processing model; Step S5: Output skin image classification data.
2. The skin image data processing and classification method according to claim 1, characterized in that: The skin image data classification processing model includes a multi-branch skin lesion classification module and a long-tail distribution processing module.
3. The skin image data processing and classification method according to claim 2, characterized in that: The steps of constructing the skin image data classification processing model include: The data set construction step includes data collection, data annotation and data set division; wherein the data set includes a four-category set and a six-category set; The model architecture step includes adjusting the attention weight through the information complementation block to obtain the multi-branch skin lesion classification module, and adjusting the parameters based on the CLIP pre-training model to obtain the long-tail distribution processing module; Model training and evaluation steps, including training process and performance evaluation process.
4. The skin image data processing and classification method according to claim 3, characterized in that: The multi-branch skin lesion classification module includes: a dermatoscope branch, a clinical branch, a fusion branch and an output layer.
5. The skin image data processing and classification method according to claim 3, characterized in that: The fusion branch includes a first information complementary block, a convolutional layer, and a second information complementary block, wherein: The first information complementation block receives the feature data D1 from the dermoscopic branch and the feature data C1 from the clinical branch to process and obtain first fused feature data; The convolution layer performs convolution processing on the first fused feature data to obtain convolution feature data; The second information complementation block receives the convolution feature data and the feature data D2 from the dermatoscope branch and the feature data C2 from the clinical branch to perform classification processing to obtain the skin lesion image classification data.
6. The skin image data processing and classification method according to claim 5, characterized in that: The skin lesion image classification data is cutaneous fibroma image data, melanocytic nevus image data, seborrheic keratosis image data or basal cell carcinoma image data.
7. The skin image data processing and classification method according to claim 1, characterized in that: The step of inputting the first skin image data into the skin image data classification processing model for classification processing specifically includes: using the first skin image data as input of the multi-branch skin lesion classification module for lesion recognition and classification processing; Obtaining skin lesion image classification data according to the output of the multi-branch skin lesion classification module; When it is determined that the skin lesion image classification data includes basal cell carcinoma image data: Using the basal cell carcinoma image data as input to the long-tail distribution processing module for long-tail identification and classification processing; Obtaining long-tail distribution image data according to the output of the long-tail distribution processing module; Obtaining classification result data to be output according to the long-tail distribution image data and the skin lesion image classification data; When it is determined that the skin lesion image classification data does not include basal cell carcinoma image data: The classification result data to be output is obtained according to the skin lesion image classification data.
8. The skin image data processing and classification method according to claim 7, characterized in that: The long-tail distribution image data includes one or more of nodular basal cell carcinoma, superficial basal cell carcinoma, a mixed form of subtypes, micronodular basal cell carcinoma, basosquamous cell carcinoma, and invasive basal cell carcinoma.
9. The skin image data processing and classification method according to claim 1, characterized in that: The skin classification data includes skin image data and classification result description data, and the outputting of the skin image classification data specifically includes displaying the skin image classification data through a display device according to a preset output mode.
10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for processing multi-volume sequence image data according to any one of claims 1 to 9 is implemented.