Intelligent melanoma screening method based on unsupervised learning
Through unsupervised learning and deep neural network technology, the blood routine data of melanoma patients are automatically marked and screened, solving the problems of diagnosis and time-consuming labeling caused by the lack of standards in the screening process, characteristic similarity in the existing technology, and achieving efficient and accurate melanoma screening.
Patent Information
- Application Number
- CN202510270204.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
The existing melanoma screening technology lacks unified standards, and melanoma images are similar to other benign skin lesions, which leads to confusion in automatic diagnosis and image labeling.
An intelligent screening method based on unsupervised learning is adopted to classify blood conventional data through unsupervised K-means clustering algorithm, and a binary classification model is constructed using deep neural networks to automatically identify the inherent laws of blood conventional data for labeling, realizing early screening of melanoma.
It reduces the cost of melanoma screening, improves the efficiency and accuracy of screening, solves the problem of time-consuming and labor-consuming labeling data, and realizes an automated screening process.
Smart Images

Figure CN120199461A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of melanoma screening, and particularly to an intelligent melanoma screening method based on unsupervised learning. Background Art
[0002] Melanoma is a highly malignant tumor derived from melanocytes. The occurrence of melanoma is mostly due to normal melanocytes being stimulated by external factors (such as ultraviolet radiation, trauma, etc.) or internal factors (nevus cell malignancy, racial factors, genetic factors, etc.), presenting an irregular pigment halo on the skin surface, mucous membrane, or internal organs of the body. This kind of tumor mostly occurs in the skin, and can also occur in the mucous membrane (including internal organ mucous membrane), uvea, pia mater, etc.
[0003] Currently, the methods for treating melanoma include surgical resection, radiotherapy, immunotherapy, targeted therapy, and chemotherapy, etc. For early-stage and localized melanoma, surgical resection is the preferred treatment method. Radiotherapy, immunotherapy, and targeted therapy, etc. can be used as adjuvant treatment means to reduce the risk of recurrence or control the progression of the disease. For advanced or metastatic melanoma, chemotherapy may be a treatment option, but its efficacy is usually limited. Malignant melanoma, due to its high invasiveness and rapid metastasis potential, indeed poses a major challenge in the field of dermatology. Therefore, early screening, accurate diagnosis, and effective treatment are the key strategies to address this threat.
[0004] Most of the existing related technologies are based on artificial intelligence algorithms for disease prediction and recognition in the image field, but this method has the following disadvantages when applied to the medical industry:
[0005] First, there is a lack of unified standards and specifications in the processes of collecting, processing, and diagnosing melanoma images;
[0006] Second, melanoma has a high similarity in image features with other benign skin lesions, making it easy to cause confusion in the automatic diagnosis process;
[0007] Third, the annotation and collection of melanoma images is a time-consuming and laborious task, requiring professional medical knowledge and experience. Summary of the Invention
[0008] The purpose of the present invention is to provide an intelligent melanoma screening method based on unsupervised learning. Aiming at the non-linear relationship between blood routine data and melanoma disease, an intelligent melanoma screening method based on unsupervised learning is provided to find valuable information from the blood, so as to provide a direction for doctors to make a preliminary judgment on melanoma.
[0009] To achieve the above purpose, the present invention provides an intelligent melanoma screening method based on unsupervised learning, including the following steps:
[0010] Step S1: Obtain the blood routine test data, age, and gender of melanoma patients and the physical examination population, and construct an initial data set;
[0011] Step S2: Use the unsupervised K-means clustering algorithm to classify the initial data set to obtain a processed data set;
[0012] Step S3: Based on the k-fold cross-validation method, divide the processed data set into k parts, and divide the test set and the training set;
[0013] Step S4: Input the data of the training set into the deep neural network classifier to complete the construction of the binary classification model, test the classification effect through the data of the test set, and count the corresponding evaluation indicators.
[0014] Preferably, in Step S1, the blood routine test data includes mean corpuscular volume, platelet distribution width, white blood cell count, neutrophil ratio, lymphocyte ratio, eosinophil ratio, basophil ratio, neutrophil count, lymphocyte count, basophil count, mean corpuscular hemoglobin, mean corpuscular hemoglobin concentration, platelet, mean platelet volume, plateletcrit, monocyte count, monocyte ratio, eosinophil count.
[0015] Preferably, in Step S2, using the unsupervised K-means clustering algorithm to classify the initial data set to obtain a processed data set, the specific operation is as follows:
[0016] Use the unsupervised K-means clustering algorithm to train the initial data set to construct a model that can automatically divide into two categories;
[0017] Use the model that can automatically divide into two categories to predict the initial data set to generate two categories, and then match it with the initial data set to generate a new data set, which is the processed data set.
[0018] Preferably, in Step S2, the parameters of the unsupervised K-means clustering algorithm are as follows:
[0019] n_clusters is 2, the initialization center method is k-means++, n_init is 10, max_iter is 300, tol is float, verbose is 0, copy_x is bool.
[0020] Preferably, in step S4, the deep neural network classifier uses the Sequential function, where layers are set to 2 - 5 layers, units are set to 32 - 512, the step size is 32, activation is set to swish, dropout is set to 0.25, min_samples_leaf is set to 2, learning_rate is set to 0.0001 - 0.01, the maximum number of experiments is set to 10, the number of training rounds is set to 50, the loss function is set to binary_crossentropy, class_weight is set to balanced, the activation function is relu, and the optimizer is adam.
[0021] Preferably, in step S4, the evaluation metrics include sensitivity, specificity, and accuracy, and the calculation formulas are as follows:
[0022]
[0023] Among them, TP represents the number of people who are truly ill and are identified as ill; FN represents the number of people who are truly ill but are identified as normal; TN represents the number of people who are truly normal and are identified as normal; FP represents the number of people who are normal but are identified as ill; TPR represents sensitivity; TNR represents specificity; ACC represents accuracy.
[0024] Therefore, the present invention adopts the above - mentioned intelligent melanoma screening method based on unsupervised learning, and the beneficial technical effects are as follows:
[0025] (1) The acquisition of blood routine data is easy, highly universal, and simple to classify, saving the cost of melanoma screening;
[0026] (2) To solve the problem of time - consuming and laborious labeled data, the present invention uses an unsupervised learning algorithm to automatically label samples according to the internal laws of the data, and uses machine learning techniques to automatically identify the potential internal laws of the blood routine data to label the data, so as to conduct early screening for melanoma. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a flowchart of an intelligent melanoma screening method based on unsupervised learning according to an embodiment of the present invention;
[0028] Figure 2 is a flowchart of Comparative Example 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.
[0030] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention pertains.
[0031] Example 1
[0032] As Figure 1 shown, a melanoma intelligent screening method based on unsupervised learning of the present invention includes the following steps:
[0033] Step S1, obtain the blood routine test data of melanoma patients and physical examination populations, and construct an initial data set;
[0034] Select the clinical data of 5623 melanoma patients and 3626 physical examination populations in a certain hospital. The clinical data information includes demographic information, pathological information, and blood routine test data. The demographic information includes gender and age, and the pathological information includes diagnosis opinions, specimen types, and clinical diagnoses. The blood routine test data selects the common indicators of the two data sets, specifically MCV (mean corpuscular volume), PDW (platelet distribution width), WBC (white blood cell count), NEU% (neutrophil ratio), LYM% (lymphocyte ratio), EO% (eosinophil ratio), BASO% (basophil ratio), NEU# (neutrophil count), LYM# (lymphocyte count), BASO# (basophil count), MCH (mean corpuscular hemoglobin), MCHC (mean corpuscular hemoglobin concentration), PLT (platelet), MPV (mean platelet volume), PCT (plateletcrit), MONO# (monocyte count), MONO% (monocyte ratio), EO# (eosinophil count). Add the age and gender dimensions to the blood routine test data to construct the initial data set.
[0035] Step S2, perform category division on the initial data set through the unsupervised K-means clustering algorithm to obtain the processed data set;
[0036] Use the unsupervised K-means clustering algorithm to train the initial data set to construct a model that can automatically divide into two categories;
[0037] Use the model that can automatically divide into two categories to predict the initial data set to generate two categories, and then match them with the initial data set to generate a new data set, which is the processed data set.
[0038] The parameters of the unsupervised K-means clustering algorithm are as follows:
[0039] The number of clusters is 2, the initialization method of the centroids is k-means++, the number of initializations is 10, the maximum number of iterations is 300, the tolerance is of type float, the verbosity level is 0, and copy_x is of type bool.
[0040] Here, the initialization method of the centroids is k-means++. In the traditional K-means algorithm, the initial centroids are randomly selected, while K-means++ selects the initial centroids through a probabilistic method to reduce the risk of falling into a poor local optimum.
[0041] The working principle of the unsupervised K-means clustering algorithm: The data is divided into K clusters through an iterative optimization method. Each data point is assigned to the nearest cluster center, and the positions of the cluster centers are continuously updated until the result converges. The specific detailed steps are as follows:
[0042] 1) Use k-means++ to select the initial centroids through a probabilistic method. The specific steps are as follows:
[0043] 1.1 Randomly select a point from the dataset as the first cluster center c1;
[0044] 1.2 For each data point x, calculate its nearest distance to the selected cluster centers: where C is the set of currently selected cluster centers; D(x) is the nearest distance value; x is the data point; c i is the i-th cluster center;
[0045] 1.3 Select a new cluster center c i with a probability proportional to the square of the distance of each data point. Specifically, the probability P(c i =x) of selecting the data point x as the next cluster center is:
[0046]
[0047] where x' is the new data point; X is the set of all data points.
[0048] 1.4 Until K cluster centers are selected, forming the initial set of cluster centers C = c1, c2,....., c k , where C is the initial set of cluster centers.
[0049] 2) Calculate the distance of each data point to each cluster center and assign it to the nearest cluster:
[0050]
[0051] The above formula is the objective function formula, which is used to calculate the sum of squared errors within clusters, that is, the sum of the squares of the distances from all data points to their corresponding cluster centers. Among them, J is the value of the sum of squared errors within clusters; c k is the cluster center of the k-th cluster; X k are all data points in the k-th cluster; K is the dimension of the cluster center set.
[0052]
[0053] The above formula is the distance metric formula, which is used to calculate the distance from a data point to the cluster center. Among them, x ij is the j-th feature of the data point x i ; c kj is the j-th feature of the cluster center c k ; n is the data dimension.
[0054] 3) For each cluster, recalculate the mean of all data points within the cluster as the new cluster center:
[0055]
[0056] Among them, c k is the new cluster center.
[0057] 4) Repeat steps 2) and 3) until the cluster centers no longer change or reach the preset number of iterations.
[0058] Step S3: Based on the k-fold cross-validation method, divide the processed data set into k parts, and divide the test set and the training set;
[0059] In the k-fold cross-validation method, k is an arbitrary constant greater than 1. Common values of k are 5 or 10.
[0060] Step S4: Input the data of the training set into the deep neural network classifier to complete the construction of the binary classification model, evaluate the model through the data of the test set, and count the corresponding evaluation indicators.
[0061] The deep neural network classifier uses the Sequential function. Among them, the number of layers is set to 2 - 5, the number of units is set to 32 - 512, the step size is 32, the activation function is set to swish, the dropout rate is set to 0.25, the minimum number of samples in a leaf node is set to 2, the learning rate is set to 0.0001 - 0.01, the maximum number of experiments is set to 10, the number of training epochs is set to 50, the loss function is set to binary_crossentropy, the class weight is set to balanced, the activation function is relu or Sigmoid, and the optimizer is adam.
[0062] Here, the Sequential function consists of three parts: an input layer, a fully connected layer, and an output layer.
[0063] The input layer converts the multi-dimensional input into a set of arrays.
[0064] The fully connected layer is used to calculate the linear combination between the input and the weight matrix, add the bias term, and finally pass through the activation function. Here, the linear combination is a weighted sum. For the i-th neuron in the l-th layer, its input is the output h of the previous layer l-1 , the weight matrix is W l , the bias vector is b l , then the weighted sum output z l,i of this neuron is:
[0065] z l,i = W l h l-1 + b l ;
[0066] The activation function f is used for the output of the weighted sum, introducing non-linearity:
[0067] h l = f(z l );
[0068] Among them, z l is the value after the weighted sum of the neurons in the l-th layer; h l is the output value of the neurons in the l-th layer after being processed by the activation function;
[0069] Here, the activation function relu is used:
[0070] f(z l ) = max(0, z l );
[0071] The output layer has only one neuron and uses the Sigmoid activation function to map probability values. The weighted sum of the output layer is:
[0072] z = W L h L + b L ;
[0073] Among them, h L is the output of the last fully connected layer; W L is the weight of the output layer; b L is the bias vector of the output layer.
[0074] The loss function quantifies the difference between the model's predicted results and the actual labels during the training process. In binary classification, the formula for using binary_crossentropy is as follows:
[0075]
[0076] Among them, y is the true label, taking values of 0 or 1; is the probability value predicted by the model, with a value range within [0, 1]; is the loss of a single sample;
[0077] The evaluation metrics include sensitivity, specificity, and accuracy, and their calculation formulas are as follows respectively:
[0078]
[0079] Among them, TP represents the number of people who are truly ill and are identified as ill; FN represents the number of people who are truly ill but are identified as normal; TN represents the number of people who are truly normal and are identified as normal; FP represents the number of people who are normal but are identified as ill; TPR represents sensitivity; TNR represents specificity; ACC represents accuracy.
[0080] Comparative Example 1
[0081] As Figure 2 shown, extract the blood routine test data of melanoma patients and physical examination populations, and add the corresponding label columns to construct a dataset; mark melanoma patients as 1 and non-melanoma patients as 0.
[0082] Based on the k-fold cross-validation method, divide the dataset into k parts to obtain a test set and a training set, where k is an arbitrary constant greater than 1;
[0083] Select a deep neural network machine learning classifier model to train the dataset to construct a binary classification model, and count the corresponding evaluation metrics.
[0084] As shown in Table 1, by comparing the evaluation metrics of the two models in Example 1 and Comparative Example 1, it is found that the performance of the new binary classification model constructed using the dataset with relabeled labels by the unsupervised K-means clustering algorithm is better.
[0085] Table 1 Comparison of evaluation metrics of two models
[0086] model TPR TNR ACC Comparative Example 1 0.949 0.957 0.952 Example 1 0.976 0.986 0.982
[0087] Therefore, the present invention adopts the above-mentioned melanoma intelligent screening method based on unsupervised learning, uses machine learning technology to automatically identify the potential internal laws of blood routine to label the data, and thus conducts early screening for melanoma.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A melanoma intelligent screening method based on unsupervised learning, characterized in that: The following steps are involved: Step S1, obtaining blood routine test data, age, and gender of melanoma patients and physical examination population, and constructing an initial data set; Step S2, classifying the initial data set by using an unsupervised K-means clustering algorithm to obtain a processed data set; Step S3: Based on the k-fold cross-validation method, the processed data set is divided into k parts, and the test set and the training set are divided; Step S4: Input the data of the training set into the deep neural network classifier to complete the construction of the binary classification model, evaluate the model performance through the data of the test set, and calculate the corresponding evaluation indicators.
2. The method for intelligent screening of melanoma based on unsupervised learning according to claim 1, characterized in that: In step S1, the routine blood test data include the mean red blood cell volume, platelet distribution width, white blood cell count, neutrophil ratio, lymphocyte ratio, eosinophil ratio, basophil ratio, neutrophil count, lymphocyte count, basophil count, mean hemoglobin amount, mean hemoglobin concentration, platelets, mean platelet volume, platelet volume, monocyte count, monocyte ratio, and eosinophil count.
3. The method for intelligent screening of melanoma based on unsupervised learning according to claim 2, characterized in that: In step S2, the initial data set is classified by an unsupervised K-means clustering algorithm to obtain a processed data set. The specific operations are as follows: The unsupervised K-means clustering algorithm is used to train the initial data set to build a model that can automatically divide the two categories; The model that can automatically divide the two categories is used to predict the initial data set to generate two categories, and then it is matched with the initial data set to generate a new data set, which is the processed data set.
4. The method for intelligent screening of melanoma based on unsupervised learning according to claim 3, characterized in that: In step S2, the parameters of the unsupervised K-means clustering algorithm are as follows: n_clusters is 2, the initialization center method is k-means++, n_init is 10, max_iter is 300, tol is float, verbose is 0, and copy_x is bool.
5. The method for intelligent screening of melanoma based on unsupervised learning according to claim 4, characterized in that: In step S4, the deep neural network classifier uses the Sequential function, where layers is set to 2 to 5 layers, unit is set to 32 to 512, step size is 32, activation is set to swish, dropout is set to 0.25, min_samples_leaf is set to 2, learning_rate is set to 0.0001 to 0.01, the maximum number of experiments is set to 10, the number of training rounds is set to 50, the loss function is set to binary_crossentropy, class_weight is set to balanced, the activation function is relu, and the optimizer is adam.
6. The method for intelligent screening of melanoma based on unsupervised learning according to claim 5, characterized in that: In step S4, the evaluation indicators include sensitivity, specificity and accuracy, and the calculation formulas are as follows: Among them, TP represents the number of people who are actually sick and identified as sick; FN represents the number of people who are actually sick but identified as normal; TN represents the number of people who are actually normal and identified as normal; FP represents the number of people who are normal but identified as sick; TPR represents sensitivity; TNR represents specificity; ACC represents accuracy.
Citation Information
Patent Citations
Melanoma detection method based on attention mechanism and transfer learning
CN112734709A
Skin basal cell carcinoma fine-grained typing method based on clustering analysis algorithm
CN118379730A