A method for classifying skin lesions
A method using predetermined scoring and machine learning models classifies skin lesions efficiently across various cancer types by combining key features, enhancing sensitivity and accuracy in identifying suspicious lesions.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2026-03-12
AI Technical Summary
Existing skin cancer risk scoring methodologies are limited to melanomas and require a large number of clinical risk factors, increasing data processing complexity and inefficiency in classifying various skin cancer subtypes.
A method for classifying skin lesions as suspicious or non-suspicious using a combination of predetermined scoring tables and machine learning models, incorporating features such as lesion color, size, inflammation, and age, to provide a lesion risk score that can distinguish between different skin cancer types with reduced data processing.
The method effectively classifies skin lesions beyond melanomas, achieving high sensitivity, specificity, and accuracy in identifying suspicious lesions, including basal cell and squamous cell carcinomas, while minimizing data processing requirements.
Smart Images

Figure EP2025075361_12032026_PF_FP_ABST
Abstract
Description
[0001] A METHOD FOR CLASSIFYING SKIN LESIONS
[0002] The present disclosure relates to methods for classifying skin lesions as suspicious or non- suspicious with regards to skin cancer risk. Related, methods systems and products are also described.
[0003] BACKGROUND
[0004] Existing skin cancer risk scoring methodologies are available based on using seven clinical risk features, but these are only able to provide risk scores for melanomas. The 7-point checklist (7PCL) has been extensively used to assess patients with skin lesions and make referrals forthose with a suspected melanoma based on a scoring method. The 7 PCL method is however only suitable for assessing the skin lesions of patients with suspected melanomas and cannot be used as an assessment for other skin cancer sub types. Another standard approach is the Williams risk score which is used to calculate the lifetime risk of melanomas in patients without any symptoms. The Williams approach also involves making calculations based on 7 meta features to provide a risk score but as with the 7 PCL, it is only suitable for providing a determination for melanomas and not a broader set of skin cancer sub types.
[0005] While melanomas have relatively similar appearances and physical characteristics this cannot be said for all skin cancers and so combining these into a single cancer risk scoring methodology relies on a large increase in the number of different types of clinical risk factors that are analysed. At present 18 different meta-features corresponding to clinical risk factors are required to correctly classify all the skin cancer types as suspicious or non-suspicious with sufficient accuracy. Having to conduct calculations involving 18 different meta-features, compared to 7 in the prior art methods for melanomas, massively increases the volume of data that needs to be collected and processed in the calculations to achieve the outcome of a determination of whether a skin lesion is suspicious or not.
[0006] There therefore exists a need for a cancer risk scoring methodology that covers all skin cancer sub-types as there isn’t one currently available. There also exists a need to overcome the problems in the prior art by providing a methodology which is able to distinguish between all skin cancer sub types while keeping the amount of data processing in performing the calculations to a reasonable level.
[0007] SUMMARY OF THE DISCLOSURE
[0008] According to a first aspect, there is provided a method for classifying skin lesions as suspicious or non-suspicious comprising; collecting and scoring a first piece of data comprising a determination of whether a skin lesion is pink in colour or not pink to provide a lesion risk score to determine the likelihood of a skin lesion being suspicious or non-suspicious. Also described according to the present aspect is a computer-implemented method for classifying a skin lesion in a patient as suspicious or non-suspicious, the method comprising: obtaining data for one or more features associated with the lesion, the one or more features comprising a determination of whether a skin lesion is pink in colour or not pink; and classifying, using said data, the lesion between a first class of suspicious lesions and a second class of non-suspicious lesions.
[0009] In embodiments, classifying the lesion comprises obtaining a score for each of said features, and combining the scores to provide a lesion risk score. Obtaining a score for each of said features may comprise using, for each of said one or more features, a predetermined scoring table to score said feature. Classifying the lesion may comprise comparing the lesion risk score to a predetermined threshold, wherein a lesion associated with a lesion risk score equal or higher than the predetermined threshold is classified as being suspicious and a lesion with a lesion risk score lower than the predetermined threshold is classified as being non-suspicious. The predetermined threshold may have been identified using training data comprising data for said features for a plurality of training lesions and associated labels indicating whether each training lesion is suspicious or non-suspicious. The predetermined threshold may be a threshold that optimises one or more of: the sensitivity of identification of suspicious lesions, the specificity of identification of suspicious lesions, the accuracy or balanced accuracy of identification of suspicious and non- suspicious lesions. The predetermined threshold may be a threshold that results in a maximum value of an accuracy metric in a training cohort comprising data for said features for a plurality of training lesions and associated labels indicating whether each training lesion is suspicious or non- suspicious. The predetermined threshold may be an integer threshold that results in a maximum value (over possible integer thresholds) of an accuracy metric in a training cohort comprising data for said features for a plurality of training lesions and associated labels indicating whether each training lesion is suspicious or non-suspicious.
[0010] Classifying the lesion may comprise encoding each of said features that is not numerical into a numerical format, thereby obtaining a set of numerical features, and providing said numerical features as input to a machine learning model trained to classify lesions between a first class of suspicious lesions and a second class of non-suspicious lesions using training data comprising data for said features for a plurality of training lesions and associated labels indicating whether each training lesion is suspicious or non-suspicious. The encoding may be as defined in Table 1 .
[0011] The machine learning model may comprise one or more machine learning models selected from: a naive Bayes classifier, a support vector machine, a logistic regression model, a random forest classifier, and a multilayer perceptron. The machine learning model may comprise an ensemble of machine learning models. Each of the machine learning models in the ensemble of machine learning models may be a different type of model. Each of the machine learning models in the ensemble of machine learning models may use the same or different (e.g. subsets of) one or more features. The outputs of the machine learning models may be combined using majority voting or using a further classifier, e.g. a multilayer perceptron.
[0012] The one or more features may be features associated with a risk factor for characterising suspicious lesions. The one or more features may further comprise one or more of: patient age, patient gender, lesion size increase, lesion age, lesion shape change, lesion colour change, lesion total size, lesion inflammation, lesion oozing, lesion itch, lesion location, prior family history, natural hair colour, number of sunburns, number of moles, density of freckles, prior history of melanoma and prior history of skin cancer, the features are as defined in Table 1 .
[0013] In embodiments, the classifying uses four to seven features each associated with a risk factor for characterising suspicious skin lesions. The features used in addition to lesion pink may be selected from lesion size increase, lesion shape change, lesion colour change, lesion inflammation, hair colour, and lesion age. The features associated with risk factors for characterising suspicious skin lesions may comprise (a) lesion colour change, lesion inflammation and lesion age, (b) lesion colour change, lesion inflammation and lesion shape change, or (c) lesion pink, lesion inflammation, lesion shape change, lesion colour change, lesion size, lesion age and natural hair colour.
[0014] Ony one or more of the variables lesion pink, lesion size increase, lesion shape change, lesion colour change, lesion inflammation and lesion age may be encoded as binary variables. Lesion pink may be defined as whether the lesion is pink or not. In embodiments, yes is associated with a score or encoding of 1 and no is associated with a score or encoding of 0. Lesion inflammation may be defined as whether the lesion is inflamed or not. In embodiments, yes is associated with a score or encoding of 1 and no is associated with a score or encoding of 0. Lesion shape change may be defined as whether the lesion has changed in shape or not. In embodiments, yes is associated with a score or encoding of 1 and no is associated with a score or encoding of 0. Lesion colour change may be defined as whether the lesion has changed in colour or not. In embodiments, yes is associated with a score or encoding of 1 and no is associated with a score or encoding of 0. Lesion size may be defined as whether the lesion has a diameter at or above a predetermined threshold. The predetermined threshold may be 7mm. In embodiments, yes is associated with a score or encoding of 1 and no is associated with a score or encoding of 0. The diameter of a lesion may be the diameter of the smallest circle that the lesion fits into. Lesion age may be defined as whether the lesion has been present for at least a predetermined amount of time or not. The predetermined amount of time may be 6 months. In embodiments, yes is associated with a score or encoding of 1 and no is associated with a score or encoding of 0. Natural hair colour may be defined as the natural hair colour of the subject. In embodiments, natural hair colour is encoded as a discrete variable where black, red, blonde and brown are associated with respective discrete natural numbers. For example, black may be associated with a score or encoding of 1 , red may be associated with a score or encoding of 2, blonde may be associated with a score or encoding of 3, and brown may be associated with a score or encoding of 4.
[0015] Classifying the lesion between a first class of suspicious lesions and a second class of non- suspicious lesions may comprise weighting numerical values associated with each if said features such that the risk factors with greater significance for distinguishing between suspicious and non- suspicious lesions are given greater relative scoring values for determining the lesion risk score. The weighing may give importance to any of the following features in the following order from greater to lower importance: lesion pink, lesion inflammation, lesion shape change, lesion colour change, lesion size, lesion age and natural hair colour. The weighting may be as defined in Table 3. In embodiments, the one or more features further comprises a 7-point checklist score and / or a Williams score, or any one or more features of the 7-point checklist and / or Williams score. Thus, the method may further comprise calculating a Williams score and / or a 7-point checklist score and combining the Williams score and / or the-point checklist score with the lesion risk score to calculate an overall risk score.
[0016] The patient may be a human subject, for example a human subject of 18 years old or more. The lesion may be at most 15mm in largest dimension (e.g. diameter). The subject may be a subject who has been identified has having a skin lesion at risk of developing into skin cancer. The skin cancer may be melanoma, squamous cell carcinoma or basal cell carcinoma. A lesion classified in the suspicious class may be at higher risk of developing melanoma, basal cell carcinoma or squamous cell carcinoma than a lesion classified in the non-suspicious class.
[0017] In embodiments, the method further comprises capturing an image of the skin lesion and processing the image to determine an image risk score by detecting shapes, colours, textures and otherfeatures associated with the development of skin cancers and combining the image riskscore with the lesion risk score to calculate an overall risk score.
[0018] In embodiments, the method further comprises obtaining an image (or one or more images) of the skin lesion. In such embodiments classifying the lesion may comprise: using a machine learning model that takes as input (a) an image (or one or more images) of the lesion and (b) the data for one or more features associated with the lesion and / or a lesion risk score derived therefrom, wherein the machine learning model comprises a deep learning model that has been trained using training data comprising, for each of a plurality of training lesions: (a) an image (or one or more images) of the lesion and (b) the data for one or more features associated with the lesion and / or a lesion risk score derived therefrom, and (c) a label classifying each training lesion as suspicious or not suspicious. In embodiments, the machine learning model takes as input: (a) one or more images of the lesion, and (b) the data for one or more features associated with the lesion and a lesion risk score derived from the one or more features associated with the lesion, wherein the one or more features comprise or consist of: lesion pink, lesion inflammation, lesion shape change, lesion colour change, lesion size, lesion age and natural hair colour. The method may further comprise subjecting the image to a pre-processing step whereby the shape and / or resolution of the image is standardised. In embodiments, the image is resized to be square shaped using image padding, Image padding may be performed using image padding with the mode skin colour, image padding with a predetermined colour, or image padding with black pixels. In embodiments, he one or more images of the training lesions in the training data have been subject to a pre-processing step whereby the shape and / or resolution of the image is standardised.
[0019] The method may further comprise subjecting the image to a pre-processing step whereby the image is cleaned to remove features not relevant for determining the likelihood of a skin lesion being suspicious or non-suspicious. In embodiments, the one or more images of the training lesions in the training data have been subject to said pre-processing step. For example, the image may be subject to one or more pre-processing steps that includes hair removal and / or lesion detection and segmentation. Lesion detection and segmentation may comprise processing the image by object detection techniques to identify the parts of the image that are most relevant for determining the likelihood of a skin lesion being suspicious or non-suspicious.
[0020] In embodiments making using of images, the machine learning model may comprise a deep learning model. The machine learning model may comprise a convolutional neural network, such as e.g. an efficientNet. In embodiments, the machine learning model comprises an ensemble of models. For example, the machine learning model may comprise an ensemble of machine learning models each taking as input (a) an image of the lesion and (b) the data for one or more features associated with the lesion and / or a lesion risk score derived therefrom. Each machine learning model may have been trained and / or tested with images acquired using a different type of camera (e.g. dermoscopic or DSLR cameras). The outputs of the machine learning models in the ensemble may be combined using majority voting.
[0021] In embodiments, the one or more images comprise or consist of an image that has been acquired using a dermoscopic camera. In embodiments, the training data comprises for each of the plurality of training lesions, an image that has been acquired using a dermoscopic camera or an image that has been acquired using a digital single-lens reflex camera. In embodiments, the training comprises for each of a first plurality of training lesions, an image that has been acquired using a dermoscopic camera, and for each of a second plurality of training lesions, an image that has been acquired using a digital single-lens reflex camera.
[0022] In embodiments, the method further comprises producing a saliency map associated with the machine learning model to identify the parts of the image that are more relevant for determining the likelihood of a skin lesion being suspicious or non-suspicious. In embodiments, classifying, using said data, the lesion between a first class of suspicious lesions and a second class of non- suspicious lesions comprises obtaining a score indicative of the likelihood that the lesion is in the second class, optionally wherein the score is a probability. In embodiments, classifying, using said data, the lesion between a first class of suspicious lesions and a second class of non-suspicious lesions comprises obtaining a score indicative of the likelihood that the lesion is in the first class, optionally wherein the score is a probability. A score indicative of the likelihood that the lesion is in the first class may be obtained as 1 minus a score indicative of the likelihood that the lesion is in the second class, or vice-versa. In embodiments, the method comprises classifying the lesion between a first clinical pathway class associated with a score above a first predetermined threshold, a second clinical pathway class associated with a score at or below the first predetermined threshold and above a second predetermined threshold, and a third clinical pathway class associated with a score at or below the second predetermined threshold. In embodiments, the first predetermined threshold is selected such that a maximum predetermined number or proportion of suspicious lesion in a training cohort is classified in the first clinical pathway class. In embodiments, the maximum predetermined number or proportion is zero, and / or wherein the first predetermined threshold is the lowest predetermined threshold such that no suspicious lesion in training dataset is classified in the first clinical pathway class. In embodiments, the method comprises selecting the patient classified in the first clinical pathway class for discharge, and / or selecting the patient classified in the second clinical pathway class for health care practitioner reporting, and / or selecting the patient classified in the third clinical pathway class for a face-to-face assessment with or without biopsy.
[0023] In embodiments, the method further comprises selecting a patient classified in the first category for one or more further diagnostic tests, such as e.g. a biopsy. Thus, also described herein according to a further aspect is a method of treating a patient who has been identified as having a skin lesion at risk of being suspicious, the method comprising: classifying the skin lesion in the patient as suspicious or non-suspicious using the method of any embodiment of the first aspect; and performing a further diagnostic test (e.g. a biopsy), on the patient when the lesion has been classified as suspicious and / or when the lesion has been classified in the third clinical pathway class. Such methods may further comprise treating the patient with a first treatment depending on the result of the further diagnostic test. The first treatment may comprise a surgical removal of the lesion when the further diagnostic test indicates that the lesion is malignant or pre-malignant.
[0024] Also described according to a further aspect is a device for implementing the methods described herein, wherein the device comprises apparatus for capturing an image and a processing unit. Also described herein is a system comprising one or more processors and one or more computer readable memories storing instructions that, when executed by the one or more processors, cause the one or more processors to implement any method described herein. The system may further comprise image acquisition means (e.g. a dermoscopic or DSLR camera).
[0025] Also described according to a further aspect is a non-transitory computer readable medium storing instructions that, when executed by a processor, cause the processor to implement any method described herein. Also described according to a further aspect is a computer program product comprising instructions that, when executed by a processor, cause the processor to implement any method described herein.
[0026] BRIEF DESCRIPTIONS OF THE DRAWINGS
[0027] Embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings.
[0028] Figure 1a shows a schematic representation of the first aspect of the disclosure.
[0029] Figure 1 b shows a flowchart illustrating schematically a method according to a general embodiment of the disclosure.
[0030] Figure 1c shows an example of a partly automated dermatology platform according to an embodiment of the disclosure.
[0031] Figure 2 shows a chart demonstrating the distribution of a skin lesion being pink in colour or not pink in colour for skin lesions determined to be suspicious, and also for those determined to be non-suspicious.
[0032] Figure 3 shows a schematic representation of an embodiment of the invention where one or more further pieces of data corresponding to risk factors for characterising skin lesions are also scored to determine the lesion risk score.
[0033] Figure 4 shows a schematic representation of an embodiment of the invention where the method further comprises calculating an image risk score from a captured image in addition to calculating the lesion risk score to determine an overall risk score.
[0034] Figure 5 shows a chart demonstrating analysis of Hu moment features which are indicators of the shapes in an image of a skin lesion that are associated with the development of skin cancers.
[0035] Figure 6 shows an example of a captured image of a skin lesion, Figure 6a shows the original captured image and Figure 6b shows a version of the image in Figure 6a which has been subjected to a pre-processing step of standardising the shape and resolution of the original image according to an embodiment of the invention. Raw image size: 3.39MB, resolution: 2848x4273, pre-processed image size: 187KB, resolution: 1024x1024.
[0036] Figure 7 shows an example of a captured image of a skin lesion subjected to a pre-processing step to remove hair. Figure 7a shows the original captured image and Figure 7b shows a version of Figure 7a which has been subjected to a pre-processing step.
[0037] Figure 8 shows an additional example of a captured image of a skin lesion subjected to a preprocessing step for skin lesion detection and hair removal. Figures 8a and 8b show the original captured image and Figure 8c shows a version of the image which has been subjected to a preprocessing step.
[0038] Figure 9 shows a schematic representation of an embodiment of the invention where different scaling techniques are employed where ensembling is utilised to enable models of different sizes to be combined together. Compound scaling of EfficientNet that can accommodate different image sizes to gain better performance: Fig. 9(a) shows a baseline network example; Fig. 9(b)-(d) show conventional scaling that only increases one dimension of network width, depth, or resolution. Fig. 9(e) shows a new version adapted from the compound scaling method in Tan & Le 2019 that uniformly scales all three dimensions with a fixed ratio.
[0039] Figure 10 shows an example of a captured image of a skin lesion subjected to processing by attention mechanism visualisation.
[0040] Figure 11 shows an embodiment of a system for performing methods of the disclosure.
[0041] Figure 12 shows a chart comparing lesion score and Williams score for suspicious and non- suspicious cases.
[0042] Figure 13 shows a graph comparing patient age distributions for suspicious and non-suspicious cases.
[0043] Figure 14 shows a schematic representation of a proposed Al framework to identify a set of new risk factors for skin lesion classification.
[0044] Figure 15 shows a schematic overview of a proposed Al framework using patient metadata to identify a set of new risk factors followed by ranking and weighting those risk factors to deduce a new risk score for skin lesion classification.
[0045] Figure 16 shows examples of suspicious and non-suspicious skin lesion images.
[0046] Figure 17 shows images of suspicious and non-suspicious skin lesions captured with two types of cameras (dermoscopic and DSLR).
[0047] Figure 18 shows padding approaches (padding with mode skin colour and padding with black colour) tested for image reshaping.
[0048] Figure 19 shows a schematic representation of a proposed Al framework for skin lesion classification into suspicious vs non-suspicious groups based on metadata (C4C risk score) and images.
[0049] Figure 20 shows a schematic representation of the fusion of metadata with image data to build an Al framework.
[0050] Figure 21 A shows the model accuracy for skin cancer and benign lesion detection for an Al model using image data alone. Figure 21 B shows the model accuracy for skin cancer and benign lesion detection for an Al model using image data and metadata.
[0051] Figure 22 shows the principles used in development of a multi-modal Al framework which combines patient metadata with skin lesion images to classify suspicious skin lesions. A total of six Al models have been adapted and fused by varying input data types: (a) DER, (b) SLR, (c) DER+metadata, (d) SLR+metadata, (e) SLR+DER, and (f) SLR+DER+metadata.
[0052] Figure 23 shows heat map generation from the last layer of EfficientNet using Grad-CAM to check whether the model focusing on the correct lesion area or not.
[0053] Figure 24 shows adoption of the scaled dot product attention mentioned in Oquab et al. (2023, preprint) to force an Al model to focus to the correct lesion area and ignore artefacts such as rulers.
[0054] Figure 25 shows the number of misclassified skin cancers with increasing confidence in non- suspicious classification in a validation experiment.
[0055] Figure 26 shows a schematic of the partial automation of skin lesion classification using an Al model and benign confidence threshold of 93%.
[0056] DETAILED DESCRIPTION
[0057] Embodiments of the present disclosure provide a method for classifying skin lesions as suspicious or non-suspicious comprising; collecting and scoring a first piece of data comprising a determination of whether a skin lesion is pink in colour or not pink to provide a lesion risk score to determine the likelihood of a skin lesion being suspicious or non-suspicious.
[0058] The person skilled in the art will realise there are various methods for scoring the skin lesions, following a determination of whether they are pink or not according to the first aspect of the invention. For example, a pink lesion can be scored positively (for example by scoring it 1) where the lesion is determined to be pink, and then a skin lesion can be scored neutrally (for by scoring it 0) where the skin lesion is not determined to be pink. In this example a risk lesion score of 1 will provide a determination that the skin lesion in question is suspicious and requires further clinical assessment, whereas a score of 0 will provide a determination that the skin lesion in question is likely non-suspicious.
[0059] In determining whether a skin lesion is pink or not the medical professional performing the assessment will consider whether it is pink or pinky-red in colour and non-pigmented, and score these lesions as pink, but this does not include deeper reds caused by ulceration or bleeding. In determining pink colour according to the present invention, the colour pink is also distinct from the pigmented brown and black moles and brown freckles that are scored in the 7PCL methodology for determining melanomas. The person skilled in the art will also understand that pink means predominantly pink as lesions are not always unform in colour and so may have some imperfections or flecks that are not pink in colour.
[0060] Preferably the method for classifying skin lesions as suspicious or non-suspicious is executed on a data processing device such as central processing unit (CPU) on a computer to determine the lesion risk score.
[0061] In embodiments, the method further comprises collecting and scoring one or more further pieces of data corresponding to risk factors for characterising skin lesions and combining with the score corresponding to the first piece of data to provide the lesion risk score.
[0062] The person skilled in the art would understand that there are various methods for scoring the risk factors. Preferably a scoring methodology will be employed where risk factors indicative of suspicious lesions will be scored positively (for example by scoring each risk factor being characterised as 1) whereas those indicative of non-suspicious lesions will be scored neutrally (for example by scoring each one as 0). For every piece of data collected and scored, a score will be given based on a determination of whether the risk factor is indicative of a suspicious skin lesion and an alternative score will be given where it is determined to be non-suspicious, these will then be combined with the score from the first piece of data (i.e. the determination of whetherthe lesion is pink in colour or not) to provide the lesion risk score. The lesion risk score in the first embodiment of the invention is a total of all the scores for all the pieces of data characterised, so for example where four pieces of data are collected in addition to the first piece of data (lesion pink or not) then five pieces of data will be scored and combined to determine the final lesion risk score.
[0063] In embodiments, the risk factors for characterising suspicious skin lesions are selected from one or more of patient age, patient gender, lesion size increase, lesion age, lesion shape change, lesion colour change, lesion total size, lesion inflammation, lesion oozing, lesion itch, lesion location, prior family history, natural hair colour (at age 15), number of sunburns, number of moles, density of freckles, prior history of melanoma and prior history of skin cancer.
[0064] In this embodiment, a piece of data is collected, categorised and scored for each of the selected risk factors. The characterisation of the piece of data involves either a numeric or categorical coding being assigned, for example for patient age this will be assigned a numerical coding corresponding to the patient’s age whereas for categorical data types a numerical coding will be assigned according to category type (examples of possible coding is given in Table 1 below). The categorisation coding is then converted into a score for each selected risk factor. Where a positive scoring methodology is used, risk factors that are more likely to provide a determination that a skin lesion is suspicious are given a higher scoring when present. The scores of the selected risk factors are then combined together, also including the score from the first piece of data (lesion pink or not pink) to determine the final lesion risk score. Table 1 below provides details of how each of these different risk factors are characterised, categorised and coded, for example the risk factor lesion size is characterised as to whether the lesion has changed in size and then categorised accordingly with a one or zero coding. Table 1 below also gives an example of possible category coding for all the different risk factors depending on how they are characterised, the person skilled in the art would understand that different categorisation and coding methodologies will be possible. Table 1. Examples risk factors and category coding.
[0065] The person skilled in the art would understand that various different scoring methodologies could be adopted for converting the coding into scores. For example, a positive scoring methodology could be employed where coding is converted into a score between zero and ten with a score of ten being a score associated with a high likelihood that the risk factor being characterised strongly correlates with a determination that the skin lesion is suspicious, and a score of zero associated with a very low or no likelihood.
[0066] In embodiments, the method further comprises collecting four to seven pieces of data wherein each piece of data corresponds to a risk factor for characterising suspicious skin lesions and combining the scores corresponding to the four to seven pieces of data to provide the lesion risk score. In such embodiments the pieces of data collected and scored will always include a determination of whether the skin lesion is pink in colour or not, and then three to six additional pieces of data before combining the scores to determine a final lesion risk score. Most preferably the method comprises collecting and scoring seven pieces of data in total, including a determination of whether the lesion is pink or not pink, and combining the scores to determine a final lesion risk score. It has been established that collecting and scoring four to seven pieces of data is an optimal number to obtain a final lesion risk score that delivers a strong indicator of whether a skin lesion is considered suspicious or not, while keeping the need to perform calculations on the dataset to a manageable number. While it is possible to collect and score more than seven pieces of data in total, the improvements in determining whether a skin lesion is suspicious or not become more and more marginal beyond seven. By capturing the data relating to more than seven risk factors the resource requirements in terms of data collection and processing outweigh the accuracy improvements which become more and more marginal.
[0067] In embodiments, the one or more further pieces of data corresponding to risk factors for characterising suspicious skin lesions are selected from lesion size increase, lesion shape change, lesion colour change, lesion inflammation, hair colour, and lesion age.
[0068] It is these six risk factors, in addition to the determination of whether a skin lesion is pink or not, that have most significance in determining whether a skin lesion is suspicious or not. By collecting metadata relating to thousands of skin lesions from patients attending skin cancer diagnosis clinics and combining scores relating to different risk factors (including all those listed in Table 1 above), the risk factors that are the most significant in providing a determination of whether a skin lesion is suspicious or not have been revealed. Regression analysis was employed to demonstrate the relative importance of each of these risk factors in reaching a final determination of whether a skin lesion is suspicious or not. The relative importance of each of these seven risk factors as determined by regression analysis is provided in Table 2 below.
[0069] Table 2. Risk factors and relative importance in example embodiment.
[0070] In determining the relative importance values, a combination of five machine learning models were used to assess the accuracy of each risk factor listed in Table 1 above, alone and in combination with other risk factors, to correctly classify the skin lesion as suspicious or non-suspicious. The relative input of the seven most important risk factors was then identified using logistic regression. As can be seen in Table 2 above that a determination of whether a skin lesion is pink in colour or not (Lesion Pink) is the most significant factor in determining whether a skin lesion is suspicious or not.
[0071] In embodiments, the method comprises collecting and scoring four to seven pieces of data, including the determination of whether a skin lesion is pink or not and also including three to six of the risk factors listed in Table 2 above, before combining the scores to provide a final lesion risk score. Most preferably the method comprises collecting and scoring seven pieces of data corresponding to the seven risk factors mentioned in Table 2 above, and combining the scores to provide a final lesion risk score to determine whether a skin lesion is suspicious or not.
[0072] In embodiments, the one or more further pieces of data corresponding to risk factors for characterising suspicious skin lesions comprise lesion colour change, lesion inflammation and lesion age.
[0073] As well as determining the seven most significant factors in combination, regression analysis was also used to determine the 10 most important in combination, the 15 most important in combination and the 20 most important in combination. From each of these combinations four risk factors were always present in these lists, these were a determination of whether a skin lesion is pink in colour or not, lesion colour change, lesion inflammation and lesion age.
[0074] In embodiments, the one or more further pieces of data corresponding to risk factors for characterising suspicious skin lesions comprise lesion colour change, lesion inflammation and lesion shape change.
[0075] After the determination of whether a skin lesion is pink in colour or not, these three further risk factors are the next most significant factors in determining whether a skin lesion is suspicious or not (see Table 2 above).
[0076] In embodiments, the scoring is weighted such that the risk factors with greater significance for distinguishing between suspicious and non-suspicious lesions are given greater relative scoring values for determining the lesion risk score. As certain features have greater significance than others in determining whether a skin lesion is suspicious or not, the person skilled in the art would understand that it makes sense to apply weighting to certain features to have a greater influence on the final lesion risk score. This is particularly the case the more risk factors are included in the analysis. For example, the risk score corresponding to a determination of whether a lesion is pink in colour or not could be scored two with a positive determination (i.e. double weighting) whereas the other risk factors could be scored one with a positive determination so recognise that the determination of whether a lesion is pink in colour or not is the most significant risk factor (see Table 2 above). The person skilled in the art would understand that multiple different weighting criteria could be applied depending on the circumstances and the particular combinations of risk factors being analysed. Table 3 below provides an example of the type of weighting proportions that can be assigned to the seven most significant risk factors and, it can be seen that the risk factors with the greatest relative importance (as determined by the regression analysis) have been assigned the greater weighting values in this example and the weighting is roughly proportional to the relative importance of the risk factors. The person skilled in the art would understand other weighting values might be assigned, for example weighting to amply the influence of certain risk factors in the final lesion risk score (i.e. where they are not broadly proportional to the relative importance values).
[0077] Table 3. Example weighting
[0078] For the example in Table 3 above, the person skilled in the art would understand that if a positive scoring methodology is used between zero and ten (with ten being a score most strongly associated with a suspicious skin lesion) then lesion pink is the only risk factor that could score as highly as ten and that would be achieved by a determination that the lesion is pink in colour as defined above. Similarly (and according to the same scoring and weighting methodology) the presence of lesion inflammation could only score a maximum score of four, and so on. The person skilled in the art would understand that other weighting values could be assigned to the various risk factors and could be applied to all risk factors being characterised, beyond just the seven most significant risk factors identified in Tables 2 and 3 above.
[0079] In embodiments, a threshold score is set at a predetermined value and where the lesion risk score is equal or higher than the threshold score, a lesion is classified as being suspicious and where the lesion risk score is lower than the threshold score, a lesion is classified as being non- suspicious. In some such embodiments where a positive determination is made that a risk factor is present according to the scoring criteria (for example see Table 1 above), a positive score is assigned and so the more positive scores there are when the scores from all the risk factors are combined, the larger the final lesion risk score will be. In this case the person skilled in the art would understand that it is possible to set a threshold score and if a particular lesion risk score is higher or equal to the threshold score then a skin lesion is classified as being suspicious. The person skilled in the art would also understand that the threshold score can be adjusted dependent on the level of sensitivity required, for example it could be lowered so more skin lesions are classified as suspicious. When adjusting the threshold score, the overall aim is to balance sensitivity against accuracy and so the threshold shouldn’t be reduced so low that accuracy in making the determination between suspicious and non-suspicious lesions is compromised. In embodiments, the method further comprises calculating a 7-point checklist score and combining the 7-point checklist score with the lesion risk score to determine an overall risk score. The 7-point checklist score (7 PCL) is well known in the art and is used to score seven risk factors: lesion size, lesion colour change, lesion shape, lesion size, lesion inflammation, lesion oozing and lesion itch. The score from the 7 PCL can be combined with the score from the lesion risk score of the present invention to determine an overall risk score. Combining the lesion risk score with the 7PCL improves the overall sensitivity in determining whether a skin lesion is suspicious or not suspicious.
[0080] In embodiments, the method further comprises calculating a Williams score and combining the Williams score with the lesion risk score to calculate an overall risk score. The Williams Score is well known in the art and is used to score seven risk factors: patient gender, patient age, number of sunburns, natural hair colour, number of moles, density of freckles and prior skin cancer. The Williams score can be combined with the score from the lesion risk score of the present invention to determine an overall risk score. Combining the lesion risk score with the Williams score improves the overall sensitivity in determining whether a skin lesion is suspicious or not suspicious. In yet a further embodiment of the invention the Williams score, 7 PCL and the lesion risk score can all be combined to obtain even greater sensitivity in determining whether a skin lesion is suspicious or not suspicious.
[0081] Table 4 below compares the different risk scoring methodologies; 7 PCL, Williams score and the lesion risk score of the present invention in terms of sensitivity, specificity and accuracy for all skin cancer types. For the lesion risk score of the present invention, seven risk factors were used in the scoring and these are the seven found to be most significant in making a determination of whether a skin lesion is suspicious or not as listed in Tables 2 and 3 above (i.e. lesion pink, lesion inflammation, lesion shape change, lesion colour change, lesion total size, lesion age and natural hair colour). Table 4 also includes performance characteristics in terms of sensitivity, specificity and accuracy where different risk scoring methodologies were calculated in combination. It can be seen that the greatest sensitivity and accuracy is achieved by combining all three risk scoring methodologies together. However, this comes at the expense of needing to collect and score data corresponding to 18 different risk factors and then perform the calculations on these. Comparing the risk scoring methodologies alone (where seven risk factors are collected and scored in each case), the lesion risk score of the present invention delivers the best performance in terms of sensitivity, specificity and accuracy.
[0082] Table 4. Performance of methods of the disclosure and comparative methods
[0083] The performance of the 7 PCL method can be improved further by analysing additional risk factors associated with the development of skin cancers (i.e. in addition to the risk factors already scored in the 7 PCL method). One would however need to analysis a further four risk factors (so 11 in total) to get to a comparable sensitivity score to the lesion risk score alone which only requires the collection and analysis of data relating to seven risk factors. Adding and analysing a further four risk factors to the 7 PCL delivers a sensitivity of 80.56%, accuracy of 61 .91% and specificity of 60.07%. Even by adding and analysing several additional risk factors to the seven that are normally considered in the 7PCL you cannot get comparable accuracy and specificity performance gains that are comparable to using the lesion risk score of the present invention where seven risk factors are considered.
[0084] Determining a lesion risk score according to embodiments of the present disclosure including the collection and analysis of seven risk factors therefore delivers significant performance benefits in terms of sensitivity, specificity and accuracy over the prior art methods without needing to include the analysis of additional risk factors in the calculations and so is therefore also more efficient in terms of data processing. Furthermore, calculating the lesion risk score according to the present disclosure also allows the classification of skin lesions as suspicious or non-suspicious beyond just melanomas. The method of the present invention provides a determination regarding other skin cancer types including basal cell carcinomas and squamous cell carcinomas, and also benign skin lesions.
[0085] In embodiments, the method further comprises capturing an image of the skin lesion and processing the image to determine an image risk score by detecting shapes, colours, textures and otherfeatures associated with the development of skin cancers and combining the image riskscore with the lesion risk score to calculate an overall risk score. The method of processing images to detect shapes, colours, textures or other features associated with the development of skin cancers may be executed on a data processing device such as central processing unit (CPU) on a computer to determine the image risk score, and also to calculate the overall risk score.
[0086] The person skilled in the art would understand that an image of a skin lesion may be captured by any suitable device, for example a camera including cameras incorporated into a mobile phone. Once the image is captured, it is then processed to detect certain shapes, colours, textures and other features associated with the development of skin cancers. The person skilled in the art would understand that the image processing is configured to detect any feature of a skin lesion that can be visually distinguished in an image by digital means and is not just limited to shapes, colours and textures.
[0087] The detection of relevant features in an image of a skin lesion can be achieved by various different image feature extraction techniques. For example, Hu moments and Zernike moments can be identified for shape detection on the basis of shape descriptors, Haralick features and binary histogram features can be identified for texture detection on the basis of texture descriptors and colour histogram features can be identified for colour detection. The person skilled in the art would understand that other image feature extraction techniques may also be utilised to detect relevant features in the captured image. ABCDT feature analysis can also be used to detect relevant features in an image on the basis of asymmetry, border irregularity, colour variegation, diameter and texture. Asymmetry of the lesions is characterised through asymmetric index (Al) and eccentricity, border irregularity is determined through compact index (Cl) and diameter is determined as the diameter of a circle with the same area of the lesion region. Colour variegation is quantified by the normalised standard deviation of red, green and blue components of the lesion and this is important because objects with higher normal skin pigmentation will also tend to have higher pigmentation in the lesion itself and this needs to be normalised when making a comparison between lesions from different individuals. Finally characterising an image based on texture may be done by a variety of techniques including techniques quantifying the Haralick texture achieved through the computation of Gray Level Co-occurrence Matrix (GLCM).
[0088] The captured image may be compared to information held in an image repository to determine an image risk score. Where certain features detected in the image correspond closely to features classified as indicative of suspicious skin lesions then this will be scored appropriately, for example by receiving a positive score. By contrast features corresponding closely to features classified as indicative of non-suspicious skin lesions are given an alternative score (to those indicative of suspicious lesions), for example by receiving a neutral score. In embodiments the image repository contains information and metadata corresponding to image features from a large number of images collected previously and where these images of skin lesions have also been previously classified as being either suspicious or non-suspicious.
[0089] The person skilled in the art will understand that the image repository can be constructed by machine learning techniques where images of skin lesions are fed into an Al model where labels are added to the images dependent on whether they were classified as suspicious or non- suspicious. The resulting Al model can then be used to perform the comparison between the image repository and the features captured in the skin lesion image being analysed, and then calculate the image risk score. The person skilled in the art would understand there are also methods of constructing an image repository that don’t involve machine learning techniques. Feature identification methods can also be utilised to build the Al model for performing the comparison between the images in the image repository and the skin lesion being analysed. These feature identification methods can be utilised to select the most relevant features from the images that are associated with the development of skin cancers. An example of a suitable technique involves using wrapper, Shapley and Pearson correlation methods in combination. In a first step a wrapper technique adapts a random forest (RF) classifier as a base learner and iteratively estimates model performance for a subset of features. Ten-fold cross-validation is used to evaluate the performance of the RF classifier as a feature selector to identify an optimal subset of metafeatures. In a second step, a Shapley method assigns high scores to meta-features which have a high success rate in classifying instances correctly and this value-based feature selection method measures the marginal contribution of each feature when combined with other features. In a final step a Pearson correlation method ranks each meta-feature with a value from 1 (highly correlated) to -1 (no correlation) to identify the features which are most relevant in correctly distinguishing whether a skin lesion is suspicious or non-suspicious, i.e. features that are highly associated with the development of skin cancers.
[0090] Combining the image risk score with the lesion risk score is particularly advantageous where the skin lesion type is a benign lesion. When analysing image data alone the model is only 63% accurate at correctly classifying benign lesions, with the remaining 37% being false positives. By analysing the image data to calculate an image risk score and combining this with a lesion risk score calculated by collecting and scoring the seven most important risk factors (those listed in Table 2 above), 74% accuracy can be achieved and so only 26% of cases were false positives. The accuracy can be improved even further by adjusting the threshold value in accordance with the embodiment of the invention described above.
[0091] In embodiments, the method further comprises subjecting the image to a pre-processing step whereby the shape and / or resolution of the image is standardised. Suitable image reshaping techniques include reshaping keeping aspect ratio, reshaping with skin tone colours, reshaping with black padding and reshaping with white padding.
[0092] The person skilled in the art would understand that to perform a meaningful comparison between the captured image and the information and metadata held in the image repository it is advantageous that the images are comparable in terms of overall shape and resolution. For example, all captured images and image repository images could be square shaped and having the same resolution to allow for direct comparison. In this example, where a raw captured image is rectangular and has a high resolution, it could be reshaped to be square shaped and then the resolution could be reduced to be the standard resolution. The person skilled in the art would understand that it is advantageous to have higher resolution images used in both the image repository and also the captured images for better detection performance however this needs to be balanced against computational resources being used to pre-process the images in performing the comparison.
[0093] Subjecting the captured images to these pre-processing steps is relatively simple and can significantly reduce the overall size and hence processing and storage requirements of the image when it is subsequently analysed for the purpose of detecting certain shapes, colours, textures and other features associated with the development of skin cancer. By standardising the images, significantly fewer computational resources are also utilised when performing the comparison between the captured image and the images in the image repository because it is much simpler to do a like for like comparison with standardised images. Performing a comparison on standardised images also results in much better performance in terms of correctly identifying comparable features for the purpose of calculating the image risk score. Where machine learning is utilised to construct the image repository and an Al model is produced for performing the comparison, fewer computational resources in terms of memory requirements and processing fortraining the Al model are used on standardised images compared to using the raw images.
[0094] In embodiments, the image is resized to be square shaped. For example, the image may be subjected to a pre-processing step where the standardised shape is square shaped. Square shaped images are advantageous for performing the comparison between the captured images and the images in the image repository when analysing the features in the images associated with the development of skin cancer. Square shaped standardised images result in much better performance in terms of correctly identifying comparable features when calculating the image risk score.
[0095] In embodiments, the method further comprises subjecting the image to a pre-processing step whereby the image is cleaned to remove features not relevant for determining the likelihood of a skin lesion being suspicious or non-suspicious. This pre-processing step may be used in conjunction with the pre-processing step where the images are standardised but also separately. By subjecting the images to this pre-processing step irrelevant features (in determining whether a skin lesion is suspicious or not) are removed, this improves the overall performance when calculating the image risk score. This also improves the computational efficiency in calculating the image risk score as there is no need to analyse features in the images that have no bearing on determining whether a lesion is suspicious or not. Examples of features not relevant for determining the likelihood of a skin lesion being suspicious or non-suspicious are rulers, which are often included in the original images to estimate lesion size, and body hairs.
[0096] In embodiments, the method further comprises processing the image by object detection techniques to identify the parts of the image that are most relevant for determining the likelihood of a skin lesion being suspicious or non-suspicious. Object detection techniques allow the differentiation of background, artifacts (such as rulers or hairs) and skin lesion features with a high degree of confidence. Even where the image has already been subjected to a pre-processing step of cleaning the image to remove features not relevant for distinguishing between suspicious and non-suspicious lesions, using object detection techniques provides an additional layer of differentiation of relevant features over features that are not relevant. When the image is subsequently analysed to determine whether there are features present that are associated with the development of skin cancers, this analysis only needs to be in relation to relevant features and so the background and other artifacts are excluded. By performing the analysis on only the most relevant objects in the image computational efficiencies are achieved as irrelevant information is not analysed. Object detection techniques typically involve the steps of i) identifying a region of interest, for example using an algorithm such as selective search, ii) generating a feature vector for each proposed region, for example by applying image processing techniques such as histogram of oriented gradient (HOG), and iii) classifying the feature vector to determine whether it belongs to the background or an object class, this is typically achieved using deep learning-based techniques for object detection such as Region-based convolutional neural network (R-CNN), Fast convolutional neural network (Fast-CNN), Faster convolutional neural network (Faster-CNN), You Only Look Once (YOLO) and Single-Shot Detector (SSD).
[0097] In embodiments, the method further comprises utilising two or more different modelling techniques to process the image and ensembling the outputs of the different modelling techniques. Ensembling enables multiple modelling techniques to be combined when analysing the feature in the image of the skin lesion. Ensembling also enables models of different sizes to be combined together. Preferably the resulting output of the different models is then merged when calculating the image risk score. Ensembling can also accommodate different image sizes, for example by employing compound scaling to improve performance. Compound scaling techniques include width scaling, depth scaling, resolution scaling and compound scaling. Compound scaling enables different images sizes to be made comparable when analysing and this improves the performance when calculating the image risk score. For example, ensembling can be achieved using EfficientNet however the person skilled in the art would realise that other image classification methods can be used as an alternative. The combined modelling techniques and merged output by ensembling is better at differentiating between true positives and false positives, and also true negatives and false negatives. In the context of the present disclosure, a true positive is defined as suspicious classified as suspicious, a false positive is defined as non-suspicious defined as suspicious, a true negative is defined as non-suspicious defined as non-suspicious and a false negative is defined as suspicious defined as non-suspicious. By filtering out incorrect information (i.e. the false positives and false negatives) at an early stage this reduces the overall amount of data that needs to be processed to provide the final image risk score, it also provides better quality outputs in terms of the accuracy of the image risk score.
[0098] In embodiments, the method further comprises processing the image by attention mechanism visualisation to identify the parts of the image that are more relevant for determining the likelihood of a skin lesion being suspicious or non-suspicious. Attention mechanism visualisation improves the ability to focus on the most relevant parts of the image during image processing. This can include focussing on the skin lesion while ignoring background elements in the image but also focussing on the most relevant features within the skin lesion itself. Examples of attention mechanism visualisation techniques include scaled dot product attention, multiplicative attention, additive attention, and custom attention, however the person skilled in the art would understand other techniques are possible as well. Attention mechanism visualisation delivers processing efficiencies as only the most relevant features are processed in the image when calculating the image risk score. Irrelevant features such as background artifacts are excluded entirely and features in the skin lesion that do not have a bearing on a determination of whetherthe skin lesion is suspicious or not are also much more likely to be excluded from processing.
[0099] In a second aspect of the disclosure, there is provided a device for implementing the methods of the disclosure, where the method further comprises capturing an image of the skin lesion and processing the image to determine an image risk score by detecting shapes, colours, textures, and otherfeatures associated with the development of skin cancers and combining the image riskscore with the lesion risk score to calculate an overall risk score, wherein the device comprises apparatus for capturing an image and a processing unit. The apparatus for capturing an image can be any device that is suitable for capturing an image of a skin lesion in sufficient detail that certain shapes, colours, textures, and other features associated with the development of skin cancers can be captured in the image. Examples include cameras (including cameras build into a mobile phone device), scanners and other digital imaging devices. Advantageously, the image capturing apparatus may be one that is suitable for capturing digital images to enable the subsequent image processing steps. The processing unit can be any unit that is suitable for processing images captured by the image capturing apparatus, and in particular processing the image to detect shapes, colours, textures and other features associated by the development of skin cancers. The processing unit also needs to be suitable for processing information to determine an image risk score, and furthermore to calculate an overall risk score. The processing unit may advantageously also be suitable for implementing any one or more of the other processing steps as defined herein including the pre-processing step whereby the shape and / or resolution of the image is standardised (including to be square shaped), the pre-processing step where the image is cleaned to remove features not relevant to determining the likelihood of a skin lesion being suspicious or non-suspicious, the object detection techniques, ensembling the outputs of different modelling techniques and carry out attention mechanism visualisation to identify parts of the image that are more relevant for determining the likelihood of a skin lesion being suspicious or non-suspicious.
[0100] In advantageous implementations the device comprising an apparatus for capturing an image and a processing unit may be a mobile phone device with an inbuilt camera. In embodiments this is utilised in combination with an App provided on the mobile phone device capable of saving copies of the one or more images of the skin lesions taken, any risk scores calculated and in particular the image risk score calculated by the processing unit and any other data processed according to the method the present invention. It is also possible to have a desktop version of the App onboard a personal computer capable of communicating with the apparatus for capturing an image and the processing unit. Preferably the App will be configured to send copies of the images and / or other data and information captured and saved to a clinician or other medical specialist for further analysis.
[0101] Figure 1a shows a schematic representation of a general embodiment of the disclosure. Figure 1 a shows a schematic representation of a method for classifying skin lesions as suspicious or non- suspicious comprising; collecting and scoring a piece of data comprising a determination of whether a skin lesion is pink in colour or not pink to provide a lesion risk score to determine the likelihood of a skin lesion being suspicious or non-suspicious. In a first step the method comprises a determination of whether a skin lesion being classified is pink in colour or not pink (11) resulting in a first piece of data being generated. In a second step the method comprises scoring the first piece of data to provide a lesion risk score (12) to determine the likelihood of a skin lesion being suspicious (13) or non-suspicious (14).
[0102] A lesion classified as “suspicious” may refer to a lesion that is more likely to develop into a skin cancer (e.g. melanoma) than a lesion classified as “non-suspicious”. Thus, the term “suspicious” may refer to a lesion being likely to develop into melanoma. The term “non-suspicious” may refer to a lesion that is unlikely to develop into a skin cancer (e.g. melanoma) and / or that is likely to be benign.
[0103] Figure 1 b shows a schematic representation of a method of classifying a lesion according to embodiments of the disclosure. At optional step 110, data for one or more features associated with a lesion in a patient are obtained. This may comprise obtaining such value from a data store or user interface, such as e.g. through a user (patient or healthcare professional) providing input data comprising said data. Step 110 may further comprise acquiring one or more images of the lesion using an image acquisition means, such as a DLSR camera or a dermoscopic camera. Step 110 is optional because the data may have been previously obtained and may simply be received by a processor implementing a method as described herein. At optional step 120, any image data obtained at step 110 may be pre-processed. This may comprise re-sizing, padding, hair removal, object detection and / or lesion segmentation. In embodiments, pre-processing the image data comprises applying a hair removal algorithm. In embodiments, pre-processing the image data comprises resizing to a predetermined size by adding padding pixels. In embodiments, the padding pixels have a value corresponding to the mode skin colour (e.g. the mode of the pixel values in the image, optionally after hair removal and / or cropping or other removal of non-skin regions in the image). At step 130, feature data obtained at step 110 in non numerical form may be encoded in numeral form for further processing. At step 140, a subset of the data obtained at step 110 may be used to calculated a risk score for the lesion, such as e.g. a 7-PCL score and / or a Williams score for the lesion, or a risk score as described herein. Any of these scores may be used as additional features for classification at step 150. Alternatively, the value of a 7-PCL score and / or a Williams score and / or a risk score as described herein may have been provided as a feature (e.g. by a user through a user interface) at step 110. At step 150, the features obtained through the preceding steps are used to classify the lesion as suspicious or non-suspicious. Classifying the lesion as suspicious or non-suspicious at step 150 may comprise obtaining a score indicative of the likelihood that the lesion is in the second class (or conversely, a score indicative of the likelihood that the lesion is in the first class). The score may be e.g. a probability. In such embodiments, step 150 may further comprise classifying the lesion between a first clinical pathway class associated with a score above a first predetermined threshold, a second clinical pathway class associated with a score at or below the first predetermined threshold and above a second predetermined threshold, and a third clinical pathway class associated with a score at or below the second predetermined threshold. In some such embodiments, the first predetermined threshold is selected such that a maximum predetermined number or proportion of suspicious lesion in a training cohort is classified in the first clinical pathway class. In embodiments, the maximum predetermined number or proportion is zero, and / or the first predetermined threshold is the lowest predetermined threshold such that no suspicious lesion (or at most a predetermined percentage of suspicious lesions) in a training dataset is classified in the first clinical pathway class. When the classification comprise obtaining a score indicative of the likelihood that the lesion is in the first class, step 150 may comprise classifying the lesion between a first clinical pathway class associated with a score below a first predetermined threshold, a second clinical pathway class associated with a score at or above the first predetermined threshold and below a second predetermined threshold, and a third clinical pathway class associated with a score at or above the second predetermined threshold. In some such embodiments, the first predetermined threshold is selected such that a maximum predetermined number or proportion of suspicious lesion in a training cohort is classified in the first clinical pathway class. For example, the first predetermined threshold may be the highest threshold such that no suspicious lesion (or at most a predetermined percentage of suspicious lesions) in a training dataset is classified in the first clinical pathway class. At optional step 160, one or more image diagnosis plots can be obtained. Image diagnosis plots are plots that identify regions in an input image that were salient for the purpose of classification by a machine learning model as described herein. These may be referred to as saliency maps. The method may further comprise displaying the one or more saliency maps. The one or more saliency maps may comprise a saliency map computed using a gradient based method, such as e.g. Grad-CAM (Selvaraju et al., 2017; arXiv:1610.02391 [cs.CV]), and Grad-CAM++ (Chattopadhyay et al., 2018; arXiv:1710.11063 [cs.CV]), a saliency map be computed using an attribution propagation method, such as LRP (Binder et al. 2016; arXiv: 1604.00825 [cs.CV]), Contrastive-LRP (Gu et al., 2018; arXiv:1812.02100 [cs.CV]), and Softmax-Gradient-LRP (Iwana et al., 2019; arXiv:1908.04351 [cs.CV]), a saliency map may computed using a perturbation based method, such as LIME (Ribeiro et al., 2016; arXiv: 1602.04938 [cs.LG]), SHAP (Lundberg et al., 2017; arXiv:1705.07874 [cs.AI]), and RISE (Petsiuk et al., 2018; arXiv:1806.07421 [cs.CV]), a saliency map computed using the model’s activations, e.g. using TACV (Kim et al., 2018; arXiv:1711 .11279 [stat.ML]), or a saliency map computed using raw attention scores of a model comprising a multihead attention mechanism, for example using Attention-Last (Hohenstein and Beinborn 2021 ; arXiv:2106.03471 [cs.CL]) or Transformer Attribution (Chefer et al., 2021 ; arXiv:2012.09838 [cs.CV]). At optional step 170, the results of any one or more of the preceding steps may be provided to a user, e.g. through a user interface. At step 180, the patient may be selected for further diagnostic testing using an invasive test, using the results of step 150. For example, step 1780 may comprise selecting the patient classified in the first clinical pathway class for a first clinical pathway (e.g. discharge, follow up scheduled at a predetermined time (e.g. 3 to 6 months after assessment), and / or no teledermatology assessment needed), and / or selecting the patient classified in the second clinical pathway class for a second clinical pathway, e.g. health care practitioner reporting / teledetermatology assessement, and / or selecting the patient classified in the third clinical pathway class for a third clinical pathway, e.g. face-to-face assessment with or without biopsy. As another example, a patient with a lesion classified as suspicious may be selected for a biopsy. At step 190, the further diagnostic test may be performed and the patient may be treated in accordance with the results of the diagnostic test, such as e.g. by removal of the lesion when the lesion was identified as malignant or pre-malignant. An example of an automated assessment platform according to the present disclosure is illustrated on Figure 1c. In the illustrated embodiment, an image of a lesion (here illustrated as a dermoscopic image) is uploaded to a client management system (or any suitable database), together with clinical data (also referred to herein as lesion data or metadata). The data is fed to one or more machine learning models as described herein, which may be executed by a processor embodied as a server or, as illustrated, a cloud computer. The output of the machine learning models is used to classify the lesion between a first class leading to autonomous use (a), a second class leading to decision aid use (c) and a third class leading to autonomous use (b). Autonomous use (a) may be selected e.g. when the lesion is classified as non-suspicious with a confidence above a threshold (e.g. at or above 97%). Autonomous use (b) may be selected when the lesion is classified as suspicious with a confidence above a threshold. Use (c) may be selected when the lesion is classified as non-suspicious with a confidence below a threshold (e.g. below 97%). In autonomous use (a) (an example of a first clinical pathway), the patient is discharged, for example by issuing a letter indicating that the lesion was identified as non-suspicious, and recommending a follow up after a predetermined period of time, here illustrated as 3 months. In decision aid use (c) (an example of a second clinical pathway), the lesion image and clinical data is reviewed by a healthcare professional, who makes a final classification as non-suspicious (leading to discharge) or suspicious (leading to referral for a face-to-face assessment, for example including a letter and / or phone call to the patient to schedule such a consultation). The machine learning based assessment is therefore used as a decision aid in this pathway (an example of a second clinical pathway). In autonomous use (b) (an example of a third clinical pathway), the patient is referred to a face-to-face consultation, optionally with a biopsy. Referral may include a letter and a phone call.
[0104] Figure 2 shows a chart demonstrating the distribution of a skin lesion being pink in colour or not pink in colour for skin lesions determined to be non-suspicious (Figure 2a) and also for those determined to be suspicious (Figure 2b). This chart was produced following the analysis of tens of thousands of skin lesions which had been determined to be suspicious or non-suspicious while also recording the frequency of these skin lesions being pink or not pink at the same time. This same exercise was repeated to determine the frequency of several other risk factors in skin lesions determined to be suspicious and non-suspicious. The determination of whether a skin lesion is pink or not was found to be the most significant factor or all risk factors in then determining the likelihood of a skin lesion being suspicious or non-suspicious. As can be seen in Figure 2a about 83% if skin lesions defined as being non-pink were found in the non-suspicious group, with only 17% being defined as pink. For suspicious lesions, and as shown in Figure 2b, more lesions were pink in colour than were non-pink. Therefore, according to a first aspect of the invention there is provided a method for classifying skin lesions as suspicious or non-suspicious comprising; collecting and scoring a piece of data comprising a determination of whether a skin lesion is pink in colour or not pink to provide a lesion risk score to determine the likelihood of a skin lesion being suspicious or non-suspicious.
[0105] Figure 3 shows a schematic representation of an embodiment of the disclosure where one or more further pieces of data corresponding to risk factors for characterising skin lesions are scored in addition to lesion risk score obtained by determining whether a skin lesion is pink in colour or not pink in colour. In a first step the method according to this embodiment comprises collecting one or more further pieces of data corresponding to risk factors for characterising skin lesions (31), this is in addition to the first piece of data where a determination is made whether the skin lesion is pink in colour or not pink. Collectively the one or more further pieces of data and the first piece of data defining the skin lesion as pink or not pink are defined as metadata. In a second step the method comprises scoring the one or more further pieces of data and combining with the score corresponding to the first piece of data (32) to provide the lesion risk score which determines the likelihood of a skin lesion being suspicious (33) or non-suspicious (34).
[0106] Figure 4 shows a schematic representation of an embodiment where the method further comprises calculating an image risk score by detecting shapes, colours, textures or other features associated with the development of skin cancers from a captured image in addition to calculating the lesion risk score to determine an overall risk score. In a first step the method according to this embodiment comprises collecting and scoring metadata corresponding to risk factors for characterising skin lesions to provide a lesion risk score (41 a) while at the same time calculating an image risk score (41 b). In a second step the method comprises combining the lesion risk score with the image risk score to calculate an overall risk score (42) which determines the likelihood of a skin lesion being suspicious (43) or non-suspicious (44). In the embodiment shown in Figure 4, the lesion risk score is determined on the basis of scoring the metadata in relation to a determination of whether a skin lesion is pink in colour or not pink and also by scoring one or more further pieces of data corresponding to risk factors for characterising skin lesions. Alternative embodiments (not shown in Figure 4) are also possible where an image risk score is calculated and combined with a lesion risk score determined solely on the basis of scoring a first piece of data comprising a determination of whether a skin lesion is pink in colour or not pink.
[0107] Figure 5 shows a chart demonstrating analysis of Hu moment features which are indicators of the shapes in an image of a skin lesion that are associated with the development of skin cancers. This is in relation to embodiments where the method further comprises calculating an image risk score by detecting relevant features in an image of a skin lesion. As can be seen in Figure 5, Hu Moment 1 (51) is by far the most significant shape feature in an image for making a determination of whether a skin lesion in the image is suspicious or non-suspicious. In processing an image to determine the image risk score the recognition of a Hu Moment 1 shape feature in a captured image of a skin lesion will make it more likely that the skin lesion is correctly classified as either suspicious or non- suspicious. Hu-moment 2 (52) is the second most significant shape factor for making a correct determination as to classification.
[0108] Figure 6 shows an example of a captured image of a skin lesion. Figure 6a shows the original captured image and Figure 6b shows a version of the image in Figure 6a which has been subjected to a pre-processing step of standardising the shape and resolution of the original image according to an embodiment of the disclosure. The original image shown in Figure 6a has a size of 3.39MB and a resolution of 2848 v 4273. The image in Figure 6b has been subjected to a preprocessing step by resizing it to be 187KB in size and 1024 v 1024 in resolution (so it is square shaped). Preferably all images being processed to determine an image risk score will be subjected to this pre-processing step so as to be a standard size and resolution.
[0109] Figure 7 shows an example of a captured image of a skin lesion. Figure 7a shows the original captured image and Figure 7b shows a version of Figure 7a which has been subjected to a preprocessing step whereby the image is cleaned to remove features not relevant to determining a skin lesion being suspicious or non-suspicious according to an embodiment of the disclosure. I this case hairs have been removed from the image in Figure 7b as these are not relevantto determining whether a skin lesion is suspicious or non-suspicious.
[0110] Figure 8 shows an additional example of a captured image of a skin lesion, Figures 8a and 8b show the original captured image and Figure 8c shows a version of the image which has been subjected to a pre-processing step whereby the image is cleaned to remove features not relevant to determining a skin lesion being suspicious or non-suspicious according to an embodiment of the disclosure. By processing the image as described herein, it is possible to differentiate between background artifacts such as rulers and skin without lesions. Figure 8a shows the original image including the skin lesion, normal skin without a lesion and a ruler in the image. In Figure 8b the skin lesion has been detected but hairs still remain, and these are also irrelevant for determining whether a skin lesion is suspicious or non-suspicious. Figure 8c shows a skin lesion detected by the image processing that has also been subjected to a pre-processing step to remove features, in this case hairs, that are not relevant for determining whether a skin lesion is suspicious or not.
[0111] Figure 9 shows a schematic representation of an embodiment of the disclosure where different scaling techniques are employed where ensembling is utilised to enable models of different sizes to be combined. Figure 9a represents a schematic of the baseline image before subjected to any scaling techniques for the purposes of ensembling. Figure 9b is a schematic representation where the baseline of Figure 9a has been subsequently subjected to width scaling. Figure 9c is a schematic representation where the baseline of Figure 9a has been subsequently subjected to depth scaling. Figure 9d is a schematic representation where the baseline of Figure 9a has been subsequently subjected to resolution scaling. Finally figure 9e is a schematic representation where the baseline of Figure 9a has been subsequently subjected to compound scaling which includes a combination of width, depth and resolution scaling.
[0112] Figure 10 shows an example of a captured image of a skin lesion, Figure 10a shows the original captured image and Figure 10b shows a version of Figure 10a which has been subjected to processing by attention mechanism visualisation according to an embodiment of the invention.
[0113] Figure 11 shows an embodiment of a system that may be used according to embodiments of the present disclosure. The system comprises a computing device 100, which comprises a processor 101 and computer readable memory 102. In the embodiment shown, the computing device 100 also comprises a user interface 103, which is illustrated as a screen but may include any other means of conveying information to a user such as e.g. through audible or visual signals, by producing a report, etc. The computing device 100 may further comprise image acquisition means 104. The computing device 100 may be communicably connected, such as e.g. through a network 106, to a computing device 107 (comprising one or more processors 107a and one or more memories 107b) (also referred to as remote computing device) and / or database 108. The image acquisition means 104 are used to acquire image data showing a lesion. Metadata associated with the lesion (stored in memory 101 , 107a or database 108) may be analysed by the processor 101 and / or the processor 107b using instructions stored on memories 101 and / or 107b. For example, metadata may be analysed by the processor 102 using instructions previously received from remote computing device 107. Alternatively, metadata and optionally image data may be received at the device 100 and communicated to device 107 for analysis. Device 107 may be configured to send to device 100 one or more results of said analysis for providing to the user on user interface 103. Any other combination of locations of processing steps and storing of instructions may be used, such as e.g. all processing being done at computing device 100 or computing device 107. The computing device 100 may be a, smartphone, tablet, personal computer or other computing device. The computing device 100 may be associated with a user or healthcare professional. While the image acquisition means 104 are illustrated here as part of computing device 100, other configurations are possible such as e.g. where images are acquired by a separate image acquisition means and provided to the device 100 and / or the device 107 and / or database 108, e.g. through a wired or wireless connection. The remote computing device 107 may be a server, cloud computer, or any other computing device (whether localized or network-based). Communication between the computing device 100 and the remote computing device 107 and / or database 108 may be through a wired or wireless connection, and may occur over a local or public network 106 such as e.g. over the public internet. Any of the steps of any method described herein may be implemented by processor 101 and / or processor 107a, executing instructions stored on memory 102 and / or memory 107b. Database 108 may store one or more of: training data (e.g. metadata and / or image data), trained models, parameters, etc. This may be used by computing device 107 to train an machine learning model as described herein. The trained machine learning model may be used by computing device 107 and / or computing device 100 to implement methods of embodiments of the disclosure.
[0114] Due to significant advancements in Al technology, researchers now use dermoscopic images for skin cancer detection (Dildar et al., 2021). Al-based early skin cancer detection is an active area of research and has achieved the state of the art performance (Strzelecki et al., 2024). However, skin cancer detection solely based on patient metadata has been hardly explored. In the early 2000’s, skin cancer was usually diagnosed using standard techniques such as the ABODE rule (Papachristou and Bosanquet, 2020) and the 7PCL method (Walter et al., 2013) which showed moderate performances (sensitivity 73.3%, specificity 57.1 %). The ABODE checklist is used for assessing cancer-like characteristics such as lesion shape, asymmetry, border irregularity, colour variegation, lesion diameter, and lesion evolution over time. The 7PCL method considers seven risk factors, i.e., change of lesion size, shape, colour, lesion > 7 mm, inflamed, oozing, and itching, to identify patients with features suspicious of melanoma and to recommend urgent referral. On the other hand, the Williams method (Williams et al., 2011) is based on a validated scoring system that includes seven risk factors: patient age, gender, sunburn history, natural hair colour, density of freckles on arms, number of moles, and prior non-melanoma history. To the best of the inventors’ knowledge, in the literature, only the 7PCL and Williams scoring methods are purely based on patient metadata for suspicious skin lesion detection and these focus on suspicious melanoma. There is limited published work on the extensive use of patient metadata for detecting all skin cancer subtypes.
[0115] Ganster et al. (2001) and Burroni et al. (2004) explored the potential of introducing computer-based methods for skin cancer diagnosis. These researchers concentrated on analysing dermoscopic skin lesion images by first segmenting the lesion and then extracting handcrafted features, which were used to develop basic machine-learning models. Image segmentation was performed using a series of thresholding algorithms. Features such as shape, texture, and colour were then calculated from the segmented lesion. These features were used to train linear classifiers, k- nearest neighbour (KNN), and support vector machine (SVM) models. In recent years, open- source resources such as the International Skin Imaging Collaboration (ISIC) datasets and deep learning-based models, such as convolutional neural networks (CNNs), have facilitated assessing skin lesion images for skin cancer detection. Brinker et al. (2018) reviewed 13 publications that had implemented deep CNNs pre-trained on millions of images. One of the key papers from this review (Esteva et al., 2017) compared the performance of CNNs versus 21 board-certified dermatologists. They employed an lnceptionV3 network pre-trained on an extensive dataset containing more than 120,000 images. The dataset consisted of about 757 different diseases but was inferenced into benign and malignant lesions. They achieved a dermatologist-level performance with their deep neural networks. In a similar work to Esteva et al. (2017), Brinker et al. (2019) also compared the performance of a CNN with dermatologists of all levels of experience (Junior to chief physicians). They used the International Skin Imaging Collaboration (ISIC) 2016 dataset (Gutman et al., 2016, preprint) and implemented a pre-trained ResNet50 model to classify atypical nevi from melanomas, which outperformed 136 out of 157 dermatologists. Another study (Lopez et al., 2017) employed VGGNet on the ISIC 2016 dataset. They achieved 81 % accuracy with the fine-tuning method. In line with the above, Zhang et al. (2019) emphasised the intra-class variations and inter-class similarities between different types of lesions, which make it extremely difficult for machines to differentiate between them. At first, they attempted to narrow down the area of analysis by randomly cropping the images from a centre in different ratios. They also implemented several image augmentation techniques to increase the size of the dataset. They proposed a variation of the ResNet model by replacing the residual blocks with attention residual blocks. Compared to the top 6 neural network submissions to the 2017 ISIC challenge, their proposed model achieved the highest overall AUC of 91.70% (Jayapriya and Jacob, 2020). In recent years, researchers have also utilised vision transformers (ViTs) for skin cancer detection, achieving the state-of-the-art performance (Shajimon et al., 2023; Heroza et al., 2024). However, training ViTs on limited datasets remains an ongoing challenge.
[0116] Pacheco and Krohling (2020) took a different approach. Instead of using dermoscopic images, they utilised skin lesion images captured via a smartphone through an in-house developed mobile app. In this research, they highlighted the importance of the patient’s clinical information by comparing the performance of models based solely on images versus those combining images with patient’s clinical details. They collected eight clinical features: patient’s age, location of the lesion on the body, whether the lesion itches, bleeds or has bled, causes pain, has recently grown, has changed in pattern, and if it has an elevation. Several CNNs, including GoogleNet, ResNet, VGGNet, and MobileNet, were used to extract features from lesion images. For a comparative analysis, these extracted features were then combined with the clinical information and fed into a machine learning (ML) classifier for inference. They observed a 7% overall increase in balanced accuracy when clinical information was included in the analysis. Similarly, Yap et al. (2018) evaluated all combinations of dermoscopic, macroscopic, and clinical metadata (age, gender, and anatomic location) for binary melanoma and multiclass cancer detection and observed that combining all three achieved the highest overall AUC of 88.80%. However, another study (Ha et al., 2020, preprint) included more limited patient metadata, such as age, gender and lesion location, and observed that combining metadata along with lesion images did not improve their model performance. Roffman et al. (2018) evaluated the use of the clinical information alone, which was based on the National Health Interview Survey (NHIS) data from 450,000 patients between 1997 and 2015, to classify non-melanoma skin cancers against the “never-cancer” skin diseases. This dataset includes patient details such as patient age, gender, body mass index (BMI), ethnicity, hypertension, heart disease, diabetes status, and several lesion characteristics from over 450,000 patients / skin lesions. They employed a basic feed-forward neural network and achieved an AUC of 81 %, with a sensitivity of 86.2% and a specificity of 62.7%. The studies by Pacheco and Krohling (2020), Yap et al. (2018) and Roffman et al. (2018) included a limited set of metadata (age, gender, and anatomic location).
[0117] The following is presented by way of example and is not to be construed as a limitation to the scope of the claims.
[0118] EXAMPLES
[0119] Example 1 - Leveraging Al and patient metadata to develop a novel risk score for skin cancer detection
[0120] The majority of skin cancer classification research has focused on image data and DL models, with relatively little work being done on skin cancer detection using metadata alone. Moreover, the performance of the standard methods used is not satisfactory as reflected by their low sensitivity scores, such as the 7PCL (sensitivity 68 ± 2.10%) and the Williams score (sensitivity 67 ± 1 .90%). In the work described in the present example, the inventors set out to fill this gap and to further improve skin cancer detection performance through utilising patient metadata alone. They investigated artificial intelligence (Al) models that utilise patient metadata consisting of 23 attributes for suspicious skin lesion detection. They also identified a new set of most important risk factors, namely “C4C risk factors”, which is not just for melanoma, but for all types of skin cancer. The performance of the C4C risk factors for suspicious skin lesion detection was compared to that of the 7PCL and the Williams risk factors that predict the lifetime risk of melanoma.
[0121] Specifically, the inventors proposed an Al framework that ensembles five machine learning models for skin lesion classification into suspicious and non-suspicious classes. Using this Al framework, C4C risk factors were identified from a pool of 22 meta-features responsible for the development of all skin cancer subtypes (melanoma, squamous cell carcinoma (SCC), basal cell carcinoma (BCC)): lesion pink, lesion size, lesion colour, lesion inflamed, lesion shape, lesion age, and natural hair colour. This set of risk factors achieved a sensitivity of 80.46 ± 2.50% and a specificity of 62.09 ± 1.90% in detecting suspicious skin lesions when evaluated using the metadata of 53,601 skin lesions collected from different skin cancer diagnostic clinics across the UK, significantly outperforming the 7PCL-based method (sensitivity 68.09 ± 2.10%, specificity 61 .07 ± 0.90%) and the Williams risk factors (sensitivity 66.32 ± 1 .90%, specificity 61 .71 ± 0.6%). Furthermore, through weighting the seven new risk factors, with weights determined by intelligent data analysis, they came up with a new risk score, namely “C4C risk score”, which alone achieved a sensitivity of 76.09 ± 1 .20% and a specificity of 61 .71 ± 0.50%, significantly outperforming the 7PCL-based risk score (sensitivity 73.91 ± 1.10%, specificity 49.49 ± 0.50%) and the Williams risk score (sensitivity 60.68 ± 1 .30%, specificity 60.87 ± 0.80%) in suspicious skin lesion detection.
[0122] Finally, the C4C risk factors were fused with the 7PCL and Williams risk factors to find the best feature combination. Fusing the C4C risk factors with the 7PCL and Williams risk factors achieved the best performance, with a sensitivity of 85.24 ± 2.20% and a specificity of 61.12 ± 0.90%. The new set of skin cancer risk factors has the potential to be used to modify current skin cancer referral guidelines for all skin cancer subtypes, including melanoma.
[0123] The work in this example made the following major contributions:
[0124] 1. Collection of metadata of 53,601 skin lesions from 25,105 patients across a national network of private UK skin diagnostic clinics.
[0125] 2. Identification of a new list of risk factors named “C4C risk factors” from a pool of 22 metafeatures responsible for the development of all skin cancer subtypes (melanoma, SCC, BCC) through an ensemble of five Al models, which significantly outperforms the existing 7PCL and Williams methods with a balanced accuracy of 71 .27 ± 1 .10% and sensitivity of 80.46 ± 2.50%.
[0126] 3. Proposal of a new skin cancer risk score named “C4C risk score”, which is based on the weighting of “C4C risk factors” with weights determined by intelligent data analysis. Using the C4C risk score alone achieves 68.90% balanced accuracy and 76.09 ± 1.20% sensitivity in classifying suspicious and non-suspicious skin lesions, significantly higher than the exiting 7PCL risk score and Williams risk score.
[0127] 4. Fusion of the “C4C risk factors” with the 7PCL and Williams risk factors to find the best feature combination, which achieves the highest overall performance with a balanced accuracy of 73.18 ± 2.10% and a sensitivity of 85.24 ± 2.20%.
[0128] Data and Methods
[0129] Metadata Collection. Anonymised clinical metadata of 53,601 skin lesions from 25,105 patients were collected and analysed. Clinical data was collected according to protocol by centrally trained nurses, with central reporting of all skin lesions by a central team of skin cancer specialists. The lesions were eligible to be assessed if they were: (1) located in adults > 18 years, (2) between 1 and 3 suspicious lesions which are not larger than the dermatoscopic lens (< 15 mm). For each lesion, 23 meta-features were included, as shown in Table 5: 7PCL (lesion size, lesion shape, lesion colour, lesion > 7 mm, lesion inflamed, lesion oozing, and lesion itching), lesion score based on the 7PCL, Williams risk factors (patient age, patient gender, hair colour, moles, sunburn, freckles, prior non-melanoma skin cancer), the overall Williams score, Williams group, prior melanoma, prior family history of skin cancer, lesion location, lesion age and whether this was a predominantly non-pigmented pink lesion. All the features except for lesion rating (which was used as ground truth for classification) were considered candidate features for identifying the new risk factor sets to classify skin lesions into suspicious versus non-suspicious categories.
[0130] The meta-features listed in Table 5 are self-explanatory except for the lesion location feature that comprises seven values according to the anatomic location of the lesion, such as (1) Head and Neck, (2) Trunk waist up (front or back), (3) Groin / Buttocks / Genitals, (4) Hand, (5) Foot, (6) Left / Right Leg (ankle up), and (7) Left / Right Arm (wrist up). The Williams score is calculated based on the method explained in Williams et al. (2011) and summarised in Table 6, where age, gender, sunburns, natural hair colour, the density of freckles on arms, number of moles, and prior nonmelanoma history features are included to estimate the final Williams score. The 7PCL lesion score is estimated based on a weighted 7PCL as mentioned in Walter et al. (2013) using the following equation: where M is a set of major lesion features (lesion size, lesion shape, lesion colour) and N is a set of minor features (lesion>7mm, lesion inflamed, lesion oozing, lesion itching).
[0131] Table 5. List of 23 clinical meta-features: a total of 53,601 skin lesions metadata from 25,105 patients were collected.
[0132] Table 6. Calculation of Williams score based on the risk factors described in Williams et al. (2011).
[0133] The 7PCL was first formulated by Mackie et al. (1986). They used seven equally weighted lesion characteristics (change in size, shape, colour, inflammation, oozing, itching, and diameter >7 mm) to prioritise pigmented skin lesions for urgent referral. Walter et al. (2013) achieved better results with a revised version, which separated lesion features into two groups: (1) major features (change in size, shape, and colour), each having a weight of 2, and (2) minor features (inflammation, oozing, itching, and diameter >7 mm) with a weight of 1 , as shown in Eq. (1). Consequently, lesions with a lesion score >3 were sent for a specialist opinion. Finally, the target variable, with a lesion rating as suspicious or non-suspicious, was assessed by the in-house skin cancer specialists. The experts classified pigmented lesions with atypical features in size, shape, colour, or dermatoscopic appearance of melanoma as suspicious. Furthermore, skin lesions suspicious of either BCC, SCC or potentially pre-malignant Actinic Keratoses were also rated as suspicious. Biopsy results were also available. However, a limited number of lesions went for biopsy (only 10% of lesions undergo biopsy). If biopsy results were used as the target variable, the data size would have reduced significantly (90%). As a result, the lesion rating was used as the target variable rather than the biopsy results. The ultimate goal was to use Al as a clinical decision aid for the classification of suspicious or non-suspicious skin lesions during teledermatology triage.
[0134] Statistical Data Analysis. Patient metadata consists of meta-features with categorical / text values, such as patient gender-taking values of male and female. There was a need to convert these categorical values into numerical as most Al models work well with numerical data rather than categorical data. All the non-numerical meta-features were encoded to convert them into numerical features using a one-hot encoding approach, e.g. lesion shape (change in shape yes / no) was encoded as ‘1 ’ for yes and ‘0’ for no, gender was encoded as ‘0’ for female and ‘1 ’ for male, and lesion rating was encoded as ‘o’ for green / not suspicious and ‘1 ’ for red / suspicious.
[0135] An effort was made to analyse the collected meta-features through explanatory data analysis. All the meta-features were examined and the findings highlighted using lesion pink and age risk factors for illustrative purposes. First, the inventors analysed the lesion score meta-feature (range 0-10) using a bar plot and showed a difference between suspicious and non-suspicious cases using a statistical t-test. Almost 50% of the lesions with a lesion score of 10 belonged to the suspicious category. In contrast, only 5% of lesions with a score of 0 belonged to the suspicious group, as highlighted in Fig. 12. Therefore, it can be inferred that lesion score and outcome variable, i.e., lesion rating, are highly correlated and the probability of a skin lesion being suspicious is likely to increase for higher values of the lesion score (p value < 0.01). For a Williams score between 56 and 61 , around 60% of cases belonged to the suspicious category, whereas only around 7% of cases belonged to the suspicious group with Williams scores between 0-6, as shown using a bar plot in Fig. 12. It is highly likely that the higher the Williams score, the higher the chance that the skin lesion is suspicious.
[0136] The inventors analysed another meta-feature of potential importance: “lesion pink”. Its statistics are shown using a bar plot in Fig. 2. About 83% of skin lesions with lesion pink value ‘no’ belonged to the non-suspicious group as compared to only about 17% skin lesions with lesion pink value ‘yes’. Therefore, it can be ascertained that there is a low probability of being a suspicious lesion if the lesion pink value is ‘no’. Conversely, there is a comparatively higher chance (54%) for a skin lesion to be suspicious if it is a pink lesion, as shown in Fig. 2. Another meta-feature analysed was “patient age”. Its distributions for suspicious and non-suspicious cases are compared in Fig. 13. The average age of patients with suspicious skin lesions was 52, markedly higher than the average age of 41 for patients with non-suspicious skin lesions (p value < 0.01).
[0137] Identification of New Risk Factors. For achieving reliable results, the inventors adopted an ensemble of five machine learning (ML) models for identifying a new set of effective risk factors from a pool of 22 skin lesion meta-features, the first 22 attributes as shown in Table 5. An overview of the Al framework used for identifying a new set of risk factors for skin lesion classification into suspicious or non-suspicious class is shown in Fig. 14. The motivation behind classifying skin lesions into suspicious versus non-suspicious categories instead of traditional melanoma versus benign classes has emerged from the fact that early suspicious skin lesion detection could was used for feature subset generation. Four potential feature subsets of different sizes were generated: Set1 , Set2, Set3, and Set4, with 7, 10, 15, and 20 meta-features, respectively. Firstly, by applying combination theory, various combinations of 7 meta-features out of 22 were generated, resulting in a total of 170,544 combinations for Set1. Feature subsets generation was repeated using the combination theory for Set2, Set3, and Set4, respectively. The best meta-feature combinations for these four sets were selected based on their overall performances of the ML models. Each model used is described below.
[0138] Naive Bayes (NB) classifier. This is a classifier based on the Bayes theorem. The class with the highest probability is considered as the predicted class for the given data tuple. NB classifiers assume that all attributes are conditionally independent of the given class label. The goal of this classifier is to learn a representative function from a given labelled training dataset. The conditional probability p(Y|X) of the target variable Y is calculated as follows: pen Perri p(T|X) = (2) where p(Y) is the prior probability of class Y, p(X|Y) is the conditional probability of data X given a particular class, and p(X) is the evidence or probability of data X regardless of its target class (suspicious or non-suspicious in the present example). Support vector machine (SVM). SVM finds a hyperplane to maximise the margin between the groups by utilising the Lagrangian optimisation technique . One of the fundamental advantages of SVM is that if the data is linearly separable, then there is a unique global maximum value of the margin. In cases of non-linear distribution of the data, where a hyperplane cannot separate the data, SVM uses a kernel function that transforms the data into a higher dimensional feature space where the data’s linear separation is possible.
[0139] Logistic regression (LR). This is a statistical method, in which log-odds of the probability of an event are linear combinations of independent variables. Although the model outputs the probability of an event, it can be used for the classification task by applying a threshold. The logistic regression approach’s outcome is binary, such as positive or 1 (suspicious) and negative or 0 (nonsuspicious). Here, the LR was implemented to represent a relationship (function) between the meta-features and outcome variables by finding the best descriptive fitting model, which directly maps meta-feature input to output target variable as follows: where p(Y|X) is the probability of a skin lesion being suspicious (Y = 1) given a meta-feature vector
[0140] X, po is the intercept, pi, ..., pp are the coefficients, and P is the total number of meta-features.
[0141] Random forest. It employs an ensembling technique that generates multiple random decision trees and combines the outcomes of the decision trees given a test sample based on majority voting or averaging. In the present example, the decision trees were built upon a bootstrap sample of the data. RF adds more randomness in selecting a subset of predictors compared to a standard decision tree, where each node is split using the best variable selected based on a node splitting criterion — gini or entropy. This randomness in selecting features makes the RF classifier more accurate and robust compared to other classifiers such as SVM, discriminative analysis, and neural networks. The RF model was optimised by finding the best hyper-parameters (number of trees, 500, max depth, 40, splitting criterion, gini, bootstrap, true) for classifying skin lesions into suspicious and non-suspicious categories.
[0142] Multi-layer perceptron (MLP). MLPs are simple neural networks consisting of an input and output layer and a hidden layer (in most cases) to transform input into some form of internal representation that the next layer can use. When training the model “backpropagation” is used, which allows changing the weights of the hidden layer if there are any errors. The fundamental advantage of MLP is that it does not require in-depth knowledge about the relationship between meta-feature input and output target variables. Instead, it tries to recognise patterns in the dataset and store those patterns in the form of weights for later use for the test cases. Here, an MLP with three hidden layers (32, 16, 8 neurons) was adopted, with a rectified linear activation function (ReLU), and trained using the adaptive moment estimation (Adam) optimizer. Majority voting was adopted for decision-making by combining the outcomes of NB, LR, SVM, RF, and MLP. In the stacking approach, NB, LR, SVM, and RF were stacked as feature extractors and MLP as meta-learners to classify input metadata into suspicious and non-suspicious classes. In the majority voting ensemble models, each model NB, LR, SVM, RF, and MLP determines whether input metadata belongs to suspicious or non-suspicious. The final outcome is deduced based on majority voting. In the stacking method, outcomes of NB, LR, SVM, and RF are fed as input to the MLP model to get the final outcome. As outcomes of NB, LR, SVM, and RF are used as input of MLP, the inventors may argue that these models work as feature extractors, those features are used as input to the final MLP model to get outcome decision (suspicious or not).
[0143] A New Skin Cancer Risk Score. An overview of the proposed Al framework for deriving a new skin cancer risk score for suspicious versus non-suspicious skin lesion detection is illustrated in Fig. 15. A LR model was used to rank the identified new risk factors. Feature ranking helps to find the most important features and provides an interpretation of the Al model on why certain features are more important than others and therefore can discard those features having the lowest / no correlation with the outcome variables. Furthermore, it can facilitate the reduction of the data collection burden in the future (e.g., instead of collecting 22 meta-features, data collection can be reduced to 7 meta-features only), as well as reduce model complexity and training time. In the present example, a set of N training samples were used, with each sample represented in the form of (X, Y), where X is a meta-feature vector and Y is the corresponding output (i.e., the target class). A classification rule is formulated using the training data to assign a class label, suspicious (Ytest = 1) or non-suspicious (Ytest = 0), to a new test input Xtest, by minimising the probability of error. Typically, a feature is deemed relevant if it aids in distinguishing between classes and is not redundant with other relevant features. As shown in Eq. (3), the LR model produces posterior probabilities through a linear function of elements in the meta-feature vector X, ensuring that the probabilities sum to one and remain within the range of [0, 1]. The LR comes with a set of diagnostic tools that allows us to quantify the goodness-of-fit of the proposed model and select the features accordingly. The performance of the model is evaluated based on the maximum value of the loglikelihood (LL) achieved for each feature from X using the deviance D defined below:
[0144] D = —2(LL of the current model — LL of the saturated model) (4)
[0145] The saturated model is the one with the number of parameters equal to the sample size, the likelihood of which is one. Low deviance values indicate a good fit, or equivalently, a high predictive value of the corresponding features. The deviance is useful for comparing two models with different numbers of features. The reduction in the deviance by adding a new feature is identical to the likelihood-ratio statistic, which has a Chi-squared (x2) distribution, provided that the sample size N is large. Hence, the likelihood-ratio test can be used to include the features sequentially in a forward-selection procedure. If the difference in the deviance of the models before and after adding a new feature is at or above the critical value, then the new feature is significant in predicting the target class, otherwise not.
[0146] Although four new sets of risk factors were investigated, where Set1 , Set2, Set3, and Set4 consist of 7, 10, 15, and 20 risk factors, respectively, the inventors wanted to benchmark the Al models’ performance using their proposed C4C risk score along with the 7PCL-based lesion score and Williams score. Therefore, to develop the C4C risk score, only the 7 risk factors from Set1 were used. The inventors believe that the use of 7 risk factors to develop their proposed C4C risk score facilitates fair comparison with the 7PCL-based lesion score (included 7 risk factors), and Williams score (included 7 risk factors). The 7 risk factors from Set1 were ranked based on their coefficient values obtained using the LR model. The highest-ranked risk factor was assigned the highest weight, and consequently, lower-ranked risk factors were assigned lower weights to calculate the weighted sum of the seven risk factors, leading to the proposed C4C risk score.
[0147] Data Split and Evaluation Metrics. The metadata for 53,601 skin lesions was divided into training and test datasets, comprising 80% and 20% of the data, respectively. To prevent data leakage during the split, all metadata corresponding to each patient was exclusively assigned to either the training or testing dataset. During the training, a tenfold cross-validation (CV) method was employed to construct the models. These models were optimised by adjusting hyperparameters, and the most effective models were chosen based on their tenfold CV results using the training data. Subsequently, the selected models were assessed using the test dataset. Using a tenfold CV for model development on the training dataset only while keeping the test dataset completely independent from model development reduces the risk of overfitting. The performance of the developed Al framework was assessed using the following evaluation metrics:
[0148] TP
[0149] Sensitivity (Sen) =pp pN(5)
[0150] TN
[0151] Specificity (Spc) =Fp + TN(6)
[0152] AUC = p (Score (TP) > Score (TN) ) (8) where TP, TN, FP, FN refer to true positive (suspicious classified as suspicious), true negative (non-suspicious classified as non-suspicious), false positive (non-suspicious misclassified as suspicious), and false negative (suspicious misclassified as non-suspicious) instances, respectively. The area under the curve (AUC) of a classifier is the probability that a randomly chosen TP case will be ranked higher than a randomly chosen TN case. Results and Discussion
[0153] Firstly, by applying combination theory, the inventors generated various combinations of 7 metafeatures out of 22, resulting in a total of 170,544 combinations. Among these, the 7 meta-features listed as Set 1 in Table 7 exhibited the highest balanced accuracy and sensitivity, which are lesion colour, size, shape, inflammation, natural hair colour, lesion age, and pinkness. Similarly, the inventors replicated the experiment for identifying risk factor sets 2, 3, and 4. Table 7 shows the best meta-features in sets 2, 3, and 4, comprising 10, 15, and 20 features, respectively. The performances of the identified sets of risk factors for skin lesion classification (suspicious vs. non- suspicious) were evaluated by training and testing the developed ML models. The performances of the four sets of skin cancer risk factors are also presented in Table 7. Using risk factor set1 with
[0154] 7 meta-features only, the best-ensembled ML model achieved a balanced accuracy of 71.27 ± 1 .10% , a sensitivity of 80.46 ± 2.50% , and a specificity of 62.09 ± 1 .90% . The balanced accuracy and sensitivity were notably improved to 73.01 % and 84.51 %, respectively, when the risk factor set3 was used. However, using the larger risk factor set4 with further meta-features added did not enhance the model’s performance, as shown in Table 7.
[0155] Table 7. New risk factor sets identified through applying Al-based model fusion.
[0156] An attempt was made to benchmark the newly proposed C4C risk factors-classifier with the 7PCL and Williams risk factors. The performance of the C4C risk factors for skin lesion classification (ensembled classifier using the indicated features) into suspicious or not suspicious is presented and compared with that of the 7PCL and Williams risk factors in Table 8, where the inventors showed that their approach outperformed the 7PCL and Williams method in terms of balanced accuracy and sensitivity scores (p value < 0.01). It is interesting to note that lesion age, lesion pink and hair colour are risk factors for all skin cancer subtypes, which are not among the 7PCL since this test is only relevant for melanoma, the pigmented type of skin cancer. Finally, the performance gain from feature fusion was evaluated (i.e. ensembled classifier models using the indicated features comprising features from set 1 in Table 7 and additional features from the 7PCL and Williams methods), with results also shown in Table 8. It is noteworthy that the highest performance with a balanced accuracy of 73.18 ± 2.10% , a sensitivity of 85.24 ± 2.20% and a specificity of 61 .12 ± 0.90% was achieved when the C4C risk factors (set 1 in Table 7) were fused with the 7PCL and Williams risk factors, which forms a set of 18 meta-features. The inventors investigated the performance gain for each feature from the 7PCL and Williams risk factors. The feature subset generator described above (i.e. using the combination formula to create all possible combinations from the set of available features) was used to find the optimal feature set, which shortlisted 11 external risk factors (patient age, patient gender, Williams score, Williams group, sunburn, freckles, moles, lesion body, lesion itch, lesion > 7 mm, and lesion oozing). Fusing them with the C4C risk factors can achieve better performance with the developed ML models, but it is at the price of collecting much more metadata.
[0157] Finally, the performance of the C4C risk score is compared with that of the 7PCL lesion score and Williams score in Table 9. The C4C risk score is calculated as: 10 * Lesion Pink + 4 * Lesion Inflamed + 2.5 * Lesion Shape + 2.5 * Lesion Color + Lesion Size + Lesion Age + 0.25 * Hair Colour, where lesion pink takes the value 0 (no) or 1 (yes), lesion inflamed takes the value 0 (no) or 1 (yes), lesion shape takes the value 0 (no change in shape) or 1 (change in shape), Lesion size takes the value 0 (no change in size) or 1 (change in size), Lesion age takes the value 0 (lesion present for 6 months or less) or 1 (lesion present for more than 6 months), and hair colour takes the value of 1 for black hair, 2 for red, 3 for blonde hair, and 4 for brown hair. As explained above, the inventors fitted a logistic regression model using these 7 predictive features to find the weights of those risk factors and rank them based on their corresponding weights. The weights and ranking from the model were then translated to those simpler weights based on expert input. The inventors then tested all possible values in the range of this score (0 to 65) to find the optimal threshold value that separates suspicious from non-suspicious lesions solely based on the C4C risk score. They identified an optimal threshold as >12. Thus, a lesion with a C4C risk score calculated as above that is at or above 12 is considered suspicious, while a lesion with a C4C risk score below 12 is considered non suspicious. With this approach, the C4C risk score alone achieved a sensitivity of 76.09 ± 1.20% and a specificity of 61.71 ± 0.6% , significantly outperforming the 7PCL-based risk score (sensitivity 73.91 ± 1.10%, specificity 49.49 ± 0.50% ) and Williams risk score (sensitivity 60.68 ± 2.10%, specificity 60.87 ± 0.80%). The inventors believe the C4C risk score has significant potential as it doesn’t require logistic regression or any other complex model to separate lesions into suspicious or non-suspicious groups (e.g. a simple calculator is sufficient). The use of the score can simply identify suspicious lesions using only seven risk factor values, reducing data collection burden, and model training time and resources.
[0158] It is noted that the study by Pacheco and Krohling (2020) included patient metadata such as patient age, gender, lesion location, lesion bleeding, and lesion pain along with patient skin images and they reported a 7% performance improvement due to the addition of metadata to image assessment. However, they did not report the contribution of metadata alone in detecting skin cancer.
[0159] Table 8. Performance gain comparison through fusing the 7PCL, Williams and C4C risk factors (i.e. combining each of the individual sets of risk features). Significant values are in bold.
[0160] Table 9. Performance comparison of the new C4C risk score with the 7PCL-based lesion score and Williams score. Significant values are in bold.
[0161] To the best of the inventors’ knowledge, the majority of the previous studies used image data only and there is limited work done on using patient metadata to classify lesions for skin cancer detection. Therefore, they have developed an Al framework solely based on metadata and observed that it can separate suspicious skin lesions from non-suspicious ones with a high sensitivity, which has the potential to support current skin cancer assessment when considered alongside image data. In the future, patients with both metadata and images classified as non- suspicious could be reassured without referral to a specialist clinic. Furthermore, the C4C risk score can be used as a decision-aid by telemedicine reporters to help with final lesion classification that is equivocal after image classification alone. This has the potential to reduce the number of referrals to a specialist clinic for possible biopsy and help reduce the waiting times for skin cancer diagnosis. With a reduction in patient referrals for possible biopsies, waiting times for skin cancer diagnosis and treatment will be shortened, resulting in improved outcomes. In the present example, the inventors have developed an Al framework based on patient metadata for skin lesion classification, which outperformed the existing 7PCL and Williams methods. This example also contributed to high-quality data collection followed by the identification of a subset of meta-features highly relevant to skin cancer diagnosis.
[0162] Example 2 - Advancing skin cancer detection through fusion of patient metadata and skin lesion images
[0163] To reduce waiting time and to make a faster decision, there is a need to develop automated methods that can be used to classify whether a skin lesion is suspicious or non-suspicious during teledermatology triage. In this example, an Al framework is proposed that utilises patient metadata together with image data to classify skin lesions into suspicious or non-suspicious categories. The present example describes data collection, data pre-processing, feature extraction and selection, Al model for classification, and explanation followed by evaluation and benchmarking. To evaluate the proposed approach, 79,246 skin lesion images along with their 22 meta-features (e.g. lesion size, lesion colour, lesion shape, patient age, and gender) were collected from 19,295 patients who attended a network of private skin cancer diagnostic centres across the UK during 2015-2022.
[0164] Three separate models for skin lesion classification were developed: 1) an Al model using only metadata that achieved 85.24% sensitivity and 61.12% specificity; b) an Al model using only images that achieved 99.72% sensitivity and 63.22% specificity; and 3) a fused model based on both metadata and images that achieved 99.66% sensitivity and 74.45% specificity. The decisions of the developed Al models were then fused through a majority voting technique, which achieved a sensitivity of 99.50% and a specificity of 82.45%, significantly outperforming the state-of-the-art methods that rely solely on image data. Results were benchmarked with the state-of-the-art, showing superior performance of the proposed models in terms of sensitivity and specificity. Furthermore, the inventors add a post-processing step in their Al framework to enhance decisionmaking transparency by generating heat maps of skin lesions, by implementing a soft-attention module that provides crucial Al explainability to support healthcare professionals in their decisionmaking process. The developed Al framework has great potential for the detection of suspicious skin lesions. With a reduction in patient referrals for possible biopsies, waiting times for skin cancer diagnosis and treatment will be shortened, resulting in improved outcomes.
[0165] In the work in Example 1 , the inventors investigated the potential of patient metadata in skin cancer detection. They identified a new list of seven risk factors named “Check4Cancer (C4C) risk factors” from a pool of 22 meta-features responsible for the development of all skin cancer subtypes (melanoma, SCC, BCC) through an ensemble of five Al models, which significantly outperforms the existing 7PCL and Williams methods with a balanced accuracy of 71.27% and sensitivity of 80.46%. They also proposed a new skin cancer risk score named “C4C risk score”, which is based on the weighting of “C4C risk factors” with weights determined by intelligent data analysis. Using the “C4C risk score” alone can achieve 68.90% balanced accuracy and 76.09% sensitivity in classifying suspicious and non-suspicious skin lesions, significant! higher than the exiting 7PCL risk score and Williams risk score. Furthermore, the inventors fused the “C4C risk factors” with the 7PCL and Williams risk factors to find the best feature combination, which achieves the highest overall performance with a balanced accuracy of 73.18% and a sensitivity of 85.24%. In this example, the inventors investigate the fusion of the newly identified skin risk factors and weighted risk score together with lesion images using deep learning models to further boost the performance of skin cancer detection.
[0166] In the literature, the majority of previous skin cancer classification research has focused solely on image data alongside DL models, with limited exploration of detecting skin cancer through the fusion of patient metadata and image data. While studies such as Pacheco et al. 2020, Yap et al. 2018 and Roffman et al. 2018 incorporated a limited set of patient metadata (age, gender, and anatomical location), they did not highlight the significance of combining patient metadata with image data for enhancing Al model performance. In an attempt to fill the mentioned research gap, and to further improve the skin cancer detection performance through utilising patient metadata and image data altogether, the inventors devised an Al framework for classifying fused metadata and skin lesion images into suspicious or non-suspicious categories. The work in the present example makes the following major contributions:
[0167] 1 . Collection and evaluation of 79,246 skin lesion images from 19,295 patients across a national network of private UK skin diagnostic clinics. For each lesion, the inventors collected 22 metafeatures and two types of images using dermoscopic camera (DER images) and digital singlelens reflex camera (DSLR) camera (SLR images). The inventors identified a new set of seven primary risk factors, including lesion pinkness, lesion size, lesion colour, lesion inflammation lesion shape, lesion age, and natural hair colour. These factors are pertinent not only for melanoma but for all types of skin cancer.
[0168] 2. A multi-modal Al framework, including deep learning models, which combines patient metadata with skin images for the classification of concerning skin lesions. Three types of Al models have been developed by varying input data types. Ultimately, the model utilising both metadata and images achieved the highest performances with 99.66% sensitivity and 74.45% specificity, significantly higher than the Al model performance using image data only (99.72% sensitivity and 63.22% specificity) Furthermore, we fused outcome decisions of the developed Al models through a majority voting technique, which achieved a sensitivity of 99.50% and a specificity of 82.45%, significantly outperforming the state-of-the-art methods that rely solely on image data.
[0169] 3. A post-processing module comprising gradient class activation map (Grad-CAM) and soft attention technique is added into the proposed Al framework to enhance decision-making transparency by generating heat maps of skin lesions, providing crucial Al explainability to support healthcare professionals in their decision-making process.
[0170] The present example is an extension of Example 1 through the fusion of the newly identified skin risk factors and weighted risk score together with lesion images using deep learning models to further boost the performance of the inventors’ Al model for skin cancer detection.
[0171] Data and Methods
[0172] Data Collection. 79,246 images were collected from 39,623 skin lesions belonging to 19,295 patients who attended Check4Cancer (C4C)’s private skin cancer diagnosis clinics between 2015 and 2022 as summarised in Table 10. For each skin lesion, two types of images were collected, one using a dermoscopic camera (39,623 images - referred to as “DER” below) and another using a DSLR camera (39,623 images - referred to as “SLR” below).
[0173] Table 10. Skin lesion image data collection summary.
[0174] Each lesion was visually assessed by the in-house skin cancer specialists during teledermatology triage. The experts classified pigmented lesions with atypical features in size, shape, colour, or dermoscopic appearance of melanoma as suspicious. Furthermore, skin lesions suspicious of either BCC, SCC, potentially pre-malignant Actinic Keratoses, Bowen’s disease, or in-situ carcinoma were also rated as suspicious. The experts’ classification (suspicious or non-suspicious) was used as ground truth while developing Al models. There were 1 ,546 Melanoma cases, 4,420 BCC cases, 530 SCC cases, and 4,762 other suspicious cases belonging to the “suspicious” category. Conversely, a significantly higher number of images (43,987) with a subcategory reported as Mole belong to the non-suspicious group. Other major subcategories comprise naevus (1 ,829 cases), actinic (1 ,918 cases) and Seborroehic keratosis (8,778 cases).
[0175] Some examples of the skin lesion images belonging to suspicious and non-suspicious categories are shown in Fig. 16, whereas pairs of skin lesion images captured by dermoscopic and DSLR cameras are illustrated in Fig. 17. There were 67,988 non-suspicious images (out of 79,246) and 11 ,258 suspicious images, which were collected from a diverse location of lesions to help build a comprehensive Al model. For each lesion, the inventors also collected 22 meta-features, with further details described in Example 1 . The skin types info of this dataset were not recorded during data collection. The dataset included UK patients which comprises Fitzpatrick skin types l-VI as mentioned in a UK-based similar study in Thomas et al. 2023. Although the dermoscopic camera provides magnified and better-quality images than the DSLR camera, both camera images were used to build a robust Al model that is capable of classifying both types of images when implemented in real-world applications. Five images were discarded due to inadequate lighting (2), image corruption (2) and low resolution (1).
[0176] Data Pre-Processing - Image reshaping. The average size of the collected raw skin lesion images was approximately 5MB with a resolution of 3000 x 4000. To build the Al model, the images were resized and converted to a square shape with a resolution of 1024 x 1024. Reshaping images in this way allows fortraining an Al model with lower memory requirements and reduced training time compared to using raw images alone. Additionally, it is recommended to provide square-shaped images as input to the Al model, as this enables more efficient convolutional computations compared to non-square images (Simonyan, 2014, preprint). However, this process distorts the original shape of the lesion, which is a crucial feature for accurately classifying lesions into suspicious or non-suspicious categories (as illustrated on Fig. 6). Therefore, two different image reshaping approaches were adapted and tested, as shown in Fig. 18: a) reshaping with skin tone colours where pixel values of padding location were replaced with mode pixel values of the corresponding images; b) reshaping with black padding, where pixel values of padding location were replaced with (0,0,0) values. These padding methods were compared to identify the right approach that yielded the best Al model performance, with the aim of achieving optimal performance in skin cancer detection. Method (a) (reshaping with skin tone colours) performed better than both (b) (reshaping with black padding) and lesion segmentation, and was selected for the results shown.
[0177] Data pre-processing - Hair removal. Some artefacts, such as hairs and rulers, are present in the skin lesion images (as shown in Fig. 17), which the inventors believed could potentially reduce the Al model’s performance. To address this, they implemented the hair removal method used by Bardou et al. (2022). While this method effectively removes hairs from the skin lesion images, it does come with a trade-off, as the quality of the reconstructed images appears to diminish upon visual inspection. After a thorough literature review, the inventors found no previous work that effectively removes rulers from skin lesion images. Therefore, they adopted two Al-based approaches for lesion detection and lesion segmentation (see below) to eliminate the rulers and present only the relevant lesion area to the Al model for classification. Hair removal was performed prior to lesion detection / segmentation, as the presence of hair can negatively impact this process. The pre-processed skin lesion images after hair removal, lesion detection and segmentation are highlighted in Fig. 7 and Fig. 8. By visual inspection, the inventors observed that hairs were removed and reconstructed images had better lesion area clarity which helped to correctly classify lesions into suspicious and non-suspicious categories.
[0178] Data pre-processing - Lesion detection. Deep learning is an effective approach for detecting objects within images by automatically learning the features crucial for object detection tasks. Advanced deep learning-based object detection models enable accurate detection and localisation of objects within an image, with the added benefit of removing artefacts such as hair and rulers from the image presented to the Al model. The inventors implemented a state-of-the-art object detector, namely Faster R-CNN, to detect lesions from input skin lesion images (after hair removal). This model was trained to estimate and draw a bounding box around any objects (lesion) if there is any and correctly classify that object class. In the present case there is only one target- whether there is any lesion present in the input image or not. If there is an object (lesion), the model draws a bounding box around it. The model was evaluated based on ground truth bounding boxes that are boxes of arbitrary size drawn by an expert around the lesions. Average precision (AP) and average recall (AR) were used as evaluation metrics to evaluate the proposed model. A confidence score is the probability that an anchor box contains an object. Intersection over Union (loU) is defined as the area of the intersection divided by the area of the union of a predicted bounding box and a ground-truth box. The inventors considered a true positive (TP) only if it satisfied three conditions: confidence score > threshold; the predicted class matches the class of a ground truth (where the only class is lesion, i.e. the model classifies whether the input image shows a lesion or not); the predicted bounding box has an loU greater than a threshold (e.g., 0.5, and 0.75) with the ground-truth. Violation of the latter two conditions makes a false positive (FP). During lesion detection training, an intersection-over-union (loU) threshold of 0.75 was used, meaning the ratio of the detected lesion area to the actual lesion area must be 0.75 or higher. An loU of 0.5 is considered a “good” score, while an loU of 1 represents a perfect score (Kim and Lee, 2020). If multiple predictions correspond to the same ground truth, only the one with the highest confidence score counts as a true positive, while the remaining are false positives. Precision is defined as the number of true positives divided by the sum of true positives and false positives. Finally, recall is defined as the number of true positives divided by the sum of true positives and false negatives. The inventors observed that their lesion detection model could differentiate between background, artefacts, rulers, and skin lesions with high confidence scores.
[0179] Al Model Development. An overview of the proposed Al framework for suspicious skin lesion classification is outlined in Fig. 19. During model development, raw images were pre-processed to re-shape to 1024 x 1024 and metadata were encoded to convert from string to nominal (using both encoded metadata encoded as described in Example 1 , and the C4C risk score calculated as described in Example 1) and then fed as input to the Al model. The C4C risk factors together with the risk score (i.e. risk factors themselves which form part of the risk score, encoded as explained above, and the risk score calculated as described above) and skin lesion images are used as inputs to the Al model and the model provides a decision whether the input belongs to the suspicious or non-suspicious group.
[0180] EfficientNet (Tan and Le, 2019; specifically, the EfficientNet B2 model) was used as the backbone Al model architecture. One of the major advantages of the EfficientNet B2 model as compared to other CNN-based models is its ability to scale up the model architectures in three dimensions (width, depth, and resolution) uniformly. This uses a compound scaling method illustrated on Fig. 9, which illustrates how the EfficientNet architecture can accommodate different image sizes to gain better performance. Fig. 9a shows a baseline network example, and Fig. 9b, c, d show how this architecture can be used with conventional scaling that only increases one dimension of network width, depth, or resolution, respectively. Fig. 9e shows the approach used here, which uses a compound scaling method adapted from Tan & Le (2019) that uniformly scales all three dimensions with a fixed ratio.
[0181] In developing the Al model, the inventors have adapted the ensembling (merging multiple model outputs to improve model performance) of the EfficientNet Al model for classifying input data into a binary class (suspicious vs non-suspicious). EfficientNet model works very well for image data classification. However, limited work has been done on the fusion of patient metadata and image data. As explained further above and below, the inventors first fused metadata and image data and fed these fused data to the EfficientNet model (adaptation 1). Second, the inventors developed multiple Efficient models based on input data type- 1) DER image alone, 2) SLR image alone, 3) DER+Metadata, 4) SLR+Medatadata, 5) DER+SLR image, and 6) DER+SLR+metadata (adaptation 2). Third, the inventors fused the outputs of those developed models to get a final decision about input data (adaptation 3), leading to significant performance improvement. The metadata and images were divided into training (80%) and test (20%) sets. During the training, 10- fold CV was used to build the Al model.
[0182] In small to medium-sized datasets, image augmentation is important to prevent overfitting. In the inventors’ pipeline, the following augmentations from the Pytorch augmentation library Albumentations (Buslaev et al., 2020) were used: Transpose, Flip, Rotate, RandomBrightness, RandomContrast, MotionBlur, MedianBlur, GaussianBlur, GaussNoise, OpticalDistortion, GridDistortion, ElasticTransform, CLAHE, HueSaturationValue, ShiftScaleRotate, and Cutout. Pytorch Albumentations is effciently implements a rich variety of image transform operations that are optimised for performance for different computer vision tasks, including object classification and detection. For the training schedule, cosine annealing was employed with a warm-up phase lasting one epoch. The models were trained for a total of 50 epochs. The initial learning rate for the cosine cycle was adjusted for each model, ranging from 1 e-4 to 3e-4. During the warm-up epoch, the learning rate was set to one-tenth of the initial rate for the cosine cycle. A batch size of 32 was used across all models. The minimum specification for Al model training was: RAM (128GB); CPU ( Intel Corei9 8 Core Processor, 3.0GHz); GPU (NVIDIA GPU with 24GB RAM). VIDIA CUDA CUDNN, Python 3.7, PyTorch, and Anaconda packages and tools were used to develop Al models.
[0183] Model validation. Establishing a reliable validation scheme is particularly important when dealing with small to medium-sized datasets or imbalanced data, as seen in the present example, where only 11 ,258 of the 79,246 images fall into the suspicious category. The disproportionate number of suspicious cases compared to non-suspicious ones can lead to instability in evaluation metrics such as accuracy and AUC (Area Under the ROC Curve). To address this, the inventors employed 10-fold CV and used balanced accuracy instead of standard accuracy. 10-fold CV was chosen to achieve a more generalised model, rather than relying on a single train-test data split. This approach allows all of the data to be used for both training and testing. By using 10-fold CV, the learning algorithm can be evaluated on examples it hasn’t encountered before, ensuring a more robust assessment. Creating five different models using 10-fold CV and testing them on five distinct test sets increases confidence in the algorithm’s performance. A single evaluation on a test set yields only one result, which could be due to chance or a biased test set. To mitigate bias and randomness, the experiment was repeated ten times with different train-test splits, averaging the results. The 79,246 skin lesion images were divided into training and test datasets, with 80% allocated to training and 20% to testing. To avoid data leakage, all data (including images and corresponding metadata) for each patient was exclusively assigned to either the training or testing dataset. During training, a 10-fold CV method was used to build the models. These models were optimised by fine-tuning hyperparameters, and the best-performing models were selected based on their 10-fold CV results using the training data. The chosen models were then evaluated using the test dataset. The performance of the Al framework was measured using the following evaluation metrics: Sensitivity (Sen), Specificity (Spc), Balanced Accuracy (Bal. Acc.), as defined in Eq. 5-7, and AUC=p(score(TP)>Score(TN)) where TP and TN refer to true positive (suspicious classified as suspicious), and true negative (non-suspicious classified as non-suspicious), respectively. The AUC of a classified is the probability that a randomly chosen TP case will be ranked higher than a randomly chosen TN case.
[0184] Metadata fusion with image data. The key advantage of the inventors’ approach to developing an Al model for skin lesion classification in the present example lies in the availability of metadata for 79,246 images. This metadata includes eight meta-features, comprising seven C4C risk factors and the C4C risk score, which can be integrated with image outputs to enhance model performance. The integration of metadata with images is illustrated in Fig. 20. In this context, ‘Swish’ is an activation function (Ramachandran et al. 2017), and ‘concat’ refers to the concatenation or fusion of the image vector with the metadata vector followed by a linear dropout layer with a ratio of 0.5 to get the final feature map to classify whether input belongs to suspicious or non-suspicious categories.
[0185] Al Model Decision Fusion. Six Al-based EfficientNet models were developed using different combinations of data, as shown in Figure 22. The inventors used the same configuration of base EfficientNet-B model but only varied input data. The fused six models are briefly summarised as: (1) DER- used only dermoscopic images to train the model; (2) SLR- used only DSLR images to train the model; (3) DER+Meta- fused dermoscopic images along with the metadata to train the model; (4) SLR+Meta- fused DSLR images along with the metadata to train the model; (5) DER+SLR- used both dermoscopic and DLSR images to train the model; (6) DER+SLR+Meta- fused metadata along with both dermoscopic and DLSR images to train the model. The outcomes from these models on test data were fused based on majority voting to get a final decision whether the test input belongs to the suspicious or not suspicious category.
[0186] Results and Discussion
[0187] Using Metadata Alone. In the previous example, the inventors found a new set of key risk factors, termed “C4C risk factors”, which include lesion pinkness, lesion size, lesion colour, lesion inflammation, lesion shape, lesion age, and natural hair colour. These factors are relevant not only for melanoma but for all types of skin cancer. In the work in Example 1 , the inventors assessed the effectiveness of the C4C risk factors in detecting suspicious skin lesions through ensembling of ML models, comparing them to the 7PCL and Williams risk factors. The inventors achieved a sensitivity of 80.46% and a specificity of 62.09% in detecting suspicious skin lesions as shown in Table 17, significantly outperforming the 7PCL-based method (sensitivity 68.09%, specificity 61.07%) and the Williams risk factors (sensitivity 66.32%, specificity 61.71 %). Furthermore, through weighting the seven new risk factors, the inventors came up with a new risk score, namely “C4C risk score”, which alone achieved a sensitivity of 76.09% and a specificity of 61.71 %, significantly outperforming the 7PCL-based risk score (sensitivity 73.91 %, specificity 49.49%) and the Williams risk score (sensitivity 60.68%, specificity 60.87%). Finally, fusing the C4C risk factors with the 7PCL and Williams risk factors (i.e. using all risk factors from the C4C, 7PCL, and Williams scores as input to NB, LR, SVM, RF, and MLP classifiers, with a final decision made based on majority voting) achieved the best performance, with a sensitivity of 85.24% and a specificity of 61.12%.
[0188] Table 17. Performance if the Al-based model using metadata alone.
[0189] Using Image Data Alone. The results of the Al model using image data alone are shown in Fig. 21A, with Al model performance for overall skin cancer detection of 99.72%, and with individual results for melanoma (99.36%), squamous cell carcinoma - SCC (100%), basal cell carcinoma - BCC (99.80%) and 63.22% for benign (non-malignant cases).
[0190] Image and Metadata Fusion. The results of the Al model using a combination of image data and metadata are shown in Fig. 21 B, with Al model performance for overall skin cancer detection of 99.44%, and with individual results for melanoma (99.37%), SCC (100%), BCC (99.50%) and 74.45% for benign (non-malignant cases). The incremental gain in model accuracy of 12.23% (75.45-63.22) for correct benign lesion classification when combining image data and metadata is significant and will substantially reduce the number of false positives and unnecessary clinical appointments for possible biopsy. This gain in model performance is novel and has the potential to be used during teledermatology triage.
[0191] Decision Fusion. The performance of individual models is summarised in Table 11. The performance of decision fusion for different data combination methods is summarised in Table 12. For each lesion, three pieces of information are available: 1) a DER image, 2) an SLR image, and 3) Metadata. Model DER+SLR+Meta was trained using DER, SLR, and Metadata. The models do not differentiate whether the input image is DER or SLR, they treat it as just an input image, but the ground truth is known (DER or SLR and its class- suspicious or non-suspicious). Therefore, all models take an image (which can be DER only, SLR only, or either of these, depending on the model) and metadata as input and provide an outcome decision. However, the performance of the models can be tested DER + Metadata or SLR+Metadata, and the image type in brackets in Table 12 indicates whether the model was tested on DER + Metadata or SLR+Metadata. The DER model, when tested on DER images, achieved a sensitivity of 99.50%, specificity of 63.06%, and balanced accuracy of 81 .28%. Sensitivity decreased to 95.48% for the SLR model when tested on SLR images. Adding metadata to the DER-Meta model significantly enhanced performance, achieving 99.50% sensitivity, 74.73% specificity, and 87.15% balanced accuracy. The best results were obtained with the DER-SLR-Meta model tested on DER images with metadata, yielding 99.83% sensitivity, 77.71 % specificity, and 88.77% balanced accuracy. Overall, models incorporating metadata with both DER and SLR images performed better on DER images than on SLR images.
[0192] A majority voting strategy was applied to determine if a test sample fell into suspicious or non- suspicious categories. The fusion of decisions from the DER-SLR and DER-SLR-Meta models, when tested on SLR images and DER images with metadata, respectively, resulted in a notable increase in specificity, reaching 82.72% (up from 77.71 %), compared to the best-performing single model, DER-SLR-Meta The second-best fusion involved combining decisions from DER-Meta, DER-SLR-Meta, DER-SLR and DER-SLR-Met models, yielding 99.66% sensitivity, 80.12% specificity, and a balanced accuracy of 89.89%. Although the fusion of-SLR-Meta, DER-SLR, and DER-SLR-Met achieved perfect sensitivity (100%), specificity fell to 75.38%. Thus, the fusion of DER-SLR and DER-SLR-Meta offered the best balance while adding more models increased computational complexity without enhancing performance.
[0193] Table 11. Performance comparison of individual models developed based on different data combinations. (4) DER- used only dermoscopic images to train the model; (6) SLR- used only DSLR images to train the model; (3) DER+Meta- fused dermoscopic images along with the metadata to train the model; (5) SLR+Meta- fused DSLR images along with the metadata to train the model; (2) DER+SLR- used both dermoscopic and DLSR images to train the model; (1) DER+SLR+Meta- fused metadata along with both dermoscopic and DLSR images to train the model.
[0194] Table 12. Performance comparison of decision fusion based on models using different data combination along with best individual models. Top 1 is the best performing single model which in this case was the model trained using dermoscopic & SLR images with metadata, then tested with dermoscopic images and metadata. Top2 is best combination of two separate models which in this case was two models each trained on DER+SLR data with and without Metadata but tested separately on SLR images and DER images+metadata, respectively. Top3 is the best combination of three models, op 4 is the best combination of 4 models, and top 5 is the best combination of 5 models The best combination overall was the top 2.
[0195] Benchmarking. In order to compare the present Al model sensitivity with Health Care Professional (HCP) reporting, data was used from two Cochrane Database Systemic Reviews (Dinnes et al., 2018a; Dinnes et al., 2018b) that provide a comparison table of visual inspection + / - dermoscopy for the detection of any skin lesion requiring excision, which is more relevant for comparison to the present example’s skin cancer diagnosis pathway, which includes patients with symptoms of all three major skin cancer types including melanoma, BCC and SCC. The Cochrane Reviews compute values of sensitivity at a fixed specificity of 80% to facilitate interpretation of results from different studies and report a 96% sensitivity for all skin cancer detection based on in-person inspection with dermoscopy. The study (Pacheco et al. 2021) proposed a CNN-based metadata processing block (Metablock) comprising 21 meta-features: age, sex, anatomical region, cancer history, skin phototype, family background along with 2298 clinical images captured using a smartphone and achieved 70% balanced accuracy that is higher then the feature concatenationbased method MetaNet (balanced accuracy, 66.20%). For comparison, the present example’s Al model has a sensitivity of 98.75%% for all skin cancer detection at a specificity of 80.37% and is therefore superior to the published data on in-person evaluation as highlighted in Table 13. Furthermore, the inventors benchmarked their approach against that of Skin Analytics (SA), a UKbased company working with the NHS for skin cancer detection. The Skin Analytics model (skinanalytics. com / ai-pathways / derm-performance / , accessed Aug 2024) only analyses a dermoscopic image. The inventors’ model outperformed Skin analytics’ in terms of accuracy for melanoma detection, BCC detection, SCC detection and correct classification of benign lesions, as highlighted in Table 13.
[0196] Table 13. Performance evaluation of the proposed Al model by benchmarking with the Skin Analytics, an UK-based company working with the NHS for skin cancer detection. HPC from Dinnes et al., 2018a; Dinnes et al., 2018b.
[0197] Explainability of Al Decision. Explainable Al (X-AI) has the potential to explain a transparent decision-making processes to the HCP (healthcare professional). The inventors set out to explain the decision-making of the Al model by adapting different attention mechanisms. To investigate whether the model was focusing on the correct lesion area or not while decision making, Grad- CAM was used to produce the heat map of the last layer of the Al model. Although the model did focus on the suspicious lesion, it was found that it may also focus on other areas or artefacts like rulers, as shown in Fig. 23. In order to tackle this issue, the inventors adapted four different attention mechanisms inspired from Dina model, a vision transformer model proposed by Facebook Al research team (Oquab et al., 2023, preprint), in order to force the Al model to focus on the correct lesion area by adoption of this attention mechanism. Some preliminary results of the Scaled Dot Product Attention tested are shown in Fig. 24 for illustrative purposes.
[0198] Conclusion
[0199] The fusion of multi-modal data comprising patient metadata and skin lesion images followed by applying advanced Al techniques has great potential to detect suspicious skin lesions at an early stage during teledermatology triage. With a reduction in patient referrals for possible biopsy, waiting times for skin cancer diagnosis and treatment will be shortened, resulting in improved outcomes. In this study, we devised an Al framework based on multi-modal data input for skin lesion classification which outperformed the existing state-of-the-art results reported by Skin Analytics, as well as published in-person evaluation. This study contributed to high-quality multimodal data collection followed by Al model development and Al decision fusion for suspicious skin cancer detection. Fusing patient metadata along with image data significant! improved the Al model performance as compared to using image data alone. We further attempted to explain Al decisions by adding a post-processing module that uses Grad-CAM to generate heatmaps, which show where an Al model focuses on during decision-making to decide whether a skin lesion belongs to suspicious or non-suspicious categories. In our future research, we are extending our investigation by adopting a soft-attention that forces an Al model to focus only on the lesion area and ignore artifacts such as rulers present in the images, which we believe will further boost the performance of skin cancer detection cost-effectively.
[0200] Example 3 - An Al model for skin cancer detection using dermoscopic images and a novel clinical risk score
[0201] Skin cancer incidence continues to rise globally, with 7.7 million new cases of non-melanoma (basal cell carcinoma - BCC; squamous cell carcinoma-SCC) skin cancer reported in 2017 (Ciuciulete et al., 2022), and melanoma cases predicted to rise to 510,000 by 2024 (Arnold et al., 2022). This rising cancer incidence, when combined with international shortages of dermatologists (Dermatology National Speciality Report, NHS, 2021 ; Hannah et al., 2021), has resulted in longer waiting times for skin cancer diagnosis and treatment. Most suspicious skin lesions referred for investigation are benign. With the rising skin cancer incidence and global shortages of dermatologists, technologies that can accelerate or automate lesion classification can reduce waiting times for skin cancer diagnosis or reassurance. A recent meta-analysis of 100 global studies has reported mean sensitivity for melanoma detection by experienced dermatologists using dermoscopy, or having access to dermoscopic images, of 85.7% (82.5%-88.3%), with mean sensitivity for combined BCC & SCC detection of 83.7% (76.6%- 89.0%) (Chen et al., 2025).
[0202] Artificial Intelligence (Al) models have already been shown to have the ability to detect melanoma on dermoscopic images with an accuracy comparable to experienced specialists (Phillips et al., 2019; Tschandl et al., 2019). To date, most studies using Al models have been focused only on melanoma detection using dermoscopic images alone. However, as illustrated by the work described in Example 1 , the inventors have highlighted the importance of clinical metadata for skin lesion classification, with the development of seven C4C risk factors and a novel C4C risk score for detection of all skin cancer subtypes including melanoma, BCC & SCC (Islam et al., 2024).
[0203] By combining the C4C risk score and risk factors with 79,246 images captured from 2015-2022, the inventors developed six Al models using the EffiicientNet-B2 architecture and different data input combinations related to all possible combinations of the dermoscopic image (DER), digital image (SLR) or metadata (see Example 2). Of these, the top-performing individual model was trained on DER, SLR & metadata, and is henceforth referred to as the C4C V1.1 model. As described in Example 2, performance using the C4C V1 .1 model for overall skin cancer detection was 99.4%, with individual results for melanoma (99.4%), SCC (100%) and BCC (99.5%) and benign (non-malignant) (74.5%) cases also reported (Islam et al., 2025). These skin cancer detection rates using the C4C V1.1 model were superior to two Cochrane Database Systematic Reviews that reported 96% sensitivity for in-person skin cancer detection at a fixed specificity of 80% (Dinnes et al., 2018a; 2018b). By fixing Al model specificity close to 80% (80.37%), skin cancer detection rates remained more than 96% for melanoma (98.1 %), SCC (100%) and BCC (98.8%). Finally, by using this fused model, the incremental gain in accuracy of 11.2% for benign cases compared to images alone, has the potential to reduce the number of false positive results and unnecessary appointments for face-to-face assessment and possible biopsy.
[0204] In the work described in the present example, the inventors validate the top-performing individual model developed in Example 2 (C4C V1 .1 , trained on DER, SLR and metadata, see Table 11) on a new dataset not used for model development. They aimed to assess the accuracy of their artificial intelligence (Al) model for detecting all skin cancer subtypes with dermoscopic images using their novel clinical risk score. This work was based on 36,094 dermoscopic images and metadata from 16,685 patients who attended UK nurse-led clinics during 2023-2024. Dermoscopic images and clinical metadata were captured from 16,685 patients aged 18+ with clinically suspicious skin lesions during 2023-2024, including 9,471 women (56.8%), with a mean (SD) age of 42.4 (11.8) years. It is widely recognised that model overfitting can occur in Al model training, resulting in reduced performance when evaluated in datasets not used for training. Therefore, in this study, the primary objective was to assess performance of skin lesion classification using the Al model and compare model assessment with prior healthcare professional skin lesion classification and histopathological diagnosis. The secondary objective was to assess the potential impact of automated Al model assessment prior to or without human reporting in images and metadata.
[0205] Sensitivity, specificity, area under the receiver operating characteristic curve (AUROC) and Negative Predictive Value (NPV) of algorithmic skin lesion classification versus specialist classification were evaluated. For the 36,094 images of clinically suspicious skin lesions from 16,685 patients, Al model sensitivity for the tested model for all skin cancer detection was 99.0% specificity = 73.2%; AUROC = 86.5%), with individual sensitivity for melanoma 96.8%, basal cell carcinoma 99.3%, and squamous cell carcinoma 100%, and with the potential to reassure 54.2% of images without human reporting. This specificity of 73.2% was achieved using a benign confidence threshold of 86% in the Al model, which resulted in 5 misclassified cancers including 2 melanomas and 3 BCCs (NPV=99.9%). By increasing the benign threshold to 93%, no skin cancers were misclassified as non-suspicious (NPV=100%). With skin cancer detection rates higher than experienced dermatologists using dermoscopy, this Al model can help tackle the rising cancer incidence and global shortages of dermatologists.
[0206] Methods
[0207] Data collection. 36,094 dermoscopic images and clinical metadata (C4C Risk Factors + Risk Score) from 16,685 patients were collected for patients with symptoms that were suspicious of skin cancer (change in colour, shape or size; >7mm; inflamed, oozing). Clinical data and images were captured by centrally trained nurses according to protocol, with teledermatology reporting by a team of skin cancer specialists. Dermoscopic images in 2023 were captured using a Canon EOS 1000-4000s camera with a dermoscopic attachment, and in 2024, a smartphone with dermoscopic attachment was used. This study followed the Standards for Reporting of Diagnostic Accuracy (STARD) and Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guidelines.
[0208] The result of skin lesion classification to suspicious (n=4,691) or non-suspicious (n=31 ,403) categories following teledermatology triage by a healthcare professional (HCP) was available, and this was accepted as the ground truth during assessment of algorithm performance. In total, there were 502 biopsy-proven skin cancers including 63 melanomas, 418 BCCs and 21 SCCs. Fitzpatrick skin type was recorded from 2024 onwards, with 14,835 images having skin type recorded during the nurse-led clinical examination.
[0209] Al framework & data fusion. Image pre-processing and development of the C4C risk factors and C4C risk score (metadata) are described in Examples 1 and 2, and in Islam et al. (2024, 2025). As described in Example 2, both images and metadata (C4C risk factors and C4C risk score) were used as inputs to the Al model to classify the skin lesion as suspicious or non-suspicious. In summary, six Al models were developed using the EffiicientNet-B2 architecture and different data input combinations related to all possible combinations of the dermoscopic image (DER), digital image (SLR) or metadata. The best single performing Al model in terms of balanced accuracy and area under the receiver operating characteristic curve (AUROC) was trained on DER, SLR & metadata which, when tested on dermoscopic images combined with metadata, had a sensitivity of 98.3%, a specificity of 77.1 %, a balanced accuracy of 88.77% and an AUROC of 92.98% (see Example 2 and Islam et al., 2025). This particular Al model (C4C V1 .1) was therefore selected for use in the validation work described in the present example.
[0210] Al algorithm performance. A total of 36,094 dermoscopic images and associated clinical metadata (C4C Risk Factors & Risk Score) collected from 2023-2024 were classified as suspicious or non- suspicious by the Al algorithm. The performance metrics evaluated were sensitivity (all skin cancers and by subtype), specificity, AUROC and Negative Predictive Value (NPV). Al model performance was compared to published results of healthcare professional reporting accuracy for skin cancer diagnosis, including the recently published meta-analysis of 100 studies that reported on the accuracy of skin cancer diagnosis according to clinical experience and examination type (Chen et al., 2025).
[0211] Partial automation of skin lesion classification by Al algorithm. A total of 36,094 images and associated clinical metadata collected from 2023-2024 were classified by the Al model. Higher confidence thresholds for a non-suspicious classification were correlated with the potential number of misclassified cancers that would result from partial automation of the non-suspicious category during teledermatology triage.
[0212] Al algorithm performance according to skin type. In 2024, a total of 14,835 dermoscopic skin lesion images had skin type recorded by the specialist skin nurse using the Fitzpatrick Scale l-VI (Fitzpatrick, 1988). Al model performance metrics including sensitivity (all skin cancers and by subtype), specificity and balanced accuracy were analysed by comparing groups Fitzpatrick skin types (l-ll) with (lll-IV) and (V-VI).
[0213] Results
[0214] Al algorithm performance. Performance of the Al model in 36,094 images and clinical metadata from 16,685 patients confirmed overall sensitivity of 99.0%, specificity of 73.2% and an AUROC of 86.5% (Table 14).
[0215] In total, 4,218 of 4,691 images (90.0%) previously classified as suspicious by the HCP were also classified as suspicious by the Al model. Of the 502 biopsy-proven skin cancers that were assessed by the Al model, correct classification as suspicious was achieved in 61 of 63 melanomas (96.8%), 415 of 418 BCCs (99.3%) and all 21 SCCs (100%).
[0216] Table 14. Al model performance, with specific sensitivity for melanoma (MEL), basal cell carcinoma (BCC) and squamous cell carcinoma (SCO) detection.
[0217] Comparison with published results of skin cancer diagnosis by healthcare professionals. To further validate the Al model, the inventors compared the accuracy of the model with published data on skin cancer diagnosis by experienced healthcare professionals using dermoscopy.
[0218] The sensitivity for overall skin cancer detection using dermoscopic images and clinical metadata (99.0%) was greater than the 82% quoted in two Cochrane Reviews for the assessment of dermoscopic images and higher than the 96% reported for in-person visual inspection using dermoscopy (Dinnes et al., 2018a; 2018b). A more relevant comparison with the 96% sensitivity reported in the Cochrane Reviews is to calculate the C4C model sensitivity at a fixed specificity of 80%, which was 96.2% for the 2023-2024 dataset. This result confirmed equivalence with results for visual inspection using dermoscopy (96%) quoted in the Cochrane Reviews.
[0219] The Al model accuracy can also be compared with the recent meta-analysis (Chen et al., 2025), whcih reported that mean sensitivity for melanoma detection by experienced dermatologists using dermoscopy, or having access to dermoscopic images, was 85.7% (82.5%-88.3%), with a mean specificity of 81 .3% (76.3%-85.4%) and an AUC of 0.90. The sensitivity for melanoma detection of 96.8% in this study exceeded the meta-analysis figure of 85.7%.
[0220] The meta-analysis also reported mean sensitivity (95% Cl) for combined BCC & SCC detection by experienced dermatologists using dermoscopy, or having access to dermoscopic images, was 83.7% (76.6%-89.0%), with a specificity (95% Cl) of 87.4% (78.9%-92.8%) and an AUROC of 0.91 . The sensitivity for BCC detection by C4C V1.1 of 99.3% exceeded the meta-analysis figure of 83.7% and the sensitivity for SCC detection by C4C V1.1 of 100% exceeded the meta-analysis figure of 83.7%.
[0221] Partial automation of non-suspicious classification. The inventors’ proposed partly automated teledermatology pathway is shown in Figure 26. In the automated pathway, images classified as suspicious can proceed straight to face-to-face assessment and possible biopsy by an appropriate clinician. Images classified as non-suspicious above a certain threshold can be reassured without a HCP teldermatology report (i.e. patients with skin lesions classified as non-suspicious by the Al model with confidence exceeding a benign confidence threshold can be discharged). All remaining images will undergo telemedicine reporting by an HCP to classify them as suspicious or non- suspicious (i.e. patients with lesions classified as non-suspicious but not exceeding the benign confidence threshold proceed to HCP reporting).
[0222] The specificity, or benign sensitivity, for the 2023-2024 dataset was 73.2% and this was achieved using a benign confidence threshold in the Al model of 86%. Figure 25 plots the number of misclassified cancers by subtype against the non-suspicious confidence threshold at 1 % intervals from 79% to 93%. At a threshold of 86%, there were 5 cancers misclassified as non-suspicious, including two melanomas and three BCCs, out of a total of 502 biopsy-proven cancers, with a NPV of 99.9% for all skin cancer subtypes. At a confidence threshold of 93%, no skin cancers were misclassified, with a NPV of 100%.
[0223] Although there are no misclassified skin cancers using a benign confidence threshold of 93%, the threshold can be increased to provide additional reassurance of a non-suspicious diagnosis. As can be seen in Table 15, as the threshold was increased incrementally from 86% to 97%, there was a corresponding drop in the number of images that could be reassured without HCP reporting.
[0224] Table 15. Number of images that can be reassured according to the benign confidence threshold.
[0225] When the 36,094 images from 2023-2024 were mapped to the inventors’ proposed and partly automated teledermatology pathway, with all skin lesion images and metadata undergoing Al model classification prior to HCP reporting and using a benign confidence threshold of >93%, 19,575 (54.2%) of all images could be reassured without HCP reporting, with no misclassified cancers (Figure 26). The 19,575 images included 228 images that were originally classified as suspicious by the HCP but were not proven to be skin cancers. A total of 11 ,828 images (32.8%), of which 245 were previously reported as suspicious by the HCP, would proceed to HCP reporting, including 5 misclassified skin cancers according to the default benign confidence threshold of 86%.
[0226] Furthermore, patients with skin lesions classified as suspicious by the Al algorithm (4,218 images (11.7%)) could proceed straight to a face-to-face assessment with or without biopsy, without HCP reporting. This group of images included 497 of the 502 biopsy-proven cancers, with the remaining 5 cancers being classified as suspicious during HCP reporting. This partly automated teledermatology triage pathway provides total efficiency savings of 65.9% (54.2% + 11.7%) for skin lesion classification.
[0227] Al model performance according to skin type. The dataset for Al model performance according to skin type comprised 14,835 dermoscopic images and clinical metadata from 6,623 patients captured during 2024, including 3779 women (57.1 %), with a mean (SD) age of 43.1 (12.1) years. This dataset included 210 biopsy-proven skin cancers (melanoma, n=27; BCC, n=177; SCC, n=6). As can be seen in Table 16, there was high sensitivity and specificity for all skin cancer types in Fitzpatrick skin type groups l-IV.
[0228] In Fitzpatrick skin types V-VI, there were small numbers of patients overall and, consequently, small numbers of biopsies and only one biopsy-proven cancer (8 “ground truth” cancers, of which only 1 was biopsy-proven and the other 7 were identified based on image data alone). As a result, Al model performance in these groups was linked to the likely diagnosis assigned by the HCP reporter rather than a histopathological diagnosis. Despite there being only one biopsy-proven cancer in skin type Group V-VI, the Al model classification agreed with the HCP reporter in all the other seven skin cancers, including four melanomas, with an overall sensitivity for skin lesion classification of 100% and a specificity of 58.8%.
[0229] Table 16. Al model performance by Fitzpatrick skin type groups in dermoscopic images with C4C risk score.
[0230] Discussion
[0231] Al model performance. The work described in the present example is the first validation study of the C4C V1 .1 Al model with a large dataset of skin lesion images that were not used for model development. Analysis of 36,094 images and associated metadata in the 2023-2024 dataset confirmed overall sensitivity of 99.0%, specificity of 73.2%, and an AUROC of 86.5%. Furthermore, of the 502 skin cancers that were assessed by the Al model, correct classification to the suspicious category was achieved in 61 of 63 melanomas (96.8%), 415 of 418 BCCs (99.3%) and all 21 SCCs (100%). As the main priority during initial Al model development was to maximise the sensitivity for skin cancer detection, in order to keep the number of misclassified cancers to a minimum, it was reassuring to see skin cancer detection rates that were at least equivalent to published results for experienced clinicians using dermoscopy, both overall and by individual skin cancer subtype.
[0232] The results for Al model performance confirmed equivalence with results for visual inspection using dermoscopy (96%) quoted in two Cochrane Reviews (Dinnes et al., 2018a; 2018b), with sensitivity results for melanoma (96.8%), BCC (99.3%) and SCC (100%) exceeding those from a recent meta-analysis for melanoma (86%) and non-melanoma skin cancer (84%) (Chen et al., 2025). When rounded up, these sensitivity results for individual skin cancer subtypes match or exceed results from the DERM Al algorithm for melanoma (97%), BCC (98%) and SCC (98%) based on 31 ,075 skin lesions assessed during 2023-2024 (https: / / skin-analytics.com / ai-pathways / derm- performance / ).
[0233] Image and metadata fusion. To date, most Al models for skin cancer detection have been developed using only images or images together with limited clinical metadata such as age, gender and lesion location (Ha et al., 2020; Pacheco et al., 2020) and are often limited to melanoma detection only (Phillips et al., 2019). The Al model performance results achieved in the work described in the present example can partly be explained by the access to high-quality image data, collected according to protocol, as well as extensive clinical metadata. As shown in the work described in Examples 1 and 2, a combination of image data and the seven C4C Risk Factors and C4C Clinical Risk Score improves Al model accuracy for the correct classification of non-suspicious lesions (specificity) by 11% from 63% to 74% (Islam et al., 2024; 2025). This increase in specificity is clinically significant and will result in a reduction in the number of false positives and a reduction in the number of patients recalled for unnecessary face-to-face appointments and possible biopsy.
[0234] Automation of lesion classification. With a rising skin cancer incidence and global dermatology workforce shortages, it is not surprising that waiting times for skin cancer diagnosis and treatment have lengthened. It is well established that teledermatology pathways can shorten waiting times to skin cancer diagnosis and improve early-stage melanoma diagnosis (Walls et al., 2024). With appropriate controls, integration of Al models into the teledermatology triage process has the potential to partly automate lesion classification and standardise the level of reporting, not just for melanoma but for all skin cancer subtypes.
[0235] The secondary objective was to assess the potential impact of automated assessment using the C4C V1.1 model prior to human reporting in 36,094 images and metadata captured from 2023- 2024, including 502 biopsy-proven cancers. A sensitivity of 99.0% and specificity of 73.2% were achieved using a benign confidence threshold of 86% by the Al algorithm, but this threshold resulted in 5 misclassified cancers including 2 melanomas and 3 BCCs. By increasing the benign confidence threshold to 93%, no skin cancers of any type were misclassified in this cohort.
[0236] Application of the 93% benign threshold to the validation dataset, with all skin lesion images and metadata undergoing Al model classification prior to HCP reporting, resulted in 19,575 (54.2%) non-suspicious images being reassured without HCP reporting, with no misclassified cancers. Further automation is also possible for patients with skin lesions classified as suspicious by the Al model (11.7% of images in the validation set), who can proceed straight to a face-to-face assessment with / without biopsy without HCP reporting, providing total efficiency savings of 65.9% (54.2% + 11 .7%) for lesion classification. The ability to partly automate lesion classification by 65.9% during teledermatology triage, by using an Al model that is based on images and clinical metadata inputs, can result in shorter waiting times to patients being reassured or diagnosed with skin cancer.
[0237] Al model performance and skin tone. It is well documented that patients with skin of colour have a worse survival rate from melanoma yet are often underrepresented in research studies (Brunsgaard et al., 2023). As this study was solely based on a UK population, only 261 of 14,835 patients (1.8%) were recorded to have Fitzpatrick skin types V or VI. Although there was high sensitivity and specificity across all skin cancer subtypes in Fitzpatrick skin types l-IV, the lack of patients and biopsy-proven cancers in skin types V-VI means that further research is required to assess the performance of this Al model in patients with skin of colour. It was reassuring however that despite there being only one biopsy-proven cancer in skin type group V-VI, the Al model agreed with the HCP reporter opinion in all the other seven skin cancers, including four melanomas. Further research in this field is already underway by the inventors.
[0238] Limitations of study. This study was limited by the small percentage of patients with Fitzpatrick skin types V-VI and as a result, the small number of biopsy proven cancers in these groups. Despite there being a strong correlation between HCP opinion and Al algorithm classification for skin cancer diagnosis, a larger dataset is required to assess Al model performance in patients with Fitzpatrick skin types V-VI.
[0239] Conclusions. Al model performance for skin lesion classification is equivalent to face-to-face assessment by experienced dermatologists. The inclusion of the C4C risk score increases benign accuracy and facilitates use of the model as a clinical decision support tool or autonomously within a clinically supervised teledermatology pathway.
[0240] Al model sensitivity is higher than published results for skin cancer detection by experienced dermatologists using dermoscopy, with potential to reassure 54.2% of images without human reporting and no misclassified cancers.
[0241] REFERENCES
[0242] All references cited herein are incorporated by reference in their entirety.
[0243] Arnold M, Singh D, Laversanne M et al. Global burden of cutaneous melanoma in 2020 and projections to 2040. JAMA Dermatol. 2022; 158(5), 495-503.
[0244] Bardou, D., Bouaziz, H., Lv, L. & Zhang, T. Hair removal in dermoscopy images using variational autoencoders. Ski. Res. Technol. 28, 445-454 (2022).
[0245] Brinker, T. J. et al. Skin cancer classification using convolutional neural networks: systematic review. J. Med. Internet Res. 20, e11936 (2018).
[0246] Brinker, T. J. et al. Deep learning outperformed 136 of 157 dermatologists in a head-to-head dermoscopic melanoma image classification task. Eur. J. Cancer 113, 47-54 (2019).
[0247] Brunsgaard EK, Jensen J, Grossman D. Melanoma in skin of color: Part II. Racial disparities, role of UV, and interventions for earlier detection. J. Am. Acad. Dermatol. 2023; 89(3): 459-468.
[0248] Burroni, M. et al. Melanoma computer-aided diagnosis: reliability and feasibility study. Clin. Cancer Res. 10, 1881-1886 (2004).
[0249] Buslaev, A. et al. Albumentations: fast and flexible image augmentations. Information 11 , 125 (2020).
[0250] Chen JY, Fernandez K, Fadadu R et al. Skin cancer diagnosis by lesion, physician and examination type. JAMA Dermatology, 2025; 161 (2): 135-146.
[0251] Ciuciulete A-R, Stepan AE, Andreiana BC & Simionescu CE. Non-melanoma skin cancer: statistical associations between clinical parameters. Curr. Heal. Sci. J. 2022; 48(1), 110-115.
[0252] Derm performance- skin analytics, https: / / skin-analytics.com / ai-pathways / derm-performance / (2024). Accessed Aug. 2024.
[0253] Dildar, M. et al. Skin cancer detection: a review using deep learning techniques. Int. J. Environ. Res. Public Heal. 18, 5479 (2021). Dinnes, J. et al. Dermoscopy, with and without visual inspection, for diagnosing melanoma in adults. Cochrane Database Syst. Rev. (2018a).
[0254] Dinnes, J. et al. Visual inspection and dermoscopy, alone or in combination, for diagnosing keratinocyte skin cancers in adults. Cochrane Database Syst. Rev. (2018b).
[0255] Esteva, A. et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature 542, 115-118 (2017).
[0256] Fitzpatrick TB. The validity & practicality of sun=reactive skin types I through VI. Arch. Dermatol. 1988; 124(6): 869-871.
[0257] Ganster, H. et al. Automated melanoma recognition. IEEE Transactions on Med. Imaging 20, 233-239 (2001).
[0258] Getting it Right First Time. Dermatology National Speciality Report. NHS, 2021 . https: / / gettingitrightfirsttime.co.uk / girft-reports. Accessed May 2024.
[0259] Gutman, D. et al. Skin lesion analysis toward melanoma detection: A challenge at the international symposium on biomedical imaging (isbi) 2016, hosted by the international skin imaging collaboration (isic). arXiv preprint arXiv:1605.01397 (2016).
[0260] Ha, Q., Liu, B. & Liu, F. Identifying melanoma images using effcientnet ensemble: Winning solution to the SIIM-ISIC melanoma classification challenge. arXiv preprint arXiv: 2010. 05351 (2020).
[0261] Hannah C, Williams V, Fuller LC & Forrestel A. The impact of the covid-19 pandemic on global health dermatology. Dermatol. Clin. 2021 ; 39(4), 619-625.
[0262] Harrell, F. E. Ordinal logistic regression. In Regression Modeling Strategies, 311-325 (Springer, 2015).
[0263] Heroza, R. I., Gan, J. Q. & Raza, H. Enhancing skin lesion classification: A self-attention fusion approach with vision transformer. In Annual Conference on Medical Image Understanding and Analysis, 309-322 (Springer, 2024).
[0264] Islam S, Wishart GC, Walls J, Hall P, Seco de Herrera AG, Gan JQ, Raza H. Leveraging Al and patient metadata to develop a novel score for skin cancer detection. 2024. Sci Rep; 14(1): 20842.
[0265] Islam S, Wishart GC, Walls J, Hall P, Seco de Herrera AG, Gan JQ, Raza H. Advancing skin cancer detection through deep learning and fusion of patient metadata and skin lesion images. 2025. Scientific Reports (resubmitted with revisions March 2025).
[0266] Jayapriya, K. & Jacob, I. J. Hybrid fully convolutional networks-based skin lesion segmentation and melanoma detection using deep feature. Int. J. Imaging Syst. Technol. 30, 348-357 (2020).
[0267] Kim, K. & Lee, H. S. Probabilistic anchor assignment with iou prediction for object detection. In Computer Vision-ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XXV 16, 355-371 (Springer, 2020).
[0268] Lopez, A. R., Giro-i Nieto, X., Burdick, J. & Marques, O. Skin lesion classification from dermoscopic images using deep learning techniques. In 2017 13th IASTED International Conference on Biomedical Engineering (BioMed), 49-54 (IEEE, 2017).
[0269] MacKie, R. M. An Illustrated Guide to the Recognition of Early Malignant Melanoma (University of Glasgow, Glasgow, 1986).
[0270] Oquab, M. et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 (2023).
[0271] Pacheco, A. G. & Krohling, R. A. The impact of patient clinical information on automated skin cancer detection. Comput. Biol. Med. 116, 103545 (2020).
[0272] Papachristou, I. & Bosanquet, N. Improving the prevention and diagnosis of melanoma on a national scale: A comparative study of performance in the united kingdom and Australia. J. Public Health Policy 41 , 28-38 (2020). Phillips M, Marsden H, Jaffe W et al. Assessment of accuracy of an artificial intelligence algorithm to detect melanoma in images of skin lesions. JAMA Netw. Open. 2019; 2(10): e1913436.
[0273] Ramachandran P., Zoph B., Le Q.V.. Swish: a self-gated activation function. arXiv:1810.05941v1 . 16 Oct 2017
[0274] Roffman, D., Hart, G., Girardi, M., Ko, C. J. & Deng, J. Predicting non-melanoma skin cancer via a multi-parameterized artificial neural network. Sci. Rep. 8, 1701 (2018).
[0275] Shajimon, G. M., Ufumaka, I. & Raza, H. An improved vision-transformer network for skin cancer classification. In 2023 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), 2213-2216 (IEEE, 2023).
[0276] Simonyan, K. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)
[0277] Skin Analytics DERM Al model performance 2023-2024. https: / / skin-analytics.com / ai- pathways / derm-performance / . Accessed April 7, 2025.
[0278] Strzelecki, M. et al. Artificial intelligence in the detection of skin cancer: state of the art. Clinics in Dermatology 42, 280-295 (2024).
[0279] Tan, M. & Le, Q. Efficientnet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning, 6105-6114 (PMLR, 2019).
[0280] Tschandl P, Codella N, Akay BN at al. Comparison of the accuracy of human readers versus machine learning algorithms for pigmented skin lesion classification: an open, web-based, international, diagnostic study. Lancet Oncol. 2019; 20(7): 938-947
[0281] Thomas, L. et al. Real-world post-deployment performance of a novel machine learning-based digital health technology for skin lesion assessment and suggestions for post-market surveillance. Front. Medicine 10, 1264846 (2023).
[0282] Walls J, Hall P, Mills L, Kilford C, Wishart GC. Teledermatology pathway performance for detection of all skin cancer categories. British Journal of Dermatology. 2024; 191 (Supplement_1): i198.
[0283] Walter, F. M. et al. Using the 7-point checklist as a diagnostic aid for pigmented skin lesions in general practice: A diagnostic validation study. Br. J. Gen. Pract. 63, e345-e353 (2013).
[0284] Williams, L. H., Shors, A. R., Barlow, W. E., Solomon, C. & White, E. Identifying persons at highest risk of melanoma using self-assessed risk factors. J. Clin. Exp. Dermatol. Res. 2, 6 (2011).
[0285] Yap, J., Yolland, W. & Tschandl, P. Multimodal skin lesion classification using deep learning. Exp. Dermatol. 27, 1261-1267 (2018).
[0286] Zhang, J., Xie, Y, Xia, Y. & Shen, C. Attention residual learning for skin lesion classification. IEEE Transactions on Med. Imaging 38, 2092-2103 (2019).
Claims
CLAIMS1 . A computer-implemented method for classifying a skin lesion in a patient as suspicious or non-suspicious, the method comprising: obtaining data for one or more features associated with the lesion, the one or more features comprising a determination of whether a skin lesion is pink in colour or not pink; and classifying, using said data, the lesion between a first class of suspicious lesions and a second class of non-suspicious lesions.
2. The method according to claim 1 , wherein classifying the lesion comprises obtaining a score for each of said features, and combining the scores to provide a lesion risk score.
3. The method of claim 2, wherein obtaining a score for each of said features comprises using, for each of said one or more features, a predetermined scoring table to score said feature; and / or wherein classifying the lesion comprises comparing the lesion risk score to a predetermined threshold, wherein a lesion associated with a lesion risk score equal or higher than the predetermined threshold is classified as being suspicious and a lesion with a lesion risk score lower than the predetermined threshold is classified as being non- suspicious.
4. The method of claim 3, wherein the predetermined threshold has been identified using training data comprising data for said features for a plurality of training lesions and associated labels indicating whether each training lesion is suspicious or non-suspicious, and / or wherein the predetermined threshold is a threshold that optimises one or more of: the sensitivity of identification of suspicious lesions, the specificity of identification of suspicious lesions, the accuracy or balanced accuracy of identification of suspicious and non-suspicious lesions and / or the wherein the predetermined threshold is a threshold, optionally an integer threshold, that results in a maximum value of an accuracy metric in a training cohort comprising data for said features for a plurality of training lesions and associated labels indicating whether each training lesion is suspicious or non-suspicious.
5. The method according to claim 1 , wherein classifying the lesion comprises encoding each of said features that is not numerical into a numerical format, thereby obtaining a set of numerical features, and providing said numerical features as input to a machine learning model trained to classify lesions between a first class of suspicious lesions and a second class of non-suspicious lesions using training data comprising data for said features for a plurality of training lesions and associated labels indicating whether each65training lesion is suspicious or non-suspicious, optionally wherein said encoding is as defined in Table 1 .
6. The method of claim 5, wherein the machine learning model comprises one or more machine learning models selected from: a naive Bayes classifier, a support vector machine, a logistic regression model, a Random forest classifier, and a multilayer perceptron, and / or wherein the machine learning model comprises an ensemble of machine learning models, optionally wherein each of the machine learning models in the ensemble of machine learning models is a different type of model and / or wherein each of the machine learning models in the ensemble of machine learning models uses the same one or more features and / or wherein the output of the machine learning models are combined using majority voting or using a further classifier, optionally a multilayer perceptron.
7. The method of any preceding claim, wherein the one or more features are features associated with a risk factor for characterising suspicious lesions, and the one or more features further comprise one or more of: patient age, patient gender, lesion size increase, lesion age, lesion shape change, lesion colour change, lesion total size, lesion inflammation, lesion oozing, lesion itch, lesion location, prior family history, natural hair colour, number of sunburns, number of moles, density of freckles, prior history of melanoma and prior history of skin cancer, optionally wherein the features are as defined in Table 1 .
8. The method according to any preceding claim, wherein the classifying uses four to seven features each associated with a risk factor for characterising suspicious skin lesions.
9. The method according to claim 8, wherein the features used in addition to lesion pink are selected from lesion size increase, lesion shape change, lesion colour change, lesion inflammation, hair colour, and lesion age, optionally wherein the features associated with risk factors for characterising suspicious skin lesions comprise: (a) lesion colour change, lesion inflammation and lesion age, (b) lesion colour change, lesion inflammation and lesion shape change, or (c) lesion pink, lesion inflammation, lesion shape change, lesion colour change, lesion size, lesion age and natural hair colour.
10. The method according to any preceding claim, wherein: (i) lesion pink is defined as whether the lesion is pink or not, optionally wherein lesion pink is encoded as a binary66variable where yes is associated with a score or encoding of 1 and no is associated with a score or encoding of 0.
11. The method according to any of claims 7 to 10, wherein: (i) lesion inflammation is defined as whether the lesion is inflamed or not, optionally wherein lesion inflammation is encoded as a binary variable where yes is associated with a score or encoding of 1 and no is associated with a score or encoding of 0, (ii) lesion shape change is defined as whether the lesion has changed in shape or not, optionally wherein lesion shape change is encoded as a binary variable where yes is associated with a score or encoding of 1 and no is associated with a score or encoding of 0, (iii) lesion colour change is defined as whether the lesion has changed in colour or not, optionally wherein lesion colour change is encoded as a binary variable where yes is associated with a score or encoding of 1 and no is associated with a score or encoding of 0, (iv) lesion size is whether the lesion has a diameter at or above a predetermined threshold, optionally wherein the predetermined threshold is7mm and / or wherein lesion size is encoded as a binary variable where yes is associated with a score or encoding of 1 and no is associated with a score or encoding of 0 and / or wherein the diameter of a lesion is the diameter of the smallest circle that fits the lesion, (v) lesion age is defined as whether the lesion has been present for at least a predetermined amount of time or not, optionally the predetermined amount of time is 6 months and / or wherein lesion age is encoded as a binary variable where yes is associated with a score or encoding of 1 and no is associated with a score or encoding of 0, and / or (vi) natural hair colour is defined as the natural hair colour of the subject, optionally wherein natural hair colour is encoded as a discrete variable where black, red, blonde and brown are associated with respective discrete natural numbers, and / or wherein black is associated with a score or encoding of 1 , red is associated with a score or encoding of 2, blonde is associated with a score or encoding of 3, and brown is associated with a score or encoding of 4.
12. The method according to any preceding claim, wherein classifying the lesion between a first class of suspicious lesions and a second class of non-suspicious lesions comprises weighting numerical values associated with each if said features such that the risk factors with greater significance for distinguishing between suspicious and non-suspicious lesions are given greater relative scoring values for determining the lesion risk score, optionally wherein the weighing gives importance to any of the following features in the following order from greater to lower importance: lesion pink, lesion inflammation, lesion shape change, lesion colour change, lesion size, lesion age and natural hair colour, optionally wherein the weighting is as defined in Table 3.6713. The method according to any preceding claim, wherein the one or more features further comprises a 7-point checklist score and / or a Williams score and / or any one or more features of the 7-point checklist and / or any one or more features of the Williams score.
14. The method according to any of claims 2 to 11 , wherein the method further comprises calculating a Williams score and / or a 7-point checklist score and combining the Williams score and / or the-point checklist score with the lesion risk score to calculate an overall risk score.
15. The method according to any preceding claim, wherein the patient is a human subject, optionally a human subject of 18 years old or more, and / or wherein the lesion is at most 15mm in largest dimension, and / or wherein the subject is a subject who has been identified has having a skin lesion at risk of developing into skin cancer, optionally wherein the skin cancer is melanoma, squamous cell carcinoma or basal cell carcinoma.
16. The method according to any preceding claim, wherein a lesion classified in the suspicious class is at higher risk of developing melanoma, basal cell carcinoma or squamous cell carcinoma than a lesion classified in the non-suspicious class.
17. The method according to any preceding claim, wherein the method further comprises obtaining an image of the skin lesion, wherein classifying the lesion comprises: using a machine learning model that takes as input (a) one or more images of the lesion and (b) the data for one or more features associated with the lesion and / or a lesion risk score derived therefrom, wherein the machine learning model comprises a deep learning model that has been trained using training data comprising, for each of a plurality of training lesions: (a) one or more images of the training lesion, (b) data for one or more features associated with the lesion and / or a lesion risk score derived therefrom, and (c) a label classifying each training lesion as suspicious or not suspicious.
18. The method according to claim 17, wherein the machine learning model takes as input: (a) one or more images of the lesion, and (b) the data for one or more features associated with the lesion and a lesion risk score derived from the one or more features associated with the lesion, wherein the one or more features comprise or consist of: lesion pink, lesion inflammation, lesion shape change, lesion colour change, lesion size, lesion age and natural hair colour.
19. A method according to claim 17 or claim 18, wherein the method further comprises subjecting the one or more images to a pre-processing step whereby the shape and / or68resolution of the image is standardised, and / or wherein the one or more images of the training lesions in the training data have been subject to a pre-processing step whereby the shape and / or resolution of the image is standardised.
20. A method according to claim 19, wherein the one or more images are resized to be square shaped using image padding, optionally using image padding with the mode skin colour, image padding with a predetermined colour, or image padding with black pixels.21 . A method according to any of claims 17 to 20, wherein the method further comprises subjecting the one or more images to a pre-processing step whereby the image is cleaned to remove features not relevant for determining the likelihood of a skin lesion being suspicious or non-suspicious, optionally wherein said pre-processing includes hair removal and / or lesion detection and segmentation, and / or wherein the one or more images of the training lesions in the training data have been subject to said preprocessing step.
22. A method according to any of claims 17 to 21 , wherein the method further comprises processing the image by object detection techniques to identify the parts of the image that are most relevant for determining the likelihood of a skin lesion being suspicious or non-suspicious.
23. A method according to any of claims 17 to 22, wherein the machine learning model comprise a deep learning model, or wherein the machine learning model comprises a convolutional neural network, optionally an efficientNet, and / or wherein the machine learning model comprises an ensemble of models, optionally wherein the machine learning model comprises an ensemble of machine learning models each taking as input (a) one or more images of the lesion and (b) the data for one or more features associated with the lesion and / or a lesion risk score derived therefrom, further optionally wherein each machine learning model has been trained and / or tested with images acquired using a different type of camera, and / or wherein the outputs of the machine learning models in the ensemble are combined using majority voting.
24. A method according to any of claims 17 to 23, wherein the one or more images comprise or consist of an image that has been acquired using a dermoscopic camera, and / or wherein the training data comprises for each of the plurality of training lesions, an image that has been acquired using a dermoscopic camera or an image that has been acquired using a digital single-lens reflex camera.6925. The method of claim 24, wherein the training comprises for each of a first plurality of training lesions, an image that has been acquired using a dermoscopic camera, and for each of a second plurality of training lesions, an image that has been acquired using a digital single-lens reflex camera.
26. A method according to any of claims 17 to 25, wherein the method further comprises producing a saliency map associated with the machine learning model to identify the parts of the image that are more relevant for determining the likelihood of a skin lesion being suspicious or non-suspicious.
27. A method according to any preceding claim, wherein classifying, using said data, the lesion between a first class of suspicious lesions and a second class of non-suspicious lesions comprises obtaining a score indicative of the likelihood that the lesion is in the second class, optionally wherein the score is a probability.
28. The method according to claim 27, wherein the method comprises classifying the lesion between a first clinical pathway class associated with a score above a first predetermined threshold, a second clinical pathway class associated with a score at or below the first predetermined threshold and above a second predetermined threshold, and a third clinical pathway class associated with a score at or below the second predetermined threshold, optionally wherein the first predetermined threshold is selected such that a maximum predetermined number or proportion of suspicious lesion in a training cohort is classified in the first clinical pathway class, optionally wherein the maximum predetermined number or proportion is zero, and / or wherein the first predetermined threshold is the lowest predetermined threshold such that no suspicious lesion in training dataset is classified in the first clinical pathway class.
29. The method according to claim 28, wherein the method comprises selecting the patient classified in the first clinical pathway class for discharge, and / or selecting the patient classified in the second clinical pathway class for health care practitioner reporting, and / or selecting the patient classified in the third clinical pathway class for a face-to-face assessment with or without biopsy.
30. A method according to any preceding claim, further comprising selecting a patient classified in the first category for one or more further diagnostic tests, optionally wherein the further diagnostic tests comprise a biopsy.31 . A method of treating a patient who has been identified as having a skin lesion at risk of being suspicious, the method comprising:70classifying the skin lesion in the patient as suspicious or non-suspicious using the method of any preceding claim; and performing a further diagnostic test, optionally a biopsy, on the patient when the lesion has been classified as suspicious and / or when the lesion has been classified in the third clinical pathway class; and optionally treating the patient with a first treatment depending on the result of the further diagnostic test, optionally wherein the first treatment comprises a surgical removal of the lesion when the further diagnostic test indicates that the lesion is malignant or pre- malignant.
32. A system comprising one or more processors and one or more computer readable memories storing instructions that, when executed by the one or more processors, cause the one or more processors to implement the method of any preceding claim.
33. The system of claim 32, further comprising image acquisition means.
Citation Information
Patent Citations
Fusion of deep learning and handcrafted techniques in dermoscopy image analysis
US20210209754A1
Method for melanoma screening and artificial intelligence scoring
WO2024102062A1