Method of classifying cancer subtypes

A deep learning model trained on FLIM images accurately distinguishes NSCLC subtypes like adenocarcinoma and squamous cell carcinoma, overcoming the limitations of traditional histology by providing rapid and cost-effective classification.

WO2025132537A9PCT designated stage expired Publication Date: 2025-12-26THE UNIV COURT OF THE UNIV OF EDINBURGH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/087040
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-18
Filing Date
2024-12-18
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Current methods for classifying non-small cell lung cancer (NSCLC) subtypes, such as adenocarcinoma and squamous cell carcinoma, rely on time-consuming and costly histological analysis of stained tissue slices, while existing fluorescence lifetime imaging (FLIM) techniques lack consistency and discriminative power for accurate classification.

Method used

A deep learning model is trained using FLIM images to differentiate between adenocarcinoma and squamous cell carcinoma without the need for histological staining, utilizing autofluorescence lifetime images and intensity-weighted lifetime images to achieve accurate classification.

Benefits of technology

The deep learning model enables rapid and cost-effective classification of NSCLC subtypes directly from FLIM images, providing accurate results in minutes, potentially replacing traditional histological analysis and enhancing clinical efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024087040_26122025_PF_FP_ABST
    Figure EP2024087040_26122025_PF_FP_ABST
Patent Text Reader

Abstract

Methods of identifying a lung cancer subtype in a patient are described, comprising classifying one or more autofluorescence lifetime images of lung tissue of the patient between a plurality of classes associated with different subtypes of lung cancer, the plurality of classes comprising a first class associated with adenocarcinoma and a second class associated with squamous cell carcinoma, using a deep learning model Related methods and products are also described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method of classifying cancer subtypes

[0002] Field of the Disclosure

[0003] The present invention relates to a computer-implemented method of classifying cancer subtypes and particularly, although not exclusively, to methods of classifying Non-Small Cell Lung Cancer (NSCLC) cancer into adenocarcinoma (ADE), squamous cell carcinoma (SCC), and other NSCLC subtypes.

[0004] Background

[0005] Autofluorescence lifetime is a unique characteristic of the inherent fluorescent signals emitted by natural fluorophores in biological samples. Fluorescence lifetime characterises the decay of a fluorophore from the excited to the ground state. It is independent of fluorescence concentration but sensitive to the surrounding biological environment. Fluorescence lifetime imaging microscopy (FLIM) is a label-free optical technique that captures fluorescence lifetime. Differences in fluorescence lifetime, referred to as lifetime contrast, have been shown to exist between healthy / unhealthy biological tissues, enabling FLIM to be used to distinguish normal from cancerous tissues (Zheng et al., 1996).

[0006] One common approach to cancer classification is to derive the average lifetime by using statistical methods, such as a histogram of FLIM images and discriminate cancers based on lifetime difference, with the assistance of histological images (Wang et al., 2022). For example, average fluorescence lifetime of resected lung cancer tissue was shown to be significantly reduced compared with that of adjacent healthy tissue (Fernandes et al., 2021; Fernandes, 2022).

[0007] Further, machine learning-based approaches (in particular, convolutional neural networks, CNNs) have been shown to effectively enable automatic cancer classification. Machine learning (ML), and particularly deep learning (DL), has transformed conventional biomedical imaging analysis in many aspects, including in cancer classification. CNNs haven been shown to successfully discriminate between cancerous lung tissue and normal lung tissue ex vivo (Wang et al., 2020; Wang et al., 2022).

[0008] Lung cancer is the most common cause of cancer related death worldwide (Perez-Moreno et al. 2012). Lung cancer is typically differentiated between non-small cell lung cancer (NSCLC) and small cell lung cancer (SCLC). Non-small cell lung cancer is the most common type of cancer (about 80-85% of lung cancers). There are different histological subtypes of NSCLC including adenocarcinoma, squamous cell carcinoma, large cell carcinoma, large cell neuroendocrine carcinoma, adenosquamous carcinoma and sarcomatoid carcinoma. Adenocarcinoma and squamous cell carcinoma are the most frequent types of NSCLC, together representing over 80% of cases. Adenocarcinoma starts in cells that secrete mucus whereas squamous cell carcinoma starts in squamous cells. A correct histologic diagnosis is becoming increasingly important because it has been found to predict response and toxicity to therapies (Dietel et al. 2016). At present, histological subtyping is performed by expert pathologist analysis of stained tissue slices (e.g. H&E stained slices). The process is relatively slow and costly due to the manual labour involved in sample collection, processing and analysis. Summary of the Invention

[0009] In this work, the present inventors investigated the capability of FLIM for distinguishing adenocarcinoma (ADE) and squamous cell carcinoma (SCC) from other non-small cell lung cancer (NSCLC) subtypes. The present inventors postulated that this may be possible using deep learning, and that if that success in this would represent an extremely valuable improvement to current clinical practice as it would avoid the need to analyse stained images. Indeed, FLIM images are label-free and an analysis of such images using deep learning could be done within minutes of collection at minimal cost, rather than days as is the case in current practice. Additionally, FLIM images can even be acquired in vivo using fluorescence lifetime endomicroscopy.

[0010] However, prior to the present work, there was considerable uncertainty as to whether such an approach would be possible, since FLIM measurements are extremely variable from one patient to another. Indeed, while it has been previously suggested that in resected lung cancer tissues a slight difference in mean fluorescence lifetime could be observed between subtypes of lung cancer, there was generally no agreement in the field as to whether such fluorescence lifetime differences could be consistently observed. Indeed, directions of change observed are not consistent across datasets, any signal that can be observed at the level of whole FLIM images is unlikely to remain sufficiently discriminant in a realistic scenario where images may include a mixture of healthy and cancerous tissue, and is unlikely to generalise to single images that may not have been acquired using the exact same instrument and protocol. Additionally, when the present inventors attempted to quantify fluorescence lifetime properties even on single, homogeneous datasets, the difference at the average image level observed was not distinct enough to enable accurate classification of samples on this basis.

[0011] The inventors postulated that a deep learning approach to lung cancer subtyping from FLIM images might be able to achieve usable accuracy in a reproducible manner. The present inventors found, surprisingly, that it was possible to train a deep learning model to accurately differentiate between Non-Small Cell Lung Cancer (NSCLC) subtypes adenocarcinoma (ADE), squamous cell carcinoma (SCC) and other NSCLC subtypes. The deep learning-based classification methods applied to FLIM images described herein may thus be used to subtype lung cancers without the need for histological staining, thus saving time and costs. In addition, FLIM is a label-free optical imaging method and may thus be carried out ex vivo and in situ (e.g. using fluorescence lifetime imaging endoscopy).

[0012] Thus, according to a first aspect, there is provided a computer implemented method of identifying a lung cancer subtype in a patient, the method comprising: obtaining one or more previously acquired autofluorescence lifetime images of lung tissue of the patient; and classifying the one or more images between a plurality of classes associated with different subtypes of lung cancer, the plurality of classes comprising a first class associated with adenocarcinoma and a second class associated with squamous cell carcinoma, using a deep learning model that has been trained using training fluorescence lifetime images of lung tissue from a plurality of patients, the training fluorescence lifetime images comprising a first set of images of lung tissue identified as adenocarcinoma and a second set of images of lung tissue identified as squamous cell carcinoma.

[0013] Thus, the deep learning model may have been trained using training fluorescence lifetime images of lung tissue from a plurality of patients, each of the training images associated with a ground truth label selected from a plurality of labels comprising at least one label associated with adenocarcinoma and at least one label associated with squamous cell carcinoma.

[0014] The method may have any one or more of the following optional features.

[0015] For each image provided as input to the deep learning model, the deep learning model may provide as output one or more probabilities that the image provided as input shows lung tissue from the one or more respective lung cancer subtypes associated with the plurality of classes. The deep learning model may provide as output a probability for each of the plurality of classes, and the patient may be identified as having a lung cancer subtype that is associated with the class that has the highest probability amongst the probabilities for the plurality of classes.

[0016] The plurality of classes may comprise or consist of: a first class associated with adenocarcinoma, a second class associated with squamous cell carcinoma, a third class associated with lung cancer subtypes that are not adenocarcinoma or squamous cell carcinoma, and optionally a fourth class associated with normal tissue. The lung cancer may be non-small cell lung cancer. The patient may be a patient that has been diagnosed as having lung cancer.

[0017] The autofluorescence lifetime images may be composites of autofluorescence intensity and lifetime images obtained from the same fluorescence lifetime image. A composite image may be an intensity-weighted lifetime image. A composite image refers to an image where the value of each pixel depends on the value of the pixel in each of the autofluorescence intensity and corresponding lifetime images. The method may comprise obtaining one or more autofluorescence intensity images and corresponding fluorescence lifetime images, and obtaining, for each pair of autofluorescence intensity and lifetime images, an intensity weighted lifetime image. An intensity weighted lifetime image may be a false-colour lifetime image with colour depending on lifetime and the corresponding intensity image as the alpha channel, an image comprising lifetime as a first channel and intensity as a second channel, or an image obtained by multiplying pixel values in an autofluorescence intensity image by the corresponding pixel values in a corresponding fluorescence lifetime image. In embodiments, an intensity weighted lifetime image is a false-colour lifetime image with colour of each pixel depending on lifetime and the saturation of each pixel corresponding to the intensity of the corresponding pixel in the corresponding intensity image.

[0018] The one or more images may be processed or may have been processed using one or more of: thresholding of the intensity information, denoising of one or both of the lifetime and fluorescence intensity information, normalising of one or both of the lifetime and fluorescence intensity information, contrast enhancing of the intensity information and contrast enhancing of the lifetime information, optionally wherein the one or more images are processed or have been processed using at least contrast enhancing of the intensity information. Thresholding of the intensity information may comprise setting all pixels in an intensity image that are below a predetermined intensity to a value of 0. The predetermined intensity may be selected such that all pixels that are not associated with sample in the image are below the threshold. Denoising one or both of the lifetime and fluorescence intensity information may be performed as described in Wang et al. 2022. Normalising of one or both of the lifetime and fluorescence information may be performed as described in Wang et al. 2022. Contrast enhancing of the intensity information may be performed using a histogram based approach as known in the art, such as e.g. histogram equalisation, or by saturating any pixels with intensity in the top predetermined percentage of pixels in an image and in the bottom predetermined percentage of pixels in the image. For example, the top 3% of pixels (by intensity) in an image may be set to the value of the highest observed intensity. Instead or in addition to this, the bottom 1% of pixels (by intensity) in an image may be set to the value of the lowest observed intensity (e.g. 0).

[0019] The lung tissue may be an ex vivo tumour tissue sample that has been previously obtained from the patient, optionally wherein the lung tissue is from a fixed tissue sample, or a tumour microarray. The fluorescence lifetime images may have been acquired using a fluorescence lifetime imaging microscope. The fluorescence lifetime images may be images of in vivo lung tissue that have been previously obtained from the patient. The fluorescence lifetime images may have been acquired using a fluorescence lifetime endomicroscope. The fluorescence lifetime images may have been acquired using a fiber-based fluorescence lifetime imaging system.

[0020] Obtaining one or more previously acquired fluorescence lifetime images of lung tissue of the patient may comprise receiving a previously acquired fluorescence lifetime image of lung tissue of the patient and obtaining a plurality of patches from the image, wherein a patch is a subset of an image of a predetermined size, and wherein the one or more images provided as input to the deep learning model are individual patches. Receiving a previously acquired fluorescence lifetime image of lung tissue of the patient and obtaining a plurality of patches from the image may comprise stitching a plurality of tiles obtained by imaging the same sample to obtain a whole sample image, and obtaining a plurality of non-overlapping patches of a predetermined size from the whole sample image. In such embodiments, individual patches are provided as input to the deep learning model. Therefore, a classification result may be obtained for each patch independently. As such, a patient may be identified as having a plurality of subtypes of lung cancer when the plurality of patches are classified in at least two different classes. This reflects clinical reality in which some patients may have different subtypes of lung cancer. Further, in embodiments using image patches, the training images may be patches that have been obtained by processing whole sample images in a similar manner. Thus, ground truth labels for training images may be associated with individual images (patches). These may be propagated from a whole sample image annotation and may be the same for all patches derived from the same image (e.g. when no region specific annotations are available), or different depending on the patch (e.g. when region specific annotations are available).

[0021] The method may further comprise excluding any patch that includes more than a predetermined threshold proportion of background pixels. The predetermined threshold may be 50%. Background pixels may be identified as pixels associated with an intensity value below a threshold, such as e.g. 0.

[0022] The training fluorescence lifetime images may comprise a plurality of image of lung tissue and images obtained from the images by image augmentation. In embodiments, image augmentation comprises creating a flipped version of one or more of the images and / or creating a randomly rotated version of one or more of the images. For example, a set of images (e.g. one or more, or all of the original images) may be rotated by a randomly selected amount between predetermined boundaries, such as e.g. -15 to +15 degrees. As another example, a set of images (e.g. one or more, or all of the original images) may be horizontally flipped. The set of images that are flipped and / or rotated may be randomly drawn from an original set of lung tissue images included in the training fluorescence lifetime images.

[0023] In embodiments, the deep learning model has been trained using training autofluorescence lifetime images of lung tissue from a plurality of patients, the training fluorescence lifetime images comprising a first set of images of lung tissue identified as adenocarcinoma, a second set of images of lung tissue identified as squamous cell carcinoma, a third set of images of lung tissue identified as lung cancer of a type other than adenocarcinoma or squamous cell carcinoma, and optionally a fourth set of images of lung tissue identified as normal tissue. In embodiments, the deep learning model is a multiclass classifier configured to classify images between at least 3 classes comprising a first class associated with adenocarcinoma and a second class associated with squamous cell carcinoma. The use of a multiclass classifier advantageously means that images that cannot be associated with adenocarcinoma or squamous cell carcinoma can be associated with another class (which can be a single “catch all” class for other subtypes, or a plurality of classes for respective other subtypes or groups of subtypes). Indeed, while it is possible to train a deep learning model for binary classification (adenocarcinoma vs squamous cell carcinoma), using only training images of adenocarcinoma and squamous cell carcinoma, this would have lower performance in practice than a multiclass model as images of unknown subtype that the model is being used on may in fact be associated with a different subtype and may be misidentified as adenocarcinoma or squamous cell carcinoma. Alternatively, the deep learning model may comprise a plurality of binary classifiers configured jointly to classify images between said at least 3 classes.

[0024] The deep learning model may have been trained using training fluorescence lifetime images of lung tissue from a plurality of patients, the training fluorescence lifetime images comprising images of lung tissue from: at least 50 patients, at least 70 patients, at least 20 patients with lung adenocarcinoma and at least 20 patients with squamous cell carcinoma, at least 30 patients with lung adenocarcinoma and at least 20 patients with squamous cell carcinoma, at least 40 patients with lung adenocarcinoma and at least 20 patients with squamous cell carcinoma, at least 20 patients with lung adenocarcinoma, at least 20 patients with squamous cell carcinoma and at least 5 patients with a type of lung cancer other than adenocarcinoma or squamous cell carcinoma, at least 30 patients with lung adenocarcinoma, at least 20 patients with squamous cell carcinoma, and at least 10 patients with a type of lung cancer other than adenocarcinoma or squamous cell carcinoma, or at least 40 patients with lung adenocarcinoma, at least 20 patients with squamous cell carcinoma, and at least 10 patients with a type of lung cancer other than adenocarcinoma or squamous cell carcinoma.

[0025] In embodiments, the deep learning model has been trained using training fluorescence lifetime images of lung tissue from a plurality of patients, the training fluorescence lifetime images comprising: a. at least 5000 image patches from images of lung tissue with adenocarcinoma and at least 3000 image patches from images of lung tissue with squamous cell carcinoma, b. at least 7000 image patches from images of lung tissue with adenocarcinoma and at least 4000 image patches from images of lung tissue with squamous cell carcinoma, c. at least 5000 image patches from images of lung tissue with adenocarcinoma, at least 3000 image patches from images of lung tissue with squamous cell carcinoma, and at least 1000 patches from images of lung tissue with a lung cancer type other than adenocarcinoma and squamous cell carcinoma, or d. at least 7000 image patches from images of lung tissue with adenocarcinoma, at least 4000 image patches from images of lung tissue with squamous cell carcinoma, and at least 1500 patches from images of lung tissue with a lung cancer type other than adenocarcinoma and squamous cell carcinoma.

[0026] In embodiments, the deep learning model has been trained to classify images in the plurality of classes with a precision of at least 85%, or at least 90%, a recall of at least 85%, or at least 90%, and a specificity of at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, or at least 90%. In embodiments, the autofluorescence lifetime images are images that have been acquired using an excitation wavelength selected between 400nm and 500nm, between 420nm and 500nm, between 440nm and 500nm, between 460nm and 500n, between 470nm and 500nm, between 480nm and 490n, between 430nm and 450nm, or about 485 nm, or about 445nm. In embodiments, the autofluorescence lifetime images are images that have been acquired using a spectral band of emission be selected to have a lower boundary between 450 and 510nm, between 460 and 500nm, between 490nm and 510nm, between 495nm and 505nm, at 490nm, at 495nm, at 500nm, at 505nm, or at 510nm, or between 440nm and 480nm, between 450nm and 460nm, or about 460nm. In embodiments, the autofluorescence lifetime images are images that have been acquired using a spectral band of emission selected to have an upper boundary between 550nm and 650nm, between 560nm and 650nm, between 570nm and 650nm, between 580nm and 650nm, between 590nm and 650nm, between 600nm and 650nm, between 610nm and 650nm, between 620nm and 650nm, at 600nm, at 610nm, at 620nm, at 630nm or at 640nm. In embodiments, the autofluorescence lifetime images are images that have been acquired using an excitation wavelength at 485nm and a spectral emission band with a lower boundary at 500nm and an upper boundary at 640nm. In embodiments, the autofluorescence lifetime images are images that have been acquired using an excitation wavelength at 445nm and a spectral emission band with a lower boundary at 460nm and an upper boundary at 640nm.

[0027] In embodiments, the deep learning model is a deep artificial neural network. In embodiments, the deep learning model is a convolutional neural network. In embodiments, the deep learning model isa CNN with residual connections. In embodiments, the deep learning model is a CNN with dense connections. For example, the deep learning model may be a ResNet or RestNet derivative model, such as a multilevel CNN with residual connections, or a DenseNet. In embodiments, the deep learning model is a transformer based model. In embodiments, the deep learning model is a model comprising a mixture of convolutional and transformer based layers.

[0028] According to a second aspect, there is provided a computer implemented method of providing a tool for identifying a lung cancer subtype in a patient, the method comprising: obtaining a plurality of training fluorescence lifetime images of lung tissue from a plurality of patients, each of the training images e associated with a ground truth label selected from a plurality of labels comprising at least one label associated with adenocarcinoma and at least one label associated with squamous cell carcinoma; and training a deep learning model to classify fluorescence lifetime images between a plurality of classes associated with different subtypes of lung cancer using the training fluorescence lifetime images and associated ground truth labels, the plurality of classes comprising a first class associated with adenocarcinoma and a second class associated with squamous cell carcinoma, using said training images.

[0029] The method of the present aspect may have any of the features described in relation to the first aspect.

[0030] The methods of any aspect may comprise providing to a user, for example through a user interface, the results of the classification, the trained deep learning model, and / or any information derived therefrom. A data store may be a public or private database. The results of the classification may comprise one or more of: a probability of belonging to the first and / or second class obtained using the deep learning model, a classification label for one or more images, a classification label for one or more patients, a trained deep learning model, the values of parameters (e.g. architecture and weights) of a trained deep learning model. Information derived from the results of the classification may comprise one or more of: a prognostic indication derived from a classification obtained using the deep learning model, a therapeutic indication derived from a classification obtained using the deep learning model, an indication of suitability for taking part in a clinical trial derived from a classification obtained using the deep learning models, etc.

[0031] According to a third aspect, there is provided a method of selecting a subject that has been diagnosed as having lung cancer for participation in a clinical trial and / or for further diagnostic testing, the method comprising: identifying a lung cancer subtype in the subject using the method of any embodiment of the first aspect; and selecting a subject identified as having a predetermined lung cancer subtype for participating in the clinical trial and / or selecting a subject identified as having a predetermined lung cancer subtype for further diagnostic testing. The further diagnostic testing may be selected from mutation analysis and histopathology using stained tissue slices.

[0032] According to a fourth aspect, there is provided a method of providing a prognosis for a subject that has been diagnosed as having lung cancer, the method comprising: identifying a lung cancer subtype in the subject using the method of any embodiment of the first aspect; and determining a prognosis for the subject, wherein a subject identified as having adenocarcinoma has better prognosis than a subject identified as having squamous cell carcinoma.

[0033] According to a fifth aspect, there is provided a method of identifying a treatment for a subject that has been diagnosed as having lung cancer, the method comprising: identifying a lung cancer subtype in the subject using the method of any embodiment of the first aspects; and selecting a treatment for the subject based on the identified lung cancer subtype and optionally one or more predetermined clinical characteristics. Selecting a treatment for the subject may comprise determining a tumour grade, wherein the determining uses a tumour grading scheme that is dependent on the identified lung cancer subtype, and selecting a treatment associated with the determined tumour grade. According to a sixth aspect, there is provided a system for identifying lung cancer subtypes and / or providing a tool for identifying lung cancer subtypes and / or selecting a subject for participating in a clinical trial and / or selecting a subject for further diagnostic testing, and / or providing a prognosis, and / or identifying a treatment for a subject, the system comprising: one or more processors and computer readable memory storing instructions that cause the processor to perform the method of any embodiment of any preceding aspect. The system may further comprise fluorescence lifetime image data acquisition means configured to obtain fluorescence lifetime imaging data relating to one or more patients. In some embodiments, the system may comprise one or more computers, servers, or cloud-based devices, for example.

[0034] According to a seventh aspect, there is provided a non-transitory computer readable storage medium containing machine executable instructions which, when executed on a processor, cause the processor to perform any method described herein, such as the method of any embodiment of the first to fifth aspects, including any one, or any combination insofar as they are compatible, of the optional features set out with reference thereto.

[0035] According to an eighth aspect, there is provided a computer program comprising executable code which, when run on a computer, causes the computer to perform any method described herein, such as the method of any embodiment of any of the first to fifth aspects, including any one, or any combination insofar as they are compatible, of the optional features set out with reference thereto.

[0036] The invention includes the combination of the aspects and preferred features described except where such a combination is clearly impermissible or expressly avoided.

[0037] Summary of the Figures

[0038] Embodiments and experiments illustrating the principles of the invention will now be discussed with reference to the accompanying figures in which:

[0039] Figure 1 shows a flow diagram of a method for identifying lung cancer subtypes in a subject according to the disclosure.

[0040] Figure 2 shows a flow diagram of a method for providing a tool for identifying lung cancer subtypes in a subject.

[0041] Figure 3 shows an embodiment of a system for analysing FLIM images and / or for providing a tool for FLIM medical images, according to the present disclosure.

[0042] Figure 4 shows lifetime images reconstructed by an exponential curve fitting raw intensity of unstained lung samples. A. shows the raw intensity image. B. shows the lifetime image exported from the Leica LAX-X software. C. shows the false-colour lifetime with the normalised intensity image as the alpha channel.

[0043] Figure 5 shows the average lifetime of NSCLC subtypes, including 3 cases of adenocarcinoma (ADC), 3 cases of squamous cell carcinoma (SCC), and 2 cases of other NSCLC subtypes.

[0044] Figure 6 illustrates schematically a method of analysing FLIM images used in examples of the disclosure. Figure 7 illustrates artificial neural network architectures usable in the context of the disclosure. A. ResNet residual block. B. ResNetZ module which performs multi-scale feature extraction on the input features as a whole, where A is an aggregation operator. C. Res2Net module which splits input features into groups and performs multi-scale feature extraction per grouped features.

[0045] Figure 8 shows comparisons of using lifetime and intensity images to train the DenseNet DNN model. Figure 8A shows results of three binary classifications using intensity weighted lifetime images, from the top: (a) cancer & non-cancer, (b) ADE & (SqCC + OS), and (c) SqCC & OS demonstrated with ROC and confusion matrices. Figure 8B shows results of three binary classifications using intensity only images, from the top: (a) cancer & non-cancer, (b) ADE & (SqCC + OS), and (c) SqCC & OS demonstrated with ROC and confusion matrices.

[0046] Figure 9 shows AUC comparisons of three models (a) DenseNet, (b) ResNet, and (c) EfficientNet for multiclassification using intensity weighted lifetime images.

[0047] Figure 10 shows accuracy comparisons using ROCs and AUC scores for multiclass classification using a DenseNet model trained with intensity weighted lifetime images (ILI-DensNet) and a DenseNet model trained with intensity only images (ITI-DenseNet). ROCs and AUC scores are used to evaluate the performance. Fig. 10a shows ROC curves for the ILI-DenseNet model. Fig. 10c shows a confusion matrix for the ILI-DenseNet model. Fig. 10b shows ROC curves for the ITI-DenseNet model. Fig. 10d shows a confusion matrix for the ITI-DenseNet model.

[0048] Figure 11 shows results of core-based performance evaluation of multiclass classification models using DenseNet trained on intensity weighted lifetime images (ILI-DenseNet) and on intensity only images (ITIs- DenseNet). Fig, 11(a) shows a radar plot with five indicators for both lifetime- and intensity-based models. Fig. 11(b) shows violin plots for averaged probability per core of the lifetime-based model. Fig. 11(c) shows violin plots for averaged probability per core of the intensity-based models. Each dot on (b) and (c) is the averaged probability predicted by the model for the core in the ground truth class associated with the core.

[0049] Detailed Description of the Invention

[0050] Aspects and embodiments of the present invention will now be discussed with reference to the accompanying figures. Further aspects and embodiments will be apparent to those skilled in the art. All documents mentioned in this text are incorporated herein by reference.

[0051] The methods described herein are computer implemented unless context indicates otherwise. Indeed, image analysis using deep learning models, and the process of training deep learning models is of a complexity, and in particular requires the analysis of large amounts of data through complex mathematics, that places the methods described herein far beyond the capability of mental investigation. Thus, any method described herein may be implemented in a computer system (in particular in computer hardware or in computer software) in addition to the structural components and user interactions described.

[0052] The term “computer system” includes the hardware, software and data storage devices for embodying a system or carrying out a method according to the above-described embodiments. For example, a computer system may comprise one or more processing units such as central processing units (CPU) and / or graphics processing units (GPU), input means, output means and data storage. Preferably the computer system has a monitor to provide a visual output display. The data storage may comprise RAM, disk drives or other computer readable media. The computer system may include a plurality of computing devices connected by a network and able to communicate with each other over that network. It is explicitly envisaged that the computer system may consist of or comprise a cloud computer.

[0053] The methods described herein may be provided as computer programs or as computer program products or computer readable media carrying a computer program which is arranged, when run on a computer, to perform the method(s) described herein. The term “computer readable media” includes, without limitation, any non-transitory medium or media which can be read and accessed directly by a computer or computer system. The media can include, but are not limited to, magnetic storage media such as floppy discs, hard disc storage media and magnetic tape; optical storage media such as optical discs or CD-ROMs; electrical storage media such as memory, including RAM, ROM and flash memory; and hybrids and combinations of the above such as magnetic / optical storage media.

[0054] The features disclosed in the following description, or in the following claims, or in the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for obtaining the disclosed results, as appropriate, may, separately, or in any combination of such features, be utilised for realising the invention in diverse forms thereof.

[0055] While the invention will be described in conjunction with the exemplary embodiments described below, many equivalent modifications and variations will be apparent to those skilled in the art when given this disclosure. Accordingly, the exemplary embodiments of the invention set forth below are considered to be illustrative and not limiting. Various changes to the described embodiments may be made without departing from the spirit and scope of the invention.

[0056] For the avoidance of any doubt, any theoretical explanations provided herein are provided for the purposes of improving the understanding of a reader. The inventors do not wish to be bound by any of these theoretical explanations. Any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0057] Throughout this specification, including the claims which follow, unless the context requires otherwise, the word “comprise” and “include”, and variations such as “comprises”, “comprising”, and “including” will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. As used in the specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by the use of the antecedent “about,” it will be understood that the particular value forms another embodiment. The term “about” in relation to a numerical value is optional and means for example + / - 10%. The expression “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example “A and / or B” is to be taken as specific disclosure of each of (i) A, (ii) B and (Hi) A and B, just as if each is set out individually herein.

[0058] Identifying lung cancer subtype from FLIM images

[0059] The present disclosure relates to the use of a deep learning model for identification of lung cancer subtypes from FLIM images.

[0060] As used herein, the term “deep learning model”, also referred to as “deep artificial neural network” or “deep neural network” refers to a machine learning model that has a neural network architecture with one or more hidden layers. Any DL model suitable for image analysis may be used in the context of the present disclosure. The DL model may be a convolutional neural network (CNN). Convolutional neural networks have been shown to perform particularly well at image recognition tasks. In embodiments, the DL model is a CNN that is trained entirely on FLI M images. The DL model may have been trained directly for the task of identifying lung cancer subtypes from FLIM images. Thus, the DL may have been initiated with random parameters that are iteratively trained for a lung cancer subtype identification task in FLIM images. Alternatively, the CNN may have been pre-trained on unrelated image data. For example, CNNs that have been pre-trained for image recognition tasks on large collections of image data such as the ImageNet database are available. The use of models trained “from scratch” for analysis of FLIM images may be particularly advantageous as FLIM images do not have the same visual features as common images used for object recognition tasks, such that pretraining on non-FLIM images may not be particularly informative. In embodiments, the CNN is a 50 layers CNN. In embodiments, the CNN is a CNN that has been trained using a deep residual learning framework. Deep residual learning is a learning framework that has been developed for image recognition, to address the problem known as “degradation” (the observation that as the network depth increases, the accuracy saturates then degrades rapidly). More detail can be found in Het et al. 2016. In embodiments, the CNN is a pre-trained network that has been trained using deep residual learning, also known as ResNets. The architecture of a residual block used in ResNet models is shown on Figure 7A. Derivatives of this model including ResNeXt (described in Xie et al. 2017), Res2Net (described in Gao et al. 2021 , architecture illustrated on Figure 7C), and ResNetZ (described in Wang et al. 2022, architecture illustrated in Figure 7B) can be used in embodiments of the disclosure. In embodiments, the CNN is ResNet50. ResNet50 is a 50-layers deep CNN comprising residual blocks as illustrated on Figure 7 A. In embodiments, the DL model is a multiscale model. A multiscale model is a model that includes processing operations applied at different resolutions of an input image. In embodiments, the DL model is a multiscale CNN. Multiscale CNNs have an architecture comprising a number of single / composite operations in parallel at different levels (resolution). Examples of multiscale CNNs include Inception, Res2Net and ResNetZ. Multiscale models are able to simultaneously extract features at different scales and integrate these features for improved performance. In embodiment, the DL model is a densely connected CNN. A densely connected CNN may be a CNN that uses dense connections in one or more or all convolutional layers. In embodiments, the DL model is a CNN that uses dense connections in all convolutional layers, such as e.g. DenseNet (Huang et al. 2018). Any of the DenseNet architectures described in Huang et al. 2018 may be used in the context of the present disclosure, such as e.g. DenseNet- 121 , DenseNet-169, DenseNet-201 , and DenseNet-264. In embodiments, the DL model is a CNN that uses a compound scaling method that uniformly scales all 3 network dimensions (width, depth and resolution) with a fixed ratio (i.e. a set of fixed scaled coefficients). In embodiments the DL model is an EfficientNet, as described in Tan & Le (2020). Any of the architectures described in Tan & Le 2020 may be used, such as e.g. any of EfficientNet-BO to EfficientNet-B7. In embodiments, the deep learning model is a transformerbased model. A transformer-based model is a model that comprises one or more transformer layers. A transformer layer is a layer that processes inputs using only attention mechanisms. In embodiments, the deep learning model is a model comprising a mixture of convolutional and transformer-based layers. Any deep neural network architecture suitable for image classification may be used, such as e.g. squeezenet (landola et al., arxiv.org / abs / 1602.07360), googlenet, inceptionv3 (Szegedy et al., Proc of IEEE conf comp vis pat recog, pp.1-9, 2015), densenet201 (Huang et al., CVPR vol. 1, no. 2, p / 3, 2017), resnet-50 or -101 (He et al., 2016), efficientnetbO (Migxing Tan and Le, Arxiv: 1905.1194, 2019), alexnet (Krizhevsky et al., Adv neur info proc sys, 2012), vgg16 (Simonyan and Ziserman, arxiv: 1409.1556, 2014), ViT (vision transformer, described in Dosovitskiy et al., 2021) and derivatives thereof such as the Swin transformer (Liu et al. 2021 ), ConvNext (Liu et al. 2022), etc. ResNet and DenseNet architectures have been shown by the inventor to perform particularly well in the present context. The DL model may be a multiclass classifier. A multiclass classifier is a machine learning model configured to classify an input image between a plurality of classes. For example, a multiclass classifier may take an image as input and provide as output a respective probability for each of a plurality of classes, indicating the likelihood that the image belongs to the respective class. Thus, the probabilities typically sum to 1. The image may be assigned to the class that is associated with the highest probability. Additional criteria may apply, such as e.g. a threshold probability, such that a classification is only deemed to have been obtained if the highest class probability is above a predetermined threshold (e.g. 0.4, 0.5 or 0.6). Images that are not associated with any probability above such a threshold may be considered as not classified with sufficient confidence. In embodiments, an image is simply assigned the class label of the class with highest probability. The plurality of classes may include a plurality of classes selected from: a class associated with normal tissue, a class associated with adenocarcinoma, a class associated with squamous cell carcinoma, and a class associated with other cancer types (e.g. lung cancer subtypes other than adenocarcinoma and squamous cell carcinoma). A multiclass classifier may classify input images between at least 4 classes including a class associated with normal tissue, a class associated with adenocarcinoma, a class associated with squamous cell carcinoma, and a class associated with other cancer types. The classifier may be a binary classifier or may comprise a plurality of binary classifiers. A binary classifier is a machine learning model configured to classify an input image between two classes. For example, a binary classifier may be configured to classify images between a first class associated with Adenocarcinoma and a second class associated with other types of lung cancer. As another example, a binary classifier may be configured to classify images between a first class associated with squamous cell carcinoma and a second class associated with other types of lung cancer. Further, a binary classifier may be configured to classify images between a first class associated with cancer and a second class associated with normal tissue. A binary classifier may output a single probability associated with one of the two classes (the probability associated with the other class being obtainable as 1-probability associated with the one of the two classes). An image or sample may be classified in the first class when the probability output by the classifier exceeds a predetermined threshold, and in the second class otherwise. The classifier may comprise a plurality of classifiers including one or more binary classifiers and / or one or more multiclass classifiers. For example, the classifier may include a first binary classifier trained to classify input images between a first class associated with normal tissue and a second class associated with cancerous tissue, and one or more further classifiers trained to classify input images between respective classes associated with different cancer types. For example, the one or more further classifiers may comprise a second binary classifier trained to classify input images between a first class associated with adenocarcinoma and a second class associated with other cancer subtypes, and a third binary classifier trained to classify input images between a first class associated with squamous cell carcinoma and a second class associated with other cancer subtypes. Alternatively, the one or more further classifiers may comprise a second multiclass classifier trained to classify input images between a first class associated with adenocarcinoma, a second class associated with squamous cell carcinoma, and a third class associated with other cancer subtypes (i.e. cancer tissue that is neither adenocarcinoma nor squamous cell carcinoma). In embodiments, an input image classified in the second class associated with cancerous tissue by the first binary classifier may be processed by the one or more further classifiers to assign a cancer subtype to the image. For example, the image may be provided as input to the second multiclass classifier. Alternatively, the image may be provided as input to the second binary classifier, and further as input to the third binary classifier if the image is classified in the second class by the second binary classifier (or alternatively using the third binary classifier then the second binary classifier).

[0061] As used herein, the terms “lung cancer type” or “lung cancer subtype” refers to histologic types of lung cancer. In embodiments, the lung cancer is a NSCLC, and therefore the lung cancer types are histological subtypes of NSCLC. Histological subtypes of NSCLC include squamous cell carcinoma (SCC) of the lung, also known as squamous cell lung cancer or lung squamous cell carcinoma (LUSC), adenocarcinoma (ADC or ADE) also known as lung adenocarcinoma (LUAD), and other subtypes that are less frequent including large cell carcinoma, large cell neuroendocrine carcinoma, adenosquamous carcinoma and sarcomatoid carcinoma. Embodiments of the methods described herein are able to identify whether a subject or sample shows evidence of a lung cancer subtype selected from a predetermined plurality of subtypes that include squamous cell carcinoma and adenocarcinoma. The plurality of subtypes may further include one or more additional subtypes and / or a “catch all” other category for subtypes that are any subtype other than squamous cell lung cancer, adenocarcinoma and any other subtype individually identified. SCC and ADC subtypes together represent the vast majority of lung cancer cases, and therefore subtyping between these types already represents an extremely useful clinical tool in the context of lung cancer as a whole.

[0062] As used herein, the term “fluorescence lifetime imaging (endo)microscopy” (FLIM) refers to a technology that acquires time resolved fluorescence spectra. In the context of the present disclosure, the fluorescence spectra are acquired in a label-free manner, and therefore the fluorescence is autofluorescence of biological tissues, rather than fluorescence associated with fluorescent labels. FLIM enables the quantification of both fluorescence intensity and fluorescence lifetime, where the latter is obtained by estimating a fluorescence decay curve from time resolved fluorescence data by exponential curve fitting. In particular, a fluorescence decay curve is typically assumed to be an exponential function lo(t)=lo Exp(-t / T) where T is the fluorescence lifetime and Io is an initial maximal fluorescence. Estimation of fluorescence lifetime from time resolved fluorescence spectra is known in the art, and commercial devices and software exist to acquire time resolved fluorescence spectra and obtain lifetime from these. A FLIM device as used herein may be a microscope, such as e.g. a confocal microscope (e.g. Leica TCS SP8 Confocal Microscope) or an endoscope, such as e.g. a fiber based FLIM system as described in Wang et al. 2022. FLIM images are typically acquired using a single excitation wavelength and a spectral band of emission. The excitation wavelength may be set to an excitation wavelength that has been identified as optimal for the particular tissue type and instrument, for example through a A-to-A scan. The excitation wavelength may be an excitation wavelength selected between 400nm and 500nm, between 420nm and 500n, between 430nm and 500nm, between 440nm and 490n, about 445nm or about 485 nm. The emission wavelengths (spectral band of emission) may be selected to have a lower boundary between 450nm and 510nm, between 460nm and 510nm, between 450nm and 500nm, between 495nm and 505nm, such as e.g. 490nm, 495nm, 500nm, 505nm, or 510nm, or between 450 and 470nm, such as e,g 450nm, 455nm, 460nm, 465nm, or 470nm. The emission wavelengths (spectral band of emission) may be selected to have an upper boundary between 550nm and 650nm, between 560nm and 650nm, between 570nm and 650nm, between 580nm and 650nm, between 590nm and 650nm, between 600nm and 650nm, between 610nm and 650nm, between 620nm and 650nm, such as e.g. 600nm, 610nm, 620nm, 630nm or 640nm. Any of the above lower boundaries are explicitly envisaged in combination with any of the above upper boundaries. For example, the lower boundary may be between 490nm and 510nm, such as 500m, and the upper boundary may be between 600nm and 650nm, such as e.g. 640nm. As another example, the lower boundary may be between 450 and 500nm, such as 460nm, and the upper boundary may be between 600nm and 650nm, such as e.g. 640nm. Further, images that have been acquired using different excitation and / or emission wavelengths in the above ranges may be used together e.g. for training a model as described herein. Similarly, an image that has been acquired using different excitation and / or emission wavelengths in the above ranges from the excitation and / or emission wavelengths in the above ranges associated with one or more image that have been used to train a model as described herein may be analysed with the trained model. In other words, models as described herein may be trained and / or used using images that have been acquired using any of the excitation wavelengths and spectral emission bands described herein. The methods described herein use FLIM images of lung cancer tissue. The images may be acquired ex vivo, such as e.g. using lung cancer tissue from biopsies, resected tumours (such as e.g. tumour microarray samples). Such images may be acquired using a conventional FLIM imaging system. Alternatively, the images may be images that have been acquired in vivo (i.e. in situ), for example using a fiber based FLIM system.

[0063] The images used herein are images of lung tissue comprising lung cancer tissue, the images having been previously obtained from a subject. The terms “subject and “patient” are used herein interchangeably. The subject is typically a human subject. The subject may be a subject who has been diagnosed as having or being likely to have lung cancer. The subject is typically a subject who has not yet been treated for lung cancer (treatment naive). The subject may be a subject who has been diagnosed as having or being likely to have lung cancer (e,g, using imaging technologies) and who has been subject to a tissue biopsy (e.g. diagnostic biopsy) for further characterisation of the lung cancer (e.g. when the methods described herein are performed using images of ex vivo tissue samples). Alternatively, the subject may be a subject who has been diagnosed as having lung cancer and who has been subject to endomicroscopy using a fiber based FLIM system.

[0064] Figure 1 shows a flow diagram of a method for identifying lung cancer subtypes in a subject according to the disclosure. At optional step 100, a lung tissue sample may be obtained from the subject, such as e.g. by performing a tissue biopsy or receiving a previously obtained tissue biopsy. This is optional as the methods described herein may start from previously obtained images and / or may use images obtained by endomicroscopy. At optional step 110, one or more fluorescence lifetime images of lung cancer tissue in the sample obtained by imaging the sample obtained at step 100. This step is optional because the methods described herein may start from previously obtained images. The method comprises a step 112 of obtaining one or more previously acquired fluorescence lifetime images of lung tissue of the patient. A fluorescence lifetime image is typically acquired as part of a dataset that comprises two sets of information for each location (pixel) on the image: fluorescence intensity (also referred to as autofluorescence intensity image, fluorescence intensity image, intensity image or intensity information), and lifetime (also referred to as autofluorescence lifetime image, fluorescence lifetime image, lifetime image or lifetime information). Thus, step 110 typically comprises obtaining both of these sets of information.

[0065] Step 112 as illustrated comprises an optional step 112A of preprocessing the images obtained. This step may be performed at the level of whole sample images, tiles or patches. Pre-processing may use one or more of: thresholding of the intensity information, denoising of one or both of the lifetime and fluorescence intensity information, normalising of one or both of the lifetime and fluorescence intensity information and contrast enhancing ofthe one or both of the lifetime and fluorescence intensity information. In embodiments, the one or more images are processed or have been processed using at least contrast enhancing of the intensity information. Thresholding ofthe intensity information may comprise setting all pixels in an intensity image that are below a predetermined intensity to a value of 0. The predetermined intensity may be selected such that all pixels that are not associated with sample in the image are below the threshold. Denoising one or both of the lifetime and fluorescence intensity information may be performed as described in Wang et al. 2022. Normalising of one or both of the lifetime and fluorescence information may be performed as described in Wang et al. 2022. Contrast enhancing of the intensity or lifetime information may be performed using a histogram-based approach as known in the art, such as e.g. histogram equalisation, or by saturating any pixels with intensity in the top predetermined percentage of pixels in an image and in the bottom predetermined percentage of pixels in the image. For example, the top 3% of pixels (by intensity) in an image may be set to the value of the highest observed intensity. Instead or in addition to this, the bottom 1% of pixels (by intensity) in an image may be set to the value of the lowest observed intensity (e.g. 0). In embodiments, step 112A only comprises contrast enhancing the intensity images. Step 112A may not include further processing of the intensity images after they have been obtained. Whether or not an image (whether the intensity or lifetime image) benefits from processing such as e.g. with contrast enhancing depends on factors such as the quality and signal to noise ratio of the images.

[0066] At optional step 112B, a composite image is obtained for each fluorescence lifetime image obtained, which is a composites of autofluorescence intensity and lifetime images obtained from the same fluorescence lifetime image (i.e. from the dataset comprising corresponding intensity and lifetime information). As illustrated, the composite is an intensity-weighted lifetime image. The intensity weighted lifetime image may be a false-colour lifetime image with colour depending on lifetime and the corresponding intensity image as the alpha channel. This is believed to be particularly advantageous as resulting in improved classification performance compared to other types of composite or non-composite lifetime images. Thus, an intensity weighted lifetime image may be a false-colour lifetime image with colour of each pixel depending on lifetime and the saturation of each pixel corresponding to the intensity of the corresponding pixel in the corresponding intensity image. Alternatively, the intensity weighted lifetime image may be an image comprising lifetime as a first channel and intensity as a second channel. Alternatively, the intensity weighted lifetime image may be an image obtained by multiplying pixel values in an autofluorescence intensity image by the corresponding pixel values in a corresponding fluorescence lifetime image.

[0067] Step 112 may comprise a step 112C of obtaining a plurality of patches from a previously acquired fluorescence lifetime image of lung tissue of the patient, wherein a patch is a subset of an image of a predetermined size. This may comprise stitching a plurality of tiles obtained by imaging the same sample or subject to obtain a whole sample image and obtaining a plurality of non-overlapping patches of a predetermined size from the whole sample image. In such embodiments, individual patches are provided as input to the deep learning model (at step 114). Therefore, a classification result may be obtained for each patch independently at step 114.

[0068] Step 112 as illustrated further comprises a step 112D of filtering patches to remove patches that satisfy one or more predetermined criteria. For example, step 112D may comprise excluding (i.e. filtering out) any patch that includes more than a predetermined threshold proportion (e.g. 50%, 60%, 70% or 80%) of background pixels. Background pixels may be identified as pixels associated with an intensity value below a threshold, such as e.g. 0. In other words, the one or more criteria may include a criterion that the proportion of background pixels (e.g. pixels with an intensity of 0) is at or above a predetermined threshold.

[0069] At step 114, the images obtained at step 112 are classified between a plurality of classes associated with different subtypes of lung cancer, the plurality of classes comprising a first class associated with adenocarcinoma and a second class associated with squamous cell carcinoma, using a deep learning model that has been trained using training fluorescence lifetime images of lung tissue from a plurality of patients, the training fluorescence lifetime images comprising a first set of images of lung tissue identified as adenocarcinoma and a second set of images of lung tissue identified as squamous cell carcinoma. For each image provided as input to the deep learning model, the deep learning model may provide as output one or more probabilities that the image provided as input shows lung tissue from the one or more respective lung cancer subtypes associated with the plurality of classes. The output may be a single value for each class for the whole image, or individual values for respective pixels. In embodiments, the deep learning model provides as output, for each image provided as input, a probability for each of the plurality of classes. In such embodiments, the patient may be identified as having a lung cancer subtype that is associated with the class that has the highest probability amongst the probabilities for the plurality of classes. The identification may be subject to one of more additional criteria, such as e.g. a criterion that the highest probability is above a predetermined threshold. The results of step 114 (e.g. classification labels or probabilities that the image provided as input shows lung tissue from the one or more respective lung cancer subtypes associated with the plurality of classes) may be summarised across a plurality of images (patches) from the same image. For example, a summarised probability may be obtained for each class based on the respective probabilities associated with the patches. A summarised probability may be e.g. an average or median, or trimmed versions thereof. The summarised probabilities may be used to classify the image from which the patches were extracted, as explained elsewhere herein (e.g. by comparing the probabilities to thresholds and / or selecting the class with highest probability). Similarly, classification labels associated with patches may also be summarised for an image from which the patches were extracted using e.g. a majority voting approach in which the classification label most represented is associated with the image. In embodiments, the proportions of patches of an image classified in each of one or more classes may be reported as a classification result for the image.

[0070] The plurality of classes may comprise or consists of: a first class associated with adenocarcinoma, a second class associated with squamous cell carcinoma, and a third class associated with lung cancer subtypes that are not adenocarcinoma or squamous cell carcinoma. Thus, the deep learning model may be a multiclass classifier configured to classify images between at least 3 classes comprising a first class associated with adenocarcinoma and a second class associated with squamous cell carcinoma. Alternatively, the deep learning model may comprise a plurality of classifiers configured to, collectively, classify images between said at least 3 classes.

[0071] Steps 100 to 114 may be used as parts of a method of selecting a subject that has been diagnosed as having lung cancer for participation in a clinical trial. Such methods may comprise step 120 of selecting a subject identified as having a predetermined lung cancer subtype for participating in the clinical trial.

[0072] Steps 100 to 114 may be used as parts of a method of selecting a subject identified as having a predetermined lung cancer subtype for further diagnostic testing. Such methods may comprise optional step 118 of identifying a further diagnostic test to be performed for the subject based on the classification result at step 114. Further, the one or more tests identified at step 119 may be performed at optional step 119. For example, the further diagnostic testing may be selected from mutation analysis and histopathology using stained tissue slices.

[0073] Steps 100 to 114 may be used as parts of a method of providing a prognosis for a subject that has been diagnosed as having lung cancer. Such a method may comprise optional step 116 of determining a prognosis for the subject, wherein a subject identified as having adenocarcinoma has better prognosis than a subject identified as having squamous cell carcinoma. Thus, step 116 may comprise identifying a first prognosis for a subject classified as having adenocarcinoma and a second prognosis different form the first prognosis for a subject classified as having squamous cell carcinoma. At optional step 117 the subject may be treated using a treatment that depends on the prognosis identified at step 116.

[0074] Steps 100 to 114 may be used as parts of a method of identifying a treatment for a subject that has been diagnosed as having lung cancer. Such a method may comprise step 116 of selecting a treatment for the subject based on the identified lung cancer subtype and optionally one or more predetermined clinical characteristics. Selecting a treatment for the subject may comprise determining a tumour grade, wherein the determining uses a tumour grading scheme that is dependent on the identified lung cancer subtype, and selecting a treatment associated with the determined tumour grade. At optional step 117 the subject may be treated using a treatment identified at step 116.

[0075] At optional step 121, the results of any one or more of the preceding steps may be provided to a user, such as e.g. through a user interface.

[0076] Current clinical practice in the context of lung cancer includes a step of stratifying patients with NSCLC subtypes using histologically stained images. This allows patients to be selected for different downstream diagnostic steps and pathological analyses, all of which ultimately enable patients to be selected to receive specific therapies. Therefore, the cornerstone of treatment decision is based on the classification between lung cancer subtypes. For example, once patients are diagnosed as having adenocarcinoma, the current clinical practice involves selecting these patients for mutation detection to determine whether they will benefit from targeted therapies. The methods of the present disclosure can be used to replace this initial histological analysis, leading to an accurate and crucially significantly faster and less labour-intensive stratification step from which all subsequent clinical steps are derived.

[0077] The methods of the present disclosure find use in the treatment, prognosis and management of patients with lung cancer, including selection of patients for participating in a clinical trial based on lung cancer subtype.

[0078] Thus, also described herein is a method of selecting a subject that has been diagnosed as having lung cancer for participation in a clinical trial, the method comprising: identifying a lung cancer subtype in the subject using a method as described herein; and selecting or excluding the subject from participation in the clinical trial depending on whether the subject was classified as having a first lung cancer subtype or a second lung cancer subtype.

[0079] Additionally, LUAD and LUSC are known to have different response to treatments and different prognostic profiles (see e.g. Relli et al. 2019). For example, patients with LUSC were found to have significantly poorer overall survival than patients with LUAD (Akikazu et al. 2012). Further, treatment of lung cancer currently depends on the identified stage, which depends on the histological subtype identified.

[0080] Thus, also described herein is a method of providing a prognosis for a patient that has been diagnosed as having lung cancer, the method comprising: identifying a lung cancer subtype in the subject using a method as described herein; and determining a prognosis based on the classification. A patient may be classified in a first class that has a good prognosis when they are classified as having adenocarcinoma, and in a second class that has poor prognosis when they are classified as having squamous cell carcinoma. The first class may be associated with an expected overall survival that is higher than the second class.

[0081] Further, also described herein is a method of selecting a patient that has been diagnosed as having lung cancer for further diagnostic tests, the method comprising: identifying a lung cancer subtype in the subject using a method as described herein; and selecting a patient that has been classified as having lung adenocarcinoma or squamous cell carcinoma for mutation testing, such as e.g. testing for the presence of one or more genetic alterations (e.g. mutations or gene expression alterations in tumour driver or suppressor genes) which may be associated with response to therapy. For example, a patient that has been classified using a method as described herein as having adenocarcinoma may be selected for further testing to identify whether the patient has one or more alterations (e.g. mutations including fusions, or gene expression alterations) in one or more genes selected from: EGFR, KRAS, ALK, MET, PIK3CA, BRAF, ROS1 , RET, NTRK, PTEN and HER2. As another example, a patient that has been classified using a method as described herein as having squamous cell carcinoma may be selected for further testing to identify whether the patient has one or more mutations in one or more genes selected from: FGFR1 , PIK3CA, and MET. Diagnostic tests for mutational status of one or more genes may comprise obtaining a tumour tissue biopsy (e.g. by bronchoscopy of transthoracic needle biopsy), and analysing the sample using sequencing, fluorescence in-situ testing (FISH), or immunohistochemistry, or obtaining a liquid biopsy such as a blood sample and analysing the sample using sequencing. There are approved treatments for lung cancer with alterations in EGFR, ALK, ROS1 , BRAF, HER2, NTRK, MET, RET and KRAS, and treatments in clinical trial for alterations in at least PIK3CA, AKT1 and PTEN.

[0082] Further, also described herein is a method of identifying a treatment or treating a patient that has been diagnosed as having lung cancer, the method comprising: identifying a lung cancer subtype in the subject using a method as described herein; and identifying a treatment based on the classification. For example, a patient identified as having adenocarcinoma may be selected for treatment with a therapy approved for the treatment of adenocarcinoma. A patient identified as having squamous cell carcinoma may be selected for treatment with a therapy approved for the treatment of squamous cell carcinoma. For example, a patient identified as having adenocarcinoma is more likely to respond to anti-VEGF therapy (e.g. bevacizumab) than a patient identified as having LUSC. As another example, targeted therapies for cancers bearing ALK rearrangements, ROS1 fusions and BRAF mutations exist, all of which are more likely to be present in adenocarcinoma, such that a patient identified as having adenocarcinoma is more likely to respond to these therapies. Further, the anti-EGFR necitumumab was identified as more likely to be effective in LUSC than LUAD. Identifying a treatment may comprise identifying a cancer stage based at least in part on the classification, and selecting a treatment recommended for the identified cancer stage. For example, the American Society of Clinical Oncology (ASCO) recommends different therapies for stage IV non-small cell lung cancer depending on the subtype. While first line therapies typically depend on mutational status, second line therapies recommended include docetaxel, erlotinib, gefitinib, or pemetrexed for patients with LUAD; docetaxel, erlotinib, or gefitinib for those with LUSC. Further, response to immune checkpoint inhibitors has been shown to be dependent on lung cancer subtype, leading to approval of therapeutics specifically in LUAD and not LUSC. For example, Atezolizumab in combination with carboplatin / paclitaxel / bevacizumab, has been granted approval in untreated LUAD patients, but not in LUSC patients. Thus, a patient identified using a method described herein as having LUAD may be selected for treatment with an immune checkpoint inhibitor.

[0083] Identifying a treatment may comprise identifying a likely genetic alteration based at least in part on the classification, and / or obtaining results of one or more genetic alteration tests based at least in part on the classification, and identifying a treatment associated with a genetic alteration identified as likely to be present in the patient. Obtaining results of one or more genetic alteration tests may comprise receiving test results from a user or computing device. Alternatively, the method may further comprise analysing a sample obtained from the patient to determine the presence of one or more genetic alterations. The one or more genetic alterations may have been selected based on the classification. The treatment may be selected based on a combination of the classification and the identified genetic alterations.

[0084] The method may further comprise administering the selected treatment.

[0085] Obtaining a tool for lung cancer subtyping

[0086] Figure 2 shows a flow diagram of a method for providing a tool for identifying lung cancer subtypes in a subject. In particular, the tool may refer to a computer implemented tool comprising a trained deep learning model for use as described herein such as e.g. by reference to Figure 1.

[0087] The method comprises step 212 of obtaining a plurality of training fluorescence lifetime images of lung tissue from a plurality of patients. Each of the training images may be associated with a ground truth label selected from a plurality of labels comprising at least one label associated with adenocarcinoma and at least one label associated with squamous cell carcinoma. The training images may have been obtained at step 212A using steps as described in Figure 1 , in particular by reference to steps 100 to 112. Step 212 may further comprise step 212B of obtaining additional images from the images obtained at step 212A, using image augmentation. For example, step 212B may comprise creating a flipped version of one or more of the images and / or creating a randomly rotated version of one or more of the images. For example, a set of images (e.g. one or more, or all of the original images) may be rotated by a randomly selected amount between predetermined boundaries, such as e.g. -15 to +15 degrees. As another example, a set of images (e.g. one or more, or all of the original images) may be horizontally flipped. The set of images that are flipped and / or rotated may be randomly drawn from an original set of lung tissue images included in the training fluorescence lifetime images. The training images may comprise a first set of images of lung tissue identified as adenocarcinoma, a second set of images of lung tissue identified as squamous cell carcinoma, and a third set of images of lung tissue identified as lung cancer of a type other than adenocarcinoma or squamous cell carcinoma. In such embodiments, the DL model may be a multiclass classifier (such as in particular a 3 classes classifier). The training images may comprise images from at least 50 or at least 70 patients. The training images may comprise images from at least 20 patients with lung adenocarcinoma and at least 20 patients with squamous cell carcinoma, at least 30 patients with lung adenocarcinoma and at least 20 patients with squamous cell carcinoma, at least 40 patients with lung adenocarcinoma and at least 20 patients with squamous cell carcinoma, at least 20 patients with lung adenocarcinoma, at least 20 patients with squamous cell carcinoma and at least 5 patients with a type of lung cancer other than adenocarcinoma or squamous cell carcinoma, at least 30 patients with lung adenocarcinoma, at least 20 patients with squamous cell carcinoma, and at least 10 patients with a type of lung cancer other than adenocarcinoma or squamous cell carcinoma, or at least 40 patients with lung adenocarcinoma, at least 20 patients with squamous cell carcinoma, and at least 10 patients with a type of lung cancer other than adenocarcinoma or squamous cell carcinoma. The training image may be image patches as described above, Thus, obtaining the training fluorescence images at step 212 may comprise obtaining patches as described by reference to Figure 1 (and optionally filtering said training patches). Thus, the training fluorescence lifetime images may comprise: at least 5000 image patches from images of lung tissue with adenocarcinoma and at least 3000 image patches from images of lung tissue with squamous cell carcinoma, at least 7000 image patches from images of lung tissue with adenocarcinoma and at least 4000 image patches from images of lung tissue with squamous cell carcinoma, at least 5000 image patches from images of lung tissue with adenocarcinoma, at least 3000 image patches from images of lung tissue with squamous cell carcinoma, and at least 1000 patches from images of lung tissue with a lung cancer type other than adenocarcinoma and squamous cell carcinoma, or at least 7000 image patches from images of lung tissue with adenocarcinoma, at least 4000 image patches from images of lung tissue with squamous cell carcinoma, and at least 1500 patches from images of lung tissue with a lung cancer type other than adenocarcinoma and squamous cell carcinoma.

[0088] At step 214, the method may comprise using the images obtained at step 212 to train a deep learning model to classify fluorescence lifetime images between a plurality of classes associated with different subtypes of lung cancer using the training fluorescence lifetime images and associated ground truth labels, the plurality of classes comprising a first class associated with adenocarcinoma and a second class associated with squamous cell carcinoma, using said training images. As described elsewhere, the deep learning model may be a deep artificial neural network, a convolutional neural network, or a CNN with residual connections. For example, the deep learning model may be a ResNet or RestNet derivative model, such as a multilevel CNN with residual connections, a DenseNet, an EfficientNet, or any other CNN or image feature extraction deep learning model. The training of a DL model for classification using images and ground truth labels is known in the art. The model may be trained using a predetermined proportion of the training images obtained at step 212, and tested using the remaining proportion of images. DL models trained as described herein are expected to be able to classify images in the plurality of classes with a precision of at least 85%, a recall of at least 85%, and a specificity of at least 65%.

[0089] Training the DL model may comprise evaluating the performance of the model using one or more test datasets. For example, training the DL model may comprise evaluating the performance of the model using a test dataset that is a subset of the training dataset, such as where the training dataset is dividing into a training set comprising 80% (or 85%, 90%) of the training images which is used for training the model, and a test set comprising the remaining e.g. 20% of the training images which is used to test the model. The training set and test set may be formed from the training set such that all images from the same sample are in the same set. The training and test sets may be formed such that all patches from the same image are in the same set. The training and test sets may be formed such that the proportions of images in the classes of the plurality of classes is similar between the test set and the training set. Performance may be evaluated using any metric known in the art such as e.g. precision, recall and specificity. Performance may be evaluated individually for a class of the plurality of classes, or collectively across all classes. In embodiments, the deep learning model has been trained to classify images in the plurality of classes with a precision of at least 85%, a recall of at least 85%, and a specificity of at least 65%, when evaluated over all classes collectively using a test set that forms part of a training dataset from which the training images used for training the model have been obtained, the test set and training set comprising non-overlapping sets of images from the training dataset.

[0090] At optional step 216, the results of any one or more of the preceding steps may be provided to a user, computing device or memory. This may include the trained model, such as e.g. for use in a method as described herein such as by reference to Figure 1 . Systems

[0091] Figure 3 shows an embodiment of a system for analysing FLIM images and / or for providing a tool for FLIM medical images, according to the present disclosure. The system comprises a computing device 1 , which comprises a processor 101 and computer readable memory 102. In the embodiment shown, the computing device 1 also comprises a user interface 103, which is illustrated as a screen but may include any other means of conveying information to a user such as e.g. through audible or visual signals. The computing device 1 is communicably connected, such as e.g. through a network, to data acquisition means 3 (also referred to as “FLIM imaging means”), such as fluorescence lifetime imaging microscope or endomicroscope or computing device associated therewith, and / or to one or more databases 2 storing FLIM data. The one or more databases 2 may further store one or more of: one or more deep learning algorithm, training data, parameters (such as e.g. parameters of a deep learning model used to diagnose lung cancer subtypes), image data acquisition parameters, clinical and / or sample related information, etc. The computing device may be a smartphone, tablet, personal computer or other computing device. The computing device is configured to implement a method for identifying lung cancer subtypes, analysing FLIM images, obtaining a tool for identifying lung cancer subtypes, and / or identifying a prognosis, treatment or selecting a subject for a clinical trial, as described herein. In alternative embodiments, the computing device 1 is configured to communicate with a remote computing device (not shown), which is itself configured to implement a method as described herein. In such cases, the remote computing device may also be configured to send the result of the method to the computing device. Communication between the computing device 1 and the remote computing device may be through a wired or wireless connection, and may occur over a local or public network 6 such as e.g. over the public internet. The data acquisition means 3 may be in wired connection with the computing device 1 , or may be able to communicate through a wireless connection, such as e.g. through WiFi and / or over the public internet, as illustrated. The connection between the computing device 1 and the data acquisition means 3 may be direct or indirect (such as e.g. through a remote computer).

[0092] The following is presented by way of example and is not to be construed as a limitation to the scope of the claims.

[0093] Example 1

[0094] Introduction

[0095] Autofluorescence lifetime is a unique characteristic of biological molecules, which is independent of their intensity but sensitive to the surrounding bio-environment. Fluorescence lifetime imaging microscopy (FLIM) can capture autofluorescence lifetime, and therefore, detect cancer by utilising the lifetime contrast between normal / cancerous tissue. Previous investigations have shown that label-free FLIM images can distinguish cancerous tissue from normal one on the lung using deep learning (DL) techniques (Wang et al., 2022). In the present examples, the inventors studied the capability of FLIM for NSCLC subtyping, and in particular for distinguishing adenocarcinoma (ADE) and squamous cell carcinoma (SCC) from other nonsmall cell lung cancer (NSCLC) subtypes. First, the inventors showed that differences in the mean lifetime over whole FLIM images are insufficient to differentiate between adenocarcinoma (ADC), squamous cell carcinoma (SCC), and other NSCLC subtypes. Then, they investigated whether deep learning could be used instead to distinguish between these three subtypes of lung cancer, demonstrating that the deep learning based analysis of FLIM images was able to distinguish between these three types of NSCLC.

[0096] Methods

[0097] The process of data collection and processing used in the present examples is illustrated on Figure 6.

[0098] Data collection. Paired cancer and non-cancerous lung biopsies were obtained from 79 (43 ADE, 25 SCC, and 11 other subtypes) NSCLC patients. The other subtypes included any lung cancer that could not be identified as either ADE or SCC, such as e.g. large cell carcinoma, large cell neuroendocrine carcinoma, and mixtures of those subtypes. These subtypes are much less common than ADE and SCC, and as such it was not possible to obtain enough data to train a model to classify them separately. Only the data from the cancer lung biopsies were used in the present examples. The data was collected using a commercial FLIM imaging system (Leica T:S SP8 Confocal Microscope). Through a A-to-A scan of the fixed unstained lung sample, the optimal excitation was set to 485 nm and the emission wavelengths are [500 nm, 640 nm] (i.e. recording all signal emitted across the 500 to 640nm band), with a 20x / 0.75NA objective. The A-to-A scan was performed once and all data was scanned using the identified optimal excitation and emission wavelengths (which provide maximal emission intensity). The optimal excitation and emission wavelengths depend on the fluorophores that are present, which is why data is not acquired at a predetermined excitation / emission wavelength but is instead acquired using settings identified through a A-to-A scan for the particular type of samples at hand (in this case, fixed unstained lung tissue samples). The image size was fixed at 512x512, resulting in a pixel size of 0.445 gm. Afterwards, lifetime images were reconstructed by an exponential curve fitting method provided by Leica LAX-X software. Intensity and lifetime images were exported separately. Examples of intensity (Figure 4A) and lifetime (Figure 4B) images are shown on Figures 4A and 4B.

[0099] Note that the intensity and lifetime images shown in Figure 4 are after the stitching of all individual tiles per each TMA (Tumour Microarray) core, which was achieved by using FIJI software.

[0100] Data processing. Images were processed using the proprietary software provided with the instrument for lifetime reconstruction. Contrast enhancement was applied to the intensity images by saturating the bottom 1% and the top 3% of all the pixel values on intensity images. No image post-processing was applied on the lifetime images. False-colour lifetime images using the corresponding intensity images as the alpha channel were generated as the input for the downstream analysis. This type of images (intensity weighted lifetime images) applies an intensity image as a soft weight to the corresponding lifetime image. This process is described in Wang et al. 2022 and McGinty et al. 2010. False colour lifetime images produce RBG images (or other colour images) by assigning each pixel in a FLIM image a colour according to its lifetime value (step “Colour mapping” on Fig. 6), with the intensity information modulating pixel saturation of the coloured lifetime image (i.e. acting as a transparency filter). An example of a resulting image is shown on Figure 4C. Mean lifetime analysis. For each image tile (512x512) FLIM image, the averaged lifetime was obtained using a conventional histogram-based method. In particular, a histogram of lifetime (i.e. counts of pixels as a function of lifetime associated with the pixel in ns was obtained) and the mode of the histogram was used as the average lifetime (see e.g. Wang et al. 2020, Fig. 1 ).

[0101] Deep learning. Tumour microarray cores from the above-mentioned 79 NSCLC patient biopsies were scanned as explained above. Since the deep learning model for classification was designed for images of size 224x224, the FLIM images were resampled to 224x224 by stitching together all tiles from the same lung tissue section then splitting the resulting images into 224 x 224 pixels patches. FLIM patches with over 50% background were disposed of. Background was identified as pixels with intensity 0 (where pixels with intensity 0 are obtained as a result of thresholding on intensity, where a fluorescence intensity threshold is selected such that all signal from outside the specimen being images is below the threshold). Any patch comprising over 50% of such pixels was discarded. The images were split between training and testing sets, as summarised in Table 1.

[0102] Table 1 : dataset for the training and testing of DL-based classification model.

[0103] A well-established model architecture was used, namely ResNet50 (He et al. 2016). The model was trained for 200 epochs with batch size 128. The initial learning rate was set to 0.1 and decayed by 10 for every 60 epochs. ADAM was used as the optimises Simple data augmentation was applied to all images used for the training, including a random rotation of 15 degrees and a random horizontal flipping (i.e. each image in the training dataset is rotated by 15 degrees and / or horizontally flipped to produce an augmented training dataset). The model was configured as a multiclass classifier, outputting a probability for each class for an input image (here a patch), the probabilities summing to 1. The class of a patch was assigned as that associated with the highest probability. The model was implemented in PyTorch (see pytorch.org / hub / pytorch_vision_resnet / ).

[0104] Model evaluation. Quantitative outcomes were derived using sensitivity (also named recall), specificity, and precision using the following equations: where TP is true positive, TN is true negative, FP is false positive, and FN is false negative. In particular, for each class, the precision, recall and specificity of the model were calculated. An aggregate over all classes was also calculated as the sum of the one-vs-all precision, recall and specificity, respectively, across the three categories. All metrics were calculated at the patch level, because some patient images showed a mixture of ADE and SCC. Alternative architectures. A ResNet model was used to generate the results below. ResNet is a type of deep artificial neural network (deep ANN, or deep learning DL model), that uses residual blocks, in which a residual connection (also referred to as skip connection) connects the input of the block to the output of the block. This is illustrated on Figure 7a. Alternative architectures based on the ResNet model have been suggested in Wang et al. 2022, including architectures including ResNetZ blocks and Res2Net blocks, illustrated on Figures 7b and 7c, respectively. Each of these architectures is expected to be usable in the methods described. Other architectures suitable for the present use include DenseNet, Inception, SqueezeNet, and ResNeXt.

[0105] Results

[0106] Figure 5 depicts the distribution of average lifetime of FLIM images of NSCLC tissues, separated by subtype, i.e. adenocarcinoma (ADC), squamous cell carcinoma (SCC), and other NSCLC subtypes. The data show that although the different subtypes have different mean lifetimes and standard deviations, (i.e. 3.32±0.2335 ns for ADC, 3.36±0.4323 ns for SCC, and 3.129±0.4887 ns for other subtypes), it is extremely challenging to differentiate individual cases using average lifetime. Indeed, there is no threshold that can be used to separate ADC, SCC and other subtypes with high specificity, and any threshold that can be identified for a specific dataset does not generalise to another dataset. Indeed, the particular values that are found depend highly on the particular data at hand and how it was processed, whereas the patterns of lifetime that can be identified by deep learning models (as demonstrated further below) do not.

[0107] These data demonstrate that the differences in the averaged / mean lifetime are insufficient to differentiate NSCLC subtypes. In other words, the variability in lifetime across patients is very high, even when looking specifically at cancer regions (let alone when looking at images that, for the purpose of automated image analysis, necessarily include cancer microenvironment and possibly even surrounding tissue). As demonstrated herein, this complexity makes simple approaches based on averaged lifetime inaccurate to the extent of being not clinically usable.

[0108] The inventors therefore tested a DL-based classification model, ResNet, to differentiate between NSCLC subtypes by analysis of intensity-weighted lifetime images from FLIM. The results of this analysis are shown in Table 2. Intensity-weighted lifetime images were shown in Wang et al. 2020 (using intensity stacked images) and Wang et al. 2022 (using both intensity stacked and intensity weighted images where the intensity is used as the alpha channel) to result in better performance for cancer vs normal classification, as they address the problem of lifetime independent from intensity, increasing the classification performance of classic CNNs. Amongst these, Intensity-weighted lifetime images where the intensity is used as an alpha channel were found to be associated with particularly good model performance. However, intensity stacked images (i.e. 2 channel images used as input to the deep learning model, one channel corresponding to intensity and the other corresponding to lifetime) were also shown to be usable in Wang et al. 2022, and are expected to be usable in the present setting as well (albeit with lower performance). Further, composite images where lifetime information is multiplied pixelwise by the intensity in a corresponding intensity image may also be used.

[0109] Table 1 : classification of NSCLC subtypes using precision, recall / sensitivity, and specificity.

[0110] Using the independent subset of the images for testing, the aggregated outcomes, that is the sum for all subtypes, are 89.2%, 87.9%, and 66.41%, for precision, recall, and specificity, respectively, which are listed in Table 1. The precision values for one-vs-all classification are 32.69% for ADE, 33.77% for SCC, and 22.72% for other subtypes. The recall values are 64.03% for ADE, 22.46% for SCC, and 1.37% for others. The one-vs-all specificities are 21.52%, 72.71%, and 98.49% for ADE, SCC, and others, respectively. These number are in lines with the expected proportions of samples as ADC and SCC represented each 2 / 5 of the test set at the patient level (noting that the patient level classification is not necessarily 100% aligned with the tile level classification since some patients may include both ADC and SCC regions). The recall for the “others” category was lower than for the other categories because the category is very small ( a single patient in the test set, 10 in the training set) and as such there are simply very few true positive slides for training and testing. Nevertheless, the high aggregated accuracies obtained indicate that deep learning can be used effectively for NSCLC subtyping.

[0111] Note that NSCLC represents over 80% of cancers, and amongst NSCLC patients, ADC and SCC represent over 80%. Therefore, the ability to achieve such high precision and recall in identifying these lung cancer subtypes from others represents a clinically useful diagnostic tool for lung cancer in general.

[0112] These results demonstrate that label-free FLIM images can be exploited for automatic lung cancer subtyping using DL techniques.

[0113] Example 2

[0114] The inventors employed deep learning techniques to classify different types of cancer from tissue images obtained from Tissue Microarrays (TMAs). The classification focused on distinguishing between adenocarcinoma, squamous carcinoma, other non-small cell lung cancer (NSCLC) subtypes, and normal tissue. Specifically, the inventors used both intensity-weighted lifetime images (ILIs) and intensity images (ITIs) to train different deep neural networks (DNNs) and compared the results between ILIs- and ITIs- based DNNs. ITIs are single-channel grayscale images that contain information solely on photon intensity (auto-fluorescence) and tissue morphology, whereas ILIs are RGB-Alpha, four-channel images, where RGB and alpha channels contain lifetime and intensity information, respectively.

[0115] 1. Methods

[0116] The numbers of cores used for training, validation, and testing are provided in Table 3. The inventors first focused on binary classification using individual models on three types of cancer subtyping: cancer versus non-cancer, adenocarcinoma (ADE) versus squamous carcinoma (SqCC) and other subtypes (OS), and SqCC versus OS. Then the inventors used a holistic model for multiple classification for the four types of cancer subtypes.

[0117] Table 3: The numbers of cores and patches of four classes.

[0118] FLIM images acquisition and processing was as described in Example 1, except as specified here. The data were collected using two FLIM imaging system: the Leica SP8 and Stellaris Falcon FLIM systems, subset of the data corresponds to the data in Example 1 and was collected with the Leica SP8, with the excitation and emission wavelengths mentioned in Example 1, i.e. the excitation wavelength was set to 485 nm and the emission wavelengths are [500 nm, 640 nm] , The rest of the data were from the Stellaris Falcon, and was acquired using excitation 445nm and emission [460, 640] nm, all with a 20x / 0.75NA objective. The pixel sizes are various from 0.445mm to 0.2mm. However, all images were rescaled to 0.3mm.

[0119] Six metrics were employed to evaluate the accuracy and robustness of the deep learning models for binary classification: accuracy, precision, specificity, recall (sensitivity), Matthews Correlation Coefficient, and ROC & AUC Score. The Matthews Correlation Coefficient measures the correlation between predicted and actual binary outcomes, ranging from [-1, 1], +1 : Perfect classification; 0: Random classification; -1: Complete disagreement between predictions and actual labels. ROC indicates the proportion of actual positives correctly identified by the mode; AUC quantifies the overall ability of the model to discriminate between positive and negative classes.

[0120] For multiclass classification the ROC and AUC scores were used to evaluate performance.

[0121] 2. Results

[0122] 2. 1 Binary classification

[0123] First, a set of different deep neural networks were evaluated for their performance in binary classification using intensity images only (specifically, DenseNet, ResNet and EfficientNet architectures). A total of 51,355 patches from 95 cores were tested following training. As shown in Table 4, six metrics were used to quantitatively evaluate the performance of the trained model (using intensity images), and the DenseNet model achieved the best performance (although ResNet and EfficientNet also performed well) for all tasks apart from binary classification of Adenocarcinoma vs Squamous Carcinoma where ResNet50 performed slightly better.

[0124] Table 4. Classification performance from different DNNs trained using lifetime images, evaluated by six metrics.1Matthews Correlation Coefficient measures the correlation between predicted and actual binary outcomes, ranging from [-1, 1], +1: Perfect classification; 0: Random classification; -1: Complete disagreement between predictions and actual labels Having established that the DenseNet architecture performs slightly better, the inventors used this architecture to evaluate the effect of training on intensity weighted lifet images (ILI) vs intensity images only (ITI). As shown on Fig. 8, ILIs-trained models demonstrates overall higher performance compared to ITIs- trained models (compared Fig. 8A which shows ILI results and Fig. 8B which shows ITI results). Specifically, IT Is-trained DenseNet already achieves high accuracy, and this is even further enhanced by the inclusion of the lifetime information, particularly when it comes to subtype classification. Indeed, in the case of cancer versus non-cancer classification, both types of models produce nearly perfect ROC and AUC scores. However, for subtype classification models, the inclusion of the lifetime information meaningfully improves the classification.

[0125] 2.2. Performance evaluation (multiple classification) In this section, the inventors focused on training a single model that attempts to classify all four types simultaneously, assessing its ability to handle multi-class classification in one unified approach (i.e. using multi-label classification instead of binary labels). The inventors evaluated the performance of the three DNN models mentioned above, i.e., DenseNet. ResNet, and EfficientNet, trained to classify intensity weighted lifetime images between: (a) normal tissue, (b) adenocarcinoma tissue, (c) squamous cell carcinoma tissue, and (d) other cancerous tissue. The ROC and AUC scores are presented in Fig. 9 ((a) DenseNet, (b) ResNet, (c) EfficientNet) and Table 5. This shows that all 3 models performed well, with DenseNet yielding the highest AUC scores, and ResNet also performing extremely well.

[0126] Table 5. Performance Evaluation of classical deep learning architectures for multiple cancer type classification, using different metrics.

[0127] The performance of multi-classification using ILI-DenseNet vs ITIs-DenseNet was also investigated. As shown in Fig. 10, ILI-DenseNet (Fig. 10a) achieves overall higher AUC scores compared to ITIs-DenseNet (Fig. 10b). Similarly, the confusion matrix of ILI-DenseNet (Fig. 10c) shows fewer misclassified images compared to that of ITIs-DenseNet (Fig. 10c). This again shows that inclusion of the lifetime information significantly improves the performance of NSCLC subtype classification.

[0128] 2.3. Core-based evaluation

[0129] In addition to patches-based evaluation (results in sections 2.1 and 2.2), the inventors also analyzed the results of core-based classification, using both ILI-DenseNet (DenseNet trained with intensity weighted lifetime images) and ITIs-DenseNet (DenseNet trained with intensity images only), using the average probability of each class provided the multiclass classification models over all patches of a core to assign a class to the core. Results are shown in Fig. 11. These show that ILI-DenseNet has higher performance than ITIs-DenseNet, albeit with slightly higher standard deviations for each subtype (the overall distribution remaining higher than with intensity only images, i.e. classification of all cores benefit from inclusion of lifetime information, although some cores benefit even more than others), compared ITIs-DenseNet. This indicates that the intensity weighted lifetime images enable more confident NSCLC subtype identification.

[0130] The inventors therefore demonstrate that a variety of existing DNN architectures can be used to effectively classify lung cancer subtypes using either lifetime or intensity images, and that the inclusion of lifetime information further enhances this classification.

[0131] References

[0132] A number of publications are cited above in order to more fully describe and disclose the invention and the state of the art to which the invention pertains. Full citations for these references are provided below. The entirety of each of these references is incorporated herein.

[0133] Fernandes S, Williams G, Williams E, et alS99 Fluorescence-lifetime imaging: a novel diagnostic tool for suspected lung cancer. Thorax 2021 ;76:A63-A64. Fernandes, ES., 2022, Photonic signatures of human lung cancer, PhD Thesis, Edinburgh Medical School, The University of Edinburgh.

[0134] K. He, X. Zhang, S. Ren and J. Sun, “Deep Residual Learning for Image Recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016, pp. 770- 778, doi: 10.1109 / CVPR.2016.90

[0135] Suhling K, Hirvonen LM, Levitt JA, Chung PH, Tregidgo C, Le Marois A, Rusakov DA, Zheng K, Ameer- Beg S, Poland S, Coelho S, Henderson R, Krstajic N (2015) Fluorescence lifetime imaging: Basic concepts and some recent developments. Med Photonics 27:3-40.

[0136] Wang, Q., Fernandes, S., Williams, G.O.S. et al. Deep learning-assisted co-registration of full-spectral autofluorescence lifetime microscopic images with H&E-stained histology images. Commun Biol 5, 1119 (2022)

[0137] Wang, Q., Hopgood, J.R., Fernandes, S. et al. A layer-level multi-scale architecture for lung cancer classification with fluorescence lifetime imaging endomicroscopy. Neural Comput & Applic 34, 18881— 18894 (2022)

[0138] Wang Q, Hopgood JR, Finlayson N, Williams GO, Fernandes S, Williams E, Akram A, Dhaliwal K, Vallejo M (2020) Deep learning in ex-vivo lung cancer discrimination using fluorescence lifetime endomicroscopic images. In: 202042ndannual international conference of the IEEE engineering in medicine & biology society (EMBC), pp 1891-1894. IEEE

[0139] Wei Zheng, Zhiwei Huang, Shusen Xie, T eck-Chee Chia, Zukang Lu, and Jin Kai Chen “Lifetimes of normal and tumorous human lung tissues by time-resolved fluorescence spectra”, Proc. SPIE 2887, Lasers in Medicine and Dentistry: Diagnostics and Treatment, (23 September 1996).

[0140] Dosovitskiy, A., Beyer, L, Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Geliy, S., Uszkoreit, J., Houlsby, N.: An image is worth 16x16 words: Transformers for image recognition at scale. In: 9thICLR. Pp. 1-21. ICLR (2021 )

[0141] Perez-Moreno P, Brambilla E, Thomas R, Soria JC. Squamous cell carcinoma of the lung: molecular subtypes and therapeutic opportunities. Clin Cancer Res. 2012 May 01;18(9):2443-51.

[0142] Dietel M, Bubendorf L, Dingemans AM, Dooms C, Elmberger G, Garcia RC, Kerr KM, Lim E, Lopez-Rios F, Thunnissen E, Van Schil PE, von Laffert M. Diagnostic procedures for non-small-cell lung cancer (NSCLC): recommendations of the European Expert Group. Thorax. 2016 Feb;71(2):177-84.

[0143] Relli V, Trerotola M, Guerra E, Alberti S. Abandoning the notion of non-small cell lung cancer. Trends Mol Med (2019) 25:585-94. Doi: 10.1016 / j.molmed.2019.04.012

[0144] Akikazu Kawase, Junji Yoshida, Genichiro Ishii, Masayuki Nakao, Keiju Aokage, Tomoyuki Hishida, Mitsuyo Nishimura, Kanji Nagai, Differences Between Squamous Cell Carcinoma and Adenocarcinoma of the Lung: Are Adenocarcinoma and Squamous Cell Carcinoma Prognostically Equal?, Japanese Journal of Clinical Oncology, Volume 42, Issue 3, March 2012, Pages 189-195

[0145] Shang-Hua Gao, Ming-Ming Cheng, Kai Zhao, Xin-Yu Zhang, Ming-Hsuan Yang, Philip Torr. Res2Net: A New Multi-scale Backbone Architecture. arXiv: 1904.01169v3 27 Jan 2021.

[0146] Saining Xie, Ross Girshick, Piotr Dollar, Zhuowen Tu, Kaiming He. Aggregated Residual Transformations for Deep Neural Networks. arXiv:1611.05431v2. 11 April 2017. James McGinty, Neil P. Galletly, Chris Dunsby, Ian Munro, Daniel S. Elson, Jose Requejo-lsidro, Patrizia Cohen, Raida Ahmad, Amanda Forsyth, Andrew V. Thillainayagam, Mark A. A. Neil, Paul M. W. French, and Gordon W Stamp, "Wide-field fluorescence lifetime imaging of cancer," Biomed. Opt. Express 1 , 627- 640 (2010).

[0147] Gao Huang, Zhuang Liu, Laurens van der Maaten, Kilian Q. Weinberger. Densely Connected Convolutional Networks. arXiv:1608.06993v5. 28 Jan 2018

[0148] Mingxing Tan, Quoc V. Le. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. arXiv: 1905.11946v5. 11 Sep 2020.

[0149] Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, Baining Guo. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. arXiv:2103.14030v2. 17 Aug 2021 Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, Saining Xie. A ConvNet for the 2020s. arXiv:2201.03545v2. 2 Mar 2022

Claims

Claims:1 . A computer implemented method of identifying a lung cancer subtype in a patient, the method comprising: obtaining one or more previously acquired autofluorescence lifetime images of lung tissue of the patient; and classifying the one or more images between a plurality of classes associated with different subtypes of lung cancer, the plurality of classes comprising a first class associated with adenocarcinoma and a second class associated with squamous cell carcinoma, using a deep learning model that has been trained using training fluorescence lifetime images of lung tissue from a plurality of patients, the training fluorescence lifetime images comprising a first set of images of lung tissue identified as adenocarcinoma and a second set of images of lung tissue identified as squamous cell carcinoma.

2. The method of claim 1 , wherein for each image provided as input to the deep learning model, the deep learning model provides as output one or more probabilities that the image provided as input shows lung tissue from the one or more respective lung cancer subtypes associated with the plurality of classes, optionally wherein the deep learning model provides as output a probability for each of the plurality of classes, and the patient is identified as having a lung cancer subtype that is associated with the class that has the highest probability amongst the probabilities for the plurality of classes.

3. The method of any preceding claim, wherein the plurality of classes comprises or consists of: a first class associated with adenocarcinoma, a second class associated with squamous cell carcinoma, a third class associated with lung cancer subtypes that are not adenocarcinoma or squamous cell carcinoma, and optionally a fourth class associated with normal tissue.

4. The method of any preceding claim, wherein the lung cancer is non-small cell lung cancer and / or wherein the patient is a patient that has been diagnosed as having lung cancer.

5. The method of any preceding claim, wherein the autofluorescence lifetime images are composites of autofluorescence intensity and lifetime images obtained from the same fluorescence lifetime image, optionally wherein a composite image is an intensity-weighted lifetime image.

6. The method of claim 5, wherein the method comprises obtaining one or more autofluorescence intensity images and corresponding fluorescence lifetime images, and obtaining, for each pair of autofluorescence intensity and lifetime images, an intensity weighted lifetime image.

7. The method of claim 5 or claim 6, wherein an intensity weighted lifetime image is a false-colour lifetime image with colour depending on lifetime and the corresponding intensity image as the alpha channel, an image comprising lifetime as a first channel and intensity as a second channel, or an image obtained by multiplying pixel values in an autofluorescence intensity image by the corresponding pixel values in a corresponding fluorescence lifetime image.

8. The method of any preceding claim, wherein the one or more images are processed or have been processed using one or more of: thresholding of the intensity information, denoising of one or both of the lifetime and fluorescence intensity information, normalising of one or both of the lifetime and fluorescence intensity information, contrast enhancing of the intensity information and contrast enhancing of the lifetime information, optionally wherein the one or more images are processed or have been processed using at least contrast enhancing of the intensity information.

9. The method of any preceding claim, wherein the lung tissue is an ex vivo tumour tissue sample that has been previously obtained from the patient, optionally wherein the lung tissue is from a fixed tissue sample, or a tumour microarray, and / or wherein the fluorescence lifetime images have been acquired using a fluorescence lifetime imaging microscope.

10. The method of any of claims 1 to 8, wherein the fluorescence lifetime images are images of in vivo lung tissue that have been previously obtained from the patient, and / or wherein the fluorescence lifetime images have been acquired using a fluorescence lifetime endomicroscope,11. The method of any preceding claim, wherein the fluorescence lifetime images have been acquired using a fiber-based fluorescence lifetime imaging system.

12. The method of any preceding claim, wherein obtaining one or more previously acquired fluorescence lifetime images of lung tissue of the patient comprises receiving a previously acquired fluorescence lifetime image of lung tissue of the patient and obtaining a plurality of patches from the image, wherein a patch is a subset of an image of a predetermined size, and wherein the one or more images provided as input to the deep learning model are individual patches.

13. The method of claim 12, wherein the method further comprises excluding any patch that includes more than a predetermined threshold proportion of background pixels, optionally wherein the predetermined threshold is 50%.

14. The method of any preceding claim, wherein the training fluorescence lifetime images comprise a plurality of image of lung tissue and images obtained from the images by image augmentation, optionally wherein image augmentation comprises creating a flipped version of one or more of the images and / or creating a randomly rotated version of one or more of the images.

15. The method of any preceding claim, wherein the deep learning model has been trained using training fluorescence lifetime images of lung tissue from a plurality of patients, the training fluorescence lifetime images comprising a first set of images of lung tissue identified as adenocarcinoma, a second set of images of lung tissue identified as squamous cell carcinoma, and a third set of images of lung tissue identified as lung cancer of a type other than adenocarcinoma or squamous cell carcinoma, and optionally a fourth set of images of lung tissue identified as normal tissue and / or wherein the deep learning model is a multiclass classifier configured to classify images between at least 3 classes comprising a first class associated with adenocarcinoma and a second class associated with squamous cell carcinoma.

16. The method of any preceding claim, wherein the deep learning model has been trained using training fluorescence lifetime images of lung tissue from a plurality of patients, the training fluorescence lifetime images comprising images of lung tissue from: a. at least 50 patients, b. at least 70 patients, c. at least 20 patients with lung adenocarcinoma and at least 20 patients with squamous cell carcinoma, d. at least 30 patients with lung adenocarcinoma and at least 20 patients with squamous cell carcinoma, e. at least 40 patients with lung adenocarcinoma and at least 20 patients with squamous cell carcinoma, f. at least 20 patients with lung adenocarcinoma, at least 20 patients with squamous cell carcinoma and at least 5 patients with a type of lung cancer other than adenocarcinoma or squamous cell carcinoma, g. at least 30 patients with lung adenocarcinoma, at least 20 patients with squamous cell carcinoma, and at least 10 patients with a type of lung cancer other than adenocarcinoma or squamous cell carcinoma, or h. at least 40 patients with lung adenocarcinoma, at least 20 patients with squamous cell carcinoma, and at least 10 patients with a type of lung cancer other than adenocarcinoma or squamous cell carcinoma.

17. The method of any preceding claim, wherein the deep learning model has been trained using training fluorescence lifetime images of lung tissue from a plurality of patients, the training fluorescence lifetime images comprising: a. at least 5000 image patches from images of lung tissue with adenocarcinoma and at least 3000 image patches from images of lung tissue with squamous cell carcinoma, b. at least 7000 image patches from images of lung tissue with adenocarcinoma and at least 4000 image patches from images of lung tissue with squamous cell carcinoma, c. at least 5000 image patches from images of lung tissue with adenocarcinoma, at least 3000 image patches from images of lung tissue with squamous cell carcinoma, and at least 1000 patches from images of lung tissue with a lung cancer type other than adenocarcinoma and squamous cell carcinoma, or d. at least 7000 image patches from images of lung tissue with adenocarcinoma, at least 4000 image patches from images of lung tissue with squamous cell carcinoma, and at least 1500 patches from images of lung tissue with a lung cancer type other than adenocarcinoma and squamous cell carcinoma.

18. The method of any preceding claim, wherein the autofluorescence lifetime images are images that have been acquired using: (i) an excitation wavelength selected between 450nm and 500nm, between 400nm and 500nm, between 420nm and 500nm, between 440nm and 500nm, between 460nm and 500n, between 470nm and 500nm, between 480nm and 490n, between 430nm and 450nm, or about 485 nm, or about 445nm; and / or (ii) a spectral band of emission be selected to havea lower boundary between 450 and 510nm, between 460 and 500nm, between 490nm and 510nm, between 495nm and 505nm, at 490nm, at 495nm, at 500nm, at 505nm, or at 51 Onm, or between 440nm and 480nm, between 450nm and 460nm, or about 460nm; and / or (iii) a spectral band of emission selected to have an upper boundary between 550nm and 650nm, between 560nm and 650nm, between 570nm and 650nm, between 580nm and 650nm, between 590nm and 650nm, between 600nm and 650nm, between 61 Onm and 650nm, between 620nm and 650nm, at 600nm, at 61 Onm, at 620nm, at 630nm or at 640nm; and / or (iv) an excitation wavelength at 485nm and a spectral emission band with a lower boundary at 500nm and an upper boundary at 640nm; and / or (v) excitation wavelength at 445nm and a spectral emission band with a lower boundary at 460nm and an upper boundary at 640nm.

19. The method of any preceding claim, wherein the deep learning model is a deep artificial neural network, a convolutional neural network, a transformer-based model, a CNN with residual connections, a CNN with dense connections, optionally wherein the deep learning model is a ResNet or RestNet derivative model, optionally a multilevel CNN with residual connections, or wherein the deep learning model is a DenseNet model.

20. A method of selecting a subject that has been diagnosed as having lung cancer for participation in a clinical trial and / or for further diagnostic testing, the method comprising: identifying a lung cancer subtype in the subject using the method of any preceding claim; and selecting a subject identified as having a predetermined lung cancer subtype for participating in the clinical trial and / or selecting a subject identified as having a predetermined lung cancer subtype for further diagnostic testing, optionally wherein the further diagnostic testing is selected from mutation analysis and histopathology using stained tissue slices.21 . A method of providing a prognosis for a subject that has been diagnosed as having lung cancer, the method comprising:Identifying a lung cancer subtype in the subject using the method of any preceding claim; and Determining a prognosis for the subject, wherein a subject identified as having adenocarcinoma has better prognosis than a subject identified as having squamous cell carcinoma.

22. A method of identifying a treatment for a subject that has been diagnosed as having lung cancer, the method comprising:Identifying a lung cancer subtype in the subject using the method of any preceding claim; and Selecting a treatment for the subject based on the identified lung cancer subtype and optionally one or more predetermined clinical characteristics.

23. The method of claim 22, wherein selecting a treatment for the subject comprises determining a tumour grade, wherein the determining uses a tumour grading scheme that is dependent on the identified lung cancer subtype, and selecting a treatment associated with the determined tumour grade.

24. A system comprising: one or more processors and computer readable memory storing instructions that cause the processor to perform the method of any of claims 1 to 23, optionally wherein the system further comprises data acquisition means configured to obtain fluorescence lifetime imaging data from lung tissue.

25. A non-transitory computer readable storage medium containing machine executable instructions which, when executed on a processor, cause the processor to perform the method of any of claims 1 to 23.