Detection of anomalies in sample image
By detecting cells that cannot be identified with predetermined confidence or uncertainty scores, the problem of difficult detection of rare morphological abnormal cells in the prior art is solved, and efficient identification without training on the appearance of rare diseases is achieved.
Patent Information
- Application Number
- CN202380071707.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-19
- Filing Date
- 2023-10-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to detect rare morphological abnormal cells with high accuracy in deep neural networks that are not trained in the appearance of rare diseases.
The presence of a population of morphological abnormal cells is detected by detecting cells in the sample image that cannot be identified by a predetermined confidence score or cannot be identified by a lower than the predetermined uncertainty score. This method does not require training of the analytical model on extradomain data, nor does it require modification of existing models.
The detection of the presence of morphological abnormal cell populations without the appearance training of rare disease is achieved, and the recognition ability of rare disease cells is improved without modifying existing models.
Smart Images

Figure BDA0005348631060000041 
Figure HDA0005348631070000011 
Figure HDA0005348631070000021
Abstract
Description
Technical Field
[0001] The present invention relates to a computer-implemented method and related system for detecting the presence of morphologically abnormal cells in a sample image. Background Art
[0002] In order to train a deep neural network to recognize entities within an image and accurately identify them, training data (consisting of cell images and type labels) is required to establish ground truth.
[0003] In the case of hematology, there is a specific goal to distinguish healthy white blood cells (normal cells) from unhealthy white blood cells (abnormal cells, which can be malignant or benign).
[0004] While it is possible to collect enough data from common abnormalities (common leukemias, infections, etc.), it is not possible to collect enough data from rare diseases. Rare conditions may result in significantly different cell appearances, so a neural network that has not been trained on these rare appearances cannot classify cells with high accuracy.
[0005] The present invention solves these problems. Summary of the invention
[0006] At a high level, the present invention provides a computer-implemented method for detecting the presence of a population of cells with morphological abnormalities by detecting a threshold number of cells in a sample image that cannot be identified with a predetermined confidence score or cannot be identified with less than a predetermined uncertainty score. An increase in a population of cells that cannot be easily identified in a sample is an indicator of the presence of cells with morphological abnormalities. The computer-implemented method provided by the present invention is advantageous because it means that the presence of a population of morphologically abnormal cells can be detected without the need to train an analytical model (e.g., a machine learning model) on out-of-domain data. The computer-implemented method of the present invention also does not require any modifications to existing "vanilla" models.
[0007] More specifically, a first aspect of the present invention provides a computer-implemented method for detecting the presence of morphologically abnormal cells in a sample image, the computer-implemented method comprising: receiving electronic image data representing a sample image, the sample image depicting a plurality of cells; applying an analysis model to each of a plurality of subsets of the image data, each subset corresponding to a respective portion of the sample image depicting a single cell, the analysis model being configured to output, for each subset of the image data: a value of a property of a parameterized cell; and a confidence score or an uncertainty score associated with the value, thereby generating output data comprising a plurality of confidence scores or a plurality of uncertainty scores; and determining, based on the output data, whether one or more morphologically abnormal cells are likely to be present in the sample image.
[0008] Here, "may exist" may be understood to mean "exist". In some cases, the computer-implemented method may include determining whether a predetermined abnormality criterion is met, rather than actively determining the possibility of the presence of morphologically abnormal cells. When the predetermined abnormality criterion is met, this may indicate that the presence of morphologically abnormal cells is possible, or that morphologically abnormal cells exist. After determining whether there are morphologically abnormal cells or whether the predetermined abnormality criterion is met, the computer-implemented method may further include generating an output indicating the determined result. Generating the output may include generating instructions that, when executed by a display component of a computing device, cause the display component to display a visual indication of the determined result. Alternatively or additionally, the computer-implemented method may include transmitting the output to a database accessible by a clinical computer system.
[0009] In the event that it is determined that the abnormality criteria are met or the presence of morphologically abnormal cells is possible, the output preferably includes a marker. The marker may be added to the electronic image data representing the sample image. The marker is not necessarily a specific indication that morphologically abnormal cells are present in the sample image, but it marks the sample image to the clinician for further attention.
[0010] However, in a clinical setting, for a variety of reasons, only the results need to be flagged to a clinician rather than an automated diagnosis.The computer-implemented method of the present invention may include generating a flag in response to determining that one or more morphologically abnormal cells may be present.
[0011] In the context of this application, the term "value" may refer not only to numerical values but also to non-numerical values, such as classification output (the details will be described in detail below), which may be in the form of a numerical value indicating a specific classification, or alternatively in the form of text output.
[0012] Herein, "morphologically abnormal cells" are cells whose physical structure is different from that of normal cells, so that their appearance in the sample image is different from that of normal or "healthy" cells. Many diseases or other conditions can be detected by the presence of such morphological abnormalities. Therefore, the present invention can be used to detect or help detect those conditions that cause morphological abnormalities. However, it should be understood that the present invention is unlikely to be applied to detect conditions that cause abnormalities other than morphological abnormalities in cells. The detection of such conditions is not within the scope of this patent application.
[0013] In the context of the present invention, a "sample image" is an image depicting a human or animal tissue. The image can be obtained, for example, from a microscope or other imaging device. The sample image is described as depicting a plurality of cells. Therefore, the sample image is preferably at a magnification that enables the cells to be individually resolved by, for example, an imaging processing algorithm. The sample image can be an image of a slide of a sample of a body fluid or tissue obtained from a human or animal subject. The body fluid can be blood, and therefore, the abnormality can be a hematological abnormality. Blood is not the only body fluid that the computer-implemented method of the present invention can be used to achieve clinically significant results. For example, a body fluid can include a sample of a tissue / organ of a subject and / or a sample of a product produced by the tissue / organ of the subject. The product produced by the tissue / organ of the subject can be, for example, a secretory product (e.g., glandular secretion, milk, colostrum, tears, saliva, sweat, cerumen, mucus), sputum, semen, vaginal fluid / cervical fluid, blood (plasma, serum), cerebrospinal fluid (CSF), excretion products, feces or urine, skin or hair.
[0014] The computer-implemented method of the first aspect of the present invention includes the application of an analysis model. Herein, the term "analysis model" may refer to a mathematical model configured to determine at least one target variable for at least one state variable. The term "target variable" may refer to a clinical value to be predicted. The target variable value to be predicted may depend on the disease or condition to be predicted for its presence or state. The target variable may be a numerical variable or a categorical variable. For example, the target variable may be a categorical variable, and may be "positive" in the presence of a disease, or may be "negative" in the absence of a disease. In other cases, the target variable may refer to the classification of a cell type (the specific content will be described in detail below). The target variable may be a numerical variable, such as at least one value and / or a scale value.
[0015] The analysis model may be a regression model or a classification model. In the context of the present application, the term "regression model" may be used to refer to an analysis model whose output is a numerical value within a certain range. For example, in this example, the output of such a regression model may be a numerical value. In other cases, the analysis model may be a segmentation model whose output is an indication of a segment of a sample image associated with, for example, a single pixel.
[0016] In the context of this application, the term "classification model" may be used to refer to an analysis model whose output is a respective classification or score indicating the type of cell depicted in each subset of the image data. For completeness, we note that when the analysis model is a classification model, the "value of the parameterized property of the cell" may be a numeric or textual output indicating the type of the cell (i.e., the "property" of the cell is the cell type).
[0017] Specifically, the analysis model can be a machine learning model trained to output corresponding results indicating the type of cells depicted in a given subset of the image data. The machine learning model can be a regression model or a classification model, as previously defined. The machine learning model is preferably trained using supervised learning. The machine learning model can include a neural network such as a convolutional neural network.
[0018] In a preferred embodiment, the analysis model is a classification model based on a neural network (such as a convolutional neural network and / or a deep neural network). In these cases, the classification model is preferably trained to classify cells into one of a plurality of cell types based on data representing an electronic image of the cell. For example, the classification model may be configured to classify cells into one of a plurality of types of white blood cells (e.g., at least as follows: neutrophils; lymphocytes; monocytes; eosinophils; and basophils). Preferably, the classification model has been trained using training data comprising a plurality of images of normal or healthy cells, each image being associated with a label indicating the type of the cell. This indicates that the classification model (or more generally, the analysis model) does not need to be trained on morphologically abnormal cells.
[0019] The analytical model may include multiple analytical sub-models, each of which may take the form of a machine learning model outlined in the previous two paragraphs.
[0020] The concept of confidence score and uncertainty score is the core of the present invention. These terms are well known in statistics and in the field of machine learning. The present invention is applicable to all situations where the analysis model can establish the distribution of normal samples (i.e., healthy or normal cells). This is because the understanding of the distribution can calculate the uncertainty score or confidence score. In fact, many analysis models can calculate uncertainty or confidence scores in addition to the usual output values.
[0021] Different analytical models (such as machine learning models) report different uncertainties. In the computer-implemented method according to the present invention, the prediction uncertainty caused by the fact that the data point is too far away from the training data to be reliably classified is called epistemic uncertainty. 1 .
[0022] In general, uncertainty can be quantified as the amount of disagreement within the distribution of predictions reported by the model. Variance or standard deviation are common ad hoc choices for this purpose. A theoretically sound approach is to define uncertainty as the mutual information between the model parameters and the output 2 .
[0023] The predictive distribution can be approximated by using an ensemble of models 3Using model ensembles can generally refer to the process of combining models that are trained to perform the same task (in this case, models with the same class). In general, using ensembles and combining the outputs of different models is often done to improve accuracy, as uncorrelated errors made by different models cancel out when the outputs are appropriately combined. However, in this case, ensembles can be used to obtain a measure of disagreement between the various sub-models in the ensemble, see below. If the ensemble predictions of the individual models in the ensemble disagree to a large extent, this indicates that the data point is farther from the training data and should therefore be assigned high epistemic uncertainty. Conversely, if the models in the ensemble agree, this indicates that the data point is close to the training set and therefore has low epistemic uncertainty.
[0024] More specifically, in this context, the use of an ensemble may refer to a process in which several analysis submodels are each used to obtain a corresponding output. The outputs may then be combined, and the variance or standard deviation (or other information-theoretic measure) of the outputs may be used as an uncertainty score. More specifically, applying the analysis model to each of a plurality of subsets of image data may include applying a plurality of analysis submodels (or an ensemble of analysis submodels) to each subset of image data. Each of the plurality of analysis submodels may then be configured to output a corresponding plurality of values (for each subset of image data), the corresponding plurality of values each parameterizing a property of the cell. Applying the analysis model to each of a plurality of subsets of image data may further include determining a statistical property indicating a confidence or uncertainty in the plurality of values for each of the output plurality of values. The statistical property may be a variance or a standard deviation, or other information-theoretic parameter. Each of the analysis submodels may be of the same type and trained on the same training data. In order to ensure that the ensembles of the analysis submodels are not identical (in which case it would be impossible to obtain a meaningful uncertainty score), it is preferred that each of the plurality of analysis submodels differ at least by an initial random seed that is randomly used to initialize the model before training.
[0025] In some cases, the various analysis sub-models may not be of the same type, for example, some analysis sub-models may be convolutional neural networks and some may be transformers, etc. The different types of sub-models may be trained on the same or different training datasets. However, the different types of sub-models are all trained on the same task (i.e., classified into the same category).
[0026] The predictive distribution can also be approximated in other ways, such as by using dropout. In the context of machine learning, the term "dropout" is used to refer to the process of randomly ignoring (i.e., dropping) some neurons. Dropout, as well as ensembles, can be used as approximations of Bayesian neural networks. 4 .
[0027]
[0028] If the uncertainty for each data point is not accurate enough, it can be determined whether the model has enough confidence to make a decision based on a semantically meaningful collection of data points, such as images of cells from the same sample. In the case of hematology, if the model reports that the uncertainty is elevated for enough cells on the slide, it will decide whether the sample should be flagged.
[0029] We now consider how to use uncertainty / confidence scores to detect the presence of morphologically abnormal cells. The core concept of the present invention is the ability to use uncertainty or confidence scores to identify morphologically abnormal cell populations, even in the absence of training data covering morphologically abnormal cells. The output data includes multiple confidence scores or multiple uncertainty scores, which can be calculated using the process explained in the previous paragraphs. It has been observed that in sample images depicting samples containing morphologically abnormal cells, the number or proportion of reduced confidence scores or increased uncertainty scores increases. This is because the analysis model has not been trained to generate output values based on such morphologically abnormal cells.
[0030] In some cases, a computer-implemented method may include detecting that one or more subsets of image data represent morphologically abnormal cells based on respective uncertainty scores or confidence scores calculated for each of the one or more subsets of image data. For example, for a given subset of data, a computer-implemented method may include determining whether an uncertainty score generated relative to the subset of image data exceeds a predetermined maximum uncertainty threshold. And, if it is determined that the uncertainty score exceeds the predetermined maximum uncertainty threshold, then it is determined that the cells represented by the given subset of image data are morphologically abnormal. Similarly, for a given subset of data, a computer-implemented method may include determining whether a confidence score generated relative to the subset of image data is less than a predetermined minimum confidence threshold. And, if it is determined that the uncertainty score is less than the predetermined minimum confidence threshold, then it is determined that the cells represented by the given subset of image data are morphologically abnormal.
[0031] When a human or animal subject suffers from a condition that results in the presence of morphologically abnormal cells, the sample image will likely include a plurality of such cells. Thus, in addition to detecting morphological abnormalities based on confidence or uncertainty scores associated with individual cells, a statistical approach may be employed in which a histogram of uncertainty scores or confidence scores for the entire population of cells depicted in the sample image is considered. In these cases, the computer-implemented method may detect the presence of a subset of morphologically abnormal cells without necessarily identifying the specific cells that display the abnormality.
[0032] Thus, determining whether an abnormality is likely to exist may include: determining a proportion of confidence scores in the output data that are less than a predetermined minimum confidence threshold; and determining whether one or more morphologically abnormal cells are likely to exist based on the proportion of confidence scores less than the predetermined minimum confidence threshold. Thus, a computer-implemented method may include: determining a proportion of confidence scores in the output data that are less than a predetermined minimum confidence threshold; and determining whether one or more morphologically abnormal cells are likely to exist based on the proportion of confidence scores less than the predetermined minimum confidence threshold. Determining the proportion of confidence scores in the output data that are less than the predetermined confidence threshold may include: determining, for each subset of the data, whether the confidence score is less than or equal to the predetermined minimum confidence threshold; counting the number of subsets of data with confidence scores less than or equal to the predetermined minimum confidence threshold; and dividing the counted number by the total number of subsets of the data. Determining whether one or more morphologically abnormal cells are likely to exist based on the proportion of confidence scores less than the predetermined minimum confidence threshold may include: determining whether the proportion of confidence scores exceeds a predetermined threshold proportion; and if it is determined that the proportion of confidence scores exceeds the predetermined threshold proportion, determining that the presence of morphologically abnormal cells in the sample image is likely.
[0033] A similar process can be performed using uncertainty scores instead of confidence scores.
[0034] Specifically, determining whether an abnormality may exist may include: determining the proportion of uncertainty scores in the output data that exceed a predetermined maximum uncertainty threshold; and determining whether one or more morphologically abnormal cells may exist based on the proportion of uncertainty scores that exceed the predetermined maximum uncertainty threshold. Therefore, the computer-implemented method may include: determining the proportion of uncertainty scores in the output data that exceed a predetermined maximum uncertainty threshold; and determining whether one or more morphologically abnormal cells may exist based on the proportion of uncertainty scores that exceed the predetermined maximum uncertainty threshold. Determining the proportion of uncertainty scores in the output data that exceed the predetermined maximum uncertainty threshold may include: determining, for each subset of the data, whether the confidence score exceeds the predetermined maximum uncertainty threshold; counting the number of data subsets whose uncertainty scores exceed the predetermined maximum uncertainty threshold; and dividing the counted number by the total number of subsets of the data. Determining whether one or more morphologically abnormal cells may exist based on the proportion of uncertainty scores that exceed a predetermined minimum confidence threshold may include: determining whether the proportion of uncertainty scores exceeds a predetermined threshold proportion; and if it is determined that the proportion of uncertainty scores exceeds the predetermined threshold proportion, determining that the presence of morphologically abnormal cells in the sample image is possible.
[0035] In some cases, the output may include an indication of a subset of data for which the uncertainty score exceeds a predetermined maximum uncertainty threshold or the confidence score is less than a predetermined minimum confidence threshold. The computer-implemented method may then further include instructions that, when executed by a display component of the computing device, display an annotated version of the sample image. Preferably, the annotated version of the sample image includes an indication (e.g., in the form of an annotation or overlay) of cells for which the uncertainty score exceeds a predetermined maximum uncertainty threshold or the confidence score is less than a predetermined minimum confidence threshold. In this way, the clinician will have an ergonomically improved display that highlights cells that require further attention, ultimately reducing the time required to analyze the sample image.
[0036] The second aspect of the present invention can provide a clinical support system including a processor, which is configured to perform the computer-implemented method of the first aspect of the present invention. The clinical support system may further include a display component. It should be understood that the optional features described above with respect to the first aspect of the present invention are equally well applicable to the second aspect of the present invention, unless clearly incompatible or the context clearly stipulates otherwise. The clinical support system may include appropriate modules, which are configured to perform each of the different operations of the computer-implemented method of the first aspect of the present invention.
[0037] The third aspect of the present invention may provide a computer program comprising instructions, which, when executed by a processor of a computer, causes the processor to perform the computer-implemented method of the first aspect of the present invention. It should be understood that the optional features set forth above with respect to the first aspect of the present invention are equally well applicable to the third aspect of the present invention, unless clearly incompatible or the context clearly stipulates otherwise. The fourth aspect of the present invention may provide a computer-readable medium storing the computer program of the third aspect of the present invention. It should be understood that the optional features set forth above with respect to the first aspect of the present invention are equally well applicable to the fourth aspect of the present invention, unless clearly incompatible or the context clearly stipulates otherwise.
[0038] The fifth aspect of the present invention can provide a computer-implemented method for detecting the presence of morphologically abnormal cells in a sample image, the computer-implemented method comprising: receiving electronic image data representing a sample image, the sample image depicting a plurality of cells; applying an analysis model to each of a plurality of subsets of the image data, each subset corresponding to a corresponding portion of the sample image depicting a single cell, the analysis model being configured to output for each subset of the image data: a value of a parameterized cell attribute; and a confidence score or uncertainty score associated with the value, thereby generating output data including a plurality of confidence scores or a plurality of uncertainty scores; any of the following: determining the proportion of confidence scores in the output data that are less than a predetermined minimum confidence threshold; or determining the proportion of uncertainty scores in the output data that exceed a predetermined maximum uncertainty threshold; determining whether the proportion exceeds a predetermined threshold proportion; and generating an output indicating the result of the determination, wherein if the determination proportion exceeds a predetermined threshold proportion, the output includes a mark. The mark preferably alerts the clinician that morphologically abnormal cells may be present. It should be understood that the optional features described above with respect to the first aspect of the present invention are equally well applicable to the fifth aspect of the present invention, unless clearly incompatible or otherwise clearly specified in the context.
[0039] The sixth aspect of the present invention may provide a computer program comprising instructions, which, when executed by a processor of a computer, causes the processor to perform the computer-implemented method of the fifth aspect of the present invention. It should be understood that the optional features set forth above with respect to the first aspect of the present invention are equally well applicable to the sixth aspect of the present invention, unless clearly incompatible or the context clearly stipulates otherwise. The seventh aspect of the present invention may provide a computer-readable medium storing the computer program of the sixth aspect of the present invention. It should be understood that the optional features set forth above with respect to the first aspect of the present invention are equally well applicable to the seventh aspect of the present invention, unless clearly incompatible or the context clearly stipulates otherwise.
[0040] The eighth aspect of the present invention can provide a clinical support system including a processor, which is configured to perform the computer-implemented method of the fifth aspect of the present invention. The clinical support system may further include a display component. It should be understood that the optional features described above with respect to the first aspect of the present invention and the fifth aspect of the present invention are equally well applicable to the eighth aspect of the present invention, unless clearly incompatible or clearly specified in the context. The clinical support system may include appropriate modules, which are configured to perform each of the different operations of the computer-implemented method of the fifth aspect of the present invention.
[0041] The present invention includes any combination of described aspects and preferred features unless such a combination is expressly impermissible or explicitly avoided. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Embodiments of the present invention will now be described with reference to the accompanying drawings, in which:
[0043] - Figure 1 A clinical support system configured to perform the computer-implemented method according to the present invention is shown.
[0044] - Figure 2 is a flow chart illustrating an example of a computer-implemented method according to the present invention.
[0045] - Figure 3 is a histogram showing the uncertainty distribution in a normal cell population.
[0046] - Figure 4A and Figure 4B are histograms showing the uncertainty distribution in the cell population containing immature granulocytes and abnormal blasts, respectively.
[0047] Detailed description with drawings
[0048] Aspects and embodiments of the present invention will now be discussed with reference to the accompanying drawings. Other aspects and embodiments will be apparent to those skilled in the art. All documents mentioned in this text are incorporated herein by reference.
[0049] Figure 1 1 is a schematic diagram of a clinical support system 1 according to, for example, the first aspect of the present invention. The clinical support system 1 includes a processor 12, a memory 14, and a display unit 16. Figure 1In the present invention, these components are all shown as part of the same system, but it should be understood that the system can be a distributed system in which various components are located on different pieces of hardware, optionally in different locations. In those cases, the components (e.g., processor 12, memory 14, and display component 16) can be connected via a network (not shown). The network can be a wired network such as a LAN or WAN, or a wireless network such as a Wi-Fi network, the Internet, or a cellular network. We now discuss the structure of the clinical support system 1 and then refer to Figure 2 4 discusses the operations it is configured to perform. Processor 12 includes a plurality of modules. As used herein, the term "module" is used to refer to a functional module that is configured or adapted to perform a specific function. These modules may be implemented in hardware (i.e., they may be separate physical components within a computer), software (i.e., they may represent separate code segments that, when executed by processor 12, cause it to perform a specific function), or a combination of both.
[0050] Specifically, Figure 1 The processor 12 includes: a cell identification module 120, an analysis module 122, an uncertainty determination module 124, an anomaly detection module 126, and an output module 128. The functions of each of these modules will be described in more detail later. The memory 14 can be in the form of a permanent memory or a temporary memory, or can include a combination of both. The memory 14 stores an analysis model 140 and a set of thresholds 142. The display component 16 is preferably in the form of a VPU, a screen or a monitor, which is configured to visually present data to a clinician to view the results.
[0051] Figure 2 A process performed by the processor 12 of the clinical support system 1 is shown. In a first step S200, image data is received at the processor 12. More specifically, the image data is electronic image data representing a sample image, the sample image depicting a plurality of cells. The electronic image data naturally includes a plurality of subsets of the image data, each subset corresponding to a respective portion of the image data depicting a single cell. In step S202, the cell identification module 120 is configured to identify a subset of data corresponding to each individual cell. This can be accomplished by applying an image analysis algorithm to the electronic image data received in step S200. The output of step S202 is a plurality of subsets of the image data, each subset corresponding to a portion of the sample image depicting a single cell.
[0052] Steps S204 and S206 are performed for each subset of image data output by step S202. However, for the sake of brevity, we will only describe the process for a single subset of data. In step S204, the analysis model 140 is retrieved from the memory 14 and applied to the subset of image data by the analysis module 122. For each subset of data, the output of step S204 is a value that parameterizes the properties of the cells represented by the subset of image data. For example, the analysis model 140 can be in the form of a classification model, and the output of step S204 is a classification of the type of cells depicted. In step S206, the uncertainty determination module 124 determines the uncertainty value associated with the output generated by the analysis module 122 for each subset of image data. Various methods can be used to determine uncertainty. It should be noted that in other embodiments, a confidence score can be generated instead of an uncertainty score. The output of step S206 is an uncertainty score for each of the subsets of image data. In general, this can be referred to as output data.
[0053] In step S208, the anomaly detection module 126 of the processor 12 determines whether morphologically abnormal cells may be present in the sample image based on the output data. There are various ways to do this, but in one embodiment, the proportion of cells whose uncertainty scores exceed a predetermined maximum uncertainty threshold is determined. This proportion is then compared to a predetermined threshold proportion, and if it exceeds the predetermined threshold proportion, this is an indication that one or more morphologically abnormal cells are present in the sample image. This is because high uncertainty is associated with unfamiliar cell types, and the analysis model is either not trained or insufficiently trained due to a lack of training data (so-called "epistemic uncertainty"). Similarly, it should be understood that similar procedures can be performed using confidence scores instead of uncertainty scores.
[0054] After the determination is made by the uncertainty determination module 126 in step S208, an output is generated by the output module 128 in step S210. In some cases, the output may be transmitted to a database so that it can be accessed by the clinical computing network. Alternatively, the output may be a visual output, which may include a mark, as explained earlier in this application.
[0055] Experimental Results
[0056] The inventors were able to test the invention using a classification model configured to identify white blood cells. A training set of 8 (normal) sample slides was used, and a validation set of 4 sample slides was used. Each slide had approximately 600 images. For the abnormal example, 10 sample images including immature granulocytes and 13 sample images including blasts were used. A table explaining this is shown below
[0057] normal
[0058] # Slide type 8 Training set 4 Validation set
[0059] abnormal
[0060] # Slide type 10 Immature granulocytes (IG) 13 Mother Cell
[0061] Abnormal slides were used for verification only.
[0062] An ensemble of 5 convolutional neural networks was trained on normal sample images to classify cells into normal types: neutrophils, lymphocytes, monocytes, eosinophils, basophils, and immature granulocytes. The overall F1 score on the validation set reached 92% (5 normal classifications plus immature granulocytes). Due to the small number of normal slides, the validation set was also used for testing.
[0063] After training, the ensemble predictions of all abnormal and normal (validation set) slides are used as uncertainty approximation. The histogram of the uncertainty distribution is used to distinguish normal samples from abnormal samples. Figure 3 , Figure 4A and Figure 4B shown.
[0064] Figure 3 A normal sample is shown. The x-axis represents the uncertainty value, and the y-axis represents the number of cells having the corresponding uncertainty level in the sample image. Figure 4A shows a histogram representing the distribution of uncertainty in a sample containing morphologically abnormal immature granulocytes, and Figure 4B Histograms representing the distribution of uncertainty in samples containing morphologically abnormal blasts are shown. In each case, it should be noted that the scale on the x-axis is Figure 3 A is different. Figure 4A and Figure 4B It is clearly shown that the number of cells with higher uncertainty scores is significantly higher than Figure 3 A, which suggests that samples with morphologically abnormal cells can be detected from considerations of uncertainty or confidence scores.
[0065] General Statement
[0066] The features disclosed in the foregoing description, or in the following claims, in terms of the manner of expressing or implementing the disclosed functions in their specific forms, or in terms of the methods or processes for obtaining the disclosed results, may be appropriately used alone in their various forms, or in any combination to implement the present invention.
[0067] Although the present invention has been described in conjunction with the above exemplary embodiments, many equivalent modifications and variations will be apparent to those skilled in the art when this disclosure is given. Therefore, the above exemplary embodiments of the present invention are considered to be illustrative rather than restrictive. Various changes may be made to the described embodiments without departing from the spirit and scope of the present invention.
[0068] For the avoidance of any doubt, any theoretical explanations provided herein are intended to improve the reader's understanding. The inventors do not wish to be bound by any of these theoretical explanations.
[0069] Any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0070] Throughout the specification, including the following claims, unless the context requires otherwise, the words "comprise" and "include" and variations such as "comprises, comprising" and "including", will be understood to imply the inclusion of stated integers or steps or groups of integers or steps but not the exclusion of any other integers or steps or groups of integers or steps.
[0071] It must be noted that, as used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from "about" one particular value and / or to "about" another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent "about", it will be understood that the particular value forms another embodiment. The term "about" in relation to a numerical value is optional and means, for example, + / - 10%.
Claims
1. A computer-implemented method for detecting the presence of morphologically abnormal cells in a sample image, the computer-implemented method comprising: receiving electronic image data representing an image of a sample, the sample image depicting a plurality of cells; applying an analysis model to each of a plurality of subsets of the image data, each subset corresponding to a respective portion of the sample image depicting a single cell, the analysis model being configured to output, for each subset of the image data: a value parameterizing a property of the cell; and a confidence score or an uncertainty score associated with the value, thereby generating output data comprising a plurality of confidence scores or a plurality of uncertainty scores; as well as A determination is made based on the output data as to whether one or more morphologically abnormal cells may exist in the sample image.
2. The computer-implemented method of claim 1 , wherein: The analysis model is a classification model; and The value of the property parameterizing the cell is a numeric or textual output indicating the type of cell.
3. The computer-implemented method of claim 2, wherein: The classification model is a neural network model that has been trained to classify cells into one of a plurality of cell types based on data representing electronic images of the cells, the training using training data comprising a plurality of images of normal or healthy cells, each image being associated with a label indicating the type of the cell.
4. The computer-implemented method of claim 3, wherein: The classification model is configured to classify a cell as one of a plurality of types of white blood cells, the plurality of types of white blood cells comprising: neutrophils; lymphocytes; monocytes; eosinophils; and basophils.
5. A computer-implemented method according to claim 2 or claim 3, wherein: Applying the analysis model to each of the plurality of subsets of image data comprises: applying a plurality of analysis sub-models to each subset of the image data, each of the plurality of analysis sub-models being configured to output a corresponding plurality of values for each subset of the image data, the corresponding plurality of values each parameterizing a property of the cell; and A variance or standard deviation is determined for each of the plurality of values outputted, the determined variance or standard deviation corresponding to the uncertainty score.
6. The computer-implemented method of claim 5, wherein: The sub-models are trained based on the same training data using respective different initial random seeds for initializing each of the plurality of analysis sub-models prior to training.
7. A computer-implemented method according to any one of claims 2 to 4, wherein: The uncertainty scores were calculated using the discard method.
8. The computer-implemented method of any one of claims 1 to 7, further comprising: determining a proportion of said confidence scores in said output data that are less than a predetermined minimum confidence threshold; as well as A determination is made as to whether one or more morphologically abnormal cells are likely to be present based on the proportion of confidence scores that are less than the predetermined minimum confidence threshold.
9. The computer-implemented method of claim 8, wherein: Determining whether one or more morphologically abnormal cells may exist in the sample image based on the proportion of the confidence scores below a predetermined confidence threshold comprises: determining whether the ratio of confidence scores exceeds a predetermined threshold ratio; and If it is determined that the ratio of the confidence scores exceeds the predetermined threshold ratio, it is possible to determine the presence of morphologically abnormal cells in the sample image.
10. The computer-implemented method of any one of claims 1 to 7, further comprising: determining a proportion of said uncertainty scores in said output data that exceeds a predetermined maximum uncertainty threshold; as well as A determination is made as to whether one or more morphologically abnormal cells are likely to be present based on the proportion of the uncertainty scores that exceed the predetermined maximum uncertainty threshold.
11. The computer-implemented method of claim 10, wherein: Determining whether one or more morphologically abnormal cells may exist in the sample image based on the proportion of the uncertainty scores exceeding a predetermined uncertainty threshold comprises: determining whether the ratio of the uncertainty scores exceeds a predetermined threshold ratio; and If the ratio of the uncertainty scores is determined to exceed the predetermined threshold ratio, it is possible to determine the presence of morphologically abnormal cells in the sample image.
12. The computer-implemented method of any one of claims 1 to 11, wherein: If it is determined that an anomaly may be present in the sample image, the computer-implemented method further includes adding a marker to the electronic image data representing the sample image.
13. The computer-implemented method of any one of claims 1 to 12, wherein: The sample image is an image of a slide of a sample of a body fluid obtained from a human or animal subject.
14. The computer-implemented method of claim 13, wherein: The body fluid is blood.
15. A computer-implemented method for detecting the presence of morphologically abnormal cells in a sample image, the computer-implemented method comprising: receiving electronic image data representing an image of a sample, the sample image depicting a plurality of cells; Applying the analysis model to each of a plurality of subsets of the image data, each subset corresponding to a respective portion of the sample image depicting a single cell, the analysis model being configured to output, for each subset of the image data: a value parameterizing a property of the cell; and a confidence score or an uncertainty score associated with the value, thereby generating output data comprising a plurality of confidence scores or a plurality of uncertainty scores; Any of the following: determining the proportion of said confidence scores in said output data that are less than a predetermined minimum confidence threshold; or determining a proportion of said uncertainty scores in said output data that exceeds a predetermined maximum uncertainty threshold; determining whether the ratio exceeds a predetermined threshold ratio; as well as An output is generated indicating a result of the determination, wherein if the ratio is determined to exceed the predetermined threshold ratio, the output includes a flag.