Classification of events in flow cytometry data
Patent Information
- Application Number
- EP2024805033
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-10-25
- Publication Date
- 2026-09-09
AI Technical Summary
Current flow cytometry systems face challenges in efficiently classifying events due to subjective gate definition in region-based gating, which can lead to overlapping populations and data loss.
The implementation of an AI-driven event classifier that uses probabilistic thresholds and multiparametric data analysis to classify events more objectively and efficiently, allowing for real-time adjustment of thresholds without retraining the model.
This approach enables more accurate and objective classification of events, reduces data loss, and allows for dynamic threshold adjustments, improving the efficiency and reliability of flow cytometry analysis.
Smart Images

Figure US2024053060_08052025_PF_FP_ABST
Abstract
Description
CLASSIFICATION OF EVENTS IN FLOW CYTOMETRY DATACROSS-REFERENCE TO RELATED APPLICATION
[0001] This application is being filed on October 25, 2024, as a PCT International application and claims the benefit of and priority to U.S. Application No. 63 / 594,165, filed on October 30, 2023, entitled CLASSIFICATION OF EVENTS IN FLOW CYTOMETRY DATA, the disclosure of which is hereby incorporated by reference in its entirety7.BACKGROUND
[0002] Flow cytometry is a technique for detecting and analyzing chemical and physical characteristics of cells or particles in a fluid sample. For example, a flow cytometer may be used to assess cells from blood, bone marrow, tumors, or other body fluids. Typically, the sample is passed through a fluid nozzle which aligns particles in a single file line within a sheath fluid. A laser beam illuminates the particles as they pass through in single file to generate radiated light including forward scattered light, side scattered light, and fluorescent light. The radiated light can then be detected and analyzed to determine one or more characteristics of the particles.SUMMARY
[0003] In general terms, the present disclosure relates to analyzing particles using flow cytometry7. In one possible configuration, an Al algorithm is used to automatically classify events using one or more control variables / thresholds that are automatically adjusted. In another possible configuration, one or more control variables / thresholds are manually adjusted to correct the classification of events and if necessary7to further augment training data sets. Various aspects are described in this disclosure, which include, but are not limited to, the following aspects.
[0004] One aspect relates to a method for analyzing particles within a flow cytometry system using a computing device. The method includes receiving flow cytometry7data from the flow cytometry7system, the flow7cytometry7data including at least one event; assigning a probability value to the at least one event, the probability value indicating a probability that the at least one event belongs to a cluster / class of events; displaying the probability value to the user, wherein the probabilistic thresholdcategorizes an event as belonging to the population of interest: allowing the user to adj ust the probabilistic threshold; and generating an output. The output is generated by comparing the probability value of the at least one event to the probabilistic threshold and assigning categorical information to at least one inclusive event, wherein the at least one inclusive event includes the probability value exceeding the probabilistic threshold.
[0005] Another aspect relates to a method for adjusting the classification of results within a flow cytometry system. The method includes inspecting flow cytometry data, the flow cytometry data comprising at least one event, the at least one event including a probability value, the probability' value indicating a probability that the at least one event is associated with a population of interest; allowing an adjustment of a probabilistic threshold on the flow cytometry system to filter out one or more of the at least one event when the at least one probability value is greater than the probabilistic threshold input; and receiving an output from the flow cy tometry’ system. The output includes at least one inclusive event that exceeds the probabilistic threshold input.
[0006] Yet another aspect relates to a flow cytometry system for analyzing particles. The flow cytometry system includes: at least one processing device; and a non-transitory computer readable storage media storing instructions which, when executed by the processing device, cause the at least one processing device to: receive flow cytometry7data from the flow cytometry system, the flow cytometry data comprising at least one event; assign a probability value to the at least one event, the at least one probability value indicating a probability that the at least one event is located within a population of interest; display the at least one probability value to a user; allowing for the adjustment of a probabilistic threshold from the user; and generate an output. The output is generated by comparing the at least one probability value to the probabilistic threshold; identifying at least one inclusive event when the at least one probability value exceeds the probabilistic threshold; and displaying the at least one inclusive event to the user.
[0007] A variety' of additional aspects will be set forth in the description that follow s. The aspects can relate to individual features and to combinations of features. It is to be understood that both the foregoing general description and the following detailed description are exemplary' and explanatory' only and are not restrictive of the broad inventive concepts upon which the examples disclosed herein are based.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The following drawings are illustrative of examples of the present disclosure and therefore do not limit the scope of the present disclosure. Examples of the present disclosure will hereinafter be described in conjunction with the appended drawings, wherein like numerals denote like elements.
[0009] FIG. 1 schematically illustrates an example of a flow cytometer system.
[0010] FIG. 2 schematically illustrates an example of a waveform analysis device of the flow cytometer system of FIG. 1.
[0011] FIG. 3 illustrates an example of a method for generating an output illustrating events that are likely to be included within a cluster / class of events adjustment based on a default probabilistic threshold or one received from the user.
[0012] FIG. 4 illustrates an example of a method for manually adjusting flow cytometry data that is autonomously produced by the artificial intelligence models of FIGS. 2-3.
[0013] FIG. 5 illustrates an example of an artificial intelligence model of FIGS. 2-4 having at least one input, at least one node in a first hidden layer, at least one node in a second hidden layer, and an output produced by the artificial intelligence model.
[0014] FIG. 6 illustrates an example of a method for training the artificial intelligence models of FIG. 2 to evaluate flow cytometry data.
[0015] FIG. 7 illustrates an example gate adjustment performed by the waveform analysis device of FIGS. 1.
[0016] FIG. 8 illustrates an example of a probabilistic event classification performed by the waveform analysis device of FIGS. 1-2 corresponding to a selection of a probabilistic threshold of fifty percent that is received.
[0017] FIG. 9 illustrates an example of a probabilistic event classification performed by the waveform analysis device of FIGS. 1-2 corresponding to a selection of a probabilistic threshold of ninety -five percent.
[0018] FIG. 10 includes example histogram plots having varying probability distributions between two cell populations, which decreases a confidence rating generated by the waveform analysis device of FIGS. 1-2.
[0019] FIG. 11 illustrates the flow cytometry data illustrated in three different resolutions.
[0020] FIG. 12 illustrates an exemplary architecture of a computing device that can be used to implement aspects of the present disclosure.
[0021] In the appended figures, similar components and / or features can have the same reference label. Further, various components of the same type can be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label.DETAILED DESCRIPTION
[0022] Various examples will be described in detail with reference to the drawings, wherein like reference numerals represent like parts and assemblies throughout the several views.
[0021] FIG. 1 schematically illustrates an example of a flow cytometer system 100. In some instances, the flow cytometer system 100 can include aspects and features described in U.S. Provisional Patent Application No. 63 / 410,984, filed September 28, 2022, which is herein incorporated by reference in its entirety.
[0022] In general, flow cytometry is a technique for measuring and analyzing properties of particles or cells when flowing in a fluid stream 1 14. Data from millions of particles or cells can be collected by the flow cytometer system 100 in a matter of minutes and displayed in a variety of formats. Illustrative example applications of flow cytometry include phenotyping to identify and count specific cell types within a population, analyzing DNA or RNA content within cells, determining presence of antigens on a surface or within cells, and assessing cell health status.
[0023] As show n in the illustrative example of FIG. 1, the flow cytometer system 100 generally includes three main component subsystems: a fluidic system 110, an optical system 120, and an electronic system 130. The fluidic system 1 10 includes a nozzle 112 which receives a sample containing particles or cells suspended in a fluid. The nozzle 112 creates and ejects a fluid stream 114 of the particles or cells arranged in a single file line. Each particle or cell passes through one or more beams of light produced by a light source 102. The point at which a particle or cell intersects with a light beam is known as an interrogation zone 116. In some examples, the light source 102 includes one or more lasers.
[0024] The optical system 120 includes the light source 102, optical elements 122, and detectors 124. At the interrogation zone 116, light from the light source 102 hits a particle or cell in the fluid stream 114 and scatters. The optical elements 122 direct thescattered light toward the detectors 124. The detectors 124 can include a forward scatter (FSC) detector to measure scatter in the path of the light source 102, a side scatter (SSC) detector to measure scatter at a ninety -degree angle relative to the light source 102, and one or more fluorescence detectors (FL1, FL2, FL3 ... FLn) to measure the emitted fluorescence intensity at different wavelengths of light.
[0025] Generally, FSC intensity is proportional to the size or diameter of a particle due to light diffraction around the particle. FSC may therefore be used for the discrimination of particles by size. SSC, on the other hand, is produced from light refracted or reflected by internal structures of the particle and may therefore provide information about the internal complexity or granulanty of the particle. By adding fluorescent labelling to a sample, different fluorescent signals / channels (e.g.. green, yellow, and red) can be analyzed for functional characteristics of a cell. For example, since T-cells present CD3 binding sites, a sample containing T-cells may be “stained” with anti-CD3 antibodies conjugated with a fluorescent molecule. As these cells pass through the interrogation zone 116, the light from the source light excites the fluorescent tag, or fluorochrome, to emit photons at a wavelength detectable by a fluorescence detector. The detectors 124 may therefore simultaneously measure several parameters and enable categorization of particles by their function based on detected wavelengths of light.
[0026] The electronic system 130 includes a waveform acquisition device 140 and a waveform analysis device 150. The waveform acquisition device 140 is communicatively coupled with the detectors 124 to receive analog waveform data 126 generated by the detectors 124. The waveform acquisition device 140 includes an analog-to-digital converter (ADC) 142 configured to digitize the waveform data.
[0027] The waveform analysis device 150 is configured to receive the digital waveform data and display it for a user of the flow cytometer system 100. In some embodiments, the waveform analysis device 150 comprises a computing device communicatively coupled with a flow cytometer 101, such as over a network. The flow cytometer 101 may include the fluidic system 110, optical system 120. and waveform acquisition device 140. In other embodiments, the waveform analysis device 150 is integrated with the flow cytometer 101.
[0028] Current flow cytometers use a field-programmable gate array (FPGA) in the waveform acquisition device 140 to obtain information about individual particles passing through the light beam. The w aveform acquisition device 140 uses a singlethreshold value to determine when the output of the detectors 124 begins conversion from analog to digital. As such, if or when a detector 124 outputs a voltage value that crosses the threshold, digitization begins, and the digital value is sent to the FPGA. As waveform data is digitized, the FPGA computes the height, width, and area of each pulse. Besides the height, width, and area of each pulse, other data relating to the waveform, including data not exceeding the voltage threshold value, is not captured, stored, or otherwise available for analysis. If a user wishes to adjust the threshold value, the experiment has to be re-run with the new threshold value, incurring costs in resources and time.
[0029] Furthermore, if a user wishes to gate the results (i.e. , select specific populations of cells or particles based on their unique characteristics), region-based gating is often used to define boundaries around specific populations on tw o- dimensional scatterplots, such as a forw ard scatter (FSC) v. side scatter (SSC) plot. Region-based gating includes several limitations, where the gate is often subjectively defined by a user based on visual inspection of the data, the defined regions may include overlapping populations that make it challenging to define separate gates for each population based on the scatterplot data, and certain subpopulations of interest may be discarded during the gating process when one subpopulation is selected over another.
[0030] To address the above issues, the example shown in FIG. 1 illustrates an event classifier 154 included as a component of the waveform analysis device 1 0. The event classifier 154 allows a user to classify events as being part of a population of interest or cluster of events more efficiently and effectively than standard region-based gating methods. In certain examples, the event classifier 154 allows the user to classify’ samples using artificial intelligence models, such as a machine learning algorithm, that define regions using multiparametric data through statistical probabilities rather than one and two dimensional subjective gating, which is described and illustrated in further detail with respect to FIGS. 3-11. Furthermore, the event classifier 154 allows a user to cluster / classify events by comparing multiple dimensions (e.g.. forward scatter, impedance, side scatter, fluorescent channels pertaining to the detection of light scatter or emission from fluorochromes or fluorescent dyes, time taken by particles to pass through the interrogation point, etc.) of data at a time rather than analyzing only one or two dimensions of data at a time. Further yet. the event classifier 154 can cluster populations without discarding specific subpopulations during the gating process.
[0031] Furthermore, the flow cytometer system 100 includes a graphics processing unit (GPU) 152. In the example illustrated in FIG. 1, the GPU 152 is shown included as a component of the waveform analysis device 150. The GPU 152 processes a continuous digital stream generated by the waveform acquisition device 140. The digital stream processed by the GPU 152 is continuous in that the waveform acquisition device 140 collects data pertaining to one or more pulses that is then utilized by the artificial intelligence models 202 (as shown in FIG. 2) to produce a probabilistic cluster / class of events.
[0032] The waveform acquisition device 140 captures an electrical signal (a.k.a. a “pulse"’) that is generated when a particle passes through the light source 102 and is detected by the detectors 124. In doing so, the waveform acquisition device 140 captures pulse attributes that are measured on a two-dimensional histogram such as the pulse peak, area, width, and half height. These parameters are transmitted to the GPU 152 for each pulse that is detected.
[0033] Given the foregoing description, the waveform analysis device 150 calculates a digitized version of the waveform data with increased data points, and the waveform data for an experiment is displayed and available in its entirety for processing by the GPU 152. The event classifier 154 allows the user to classify one or more events in several dimensions while performing objective analysis and limiting data loss during the event classification process. Because the probabilistic data is maintained, the user has the ability to dynamically adjust thresholds and update graphical plots in real-time without re-training the artificial intelligence models 202. Further details of operation and advantages are discussed below-.
[0034] The flow cytometer system 100 includes elements which are shown and described for purposes of discussion, and it will be appreciated that numerous variations in components and functions are possible. The optical elements 122 may include a series of filters, dichroic mirrors, and / or beam splitters to select out different wavelengths of light and provide the wavelength to the appropriate detector 124. The detectors 124 may comprise, for example, photomultiplier tubes (PMTs) or avalanche photodiodes (APDs) or single photon counting devices.
[0035] FIG. 2 schematically illustrates an example of a waveform analysis device 150 of the flow- cytometer system 100 of FIG. 1. The waveform analysis device 150 receives, stores, and displays waveform data. The waveform analysis device 150 includes an interface 210 to receive digitized raw waveform data 232, a persistentstorage 230 to store the digitized raw waveform data 232. and can include a graphical user interface (GUI) 220 to display the digitized raw waveform data 232. The persistent storage 230 may also store a plurality of event probability values that allow for thresholding and real-time updating and displaying of applied thresholds as further described below. The persistent storage 230 may comprise system memory such as random-access memory (RAM) and / or long-term non-volatile memory such as a hard drive.
[0036] The waveform analysis device 150 may further include the event classifier 154 comprising a software application or a set of related software applications configured to instruct the GPU 152 to process the digitized raw waveform data 232. The event classifier 154 may execute on one or more processors or in the cloud to provide the functionality described herein in conjunction with the GPU 152 such as receiving user input via the GUI 220. The event classifier 154 may include one or more artificial intelligence models 202, such as a machine learning algorithm, to analyze data received from the waveform acquisition device 140. One or more components of the waveform analysis device 150 may reside in a cloud computing application in a network distributed system. In that regard, the waveform analysis device 150 may be any of a variety7of computing devices, including, but not limited to, a personal computing device, a server computing device, or a distributed computing device.
[0037] FIG. 3 illustrates an example of a method 300 for generating an output illustrating events that are likely to be clustered based on a probabilistic threshold received from the user.
[0038] At step 302, the event classifier 154 receives flow cytometry data including at least one event. The flow cytometry data is collected by the flow cytometer system 100 as illustrated and described above with reference to FIGS. 1 and 2. The flow cytometry data includes at least one event. An event is a single measurement or detection of a particle (e.g., a cell, microorganism, compensation bead, protein, etc.) as it passes through the light source 102 at the interrogation zone 116. In certain examples, the at least one event measures the characteristics of blood cells, such as leukocytes and erythrocytes. The event is detected by the detectors 124 before it is received by the waveform acquisition device 140. The waveform acquisition device 140 digitizes the waveform data using the ADC 142 and transmits the waveform data to the waveform analysis device 150. The waveform analysis device 150 receives the waveform data andprepares the waveform data for further analysis, which can be performed by components such as the event classifier 154.
[0039] At step 304 the event classifier 154 assigns a probability value to each of the events received within the flow cytometry data. The probability value relates to a probability that each event is located within a desired class (i.e., the probability that a particle detected by the detector 124 is part of a population of interest, such as a specific cell type contained within a sample). In certain examples, the probability value is a numerical value within a range from zero to one. In certain examples, the probability value is a percentage within a range from zero to one hundred or can be formatted for a flow cytometry histogram classification on a 10-bit resolution of 0 to 1023.
[0040] In certain examples, the artificial intelligence models 202 are used to analyze the flow cytometry data and assign the probability value to each event within the flow cytometry data. As described above with reference to FIG. 1, the artificial intelligence models 202 may assign the probability value to each event by analyzing the flow cytometry data and determining a likelihood that each event is included within the desired class (i.e.. a cluster / class of events with one or more characteristics pertaining to a population of interest). In certain examples, the artificial intelligence models 202 analyze whether each event is included within the desired class by analyzing flow cytometry data that includes multiple dimensions, such as data relating to the forward scatter, the side scatter, fluorescent channels pertaining to the detection of light emitted from fluorochromes or fluorescent dyes, or the time taken by particles to pass through the interrogation point. Analyzing the data in multiple dimensions allows the artificial intelligence models 202 to assign a more accurate probability value to each of the events compared to a process by which a flow cytometry data is analyzed using only one or two dimensions. Furthermore, the artificial intelligence models 202 may objectively assign a probability value to one or more events within the flow cytometry data without relying upon human intervention, which would have included drawing a region around certain events within a two-dimensional histogram (e.g., a plot of side scatter vs. forward scatter).
[0041] The artificial intelligence models 202 (e.g., a machine learning algorithm) can be trained using flow cytometry data to assign accurate and precise probability values to each event. In certain examples, the artificial intelligence models 202 are trained using a supervised learning process. The supervised learning process includes teaching the artificial intelligence models 202 how to correctly identify a desired cluster of eventswithin the input flow cytometry data. In certain examples, this includes a process by which a user manually identifies the desired cluster of events and the artificial intelligence models 202 develop a method for mapping the input data to the cluster of events identified by the user. Once the artificial intelligence models 202 are properly trained (z. e. a confidence rating is generated based on a relatedness between the output of the artificial intelligence models 202 and the desired cluster manually identified by the user, and that confidence rating is above a threshold value), the artificial intelligence models 202 generate probabilistic event classifications. In certain examples, the probabilistic event classifications can be manually adjusted by the user. A process by which the artificial intelligence models 202 are trained is described and illustrated in further detail with respect to FIG. 6.
[0042] At step 306 the at least one probability value is displayed to the user. In certain examples, the at least one probability value is displayed to the user on a display device 1242 (not shown) as described and illustrated in further detail with respect to FIG. 12.
[0043] At step 308 the event classifier 154 receives classifies at least one event based on a probabilistic threshold. The probabilistic threshold may be a default threshold or it may be received from the user. In certain examples, the probabilistic threshold is received from the user through one or more input devices 1226, which is then transmitted to a processing device 1202 via an input / output interface 1236 and a system bus 1206 (1226, 1202, 1236, and 1206 not shown) as illustrated and described in further detail with respect to FIG. 12.
[0044] The probabilistic threshold is a threshold value that the event classifier 154 compares to each of the assigned probability values to determine whether an event should be classified as being within a population of interest (z.e., any probability value that is greater than the probabilistic threshold is determined to be within the population of interest and can be assigned a value of one if binary categorical information is assigned, and any probability’ value that is less than the probabilistic threshold is determined to be outside of the region of interest and can be assigned a value of zero if binary categorical information is assigned). The user can adjust the probabilistic threshold to require a greater or lesser probability’ that the event is included within the population of interest. A higher probabilistic threshold will output fewer events that are more likely to be included in that population of interest. Conversely, a lower probabilistic threshold will output more events, but the events may be less likely to be included within the population of interest.Ill certain examples, the user may optimize the probabilistic threshold based on visual histogram appearance or to include a desired number of events with the greatest probability of being included within the population of interest.
[0045] At step 310, the event classifier 154 generates an output. The output is generated by comparing each probability value to the probabilistic threshold, identifying at least one inclusive event (i.e., an event that is within the population of interest) when the probability value exceeds the probabilistic threshold (or some function of the threshold), and displaying the at least one inclusive event to the user. In certain examples, the event classifier 154 identifies at least one non-inclusive event as not being within the population of interest when the probability value does not exceed the probabilistic threshold. The event classifier 154 may filter out the at least one non-inclusive event from the at least one event.
[0046] Each probability value is compared to the probabilistic value to produce an output that is a binaiy characterization. Events that are determined to be within the population of interest (i.e., those that include an assigned probability value greater than the probabilistic threshold, are assigned an output value of one. Conversely, events that are determined to be outside of the population of interest (i.e., those that include an assigned probability value less than the probabilistic threshold) are assigned an output value of zero. After the comparison is made between each probability value and the probabilistic threshold, one or more inclusive events are identified that includes an output value of one. Each inclusive event can then be displayed to the user (e.g., in a table, scatter plot, or histogram) to illustrate the events that are determined to be within the population of interest based on the probabilistic threshold supplied by the user.
[0047] Furthermore, upon inspection of the output, a user might elect to optimize the output by adjusting the probabilistic threshold that is provided to the event classifier 154. If the user chooses to adjust the probabilistic threshold, the event classifier 154 will revert to step 306 to display the at least one probability value to the user and allow the user to input a new probabilistic threshold. In certain examples, the event classifier 154 receives one or more additional probabilistic thresholds from the user and generates one or more additional outputs corresponding to the one or more additional probabilistic thresholds.
[0048] Once an adjusted probabilistic threshold is received from the user, an adjusted output can be generated based on the adjusted probabilistic threshold. The user may adjust the probabilistic threshold without re-running the experiment with a new threshold value, which reduces costs in resources and time of retraining the Al model. In certainexamples, the binary characterization relating to one or more events may be manually adjusted by the user.
[0049] FIG. 4 illustrates an example of a method 400 for manually adjusting flow cytometry data that is autonomously produced by the artificial intelligence models 202 of FIGS. 2-3.
[0050] The method 400 includes a step 402 of inspecting flow cytometry data. The user inspects flow cytometry data including at least one event. The event includes a probability value that is assigned to the event by the event classifier 154 that indicates the probability that the at least one event is located within the population of interest. The user inspects the flow cytometry data to determine various characteristics of the data, such as the number of events and the probabilities assigned to each event. The process of receiving flow cytometry data, assigning a probability value to each event, and displaying the event to the user is described and illustrated above with reference to FIG. 3.
[0051] The method 400 includes a step 404 of providing a probabilistic threshold to the event classifier 154. The user may provide the probabilistic threshold input to the event classifier 154 based on observations made during step 402. In certain examples, the user may observe many events that include a high probability value and elect to provide a high probabilistic threshold to correspond with the observed high probability. In another example, the user may observe few events that include a high probability value and elect to provide a lower probabilistic threshold to increase the number of inclusive events having a probability value greater than the probabilistic threshold.
[0052] The method 400 includes a step 406 of receiving an output from the event classifier 154. As described above with reference to FIG. 3, the output includes the binary characterization with events including a non-inclusive events being assigned a value of zero and inclusive events being assigned a value of one. The output is displayed to the user, where the user can repeat the method 400 by reinspecting the flow cytometry data and output, providing an adjusted probabilistic threshold to the event classifier 154, and receiving an adjusted output from the event classifier 154.
[0053] FIG. 5 illustrates an example of an artificial intelligence model 202 of FIGS. 2-4 having at least one input 500, at least one node 512 in a first hidden layer 510, at least one node 522 in a second hidden layer 520, and an output 530 produced by the artificial intelligence model 202. In certain examples, the artificial intelligence model 202 includes a machine learning algorithm.
[0054] In certain examples, the at least one input 500 includes a first input 502, a second input, 504, or any number of inputs 506. The at least one input 500 includes flow cytometry data collected from the flow cytometer system 100.
[0055] The first hidden layer 510 includes at least one node 512. In certain examples, the first hidden layer includes a plurality of nodes 512, 514, 516, or any number of nodes 518. The at least one node 512 receives the at least one input 500 and uses a mathematical function to transform the at least one input 500 into an output that is transmitted to the second hidden layer 520. In certain examples, the specific mathematical function used by the at least one node 512 is manually inputted by a user. In certain examples, the specific mathematical function used by the at least one node 512 is determined by the architecture of the artificial intelligence models 202 and a machine learning algorithm used to train the artificial intelligence models 202. In certain examples, the machine learning algorithm iteratively adjusts weights and biases of connections between nodes within a hidden layer to identify and extract desired features of the flow cytometry data. In certain examples, the number of hidden layers, the number of nodes in each hidden layer, and other hyperparameters can be adjusted to optimize the accuracy and computational efficiency of the artificial intelligence models 202.
[0056] The second hidden layer 520 includes at least one node 522. In certain examples, the second hidden layer 520 includes a plurality of nodes 522, 524, 526, or any number of nodes 528. The second hidden layer 520 transforms a mathematical function to transform data received from the first hidden layer 510 and generate an output 530 (e.g., the output described and illustrated in FIGS. 3-4). In certain examples, the mathematical function performed by the second hidden layer 520 is the same or substantially similar to the mathematical function performed by the first hidden layer 510. In other examples, the mathematical function performed by the second hidden layer 520 is different than the mathematical function performed by the first hidden layer 510. The artificial intelligence models 202 may include any number of hidden layers (i.e., one hidden layer, two hidden layers, or more than two hidden layers).
[0057] FIG. 6 illustrates an example of a method 600 for training the artificial intelligence models 202 of FIGS. 2-5 to evaluate flow cytometry data.
[0058] The method 600 includes a step 602 of collecting flow cytometry data. The flow cytometry data is collected using the flow cytometer system 100 illustrated and described above with reference to FIG. 1.
[0059] The method 600 includes a step 604 of reviewing the flow cytometry data. At step 604, collected flow cytometry data is reviewed by a user. In certain examples, the user is an expert in analyzing flow cytometry data. The user may review and analyze the flow cytometry data by creating one-dimensional and / or two-dimensional scatterplots that visually represent the flow cytometry data.
[0060] The method 600 includes a step 606 of identifying one or more populations of interest (e.g., a white blood cell) within flow cytometry data by labeling events that correspond to the one or more populations of interest. In certain examples, the one or more populations of interest are manually identified by a user. The flow cytometry data may be arbitrated, or compared to an independent data set, to ensure the one or more population of interests are accurately provided to the event classifier 154.
[0061] The method 600 includes a step 608 of generating a machine learning algorithm. The event classifier 154 creates the machine learning algorithm using the flow cytometry data, the one or more populations of interest are provided, and an output is created that includes an indication of whether each event should be classified within one or more populations of interest.
[0062] In certain examples, the machine learning algorithm is enhanced over one or more training iterations using supervised learning with a labeled data set corresponding to one or more correct outputs. The user may provide flow cytometry data including a scatter plot having one or more populations of interest and an indication of whether the events within the scatter plot are included in the one or more populations of interest. The machine learning algorithm can then be generated by mapping the input data (including scatter plots and populations of interest) to the output data (the indication of whether each event is within the population of interest) to identify patterns and relationships within the data that can be used to predict whether events are within a population of interest. The supervised training process includes iteratively adjusting the machine learning algorithm’s parameters based on differences between the machine learning algorithm’s output and a correct output provided by the user. Once the machine learning algorithm is trained, it can be tested on new data as discussed below in step 610.
[0063] In other examples, the machine learning algorithm is enhanced over one or more training iterations using unsupervised learning by finding inherent patterns or structures in the data provided without explicit guidance from a user.
[0064] The method 600 includes a step 610 of evaluating new flow cytometry data using the machine learning algorithm, generating an output prediction of whether eventsare within one or more populations of interest, and comparing the output prediction from the machine learning algorithm to the labeled events corresponding to the populations of interest to generate a confidence rating. The confidence rating measures the machine learning algorithm’s ability to analyze the events and correctly predict whether the events are within a population of interest. In certain examples, the confidence rating is a numerical value between zero and one. In certain examples, the confidence rating is a percentage between zero and one hundred.
[0065] The method 600 includes a step 612 of evaluating whether the confidence rating exceeds a threshold value. This is done by comparing the confidence rating generated in step 610 to a threshold value inputted by the user. The method 600 includes a step 616 of locking the machine learning algorithm (z.e., the event classifier 154 will not make any additional adjustments to the machine learning algorithm) if the results produced by the confidence rating exceeds or achieves accuracy specifications. The method 600 includes a step 614 of allowing a user to collect more training data and repeat.
[0066] Furthermore, once additional training data is collected, any of steps 602, 604, 606, and 608 can be repeated to improve the training of the machine learning algorithm and better predict the correct output.
[0067] FIG. 7 illustrates an example gate adjustment performed using manual gating where an amorphous region is adjusted. Figure 7 includes a graph of the emitted fluorescent intensity of a fluorescent marker (“CD45-FITC”) 704 along the x-axis and the intensity of side-scattered light 702 measured by a side scatter detector along the y- axis. The graph includes many events 708 that are detected by detectors 124 within the flow cytometer system 100. Furthermore, the graph includes a gated region 706 identified by a polygonal shape drawn in the lower-righthand comer of the graph.
[0068] In certain examples, and as described and illustrated above with respect to FIG. 6, the gated region is provided by a user to train the machine learning algorithm. In certain examples, the machine learning algorithm identifies the events within the gated region 706 autonomously based on unique characteristics of the events from the waveform acquisition device 140.
[0069] FIG. 8 illustrates an example of a probabilistic event classification performed by the waveform analysis device 150 corresponding to a selection of a probabilistic threshold of fifty percent. In certain examples, the probabilistic threshold can include a default setting (e.g., 50%) that is adjustable by a user and may be overridden. FIG. 8shows a scatter plot graph 800 and a histogram plot 810. The scatter plot graph illustrates the intensity of side scattered light 802 on the x-axis and the intensity of forward- scattered light 804 on the y-axis. Furthermore, the scatter plot graph 800 shows a plurality of events 806 that the event classifier 154 has selected as being within the gated region 706 shown in FIG. 7.
[0070] The histogram plot 810 includes a probability 812 that an event lies within the population of interest on the x-axis and a number of events having a certain probability value 814 on the y-axis. Furthermore, the histogram plot 810 shows the probabilistic threshold 814 (which is set at 50%). In certain examples, the probability that an event lies within the population of interest is a numerical value between zero and one. In certain examples, the probability that an event lies within the population of interest is a percentage between zero and one hundred. In other examples, the probability that an event lies within the population of interest is scaled to a range including a minimum value and a maximum value to correspond with existing flow cytometry software (e.g., a range from 0-1023 as shown in FIGS. 8-10).
[0071] As shown in the histogram plot 810, and by non-limiting example, 25,939 total events were included in the flow cytometry data. Adjusting the probabilistic threshold 814 to 50% filtered out all events that the machine learning algorithm determined to have a probability value of less than 50% while displaying all of the inclusive events. As shown in the histogram plot 810, and by non-limiting example, the number of inclusive events is 10,134 events, which is approximately 39.1 % of the total events included in the flow cytometry data. In general, increasing the probabilistic threshold 814 will reduce the number of inclusive samples and provide a greater likelihood that the inclusive samples displayed to the user are within the population of interest. In contrast, decreasing the probabilistic threshold will increase the number of inclusive samples while providing a decreased likelihood that the inclusive samples displayed to the user are within the population of interest. An example of increasing the probabilistic threshold is illustrated and described in FIG. 9.
[0072] FIG. 9 illustrates an example of a probabilistic event classification performed by the waveform analysis device 150 corresponding to a selection of a probabilistic threshold of ninety-five percent. FIG. 9 illustrates an example scatter plot graph 900 and an example histogram plot 910 that is similar to the example scatter plot graph 800 and the example histogram plot 810 shown in FIG. 8. In the example histogram plot 910, the probabilistic threshold 912 has been increased to a 95% thresholdprobability, and inclusive events 914 are shown that include assigned probability values that exceed the probabilistic threshold 912. As shown in FIG. 9, and by non-limiting example, 25,939 total events were included in the flow cytometry data. Adjusting the probabilistic threshold 912 to 95% displayed the 6,693 events that the machine learning algorithm determined to have a probability value of greater than 95%. Thus, by comparing the probability value of each event to the 95% probabilistic threshold 912, the machine learning algorithm determined that approximately 25.8% of the events 708 shown in FIG. 7 (6,693 inclusive events out of 25,939 total events) are located within the gated region 706.
[0073] FIG. 10 includes example histogram plots 1000, 1010 from to different population classifiers. The two classifiers have varying probability distributions. When the populations are more difficult to separate, the probability distributions are more spread as is the case in 1010, which decreases a confidence rating generated by the waveform analysis device of FIGS. 1-2. FIG. 10 includes a high-confidence histogram plot 1000 and a low-confidence histogram plot 1010. The high-confidence histogram plot 1000 and the low-confidence histogram plot 1010 illustrate the probability that an event lies within the population of interest on the x-axis and the number of inclusive events 816 having a probability value that exceeds the probabilistic threshold 814 on the y-axis similarly to FIG. 8. Furthermore, the high-confidence histogram plot 1000 and the low- confidence histogram plot 1010 include the fifty-percent probabilistic threshold 814 of FIG. 8. The high-confidence histogram plot 1000 and the low-confidence histogram plot 1010 include a distribution of probability values assigned by the machine learning algorithm to each event within the flow cytometry data. A concentration of probabilities within a limited range of probability values (as shown in the high-confidence histogram plot 1000) at the upper and lower end of the histogram indicates a greater confidence rating that the outputs generated by the machine learning algorithm (i.e., the determination of which events are inclusive events and which events are not inclusive events) are correct. Conversely, a greater distribution of probability values (as shown in the low-confidence histogram plot 1010) indicates a lower confidence rating that the outputs generated by the machine learning algorithm are correct.
[0074] FIG. 11 illustrates the flow cytometry data illustrated in three different resolutions. FIG. 11 includes a 6-bit display 1100 showing events in a 6-bit resolution 1102, an 8-bit display 1110 showing events in an 8-bit resolution 1104, and a 9-bit display 1 120 showing events in a 9-bit resolution 1 106. Each display includes the samenumber of events, but the resolution effects the size of the clusters of events and the accuracy by which a user can classify an event as being within a population of interest. When performing visual inspection of the data, higher resolutions tend to increase the accuracy of defining an event as being within a population of interest while lower resolutions tend to decrease the accuracy of defining an event as being within the population of interest. Determining whether an event is included within a population of interest using the probabilistic threshold eliminates potential errors that are introduced by the subjectivity of visual inspection. Furthermore, the event classifier 154 can interpret data in greater than two dimensions to perform this determination, whereas visual inspection of the graphs shown in FIG. 11 is limited to two dimensions.
[0075] FIG. 12 illustrates an exemplary architecture of a computing device that can be used to implement aspects of the present disclosure. The computing device 1200 can be used to implement aspects of the present disclosure, including the aspects of the waveform acquisition device 140 and the waveform analysis device 150, as described above. Furthermore, the computing device 1200 can be used to execute the operating system, application programs, and software modules (including the software engines) described herein.
[0077] The computing device 1200 includes at least one processing device 1202, such as a central processing unit (CPU). In this example, the computing device 1200 also includes a system memory 1204, and a system bus 1206 that couples various system components including the system memory 1204 to the at least one processing device 1202. The system bus 1206 is one of any number of types of bus structures including a memory bus, or memory controller; a peripheral bus; and a local bus using any of a variety of bus architectures.
[0078] The system memory 1204 includes read only memory (ROM) 1208 and random-access memory (RAM) 1210. A basic input / output system 1212 containing the basic routines that act to transfer information within computing device 1200, such as during start up, is typically stored in the read only memory 1208. In some examples, the system memory 1204 has a large memory capacity, such as equal to or greater than one Terabyte of RAM. The RAM can be used to load and subsequently analyze the waveform data (e.g., the raw waveform data, such as stored in a raw waveform data file, which can include digitalized waveform data).
[0079] The computing device 1200 also includes a secondary storage device 1214 in some embodiments, such as a hard disk drive, for storing digital data. The secondarystorage device 1214 is connected to the system bus 1206 by a secondary storage interface 1216. In some examples, the secondary storage devices 1214 and their associated computer readable media provide nonvolatile storage of computer readable instructions (including application programs and program modules), data structures, and other data for the computing device 1200.100801 Although the exemplary environment described herein employs a hard disk drive as a secondary storage device, other types of computer readable storage media are used in other embodiments. Examples of these other types of computer readable storage media include magnetic cassettes, flash memory cards, digital video disks, Bernoulli cartridges, compact disc read only memories, digital versatile disk read only memories, random access memories, or read only memories. Some embodiments include non- transitory media. Additionally, such computer readable storage media can include local storage or cloud-based storage.
[0081] Several program modules can be stored in secondary storage device 1214 or the system memory 1204, including an operating system 1218, one or more application programs 1220, other program modules 1222 (e.g., software engines described herein), and program data 1224. The computing device 1200 can utilize any suitable operating system, such as Microsoft Windows™, Google Chrome™, Apple OS, and any other operating system suitable for a computing device.
[0082] In some examples, a user provides inputs to the computing device 1200 through one or more input devices 1226. Examples of input devices 1226 include a keyboard 1228, mouse 1230, microphone 1232, and touch sensor 1234 (such as a touchpad or touch sensitive display). Additional examples include additional types of input devices 1226, or fewer types of input devices 1226. The input devices 1226 are connected to the at least one processing device 1202 through an input / output interface 1236 coupled to the system bus 1206. The input / output interface 1236 can include any number of input / output interfaces, such as a parallel port, serial port, game port, or a universal serial bus. Wireless coupling between input devices 1226 and the input / output interface 1236 is possible as well, such as through infrared, BLUETOOTH®, 802.1 la / b / g / n, cellular, or other radio frequency communication systems in some possible embodiments.
[0083] In this example embodiment, a display device 1242, such as a monitor, liquid crystal display device, projector, or touch sensitive display device, is also connected to the system bus 1206 via a video adapter 1240. In addition to the display device 1242, thecomputing device 1200 can include various other peripheral devices (not shown), such as speakers or a printer.
[0084] When used in a local area networking environment or a wide area networking environment (such as the Internet), the computing device 1200 is typically connected to a network such as through a network interface 1238, such as an Ethernet interface. Other possible embodiments use other communication devices. For example, some embodiments of the computing device 1200 include a modem for communicating across the network.
[0085] The computing device 1200 typically includes at least some form of computer readable media. Computer readable media includes any available media that can be accessed by the computing device 1200. By way of example, computer readable media include computer readable storage media and computer readable communication media.
[0086] Computer readable storage media includes volatile and nonvolatile, removable, and non-removable media implemented in any device configured to store information such as computer readable instructions, data structures, program modules or other data. Computer readable storage media includes, but is not limited to, random access memory, read only memory, electrically erasable programmable read only memory, flash memory, compact disc read only memory, digital versatile disks or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device. Computer readable storage media does not include computer readable communication media.
[0087] Computer readable communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. The term “modulated data signal” refers to a signal that has one or more of its characteristics set or changed in such a maimer as to encode information in the signal. By way of example, computer readable communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency, infrared, and other wireless media. Combinations of any of the above are also included within the scope of computer readable media.
[0088] The computing device 1200 illustrated in FIG. 12 is also an example of programmable electronics, which may include one or more such computing devices, andwhen multiple computing devices are included, such computing devices can be coupled together with a suitable data communication network to collectively perform the various aspects disclosed herein.
[0089] The present disclosure includes the subject matter set forth in the following numbered clauses:
[0090] Clause 1. A method for analyzing particles within a flow cytometry system using a computing device, the method comprising: receiving flow cytometry data from the flow cytometry system, the flow cytometry' data including at least one event; assigning a probability7value to the at least one event, the probability' value indicating a probability that the at least one event is associated with a population of interest; displaying a probabilistic threshold to a user, wherein the probabilistic threshold categorizes an event as belonging to the population of interest; allowing the user to adjust the probabilistic threshold; and generating an output by: comparing the probability value of the at least one event to the probabilistic threshold; and assigning categorical information to at least one inclusive event, wherein the at least one inclusive event includes the probability value exceeding the probabilistic threshold.
[0091] Clause 2. The method of clause 1, further comprising identifying at least one non-inclusive event as not being within the population of interest when the probability' value does not exceed the probabilistic threshold.
[0092] Clause 3. The method of any one of clauses 1-2, further comprising filtering the at least one non-inclusive event from the at least one event.
[0093] Clause 4. The method of clause 2, wherein the at least one inclusive event and the at least one non-inclusive event are classified using a binary characterization.
[0094] Clause 5. The method of any one of clauses 1-4, wherein the probabilistic threshold includes a value between zero and one.
[0095] Clause 6. The method of any' one of clauses 1-5, w herein the at least one inclusive event is identified by applying one or more filters.
[0096] Clause 7. The method of clause 6, wherein the at least one inclusive event is identified through a series of filters including at least one of a bandpass filter, a low-pass filter, and a high-pass filter.
[0097] Clause 8. The method of clause 6, wherein the at least one inclusive event is categorized into a blood cell population through a Boolean logic binary characterization of events, wherein the at least one event is characterized as either the inclusive event or the non-inclusive event.
[0098] Clause 9. The method of any one of clauses 1-8, wherein the output is generated using a machine learning algorithm.
[0099] Clause 10. The method of clause 9, further comprising training the machine learning algorithm by a supervised approach by comparing the output to an independent data set to generate a confidence rating.
[0100] Clause 11. The method of clause 10, further comprising: adjusting output generated by the machine learning algorithm by changing the probabilistic threshold; and generating at least one adjusted output.
[0101] Clause 12. The method of clause 10, wherein the independent data set is manually adjusted by the user.
[0102] Clause 13. The method of clause 10, further comprising locking the machine learning algorithm for use when the confidence rating exceeds a threshold value.
[0103] Clause 14. The method of clause 1, further comprising: receiving one or more probability value events; and generating one or more additional outputs corresponding to one or more additional probabilistic thresholds.
[0104] Clause 15. The method of any one of clauses 1-14. further comprising isolating a desired quantity of events for further analysis by: generating a histogram including the at least one event, wherein the histogram corresponds to a quantity7of events having the probability value; identifying the desired quantity of events to analyze; and adjusting the probabilistic threshold to include the desired quantity of events while excluding other events.
[0105] Clause 16. The method of any one of clauses 1-15, further comprising displaying the at least one inclusive event to the user.
[0106] Clause 17. A method classifying events within a flow cytometry system, the method comprising: inspecting flow cytometry data, the flow cytometry data comprising at least one event, the at least one event including a probability value, the probability value indicating a probability that the at least one event is associated with a population of interest; allowing an adjustment of a probabilistic threshold on the flow cytometry system to filter one or more of the at least one event when the probability value is greater than a probabilistic threshold; and receiving an output from the flow cytometry system, the output including at least one inclusive event that exceeds the probabilistic threshold.
[0107] Clause 18. The method of clause 17, further comprising training the flow cytometry system, wherein the flow cytometry system includes a machine learningalgorithm, by: generating a confidence rating by comparing the output to an expected output; adjusting the probabilistic threshold and retraining the machine learning algorithm when the confidence rating is less than a threshold value; instructing the machine learning algorithm to provide at least one adjusted output; and inspecting the at least one adjusted output.
[0108] Clause 19. The method of clause 18, wherein the at least one adjusted output is inspected by: comparing the at least one adjusted output to the independent data set;
[0109] generating at least one adjusted confidence rating; and comparing the at least one adjusted confidence rating to the threshold value.
[0110] Clause 20. The method of clause 18, further comprising locking the machine learning algorithm when the confidence rating exceeds the threshold value.
[0111] Clause 21. A flow cytometry system for analyzing particles, the flow cytometry system comprising: at least one processing device; and a non-transitory computer readable storage media storing instructions which, when executed by the processing device, cause the at least one processing device to: receive flow cytometry data from the flow cytometry system, the flow cytometry data comprising at least one event; assign a probability value to the at least one event, the probability value indicating a probability that the at least one event is located within a population of interest; display the at least one probability value to a user; receive a selection of a probabilistic threshold from the user; and generate an output by: comparing the at least one probability’ value to the probabilistic threshold; and identifying at least one inclusive event when the at least one probability value exceeds the probabilistic threshold.
[0112] Clause 22. The flow’ cytometry system of clause 21, further comprising the flow cytometry system, the flow cytometry system including: a light source for generating a light beam toward an interrogation zone; and an optical system including detectors for detecting radiated light from particles passing through the light beam in the interrogation zone.
[0113] Clause 23. The flow cytometry system of clause 22, wherein the non- transitory computer readable storage media stores instructions which, when executed by the processing device, further cause the processing device to detect flow cytometry data, using the light source and the optical system, from particles passing through the interrogation zone.
[0114] Clause 24. The flow cytometry system of any one of clauses 21-23, further comprising one or more electrodes for measuring an electrical impedance of particles passing through an interrogation zone.
[0115] Clause 25. The flow cytome y7system of clause 24, wherein the non- transitory computer readable storage media stores instructions which, when executed by the processing device, further cause the processing device to detect flow cytometry data, using the one or more electrodes, from particles passing through the interrogation zone.
[0116] Clause 26. The flow cytometry' system of any one of clauses 21-25, wherein the non-transitory computer readable storage media stores instructions which, when executed by the processing device, further cause the processing device to display the at least one inclusive event to the user.
[0117] The above description is illustrative and is not restrictive. Many variations of the invention will become apparent to those skilled in the art upon review of the disclosure. The scope of the invention should, therefore, be determined not with reference to the above description, but instead should be determined with reference to the pending claims along with their full scope or equivalents.
[0118] One or more features from any embodiment may be combined with one or more features of any other embodiment w ithout departing from the scope of the invention.
Claims
WHAT IS CLAIMED IS:
1. A method for analyzing particles within a flow cytometry system using a computing device, the method comprising: receiving flow cytometry data from the flow cytometry system, the flow cytometry7data including at least one event; assigning a probability value to the at least one event, the probability value indicating a probability that the at least one event is associated with a population of interest; displaying a probabilistic threshold to a user, wherein the probabilistic threshold categorizes an event as belonging to the population of interest; allowing the user to adjust the probabilistic threshold; and generating an output by: comparing the probability value of the at least one event to the probabilistic threshold; and assigning categorical information to at least one inclusive event, wherein the at least one inclusive event includes the probability value exceeding the probabilistic threshold.
2. The method of claim 1, further comprising identifying at least one non-inclusive event as not being within the population of interest when the probability value does not exceed the probabilistic threshold.
3. The method of claim 2, further comprising filtering the at least one non- inclusive event from the at least one event.
4. The method of claim 2, wherein the at least one inclusive event and the at least one non-inclusive event are classified using a binary7characterization.
5. The method of claim 1, wherein the probabilistic threshold includes a value between zero and one.
6. The method of claim 1, wherein the at least one inclusive event is identified by applying one or more filters.
7. The method of claim 6, wherein the at least one inclusive event is identified through a series of filters including at least one of a bandpass filter, a low-pass filter, and a high-pass filter.
8. The method of claim 6, wherein the at least one inclusive event is categorized into a blood cell population through a Boolean logic binary characterization of events, wherein the at least one event is characterized as either the inclusive event or the non- inclusive event.
9. The method of claim 1, wherein the output is generated using a machine learning algorithm.
10. The method of claim 9, further comprising training the machine learning algorithm by a supervised approach by comparing the output to an independent data set to generate a confidence rating.
11. The method of claim 10, further comprising: adjusting output generated by the machine learning algorithm by changing the probabilistic threshold; and generating at least one adjusted output.
12. The method of claim 10, wherein the independent data set is manually adjusted by the user.
13. The method of claim 10, further comprising locking the machine learning algorithm for use when the confidence rating exceeds a threshold value.
14. The method of claim 1, further comprising: receiving one or more probability value events; and generating one or more additional outputs corresponding to one or more additional probabilistic thresholds.
15. The method of claim 1, further comprising isolating a desired quantity of events for further analysis by:generating a histogram including the at least one event, wherein the histogram corresponds to a quantity of events having the probability value; identifying the desired quantity of events to analyze; and adjusting the probabilistic threshold to include the desired quantity of events while excluding other events.