Particle type identification using flow cytometry
The method and system for flow cytometry enable label-free identification of particle types by analyzing waveform data with a trained model, addressing the limitations of costly fluorescent markers and improving accuracy and sample preservation.
Patent Information
- Application Number
- PCT/US2024/059005
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-12-06
- Publication Date
- 2025-06-19
AI Technical Summary
Current flow cytometry methods rely heavily on fluorescent markers for particle type identification, which are costly and can lead to cell loss and variability in results.
A method and system for recognizing particle types in a sample using flow cytometry that involves obtaining waveform data from particles, segmenting this data into constituent waveforms, and using a trained model to identify particle types without the need for labeling.
This approach reduces the cost and complexity of particle type identification, preserves the sample for further use, and potentially offers improved accuracy over traditional label-based methods.
Smart Images

Figure US2024059005_19062025_PF_FP_ABST
Abstract
Description
PARTICLE TYPE IDENTIFICATION USING FLOW CYTOMETRYCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application is being filed on December 6, 2024, as a PCT International application and claims the benefit of and priority' to U.S. Provisional Application No. 63 / 610,216, filed December 14, 2023; the disclosure of which is hereby incorporated by reference in its entirety’.BACKGROUND
[0002] Floyv cytometry is a technique for detecting and analyzing the chemical and physical characteristics of cells or particles in a fluid sample. For example, a flow cytometer may be used to assess cells from blood, bone marrow, tumors, or other body fluids. Typically, the sample is passed through a fluid nozzle which aligns particles in a single file line within a sheath fluid. A laser beam illuminates the particles as they pass through in single file to generate radiated light including forward scattered light, side scattered light, and fluorescent light. The radiated light can then be detected and analyzed to determine one or more charactenstics of the particles.SUMMARY
[0003] Examples presented herein relate to a method of recognizing one or more particle types in a sample using flow cytometry. The method includes obtaining a set of waveform data, the waveform data generated by interrogation of the sample to produce light signals from a plurality of particles within the sample, the plurality of particles including one or more particle types and segmenting the set of yvaveform data into a plurality of constituent particle waveforms. The method further includes submitting the constituent particle waveforms to a model and identifying, using the model, a particle type of the one or more particle types for the constituent particle waveforms.
[0004] In other examples presented herein, the method further includes generating a graphical output marking the constituent particle waveforms. In still other examples presented herein, one or more particles of the particle type are unlabeled. In further examples presented herein, each particle of the plurality of particles are unlabeled.
[0005] In yet other examples presented herein, the model is trained using a set of training yvaveform data, the set of training waveform data generated by interrogation of a stained particle type to produce light signals from the plurality of particles within the stained particle type. In further examples presented herein, the training waveform data including at least one of a forward scatter waveform, a side scatter yvaveform, anautofluorescence waveform, and a fluorescent waveform associated with the stain. In still further examples presented herein, the training waveform data including each of the forward scatter waveform, the side scatter waveform, the autofluorescence waveform, and the fluorescent waveform associated with the stain.
[0006] In other examples presented herein, the set of waveform data including at least one of a forward scatter waveform, a side scatter waveform, and an autofluorescence waveform. In further examples presented herein, the set of waveform data including each of the forward scatter waveform, the side scatter waveform, and the autofluorescence waveform. In yet other examples presented herein, the model is a binary7classifier. In still other examples presented herein, the model is a first model and the particle ty pe is a first particle type and the method further comprises identifying, using a second model, a second particle type of the one or more particle types for another constituent particle waveform of the plurality of constituted particle waveforms.
[0007] In other examples presented herein, the sample is an unprocessed sample. In further examples presented herein, the sample is one of a peripheral blood sample or a bone marrow sample. In still further examples presented herein, the particle type is at least one of erythrocytes, thrombocytes and leukocytes. In yet still further examples presented herein, the particle type is leukocytes and is further identified as one of lymphocytes, granulocytes, monocytes and blasts. In yet further examples presented herein, the sample remains viable for further processing and use after obtaining the waveform data. In other examples presented herein, the set of waveform data includes time data between each particle of the plurality' of particles in which the particle is not interrogated.
[0008] Other examples presented herein relate to a flow cytometry system for recognizing one or more particle types in a sample using flow cytometry. The system includes a laser configured to emit light toward an interrogation location to produce light signals from a plurality of particles in a sample directed through the interrogation location in a fluid stream, one or more detectors configured to convert the light signals to waveform data, and a processor in communication with a memory storing instructions. When executed by the processor, the instructions cause the flow cytometry system to: obtain a set of waveform data, the waveform data generated by interrogation of the sample to produce light signals from a plurality' of particles within the sample; submit the set of waveform data to a model; and identify, using the model, one or moreconstituent waveforms in the set of waveform data as being associated with a particle type.
[0009] Other examples presented herein relate to a system for recognizing one or more particle types in a sample using flow cytometry. The system includes a model trained using a set of training waveform data, the set of training waveform data generated by interrogation of a stained particle type to produce light signals from the plurality of particles within the stained particle type and a processor in communication with a memory storing instructions. When executed by the processor, the instructions cause the processor to: obtain a set of waveform data, the waveform data generated by interrogation of the sample to produce light signals from a plurality' of particles within the sample; submit the set of waveform data to the model; and identify, using the model, one or more constituent waveforms in the set of waveform data as being associated with a particle type, wherein the particle type comprises unlabeled particles of a same type as the stained particle type.
[0010] Other examples presented herein relate a method of training a model for recognition of one or more particle types in a sample using flow cytometry. The method includes obtaining a sample of a plurality of particles, the plurality of particles including a particle type, marking the particle type with a stain, and obtaining a set of w aveform data, the waveform data generated by interrogation of the sample to produce light signals from the plurality of particles within the sample. The method further includes submitting the set of waveform data to the model as training data, and identifying one or more constituent waveforms in the set of w aveform data as being associated with the particle type using one or more characteristics of the stain.
[0011] In other examples presented herein, the method further includes excluding one or more constituent w aveforms in the set of waveform data from being associated with the particle type using an absence of one or more characteristics of the stain. In still other examples presented herein, the set of waveform data includes at least one of a forward scatter waveform, a side scatter waveform, an autofluorescence waveform, and a fluorescent waveform associated with the stain. In further examples presented herein, the set of waveform data includes each of the forward scatter waveform, the side scatter waveform, the autofluorescence waveform, and a fluorescent w aveform associated with the stain.
[0012] In yet other examples presented herein, the sample is one of a peripheral blood sample or a bone marrow sample. In further examples presented herein, the particletype is at least one of erythrocytes, thrombocytes and leukocytes. In still further examples presented herein, the particle type is leukocytes and is further identified as one of lymphocytes, granulocytes, monocytes and blasts. In still other examples presented herein, the stain is a CD- 19 stain.
[0013] A variety' of additional inventive aspects will be set forth in the description that follows. The inventive aspects can relate to individual features and to combinations of features. It is to be understood that both the forgoing general description and the following detailed description are exemplary- and explanatory only and are not restrictive of the broad inventive concepts upon yvhich the embodiments disclosed herein are based.BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings, which are incorporated in and constitute a part of the description, illustrate several aspects of the present disclosure. A brief description of the drawings is as follows:
[0015] FIG. 1 is a schematic block diagram illustrating an example of a flow cytometer system.
[0016] FIG. 2A shows a particle entering a laser beam.
[0017] FIG. 2B shows the particle passing through a center area of the laser beam.
[0018] FIG. 2C shows the particle exiting the laser beam.
[0019] FIG. 3 is a block diagram of a waveform analysis device in accordance with some embodiments.
[0020] FIG. 4 is a floyvchart of an example method of recognizing one or more particle ty pes in a sample using flow cytometry.
[0021] FIG. 5 is a flowchart of an example method of training a model for recognition of one or more particle types in a sample using flow cytometry.
[0022] FIG. 6 illustrates an exemplary architecture of a computing device that can be used to implement aspects of the flow cytometry system of FIG. 1.DETAILED DESCRIPTION
[0023] The figures and the following description illustrate specific example embodiments of the disclosure. It will thus be appreciated that those skilled in the art will be able to devise various arrangements that, although not explicitly described or shown herein, embody the principles of the disclosure and are included within the scope of the disclosure. Furthermore, any examples described herein are intended to aid inunderstanding the principles of the disclosure and are to be construed as being without limitation to such specifically recited examples and conditions. As a result, the disclosure is not limited to the specific embodiments or examples described below, but by the claims and their equivalents.
[0024] FIG. 1 is a schematic block diagram illustrating an example of a flow cytometer system 100. In general, flow cytometry’ is a technique for measuring and analyzing the physical and chemical properties of a sample of particles or cells. Data from millions of cells can be collected in a matter of minutes and displayed in a variety of formats for researchers or clinicians. Some example applications include phenotyping to identify and count specific cell types within a population, analyzing DNA or RNA content, determining the presence of antigens on the surface or within cells, and assessing cell health status.
[0025] Flow cytometry can be used to analyze and sort cells based on their physical and chemical properties. It is commonly employed in various fields, including immunology, oncology, hematology, and microbiology. Flow cytometers measure characteristics of individual cells as they pass through a fluidic system. As will be discussed in further detail below, several features of the flow cytometer contribute to differentiating between different cell and particle ty pes, including fluorescent labels, cell sorting, use of both forward scatter (FSC) and side scatter (SSC), multi-parameter analysis, and functional and quantitative assays.
[0026] For example, one implementation of flow cytometry- is identification of particle subsets within a sample, such as types of blood cells within a blood sample. Using a hematology analyzer, the three major lineages of white blood cells — erythrocytes, thrombocytes, and leukocytes — can be distinguished, but further distinction into subsets is limited. These lineages can be assessed separately with analysis performed by flow cytometry-. To further identify leukocyte subsets, erythrocytes are ty pically first eliminated by lysis. Alternatively, nucleated cells may be stained using a DNA dye. Fluorochrome-coupled antibodies may be used to identify additional subsets.
[0027] The flow cytometer system 100 generally includes three main component subsystems: a fluidic system 110, an optical system 120, and an electronic system 130. The fluidic system 110 includes a nozzle 112 which receives a sample containing particles or cells suspended in a fluid. The nozzle 112 creates and ejects a fluid streambeams of light produced by a laser 102. The point at which a particle intersects with a light beam is known as an interrogation location 116.
[0028] The optical system 120 includes the laser 102, optical elements 122, and detectors 124. At the interrogation location 1 16, light from the laser 102 hits a particle and scatters. The optical elements 122 direct the scattered light toward the detectors 124. The detectors 124 may include a forward scatter (FSC) detector to measure scatter along the path of the laser 102, a side scatter (SSC) detector to measure scatter at a ninetydegree angle relative to the laser 102, and / or one or more fluorescence (FL1, FL2, and FL3) detectors to measure the emitted fluorescence intensity7of different wavelengths of light.
[0029] FSC provides information about cell size, while SSC gives insights into cell granularity or internal complexity. By analyzing FSC and SSC, different cell types can be distinguished based on their size and internal structure. Generally, FSC intensity is proportional to the size or diameter of the particle due to light diffraction around the particle. FSC may therefore be used for the discrimination of particles by size. SSC, on the other hand, is produced from light refracted or reflected by internal structures of the particle and may therefore provide information about the internal complexity or granularity of the particle.
[0030] Different fluorochromes can be used to label different cell types or cellular markers. For example, antibodies conjugated with fluorochromes can specifically bind to cell surface markers, allowing for the identification of particular cell populations. By adding fluorescent labelling to a sample, different fluorescent signals / channels (e.g., green, orange, and red) can be analyzed for functional characteristics of a cell. For example, since T-cells present CD3 binding sites, a sample containing T-cells may be "stained" with anti-CD3 antibodies conjugated with a fluorescent molecule. As these cells pass through the interrogation location 116, the laser light excites the fluorescent tag, or fluorochrome, to emit photons at a wavelength detectable by a fluorescence detector. The detectors 124 may therefore simultaneously measure a number of parameters and enable categorization of particles by their function based on detected wavelengths of light.
[0031] The number of detectors 124, and in particular the number of fluorescence detectors (FL1, etc.), may be determined based on the particular parameters of a given experiment. For example, polychromatic flow cytometry, which is used to analyze and sort multiple characteristics of individual cells simultaneously, involves the use ofmultiple fluorescent markers. Each marker has a distinct emission spectrum and labels specific molecules or cellular components within a cell. By measuring the emitted light at different wavelengths, researchers can gather a wealth of information about a cell's properties, such as surface markers, intracellular proteins, and DNA content. This technique is particularly valuable for studying complex cell populations, including immune cells, and is commonly used in immunology, hematology, and cancer research to characterize and sort cells based on their unique molecular profiles.
[0032] However, there are limitations to the relying on fluorescent markers to distinguish cell and particle types. The staining reagents used are expensive and the processing steps required can lead to cell loss, impacting the sensitivity of the detection, particularly for rare cells. Also, these processing steps are challenging to standardize. Lot-to-lot comparisons for antibodies alone are time and resource intensive. Further variability is introduced by fluorescence spillover correction (compensation), which is required to unequivocally identify signals from different markers and thereby assign cell population identities. As is disclosed herein, these and other limitations of label-based subset identification are solved by methods and system capable of label-free subset identification.
[0033] The electronic system 130 includes a waveform acquisition device 140 and a waveform analysis device 150. The waveform acquisition device 140 is communicatively coupled with the detectors 124 and is configured to receive analog waveform data 126 generated by the detectors 124. The waveform acquisition device 140 includes an analog-to-digital converter (ADC) 142 configured to digitize the waveform data. The waveform analysis device 150 is configured to receive the digital waveform data and display it for a user of the flow cytometer system 100. In some embodiments, the waveform analysis device 150 comprises a computing device communicatively coupled with a flow cytometer 101 over a network, and the flow cytometer 101 may include the fluidic system 110, optical system 120, and waveform acquisition device 140. In other embodiments, the waveform analysis device 150 is integrated with the flow cytometer 101.
[0034] A graphics processing unit (GPU) 152 is a component of the waveform analysis device 150. The GPU 152 is configured to process a continuous digital stream generated by the waveform acquisition device 140 and provided to the waveform analysis device 150. The digital stream is continuous in that the waveform acquisition device 140 does not threshold the waveform data produced by the detectors 124. During anexperiment, the waveform acquisition device 140 continuously digitizes the analog waveform data 126 at a high rate (e.g., 1 GHz) without thresholding. A field programmable gate array (FPGA) may also be included or excluded from the waveform acquisition device 140.
[0035] The waveform analysis device 150 may thus receive a digitized version of the waveform data with increased data points, and the waveform data for an experiment is unthresholded and available in its entirety for processing by the GPU 152. In addition to having the capability of processing a large stream or file of waveform data, the GPU 152 enables thresholding the waveform at the post-processing step as opposed to the w aveform acquisition step. This in turn provides several technical benefits including the ability to dynamically adjust thresholds and update graphical plots in real-time without re-running an experiment. The GPU 152 may also measure and extract biologically relevant information present in the wav eform data beyond the three parameters of height, width, and area traditionally analyzed. Further details of operation and advantages are discussed below.
[0036] The flow7cytometer system 100 of FIG. 1 includes elements which are show n and described for purposes of discussion and it will be appreciated that numerous variations in components and functions are possible. The optical elements 122 may include a series of filters, dichroic mirrors, and / or beam splitters to select out different wavelengths of light and provide the wavelength to the appropriate detector 124. The detectors 124 may comprise, for example, photomultiplier tubes (PMTs) or avalanche photodiodes (APDs).
[0037] FIGS. 2A-2C illustrate waveform data generated by the detection of a particle 201 passing through a laser beam 202. As the particle 201 passes through the interrogation point of the light source, a pulse is generated in one or more of the detectors 124. FIG. 2A shows the particle 201 entering the laser beam 202. As the particle 201 starts to intersect with the laser beam 202 it begins to generate scattered light and fluorescence signals. The detector 124 produces a current or voltage that is proportional to the number of photons that hit the photocathode. As such, the output of the detector 124 begins to rise as shown in plot 212 due to current flowing in the detector 124.
[0038] FIG. 2B shows the particle 201 passing through a center area of the laser beam 202. As the particle 201 continues to move downward into a center of the laser beam 202 the particle 201 is fully illuminated. Since photon density of the laser beam202 is highest in the center a maximum amount of optical signal is produced. The current or voltage of the detector 124 therefore peaks as shown in plot 232.
[0039] FIG. 2C shows the particle 201 exiting the laser beam 202. As the particle flows out of the laser beam 202 the current or voltage output of the detector 124 returns to baseline, as shown in plot 252. This generation of a pulse is called an event. The height is the maximum current / voltage output by the detector 124, the width is the time interval during which the pulse occurs, and the area is the integral of the pulse. Generally speaking, the height and area correspond with signal intensity and the width corresponds with the time the particle is illuminated by the laser beam 202. Accordingly, as pulses are generated they may be quantified by height, width, and area. This information may be used to distinguish between particles and fluorescence signals may be displayed on plots, analyzed, and interpreted.
[0040] The waveform of signals elicited by laser light scattered in the direction of the laser beam (FSC), at a 90 degree angle to the laser (SSC), and generated by the autofluorescence of cells can be detected. The additional information contained within the full waveform, compared to the height or area measurements of traditional flow cytometry data collection, enables subsets among the particles analyzed to be distinguished. In particular, cellular subsets that cannot be differentiated based on area and height measurements only can now be sorted. This will allow the distinguishing among, for example, eosinophils, neutrophils and basophils within the granulocytic lineage.
[0041] As will be discussed in further detail below, a machine learning model is trained to identify particle subsets based on characteristic waveforms. For example, using machine learning, the waveform data collected will enable identification of different lymphocyte subsets based on the waveform only (without labeling). In embodiments, a machine learning model is first trained to classify relevant particle and / or cellular subsets based on their waveform after using labeled training data, such as antibody labeled training data. The trained model can then be applied to a collection of particles or cells to be classified.
[0042] For example, B-cells may be stained using an anti-CDl 9 reagent. The model is trained to recognize the population staining positive for CD19 based on the waveform of FSC, SSC and autofluorescence signals. The trained model will then be able to identify B-cells in unstained samples based on their waveform only.
[0043] FIG. 3 is a block diagram of a waveform analysis device 300 in accordance with some embodiments. The waveform analysis device 300 is configured to receive, store, and display waveform data that has been continuously sampled. The waveform analysis device 300 may include an interface 310 to receive digitized raw waveform data 332, persistent storage 330 to store the digitized raw waveform data 332, and a graphical user interface (GUI) 320 to display the digitized raw waveform data 332. The persistent storage 330 may also include subset characteristics model 334. Waveform analysis device 300 may further include waveform segmenter 340, flow cytometry analysis application 350, and a graphics processing unit (GPU) 152. The persistent storage 330 may comprise system memory such as random access memory (RAM) and / or long term non-volatile memory such as a hard drive.
[0044] Subset characteristics model 334 embodies of set of characteristics defining one or more subsets based on their waveforms. In embodiments, subset characteristics model 334 is one or more classifiers, such as binary classifiers. For example, each subset may be defined by an independent classifier and a series of binary classifiers may be used to characterize multiple subsets within a sample. In some cases, a multi-dimensional classifier may be used to identify subsets characterized by multiple markers, e.g., expression patterns or, in cases related to characterizing blood cells, immunophenotypes. In embodiments, subset characteristics model 334 is configured to provide one or more of a probability or a density output.
[0045] Subset characteristics model 334 is trained to infer a particle subset associated with characteristic waveform patterns based on being trained using labeled waveform patterns for the particle subset. In embodiments, subset characteristics model 334 is trained using a set of training waveform data. In embodiments, the set of training waveform data is generated by interrogation of a stained particle type to produce light signals from the plurality of particles within the stained particle type. In some cases, the set of training waveform data includes both positive and negative examples.
[0046] For example, a binary classifier may be trained to identify B-cells in a blood sample. For the training data set, B-cells are labeled, such as with a CD-19 marker, and set as a positive example. Other cells may be set as negative examples. The CD- 19 marker triggers on a dedicated fluorescence channel indicating a positive example. Data from other channels at the time the CD- 19 marker triggers, associated with the particle and not the marker, such as FSC, SSC, and autofluorescence, is characterized as associated with a B-cell. Data from the time the marker triggers, as well as data apredetermined amount of time forward and / or back in some cases, is taken and characterized as a B-cell characteristic waveform. In embodiments, the trigger from the marker is used as a center point, and a waveform from around that center point is characterized as the target particle. This example focuses on identifying a subset using a CD- 19 marker, but those of skill in the art will understand the principles discussed are readily applicable to other markers and cell and / or particle subsets. Some non-limiting examples of markers includes CD-4, CD-8, TBNK panels, etc.
[0047] In another example, nanoparticles may be labeled and the model trained to distinguish the nanoparticles from background and / or noise present in the data. In embodiments, training w aveform data may be unlabeled and other criteria may be used to characterize a target subset.
[0048] Using subset characteristics model 334 to identify unlabeled particle types provides a number of advantages over subset identification using labeling. Advantages include reducing the cost of reagents and labor. Particle labels and associated preparation reagents can be quite expensive, as well as the time required by an experienced technician to prepare and run the labeled sample. Another advantage is that by enabling subset identification without requiring introducing a labeling reagent is that the original sample, with the identified subset, can be preserved for further processing and use. For example, in the case of analyzing a blood sample, the blood sample may be reintroduced or otherwise reused due to being preserved in its original state throughout the analysis. Unlabeled particle type identification using machine learning may also provide improved accuracy over the present label-based methods of identification. For example, traditional label-based methods rely on the application of gating and compensation for spillover to evaluate the presence of the subset. The unlabeled systems and methods for evaluating the presence of a subset, as disclosed herein, directly provide an output of the subset, without requiring additional analysis which introduces further opportunities for error.
[0049] Waveform segmenter 340 divides a continuous waveform into distinct segments or sections based on certain criteria. Segmentation isolates specific features or events within the waveform for further analysis or processing. Waveform segmenter may divide individual waveforms of the digitized raw waveform data 332 into segments associated with individual particles or may detect anomalies in the data. Segmentation can be employed to isolate sections of a waveform containing desired signals while excluding noisy segments.
[0050] Methods for waveform segmentation vary' depending on the specific application and characteristics of the waveform. Common techniques include thresholding, peak detection, template matching, and machine learning algorithms. The choice of method depends on factors such as the desired features to be extracted and the noise characteristics of the waveform. In embodiments, waveform segmenter 340 segments the waves in response to analysis based on subset characteristics model 334. In some cases, waveforms are initially segmented into individual particle waveforms and then evaluated for subset characteristics. Waveform segmentation can be a preprocessing step for training machine learning models. By segmenting waveforms into meaningful sections, labeled datasets can be created for supervised learning tasks or peaks for inferencing analysis.
[0051] The waveform analysis device 300 may further include a cytometry analysis application 350 comprising a software application or a set of related software applications configured to instruct the GPU 152 to process the digitized raw waveform data 332. The cytometry analysis application 350 may execute on one or more processors (not shown) to provide other functions described herein in conjunction with the GPU 152 such as receiving user input via the GUI 320. One or more components of the waveform analysis device 300 may reside in a cloud computing application in a network distributed system. In that regard, the waveform analysis device 300 may be any of a variety of computing devices, including, but not limited to, a personal computing device, a server computing device, or a distributed computing device.
[0052] In embodiments, cytometry analysis application 350 submits a set of waveform data to subset characteristics model 334 and identifies, using subset characteristics model 334. one or more constituent waveforms in the set of waveform data as being associated with a particle type. In some cases, the particle type includes unlabeled particles of a same type as the stained particle type.
[0053] In an example case, a user may submit a sample to identify progenitor cells. A subset characteristics model trained to identify viable stem cells is used and identifies a quantity or sample percent of viable stem cells present. The subset characteristics model and / or the cytometry analysis application may contain automatic parameters for the analysis, or may accept user defined parameters. For example, in this case, it may be important to underestimate a number of viable stem cells present, to ensure a value that can be relied upon for further analysis and / or use of the sample.
[0054] FIG. 4 is a flowchart of a method 400 of recognizing one or more particle types in a sample using flow cytometry. In embodiments, the sample is an unprocessed sample. The sample may remain viable for further processing and use after obtaining the waveform data. In some cases, the sample is one of a peripheral blood sample or a bone marrow sample. For example, the particle type may be at least one of erythrocytes, thrombocytes and leukocytes. In some cases, the particle type is leukocytes and is further identified as one of lymphocytes, granulocytes, monocytes and blasts.
[0055] In embodiments, method 400 is executed by waveform analysis device 300, such as by cytometry' analysis application 350 using subset characteristics model 334. In embodiments, method 400 is executed by an independent device which receives processed waveform data from the flow cytometer or a downstream data processing system.
[0056] At operation 402, a set of waveform data is obtained. The waveform data may be generated by interrogation of a sample to produce light signals from a plurality of particles within the sample. In embodiments, the plurality of particles includes one or more particle types. In embodiments, the set of waveform data includes at least one of a forward scatter waveform, a side scatter waveform, and an autofluorescence waveform. In some cases, the set of waveform data including each of a forward scatter waveform, a side scatter waveform, and an autofluorescence waveform. In some cases, the set of waveform data includes time data between each particle of the plurality of particles in which the particles are not interrogated.
[0057] At operation 404, the set of waveform data is segmented into a plurality of constituent particle waveforms. Waveform segmentation refers to the process of dividing a continuous waveform into distinct segments or sections based on certain criteria to isolate specific features or events within the waveform for further analysis or processing. Common techniques include thresholding, peak detection, template matching, and machine learning algorithms.
[0058] At operation 406. the constituent particle waveforms are submitted to a model. The model is a trained model, having been trained using a set of training waveform data. The set of training waveform data may be generated by interrogation of a stained particle type to produce light signals from a plurality of particles of the stained particle type. In embodiments, the model is a binary classifier. In embodiments, the training waveform data includes at least one of a forward scatter waveform, a side scatter waveform, an autofluorescence waveform, and a fluorescent waveform associated withthe stain. In some cases, the training waveform data including each of forw ard scatter waveform, a side scatter waveform, an autofluorescence waveform, and a fluorescent waveform associated with the stain.
[0059] For example, a fluorescent waveform associated with a stain or other marker is used to indicate a positive example associated with a waveform. The marker indicator may be used as a center point, and waveform data from FSC, SSC, and autofluorescence channels, taken a predetermined amount ahead and / or behind the center point is used to characterize the target subset particle as a waveform. In embodiments, the predetermined amount is determined according to segmentation of the waveform. For example, a marker indication triggers and is used as a center point. Data from either side of the center point is used collectively to represent a complete particle waveform as the training waveform indicating a target subset particle.
[0060] At operation 408, the model is used to identify a particle type of the one or more particle types for the constituent particle waveforms. In embodiments, one or more particles of the particle type are unlabeled in the sample. In some cases, each particle of the plurality of particles in the sample are unlabeled. In embodiments, one or more unlabeled particle types may be identified by independent models. For example, a first model is trained to identify7a first particle type and a second model identifies a second particle type.
[0061] At operation 410. a graphical output is generated marking or otherwise identifying the subset and / or the constituent particle waveforms. The graphical output may provide an overview of the model's predictions on a given dataset. Examples of possible outputs include a confusion matrix, which visually depicts true positives, true negatives, false positives, and false negatives, offering a snapshot of classification accuracy; a Receiver Operating Characteristic (ROC) curve, which illustrates the tradeoff between sensitivity and specificity; a Precision-Recall curve, which provides insights into precision and recall dynamics at different classification thresholds; a decision boundary visualization, which showcases how the classifier demarcates between classes in the feature space; and a class probability distribution, which presents confidence levels associated with each prediction.
[0062] FIG. 5 is a flowchart of an example method 500 of training a model for recognition of one or more particle types in a sample using flow cytometry. In embodiments, method 500 is executed by waveform analysis device 300, such as by cytometry analysis application 350 using subset characteristics model 334. Inembodiments, method 500 is executed by an independent device which received processed waveform data from the flow cytometer or a downstream data processing system.
[0063] At operation 502, a sample of a plurality of particles is obtained. Each particle of the plurality of particles includes a particle ty pe. For example, the sample may be one of a peripheral blood sample or a bone marrow sample. In this example, the particle type may be at least one of erythrocytes, thrombocytes and leukocytes. Further to this example, the particle type may be leukocytes and is further identified as one of lymphocytes, granulocytes, monocytes and blasts. In another example, nanoparticles may be distinguishes from other particle ty pes and / or sample noise.
[0064] At operation 504, at least one particle type is marked. In embodiments, the particle type is marked with a stain. For example, in the case where in the sample is a blood sample, the stain may be a CD-19 stain marking B-cells. Stains or antibody markers may be preferred label or marker used in some cases but other labels or markers are contemplated.
[0065] At operation 506. a set of waveform data is obtained. In embodiments, the waveform data is generated by interrogation of the sample to produce light signals from the plurality7of particles within the sample. In embodiments, the set of waveform data includes at least one of forward scatter waveform, a side scatter waveform, an autofluorescence waveform, and a fluorescent waveform associated with the stain. In some cases, the set of waveform data includes each of forw ard scatter waveform, a side scatter w aveform, an autofluorescence waveform, and a fluorescent waveform associated with the stain.
[0066] At operation 508. the set of waveform data is submitted to the model as training data. During training, the model learns to discern patterns within the provided training data, aiming to distinguish between at least two distinct classes (e.g., a target particle and not the target particle). The process involves adjusting internal parameters iteratively to minimize the disparity between its predictions and the actual labels of the training instances. As the model refines its understanding of the data, in embodiments, it fine-tunes a decision boundary in a feature space, improving the model’s ability to categorize new7, unseen examples. The iterative adjustments may be guided by a predefined objective function, including techniques, such as gradient descent, to optimize the model's parameters. Through this learning process, the model learns to generalize itsacquired knowledge to effectively classify instances beyond the training set, developing its predictive capabilities.
[0067] At operation 510, one or more constituent waveforms in the set of waveform data are identified as being associated with the particle type using one or more characteristics of the stain. In embodiments, identifying waveforms associated with the marked particle type includes excluding one or more constituent waveforms in the set of waveform data from being associated with the particle type using an absence of one or more characteristics of the stain.
[0068] FIG. 6 illustrates an exemplary architecture of a computing device that can be used to implement aspects of the present disclosure, including the waveform analysis device 150 / 300. The computing device illustrated in FIG. 6 can be used to execute the operating system, application programs, and software modules (including the software engines) described herein.
[0069] The computing device 900 includes, in some embodiments, at least one processing device 902, such as a central processing unit (CPU). A variety of processing devices are available from a variety of manufacturers, for example. Intel or Advanced Micro Devices. In this example, the computing device 900 also includes a system memory 906, and a system bus 904 that couples various system components including the system memory 906 to the processing device 902. The system bus 904 is one of any number of types of bus structures including a memory bus, or memory controller; a peripheral bus; and a local bus using any of a variety of bus architectures.
[0070] Examples of computing devices suitable for the computing device 900 include a server computer, a desktop computer, a laptop computer, a tablet computer, a mobile computing device (such as a smart phone, an iPod® or iPad® mobile digital device, or other mobile devices), or other devices configured to process digital instructions.
[0071] The system memory 906 includes read only memory7908 and random access memory (RAM) 910. A basic input / output system 912 containing the basic routines that act to transfer information within computing device 900. such as during start up, is typically stored in the read only memory 908. In some embodiments the waveform analysis device 150 (FIG. 1) has a large memory' capacity7, such as equal to or greater than one Terabyte of RAM. The RAM can be used by the GPU 152 for loading and subsequently analyzing the waveform data (e.g., the raw waveform data, such as stored in a raw waveform data file, which can include digitalized waveform data).
[0072] The computing device 900 also includes a secondary storage device 914 in some embodiments, such as a hard disk drive, for storing digital data. The secondary storage device 914 is connected to the system bus 904 by a secondary storage interface 916. The secondary storage devices 914 and their associated computer readable media provide nonvolatile storage of computer readable instructions (including application programs and program modules), data structures, and other data for the computing device 900.
[0073] Although the exemplary environment described herein employs a hard disk drive as a secondary storage device, other types of computer readable storage media are used in other embodiments. Examples of these other types of computer readable storage media include magnetic cassettes, flash memory cards, digital video disks, Bernoulli cartridges, compact disc read only memories, digital versatile disk read only memories, random access memories, or read only memories. Some embodiments include non- transitory media. Additionally, such computer readable storage media can include local storage or cloud-based storage.
[0074] A number of program modules can be stored in secondary storage device 914 or memory 906, including an operating system 918, one or more application programs 920, other program modules 922 (such as the software engines described herein), and program data 924. The computing device 900 can utilize any suitable operating system, such as Microsoft Windows™, Google Chrome™. Apple OS. and any other operating system suitable for a computing device.
[0075] In some embodiments, a user provides inputs to the computing device 900 through one or more input devices 926. Examples of input devices 926 include a keyboard 928, mouse 930, microphone 932, and touch sensor 934 (such as a touchpad or touch sensitive display). Other embodiments include other input devices 926. The input devices 926 are often connected to the processing device 902 through an input / output interface 936 that is coupled to the system bus 904. These input devices 926 can be connected by any number of input / output interfaces, such as a parallel port, serial port, game port, or a universal serial bus. Wireless communication between input devices and the interface 936 is possible as well, and includes infrared, BLUETOOTH® wireless technology7, 802.11a / b / g / n, cellular, or other radio frequency communication systems in some possible embodiments.
[0076] In this example embodiment, a display device 938, such as a monitor, liquid crystal display device, projector, or touch sensitive display device, is also connected tothe system bus 904 via an interface, such as a video adapter 940. In addition to the display device 938, the computing device 900 can include various other peripheral devices (not shown), such as speakers or a printer.
[0077] When used in a 1 ocal area networking environment or a wide area networking environment (such as the Internet), the computing device 900 is typically connected to a network through a network interface 942. such as an Ethernet interface. Other possible embodiments use other communication devices. For example, some embodiments of the computing device 900 include a modem for communicating across the network.
[0078] The computing device 900 typically includes at least some form of computer readable media. Computer readable media includes any available media that can be accessed by the computing device 900. By way of example, computer readable media include computer readable storage media and computer readable communication media.
[0079] Computer readable storage media includes volatile and nonvolatile, removable and non-removable media implemented in any device configured to store information such as computer readable instructions, data structures, program modules or other data. Computer readable storage media includes, but is not limited to, random access memory, read only memory, electrically erasable programmable read only memory, flash memory or other memory technology7, compact disc read only memory7, digital versatile disks or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 900. Computer readable storage media does not include computer readable communication media.
[0080] Computer readable communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery7media. The term “modulated data signal” refers to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, computer readable communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency, infrared, and other wireless media. Combinations of any of the above are also included within the scope of computer readable media.
[0081] The computing device illustrated in FIG. 6 is also an example of programmable electronics, which may include one or more such computing devices, andwhen multiple computing devices are included, such computing devices can be coupled together with a suitable data communication network so as to collectively perform the various functions, methods, or operations disclosed herein.
[0082] Illustrative examples of the systems and methods described herein are provided below. An embodiment of the system or method described herein may include any one or more, and any combination of, the clauses described below.
[0083] Clause 1. A method of recognizing one or more particle types in a sample using flow cytometry. The method includes obtaining a set of waveform data, the waveform data generated by interrogation of the sample to produce light signals from a plurality of particles within the sample, the plurality of particles including one or more particle types; segmenting the set of waveform data into a plurality of constituent particle waveforms; submitting the constituent particle waveforms to a model; and identifying, using the model, a particle type of the one or more particle types for the constituent particle waveforms.
[0084] Clause 2. The method of clause 1, further including generating a graphical output marking the constituent particle waveforms.
[0085] Clause 3. The method of clause 1 or 2, wherein one or more particles of the particle type are unlabeled.
[0086] Clause 4. The method of clause 3, wherein each particle of the plurality of particles are unlabeled.
[0087] Clause 5. The method of any of clauses 1 -4, wherein the model is trained using a set of training waveform data, the set of training waveform data generated by interrogation of a stained particle type to produce light signals from the plurality of particles within the stained particle type.
[0088] Clause 6. The method of clause 5, wherein the training waveform data including at least one of a forward scatter waveform, a side scatter waveform, an autofluorescence waveform, and a fluorescent waveform associated with the stain.
[0089] Clause 7. The method of clause 6, wherein the training waveform data including each of the forward scatter waveform, the side scatter waveform, the autofluorescence waveform, and the fluorescent waveform associated with the stain.
[0090] Clause 8. The method of any of clauses 1-7, wherein the set of waveform data including at least one of a forward scatter waveform, a side scatter waveform, and an autofluorescence waveform.
[0091] Clause 9. The method of clause 8, wherein the set of waveform data including each of the forward scatter waveform, the side scatter waveform, and the autofluorescence waveform.
[0092] Clause 10. The method of any of clauses 1-9, wherein the model is a binary classifier.
[0093] Clause 11. The method of any of clauses 1-10. wherein the model is a first model and the particle type is a first particle type and the method further comprises identifying, using a second model, a second particle type of the one or more particle types for another constituent particle waveform of the plurality7of constituted particle waveforms.
[0094] Clause 12. The method of any of clauses 1-11. wherein the sample is an unprocessed sample.
[0095] Clause 13. The method of clause 12, wherein the sample is one of a peripheral blood sample or a bone marrow sample.
[0096] Clause 14. The method of clause 13, wherein the particle ty pe is at least one of erythrocytes, thrombocytes and leukocytes.
[0097] Clause 15. The method of clause 14, wherein the particle type is leukocytes and is further identified as one of lymphocytes, granulocytes, monocytes and blasts.
[0098] Clause 16. The method of any of clauses 13-15. wherein the sample remains viable for further processing and use after obtaining the waveform data.
[0099] Clause 17. The method of any of clauses 1-1 , wherein the set of waveform data includes time data between each particle of the plurality7of particles in which the particle is not interrogated.
[0100] Clause 18. A flow cytometry system for recognizing one or more particle types in a sample using flow cytometry. The system includes a laser configured to emit light tow ard an interrogation location to produce light signals from a plurality of particles in a sample directed through the interrogation location in a fluid stream; one or more detectors configured to convert the light signals to waveform data; and a processor in communication with a memory. The memory storing instructions which, when executed by the processor, cause the flow cytometry system to: obtain a set of waveform data, the waveform data generated by interrogation of the sample to produce light signals from a plurality of particles within the sample; submit the set of waveform data to a model; and identify, using the model, one or more constituent waveforms in the set of waveform data as being associated with a particle ty pe.
[0101] Clause 19. A system for recognizing one or more particle types in a sample using flow cytometry. The system includes a model trained using a set of training waveform data, the set of training waveform data generated by interrogation of a stained particle type to produce light signals from the plurality of particles within the stained particle ty pe; and a processor in communication with a memory'. The memory storing instructions which, when executed by the processor, cause the processor to: obtain a set of waveform data, the waveform data generated by interrogation of the sample to produce light signals from a plurality7of particles within the sample; submit the set of waveform data to the model; and identify, using the model, one or more constituent waveforms in the set of waveform data as being associated with a particle ty pe, wherein the particle type comprises unlabeled particles of a same type as the stained particle type.
[0102] Clause 20. A method of training a model for recognition of one or more particle types in a sample using flow cytometry. The method including obtaining a sample of a plurality7of particles, the plurality7of particles including a particle ty pe; marking the particle type with a stain; obtaining a set of waveform data, the waveform data generated by interrogation of the sample to produce light signals from the plurality of particles within the sample; submitting the set of waveform data to the model as training data; and identifying one or more constituent waveforms in the set of waveform data as being associated with the particle type using one or more characteristics of the stain.
[0103] Clause 21 . The method of clause 20, further including excluding one or more constituent waveforms in the set of waveform data from being associated with the particle ty pe using an absence of one or more characteristics of the stain.
[0104] Clause 22. The method of clause 20 or 21, wherein the set of waveform data includes at least one of a forward scatter waveform, a side scatter waveform, an autofluorescence waveform, and a fluorescent waveform associated with the stain.
[0105] Clause 23. The method of clause 22, wherein the set of waveform data includes each of the forward scatter waveform, the side scatter waveform, the autofluorescence waveform, and a fluorescent waveform associated with the stain.
[0106] Clause 24. The method of any of clauses 20-23, wherein the sample is one of a peripheral blood sample or a bone marrow sample.
[0107] Clause 25. The method of clause 24, wherein the particle ty pe is at least one of erythrocytes, thrombocytes and leukocytes.
[0108] Clause 26. The method of clause 25, wherein the particle ty pe is leukocytes and is further identified as one of lymphocytes, granulocytes, monocytes and blasts.
[0109] Clause 27. The method of any of clauses 20-26, wherein the stain is a CD- 19 stain.
[0110] Having described the preferred aspects and implementations of the present disclosure, modifications and equivalents of the disclosed concepts may readily occur to one skilled in the art. However, it is intended that such modifications and equivalents be included within the scope of the claims which are appended hereto.
Claims
WHAT IS CLAIMED IS:
1. A method of recognizing one or more particle ty pes in a sample using flow cytometry, the method comprising: obtaining a set of waveform data, the waveform data generated by interrogation of the sample to produce light signals from a plurality of particles within the sample, the plurality of particles including one or more particle types; segmenting the set of waveform data into a plurality of constituent particle waveforms; submitting the constituent particle waveforms to a model; and identifying, using the model, a particle fype of the one or more particle types for the constituent particle waveforms.
2. The method of claim 1 , further comprising generating a graphical output marking the constituent particle waveforms.
3. The method of claim 1 or 2, wherein one or more particles of the particle type are unlabeled.
4. The method of claim 3, wherein each particle of the plurality of particles are unlabeled.
5. The method of any of claims 1-4, wherein the model is trained using a set of training waveform data, the set of training waveform data generated by interrogation of a stained particle type to produce light signals from the plurality of particles within the stained particle type.
6. The method of claim 5, wherein the training waveform data including at least one of a forward scatter waveform, a side scatter waveform, an autofluorescence waveform, and a fluorescent waveform associated with the stain.
7. The method of claim 6, wherein the training waveform data including each of the forward scatter waveform, the side scatter waveform, the autofluorescence waveform, and the fluorescent waveform associated with the stain.
8. The method of any of claims 1-7, wherein the set of waveform data including at least one of a forward scatter waveform, a side scatter waveform, and an autofluorescence waveform.
9. The method of claim 8, wherein the set of waveform data including each of the forward scatter waveform, the side scatter waveform, and the autofluorescence waveform.
10. The method of any of claims 1-9, wherein the model is a binary classifier.1 1. The method of any of claims 1-10, wherein the model is a first model and the particle type is a first particle type and the method further comprises identifying, using a second model, a second particle type of the one or more particle types for another constituent particle waveform of the plurality of constituted particle waveforms.
12. The method of any of claims 1-11, wherein the sample is an unprocessed sample, and / or is one of a peripheral blood sample or a bone marrow sample; and wherein the sample remains viable for further processing and use after obtaining the waveform data.
13. The method of claim 12, wherein the particle type is at least one of erythrocytes, thrombocytes and leukocytes.
14. The method of any of claims 1-13, wherein the set of waveform data includes time data between each particle of the plurality of particles in which the particle is not interrogated.
15. A system for recognizing one or more particle types in a sample using flow cytometry, the system comprising: a model trained using a set of training waveform data, the set of training waveform data generated by interrogation of a stained particle type to produce light signals from the plurality of particles within the stained particle type; anda processor in communication with a memory storing instructions which, when executed by the processor, cause the processor to: obtain a set of waveform data, the waveform data generated by interrogation of the sample to produce light signals from a plurality of particles within the sample; submit the set of waveform data to the model; and identify, using the model, one or more constituent waveforms in the set of waveform data as being associated with a particle type, wherein the particle ty pe comprises unlabeled particles of a same type as the stained particle type.
Citation Information
Patent Citations
Method for label-free image cytometry
US20170052106A1
Cell analysis method, training method for deep learning algorithm, cell analyzer, training apparatus for deep learning algorithm, cell analysis program, and training program for deep learning algorithm
US20220003745A1
Cell analysis method and cell analyzer
US20230314300A1