Subsampling of flow cytometry event data
Patent Information
- Application Number
- JP2024169157
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-04-19
- Filing Date
- 2024-09-27
- Publication Date
- 2026-02-17
AI Technical Summary
Existing flow cytometry methods face challenges in efficiently managing and visualizing large, multidimensional datasets, particularly in identifying and isolating rare cell populations, leading to inefficiencies in data processing and sample analysis.
The implementation of dimensionality reduction techniques, such as t-distributed Stochastic Neighbor Embedding (t-SNE), to transform high-dimensional flow cytometry event data into lower-dimensional spaces, allowing for selective subsampling and preservation of rare cell populations, while discarding common events.
This approach enables efficient representation of rare cell populations in subsampled datasets, reducing data volume and computational load while maintaining critical information, thereby enhancing the analysis and isolation of rare cell types.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates generally to the field of automated particle evaluation, and more specifically to sample analysis and particle characterization methods. [Background technology]
[0002] Particle analyzers such as flow cytometers allow characterization of particles based on electro-optical measurements such as light scattering and fluorescence. In a flow cytometer, particles such as molecules in a fluid suspension, analyte-bound beads, or individual cells, typically pass through a detection region where the particles are exposed to excitation light from one or more lasers, and the light scattering and fluorescence properties of the particles are measured. The particles or their components are typically labeled with fluorescent dyes to facilitate detection. A large number of different particles or components can be detected simultaneously by labeling different particles or components using spectrally distinct fluorescent dyes. Different cell types can be identified by their light scattering properties and fluorescent emissions resulting from labeling various cellular proteins or other components with fluorescent dye-labeled antibodies or other fluorescent probes. The data obtained from the analysis of cells (or other particles) by multicolor flow cytometry is multidimensional, with each cell corresponding to a point in a multidimensional space defined by the measured parameters. A population of cells or particles can be identified as a cluster of points in the data space. Summary of the Invention
[0003] Disclosed herein are systems, devices, computer readable media, and methods for subsampling flow cytometry event data. In some embodiments, the method includes, under control of a processor, converting first flow cytometry event data associated with a first event of a first plurality of events of a flow cytometry event data set in a higher dimensional space to first transformed flow cytometry event data associated with the first event in a first lower dimensional space. The first event may be associated with a positive subsampling requirement. The first lower dimensional space may be associated with a first plurality of bins. The first transformed flow cytometry event data may be associated with a first bin of the first plurality of bins. The method may include converting second flow cytometry event data associated with a second event of the first plurality of events of the flow cytometry event data set in the higher dimensional space to second transformed flow cytometry event data associated with the second event in the first lower dimensional space. The second event may be associated with a positive subsampling requirement. The second converted flow cytometry event data may be associated with a second bin of the first plurality of bins. The method may include determining that a first bin associated with the first converted flow cytometry event data and a second bin associated with the second converted flow cytometry event data are different. The method may include generating a sub-sampled flow cytometry event data set of flow cytometry event data including first flow cytometry event data associated with the first event and second flow cytometry event data associated with the second event.
[0004] In some embodiments, the method can include receiving flow cytometry event data including first flow cytometry event data and second flow cytometry event data. The method can include determining that the first flow cytometry event data of a first event of the first plurality of events is associated with a positive subsampling requirement and / or determining that the second flow cytometry event data of a second event of the first plurality of events is associated with a positive subsampling requirement. The method can include determining that the first converted flow cytometry event data is associated with a first bin of the first plurality of bins and / or determining that the second converted flow cytometry event data is associated with a second bin of the first plurality of bins. The method can include determining a first descriptor of the first converted flow cytometry event data based on the first bin of the first plurality of bins and / or determining a second descriptor of the second converted flow cytometry event data based on the second bin of the first plurality of bins. The first descriptor of the first transformed flow cytometry event data associated with the first bin may be a first bin number of the first bin of the first plurality of bins, and / or the second descriptor of the second transformed flow cytometry event data associated with the second bin may be a second bin number of the first bin of the first plurality of bins. The first flow cytometry event data may be associated with a first rare cell, and / or the second flow cytometry event data may be associated with a second rare cell. The first rare cell and the second rare cell may be cells of different cell types. The method may include adding the first bin, the first descriptor, and / or the first bin number to the memory data structure, and / or adding the second bin, the second descriptor, and / or the second bin number to the memory data structure.
[0005] In some embodiments, the method includes transforming third flow cytometry event data associated with a third event of the first plurality of events of the flow cytometry event data set in the higher dimensional space to third transformed flow cytometry event data associated with the third event in the first lower dimensional space. The third event may be associated with a positive subsampling requirement. The third transformed flow cytometry event data may be associated with a third bin of the first plurality of bins. The method may include determining that the third bin associated with the third transformed flow cytometry event data is a first bin associated with the first transformed flow cytometry event data or a second bin associated with the second transformed flow cytometry event data. The third flow cytometry event data may not be within the subsampled flow cytometry event data of the flow cytometry event data. The method may include determining a third descriptor of the third transformed flow cytometry event data based on the third bin of the first plurality of bins. A third descriptor of the third transformed flow cytometry event data associated with the third bin can be a third bin number of the third bin of the first plurality of bins. The method can include determining that the third bin, the third descriptor, and / or the third bin number are not within the memory data structure.
[0006] In some embodiments, the method includes determining that fourth flow cytometry event data associated with a fourth event of the first plurality of events is associated with a negative subsampling requirement. The generating may include generating a subsampled flow cytometry event data set of flow cytometry event data including the fourth flow cytometry event data associated with the fourth event. The method may include receiving a plurality of gates defining a plurality of cells of interest, the fourth flow cytometry event data being associated with a cell of interest of the plurality of cells of interest. The fourth flow cytometry event data may be associated with the sorted cells.
[0007] In some embodiments, the method includes transforming second flow cytometry event data associated with a second event of the second plurality of events of the flow cytometry event dataset in the higher dimensional space to second transformed flow cytometry event data associated with a second event of the second plurality of events in the first lower dimensional space. The second event of the second plurality of events may be associated with a positive subsampling requirement. The second transformed flow cytometry event data associated with the second event of the second plurality of events may be associated with a second bin of the first plurality of bins. The second bin associated with the second transformed flow cytometry event data associated with the second event of the second plurality of events and the first bin associated with the first transformed flow cytometry event data associated with the first event of the first plurality of events may be identical. The generating may include generating a subsampled flow cytometry event dataset of the flow cytometry event data including the second flow cytometry event data associated with the second event of the second plurality of events. The method can include determining that a last event of the first plurality of events is associated with a time parameter or an event number that exceeds a predetermined threshold. The method can include resetting the memory data structure. The method can include adding to the memory data structure a second bin associated with second transformed flow cytometry event data associated with a second event of the second plurality of events. In some embodiments, the method can include receiving a degree of subsampling parameter. The method can include determining the predetermined threshold based on the degree of subsampling parameter.
[0008] In some embodiments, transforming the first flow cytometry event data includes transforming the first flow cytometry event data using a first dimensionality reduction function. Transforming the second flow cytometry event data can include transforming the second flow cytometry event data using the first dimensionality reduction function. The first dimensionality reduction function and / or the second dimensionality reduction function can be a linear dimensionality reduction function. The first dimensionality reduction function and / or the second dimensionality reduction function can be a non-linear dimensionality reduction function. The non-linear dimensionality reduction function can be a t-SNE (Stochastic Neighbor Embedding). The method can include first receiving a dimensionality reduction function, or an identification thereof.
[0009] In some embodiments, transforming the first flow cytometry event data includes transforming the first flow cytometry event data into first transformed flow cytometry event data associated with the first event in a second, lower dimensional space using a second dimensionality reduction function. The second, lower dimensional space may be associated with a second plurality of bins. The first transformed flow cytometry event data in the second, lower dimensional space may be associated with a first bin of the second plurality of bins. Transforming the second flow cytometry event data may include transforming the second flow cytometry event data into second transformed flow cytometry event data associated with the second event in a second, lower dimensional space using a second dimensionality reduction function. The second transformed flow cytometry event data in the second, lower dimensional space may be associated with a second bin of the second plurality of bins. The first bin of the first plurality of bins may be associated with a first type of cell of interest. The second bin of the second plurality of bins may be associated with a second type of cell of interest. The second bin of the first plurality of bins may not be associated with the first type of target cells. The second bin of the first plurality of bins may not be associated with the second type of target cells. The first bin of the second plurality of bins may not be associated with the second type of target cells. The first bin of the second plurality of bins may not be associated with the first type of target cells. A combination of a first bin of the first plurality of bins and a first bin of the second plurality of bins may be associated with a first type of target cells. A combination of a second bin of the first plurality of bins and a second bin of the second plurality of bins may be associated with a second type of target cells. A combination of a first bin of the first plurality of bins and a second bin of the second plurality of bins may not be associated with the first type of target cells and the second type of target cells. A combination of a second bin of the first plurality of bins and a first bin of the second plurality of bins may not be associated with the first type of target cells and the second type of target cells.
[0010] In some embodiments, two bins of the first plurality of bins have the same size. Each bin of the first plurality of bins can have the same size. Two bins of the first plurality of bins can have different sizes. Two bins of the first plurality of bins can include (approximately) the same number of transformed flow cytometry event data. Each of the first plurality of bins can include approximately the same number of transformed flow cytometry event data. The method can include determining a size of each of the first plurality of bins. The method can include determining a size of each of the first plurality of bins based on a plurality of gates. The method can include determining a size of each of the first plurality of bins based on transformed flow cytometry event data associated with a plurality of cells of interest.
[0011] Disclosed herein include embodiments of a computing system for subsampling flow cytometry event data. In some embodiments, the computing system may include a non-transitory memory configured to store executable instructions, and a processor (e.g., a hardware processor or a virtual processor) in communication with the non-transitory memory, the processor being programmed by the executable instructions to convert a first flow cytometry event data associated with a first event of a first plurality of events in a higher dimensional space into a first transformed flow cytometry event data associated with the first event of a flow cytometry event data set in a first lower dimensional space, the first event being associated with a positive sampling requirement, the first lower dimensional space being associated with a first plurality of bins, and the first transformed flow cytometry event data being associated with the first bin of the first plurality of bins. The processor may be programmed with the executable instructions to convert second flow cytometry event data associated with a second event of the first plurality of events of the flow cytometry event dataset in the higher dimensional space to second transformed flow cytometry event data associated with the second event in the first lower dimensional space, the second event being associated with a positive subsampling requirement, the second transformed flow cytometry event data being associated with a second bin of the first plurality of bins. The processor may be programmed with the executable instructions to determine that a first bin associated with the first transformed flow cytometry event data and a second bin associated with the second transformed flow cytometry event data are different. The processor may be programmed with the executable instructions to generate a subsampled flow cytometry event dataset of flow cytometry event data including the first flow cytometry event data associated with the first event and the second flow cytometry event data associated with the second event.
[0012] In some embodiments, the processor is programmed by the executable instructions to receive flow cytometry event data including first flow cytometry event data and second flow cytometry event data. The processor can be programmed by the executable instructions to determine that the first flow cytometry event data of a first event of the first plurality of events is associated with a positive subsampling requirement. The processor can be programmed by the executable instructions to determine that the second flow cytometry event data of a second event of the first plurality of events is associated with a positive subsampling requirement. The processor can be programmed by the executable instructions to determine that the first converted flow cytometry event data is associated with a first bin of the first plurality of bins. The processor can be programmed by the executable instructions to determine that the second converted flow cytometry event data is associated with a second bin of the first plurality of bins.
[0013] In some embodiments, the processor is programmed by the executable instructions to determine a first descriptor of the first transformed flow cytometry event data based on a first bin of the first plurality of bins. The processor can be programmed by the executable instructions to determine a second descriptor of the second transformed flow cytometry event data based on a second bin of the first plurality of bins. The first descriptor of the first transformed flow cytometry event data associated with the first bin can be a first bin number of the first bin of the first plurality of bins, and / or the second descriptor of the second transformed flow cytometry event data associated with the second bin can be a second bin number of the first bin of the first plurality of bins. The first flow cytometry event data can be associated with a first rare cell, and / or the second flow cytometry event data can be associated with a second rare cell. The first rare cell and the second rare cell can be cells of different cell types.
[0014] In some embodiments, the processor is programmed with the executable instructions to add the first bin, the first descriptor, and / or the first bin number to the memory data structure, and / or add the second bin, the second descriptor, and / or the second bin number to the memory data structure. In some embodiments, the processor is programmed with the executable instructions to convert third flow cytometry event data associated with a third event of the first plurality of events of the flow cytometry event data set in the higher dimensional space to third transformed flow cytometry event data associated with the third event in the first lower dimensional space. The third event may be associated with a positive subsampling requirement. The third transformed flow cytometry event data may be associated with the third bin of the first plurality of bins. The processor may be programmed with the executable instructions to determine that a third bin associated with the third transformed flow cytometry event data is a first bin associated with the first transformed flow cytometry event data or a second bin associated with the second transformed flow cytometry event data. The third flow cytometry event data may not be within the subsampled flow cytometry event data of the flow cytometry event data. The processor may be programmed with the executable instructions to determine a third descriptor of the third transformed flow cytometry event data based on a third bin of the first plurality of bins. The third descriptor of the third transformed flow cytometry event data associated with the third bin may be a third bin number of the third bin of the first plurality of bins. The processor may be programmed with the executable instructions to determine that the third bin, the third descriptor, and / or the third bin number are not within the memory data structure.
[0015] In some embodiments, the processor is programmed by the executable instructions to determine that a fourth flow cytometry event data associated with a fourth event of the first plurality of events is associated with a negative subsampling requirement. To generate a subsampled flow cytometry event data set, the processor can be programmed by the executable instructions to generate a subsampled flow cytometry event data set of the flow cytometry event data set including a fourth flow cytometry event data associated with the fourth event. The processor can be programmed by the executable instructions to receive a plurality of gates defining a plurality of target cells. The fourth flow cytometry event data can be associated with a target cell of the plurality of target cells. The fourth flow cytometry event data can be associated with a sorted cell.
[0016] In some embodiments, the processor is programmed with executable instructions to convert second flow cytometry event data associated with a second event of the second plurality of events of the flow cytometry event data set in the higher dimensional space to second transformed flow cytometry event data associated with a second event of the second plurality of events in the first lower dimensional space. The second event of the second plurality of events may be associated with a positive subsampling requirement. The second transformed flow cytometry event data associated with the second event of the second plurality of events may be associated with a second bin of the first plurality of bins. The second bin associated with the second transformed flow cytometry event data associated with the second event of the second plurality of events and the first bin associated with the first transformed flow cytometry event data associated with the first event of the first plurality of events may be identical. To generate the subsampled flow cytometry event data set, the processor can be programmed by the executable instructions to generate a subsampled flow cytometry event data set of the flow cytometry event data set including a second flow cytometry event data associated with a second event of the second plurality of events. The processor can be programmed by the executable instructions to include determining that a last event of the first plurality of events is associated with a time parameter or event number that exceeds a predetermined threshold. The processor can be programmed by the executable instructions to reset the memory data structure. The processor can be programmed by the executable instructions to add to the memory data structure a second bin associated with the second transformed flow cytometry event data associated with the second event of the second plurality of events.
[0017] In some embodiments, the processor is programmed by the executable instructions to receive the degree of the sub-sampling parameter. The processor may be programmed by the executable instructions to determine the predetermined threshold based on the degree of the sub-sampling parameter.
[0018] In some embodiments, to transform the first flow cytometry event data, the processor may be programmed by the executable instructions to transform the first flow cytometry event data using a first dimensionality reduction function, and / or to transform the second flow cytometry event data, the processor may be programmed by the executable instructions to transform the second flow cytometry event data using the first dimensionality reduction function. The first dimensionality reduction function and / or the second dimensionality reduction function may be a linear dimensionality reduction function. The first dimensionality reduction function and / or the second dimensionality reduction function may be a non-linear dimensionality reduction function. The non-linear dimensionality reduction function may be t-SNE (Stochastic Neighbor Embedding). The processor may be programmed by the executable instructions to first receive the dimensionality reduction function, or an identification thereof.
[0019] In some embodiments, to transform the first flow cytometry event data, the processor is programmed by the executable instructions to transform the first flow cytometry event data into first transformed flow cytometry event data associated with the first event in a second, lower dimensional space using a second dimensional reduction function. The second, lower dimensional space may be associated with a second plurality of bins. The first transformed flow cytometry event data in the second, lower dimensional space may be associated with a first bin of the second plurality of bins. To transform the second flow cytometry event data, the processor can be programmed by the executable instructions to transform the second flow cytometry event data into second transformed flow cytometry event data associated with the second event in a second, lower dimensional space using a second dimensional reduction function. The second transformed flow cytometry event data in the second, lower dimensional space may be associated with a second bin of the second plurality of bins. A first bin of the first plurality of bins may be associated with a first type of target cell, a second bin of the second plurality of bins may be associated with a second type of target cell, a second bin of the first plurality of bins may not be associated with a first type of target cell, a second bin of the first plurality of bins may not be associated with a second type of target cell, a first bin of the second plurality of bins may not be associated with a second type of target cell, and / or a first bin of the second plurality of bins may not be associated with a first type of target cell. A combination of a first bin of the first plurality of bins and a first bin of the second plurality of bins may be associated with a first type of target cell, and / or a combination of a second bin of the first plurality of bins and a second bin of the second plurality of bins may be associated with a second type of target cell.The combination of a first bin of the first plurality of bins and a second bin of the second plurality of bins is not associated with a first type of target cell and a second type of target cell, and / or the combination of a second bin of the first plurality of bins and a first bin of the second plurality of bins is not associated with a first type of target cell and a second type of target cell.
[0020] In some embodiments, two bins of the first plurality of bins have the same size. Each bin of the first plurality of bins can have the same size. Two bins of the first plurality of bins can have different sizes. Two bins of the first plurality of bins can include approximately the same number of transformed flow cytometry event data. Each of the first plurality of bins can include approximately the same number of transformed flow cytometry event data. The processor can be programmed by the executable instructions to determine a size of each of the first plurality of bins. The processor can be programmed by the executable instructions to determine a size of each of the first plurality of bins based on a plurality of gates. The processor can be programmed by the executable instructions to determine a size of each of the first plurality of bins based on transformed flow cytometry event data associated with a plurality of cells of interest. [Brief description of the drawings]
[0021] [Figure 1] FIG. 1 shows a functional block diagram for an example of a sorting control system for analyzing and displaying biological events. [Figure 2A] FIG. 1 is a schematic diagram of a particle sorting system according to one embodiment presented herein. [Figure 2B] FIG. 2 is a schematic diagram of another particle sorting system according to an embodiment presented herein. [Diagram 3] FIG. 1 shows a functional block diagram of a particle analysis system for computation-based sample analysis and particle characterization. [Figure 4]FIG. 1 is a flow diagram illustrating an exemplary method for subsampling flow cytometry event data. [Diagram 5] FIG. 1 is a block diagram of an exemplary computing system configured to implement a method for subsampling flow cytometry event data. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0022] In the following detailed description, reference is made to the accompanying drawings, which form a part of this specification. In the drawings, similar symbols typically identify similar components unless the context dictates otherwise. The exemplary embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments can be utilized and other changes can be made without departing from the spirit or scope of the subject matter presented herein. It is readily understood that the aspects of the present disclosure, as generally described herein and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations. All of these are expressly contemplated and made a part of the disclosure herein.
[0023] Particle analyzers, such as flow and scanning cytometers, are analytical tools that allow characterization of particles based on electro-optical measurements such as light scattering and fluorescence. In a flow cytometer, particles, such as molecules in a fluid suspension, analyte-bound beads, or individual cells, are typically passed through a detection region where the particles are exposed to excitation light from one or more lasers, and the light scattering and fluorescence properties of the particles are measured. The particles or their components are typically labeled with fluorescent dyes to facilitate detection. A number of different particles or components can be detected simultaneously by labeling the different particles or components using spectrally distinct fluorescent dyes. In some implementations, the analyzer includes multiple photodetectors, one for each of the scattering parameters to be measured and one or more for each of the different dyes to be detected. For example, some embodiments include a spectral configuration in which more than one sensor or detector per dye is used. The acquired data includes the measured signals for each of the light scattering detectors and the fluorescent emission.
[0024] The particle analyzer may further include a means for recording the measured data and analyzing the data. For example, data storage and analysis may be performed using a computer connected to the detection electronics. For example, the data may be stored in a table format, with each row corresponding to data for one particle and columns corresponding to each of the measured features. The use of a standard file format, such as the Flow Cytometry Standard ("FCS") file format, for storing data from the particle analyzer facilitates the analysis of the data using a separate program and / or machine. Using current analysis methods, the data is typically displayed in a one-dimensional histogram or two-dimensional (2D) plot for ease of visualization, although other methods may be used to visualize multi-dimensional data.
[0025] For example, parameters measured using a flow cytometer typically include light scattered by particles at a narrow angle along the approximate forward direction (termed forward scatter (FSC)), light scattered by particles in a direction orthogonal to the excitation laser (termed side scatter (SSC)), and light emitted from fluorescent molecules at one or more detectors that measure signals over a range of spectral wavelengths, or light emitted by fluorescent dyes that are primarily detected at that particular detector or array of detectors. Different cell types can be distinguished by their light scattering properties and fluorescent emissions resulting from labeling various cellular proteins or other components with fluorochrome-labeled antibodies or other fluorescent probes.
[0026] Both flow cytometers and scanning cytometers are commercially available, for example, from BD Biosciences (San Jose, Calif.). Flow cytometry is described, for example, in Landy et al. (eds.), Clinical Flow Cytometry, Annals of the New York Academy of Sciences Volume 677 (1993); Bauer et al. (eds.), Clinical Flow Cytometry: Principles and Applications, Williams & Wilkins (1993); Ormerod (ed.), Flow Cytometry: A Practical Approach, Oxford Univ. Press (1994); Jaroszeski et al. (eds.), Flow Cytometry Protocols, Methods in Molecular Biology No. 91, Humana Press (1997); and Practical Shapiro, Flow Cytometry, 4th ed., Wiley-Liss (2003), each of which is incorporated herein by reference. Fluorescence imaging microscopy is described, for example, in Pawley (ed.), Handbook of Biological Confocal Microscopy, 2nd Edition, Plenum Press (1989), incorporated herein by reference.
[0027] The data obtained from the analysis of cells (or other particles) by multicolor flow cytometry is multidimensional, with each cell corresponding to a point in a multidimensional space defined by the measured parameters. A population of cells or particles can be identified as a cluster of points in the data space. Identification of the cluster, and thereby the population, can be performed manually by drawing a gate around the population displayed in one or more two-dimensional plots, referred to as "scatter plots" or "dot plots" of the data. Alternatively, the cluster can be identified and a gate that defines the limit of the population can be determined automatically. Examples of methods for automatic gating are described, for example, in U.S. Pat. Nos. 4,845,653, 5,627,040, 5,739,000, 5,795,727, 5,962,238, 6,014,904, 6,944,338, and 8,990,047, each of which is incorporated herein by reference.
[0028] Flow cytometry is a valuable method for the analysis and isolation of biological particles such as cells and constituent molecules. It therefore has a wide range of diagnostic and therapeutic applications. The method utilizes a fluid stream to linearly separate particles so that they can pass a detection device in single file. Individual cells can be differentiated according to their position within the fluid stream and the presence of detectable markers. Thus, flow cytometers can be used to characterize and generate diagnostic profiles of populations of biological particles.
[0029] Isolation of biological particles has been accomplished by adding sorting or collection capabilities to flow cytometers. Particles in the separated stream that are detected as having one or more desired properties can be individually isolated from the sample stream by mechanical or electrical separation. This method of flow sorting has been used to sort different types of cells, separate sperm carrying X and Y chromosomes for animal breeding, sort chromosomes for genetic analysis, and isolate specific organisms from complex populations of organisms.
[0030] The use of gating can help sort through and make sense of the large amounts of data that may be generated from a sample. Given the large amounts of data presented for a given sample, there is a need to efficiently control the graphical display of the data.
[0031] Fluorescence-activated particle sorting or cell sorting is a specialized type of flow cytometry. Fluorescence-activated particle sorting or cell sorting provides a method for sorting a heterogeneous mixture of particles, one cell at a time, into one or more containers based on the specific light scattering and fluorescent properties of each cell. Fluorescence signals from individual cells are thereby recorded and specific cells of interest are physically separated. The acronym FACS is trademarked and owned by Becton, Dickinson and Company (Franklin Lakes, NJ), and can be used to refer to devices for performing fluorescence-activated particle sorting or cell sorting.
[0032] The particle suspension is placed near the center of a narrow, rapidly flowing liquid stream. The flow is arranged such that when particles arrive stochastically (e.g., Poisson process) at the detection region, on average there is a large separation relative to their diameter. A vibration mechanism can cause the newly emerged fluid stream to separate into individual droplets containing particles previously characterized at the detection region. The system can generally be tuned such that the probability of two or more particles being present in a droplet is low. If particles are to be sorted to be collected, a charge can be applied to the flow cell and the newly emerged stream during the period when one or more droplets form and separate from the stream. These charged droplets then travel through an electrostatic deflection system that deflects the droplets into a target vessel based on the charge applied to the droplets.
[0033] A sample can contain thousands, if not millions, of cells. The cells can be sorted to purify the sample into cells of interest. The sorting process can generally distinguish three types of cells: cells of interest, non-cells of interest, and cells that cannot be distinguished. To sort cells at high purity (e.g., high concentration of cells of interest), the droplet generating particle sorter can electronically stop sorting if a desired cell is too close to another unwanted cell, thereby reducing contamination of the sorted population by any unintentional inclusion of unwanted particles in the droplets containing the particles of interest.
[0034] Disclosed herein are systems, devices, computer readable media, and methods for subsampling flow cytometry event data. In some embodiments, the method includes, under control of a processor, converting first flow cytometry event data associated with a first event of a first plurality of events in a higher dimensional space to first transformed flow cytometry event data associated with the first event in a first lower dimensional space. The first event may be associated with a positive subsampling requirement. The first lower dimensional space may be associated with a first plurality of bins. The first transformed flow cytometry event data may be associated with a first bin of the first plurality of bins. The method may include converting second flow cytometry event data associated with a second event of the first plurality of events in the higher dimensional space to second transformed flow cytometry event data associated with the second event in the first lower dimensional space. The second event may be associated with a positive subsampling requirement. The second transformed flow cytometry event data may be associated with a second bin of the first plurality of bins. The method can include determining that a first bin associated with the first transformed flow cytometry event data and a second bin associated with the second transformed flow cytometry event data are distinct. The method can include generating a sub-sampled flow cytometry event data set of flow cytometry event data including first flow cytometry event data associated with the first event and second flow cytometry event data associated with the second event.
[0035] Disclosed herein include embodiments of a computing system for subsampling flow cytometry event data. In some embodiments, the computing system may include a non-transitory memory configured to store executable instructions, and a processor (e.g., a hardware processor or a virtual processor) in communication with the non-transitory memory, the processor being programmed by the executable instructions to convert a first flow cytometry event data associated with a first event of a first plurality of events in a higher dimensional space into a first transformed flow cytometry event data associated with the first event of a flow cytometry event data set in a first lower dimensional space, the first event being associated with a positive sampling requirement, the first lower dimensional space being associated with a first plurality of bins, and the first transformed flow cytometry event data being associated with the first bin of the first plurality of bins. The processor may be programmed with the executable instructions to convert second flow cytometry event data associated with a second event of the first plurality of events of the flow cytometry event dataset in the higher dimensional space to second transformed flow cytometry event data associated with the second event in the first lower dimensional space, the second event being associated with a positive subsampling requirement, the second transformed flow cytometry event data being associated with a second bin of the first plurality of bins. The processor may be programmed with the executable instructions to determine that a first bin associated with the first transformed flow cytometry event data and a second bin associated with the second transformed flow cytometry event data are different. The processor may be programmed with the executable instructions to generate a subsampled flow cytometry event dataset of flow cytometry event data including the first flow cytometry event data associated with the first event and the second flow cytometry event data associated with the second event.
[0036] definition As used herein, the terms specifically set forth below have the following definitions: Unless otherwise defined in this section, all terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs.
[0037] As used herein, "system," "instrument," "apparatus," and "device" generally encompass both hardware (e.g., mechanical and electronic) and, in some implementations, associated software (e.g., specialized computer programs for graphics control) components.
[0038] As used herein, "event" or "event data" generally refers to data (e.g., an assembled packet of data) measured from a single particle, such as a cell or synthetic particle. Typically, data measured from a single particle includes multiple parameters or features, including one or more light scattering parameters or features, and at least one other parameter or feature derived from fluorescence detected from the particle, such as the intensity of fluorescence. Thus, each event can be represented as a vector of parameter and feature measurements, with each measured parameter or feature corresponding to one dimension of the data space. In some embodiments, data measured from a single particle includes image data, electrical data, time data, or acoustic data. An event can be associated with an experiment, an assay, or a sample source that can be identified in association with the measurement data.
[0039] As used herein, a "population," or "subpopulation," of particles, such as cells or other particles, generally refers to a group of particles that possess characteristics (e.g., optical properties, impedance properties, temporal properties, etc.) with respect to one or more measured parameters such that the measured parameter data form a cluster in the data space. Thus, a population may be recognized as a cluster in the data. Conversely, each data cluster is generally interpreted as corresponding to a population of a particular type of cell or particle, although clusters that typically correspond to noise or background are also observed. Clusters may be defined, for example, with respect to a subset of measured parameters, with a subset of dimensions corresponding to populations that differ only in a subset of measured parameters or features extracted from cell or particle measurements.
[0040] As used herein, "gate" generally refers to a classifier boundary that identifies a subset of data of interest. In cytometry, a gate can bind a specific group of events of interest. As used herein, "gating" generally refers to the process of classifying data using a gate defined for a given set of data, where a gate can be one or more regions of interest combined with Boolean logic.
[0041] Various embodiments and specific examples of systems in which they may be implemented are described further below.
[0042] Sorting Control System 1 shows a functional block diagram for an example of a sorting control system for analyzing and displaying biological events, such as an analysis controller 100. Analysis controller 100 can be configured to implement various processes for controlling the graphical display of biological events.
[0043] The particle analyzer or sorting system 102 can be configured to acquire biological event data. For example, a flow cytometer can generate flow cytometry event data. The particle analyzer 102 can be configured to provide the biological event data to the analysis controller 100. A data communication channel can be included between the particle analyzer 102 and the analysis controller 100. The biological event data can be provided to the analysis controller 100 via the data communication channel.
[0044] The analysis controller 100 can be configured to receive biological event data from the particle analyzer 102. The biological event data received from the particle analyzer 102 can include flow cytometry event data. The analysis controller 100 can be configured to provide a graphical display including a first plot of the biological event data on the display device 106. The analysis controller 100 can be further configured to render a region of interest, e.g., as a gate around a population of the biological event data shown by the display device 106, overlaid on the first plot. In some embodiments, the gate can be a logical combination of one or more graphical regions of interest plotted in a single parameter histogram or bivariate plot.
[0045] Analysis controller 100 can further configure biological event data on display device 106 within a gate to be different from other events within the biological event data outside the gate. For example, analysis controller 100 can be configured to render the color of the biological event data contained within the gate to be different from the color of the biological event data outside the gate. Display device 106 can be implemented as a monitor, tablet computer, smartphone, or other electronic device configured to present a graphical interface.
[0046] The analysis controller 100 can be configured to receive a gate selection signal identifying a gate from a first input device. For example, the first input device can be implemented as a mouse 110. The mouse 110 can initiate a gate selection signal to the analysis controller 100 identifying a gate to be displayed on or manipulated via the display device 106 (e.g., by clicking the desired gate when a cursor is positioned there or by clicking within the desired gate). In some implementations, the first device can be implemented as a keyboard 108 or other means for providing an input signal to the analysis controller 100, such as a touch screen, a stylus, an optical detector, or a voice recognition system. Some input devices can include multiple input functions. In such implementations, each input function can be considered an input device. For example, as shown in FIG. 1, the mouse 110 can include a right mouse button and a left mouse button, each of which can generate a trigger event.
[0047] The trigger event may cause the analysis controller 100 to change the way the data is displayed, what portions of the data are actually displayed on the display device 106, and / or provide input for further processing, such as selection of a subject population for particle sorting.
[0048] In some embodiments, the analysis controller 100 can be configured to detect when a gate selection is initiated by the mouse 110. The analysis controller 100 can be further configured to automatically modify the visualization of the plot to facilitate the gating process. The modification can be based on a particular distribution of the biological event data received by the analysis controller 100.
[0049] The analysis controller 100 may be connected to a storage device 104. The storage device 104 may be configured to receive and store biological event data from the analysis controller 100. The storage device 104 may also be configured to receive and store flow cytometry event data from the analysis controller 100. The storage device 104 may be further configured to enable retrieval of biological event data, such as flow cytometry event data, by the analysis controller 100.
[0050] The display device 106 can be configured to receive display data from the analysis controller 100. The display data can include plots of the biological event data and gates outlining sections of the plots. The display device 106 can be further configured to modify the information presented according to input received from the analysis controller 100 in conjunction with input from the particle analyzer 102, the storage device 104, the keyboard 108, and / or the mouse 110.
[0051] In some implementations, the analysis controller 100 can generate a user interface for receiving exemplary events for sorting. For example, the user interface can include a control for receiving exemplary events or exemplary images. The exemplary events or images, or exemplary gates, can be provided prior to collection of event data for the sample or based on an initial set of events for a portion of the sample.
[0052] Particle Sorting System A common flow sorting technique, which may be referred to as "electrostatic cell sorting", utilizes droplet sorting, where a stream or moving fluid column containing linearly separated particles is broken into droplets, and droplets containing particles of interest are electrically charged and deflected into a collection tube by passing through an electric field. Droplet sorting systems can form droplets at a rate of 100,000 drops / second in a fluid stream passing through a nozzle having a diameter of less than 100 micrometers. Droplet sorting typically requires that the droplets separate from the stream at a certain distance from the nozzle tip. This distance is usually on the order of a few millimeters from the nozzle tip, and can be stabilized and maintained for an undisturbed fluid stream by vibrating the nozzle tip at a given frequency with an amplitude to keep the breakoff constant. For example, in some embodiments, adjusting the amplitude of a sinusoidally shaped voltage pulse at a given frequency keeps the breakoff stable and constant.
[0053] Typically, linearly associated particles in the stream are characterized as they pass through an observation point located within a flow cell or cuvette, or just below the nozzle tip. Once a particle is identified as meeting one or more desired criteria, it can be predicted when the particle will reach the droplet break-off point and leave the stream in a droplet. Ideally, a brief charge is applied to the fluid stream just before the droplet containing the selected particle leaves the stream, and is grounded just after the droplet leaves. The droplet to be sorted maintains its charge as it leaves the fluid stream, while all other droplets remain uncharged. The charged droplets are deflected laterally from the downward trajectory of the other droplets by the electric field and are collected in a sample tube. The uncharged droplets fall directly into a drain.
[0054] FIG. 2A is a schematic diagram of a particle sorting system 200 (e.g., particle analyzer 102) according to one embodiment presented herein. In some embodiments, the particle sorting system 200 is a cell sorting system. As shown in FIG. 2A, a droplet forming transducer 202 (e.g., a piezoelectric oscillator) is coupled to a fluid conduit 201, which may be coupled to, may include, or may be a nozzle 203. Within the fluid conduit 201, a sheath fluid 204 hydrodynamically focuses a sample fluid 206 containing particles 209 into a moving fluid column 208 (e.g., a stream). Within the moving fluid column 208, the particles 209 (e.g., cells) are aligned in a single file to cross a monitoring area 211 (e.g., where a laser stream intersects) and are illuminated by an illumination source 212 (e.g., a laser). The vibration of the droplet forming transducer 202 causes the moving fluid column 208 to break up into a number of droplets 210 , some of which contain particles 209 .
[0055] In operation, the detection station 214 (e.g., an event detector) identifies when a particle (or cell) of interest crosses the monitoring area 211. The detection station 214 powers a timing circuit 228, which in turn powers a flash charging circuit 230. At a droplet break-off point, signaled by a timed droplet delay (Δt), a flash charge can be applied to the moving fluid column 208 such that the droplet of interest carries a charge. The droplet of interest can contain one or more particles or cells to be sorted. The charged droplet can then be sorted by activating a deflection plate (not shown) to deflect the droplet into a collection tube or a container such as a multi-well or microwell sample plate in which a well or microwell can be specifically associated with the droplet of interest. As shown in FIG. 2A, the droplet can be collected in a drain container 238.
[0056] The detection system 216 (e.g., a droplet boundary detector) serves to automatically determine the phase of the droplet drive signal as a particle of interest passes through the monitoring area 211. An exemplary droplet boundary detector is described in U.S. Pat. No. 7,679,039, which is incorporated herein by reference in its entirety. The detection system 216 allows the instrument to accurately calculate the location of each detected particle in the droplet. The detection system 216 can feed an amplitude signal 220 and / or a phase signal 218, which in turn feeds an amplitude control circuit 226 and / or a frequency control circuit 224 (via an amplifier 222). The amplitude control circuit 226 and / or the frequency control circuit 224 then control the droplet forming transducer 202. The amplitude control circuit 226 and / or the frequency control circuit 224 can be included in a control system.
[0057] In some implementations, the sorting electronics (e.g., detection system 216, detection station 214, and processor 240) may be coupled with a memory configured to store the detected events and a sorting decision based thereon. The sorting decision may be included in the particle's event data. In some implementations, the detection system 216 and detection station 214 may be implemented as a single detection unit or may be communicatively coupled such that event measurements may be collected by one of the detection system 216 or detection station 214 and provided to a non-collection element.
[0058] FIG. 2B is a schematic diagram of a particle sorting system according to one embodiment presented herein. The particle sorting system 200 shown in FIG. 2B includes deflection plates 252 and 254. An electric charge can be applied via a stream charging wire in the barb. This creates a stream of droplets 210 containing the particles 210 for analysis. The particles can be illuminated with one or more light sources (e.g., lasers) to generate light scattering and fluorescence information. Information about the particles is analyzed, for example, by sorting electronics or other detection systems (not shown in FIG. 2B). The deflection plates 252 and 254 can be independently controlled to attract or repel the charged droplets and direct the droplets toward a desired collection vessel (e.g., one of 272, 274, 276, or 278). As shown in FIG. 2B, the deflection plates 252 and 254 can be controlled to direct the particles along the first path 262 toward the vessel 274 or along the second path 268 toward the vessel 278. If it is not a particle of interest (e.g., does not exhibit scattering or illumination information within a specified sorting range), the deflector may allow the particle to continue along flow path 264. Such uncharged droplets may be passed via aspirator 270, for example to a waste container.
[0059] Sorting electronics may be included to initiate measurement collection, receive the fluorescent signal of the particles, and determine how to adjust the deflection plates to cause sorting of the particles. Exemplary implementations of the embodiment shown in Figure 2B include the BD FACSAria™ line of flow cytometers commercially offered by Becton, Dickinson and Company (Franklin Lakes, NJ).
[0060] In some embodiments, one or more components described for particle sorting system 200 may be used to analyze and characterize particles whether or not the particles are physically sorted into a collection vessel. Similarly, one or more components described below for particle analysis system 300 (FIG. 3) may be used to analyze and characterize particles whether or not the particles are physically sorted into a collection vessel. For example, particles may be grouped or displayed in a tree including at least three groups as described herein using one or more of the components of particle sorting system 200 or particle analysis system 300.
[0061] FIG. 3 shows a functional block diagram of a particle analysis system for computation-based sample analysis and particle characterization. In some embodiments, the particle analysis system 300 is a flow system. The particle analysis system 300 shown in FIG. 3 can be configured, in whole or in part, to perform the methods described herein. The particle analysis system 300 includes a fluid system 302. The fluid system 302 can include or be coupled to a sample tube 310 and a moving fluid column within the sample tube through which particles 330 (e.g., cells) of the sample move along a common sample path 320.
[0062] The particle analysis system 300 includes a detection system 304 configured to collect a signal from each particle as it passes through one or more detection stations along a common sample path. The detection station 308 generally refers to a monitoring area 340 of the common sample path. Detection, in some implementations, can include detecting light or one or more other characteristics of the particle 330 as it passes through the monitoring area 340. In FIG. 3, one detection station 308 is shown with one monitoring area 340. Some implementations of the particle analysis system 300 can include multiple detection stations. Additionally, some detection stations may monitor more than one area.
[0063] Each signal is assigned a signal value to form a data point for each particle. As explained above, this data may be referred to as event data. The data points may be multi-dimensional data points that include values of each property measured for the particle. The detection system 304 is configured to collect a series of such data points over a first time interval.
[0064] The particle analysis system 300 may also include a control system 306. The control system 306 may include one or more processors, amplitude control circuitry 226, and / or frequency control circuitry 224, as shown in FIG. 2B. The illustrated control system 206 may be operatively associated with the fluid system 302. The control system 206 may be configured to generate a calculated signal frequency for at least a portion of the first time interval based on the Poisson distribution and the number of data points collected by the detection system 304 during the first time interval. The control system 306 may further be configured to generate an experimental signal frequency based on the number of data points in the portion of the first time interval. The control system 306 may further compare the experimental signal frequency to a calculated signal frequency or a predetermined signal frequency.
[0065] Subsampling of flow cytometry event data Disclosed herein include systems, devices, computer readable media, and methods for subsampling a dataset (e.g., a large, higher dimensional dataset) that allow rare events and populations to be weighed so that they are appropriately represented in the resulting subset. In some embodiments, when it is undesirable to store the entire dataset, a subset of the dataset that stores all populations (e.g., all populations including rare cells and populations) may be represented in a subset of the dataset or subsampled dataset. In some embodiments, the subset or subsampled dataset stores representative samples from rare subpopulations. In some embodiments, a dataset may be subsampled without discarding rare events or events of interest (e.g., corresponding to rare cells or cells of interest). The system automatically detects and stores rare events while more aggressively discarding common events.
[0066] In some embodiments, the data may be subsampled non-randomly (e.g., semi-randomly). A desired subsampling rate may be selected and the data may then be fed sequentially through the subsampling method. The method may decide to store or discard events (or multi-dimensional event data associated with the events) on a single event basis. The ability to discard events without analyzing the overall distribution of the events eliminates the need to store and analyze large amounts of data.
[0067] The user can select the degree of subsampling parameters, which can determine the duration of the algorithm "memory".
[0068] The user can select one or more transformation or "fingerprint" functions. A transformation or fingerprint function can be a mathematical formula that transforms the data in some way, such as from a higher dimensional space to a lower dimensional space. For example, a transformation or fingerprint function can be a Stochastic Neighbor Embedding (t-SNE). Events can be transformed into a lower dimensional space divided into bins. The bin number can serve as a descriptor of the event.
[0069] In some embodiments, binning may be uniform or based on event density. Binning may be based on automatic population detection. Binning may be based in part on arbitrarily drawn gates (e.g., drawn by a user). In some embodiments, a transformation or fingerprint function may transform events such that similar events have the same identifier. The identifier may be smaller than the data used to generate the identifier. The transformation may be computationally inexpensive to compute. An inverse of the transformation or function may or may not exist. In some embodiments, multiple fingerprint functions may be used. For example, different target populations may be defined using different fingerprint functions. As another example, a target population may be defined based on the combined output of multiple fingerprint functions.
[0070] Third, the user can describe events that should not be subsampled. For example, a gate around a region of interest can be drawn automatically or by the user. Events within the gate around the region of interest may not be subsampled. As another example, any events that are sorted (e.g., unsorted cells) may not be subsampled.
[0071] The event data can be subsampled using the subsampling methods disclosed herein. For example, for each event: Check if the event needs to be subsampled, if the answer is no, save the event. ii. If events are subsampled, use a fingerprint function to generate the descriptors. 1. Compare the descriptor to algorithm "memory". Have you seen this descriptor before? Yes, discard the event. b. No. Save the event and store the descriptor in memory. iii. Check the time or event number: If the time and / or event number exceeds a corresponding threshold generated by the degree of the user's subsampling parameters, reset the algorithm memory.
[0072] The subsampling methods disclosed herein, non-random subsampling methods, can complement or supplement random subsampling used to subsample large data sets. When randomly sampling data, rare populations may be excluded. Subsampling methods may include some, most, or all of the rare population. In particle analysis, such as flow cytometry analysis, rare events can potentially be very valuable. Preserving rare populations can be useful as rare populations are detected when reduced data sets are analyzed. Non-random subsampling methods can intentionally bias the random sampling process, making it more likely that rare populations are represented in the final subsampled data set.
[0073] A naive segmentation of the data space without dimensionality reduction may result in sparsely populated bins due to the so-called "curse of dimensionality." The dimensionality reduction transform or function used may be a relationship-preserving embedding, which allows for binning in a lower dimensional space and more efficient grouping of the data before subsampling.
[0074] Subsampling particle analysis event data method FIG. 4 is a flow diagram illustrating an exemplary method 400 for subsampling particle analysis event data, such as flow cytometry event data. Method 400 may be embodied in a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives of a computing system. For example, a computing system 500, shown in FIG. 5 and described in more detail below, may execute a set of executable program instructions to implement method 400. When method 400 is initiated, the executable program instructions may be loaded into a memory, such as a RAM, and executed by one or more processors of computing system 500. Although method 400 is described with respect to computing system 500 shown in FIG. 5, the description is by way of example only and is not intended to be limiting. In some embodiments, method 400, or portions thereof, may be executed serially or in parallel by multiple computing systems.
[0075] After the method 400 starts at block 404, the method 400 proceeds to block 408, where the computing system can convert first flow cytometry event data associated with a first event of a first plurality of events of the flow cytometry event data set in a higher dimensional space to first transformed flow cytometry event data associated with the first event in a first lower dimensional space. The first event can be associated with a positive subsampling requirement. For example, when subsampling flow cytometry event data including the first flow cytometry event data, the subsampled flow cytometry event data may not include the first flow cytometry event data. The first lower dimensional space can be associated with a first plurality of bins. The first transformed flow cytometry event data can be associated with a first bin of the first plurality of bins. The computing system can indicate (e.g., in a data structure) that the first flow cytometry event data should be included when generating the subsampled flow cytometry event data.
[0076] In some embodiments, a computing system can receive flow cytometry event data including a first flow cytometry event data. The computing system can determine that the first flow cytometry event data of a first event of the first plurality of events is associated with a positive subsampling requirement. The computing system can determine that the first transformed flow cytometry event data is associated with a first bin of the first plurality of bins.
[0077] The processor can be programmed with the executable instructions to determine a first descriptor of the first transformed flow cytometry event data based on a first bin of the first plurality of bins. The first descriptor of the first transformed flow cytometry event data associated with the first bin can be a first bin number of the first bin of the first plurality of bins. The computing system can add the first bin, the first descriptor, and / or the first bin number to a memory data structure.
[0078] In some embodiments, two bins of the first plurality of bins have the same size. Each bin of the first plurality of bins can have the same size. Two bins of the first plurality of bins can have different sizes. Two bins of the first plurality of bins can include approximately the same number of transformed flow cytometry event data. Each of the first plurality of bins can include approximately the same number of transformed flow cytometry event data. The computing system can determine a size of each of the first plurality of bins. The processor can be programmed by the executable instructions to determine a size of each of the first plurality of bins based on a plurality of gates. The computing system can determine a size of each of the first plurality of bins based on transformed flow cytometry event data associated with a plurality of cells of interest.
[0079] In some embodiments, to transform the first flow cytometry event data, the computing system can transform the first flow cytometry event data using a first dimensionality reduction function. The first dimensionality reduction function can be a linear dimensionality reduction function. The first dimensionality reduction function can be a non-linear dimensionality reduction function. The non-linear dimensionality reduction function can be a t-SNE (Stochastic Neighbor Embedding). The computing system can include first receiving the dimensionality reduction function, or an identification thereof.
[0080] The method 400 proceeds to block 412, where the computing system can convert second flow cytometry event data associated with a second event of the first plurality of events of the flow cytometry event data set in the higher dimensional space to second transformed flow cytometry event data associated with the second event in the first lower dimensional space. The second event can be associated with a positive subsampling requirement. The second transformed flow cytometry event data can be associated with a second bin of the first plurality of bins. In some embodiments, to convert the second flow cytometry event data, the computing system can transform the first flow cytometry event data using a second dimensionality reduction function. The first dimensionality reduction function and the second dimensionality reduction function can be identical.
[0081] In some embodiments, the computing system can receive flow cytometry event data including second flow cytometry event data. The computing system can determine that the second flow cytometry event data of a second event of the first plurality of events is associated with a positive subsampling requirement. The computing system can determine that the second transformed flow cytometry event data is associated with a second bin of the first plurality of bins.
[0082] The processor can be programmed with the executable instructions to determine a second descriptor of the second transformed flow cytometry event data based on a second bin of the first plurality of bins. The second descriptor of the second transformed flow cytometry event data associated with the second bin can be a second bin number of the first bin of the first plurality of bins. The computing system can add the second bin, the second descriptor, and / or the second bin number to the memory data structure.
[0083] The first flow cytometry event data can be associated with a first rare cell and / or the second flow cytometry event data can be associated with a second rare cell. The first rare cell and the second rare cell can be cells of different cell types.
[0084] The method 400 proceeds from block 412 to block 416, where the computing system may determine that a first bin associated with the first transformed flow cytometry event data and a second bin associated with the second transformed flow cytometry event data are different. The computing system may indicate (e.g., in the data structure) that the second flow cytometry event data should be included when generating the subsampled flow cytometry event data.
[0085] In some embodiments, to transform the first flow cytometry event data, the computing system can transform the first flow cytometry event data into first transformed flow cytometry event data associated with the first event in a second, lower dimensional space using a second dimensional reduction function. The second, lower dimensional space can be associated with a second plurality of bins. The first transformed flow cytometry event data in the second, lower dimensional space can be associated with a first bin of the second plurality of bins. To transform the second flow cytometry event data, the computing system can transform the second flow cytometry event data into second transformed flow cytometry event data associated with the second event in a second, lower dimensional space using a second dimensional reduction function. The second transformed flow cytometry event data in the second, lower dimensional space can be associated with a second bin of the second plurality of bins. A first bin of the first plurality of bins may be associated with a first type of target cell, a second bin of the second plurality of bins may be associated with a second type of target cell, a second bin of the first plurality of bins may not be associated with the first type of target cell, a second bin of the first plurality of bins may not be associated with the second type of target cell, a first bin of the second plurality of bins may not be associated with the second type of target cell, and / or a first bin of the second plurality of bins may not be associated with the first type of target cell.
[0086] A first combination of a first bin of the first plurality of bins and a first bin of the second plurality of bins may be associated with a first type of target cell, and / or a second combination of a second bin of the first plurality of bins and a second bin of the second plurality of bins may be associated with a second type of target cell. A first combination of a first bin of the first plurality of bins and a second bin of the second plurality of bins is not associated with the first type of target cell and the second type of target cell, and / or a second combination of a second bin of the first plurality of bins and a first bin of the second plurality of bins is not associated with the first type of target cell and the second type of target cell. The computing system can determine that the first combination and the second combination are different.
[0087] The method 400 proceeds to block 420, where the computer system can generate sub-sampled flow cytometry event data of the flow cytometry event data, including first flow cytometry event data associated with the first event and second flow cytometry event data associated with the second event.
[0088] In some embodiments, the computing system can convert third flow cytometry event data associated with a third event of the first plurality of events of the flow cytometry event data set in the higher dimensional space into third transformed flow cytometry event data associated with a third event in the first lower dimensional space. The third event can be associated with a positive subsampling requirement. The third transformed flow cytometry event data can be associated with a third bin of the first plurality of bins. The processor can be programmed with the executable instructions to determine that the third bin associated with the third transformed flow cytometry event data is the first bin associated with the first transformed flow cytometry event data or the second bin associated with the second transformed flow cytometry event data. The third flow cytometry event data may not be within the subsampled flow cytometry event data of the flow cytometry event data. The computing system can determine a third descriptor of the third transformed flow cytometry event data based on the third bin of the first plurality of bins. A third descriptor of the third transformed flow cytometry event data associated with the third bin may be a third bin number of the third bin of the first plurality of bins. The computing system may determine that the third bin, the third descriptor, and / or the third bin number are not within the memory data structure.
[0089] In some embodiments, the computing system can determine that fourth flow cytometry event data associated with a fourth event of the first plurality of events is associated with a negative subsampling requirement. The computing system can generate a subsampled flow cytometry event data set of flow cytometry event data including the fourth flow cytometry event data associated with the fourth event. The computing system can receive a plurality of gates defining a plurality of cells of interest. The fourth flow cytometry event data can be associated with a cell of interest of the plurality of cells of interest. The fourth flow cytometry event data can be associated with the sorted cells.
[0090] In some embodiments, the computing system can convert second flow cytometry event data associated with a second event of the second plurality of events of the flow cytometry event dataset in the higher dimensional space to second transformed flow cytometry event data associated with a second event of the second plurality of events in the first lower dimensional space. The second event of the second plurality of events can be associated with a positive subsampling requirement. The second transformed flow cytometry event data associated with the second event of the second plurality of events can be associated with a second bin of the first plurality of bins. The second bin associated with the second transformed flow cytometry event data associated with the second event of the second plurality of events and the first bin associated with the first transformed flow cytometry event data associated with the first event of the first plurality of events can be identical. The computing system can generate a subsampled flow cytometry event dataset of flow cytometry event data including the second flow cytometry event data associated with the second event of the second plurality of events.
[0091] The computing system can determine that a last event of the first plurality of events is associated with a time parameter or an event number that exceeds a predetermined threshold. The computing system can reset the memory data structure. The processor can be programmed with the executable instructions to add to the memory data structure a second bin associated with second transformed flow cytometry event data associated with a second event of the second plurality of events. In some embodiments, the computing system can include receiving a degree of subsampling parameter. The computing system can determine the predetermined threshold based on the degree of subsampling parameter.
[0092] The method 400 ends at block 424.
[0093] Execution environment In FIG. 5, the general architecture of an exemplary computing device 500 configured to implement the metabolite, annotation, and gene integration system disclosed herein is shown. The general architecture of the computing device 500 shown in FIG. 5 includes an arrangement of computer hardware and software components. The computing device 500 may include more (or less) elements than those shown in FIG. 5. However, not all of these generally conventional elements need to be shown to provide an enabling disclosure. As illustrated, the computing device 500 includes a processing unit 510, a network interface 520, a computer-readable medium drive 530, an input / output device interface 540, a display 550, and an input device 560, all of which may communicate with each other by a communication bus. The network interface 520 may provide connectivity to one or more networks or computing systems. Thus, the processing unit 510 may receive information and instructions from other computing systems or services via a network. The processing unit 510 may also communicate with a memory 570 and further provide output information for an optional display 550 via the input / output device interface 540. The input / output device interface 540 may also accept input from optional input devices 560, such as a keyboard, a mouse, a digital pen, a microphone, a touch screen, a gesture recognition system, a voice recognition system, a game pad, an accelerometer, a gyroscope, or other input device.
[0094] Memory 570 may include computer program instructions (grouped in some embodiments as modules or components) that processing unit 510 executes to implement one or more embodiments. Memory 570 generally includes RAM, ROM, and / or other persistent, secondary, or non-transitory computer-readable media. Memory 570 may store an operating system 572 that provides computer program instructions for use by processing unit 510 in the general management and operation of computing device 500. Memory 570 may further include computer program instructions and other information for implementing aspects of the present disclosure.
[0095] For example, in one embodiment, memory 570 includes a subsampling module 574 for subsampling particle analysis event data, such as the subsampling method 400 described with reference to Figure 4. Additionally, memory 570 may include or be in communication with a data store 590 and / or one or more other data stores that store the flow cytometry event data set or the generated subsampled flow cytometry event data set.
[0096] term As used herein, the term "determining" or "determining" encompasses a wide variety of actions. For example, "determining" can include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, database, or another data structure), ascertaining, and the like. "Determining" may also include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and the like. "Determining" can also include resolving, selecting, establishing, and the like.
[0097] As used herein, the terms "providing" or "providing" encompass a wide variety of actions. For example, "providing" can include storing a value at a location on a storage device for subsequent retrieval, transmitting a value directly to a recipient via at least one wired or wireless communication medium, transmitting or storing a reference to a value, etc. "Providing" can also include encoding, decoding, encryption, decryption, validation, verification, etc. via hardware elements.
[0098] As used herein, the term "selectively" or "selective" may encompass a wide variety of actions. For example, a "selective" process may include determining one option from a number of options. A "selective" process may include one or more of dynamically determined inputs, pre-configured inputs, or user-initiated inputs to make a decision. In some implementations, n input switches may be included to provide the selective function, where n is the number of inputs used to make a selection.
[0099] As used herein, the term "message" encompasses a wide variety of formats for conveying (e.g., sending or receiving) information. A message may include a machine-readable aggregation of information, such as an XML document, a fixed-field message, a comma-separated message, etc. A message may, in some implementations, include a signal utilized to transmit one or more representations of information. Though enumerated in the singular, it is understood that a message may be composed of, transmitted, stored, and received in multiple parts.
[0100] As used herein, a "user interface" (also referred to as an interactive user interface, graphical user interface, or UI) may refer to a network-based interface that includes data fields, buttons, or other interactive controls for receiving input signals or providing electronic information or providing information to a user in response to any received input signal. The UI may be implemented in whole or in part using technologies such as HyperText Markup Language (HTML), JAVASCRIPT™, FLASH™, JAVA™, NET™, WINDOWS OS™, macOS™, web services, or rich site summary (RSS). In some implementations, the UI may be included in a standalone client (e.g., thick client, fat client) configured to communicate (e.g., send or receive data) according to one or more of the described aspects.
[0101] As used herein, a "data storage" may be embodied in a hard disk drive, solid state memory, and / or any other type of non-transitory computer-readable storage medium accessible from or by a device such as an access device, server, or other computing device described. The data storage may also, or alternatively, be distributed or divided among multiple local and / or remote storage devices, as known in the art, without departing from the scope of this disclosure. In yet other embodiments, the data storage may include or be embodied in a data storage web service.
[0102] Those skilled in the art will appreciate that information, messages, and signals may be represented using any of a variety of different technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.
[0103] Those skilled in the art will further appreciate that the various exemplary logic blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various exemplary components, blocks, modules, circuits, and steps are described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the invention.
[0104] The techniques described herein may be implemented in hardware, software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as specifically programmed event processing computers, wireless communication devices, or integrated circuit devices. Any features described as modules or components may be implemented together in an integrated logic device, or separately as separate but interoperable logic devices. When implemented in software, these techniques may be realized at least in part by a computer-readable data storage medium that includes program code that includes instructions that, when executed, perform one or more of the above-described methods. The computer-readable data storage medium may form part of a computer program product, which may include packaging material. The computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM), such as a synchronous dynamic random access memory (SDRAM), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, or the like. The computer-readable medium may be a non-transitory storage medium. These techniques may additionally, or alternatively, be realized, at least in part, by a computer-readable communications medium that carries or communicates program code in the form of instructions or data structures and can be accessed, read, and / or executed by a computing device, such as a propagated signal or wave.
[0105] The program code may be executed by a specifically programmed sorting strategy processor, which may include one or more processors, such as one or more digital signal processors (DSPs), configurable microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such graphics processors may be specifically configured to perform any of the techniques described in this disclosure. A combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration in at least partial data connection, may implement one or more of the described features. In some aspects, the functionality described herein may be provided in a dedicated software or hardware module configured for encoding and decoding, or may be incorporated into a dedicated sorting control card.
[0106] In general, it will be understood by those skilled in the art that the terms used herein, and particularly in the appended claims (e.g., the body of the appended claims), are generally intended as "open" terms (e.g., the term "including" should be interpreted as "including, but not limited to," the term "having" should be interpreted as "having at least," and the term "includes" should be interpreted as "including, but not limited to"). It will be further understood by those skilled in the art that where a particular number of introduced claim recitations are intended, such intent will be expressly recited in the claims, and in the absence of such recitation, no such intent exists. For example, as an aid to understanding, the following appended claims include the use of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, even if the same claim includes the introductory phrase "one or more" or "at least one" and an indefinite article such as "a" or "an," the use of such phrases should not be interpreted as meaning that the introduction of a claim recitation with the indefinite article "a" or "an" limits any particular claim that includes such an introduced claim recitation to an embodiment that includes only one such recitation (e.g., "a" and / or "an" shall be interpreted as meaning "at least one" or "one or more"). The same applies to the use of definite articles used to introduce claim recitations. In addition, even if a specific number of introduced claim recitations are explicitly recited, one of ordinary skill in the art will recognize that such recitation should be interpreted as meaning at least the number recited (e.g., a bare recitation of "two recitations" without any other modifiers means at least two recitations, or more than two recitations).Furthermore, when a convention similar to "at least one of A, B, and C, etc." is used, such a configuration is generally intended in the sense that one of ordinary skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Furthermore, when a convention similar to "at least one of A, B, or C, etc." is used, such a configuration is generally intended in the sense that one of ordinary skill in the art would understand the convention (e.g., "a system having at least one of A, B, or C" includes, but is not limited to, systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those skilled in the art that substantially any disjunctive word and / or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" is understood to include the possibilities of "A" or "B" or "A and B."
[0107] Additionally, when features or aspects of the disclosure are described in terms of a Markush group, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual members or subgroups of members of the Markush group.
[0108] As will be appreciated by those skilled in the art, for any and all purposes, including in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges, as well as combinations of those subranges. Any range listed can be readily recognized as fully descriptive and allowing for the same range to be divided into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily divided into lower third ranges, intermediate third ranges, and upper third ranges, etc. As will also be appreciated by those skilled in the art, all language such as "up to," "at least," "greater than," "less than," etc. refers to a range that includes the numbers recited and can be divided into subranges as described above. Finally, as will be appreciated by those skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 items refers to a group having 1, 2, or 3 items. Similarly, a group having 1-5 items refers to a group having 1, 2, 3, 4, or 5 items, etc.
[0109] The methods disclosed herein include one or more steps or actions for achieving the described method. The method steps and / or actions may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims.
[0110] While various aspects and embodiments are disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting. The true scope and spirit are indicated by the following claims.
Claims
1. 1. A method for subsampling flow cytometry event data, comprising: Under the control of the processor, transforming first flow cytometry event data associated with a first event of a first plurality of events of a flow cytometry event dataset in a higher dimensional space into first transformed flow cytometry event data associated with the first event in a first lower dimensional space, wherein the first event is associated with a positive subsampling requirement indicating that event data should be subsampled, the first lower dimensional space is associated with a first plurality of bins, and the first transformed flow cytometry event data is associated with a first bin of the first plurality of bins; transforming second flow cytometry event data associated with a second event of the first plurality of events of the flow cytometry event dataset in the higher dimensional space into second transformed flow cytometry event data associated with the second event in the first lower dimensional space, wherein the second event is associated with the positive subsampling requirement and the second transformed flow cytometry event data is associated with a second bin of the first plurality of bins; determining that the first bin associated with the first transformed flow cytometry event data and the second bin associated with the second transformed flow cytometry event data are different; generating a sub-sampled flow cytometry event data set of the flow cytometry event data, the sub-sampled flow cytometry event data including the first flow cytometry event data associated with the first event and the second flow cytometry event data associated with the second event; A method comprising:
2. The method of claim 1, comprising receiving flow cytometry event data including the first flow cytometry event data and the second flow cytometry event data. determining that the first flow cytometry event data for the first event of the first plurality of events is associated with the positive subsampling requirement; and determining that the second flow cytometry event data of the second event of the first plurality of events is associated with the positive sub-sampling requirement; and 3. The method of claim 1 or 2, comprising: determining that the first transformed flow cytometry event data is associated with the first bin of the first plurality of bins; determining that the second transformed flow cytometry event data is associated with the second bin of the first plurality of bins; 3. The method of claim 1 or 2, comprising: determining a first descriptor of the first transformed flow cytometry event data based on the first bin of the first plurality of bins; and determining a second descriptor of the second transformed flow cytometry event data based on the second bin of the first plurality of bins; 3. The method of claim 1 or 2, comprising:
6. The method described in claim 5, wherein the first descriptor of the first converted flow cytometry event data associated with the first bin is a first bin number of the first bin of the first plurality of bins, and the second descriptor of the second converted flow cytometry event data associated with the second bin is a second bin number of the first bin of the first plurality of bins.
7. The method described in claim 1 or 2, wherein the first flow cytometry event data is associated with a first rare cell and / or the second flow cytometry event data is associated with a second rare cell, and optionally, the first rare cell and the second rare cell are cells of different cell types. adding the first bin, the first descriptor, and / or the first bin number to a memory data structure; adding the second bin, the second descriptor, and / or the second bin number to the memory data structure; 3. The method of claim 1 or 2, comprising:
9. Transforming third flow cytometry event data associated with a third event of the first plurality of events of the flow cytometry event data set in the higher dimensional space into third transformed flow cytometry event data associated with the third event in the first lower dimensional space, wherein the third event is associated with the positive subsampling requirement and the third transformed flow cytometry event data is associated with a third bin of the first plurality of bins; determining that the third bin associated with the third transformed flow cytometry event data is the first bin associated with the first transformed flow cytometry event data or the second bin associated with the second transformed flow cytometry event data, wherein the third flow cytometry event data is not included in the sub-sampled flow cytometry event data of the flow cytometry event data; 3. The method of claim 1 or 2, comprising:
10. The method of claim 9, comprising determining that the third bin, third descriptor, and / or third bin number is not in a memory data structure.
11. The method described in claim 1 or 2, wherein the fourth flow cytometry event data is associated with the selected cells.
12. Transforming second flow cytometry event data associated with a second event of a second plurality of events of the flow cytometry event dataset in the higher dimensional space into second transformed flow cytometry event data associated with the second event of the second plurality of events in the first lower dimensional space, wherein the second event of the second plurality of events is associated with the positive subsampling requirement, the second transformed flow cytometry event data associated with the second event of the second plurality of events is associated with a second bin of the first plurality of bins, and wherein the second bin associated with the second transformed flow cytometry event data associated with the second event of the second plurality of events and the first bin associated with the first transformed flow cytometry event data associated with the first event of the first plurality of events are identical; generating the sub-sampled flow cytometry event dataset of the flow cytometry event dataset, the sub-sampled flow cytometry event dataset including the second flow cytometry event data associated with the second event of the second plurality of events; 3. The method according to claim 1 or 2.
13. A method of implementing a non-transitory memory configured to store executable instructions; a processor in communication with the non-transitory memory; Including, The executable instructions cause the processor to: transforming first flow cytometry event data associated with a first event of a first plurality of events of a flow cytometry event dataset in a higher dimensional space into first transformed flow cytometry event data associated with the first event in a first lower dimensional space, wherein the first event is associated with a positive subsampling requirement indicating that event data should be subsampled, the first lower dimensional space is associated with a first plurality of bins, and the first transformed flow cytometry event data is associated with a first bin of the first plurality of bins; transforming second flow cytometry event data associated with a second event of the first plurality of events in the higher dimensional space into second transformed flow cytometry event data associated with the second event of the flow cytometry event dataset in the first lower dimensional space, wherein the second event is associated with the positive subsampling requirement and the second transformed flow cytometry event data is associated with a second bin of the first plurality of bins; determining that the first bin associated with the first transformed flow cytometry event data and the second bin associated with the second transformed flow cytometry event data are different; generating a sub-sampled flow cytometry event data set of flow cytometry event data, the sub-sampled flow cytometry event data including the first flow cytometry event data associated with the first event and the second flow cytometry event data associated with the second event; It is programmed to A computing system for subsampling flow cytometry event data.