Adaptive Data Subsampling and Computation
By adaptively subsampling flow cytometry data based on performance criteria, the method reduces analysis computation time and maintains accurate representation of the dataset, addressing the challenges of analyzing large multi-parameter datasets in biological experiments.
Patent Information
- Application Number
- JP2022544381
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-01-24
- Filing Date
- 2021-01-08
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2041-01-08
AI Technical Summary
Analyzing large multi-parameter datasets from biological experiments, such as flow cytometry data, is computationally expensive and time-consuming, especially when recalculations are required during iterative analysis development.
The method involves adaptively subsampling flow cytometry data based on a performance criterion, selecting a subsample of event data according to a determined subsampling ratio, performing an analysis on the subsample, and providing an indication of the analysis results.
This approach significantly reduces analysis computation time while ensuring that the results accurately represent the original dataset, facilitating real-time analysis and iterative development without excessive time delays.
Smart Images

Figure 0007691430000001 
Figure 0007691430000002 
Figure 0007691430000003
Abstract
Description
Technical Field
[0001] Cross-reference This application claims the benefit of priority of U.S. Application No. 16 / 751,983, filed Jan. 24, 2020, which is incorporated herein by reference in its entirety.
Background Art
[0002] Various biological experiments involve the analysis of a very large number of samples, each of which can be associated with several parameters or other information generated through the measurement or evaluation of the sample. Such samples may include cells or other biological contents, and each sample may differ with respect to growth medium (e.g., hormones, cytokines, drugs, or other substances in the growth medium), source (e.g., cultured from natural tissue, biopsied, or transplanted by other means), incubation conditions (e.g., temperature, pH, light level or spectrum, ionizing radiation), exposure to viruses, bacteria, or other microorganisms, or any other controlled conditions for observing the reaction of the cells or other biological contents to the applied conditions. This can be done, for example, to evaluate the response of the sample to a putative therapy, to elucidate some biological process, or to investigate any other question of interest.
[0003] Evaluating such samples can involve conducting a variety of different investigations. In some examples, cells from the sample can be counted, identified, and / or sorted via flow cytometry. Additionally or alternatively, the sample can be imaged to evaluate the morphology or other characteristics of the sample at one or more time points. Imaging the entire sample can enable the sample to be evaluated at multiple time points without significantly disturbing the development of the sample. To facilitate such evaluations (e.g., to enable the analysis and / or visualization of proteins or other contents of interest within the sample, to identify or sort cells within the sample, etc.), fluorescent dyes or other substances can be added to the sample.
[0004] Performing an analysis on a large multi-parameter dataset such as those described above can be computationally expensive. Thus, it can take a long time to complete such an analysis. However, this long time can be disadvantageous, particularly when the analysis is recalculated many times as part of the iterative development of the analysis structure and parameters. In such examples, it can be beneficial to reduce the size of the dataset (e.g., by deleting data corresponding to a set of wells in a multi-well plate) to reduce the time required to complete the analysis. However, existing methods for dataset reduction often result in analysis results that do not represent the results of the analysis of the non-reduced dataset. Summary of the Invention Means for Solving the Problems
[0005] One aspect of the present disclosure relates to a method for adaptively subsampling flow cytometry data to reduce analysis computation time, the method comprising: (i) receiving flow cytometry data during a first time period, the received flow cytometry data including event data for a plurality of flow cytometry events; (ii) determining a data subsampling ratio based on a performance criterion; (iii) selecting a subsample of the event data from the flow cytometry data based on the subsampling ratio, such that the selected subsample of the event data represents a portion of the flow cytometry events corresponding to the data subsampling ratio; (iv) performing an analysis on the subsample of the event data, the step of determining the data subsampling ratio including determining the data subsampling ratio such that the step of performing an analysis on the subsample of the event data meets the performance criterion; and (v) providing an indication of the result of the analysis.
[0006] Another aspect of the present disclosure relates to a method for reducing flow cytometry data analysis computation time, the method comprising: (i) receiving an indication of a data subsampling ratio via a user interface; (ii) receiving flow cytometry data, wherein the received flow cytometry data includes event data of a plurality of flow cytometry events; (iii) selecting a subsample of the event data from the flow cytometry data based on the subsampling ratio, such that the selected subsample of the event data represents a portion of the flow cytometry events corresponding to the data subsampling ratio; (iv) performing a reduced analysis on the subsample of the event data; (v) providing an indication of the result of the reduced analysis via the user interface; (vi) receiving an instruction to perform a full analysis via the user interface; (vii) in response to receiving the instruction to perform a full analysis, performing a full analysis on the flow cytometry data; and (viii) providing an indication of the result of the full analysis via the user interface.
[0007] Yet another aspect of the present disclosure relates to a method for adaptively subsampling multi-parameter data to reduce analysis computation time, the method comprising: (i) receiving, during a first time period, multi-parameter data, the received multi-parameter data including event data for a plurality of events; (ii) determining a data subsampling ratio based on a performance criterion; (iii) selecting, from the multi-parameter data, a subsample of the event data based on the subsampling ratio, such that the selected subsample of the event data represents a portion of the events corresponding to the data subsampling ratio; (iv) performing an analysis on the subsample of the event data, the step of determining the data subsampling ratio including determining the data subsampling ratio such that the step of performing an analysis on the subsample of the event data meets the performance criterion; and (v) providing an indication of the result of the analysis.
[0008] Yet another aspect of the present disclosure relates to a method for reducing multi-parameter data analysis computation time. The method includes: (i) receiving an indication of a data subsampling ratio via a user interface; (ii) receiving multi-parameter data, where the received multi-parameter data includes event data of a plurality of events; (iii) selecting a subsample of the event data from the multi-parameter data based on the subsampling ratio, such that the selected subsample of the event data represents a portion of the events corresponding to the data subsampling ratio; (iv) performing a reduced analysis on the subsample of the event data; (v) providing an indication of the result of the reduced analysis via the user interface; (vi) receiving an instruction to perform a full analysis via the user interface; (vii) in response to receiving the instruction to perform a full analysis, performing a full analysis on the multi-parameter data; and (viii) providing an indication of the result of the full analysis via the user interface.
[0009] Yet another aspect of the present disclosure relates to a computer-readable medium configured to store at least computer-readable instructions that, when executed by one or more processors of a computing device, cause the computing device to perform one or more of the methods described herein. Such a computer-readable medium can be a non-transitory computer-readable medium.
[0010] Yet another aspect of the present disclosure relates to a system including: (i) one or more processors; and (ii) a non-transitory computer-readable medium configured to store at least computer-readable instructions that, when executed by the one or more processors, cause the system to perform one or more of the methods described herein.
[0011] These and other aspects, advantages, and alternatives will become apparent to those of ordinary skill in the art by reading the following detailed description, with reference to the accompanying drawings as appropriate. Further, it is to be understood that the descriptions provided in this summary section and elsewhere in this specification are intended to be illustrative examples and not limiting as to the claimed subject matter.
Brief Description of the Drawings
[0012]
Fig. 1A
Fig. 1B
Fig. 2
Fig. 3
Fig. 4
Fig. 5
Best Mode for Carrying Out the Invention
[0013] Examples of methods and systems are described herein. It should be understood that the terms "exemplary," "example," and "illustrative" are used herein to mean "serving as an example, instance, or illustration." Any embodiment or feature described herein as "exemplary," "example," or "illustrative" should not necessarily be construed as preferred or advantageous over other embodiments or features. Further, the exemplary embodiments described herein are not intended to be limiting. It will be readily understood that some aspects of the disclosed systems and methods can be arranged and combined in a wide variety of different configurations.
[0014] I. Overview Experiments related to biological systems can result in the generation of large amounts of multi-parameter data. For example, data can be generated for each of a number of biological samples, and the data generated for each sample itself can include a number of parameters, images, events, or other information. Such experiments can include other factors and extensive data (e.g., fluorescence images or other images, flow cytometry data) that can vary with respect to a very large number of directly measured and / or derived parameters that can be generated for each of the growth medium, source, cell type, applied drug, or sample, and can include dozens or hundreds of samples. Thus, performing data analysis on such experiments can be expensive in terms of time and / or computational resources.
[0015] In one example, flow cytometry can be performed to evaluate each sample in a set of samples (e.g., the contents of the wells of a multi-well sample container). Cells or other particles can be taken from each sample and evaluated by a flow cytometer to identify and count the cells (or other particles) within the sample. The flow cytometer can output a plurality of parameters regarding forward scattered light and / or side scattered light, absorption, emission, and / or transmission of light at one or more wavelengths, or any other information regarding each cell / particle detected from the sample. The plurality of output parameters can include, for each detected cell / particle or other detected event, pulse amplitude, width, shape, or other parameters regarding each detected wavelength of light and / or each wavelength of light used to illuminate the cell / particle (e.g., multiple wavelengths used to illuminate different fluorophores within a sample having respective different excitation spectra and / or emission spectra, and to detect the light fluoresced from those fluorophores). Such parameters can be directly detected as part of detecting an event (e.g., the light intensity detected at a particular wavelength) or can be derived from such directly detected parameters (e.g., an emission peak determined from multiple detected light intensities, predicted identification information of the detected cell / particle, a ratio of one parameter to another parameter).
[0016] When specifying the analysis of such multi-parameter data, the user may select or deselect some of the parameters (e.g., because a particular parameter does not appear to vary by a factor of interest due to the sample lacking a fluorophore corresponding to the wavelength associated with that parameter), leading to a corresponding increase or decrease in the cost of computing the analysis.
[0017] The computational and time costs of analyzing multi-parameter data generated from experiments (e.g., flow cytometry experiments as described above) can have various adverse effects. For example, an analysis with too high a computational cost can make it difficult or impossible to provide the results of the analysis of experimental data in real time when the data is generated (e.g., when the samples of the experiment are imaged, evaluated via flow cytometry, and / or evaluated in some other way). In some applications, a human user can iteratively develop the analysis by adjusting the input parameters or other configuration data for the analysis in order to explore the data and / or improve the analysis prior to publication or some other use. In such cases, the analysis must be recomputed after such adjustments so that the human user can evaluate the impact of the adjustments and make further adjustments. However, such recomputation can sometimes take a significant amount of time, which can reduce the ability of the human user to intuit the impact of their adjustments or allocate sufficient time to fully develop the final analysis.
[0018] In some applications, it may be desirable to perform a reduced analysis, e.g., perform an analysis on less than the entire available dataset or perform an analysis that is reduced in some other way with respect to computational cost. For example, the analysis can be reduced in a way that reduces the time and / or computational resources required to perform the analysis while also reducing the completeness or accuracy of the output of the analysis. Such a reduced analysis can be performed to provide a rough analysis of experimental data when the experimental data is generated by one or more instruments, e.g., to provide an ongoing analysis almost "in real time".
[0019] In another example, such a reduced analysis can be performed after user adjustments to the composition and / or characteristics of the analysis (e.g., adding or removing inputs to the analysis, adding or removing specific analysis steps or types, adjusting p-values, learning rates, or other analysis parameters) to provide a rough analysis of experimental data. This can be done to provide feedback on the impact of the adjustments without requiring the full analysis to be re-performed every time. The time saved in such an exemplary application by performing the reduced analysis can be significant, especially in cases where the user updates the analysis multiple times. Once the user has finished adjusting the analysis, the analysis can be performed in a non-reduced manner (a "full analysis") to provide results, and the results can then be published or applied in some other way.
[0020] Such a low-cost reduced analysis can be performed by removing some samples, parameters, or other aspects of the input data from the analysis. For example, if the data is generated from samples in different wells of a multi-well plate, data from a set of wells can be removed from the data set to reduce the size of the data set on which the analysis is performed. The time and / or cost of computational resources for performing such an analysis can then be reduced in proportion to the number of samples / parameters omitted from the analysis. However, a reduced analysis performed in this way is also likely to become inaccurate in a biased way. This can be due to the lack of the samples and / or parameters whose analysis was omitted. The output of such a reduced analysis can be particularly inaccurate if the omitted samples / parameters are particularly "important" for the overall output of the analysis, for example, in cases where the omitted samples and / or parameters exhibit behavior that significantly deviates from the behavior of the samples and / or parameters that were not omitted.
[0021] Instead, when the data includes a plurality of individual "events", the event data can be subsampled in order to provide a reduced analysis of the data that is improved compared to omitting data from the entire sample or input parameters. For example, if the input data includes flow cytometry data, each detected cell or other particle can be such an individual event. Thus, the analysis can be reduced with respect to computation time and / or computational cost by performing the analysis using only data from a subset of the events (e.g., data from a subsample of the cells or other particles detected via flow cytometry). By reducing the analysis in this way, the output of the reduced analysis is less likely to be skewed as a result of subsampling of the individual events than if the analysis were reduced by omitting the entire sample, the parameters, or other aspects of the input data. Thus, the output of the analysis reduced via event subsampling is likely to be more accurate compared to an unreduced (or "complete") analysis of the entire input data set.
[0022] The multi-parameter data sets subject to analysis described herein can include various individual events that can be subsampled to reduce the cost of analysis. As described above, the data set can include flow cytometry data, in which case the events can be individual detected cells, particles, or other flow cytometry events. If the data set includes multiple images of a sample (e.g., fluorescence images taken at regular intervals), each image can be an event, enabling that image to be subsampled prior to analysis. Additionally or alternatively, the individual cells or other contents of the sample represented in the image (e.g., cells, bacteria, or other contents identified via automated image segmentation techniques) can represent individual events within the data set that can be subsampled to reduce the cost of performing the analysis. This can include performing the analysis only on a subset of the cells identified in the image of the sample, e.g., based on the image, only determining the concentration or amount of fluorophores contained within a subset of the cells represented in the image. In yet another example, the events can be particle detection events detected by a mass spectrometer, action potentials detected from electrical ones of other activities of muscle or nerve cells within the sample, individual steps or other individual elements of movement or other motion data, or any other individual events present in and / or represented by the multi-parameter data set.
[0023] Events represented in a multi-parameter dataset can be identified during the generation of the dataset or detected by other means. For example, the dataset can include flow cytometry data, and the events can be individual detected cells or other particles passing through the flow cell. Detection of particles flowing through the flow cell can result in the generation of forward scatter light, side scatter light, amplitude, width, or other parameters related to transmitted light and / or fluorescence emission light at one or more wavelengths, or any other parameter related to the particle detection event. Additionally or alternatively, events can be detected within the dataset after dataset generation. For example, the events can include action potentials (detected electrophysiologically via imaging of a calcium-sensitive dye or by some other method) or other transient processes or behaviors observed in the dataset. Performing reduced analysis on such a dataset can include, following identification of such events within the dataset, performing additional analysis (e.g., statistics on the amplitude, width, timing, or other parameters of the action potentials) on only a subset of the events identified within the dataset.
[0024] To perform analysis on a subsampled portion of the events within a multi-parameter dataset, a data subsampling ratio can be determined, and then event data can be subsampled according to the data subsampling ratio. For example, a 4:1 subsampling ratio can be determined and used to select a subsample representing one quarter of the available event data from the available event data.
[0025] Such a data subsampling ratio can be manually selected by a user. For example, a user can select a 2:1 data subsampling ratio to approximately halve the time required to analyze a dataset, while experiencing a corresponding reduction in the accuracy of the output of the reduced analysis. Alternatively, the data subsampling ratio can be adaptively determined to meet performance criteria. For example, the data subsampling ratio can be specified such that analysis can be performed on a portion of the events within the dataset corresponding to the data subsampling ratio for a specified duration less than a specified duration. Such an adaptive determination of the data subsampling ratio can be performed once (e.g., at the beginning of the performance of data analysis), or can be performed repeatedly. The repeated determination of the data subsampling ratio can be performed in the face of changes in the configuration of the analysis (e.g., addition or reduction of parameters included in the analysis by the user), changes in the rate of generation of data being analyzed in a real-time analysis scenario (e.g., an increase in the detection rate of cells by a flow cytometer), changes in the operation of the hardware used to perform the analysis (e.g., failure of a hard drive of another component or an increase in data serving, Internet streaming or communication, or other computational tasks being performed by the server in addition to performing the analysis), or changes in other factors that may affect the calculation of the analysis of the multi-parameter dataset described herein, to enable the performance criteria to be met.
[0026] Such a specified duration itself can be manually selected. For example, the specified duration can be selected according to the user's preference regarding waiting for the results of repeated analysis after adjustment of the analysis specifications. Alternatively, the specified duration can be selected or determined such that the analysis can be repeatedly updated based on newly generated data by a laboratory instrument so as to appear to be updated in real-time with the newly generated data (e.g., the duration can be less than 100 milliseconds).
[0027] The data subsampling ratio can be determined in various ways. To perform an analysis within a specified duration, the data subsampling ratio can be determined based on the number of events in the data set such that it requires less time than the specified duration for performing the analysis. To perform an analysis when the data set is generated (e.g., almost in real time so that the results of the analysis can be provided and updated when additional data is generated), the data subsampling ratio can be determined based on the average generation rate or occurrence rate of events such that it requires less time than the time required to generate the data set for performing the analysis (e.g., within a specified data capture period).
[0028] The data subsampling ratio can be determined based on the expected time and / or computational cost for performing the analysis, e.g., for performing part or all of the analysis for a single event in the event data. For example, the data subsampling ratio can be determined by dividing the target duration by the number of events in the data set and the computational cost per event. Such computational cost for the analysis can be determined in various ways. In some examples, the computational cost can be determined based on past implementations of the analysis and / or parts of the analysis. Additionally or alternatively, the computational cost can be determined based on a function or other algorithm acting on information about the analysis. For example, the computational cost can be determined using a function or algorithm that acts based on information about the number of input parameters for the analysis, identification information or other information about steps or sub-analyses of the analysis, or any other information that can be used to estimate the computational cost of performing the analysis. Such a function or other algorithm can include a bias term or other parameters that can be determined and / or updated based on past implementations of the analysis and / or parts of the analysis.
[0029] Determining a data subsampling ratio based on performance metrics (e.g., based on a specified duration during which an analysis must be completed) can enable the data subsampling ratio to be automatically adjusted to accommodate changes in the analysis (e.g., when parameters are added to or removed from the analysis as the analysis iteratively evolves), changes to the system used to perform the analysis (e.g., changes in available memory in an unrelated process also being run by the system), or other changes related to the performance of the analysis over time.
[0030] For example, a first data subsampling ratio can be determined for a first analysis (e.g., a first analysis that performs a first set of analysis tasks for a first selection of input parameters from a data set), and can be used to select event data for a first subset of events represented in the data set (e.g., a flow cytometry data set). The first data subsampling ratio can be determined based on the details of the first analysis such that the first analysis can be performed on the first subset of events according to performance criteria (e.g., within a specified duration). An update to the first analysis can be received to define a second analysis. For example, a human user can add or subtract one or more available input parameters to the analysis. Then, a second data subsampling ratio can be determined based on the details of the second analysis such that the second analysis can be performed on event data for a second subset of events according to performance criteria (which can be the same as or different from the performance criteria used to determine the first data subsampling ratio), and can be used to select a second subset of events represented in the data set. Thus, the data subsampling ratio used to perform the analysis can be adapted over time such that the apparent performance of the analysis from the perspective of a human user updating the analysis remains substantially the same and / or varies by no more than a specified amount.
[0031] If the performance criteria include analyzing newly generated data as it is generated, the data subsampling ratio can be updated over time based on whether the updated analysis is completed at the same rate as the rate at which the newly generated data is being generated. If the analysis is lagging, the data subsampling ratio can be increased to reduce the number of events being analyzed per unit of time. Conversely, if the analysis is being completed more quickly, the data subsampling ratio can be decreased to increase the portion of newly generated events included in the analysis while still providing the updated analysis substantially in real time.
[0032] Once the data subsampling ratio is determined, data corresponding to the corresponding percentage of events within a data set (e.g., a data set including multiple flow cytometry events) can be selected and analyzed. This selection can be performed in various ways. In some examples, events can be randomly selected. This can include operating a random number generator or other hardware source of randomness to select which events to analyze. Alternatively, events can be randomly selected using a pre-generated pseudo-random number sequence and / or the output of a pseudo-random number generator to select which events to analyze.
[0033] Alternatively, events can be selected according to a predetermined pattern. For example, if the data subsampling ratio is 4:1, one out of every four events within the data set can be selected for analysis (the other three out of the four events are omitted from the analysis). In another example, if the data subsampling ratio is 4:1, the first and second events out of eight events within the data set can be selected for analysis (the other third through eighth events out of the eight events are omitted from the analysis).
[0034] Event data selected according to the data subsampling ratio can be selected from the set of all available events within the dataset. Alternatively, the events can be selected from a subset preselected according to some criteria or other process. For example, a noise gate filter (or other type of filter) can be applied to all of the events within the dataset to determine which of the events are affected by noise more than a threshold or for some other reason (e.g., due to events corresponding to cell types outside the scope of the experiment) and should be omitted from subsequent analysis. Applying such a noise gate filter can include determining some other parameter related to the signal power, signal-to-noise ratio, maximum and / or minimum signal values, or the presence and magnitude of noise present in the event data. Then, applying the noise gate filter can include and can determine whether to include the event data in subsequent analysis or to omit / discard the event data based on the determined parameter. The data subsampling ratio can then be used to select a subsample of the events already selected according to the noise gate filter and / or according to some other preselection process.
[0035] The data analysis described herein can be reduced by additional or alternative ways for subsampling events as described above. As described above, performing such reduced analysis can allow the accuracy of the analysis to be traded for a more rapid performance of the analysis or a lower computational cost. This can be a beneficial tradeoff when attempting to provide analysis in real time as source data is generated, when repeatedly recalculating the analysis while iteratively evolving the analysis, or in other situations. Such alternative ways for reducing the computational cost of the analysis can be implemented in addition to or as an alternative to adaptive data subsampling. For example, a static data subsampling rate may be received from a user and used to subsample event data, which is then applied to a reduced analysis in one of the alternative analysis reduction methods described herein. After one or more reduced analyses are performed (e.g., during iterative evolution of a finalized analysis configuration), an unreduced analysis can be performed on an unsampled set of event data to generate an analysis output that can be published or applied in some other way.
[0036] Such reduced analysis can be reduced by omitting some data preprocessing steps or sub-analyses compared to non-reduced analysis. For example, event data can include spectral information (e.g., for detected cells, particles, or other flow cytometry events) that can include crosstalk between different spectral channels. Such crosstalk can be due to overlap between the excitation spectra and / or emission spectra of different fluorophores (e.g., dyes incorporated into the sample) present in the test sample. In non-reduced analysis, such crosstalk can be removed or reduced by applying spectral compensation to the spectra (e.g., based on the observed level of crosstalk between channels, mixing coefficients, etc.). Reduced analysis can omit such spectral compensation steps. Other types of compensation steps, normalization steps, or other preprocessing steps can be omitted when performing reduced analysis compared to performing the corresponding non-reduced analysis.
[0037] In another example, the un-reduced analysis can include performing a process for detecting noisy data to determine whether each event, sample, well, or other sub-section of the multi-parameter dataset should be omitted from further processing. Such samples, wells, or other sub-sections of the multi-parameter dataset can be omitted due to an omitted portion that represents an artifact (e.g., a bubble rather than an interesting cell or other particle passing through a flow cytometer), out-of-range data, data that is irrelevant or otherwise not worthy of analysis, wells or samples that failed to incubate properly and / or were infected with an external pathogen, or due to some other factor or consideration. The process for detecting such noisy data can be computationally expensive to perform compared to simply performing the remainder of the reduced analysis on all of the data (including the events or other aspects of the data that would have been omitted). Thus, the reduced analysis can omit the process for detecting such noisy data.
[0038] Reduced analysis can be reduced by omitting the determination of summary statistics about data parameters compared to non-reduced analysis. Such summary statistics can be computationally expensive to determine (e.g., due to acting on the entire dataset, including sorting, summing, or other computationally expensive steps, or due to some other factor). Reduced analysis can omit determining the mean, median, average, standard deviation, third or higher-order cumulants, entropy, divergence, fitting a parametric distribution (e.g., Gaussian, exponential, binomial, logarithmic, Poisson, or some other discrete or continuous distribution), or any other statistical information about parameters determined from event data as part of the reduced analysis. For example, reduced analysis can include determining bin boundaries for a parameter, the count of the parameter within each bin, and a histogram of the parameter for specific parameters of subsampled event data. The corresponding non-reduced analysis can include these steps, as well as determining the mean, median, standard deviation, and fitting a Gaussian distribution of the parameter.
[0039] Reduced analysis can be reduced by omitting some complex analysis decisions as compared to non-reduced analysis. For example, reduced analysis can omit decisions such as plate maps, heat maps, dose response curves, complex user-defined numerical calculations, fitting of complex models to datasets, and / or any other analysis. Additionally or alternatively, reduced analysis can be reduced by performing certain analysis at a reduced level, for example, at a reduced resolution, as compared to non-reduced analysis. For example, reduced analysis can generate a histogram having fewer bins and / or a wider bin width as compared to a histogram determined as part of a non-reduced (or “full”) analysis. In another example, reduced analysis can perform an iterative analysis process (e.g., gradient descent fitting of model parameters to a dataset) using fewer iterations and / or until the change in model parameters or output accuracy from iteration to iteration varies by less than an end threshold value used to terminate the iterative analysis process when performed as part of a non-reduced analysis.
[0040] II. Exemplary Systems To implement the various embodiments described herein, various systems can be used (e.g., programmed). Such systems can include a desktop computer, laptop computer, tablet, or other single-user workstation. Additionally or alternatively, the embodiments described herein can be implemented by a server, cloud computing environment, or other multi-user system.
[0041] Such a system can analyze data received from other systems, such as data received from remote data storage on a server, from a remote cell counter or other instrument, or from some other source. Additionally or alternatively, a system configured to implement the embodiments described herein can include and / or be coupled to an automated incubator, a sample imaging system, a cell sorter, a flow cell, or other elements of a flow cytometry device, or any other instrument capable of generating experimental data for analysis. For example, such an instrument can include an incubator that includes multi-well sample containers. Samples within such multi-well sample containers can vary with respect to the sample's genome, the source of the sample, the growth medium applied to the sample, the agents applied to the sample, biologics or microorganisms, or any other conditions applied to the sample.
[0042] Samples within such an apparatus can be experimentally evaluated in a variety of ways. The sample can be imaged (e.g., using visible light, infrared light, and / or ultraviolet light). Such imaging can include fluorescence imaging of the contents of the sample, e.g., imaging a fluorescent dye or fluorescent reporter added to the sample and / or produced by the cells of the sample (e.g., after insertion of a gene encoding a fluorophore). Additionally or alternatively, materials (e.g., cells, growth media) can be extracted from the sample for analysis via chromatography, mass spectrometry, cell counting or other flow cytometry methods, or some other method. An automated gantry can be placed within such an incubator to facilitate imaging of various samples within respective wells of a sample container (e.g., by operating a pipette or other device configured to selectively extract cells or other contents from a designated well of the sample container), to facilitate cell counting or other flow cytometry analysis of various samples within respective wells of a sample container, or to facilitate measurement and analysis of various samples within respective wells of a sample container.
[0043] FIG. 1A shows an exemplary flow cytometry apparatus 100 for use in connection with a well plate 110 or other various multi-well sample containers. The flow cytometry apparatus 100 can be disposed, in whole or in part, within an incubator to facilitate control of the temperature or other environmental parameters applied to the samples within the multi-well sample container. The flow cytometry apparatus 100 includes an autosampler 102 having an adjustable arm 101 to which a hollow probe 106 is attached. As the arm 104 moves back and forth (left and right in FIG. 1A) and left and right (front and back in the plane of FIG. 1A), the probe 106 is lowered into the individual source wells 108 of the well plate 110 to acquire a sample containing particles (which can be tagged with a fluorescent tag (not shown in FIG. 1A)) to be analyzed using the flow cytometry apparatus 100. While taking in sample material from each of the source wells 108, the probe 106 can be operated to take in a fixed amount of a stripping fluid (such as air), thereby forming stripping bubbles between successive samples in the fluid flow stream.
[0044] When the sample is lifted by the probe 106, the sample is taken into the fluid flow stream, and a pump 112 (e.g., a peristaltic pump) passes the sample through a conduit 114 extending from the autosampler 102 through the pump 112 to a flow cytometer 116 that includes a flow cell 118 and a laser irradiation device 120. The flow cell 118 focuses the fluid flow stream and can be continuously operated to analyze the particles (e.g., cells) in each of the plurality of samples as the fluid flow stream passes through the flow cytometer. The laser irradiation device 120 inspects the individual samples flowing from the flow cell 118 at a laser irradiation point 122.
[0045] Figure 1B shows a series of samples 130, 132, and 134 separated from each other by detachment bubbles 136 and 138 in conduit 114 that form a fluid flow stream peeled off by bubbles. In Figure 1B, sample 130 is adjacent to sample 132, and sample 132 is adjacent to sample 134. When samples 130, 132, and 134 pass through laser irradiation point 122, the particles in the samples are sensed by flow cytometer 116. Forward scattered light is detected by forward scatter detector 124. Fluorescence emitted from tagged particles in the flow cell is detected by fluorescence detector 126. Side scattered light can also be detected (e.g., by a side scatter detector not shown). In contrast, when bubbles 136 and 138 pass through laser irradiation point 122, no particles are sensed. Thus, a graph of the sensed fluorescence data points versus the time at which a series of samples were analyzed using a flow cytometer forms distinct groups that each coincide with the time at which a sample containing particles passes through the laser irradiation point. Such a graph can be generated by the output of both forward scatter detector 124, fluorescence detector 126, and / or other sensors entrained on laser irradiation point 122.
[0046] Additionally or alternatively, an automated imaging system can be used to automatically acquire images (e.g., fluorescence activity images) of a plurality of biological samples in respective wells of a sample container during a plurality of different scan periods over time. A set of images of each sample is taken by the automated imaging system during each of the scan periods, and for example, a set of images can be taken at a rate of three images per second over a three-minute scan period. The images can then be analyzed to determine some information about the samples, for example, according to the methods described herein.
[0047] The use of such an automated imaging system can not only enhance consistency with respect to the timing, positioning, and image parameters of the generated images compared to manual imaging, but also significantly reduce the labor costs of imaging biological samples. Further, such an automated imaging system can be configured to operate within an incubator, eliminating the need to remove samples from the incubator for imaging. Thus, the growth environment for the samples can be maintained more consistently. Additionally, when the automated imaging system acts to move the microscope or other imaging device relative to the sample container (instead of moving the sample container to be imaged by, for example, a static imaging device), perturbations associated with the movement of the sample can be reduced. This can improve the growth and development of the sample and reduce movement-related perturbations.
[0048] Such an automated imaging system can be operative to acquire one or more images during a scan that is segmented over a period of more than 24 hours, more than 3 days, more than 30 days, or some other longer time period. The scan can be specified to occur at a designated rate, for example, once a day, more than twice a day, more than three times a day, or more than four times a day. The scan can be specified such that at least two, at least three, or some greater number of scans occur within a 24-hour period. In some examples, data from one or more scans is analyzed (e.g., according to the methods described herein) and used to determine the timing of additional scans (e.g., to increase the speed, duration, image capture rate, or some other characteristic of the scan to detect the occurrence of an individual event predicted to occur within the sample).
[0049] The use of such an automated imaging system can facilitate imaging of the same biological sample at multiple time points over a long time period. Thus, the development and / or behavior of individual cells and / or cell networks can be analyzed over time. For example, a set of cells, a portion of a cell, or other active objects can be identified within a single sample within scans performed over time periods with different wide intervals. These sets of identified objects can then be compared between scans to identify the same active objects across the scans. Thus, the behavior of individual cells or portions of cells can be tracked and analyzed over hours, days, weeks, or months.
[0050] Figure 2 shows the elements of such an automated imaging system 200. The automated imaging system 200 includes a frame 210 to which other elements of the automated imaging system 200 are attached. The frame 210 can be configured (e.g., sized) to fit within an incubator. The automated imaging system 200 includes a sample container 220 removably disposed within a sample container tray 230 coupled to the frame 210. The sample container tray 230 can be removable and / or can include a removable insert to facilitate holding various different sample containers (e.g., various industry standard sample containers). The system 200 additionally includes an actuating gantry 250 configured to position an imaging device 240 relative to the sample container 220 such that the imaging device 240 can operate to acquire images of the contents of individual wells (e.g., exemplary well 225) of the sample container 220.
[0051] The imaging device 240 can include a microscope, a fluorescence imager, a two-photon imaging system, a phase contrast imaging system, one or more illumination sources, one or more optical filters, and / or other elements configured to facilitate imaging of a sample contained within the sample container 220. In some examples, the imaging device 240 includes elements disposed on both sides of the sample container 220 (e.g., a source of coherent, polarized, monochromatic, or otherwise specified illumination light to facilitate phase contrast imaging of a biological sample). In such examples, the elements on both sides of the sample container 220 may be coupled to different gantries, to the same gantry, and / or the elements on one side of the sample container 220 may not be movable relative to the sample container 220.
[0052] The actuating gantry 250 is coupled to the frame 210 and the imaging device 240 and is configured to control the position of the device 240 in at least two directions relative to the sample container 220 to facilitate imaging of a plurality of different samples within the sample container 220. The actuating gantry 250 may also be configured to control the position of the imaging device 240 in a third direction toward and away from the sample container 220 to facilitate controlling the focus of an image acquired using the imaging device 240 and / or to control the depth of material within the sample container 220 that can be imaged using the imaging device 240. Additionally or alternatively, the imaging device 240 may include one or more actuators for controlling the focal length of the imaging device 240. The imaging device 240 can include one or more motors, piezo elements, liquid lenses, or other actuators to facilitate controlling the focus setting of the imaging device 240. For example, the imaging device 240 can include an actuator configured to control the distance between the imaging device 240 and the sample being imaged. This can be done to ensure that an image is taken in focus and / or to enable images to be taken at various different focus settings to facilitate an image correction method.
[0053] The actuating gantry 250 may include elements configured to facilitate detection of the absolute and / or relative position of the imaging device 240 with respect to the sample container 220 (e.g., with respect to a particular well of the sample container 220). For example, the actuating gantry 250 may include an encoder, limit switches, and / or other position sensing elements. Additionally or alternatively, the imaging device 240 or other elements of the system may be configured to detect fiducial marks or other features of the sample container 220 and / or the sample container tray 230 to determine the absolute and / or relative position of the imaging device 240 with respect to the sample container 220.
[0054] Computing functions (e.g., functions for operating the actuating gantry 250 and / or the imaging device 240 to image a sample within the sample container 220 during a specified time period, functions for operating the autosampler 102, the flow cytometer 116, or other elements of the flow cytometry device 100, and / or functions for implementing any other method described herein) may be implemented by one or more computing systems. Such computing systems may be integrated into a laboratory instrument system (e.g., 100, 200), associated with such a system (e.g., by being connected via a direct wired or wireless connection, via a local network, and / or via a secure connection over the Internet), and / or may take any other form (e.g., a cloud computing system communicating with an automated imaging system and / or having access to a store of images of biological samples).
[0055] FIG. 3 shows an example of such a computing system 300 that can be used to implement the methods described herein. The exemplary computing system 300 includes a communication interface 302, a user interface 304, a processor 306, one or more sensors 307 (e.g., a photodetector, a camera, or other sensors of a flow cytometry device, a microscope, a mass spectrometer, or any other instrumented laboratory device), and data storage 308, all of which are communicatively linked together by a system bus 310.
[0056] The communication interface 302 may function to enable the computing system 300 to communicate with other devices, access networks, and / or transport networks using analog or digital modulation of electrical, magnetic, electromagnetic, optical, or other signals. Thus, the communication interface may facilitate circuit-switched communication and / or packet-switched communication, such as basic telephone service (POTS) communication and / or Internet Protocol (IP) or other packetized communication. For example, the communication interface 302 may include a chipset and antenna arranged for wireless communication with a wireless access network or access point. Also, the communication interface 302 may take the form of, or include, wireline interfaces such as Ethernet, Universal Serial Bus (USB), or High-Definition Multimedia Interface (HDMI) (registered trademark) ports. The communication interface may also take the form of, or include, wireless interfaces such as WiFi, BLUETOOTH (registered trademark), Global Positioning System (GPS), or Wide Area Wireless Interface (e.g., WiMAX or 3GPP (registered trademark) Long Term Evolution (LTE)). However, other forms of physical layer interfaces and other types of standard communication protocols or proprietary communication protocols may be used on the communication interface 302. Further, the communication interface 302 may comprise a plurality of physical communication interfaces (e.g., a WiFi interface, a BLUETOOTH (registered trademark) interface, and a wide area wireless interface).
[0057] In some embodiments, the communication interface 302 may be operative to enable the computing system 300 to communicate with other devices, remote servers, access networks, and / or transport networks. For example, the communication interface 302 may be operative to send and / or receive instructions for flow cytometry information or some other information regarding one or more biological samples so as to send and / or receive instructions for an image of a biological sample (e.g., a fluorescence-activated image).
[0058] The user interface 304 of such a computing system 300 may be operative to enable the computing system 300 to interact with a user, for example, to receive input from the user and / or to provide output to the user. Accordingly, the user interface 304 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, etc. The user interface 304 may also include one or more output components such as, for example, a display screen, which may be combined with a presence-sensitive panel. The display screen may be based on CRT, LCD, and / or LED technology, or other technologies now known or later developed. The user interface 304 may also be configured to generate an audible output via a speaker, speaker jack, audio output port, audio output device, earphone, and / or other similar devices.
[0059] In some embodiments, the user interface 304 may include a display useful for presenting a user with video or other images (e.g., video of images generated during a particular scan of a particular biological sample). Additionally, the user interface 304 may include one or more buttons, switches, knobs, and / or dials that facilitate the configuration and operation of the computing device. Some or all of these buttons, switches, knobs, and / or dials may be implemented as functions on a touch-sensitive or presence-sensitive panel. The user interface 304 may enable a user to specify the type of sample included within the automated imaging system, specify a schedule for imaging or other evaluation of the sample, specify parameters for image segmentation, event analysis, and / or any other analysis to be performed by the system 300, or input any other commands or parameters for the operation of the automated laboratory system and / or for the analysis of data generated thereby.
[0060] The processor 306 may comprise one or more general-purpose processors, such as a microprocessor, and / or one or more dedicated processors, such as a digital signal processor (DSP), a graphics processing unit (GPU), a floating point unit (FPU), a network processor, a tensor processing unit (TPU), or an application specific integrated circuit (ASIC). In some instances, the dedicated processor may be capable of image processing, image alignment, statistical analysis, filtering, or noise reduction, among other applications or functions. The data storage 308 may include one or more volatile and / or non-volatile storage components, such as magnetic storage, optical storage, flash storage, or organic storage, and may be integrated, in whole or in part, with the processor 306. The data storage 308 may include removable and / or non-removable components.
[0061] Processor 306 may be capable of executing program instructions 318 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 308 to perform the various functions described herein. Thus, data storage 308 may include a non-transitory computer-readable medium storing program instructions that, when executed by computing device 300, cause computing device 300 to perform any of the methods, processes, or functions disclosed herein and / or in the accompanying drawings. As a result of the execution of program instructions 318 by processor 306, processor 306 may use data 312.
[0062]
[0063] By way of example, program instructions 318 may include an operating system 322 (e.g., an operating system kernel, device drivers, and / or other modules) and one or more application programs 320 (e.g., filtering functions, data processing functions, statistical analysis functions, image processing functions, event subsampling functions) installed on computing device 300. Data 312 may include flow cytometry data, microcopy images, or other multi-parameter data including information about a plurality of events (e.g., detected cells or other particles or other flow cytometry events). Application programs 320 may communicate with operating system 322 through one or more application programming interfaces (APIs). These APIs may facilitate, for example, an application program 320 receiving information via communication interface 302, receiving and / or displaying information on user interface 304, and performing a complete or reduced analysis on events within and / or on a sampled subset of events within multi-parameter data 314.
[0064] The application program 320 can take the form of an "app" that can be downloaded to the computing device 300 through one or more online application stores or application markets (e.g., via the communication interface 302). However, the application program can also be installed on the computing device 300 in other ways, such as via a web browser or through a physical interface (e.g., USB port) of the computing device 300.
[0065] In some examples, portions of the methods described herein can be performed by different devices according to an application. For example, different devices of a system may have different amounts of computing resources (e.g., memory, processor cycles) and different information bandwidths for communication between the devices. For example, a first device can be an embedded processor that can operate an operating gantry, an imaging device, a flow cytometry device, or other elements to generate information about a biological sample during and / or over a plurality of different periods. A second device can then receive information (e.g., image information, flow cytometry data, and / or event information) from the first device (e.g., via the Internet, via a dedicated wired link) and perform the processing and analysis methods described herein on the received data. Different portions of the methods described herein can be distributed according to such considerations.
[0066] V. Exemplary Methods FIG. 4 is a flowchart of a method 400 for adaptively subsampling multi-parameter data (e.g., flow cytometry data) to reduce analysis computation time. Such a method can be advantageous in situations where a reduced-accuracy analysis result is acceptable, for example, to enable substantially real-time analysis when data is generated or to reduce analysis time when iteratively specifying contents to define analysis parameters, where a reduced-accuracy result can be completed in less time.
[0067] Method 400 includes, during a first time period, a step (410) of receiving multi-parameter data, where the received multi-parameter data includes event data for a plurality of events. Such multi-parameter data can be flow cytometry data, in which case the events can be flow cytometry events (e.g., individual detected cells or other particles).
[0068] Method 400 additionally includes a step (420) of determining a data subsampling ratio based on performance criteria. Such performance criteria can be a real-time latency criterion, in which case the step of determining a data subsampling ratio based on performance criteria can include determining the data subsampling ratio based on the average event occurrence rate such that the analysis takes less time than the duration of the first time period to perform (e.g., such that the result of the analysis can be determined in real time or near real time). Such performance criteria can be an analysis duration criterion, in which case the step of determining a data subsampling ratio based on performance criteria can include determining the data subsampling ratio based on the number of events in the multi-parameter data such that the analysis can be performed in less than a specified duration.
[0069] Method 400 further includes, as an addition, a step (430) of selecting subsamples of event data from multi-parameter data based on a subsampling ratio, such that the selected subsamples of event data represent a portion of an event corresponding to the data subsampling ratio. This can include a step of selecting events according to a pre-specified pattern (e.g., every other event for a 50% ratio, or some other regular pattern), or a step of randomly selecting events.
[0070] Method 400 further includes a step (440) of performing an analysis on the subsamples of event data, including a step of determining a data subsampling ratio such that the step of performing an analysis on the subsamples of event data meets a performance criterion.
[0071] Method 400 further includes a step (450) of providing an indication of the result of the analysis.
[0072] Method 400 can include additional elements or features.
[0073] FIG. 5 is a flowchart of a method 500 for reducing the analysis calculation time of multi-parameter data (e.g., flow cytometry data). Such a method can be advantageous in situations where a reduced-accuracy analysis result is acceptable, for example, to enable substantially real-time analysis when data is generated, or to reduce analysis time when iteratively specifying contents to define analysis parameters, such that a reduced-accuracy result can be completed in less time.
[0074] Method 500 includes a step (510) of receiving an indication of a data subsampling ratio via a user interface. Method 500 additionally includes a step (520) of receiving multi-parameter data, where the received multi-parameter data includes event data for a plurality of events. Such multi-parameter data can be flow cytometry data, in which case the events can be flow cytometry events (e.g., individual detected cells or other particles).
[0075] Method 500 additionally includes a step (530) of selecting a subsample of the event data from the multi-parameter data based on the subsampling ratio, such that the selected subsample of the event data represents a portion of the events corresponding to the data subsampling ratio. This can include selecting events according to a pre-specified pattern (e.g., every other event in the case of a 50% ratio, or some other regular pattern), or randomly selecting events.
[0076] Method 500 further includes a step (540) of performing a reduced analysis on the subsample of the event data. The step of performing a reduced analysis can include an analysis that omits some data preprocessing steps (e.g., noise removal or filtering steps, false event detection and rejection steps), an analysis that omits the determination of some sample statistics (e.g., mean, median, mode, or other summary statistics for a sample of the events in the data), an analysis that omits the determination of a histogram and / or an analysis that has a reduced resolution and / or fewer bins in the histogram compared to an analysis that is not reduced, or an analysis that performs a reduced analysis in some other way compared to an analysis that is not reduced or a "full" analysis.
[0077] Method 500 further includes a step (550) of providing an indication of the results of the reduced analysis via the user interface.
[0078] Method 500 further includes a step (560) of receiving instructions for performing a complete analysis via a user interface.
[0079] Method 500 further includes a step (570) of performing a complete analysis on multi-parameter data in response to receiving instructions for performing a complete analysis. Such a complete analysis may be "complete" in the sense that it is not reduced compared to a reduced analysis. For example, the step of performing a complete analysis may include steps of performing some data preprocessing steps, steps of determining some sample statistics, steps of determining a histogram, and / or steps of determining a histogram with a higher resolution and / or more bins compared to a reduced analysis, or steps of performing any other analysis omitted from the reduced analysis and / or steps of performing any other analysis in a relatively more complex manner than a reduced version of the analysis performed as part of the reduced analysis.
[0080] Method 500 further includes a step (580) of providing an indication of the results of the complete analysis via a user interface.
[0081] Method 500 can include additional elements or features.
[0082] VI. Conclusion The foregoing detailed description has described various features and functions of the disclosed systems, devices, and methods with reference to the accompanying drawings. In the drawings, like reference numerals generally identify like components unless the context indicates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized and other changes may be made without departing from the scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure generally described herein and illustrated in the drawings can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.
[0083] With respect to any or all of the message flow diagrams, scenarios, and flowcharts shown in the figures and discussed herein, each step, block, and / or communication may represent the processing and / or transmission of information according to an illustrative embodiment. Alternative embodiments are within the scope of these illustrative embodiments. In these alternative embodiments, for example, the functions described as steps, blocks, transmissions, communications, requests, responses, and / or messages may be performed in an order different from the order shown or discussed, including substantially simultaneous or reverse orders, depending on the functionality involved. Further, more or fewer steps, blocks, and / or functions may be used with any of the message flow diagrams, scenarios, and flowcharts discussed herein, and these message flow diagrams, scenarios, and flowcharts may be combined with each other, in part or in whole.
[0084] Steps or blocks representing the processing of information may correspond to circuitry configured to perform specific logical functions of the methods or techniques described herein. Alternatively or additionally, steps or blocks representing the processing of information may correspond to a module, a segment, or a portion of program code (including associated data). The program code may include one or more instructions executable by a processor for implementing specific logical functions or actions in a method or technique. The program code and / or associated data may be stored on any type of computer-readable medium, such as a storage device including a disk drive, a hard drive, or other memory medium.
[0085] Computer-readable media may also include non-transitory computer-readable media such as register memory, processor cache, and / or random access memory (RAM) that stores data for short periods of time. Computer-readable media may also include non-transitory computer-readable media that store program code and / or data for longer periods of time, such as secondary or persistent long-term storage, like read-only memory (ROM), optical or magnetic disks, and / or compact disc read-only memory (CD-ROM). Computer-readable media may also be any other volatile or non-volatile storage system. Computer-readable media may be regarded as, for example, a computer-readable storage medium or a tangible storage device.
[0086] Furthermore, steps or blocks representing one or more information transmissions may correspond to information transmissions between software modules and / or hardware modules within the same physical device. However, other information transmissions may be between software modules and / or hardware modules in different physical devices.
[0087] Although various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for illustrative purposes only and are not intended to be limiting, and the true scope is indicated by the appended claims.
Explanation of Signs
[0088] 100 Flow cytometer 102 Auto sampler 104 Arm 106 Probe 108 Source well 110 Well plate 112 Pump 114 Conduit 116 Flow cytometer 118 Flow cell 120 Laser irradiation device 122 Laser irradiation point 124 Forward scatter detector 126 Fluorescence detector 130 Sample 132 Sample 134 Sample 136 Detached bubble, air bubble 138 Detached bubble, air bubble 200 Automatic imaging system 210 Frame 220 Sample container 225 Well 230 Sample container tray 240 Imaging device 250 Operating gantry 300 Computing system, computing device 302 Communication interface 304 User interface 306 Processor 307 Sensor 308 Data storage 310 System bus 312 Data 314 Multi-parameter data 318 Program instructions 320 Application program 322 Operating system 400 Method 500 Method
Claims
1. A method for adaptively subsampling flow cytometry data to reduce analysis calculation time, comprising: receiving, during a first time period, flow cytometry data, wherein the received flow cytometry data includes event data of a plurality of flow cytometry events; determining a data subsampling ratio based on performance criteria; selecting, from the flow cytometry data, a subsample of the event data based on the subsampling ratio, such that the selected subsample of the event data represents a portion of the flow cytometry events corresponding to the data subsampling ratio; performing an analysis on the subsample of the event data, wherein the step of determining the data subsampling ratio includes determining the data subsampling ratio such that the step of performing the analysis on the subsample of the event data meets the performance criteria; providing an indication of the result of the analysis; and a method comprising the above steps.
2. The method according to claim 1, wherein the step of performing the analysis and the step of providing the indication of the result of the analysis are performed during the first time period.
3. The method according to claim 2, wherein the performance criteria is a real-time latency criteria, and the step of determining the data subsampling ratio based on the performance criteria includes determining the data subsampling ratio based on the average occurrence rate of the flow cytometry events such that the analysis takes less time than the duration of the first time period to perform.
4. The method according to claim 2 or 3, wherein the step of performing the analysis on the subsample of the event data includes updating the analysis a plurality of times during the first time period when the flow cytometry data is received, and the step of providing the indication of the result of the analysis includes providing updated indications a plurality of times during the first time period when the analysis is updated.
5. The performance criterion is an analysis duration criterion, and the step of determining the data subsampling ratio based on the performance criterion includes determining the data subsampling ratio based on the number of the flow cytometry events in the flow cytometry data so that the analysis can be performed in less than a specified duration. The method according to claim 1 or 2.
6. The analysis is characterized by the number of parameters, and the step of determining the data subsampling ratio further includes determining the expected computational cost of the analysis based on the number of parameters. The method according to claim 5.
7. A step of selecting a noise-removed subsample of the event data from the flow cytometry data, as a result of which the selected noise-removed subsample of the event data does not represent a flow cytometry event that meets a specified noise criterion. Further including the step of, The step of selecting the subsample of the event data from the flow cytometry data based on the subsampling ratio includes selecting the subsample of the event data based on the subsampling ratio from the noise-removed subsample of the event data. The method according to any one of claims 1 to 6.
8. The step of selecting the subsample of the event data from the flow cytometry data based on the subsampling rate includes selecting data corresponding to flow cytometry events selected according to a predetermined pattern. The method according to any one of claims 1 to 7.
9. The data subsampling ratio is a first data subsampling ratio, the subsample of the event data is a first subsample of the event data, the analysis is a first analysis, and the method is, Receiving a representation of the first analysis, wherein performing the first analysis on the first subsample of the event data and providing an indication of the result of the first analysis are performed in response to receiving the representation of the first analysis, and performing the first analysis on the first subsample of the event data includes performing the first analysis according to the received representation of the first analysis, a step; Receiving a representation of a second analysis; Selecting a second subsample of the event data from the flow cytometry data based on a second subsampling ratio, such that the selected second subsample of the event data represents a portion of the flow cytometry events corresponding to the second data subsampling ratio, a step; Performing the second analysis on the second subsample of the event data according to the received representation of the second analysis; Providing an indication of the result of the second analysis The method according to any one of 1 to 8, further comprising.
10. Determining the second data subsampling ratio based on the performance criteria, wherein the performance criteria is an analysis duration criterion, and determining the second data subsampling ratio based on the performance criteria includes determining the second data subsampling ratio based on the number of flow cytometry events in the flow cytometry data such that performing the second analysis on the second subsample of the event data can be performed in less than the specified duration, a step The method according to claim 9, further comprising.
11. A method for reducing the analysis calculation time of flow cytometry data, Receiving an indication of a data subsampling ratio via a user interface; Receiving flow cytometry data, wherein the received flow cytometry data includes event data of a plurality of flow cytometry events, a step; A step of selecting a subsample of the event data based on the subsampling ratio from the flow cytometry data, such that as a result, the selected subsample of the event data represents a portion of the flow cytometry events corresponding to the data subsampling ratio. A step of performing a reduced analysis on the subsample of the event data. A step of providing an indication of the result of the reduced analysis via the user interface. A step of receiving an instruction for performing a complete analysis via the user interface. A step of performing the complete analysis on the flow cytometry data in response to receiving the instruction for performing the complete analysis. A step of providing an indication of the result of the complete analysis via the user interface. A method comprising the above steps. Claim 12 The method according to claim 11, wherein the step of performing the complete analysis further includes a step of applying spectral compensation to the flow cytometry data, and the step of performing the reduced analysis does not include the step of applying spectral compensation to the flow cytometry data. Claim 13 A step of selecting a noise-removed subsample of the event data from the flow cytometry data, such that as a result, the selected noise-removed subsample of the event data does not represent flow cytometry events meeting a specified noise criterion. The method further includes this step. The step of selecting the subsample of the event data based on the subsampling ratio from the flow cytometry data includes a step of selecting the subsample of the event data based on the subsampling ratio from the noise-removed subsample of the event data. The step of performing the complete analysis on the flow cytometry data includes a step of performing the complete analysis on the noise-removed subsample of the event data. The method according to claim 11 or 12. Claim 14 A non-transitory computer-readable medium configured to store at least computer-readable instructions that, when executed by one or more processors of a computing device, cause the computing device to perform controller operations, the controller operations comprising: Receiving flow cytometry data during a first time period, the received flow cytometry data including event data of a plurality of flow cytometry events; Determining a data subsampling ratio based on performance criteria; Selecting a subsample of the event data from the flow cytometry data based on the subsampling ratio, such that the selected subsample of the event data represents a portion of the flow cytometry events corresponding to the data subsampling ratio; Performing an analysis on the subsample of the event data, wherein determining the data subsampling ratio includes determining the data subsampling ratio such that performing the analysis on the subsample of the event data meets the performance criteria; Providing an indication of the result of the analysis; A non-transitory computer-readable medium comprising the above.
15. The non-transitory computer-readable medium according to claim 14, wherein performing the analysis and providing the indication of the result of the analysis are performed during the first time period.
16. The non-transitory computer-readable medium according to claim 15, wherein the performance criteria is a real-time latency criteria, and determining the data subsampling ratio based on the performance criteria includes determining the data subsampling ratio based on an average occurrence rate of the flow cytometry events such that performing the analysis requires less time than a duration of the first time period.
17. Performing the analysis on the subsample of the event data includes updating the analysis multiple times during the first time period when the flow cytometry data is received, and providing the indication of the result of the analysis includes providing an updated indication multiple times during the first time period when the analysis is updated. The non-transitory computer-readable medium according to claim 15 or 16.
18. The performance criterion is an analysis duration criterion, and determining the data subsampling ratio based on the performance criterion includes determining the data subsampling ratio based on the number of flow cytometry events in the flow cytometry data so that the analysis can be performed in less than a specified duration. The non-transitory computer-readable medium according to claim 14 or 15.
19. The analysis is characterized by the number of parameters, and determining the data subsampling ratio further includes determining the expected computational cost of the analysis based on the number of parameters. The non-transitory computer-readable medium according to claim 18.
20. The controller operation is selecting a noise-removed subsample of the event data from the flow cytometry data, such that the selected noise-removed subsample of the event data does not represent a flow cytometry event that meets a specified noise criterion, and further includes selecting the subsample of the event data from the flow cytometry data based on the subsampling ratio includes selecting the subsample of the event data based on the subsampling ratio from the noise-removed subsample of the event data. The non-transitory computer-readable medium according to any one of claims 14 to 19.
21. Selecting the subsample of the event data from the flow cytometry data based on the subsampling rate includes selecting data corresponding to flow cytometry events selected according to a predetermined pattern. The non-transitory computer-readable medium according to any one of claims 14 to 20.
22. The data subsampling ratio is a first data subsampling ratio, the subsample of the event data is a first subsample of the event data, the analysis is a first analysis, and the controller operation is receiving a representation of the first analysis, wherein performing the first analysis on the first subsample of the event data and providing an indication of the result of the first analysis are performed in response to receiving the representation of the first analysis, and performing the first analysis on the first subsample of the event data includes performing the first analysis according to the received representation of the first analysis, receiving, receiving a representation of a second analysis, selecting, from the flow cytometry data, a second subsample of the event data based on a second subsampling ratio, such that the selected second subsample of the event data represents a portion of the flow cytometry event corresponding to the second data subsampling ratio, selecting, performing a second analysis on the second subsample of the event data according to the received representation of the second analysis, and providing an indication of the result of the second analysis The non-transitory computer-readable medium according to any one of claims 14 to 21, further comprising. **Claim 23** The controller operation is determining the second data subsampling ratio based on the performance criterion, wherein the performance criterion is an analysis duration criterion, and determining the second data subsampling ratio based on the performance criterion includes determining the second data subsampling ratio based on the number of flow cytometry events in the flow cytometry data such that the second analysis on the second subsample of the event data can be performed in less than a specified duration. determining The non-transitory computer-readable medium according to claim 22, further comprising. **Claim 24** A non-transitory computer-readable medium configured to store at least computer-readable instructions that, when executed by one or more processors of a computing device, cause the computing device to perform controller operations, wherein the controller operations include receiving an indication of a data subsampling ratio via a user interface; receiving flow cytometry data, wherein the received flow cytometry data includes event data of a plurality of flow cytometry events; selecting a subsample of the event data from the flow cytometry data based on the subsampling ratio, such that the selected subsample of the event data represents a portion of the flow cytometry events corresponding to the data subsampling ratio; performing a reduced analysis on the subsample of the event data; providing an indication of the result of the reduced analysis via the user interface; receiving instructions for performing a full analysis via the user interface; performing the full analysis on the flow cytometry data in response to receiving the instructions for performing the full analysis; providing an indication of the result of the full analysis via the user interface comprising a non-transitory computer-readable medium. **Claim 25** The non-transitory computer-readable medium according to claim 24, wherein performing the full analysis additionally includes applying spectral compensation to the flow cytometry data, and performing the reduced analysis does not include applying spectral compensation to the flow cytometry data. **Claim 26** The controller operations further include selecting a noise-removed subsample of the event data from the flow cytometry data, such that the selected noise-removed subsample of the event data does not represent a flow cytometry event that meets a specified noise criterion. Selecting the sub-samples of the event data based on the sub-sampling ratio from the flow cytometry data includes selecting the sub-samples of the event data based on the sub-sampling ratio from the noise-removed sub-samples of the event data. Performing the complete analysis on the flow cytometry data includes performing the complete analysis on the noise-removed sub-samples of the event data. The non-transitory computer-readable medium according to claim 24 or 25. **Claim 27** A system comprising: A controller including one or more processors; A memory storing computer-readable instructions that, when executed by the one or more processors of the controller, cause the controller to perform controller operations, the controller operations including: Receiving flow cytometry data during a first time period, the received flow cytometry data including event data of a plurality of flow cytometry events; Determining a data sub-sampling ratio based on performance criteria; Selecting sub-samples of the event data from the flow cytometry data based on the sub-sampling ratio, such that the selected sub-samples of the event data represent a portion of the flow cytometry events corresponding to the data sub-sampling ratio; Performing an analysis on the sub-samples of the event data, wherein determining the data sub-sampling ratio includes determining the data sub-sampling ratio such that performing the analysis on the sub-samples of the event data meets the performance criteria; Providing an indication of the result of the analysis; And a system. **Claim 28** The system according to claim 27, wherein performing the analysis and providing the indication of the result of the analysis are performed during the first time period. **Claim 29** The performance criterion is a real-time latency criterion, and determining the data subsampling ratio based on the performance criterion includes determining the data subsampling ratio based on an average occurrence rate of the flow cytometry events such that it takes less time than a duration of the first time period to perform the analysis. The system according to claim 28.
30. Performing the analysis on the subsamples of the event data includes updating the analysis a plurality of times during the first time period when the flow cytometry data is received, and providing the indication of the result of the analysis includes providing an updated indication a plurality of times during the first time period when the analysis is updated. The system according to claim 28 or 29.
31. The performance criterion is an analysis duration criterion, and determining the data subsampling ratio based on the performance criterion includes determining the data subsampling ratio based on a number of the flow cytometry events in the flow cytometry data such that the analysis can be performed in less than a specified duration. The system according to claim 27 or 28.
32. The analysis is characterized by a number of parameters, and determining the data subsampling ratio further includes determining an expected computational cost of the analysis based on the number of parameters. The system according to claim 31.
33. The controller operation is selecting a noise-removed subsample of the event data from the flow cytometry data, such that the selected noise-removed subsample of the event data does not represent a flow cytometry event that meets a specified noise criterion, and further includes selecting the subsample of the event data based on the subsampling ratio from the flow cytometry data includes selecting the subsample of the event data based on the subsampling ratio from the noise-removed subsample of the event data. The system according to any one of claims 27 to 32.
34. Selecting the subsamples of the event data based on the subsampling rate from the flow cytometry data includes selecting data corresponding to flow cytometry events selected according to a predetermined pattern. The system according to any one of claims 27 to 33. [
35. ] The data subsampling ratio is a first data subsampling ratio, the subsample of the event data is a first subsample of the event data, the analysis is a first analysis, and the controller operation is Receiving a representation of the first analysis, performing the first analysis on the first subsample of the event data, and providing an indication of the result of the first analysis are performed in response to receiving the representation of the first analysis. Performing the first analysis on the first subsample of the event data includes performing the first analysis according to the received representation of the first analysis. Receiving, Receiving a representation of a second analysis Selecting a second subsample of the event data from the flow cytometry data based on a second subsampling ratio, such that the selected second subsample of the event data represents a portion of the flow cytometry events corresponding to the second data subsampling ratio. Selecting, Performing the second analysis on the second subsample of the event data according to the received representation of the second analysis Providing an indication of the result of the second analysis The system according to any one of claims 27 to 33, further comprising [
36. ] The controller operation is Determining the second data subsampling ratio based on the performance criteria, the performance criteria being an analysis duration criterion, and determining the second data subsampling ratio based on the performance criteria includes determining the second data subsampling ratio based on the number of flow cytometry events in the flow cytometry data such that the second analysis on the second subsample of the event data can be performed in less than the specified duration. Determining The system according to claim 35, further comprising
37. A system comprising: A controller comprising one or more processors; and A memory storing computer-readable instructions that, when executed by the one or more processors of the controller, cause the controller to perform controller operations, the controller operations comprising: Receiving an indication of a data subsampling ratio via a user interface; Receiving flow cytometry data, wherein the received flow cytometry data includes event data of a plurality of flow cytometry events; Selecting subsamples of the event data based on the subsampling ratio from the flow cytometry data, such that the selected subsamples of the event data represent a portion of the flow cytometry events corresponding to the data subsampling ratio; Performing a reduced analysis on the subsamples of the event data; Providing an indication of the result of the reduced analysis via the user interface; Receiving instructions for performing a full analysis via the user interface; Performing the full analysis on the flow cytometry data in response to receiving the instructions for performing the full analysis; and Providing an indication of the result of the full analysis via the user interface. A system comprising
38. The system according to claim 37, wherein performing the full analysis additionally includes applying spectral compensation to the flow cytometry data, and performing the reduced analysis does not include applying spectral compensation to the flow cytometry data.
39. The controller operations further include: Selecting noise-reduced subsamples of the event data from the flow cytometry data, such that the selected noise-reduced subsamples of the event data do not represent flow cytometry events that meet a specified noise criterion. Selecting the sub-samples of the event data based on the sub-sampling ratio from the flow cytometry data includes selecting the sub-samples of the event data based on the sub-sampling ratio from the noise-removed sub-samples of the event data, Performing the complete analysis on the flow cytometry data includes performing the complete analysis on the noise-removed sub-samples of the event data, The system according to claim 37 or 38.
40. A method for adaptively sub-sampling multi-parameter data to reduce analysis calculation time, Receiving multi-parameter data during a first time period, wherein the received multi-parameter data includes event data of a plurality of events, Determining a data sub-sampling ratio based on a performance criterion, Selecting sub-samples of the event data from the multi-parameter data based on the sub-sampling ratio, such that the selected sub-samples of the event data represent a portion of the events corresponding to the data sub-sampling ratio, Performing an analysis on the sub-samples of the event data, wherein determining the data sub-sampling ratio includes determining the data sub-sampling ratio such that performing the analysis on the sub-samples of the event data meets the performance criterion, Providing an indication of the result of the analysis A method comprising.
41. A method for reducing multi-parameter data analysis calculation time, Receiving an indication of a data sub-sampling ratio via a user interface, Receiving multi-parameter data, wherein the received multi-parameter data includes event data of a plurality of events, Selecting sub-samples of the event data from the multi-parameter data based on the sub-sampling ratio, such that the selected sub-samples of the event data represent a portion of the events corresponding to the data sub-sampling ratio, Performing a reduced analysis on the subsample of the event data; Providing an indication of the result of the reduced analysis via the user interface; Receiving an instruction to perform a full analysis via the user interface; Performing the full analysis on the multi-parameter data in response to receiving the instruction to perform the full analysis; Providing an indication of the result of the full analysis via the user interface A method comprising.
Citation Information
Patent Citations
Big-data clustering algorithm based on cloud computing platform
CN103838863A
Big data calculation method and system, program and recording medium
JP2018521391A
Methods for using artificial neural network analysis on flow cytometry data for cancer diagnosis
US20180247715A1