Efficient cross-platform data synchronization and analysis
Patent Information
- Application Number
- PCT/US2026/016607
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-27
- Filing Date
- 2026-02-25
- Publication Date
- 2026-09-03
Smart Images

Figure 00000044_0000 
Figure 00000044_0001 
Figure 00000045_0000
Abstract
Description
MBHB No. 24-2390- WO Efficient Cross-Platform Data Synchronization and AnalysisCross-Reference to Related Application
[0001] This application calls priority' to U. S. provisional application no. 63 / 764,349, filed Feb. 27, 2025, the contents of which are hereby incorporated by reference.Background
[0002] A wide variety of different laboratory apparatus can be employed to generate large amounts of data about a set of biological samples or some other experimental subject. However, different instruments often generate data of different types, in different formats, and at different timings, making data generated by different instruments difficult to analyze together. Additionally, such large datasets are difficult to analyze, with conventional data and analysis visualization methods often overlooking significant patterns in the data, preventing development of additional insights about an experimental subject.Summary
[0003] In one aspect, an example method for efficiently merging multiple datasets from different source instruments is provided that includes: (i) obtaining a first dataset comprising a first plurality of entries, wherein each entry of the first dataset includes a respective value of a categorical variable, a respective value of a continuous variable, and a respective value of a first measured variable; (ii) obtaining a second dataset comprising a second plurality of entries, wherein each entry of the second dataset includes a respective value of the categorical variable, a respective value of the continuous variable, and a respective value of a second measured variable; and (iii) merging the first dataset and second dataset to generate a third dataset comprising a third plurality of entries by: (a) identifying, for a first entry' of the first plurality of entries, a corresponding second entry of the second plurality of entries that (1) has a value for the categorical variable that is identical to a value for the categorical variable of the first entry’ and (2) has a value for the continuous variable that is similar to a value for the continuous variable of the first entry; and (b) responsively creating a third entry' of the third plurality of entries based on the first entry and second entry.
[0004] In another aspect, an example method for improved simultaneous visualization of large, high-dimensional datasets is provided that includes: (i) determining, for a plurality of samples, a reduced-dimensional representation thereof comprising at least two dimensionssuch that each sample of the plurality of samples is represented by a respective two-dimensional location with respect to the at least two dimensions of the reduced-dimensional representation, wherein each sample of the plurality of samples is represented by a categorical variable and a plurality of measured variables; (ii) displaying, in a first pane of a user interface, a colored two-dimensional plot of tire plurality of samples such that a given sample of the plurality of samples is represented by a spot whose location corresponds, within the first pane, to the two-dimensional location of the given sample and whose color is indicative of the value of the categorical variable of the given sample; and (iii) displaying, in respective additional panes of a user interface, respective colored two-dimensional plots of the plurality of samples such that, for a given one of the additional panes, a given sample of the plurality of samples is represented by a spot whose location corresponds, within the given pane, to the two- dimensional location of the given sample and whose color is indicative of the value of a corresponding one of the plurality of variables of the given sample
[0005] In another aspect, an example method for improved simultaneous visualization of large, high-dimensional datasets is provided that includes: (i) obtaining a dataset comprising a plurality of entries, wherein each entry of the dataset includes a respective value of a first categorical variable, a respective value of a second categorical vanable that is below the first categorical variable in a hierarchical structure, and respective values of a plurality of measured variables; and (ii) displaying, in a user interface, a two-dimensional plot of the values of the plurality of measured variables such that the value of a given measured variable for a given entry of the plurality of entries is represented at a location within the two-dimensional plot having a location along the first dimension that corresponds to the identity of the given measured variable and along the second dimension that corresponds to the values of the first and second categorical variables of the given entry such that the data for the plurality of entries is arranged, across the second dimension, according to tire hierarchical structure.
[0006] In another aspect, an example method for corresponding cells between in-sample microscopy image data and cytometer event data is provided that includes: (i) generating a first image of a biological sample that contains a first cell; (ii) determining, based on the first image, at least one of a location, an optical property, an image, or a morphological property of the first cell; (iii) using a high-throughput flow cytometer, generating event data for a plurality of particles extracted from the biological sample, wherein the plurality of particles includes the first cell, and wherein the event data includes, for each particle of the plurality of particles, a respective at least one of an optical property or a morphological property; and (iv) identifying,based on the location, optical property, image, or morphological property of the first cell and the event data, the first cell within the plurality of particles.
[0007] In another aspect, an example non-transitory computer-readable medium is disclosed. The computer readable medium has stored thereon program instructions that upon execution by a processor, cause performance of the above method.
[0008] In a still further aspect, a system is provided that includes: (i) at least one processor; and (ii) a non-transitory computer-readable medium, having stored therein instructions executable by the at least one processor to cause the system to perform the above method.
[0009] The features, functions, and advantages that have been discussed can be achieved independently in various examples or may be combined in yet other examples further details of which can be seen wdtli reference to the following description and drawings.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the Office upon request and payment of the necessary' fee.
[0011] The accompanying drawings are included to provide a further understanding of the system and methods of the disclosure and are incorporated in and constitute a part of this specification. The drawings illustrate one or more embodiment(s) of the disclosure, and together with the description serve to explain the principles and operation of the disclosure.
[0012] Figure 1 depicts aspects of a set of data, according to an example implementation;
[0013] Figure 2 depicts aspects of images of an example biological sample at different points in time, according to an example implementation;
[0014] Figure 3 A depicts an example display of high-dimensional data, according to an example implementation;
[0015] Figure 3B depicts an example display of high-dimensional data, according to an example implementation;
[0016] Figure 3C depicts an example display of high-dimensional data, according to an example implementation;
[0017] Figure 4A depicts an example display of high-dimensional data, according to an example implementation;
[0018] Figure 4B depicts an example display of high-dimensional data, according to an example implementation;
[0019] Figure 5 is a functional block diagram of a system, according to an example implementation;
[0020] Figure 6 depicts a block diagram of a computing device and a computer network, according to an example implementation;
[0021] Figure 7 A shows a flowchart of a method, according to an example implementation;
[0022] Figure 7B shows a flowchart of a method, according to an example implementation;
[0023] Figure 7C shows a flowchart of a method, according to an example implementation;
[0024] Figure 7D shows a flowchart of a method, according to an example implementati n.
[0025] The drawings are for the purpose of illustrating examples, but it is understood that the inventions are not limited to the arrangements and instrumentalities shown in the drawings.Detailed DescriptionI. Overview
[0026] A given experiment may employ a variety of different instruments, generating a large amount of data. For example, an experiment to assess the effects of various treatments (e.g., applied substances, genetic modifications, admixture of additional cell types / pathogens) on various cell types (e.g., human immune cells, blood cells, epithelial cells, cancer cells, or endocrine cells) can include incubating a large number of biological samples, applying various treatments thereto (e.g., at a variety of different concentrations or varying in some other manner) and then experimentally evaluating each sample at a plurality of points in time. Such evaluation can include extracting cells from the samples to perform high-throughput flow cytometry, microscopically imaging the samples in order to determine from images thereof various cell behaviors (e.g., patterns of cell death, cell proliferation, changes in cell morphology, type, or behavior, changes in the pattern of colonies of other groups of cells), aspirating supernatant or other substances generated thereby (e.g., using a bead-based assay to quantify the identity and / or quantity of proteins or other cell products generated by the samples), measuring pH or other variables about the composition of the samples, and / or generating some other data via some other apparatus or instrument.
[0027] Different instruments can generate different types of data, on different schedules, in different formats, and / or differing in some other manner, making it difficult to synchronizedata from multiple instruments for comprehensive analysis. Failure to do so can result in important patterns or insights from the combined dataset being missed. Various manual methods exist to combine datasets from different instruments, however, such manual methods are time-intensive, prone to error, and may fail to determine the proper correspondences between samples (e.g., entries in a database representing individual measurements generated by an instrument) generated by different instruments. Such a combination of data from different instruments can include the combination of two or more of a flow cytometer (e.g., an iQue instrument), an in-incubator live cell sample imaging instrument (e.g., an Incucyte instrument), an instrument for imaging / assessing and picking individual cells, colonies, clusters, spheroids, organoids, or other sub-portions of cell samples (e.g., a CellCelector instrument), and / or a bioreactor.
[0028] For example, an Incucyte instrument could be used to generate a dataset about kinetics and other behaviors of cell samples across different treatments / cell types and an iQue instrument could be used to generate snapshots of the cell type populations of the cell samples across different the treatments / cell types, and the datasets generated by these two instniments could be linked based on the well ID of the cell sample used to generate each data point in each of the datasets. In another example, image processing or other techniques could be used to identify individual cells of a sample from image to image over time, and also to correspond such identified cells with specific cytometer-detected particles, in order to link data (e.g., optical properties of particles, microscopic morphology or other image-related data of individual adherent or other cells while in the sample) down to the individual -cell level.
[0029] Die embodiments described herein provide an efficient, reliable, and fast method to merge datasets from different instniments, resulting in a merged dataset that allows the complete set of data for a given experiment to be analyzed together. This can enable patterns or insights that depend on interactions between measurements made by different instruments to be detected. These embodiments utilize the hierarchical structure present in datasets generated by different instruments. For example, a given sample (e.g,, entry in a dataset) generated by an instrument could include categorical variables representing the identity of a particular biological sample (e.g., a well of a particular multi -we 11 sample plate) measured to create the sample, an identity of a treatment applied to the biological sample, a cell type of cells in the biological sampl e, a concentration of the treatment appli ed to the biol ogical sample, a timing at with the biological sample was assessed (e.g., imaged, subjected to flow cytometry) to generate the sample, and other information related to the provenance of the sample, in addition to various other measured variables (e.g., cell counts, cell death rates, cellgrowth rate, cell morphology, cellular protein expression level, morphology, location or other individually-identifying information about specific cells in a sample, a determined identity of an individual cell in a sample, or other information about cell behavior determined, e.g., from microscopic image data of the biological sample) generated by the instrument.
[0030] An overall data processing workflow that incorporates the embodiments described herein can include a number of steps or aspects, which may optionally be performed in order:
[0031] Data export and import (e.g., from one instrument to another that is to perform the methods described herein, and / or from one instrument to a server, database, personal computer, or other system perform the methods described herein).
[0032] Cell identification (e.g., identifying tire location, size, morphology, optical properties, or other information about individual cells in a microscopic image of a sample containing many cells). Cell identification can also include using such information to associate identified cells from one image to another (e.g,, based on similarity from one image to the next with respect to location, size, morphology, optical properties, etc.). Additionally or alternatively, such association to identify individual cells over time, between different images or other measurement, may be performed as part of ‘data merging’ process described herein, in order to link together multiple measurements of the same cell (e.g,, data from the same instrument, like size or other morphology data generated from images taken of a sample at different times by the same instrument, and / or measurements from multiple instruments, like the size or other morphology data generated from images generated by a microscope and optical properties generated of the cell by a flow cytometer).
[0033] Data merging (e.g., merging two or more datasets generated by corresponding two or more instruments into a single dataset).
[0034] Data scaling and wrangling (e.g., for each measurement of a dataset, performing an optionally independent scaling, translation, or other mapping in order to facilitate analysis of the various measurements with each other, so as to avoid biasing downstream analy sis by difference in variance between the population of values for different measurements),
[0035] Advanced algorithm -based data analysis including partition-based clustering, dimension reduction and hierarchical clustering (e.g., performing dimensionality reduction based on the measurements and / or hierarchical categorical data for every data point in the dataset and then performing analyses and / or plotting the data points in the dataset in the reduced-dimensionality space).
[0036] Data visualization using dimension reduction plot and all-in-one data table (e.g., plotting the data points in the dataset in the reduced-dimensionality space to facilitate visualdepiction of the relationship between a variety of different measurements and / or categorical variables across the dataset and / or plotting all available data (e.g., as a heatmap) with one plotted dimension vary ing with respect to measured variables (optionally ordered according to a hierarchical of other clustering thereof) and w ith the other dimension varying with respect to categorical variables like applied treatment / cell type, treatment concentration, and / or well ID (optionally ordered according to a hierarchical of other clustering thereof)).
[0037] To correspond samples from different instruments together (or cells or other identified portions of samples from one image or other measurement to the next), exact matches could be required for certain of the variables in the hierarchy, while ‘approximate’ matches could be acceptable for others. For example, to correspond a sample from a flow' cytometer w ith a sample from an in-incubator, label-free microscopic imager, categorical variables representing the identity of the biological sample or the identify of individual cells therein could be required to match exactly, whereas a continuous variable representing the time at which the samples 'ere generated could only be required to be similar to some specified degree. A ‘cell identity’ categorical variable could be used to uniquely identify individual cells within a sample, e.g., from image to image of the same sample, and could be determined separately from a later data merge process or as part of such a process (e.g., linking image data, morphological data, or other information about individual cells from one image to the next based on the categorical ‘sample’ variable matching exactly and the ‘location’ continuous variable matching approximately (e.g., within a threshold distance, according to a ‘closest match,’ ‘optical flow',’ or other matching algorithm) from one image to the next.
[0038] Note that a variable being described herein as ‘continuous’ does not require values thereof to be stored as continuous values, only that the vanable itself represents a continuous factor or phenomenon (e.g., time, location in space). Values of such continuous variables could be stored as digital values in a computer (and thus, whether integer, floating point, or represented in some other format, having only a discrete set of possible values), rounded, truncated, or otherwise modified such that they functionally assume only a discrete set of possible values while still representing a continuous variable.
[0039] Determining that a time or other continuous variable is similar between two samples could include determining, for a particular sample in one dataset, the sample of another dataset that is nearest in value to the particular sample with respect to the continuous variable. Such similarity could be subjected to some threshold, e.g., the difference in time or location must be less than a specified maximum value in order to correspond tw o samplestogether, even if they are the closest to each other between the two datasets w ith respect to the continuous variable. Associating samples between two datasets could be performed one- to-one (i.e., any particular sample in one dataset can be associated with at most one sample from the other dataset), one-to-multiple (e.g., a number of samples of a first dataset that are closer to a particular sample of a second dataset than to any other sample of the second dataset could all be augmented with the measurements of the particular sample of the second dataset), or in some other manner. Where the association is one-to-one, the process could be made more efficient by, once a set of samples are associated with each other, removing them from the set of possible samples to associate, thereby reducing the computational cost to associate subsequent samples.|00040| Samples from one dataset that are un-associated after such a merger process could be discarded. Alternatively, they could be assigned ‘default’ values for variables that would have been assumed from associated samples from another dataset. This could include setting such values to 0, 1, inf, or some other constant. Additionally or alternatively, the default value could be determined as a mean, median, mode, or other statistic determined from the samples of the other dataset(s). Such a determination could be done based on the hierarchical variables identifying the samples, e.g., only measurements that correspond to samples from the same biological sample, same treatment, or that are the same with respect to some other categorical variable could be used to determine the default values for un-associated samples that match with respect to the categorical variable.
[0041] Once a synchronized dataset has been created by merging multiple datasets, from respective different instruments or other sources, as described herein, a variety of analyses, and displayed results thereof, can be performed. Such analyses could include dimensionality reduction, clustering, or other processes to assist human researchers in realizing relevant patterns or gaining insights from a large dataset. However, very large datasets (e.g., created via the merger processes described herein) can be difficult for a human to apprehend in their entirety and at multiple levels of organization. Accordingly, improved data visualization techniques described herein can be employed to provide improved user interfaces that display more of the structure of large datasets without artificially "hiding’ aspects thereof or biasing human review thereof.
[0042] These improved user interface display methods can include using the hierarchical data of samples of the dataset to organize the display of information, thereby emphasizing patterns in the data that may run along the structure of the hierarchy. For example, a heat map of the measured variables of a dataset could be arranged along the y-axis according to theidentity of the measured variables, and along the x-axis according to two or more categorical variables, organized per their hierarchical arrangement. For example, a top level of the hierarchy could be a treatment applied to biological samples and / or a cell type of contents of the biological samples, a second level could be concentration of a substance (e.g., the treatment) applied to the biological samples, and at third level could be the identity of specific biological samples (e.g,, different replicates of the same applied treatment and concentration thereof). Thus, display of the data would be split up, across the x-axis, first with respect to the applied treatment, and then within each treatment arranged (optionally in order) according to concentration, and within each concentration level according to sample ID. To further emphasize the display of patterns within such large and multi-level datasets, the data could be clustered (e.g., hierarchically clustered) with respect to similarity across the samples between the measured variables and / or similarity across the measured variables between different treatments, concentrations, etc. Display of the data could then be done in order of the clustering, e.g., such that measured variables that exhibit similar patterns across the samples are displayed closer to each other along the y-axis.
[0043] These improved user interface display methods can additionally or alternatively include using dimensionality reduction techniques to allow patterns across various different measured variables to be quickly and accurately visually corresponded to changes in treatment type, cell time, time, or other high-level categorical or continuous sample variables. To do so, a dimensionality-reduction technique can be used, based on the measured or other variables of the samples of a dataset, to determine an at least two-dimensional reduced-dimensionality representation of the samples in tire dataset, such that more similar samples are closer to each other in the reduced-dimensional space. Two of the representative dimensions (e.g., the two highest-variance dimensions, the two dimensions that represent most of the information content of the dataset) could then be used to plot each sample of the dataset at a respective location in a number of two-dimensional plots. A first one of the plots could color-code the samples based on one or more categorical variables thereof (e.g., according to a treatment applied or a cell type); the remainder could be color-coded to represent respective different measured variables of the dataset. In this way, a viewer can quickly and accurately correspond the pattern of the samples w ith respect to category across the two-dimensional plot space to the pattern of the measured variables and / or correspond patterns across the two-dimensional plot space of different ones of the measured variables. The ordering of the measured variable plots could correspond to a clustering thereof. In some examples, such plotting could be replicated across one or more categorical variables (e.g., a first row of plots could depict the two-dimensionalpattern of a subset of the samples for a first day of an experiment, a second row could depict as subset of the samples for a second day of the experiment, etc.).II. Data Merging and Cross-Time / Cross-Instrument Cell Identification |00044| As noted above, data generated by multiple different instruments (e.g., by an in¬ incubator microscopic sample imaging apparatus and by a high-throughput flow cytometer) for the same biological sample(s) can include considerable amounts of data (e.g., a great many images, detected events, identified cells, or other entries for many samples at many points in time), making linkage of relevant entries thereof together extremely expensive, either computationally or with respect to experimental time and effort to manually link the entries together across datasets.|00045| These difficulties can be related, in part, to the fact that different instruments may generate measurements or other data for the samples on different timelines or asynchronously in some other manner, making it difficult to unambiguously correlate individual measurements for the same target (e.g., sample, cell, set of samples) from different instruments that are similar, but not identical, with respect to timing or some other variable that indicates whether such measurements should be linked together. For example, the schedule at which a microscopic imager generates images of a sample and the schedule at which a high-throughput flow cytometer is used to quantify supernatant extracted from the sample could be at approximately the same frequency, but the exact timing of the measurements by each instrument could vary’,
[0046] Uris could be addressed by, e.g., controlling the operation of the different instruments to command measurements of the same target (e.g., same sample) to occur at approximately the same time. However, this level of control is not always possible and, if possible, may be difficult to implement, with unintended latencies being introduced by, e.g., limitations of the physical instruments. Further, certain measurements may be inherently- limited to significantly different frequencies (e.g., to avoid photobleaching samples byilluminating them for imaging too frequently), preventing such a solution from being implemented naively due to the mismatch between the number of measurements available for each subject over time.
[0047] Additionally or alternatively, automated methods could be used to link entries from such different datasets. However, it is difficult to generate a generic algorithm that can link together different datasets that differ not just with respect to the timing or other information about the schedule of generation of the data, but also w ith respect to w hat data is measured and also with respect to how individual elements of the data map to different features of theexperimental subject. For example, imaging data created for a set of biological samples could correspond, at the image level, to individual biological samples (e.g., wells of a multi-well sample plate) while event data generated by a flow' cytometer could correspond, at the event level, to individual cells or other particles from such a individual biological sample. Such a scenario is further complicated by tire fact that the same instrument could generate data that corresponds to different targets at different points in time. For example, event data generated by a How cytometer after terminally sampling the contents of a biological sample could correspond, at the event level, to individual cells or other particles from the biological sample while event data taken earlier, based on supernatant extracted from the biological sample, could correspond, as a w'hole population of events to the biological sample as a w hole.|00048| The embodiments described herein provide a solution to the above problems by¬ separating the data that defines entries of multiple different datasets (e.g., generated by different instruments) into “categorical” variables (e.g,, biological sample ID, identity and / or concentration of applied treatment, type of cells present in a sample) and “continuous” variables (e.g., time at which a data entry was generated, location of a cell or other particle within a biological sample, size or other morphological variables of a cell, organoid, or other contents of a sample, brightness, reflectivity, absorptivity, fluorescent excitability, fluorescent emissivity, or other optical properties of a cell, particle, or other target at one or more wavelengths).
[0049] Categorical values can also be hierarchical. For example, a given set of samples, sample, cell, or other subject could have a lowest-level categorical variable that represents the given subject as well as one or more higher-level categories according to a hierarchy of categorical variables. For example, a lowest-level categorical variable that represents an ID of a biological sample, a one-higher categorical variable that represents a concentration of a treatment applied to a set of samples that includes the biological sample, and a two-higher categorical variable that represents an identity of the applied treatment. Such hierarchical arrangements of categorical variables can be used as described herein to enhance or otherwise modify the dataset merging methods described herein.
[0050] Similarity with respect to continuous variable(s) can include similarity with respect to multiple variables in combination. For example, with respect to a weighted combination of differences with respect to multiple continuous variables, with respect to a distance (e.g., an LI distance, an L2 distance) in a vector space defined by the multiple continuous variables, or with respect to some other similarity algorithm defined on multiple continuous variables. For example, an algorithm could be used to, based on optical and / or morphological propertiesdetermined and / or measured for each adherent cell microscopically imaged in a sample and for each particle extracted therefrom and detected in a high-throughput flow cytometer, determine a similarity between each particle event and each identified cell and based on that similarity to determine a one-to-one mapping between the events and the identified cells. In another example, an algorithm could be used to, based on location, optical, and / or morphological properties determined and / or measured for each adherent cell microscopically imaged in a sample and / for the sample itself (e.g., optical flow patterns) at two different times, determine a similarity between each cell identified in an image from a first time and each cell identified in an image from a second time and, based on that similarity, determine a one-to-one mapping between the identified cells at the two different times in the two different images.|00051| To be linked together (e.g., into a single composite entry that includes all of the data of both source entries), two different dataset entries from two different datasets (e.g., an imagebased dataset from an in-incubator sample imager and an event-based dataset from a high-throughput flow cytometer) must match exactly with respect to the categorical variable(s) and must match approximately (e.g., within a specified maximum difference, to the nearest entry, to the nearest entry that is not more than a maximum difference) with respect to the continuous variable(s). So, for example, an image of a biological sample could be linked to a set of cytometer data (e.g., entries for one or more detected events, or a composite entry representing population data for a set of events) only if (i) the value of the “sample ID” categorical value (which identifies which biological sample of a set of samples was the target / source of the data) matches exactly between a putative pair of entries of the two datasets and (ii) the value of the “time” continuous value (which identifies when the measurement(s) composing a given data entry were generated) is similar between the putative pair of entries. In such cases, a composite entry’ of a third combined dataset can be generated based on the entries of the putative pair. Such a third entry could include all of the data of the underlying pair of linked entries, could include an identification of the two linked entries, or could contain some oilier data related to the linked pair of entries. In another example, data related to an image of a single cell in a biological sample (e.g., a patch of a microscope image that represents the cell, optical, morphological, or other measurements generated from such an image patch) could be linked to data for an event within cytometer data (e.g., optical returns, size, or other data generated for an event) only if (i) the value of the “sample ID” categorical value (which identifies which biological sample of a set of samples was the target / source of the data) matches exactly between a putative pair of entries of the two datasets and (ii) the value of the optical, morphological, and / or other continuous value(s) (which identify the size, shape, color, or other properties ofthe underlying cell or event) is similar between the putative pair of entries, indicating that the event likely corresponds to the specific cell in the sample.
[0052] Such a two-level method for merging datasets can provide a variety of benefits. For example, such a method can be implemented in a computationally efficient manner, with, e.g., the categorical variable matching computed first and the continuous variable match computed second and only for entries that have already satisfied the categorical variable match, thereby avoiding the potentially higher computational cost to evaluate the ‘approximate’ similarity between the continuous variables. Additionally, such a system only requires manual annotation of which variables are ‘categorical’ and which ‘continuous’ for puiposes of dataset merging, significantly reducing the time and effort needed to, e.g., identify specific variables and related algorithms for each new set of data to be merged.
[0053] To provide further computational efficiency, in examples wherein a one-to-one mapping is required, entries that have already been matched together can be removed from the population of entries to be searched. This can have the effect of steadily reducing the population of entries in tire search set, reducing the computational cost of each additional match.
[0054] Entries between datasets can be merged in a one-to-one manner (i.e., each entry from one dataset matched to, at most, one entry of the opposite dataset), a one-to-many manner (e.g., an entry from a first dataset can match to one or more entry in the second dataset, but any- given entry' in the second can match to, at most, one entry' in the first dataset), a many-to-one manner, a many-to-many manner, or in some other fashion. For example, an entry- in a first dataset could be matched with all of the entries in the second dataset that are identical with respect to the categorical variable(s) and for which the entry in the first dataset is the closest, within the set of categorical variable -matching entries, with respect to the continuous variable (optionally subject to a maximum difference constraint),.
[0055] Figure 1 depicts an example of entries of a first dataset (e.g., microscopic images and related measurements of samples in a multi-well sample plate), represented by circles in Figure 1, and a second dataset (e.g,, measured presences or quantities of various substances, measured via bead-based ELISA or other bead-based assay using a high-throughput flow cytometer, within supernatant extracted from the samples in the multi-well sample plate), represented by squares in Figure 1. lire entries are plotted w ith respect to their values of a first categorical variable that represents the ID (“SAMPLE ID”) of the biological sample that the entries respectively represent and their values of a second continuous variable that represents the time (“TIME”) at which the measurements used to generate the entries were respectively-made. As shown, both the sample imaging and the supernatant extraction occur at respectiveroughly regular timing, with the timing sufficiently offset for either measurement that a single instniment can perform all of the imaging (e.g., a robotic in-incubator microscope) and another single instrument able to perform all of the supernatant extraction and quantification (e.g., a high-throughput flow' cytometer and an associated robotic cell picker or oilier sample extraction apparatus).
[0056] As shown, entries from both datasets are available for four different values of the sample ID. This includes a number of entries from the first dataset for the first sample ID (including 101a, 102a, 130a), from the second dataset for the first sample ID (including 105a, 106a), from the first dataset for the second sample ID (including 101b), and from the second dataset for the second sample ID (including 105b). According to tire dataset merger embodiments described herein, such entries can be linked or otherwise merged (e.g., to generate entries of a third combined dataset) in a manner that requires any pair of merged entries between datasets to be identical with respect to sample ID and similar with respect to time. Thus, entries 101b and 105a would not be merged since, despite their being very similar with respect to time, they have different sample IDs. Instead, entry 101a might be merged with entry 105a while entry 101b might be merged with entry' 105b.
[0057] Similarity with respect to continuous variable(s) can be defined in a variety of ways. For example, it could include a difference less than a specified maximum difference (in which case entry 105a might be linked to entry' 101a but not to entry 102a). In another example, similarity' could be evaluated as a nearest neighbor which, in a one-to-many merger context, might allow? entry' 105a to be linked to both 101a and 102a, while entry 103a might be linked to entry 106a. Alternatively, in a one-to-one merger context, entries 102a and 103a might not be linked to any entries in the second dataset.
[0058] In some examples, where an entry' in one dataset is not linked to an entry’ in another dataset (e.g., due to being further than a maximum difference criterion from any second dataset entry with respect to the continuous variable(s)), a corresponding entry in the merged dataset may be created therefor using the measured or other variables of the non-linked entiy of the first dataset entry- and keeping the variables corresponding to the second dataset empty. For example, if the first dataset is sample images (and values derived therefrom, like cell death / growth parameters) and the second dataset is supernatant composition data, the third dataset entry’ for a non-linked entiy of the first dataset could include the sample image and related data from the non-linked entry' of the first dataset and could include null values or otherwise indicate that those measurements (e.g., supernatant composition measurements) were not available. Alternatively, default values of such measurements could populate suchmerged entries. For example, zeros, ones, maximum values, minimum values, or average values (e.g., mean, median, mode) could populate such merged entries. Average values could be calculated over the entire dataset or across a subset (e.g., across only those entries that correspond to the sample ID, treatment, treatment concentration, or other categorical variable of the particular non-linked entries).
[0059] The embodiments described herein can allow data from multiple instruments or other dataset sources to be specifically merged, allowing the detection of insights that might not be possible (or that might only be possible fooling extensive manual data linking) when analyzed separately. For example, data can be linked at the treatment, concentration, cell line, sample plate, sample, or even lower level. For example, the embodiments described herein allow' data to be linked at the sub-sample level, e.g., at the level of individual cell colonies, individual organoids, or even individual cells. For example, individual cells can be linked from image to image of the same biological sample, or from instrument to instrument (e.g., cells identified from sample images while adherent or otherwise incubating could be linked to individual events generated after a terminal sample extraction into a high throughput flow' cytometer). This could facilitate cell-level data analysis, revealing patterns of growth, response to treatments, or other insights that would be difficult or impossible to obtain without such linkage of per-cell data over time and / or between instruments.
[0060] Linking measurements of individual cells across multiple instruments in this way can allow the relative strengths of different instruments can support each other. For example, the relatively greater spectral resolution and sensitivity of a high-throughput flow cytometer (along with the ability to use harsher and / or more perturbative dyes for a terminal assessment than w'ould be possible throughout earlier stages of incubation) can be combined with the ability of microscopic in-sample imaging to capture the ‘natural’ morphology of cells while adherent or otherwise being incubated (in contrast with the significantly perturbed morphology of cells that have been extracted from a sample and passed through the flow cell of a flow cytometer).
[0061] Such cell-to-cell linking can be accomplished in a variety of w'ays. Where the linking is between images of a sample, a variety of image processing techniques may be applied. As an illustrative example, Figure 2 depicts tw o microscopic images 201a, 201b of a biological sample that includes a first cell, which is represented in the first image 201a by a first patch 21 Oa of the first image 21 Oa and in the second image 201b by a second patch 21 Ob of the second image 210b. As shown in Figure 2, two images 201a, 201b of the same sample may \ ary due to a variety of processes, including growth or other biological processes of thesample, shifting of the sample contents within a sample container (e.g., due to motion of the sample container, extraction of material from the sample or other external perturbations, evaporation of fluid from the sample, addition of fluid to the sample), differences in alignment of the sample relative to a camera, or other processes. Accordingly, Figure 2 depicts an amount of translation and rotation of the depicted population of cells by an amount in common, as well as an amount of translation, rotation, and change in size (e.g,, representing growth or cellular motion) that is unique to each cell.
[0062] Tire patches 210a, 210b can be used to determine various properties of the cell, e.g., size, morphology, type, optical properties (e.g., reflectivity, absorptivity, fluorescent receptivity, fluorescent emissivity at one or more wavelengths) which may be related to the presence or absence of dyes whose presence is indicative of certain substances on / in the cell, or other properties. These properties (along with, optionally, the image patches themselves) can be associated with the cell in each image and, once the patches are identified as representing the same cell, merged into a single entry' for the identified cell at the two different points in time of the two different images 201a, 201b.
[0063] ITie patches 210a, 210b can be used to determine which patches represent the same cell from image to image, allowing the data thereof to be associated together at the individual cell level. A variety of techniques may be used, alone or in combination, to associate the patches together. Such techniques can begin with image segmentation to divide each image into patches thereof that represent respective different cells. Then, information about each patch in each image can be used to associate the patches together. This can include comparing the location (optionally normalized to a rigid or nonrigid optical flow calculation between images to compensate for translation / rotation of cells of the sample in common) of individual patches, linking patches together whose locations within the images are sufficiently similar. Additionally or alternatively, the visual similarity (e.g., pattern, shape, size, or other morphological characteristics) or other characteristics betw een the patches as a w hole could be evaluated and used to link patches together. This could include evaluating correspondences (e.g., convolutions, correlations, or other areal computations in two dimensions) between the patches, using attained neural network to predict whether candidate input patches are the same, projecting the patches into a representative space (e.g., a vector space) to evaluate their similarity’, and / or some other method of evaluating the similarity^ of two patches of two images,
[0064] Where the linking is between an image of a sample and flow cytometry events (e.g., corresponding to cells or oilier particles) detected after being extracted from the sample, a variety of similarity evaluation techniques may be applied. This can include determining, fromthe image patch, information about the size, morphology, optical properties (e.g., reflectivity, absorptivity, fluorescent receptivity', and / or fluorescent emissivity at one or more wavelengths) or other properties of a cell represented by the patch and comparing those to corresponding measured or derived information about tire event (e.g., size, reflectivity, absorptivity, fluorescent receptivity, and / or fluorescent emissivity at one or more wavelengths). Bayesian or other probabilistic methods can be used to leverage the greater resolution and lesser noise characteristics of a flow cytometer flow cell with respect to measured optical properties relative to, e.g., bright field microscopes or other imaging apparatus configured to microscopically image individual cells in a multi-well sample plate or other sample container. Such techniques could include performing iterative matching to first correspond patches and events that have a higher likelihood of correspondence. Lateriterations, lacking these ‘high likelihood’ pairs, may¬ be able to more accurately correspond the remaining lower-likelihood matches. Where multiple images are available, cell patch correspondences between may first be performed from image to image, with the composite multi-image representation of an identified cell (optionally time- weighted to emphasize image(s) closer in time to the event data) then used to correspond the cell to an event.
[0065] Such cell-level linkage processes (e.g., image-to-image or image-to-cytometer event) could be performed within the context of the dataset merger embodiments described herein. For example, tire location or other information (e.g., optical properties) of identified cell patches within two images could be used as the continuous variable whose similarity (along with, e.g., the “sample ID” of the patches being identical) is used to link database entries for the patches together into a common entry that represents the cell that is represented by both patches. In another example, tire size, optical properties (e.g., reflectivity, absorptivity, fluorescent receptivity, and / or fluorescent emissivity at one or more wavelengths), or other information of identified cell patches and cytometer events could be used as tire continuous variable whose similarity (along with, e.g., the “sample ID” of the patch and event being identical) is used to link database entries for the patches and events together into a common cmiy that represents the cell that is represented by both data entries. Alternatively, a patch and / or event correspondence algorithm could be executed separately, assigning, e.g., “cell ID” values to each patch or event that must, in a later dataset merger process, be identical in order for the underlying patch data to be merged together to represent the same cell. In some examples, the image-to-image or image-to-event correspondence methods described herein could be performed outside of the dataset merger context, e.g., as a standalone process todetermine, for each cell in a sample, information therefor across time and instruments in order to generate cell-level insights.
[0066] Linkage of microscopic image data and flow cytometer event data can also allow the relatively greater spectral resolution and lesser noise of the optical measurements generated by the flow cell to be combined with the relatively greater ability to resolve unperturbed ‘natural’ morphology of cells (e.g., adherent cells) while incubating in a sample container or other incubation environment which more accurately reflect the behavior of the target cells without the significant mechanical, chemical, or other perturbations associated with extraction from a sample container and optionally being exposed to additional dyes or other substances which may perturb the biological behavior of the target cells. Indeed, dyes used to perform flow cell measurements of a sample may be significantly more toxic or otherwise perturbative, since analysis through a flow cytometer is generally a ‘terminal’ evaluation at the end of incubation of a sample. Tins greater optical characterization, optionally combined with the use of more and more probative dyes or other optical contrast agents, can lead to additional insights that might not be available without the image-to-event linkage methods described herein.
[0067] For example, the relatively more limited spectral resolution, noise characteristics, or other properties of a microscopic imaging system used to image samples (e.g., in-incubator) at multiple points in time may not be sufficient to fully and with a high likelihood resolve particular optical characteristics of interest (e.g., to quantify the presence or amount of a dye that indicates the presence or location of a physiological substance of interest), to determine the type or sub-type of a cell, or to determine some other information of interest. However, the information made available by the image data may be sufficient to accurately correspond image data for a target cell to a specific cydometer event whose optical information is sufficient to determine the desired information. Tins can also allow the non-perturbed morphology of the cell (e.g., adherent cell) while in the sample container to be determined from the image data and corresponded to the higher-resolution optical data from the flow cytometer which, due to sample handling, may significantly modify’ the morphology of such cells.
[0068] In another example, the same optical properties (e.g., reflectivity, absorptivity, fluorescent receptivity, and / or fluorescent emissivity at one or more wavelengths) could correspond to different dyes (and thus different target substances) in different cells. By using the methods described herein to correspond cells within image data to event data, the relatively limited spectral resolution / number of optical channels of the image data can be multiplexed across multiple different cell types. Tire greater spectral resolution of the corresponding event data can then be used (e.g., by using barcoding or other methods) to identify the type of celland / or to otherwise determine the ‘mapping’ between the image-derived optical data and the underlying dyes / target substances. This can allow for a greater amount of information (e.g., more target substances) to be detected / quantified in the image data, since the selection and multiplexing of the dyes need not also facilitate cell type / dye type determination (as that determination can be accomplished using the higher-resolution event data).III. Data Visualization
[0069] Once a dataset has been merged in this manner or obtained in some other way, it can be displayed in its entirety in a variety of ways, allowing a user to easily view the entirety of the dataset without summarization or omission. Very large datasets can be difficult for a human to apprehend in their entirety and at multiple levels of organization. Tirus, tire embodiments described herein can include plotting the data in a variety of ways (e.g., arranged according to a clustering or ordering of the data in a manner that respects the hierarchical arrangement of various categorical variables thereof) in order to facilitate the determination of patterns within the data. Tire improved data visualization techniques described herein can be employed to provide improved user interfaces that display more of the structure of large datasets without artificially ‘hiding’ aspects thereof or biasing human review thereof.
[0070] Human visual perception is generally very competent at detecting and correlating spatial patterns. Taking advantage of this innate skill, some of the data display embodiments described herein project an available large set of data points into a two-dimensional space (e.g., based on the similarity of the individual data points to each other, such that more similar points are closer together and less similar points are farther apart in the two-dimensional space). The points, displayed in a plot according to their locations in the two-dimensional space, can then be color-coded to reflect various properties thereof (e.g., categorical or continuous variables as described herein or other measured variables). Tins can allow patterns in the data that are wholly or partially conserved between different variables to be easily apprehended by a human user, since that user will easily perceive the similarity in the patterns of the tw o (or more) plots of the different variables within the two-dimensional space.
[0071] Such improved user interface display methods described herein can additionally or alternatively include using dimensionality reduction techniques to allow patterns across various different measured variables to be quickly and accurately visually corresponded to changes in treatment ty e, cell type, treatment concentration, or other high-level categorical variables. To do so, a dimensionality-reduction technique can be used, based on the measured or other variables of the samples of a dataset, to determine an at least two-dimensional reduced- dimensionality representation of the samples in the dataset, such that more similar samples arecloser to each other in the reduced-dimensional space. Two of the representative dimensions (e.g., the two highest-variance dimensions, the two dimensions that represent most of the information content of the dataset) could then be used to plot each sample of the dataset at a respective location in a number of two-dimensional plots.
[0072] A first one of the plots could color-code the samples based on a categorical variable thereof (e.g., according to a treatment applied, a cell type, and / or a date). In this way, a viewer can quickly and accurately correspond the pattern of the samples with respect to the categorical variable across the two-dimensional plot space to the pattern of the categories across the two-dimensional plot space of different ones of the plots.
[0073] Figures 3A, 3B, and 4C depict examples of such plots. As shown, the high¬ dimensional measurement data for each sample in a dataset is projected into the reduced two- dimensional space and then plotted, with the coloration thereof depending on the particular plot. So, in a first example plot (Figure 3A), the samples are color-coded according to their date assignment. Such plotting could be replicated for different subsets of the data (e.g., for different days), with each plot applying the same color coding (e.g., with respect to a categorical variable, with respect to a measured variable). Figure 3B depicts such a display, with three different plots depicting respective day-based subsets of a whole dataset, with the same color coding applied to each (in the example of Figure 3B, the color coding depicting tire type of sample of the data points). In addition to one of more such category-based plots, further plots (e.g., replicated across day-based or other subsets of the dataset) could be color-coded based on additional measured or categorical variables of the dataset, with each data point plotted at its respecti ve location in the two-dimensional space in each of the plots. Figure 3C depicts and example of such a display, with each column of plots representing a respective day- based subset of the complete dataset and each row representing a plot of a different measured variable of the dataset.
[0074] In some examples, the hierarchical structure of the multiple categorical variables of a dataset as describe herein could be used to facilitate the display of large datasets. This can have the effect of emphasizing patterns in the data that may run along the structure of the hierarchy. For example, bar plots, heat maps, dot plots, or other plots of the measured time variables of a dataset could be arranged along the y-axis according to the identity of the measured variables (optionally according to a hierarchical or other clustering thereof according to their similarity), and along the x-axis according to the to two or more categorical variables, organized per their hierarchical arrangement.
[0075] For example, a top level of the hierarchy could be a treatment applied to biological samples and / or a cell type of contents of the biological samples, a second level could be concentration of a substance (e.g., the treatment) applied to the biological samples and / or a treatment applied to the biological samples, and at a third level could be the identity of the individual samples. Thus, display of the data would be split up across the x-axis, first with respect to the applied treatment, and then within each treatment arranged (optionally in order) according to concentration, and within each concentration level according to the individual samples.
[0076] To further emphasize the display of patterns within such large and multi-level datasets, the data could be clustered (e.g., hierarchically clustered) with respect to similarity across the samples between the measured variables and / or similarity across the measured variables between different treatments, concentrations, etc. Display of the data could then be done in order of the clustering, e.g., such that measured variables that exhibit similar patterns across the samples are displayed closer to each other along the y-axis.
[0077] Figures 4A and 4B illustrate examples of such display techniques. As shown in Figure 4A, plots of four different measured variables are displayed, with the ordering of plotting along the X-axis is arranged first by day (in three days), then by treatment category’, then by treatment concentration, and finally by sample. As shown, within each variable, the values of the measured varia ble are indicated by both color and y-axis location within the row; in other embodiments, only one or the other method could be used alone to indicate the data. For example. Figure 4B uses only color to display the measured variable values, allowing more measured variables to be plotted, in Figure 4A, the ordering of plotting along the X-axis is arranged first by treatment category, then by treatment concentration, and finally by sample.IV. Example Architecture
[0078] Figure 5 is a block diagram showing an operating environment 100 that includes or involves, for example, automated laboratory equipment like phase contrast imager(s), multi¬ color microscopes, flow cytometers, cell pickers, plate manipulating robots, in-incubator live cell imaging systems, and / or other imaging 105 or other sample measurement apparatus configured to facilitate the measurement of large numbers of biological samples 110 at high temporal resolution (e.g., one or more times per day), thereby generating large amounts of high¬ dimensional kinetic or other data about the samples that can be analyzed, displayed, or otherwise processed or manipulated as described herein. Methods 700A-D in Figures 7A-D described below show embodiments of methods that can be implemented within this operating environment 100.
[0079] Figure 6 is a block diagram illustrating an example of a computing device 200, according to an example implementation, that is configured to interface with operating environment 100, either directly or indirectly. The computing device 200 may be used to perform functions of methods described herein, e.g., those shown in Figures 7A-D and described below. Computing device 200 can be configured to perform one or more functions, including image processing, executing machine learning models (e.g., to process input images), performing dimensionality-reduction methods on sample data (e.g., to project data for a set of biological samples into a two-dimensional or other low-dimensional space to facilitate clustering, display, or other processes), clustering data (either natively, in in dimension-reduced format) with respect to sample category and / or measurement, merging datasets, associating identified cells between microscopic images and / or between such images and flow cytometer event data, classification, model training, display, or other functions that are based on measurements made of the biological samples over time, e.g., based on phase contrast or other images of cells obtained by the imaging system 105. The computing device 200 has a processor(s) 202, and also a communication interface 204, data storage 206, an output interface 208, and a display 210 each connected to a communication bus 212. The computing device 200 may also include hardware to enable communication within the computing device 200 and between the computing device 200 and other devices (e.g. not shown). The hardware may include transmitters, receivers, and antennas, for example.
[0080] Hie communication interface 204 may be a wireless interface and / or one or more wired interfaces that allow for both short-range communication and long-range communication to one or more networks 214 or to one or more remote computing devices 216 (e.g., a tablet 216a, a personal computer 216b, a laptop computer 216c and a mobile computing device 216d, for example). Such wireless interfaces may provide for communication under one or more wireless communication protocols, such as Bluetooth, Wi-Fi (e.g., an institute of electrical and electronic engineers (IEEE) 802.11 protocol), Long-Term Evolution (LTE), cellular communications, near-field communication (NFC), and / or other wireless communication protocols. Such wired interfaces may include Ethernet interface, a Universal Serial Bus (USB) interface, or similar interface to communicate via a wire, a twisted pair of wires, a coaxial cable, an optical link, a fiber-optic link, or other physical connection to a wired network. Thus, the communication interface 204 may be configured to receive input data from one or more devices and may also be configured to send output data to other devices.
[0081] Tlie communication interface 204 may also include a user-input device, such as a keyboard, a keypad, a touch screen, a touch pad, a computer mouse, a track ball and / or other similar devices, for example.|00082| The data storage 206 may include or take the form of one or more computer- readable storage media that can be read or accessed by tire processor(s) 202. The computer-readable storage media can include volatile and / or non-volatile storage components, such as optical, magnetic, organic or other memory or disc storage, which can be integrated in whole or in part with tire processor(s) 202. The data storage 206 is considered non-transitory computer readable media. In some examples, the data storage 206 can be implemented using a single physical device (e.g., one optical, magnetic, organic or other memory / or disc storage unit), while in other examples, the data storage 206 can be implemented using two or more physical devices.
[0083] The data storage 206 is a non-transitory computer readable storage medium, and executable instructions 218 are stored thereon. Tire executable instructions 218 include computer executable code. When the instructions 218 are executed by the processor(s) 202, the processor(s) 202 are caused to perform functions. Such functions include, but are not limited to, operating an imaging system 105 or other laboratory' equipment (e.g,, high-throughput flow cytometry / instrument(s)) to obtain phase contrast or other images or other measurements of the biological specimens 110 (e.g., while the biological specimens 110 are located within an incubator). Such functions could additionally or alternatively include functions to apply such imaging data to dimensionality reduction algorithms, clustering algorithms, data merging algorithms, cell identification and association / linking algorithms, and / or perform some other computational tasks as described herein.
[0084] The processor(s) 202 may be a general -purpose processor or a special purpose processor (e.g., digital signal processors, application specific integrated circuits, etc.). Tire processor(s) 202 may receive inputs from the communication interface 204 and process the inputs to generate outputs that are stored in the data storage 206 and output to the display 210. The processor(s) 202 can be configured to execute the executable instructions 218 (e.g., computer-readable program instructions) that are stored in the data storage 206 and are executable to provide the functionality of tire computing device 200 described herein.
[0085] The output interface 208 outputs information to the display 210 or to other components as well. Thus, the output interface 208 may be similar to the communication interface 204 and can be a wireless interface (e.g., transmitter) or a wired interface as well. Tlie output interface 208 may send commands to one or more controllable devices, for example.
[0086] The computing device 200 shown in Figures 5 and 6 may also be representative of a local computing device 200a in operating environment 100, for example, in communication with imaging system 105 or other laboratory'- apparatus. This local computing device 200a may perform one or more of the steps of the methods 700A-D described below, may receive input from a user and / or may send biological sample measurement, clustering, dimensionalityreduction, and / or other data or results of methods described herein and user input to computing device 200 to perform all or some of the steps of methods 700A-D.
[0087] Figures 7A-D show flowcharts of example methods 700A-D, according to example implementations. Method 700A-D shown in Figures 7A-D presents exemplary' methods that can be used with the computing device 200 of Figure 6, for example. Additionally or alternatively, some or all of the functionality of the methods described herein (e.g., methods 700A-D) could be performed by' the servers, processors, or other elements of a cloud computing service, remote server, or other computing system that is remote from but in communication with the computing device 200 and / or imaging system 105 (and optionally additional such computing devices, imaging systems, high-throughput flow cytometers, and / or other laboratory instrumentation) via the network 214 (e.g., via the Internet). This could be beneficial in that such a remote computing system could have more extensive storage, database systems, memory, processing resources, or other computing resources to facilitate the performance of dataset merger, cell identification and association (e.g., between images and / or between images and events within flow cytometer-generated data), or other computational tasks on highdimensional (e.g., high temporal resolution) measurements for large sets of biological samples, or performing some oilier computation or process as described herein. Such a remote (e.g., cloud-based) computing system could also have the benefit of access to many' laboratory' instruments or other sources of sample data, allowing the remote system to identify clusters from a larger sample of data or to otherwise improve the performance of the methods described herein.
[0088] Further, devices or systems may be used or configured to perform logical functions presented in Figures 7A-D. In some instances, components of tire devices and / or systems may be configured to perform the functions such that the components are configured and structured with hardware and / or software to enable such performance. Components of the devices and / or systems may be arranged to be adapted to, capable of, or suited for performing the functions, such as when operated in a specific manner. Methods 700A-D may include one or more operations, functions, or actions as illustrated by one or more of the blocks thereof. Although the blocks are illustrated in a sequential order, some of these blocks may also be performed inparallel, and / or in a different order than those described herein. Also, the various blocks may be combined into fewer blocks, divided into additional blocks, and / or removed based upon the desired implementation.|00089| It should be understood that for this and other processes and methods disclosed herein, flowcharts show functionality and operation of one possible implementation of the present examples. In this regard, each block may represent a module, a segment, or a portion of program code, which includes one or more instructions executable by a processor for implementing specific logical functions or steps in the process. The program code may be stored on any type of computer readable medium or data storage, for example, a storage device including a disk or hard drive. Further, the program code can be encoded on a computer- readable storage media in a machine-readable format, or on other non-transitory media or articles of manufacture. The computer readable medium may include non-transitory computer readable medium or memory’, for example, such as computer-readable media that stores data for short periods of time such as register memory, processor cache and Random Access Memory (RAM). The computer readable medium may also include non-transitory media, such as secondary or persistent long-term storage, like read only memory (ROM), optical or magnetic disks, compact disc read only memory’ (CD-ROM), for example. The computer readable media may also be any other volatile or non-volatile storage systems. Tire computer readable medium may be considered a tangible computer readable storage medium, for example,
[0090] In addition, each block in Figures 7A-D, and within other processes and methods disclosed herein, may represent circuitry that is wired to perform the specific logical functions in the process. Alternative implementations are included within the scope of the examples of the present disclosure in which functions may be executed out of order from that shown or discussed, including substantially concurrent or in reverse order, depending on the functionality- involved, as would be understood by those reasonably skilled in the art.V. Example Methods
[0091] Referring now to Figure 7A, a method for efficiently merging multiple datasets from different source instruments 700A is illustrated, optionally using the computing device(s) of Figures 5-6. Method 700A includes, at block 710a, obtaining a first dataset comprising a first plurality of entries, wherein each entry' of the first dataset includes a respective value of a categorical variable, a respective value of a continuous variable, and a respective value of a first measured variable.
[0092] Then, at block 720a, the method 700A includes obtaining a second dataset comprising a second plurality of entries, wherein each entry of tire second dataset includes a respective value of the categorical variable, a respective value of the continuous variable, and a respective value of a second measured variable
[0093] Next, at block 730a, the method 700A additionally includes merging the first dataset and second dataset to generate a third dataset comprising a third plurality of entries. Tliis can include identifying, for a first entry of the first plurality of entries, a corresponding second entry of the second plurality of entries that (i) has a value for the categorical variable that is identical to a value for the categorical variable of the first entry and (ii) has a value for the continuous variable that is similar to a value for the continuous variable of the first entry'. This can also include responsively creating a third entry of the third plurality of entries based on the first entry' and second entry'.
[0094] Referring now to Figure 7B, a method for improved simultaneous visualization of large, high-dimensional datasets 700B is illustrated, optionally using the computing device(s) of Figures 5-6. Method 700B includes, at block 710b, determining, for a plurality of samples, a reduced-dimensional representation thereof comprising at least two dimensions such that each sample of the plurality of samples is represented by a respective two-dimensional location with respect to the at least two dimensions of the reduced-dimensional representation, wherein each sample of the plurality of samples is represented by a categorical variable and a plurality of measured variables,
[0095] Then, at block 720b, the method 700B includes displaying, in a first pane of a user interface, a colored two-dimensional plot of the plurality of samples such that a given sample of the plurality of samples is represented by a spot whose location corresponds, within the first pane, to the two-dimensional location of the given sample and whose color is indicative of the value of the categorical variable of the given sample.
[0096] Next, at block 730b, the method 700B additionally includes displaying, in respective additional panes of a user interface, respective colored two-dimensional plots of the plurality of samples such that, for a given one of the additional panes, a given sample of the plurality of samples is represented by a spot whose location corresponds, within the given pane, to the two-dimensional location of the given sample and whose color is indicative of the value of a corresponding one of the plurality of variables of the given sample.
[0097] Referring now to Figure 7C, a method for improved simultaneous visualization of large, high-dimensional datasets 700C is illustrated, optionally using the computing device(s) of Figures 5-6. Method 700C includes, at block 710c, obtaining a dataset comprising a pluralityof entries, wherein each entry of the dataset includes a respective value of a first categorical variable, a respective value of a second categorical variable that is below the first categorical variable in a hierarchical structure, and respective values of a plurality of measured variables.|00098| Then, at block 720c, the method 700C includes displaying, in a user interface, a two-dimensional plot of the values of the plurality of measured variables such that the value of a given measured variable for a given entry’ of the plurality of entries is represented at a location within the two-dimensional plot having a location along the first dimension that corresponds to the identity of the given measured variable and along the second dimension that corresponds to the values of the first and second categorical variables of the given entry’ such that the data for the plurality of entries is arranged, across the second dimension, according to the hierarchical structure.
[0099] Referring now to Figure 7D, a method for corresponding cells between in-sample microscopy^ image data and cytometer event data 700D is illustrated, optionally using the computing device(s) of Figures 5-6. Method 700D includes, at block 710d, generating a first image of a biological sample that contains a first cell.[000100] Then, at block 720d, the method 700D includes determining, based on the first image, at least one of a location, an optical property’, an image, or a morphological property’ of the first cell.[000101] Next, at block 730d, the method 700D additionally includes, using a high-throughput flow cytometer, generating event data for a plurality’ of particles extracted from the biological sample, wherein the plurality of particles includes the first cell, and wherein the event data includes, for each particle of the plurality of particles, a respective at least one of an optical property or a morphological property.[000102] Next, at block 740d, the method 700D additionally includes identifying, based on the location, optical property, image, or morphological property of the first cell and the event data, the first cell within the plurality of particles.[000103] The methods 700A, 700B, 700C, and / or 700D could include additional or alternative steps or features.[000104] As discussed above, a non-transitory computer-readable medium having stored thereon program instructions that upon execution by a processor (e.g., 202) may be utilized to cause performance of any of the functions of the foregoing methods.VI. Conclusion[000105] The above detailed description describes various features and functions of the disclosed systems, devices, and methods with reference to the accompanying figures. In thefigures, similar symbols typically identify similar components, unless the context indicates otherwise, lire illustrative embodiments described in the detailed description, figures, and claims are not meant to be limiting. Other embodiments can be utilized, and other changes can be made, without departing from the scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.[000106] With respect to any or all of the message flow diagrams, scenarios, and flowcharts in the figures and as discussed herein, each step, block and / or communication may represent a processing of information and / or a transmission of information in accordance with example embodiments. Alternative embodiments are included within the scope of these example embodiments. In these alternative embodiments, for example, functions described as steps, blocks, transmissions, communications, requests, responses, and / or messages may be executed out of order from that shown or discussed, including in substantially concurrent or in reverse order, depending on the functionality involved. Further, more or fewer steps, blocks and / or functions may be used with any of the message flow diagrams, scenarios, and flow charts discussed herein, and these message flow diagrams, scenarios, and flow charts may be combined with one another, in part or in whole.[000107] A step or block that represents a processing of information may correspond to circuitry’ that can be configured to perform the specific logical functions of a herein-described method or technique. Alternatively or additionally, a step or block that represents a processing of information may correspond to a module, a segment, or a portion of program code (including related data). The program code may include one or more instructions executable by a processor for implementing specific logical functions or actions in the method or technique, Tire program code and / or related data may be stored on any type of computer-readable medium, such as a storage device, including a disk drive, a hard drive, or other storage media.[000108] The computer-readable medium may also include non-transitory computer-readable media such as computer-readable media that stores data for short periods of time like register memory, processor cache, and / or random-access memory (RAM). The computer-readable media may also include non-transitory computer-readable media that stores program code and / or data for longer periods of time, such as secondary' or persistent long-term storage, like read only memory (ROM), optical or magnetic disks, and / or compact-disc read only memory (CD-ROM), for example. The computer-readable media may also be any other volatile or non-volatile storage systems. A computer-readable medium may be considered a computer-readable storage medium, for example, or a tangible storage device.[000109] Moreover, a step or block that represents one or more information transmissions may correspond to information transmissions between software and / or hardware modules in the same physical device. However, other information transmissions may be between software modules and / or hardware modules in different physical devices.[000110] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. Tire various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope being indicated by the following claims.
Claims
CLAIMS1. A method for efficiently merging multiple datasets from different source instruments, the method comprising:obtaining a first dataset comprising a first plurality of entries, wherein each entry of the first dataset includes a respective value of a categorical variable, a respective value of a continuous variable, and a respective value of a first measured variable;obtaining a second dataset comprising a second plurality of entries, wherein each entry of the second dataset includes a respective value of the categorical variable, a respective value of the continuous variable, and a respective value of a second measured variable; and merging the first dataset and second dataset to generate a third dataset comprising a third plurality of entries by:identifying, for a first entry of the first plurality of entries, a corresponding second entry of the second plurality of entries that (i) has a value for the categorical variable that is identical to a value for the categorical variable of the first entry and (ii) has a value for the continuous variable that is similar to a value for the continuous variable of the first entry; andresponsively creating a third entry of the third plurality of entries based on the first entry and second entry.
2. The method of claim 1, wherein creating a third entry of the third plurality of entries based on the first entry and second entry comprises creating an entry in the third plurality’ of entries that includes the value of a first measured variable of the first entry and the value of a second measured variable of the second entry'.
3. The method of claim 1, wherein creating a third entry of the third plurality of entries based on the first entry and second entry' comprises creating an entry in the third plurality of entries indicating that the first entry and the second entry correspond to the third entry'.
4. The method of any of claims 1-3, wherein the value of the categorical variable of the first entry' represents an identity of a biological sample, of a plurality of biological samples, that is represented by data of the first entry, and wherein the v'alue of the continuous variable of the first entry represents a time when data of the first entry was measured.
5. The method of claim 4, wherein each entry of the first dataset includes a respective value of an additional categorical variable that is above the categorical variable in a hierarchical structure, and wherein the value of the additional categorical variable of the first entry represents at least one of a treatment applied to the biological sample that is represented by data of the first entry or a cell type of cells in the biological sample that is represented by data of the first entry'.
6. The method of any preceding claim, wherein merging the first dataset and second dataset to generate a third dataset further comprises:responsive to identifying, for the first entry’, the corresponding second entry, invalidating tire first entry' and the second entry' as valid entries for the determination of further correspondences between the first dataset and the second dataset; and subsequently identifying, for a fourth entry of the first plurality of entries that is not the first entry', a corresponding fifth entry' of the second plurality of entries that (i) is not the second entry, (ii) has a value for the categorical variable that is identical to a value for thecategorical variable of the fourth entry and (ii) has a value for the continuous variable that is similar to a value for the categorical variable of the fourth cnuy7. The method of any preceding claim, wherein entries of the first dataset are generated based on data from a high-throughput flow cytometer, and wherein entries of the second dataset are generated based on data from a live-cell sample imaging microscope.
8. The method of claim 7, wherein the first measured variable represents at least one of a cell count or a measured property of a particular cell type in a respective biological sample, and wherein the second measured variable represents a measurement of a particular cell behavior determined from a respective image of a respective biological sample.
9. The method of claim 8, wherein the first dataset and second dataset represent the effect of a plurality of different treatments on a plurality of different types of cells including at least one of immune cells, blood cells, epithelial cells, cancer cells, endocrine cells, bacteria, or yeast.
10. The method of claim 7, wherein the first measured variable represents at least one of a presence or a quantity of a substance present in a supernatant extracted from a respective biological sample and detected by the high-throughput flow cytometer using a beadbased assay, and wherein the second measured variable represents a measurement of a particular cell behavior determined from a respective image of a respective biological sample11. The method of any of claims 1-10, wherein identifying, for the first entry of the first plurality of entries, the corresponding second entry of the second plurality of entries thathas a value for the continuous variable that is similar to a value for the categorical variable of the first entry comprises determining that the value of the continuous variable of the second entry is (i) less than a threshold difference from the value of the continuous variable of the first entry and (ii) closer to the value of the continuous variable of the first entry than the value of the continuous variable of any other entry in the second plurality of entries.
12. The method of any of claims 1-11, wherein merging the first dataset and second dataset to generate a third dataset further comprises:determining that no entry of tire second plurality of entries corresponds to a sixth entry of the first plurality of entries; andresponsively creating a seventh entry of the third plurality of entries based on the sixth entry', wherein a value of the first measured variable of the seventh entry' is the value of the first measured variable of the sixth entry-, and wherein a value of the second measured variable of the seventh entry is a default value for the second measured variable.
13. The method of claim 12, wherein the default value for the second measured variable is at least one of a mean, a median, or a mode of values of the second measured variable across a subset of entries of the second plurality of entries that match the sixth entry with respect to the categorical variable.
14. The method of any of claims 1-13, further comprising:normalizing values of the first measured variable of the third plurality of entries; and normalizing values of the second measured vanable of the third plurality of entries.
15. The method of any of claims 1-14, wherein the first dataset and second dataset represent information generated from microscopic images taken of one or more biological samples at respective different points in time, wherein the values of the categorical vanable of the first entry and second entry represent an identity of a biological sample, of a plurality of biological samples, that is represented by data of the first entry and second entry, wherein the value of the continuous variable of the first entry represents at least one of a location, an optical property, an image, or a morphological property of a first cell in the biological sample as determined from a first image of the biological sample, and wherein the value of the continuous variable of the second entry represents at least one of a location, an optical property, an image, or a morphological property of the first cell in the biological sample as determined from a second image of the biological sample.16, The method of any of claims 1-14, wherein the first dataset represents information generated from a microscopic images taken of a plurality of biological samples, wherein the second dataset represents information generated by a high-throughput flow cytometer for individual particles extracted from the plurality of biological samples, wherein the values of the categorical variable of the first entry and second entry represent an identity of a biological sample, of the plurality of biological samples, that is represented by data of the first entry and second entry'-, wherein the value of the continuous variable of the first entry represents at least one of an optical property, an image, or a morphological property- of a first cell in the biological sample as determined from an image of the biological sample, and wherein the value of the continuous variable of the second entry represents at least one of an optical property or a morphological property of the first cell in the biological sample as detected by the high-throughput flow cytometer.
17. The method of claim 16, further comprising:subsequent to generating the image of the biological sample and prior to the high- throughput flow cytometer detecting the continuous value of the second entry’, adding, to the biological sample or to material extracted therefrom, a dy'e, wherein the second measured variable represents an optical property of the dye.
18. The method of claim 17, wherein the biological sample contains a first population of cells having a first optical property that is indicative of a first substance, wherein the biological sample contains a second population of cells having the first optical property that is indicative of a second substance that differs from the first substance, and wherein the first measured variable represents measurements of the first optical property for individual cells of the biological sample, the method further comprising:based on the second measured variable, identifying the first cell as being of the first population; andresponsively determining a presence or quantity of the first substance on or within the first cell based on the value of a first measured variable of the first entry'.
19. The method of any preceding claim, further comprising:performing analysis of the third dataset to identify at least one mechanism of action of a treatment applied to samples represented by entries in both the first dataset and the second dataset; andindicating, on a display by way of at least one of a graph or a heatmap, the identified at least one mechanism of action.
20. The method of any preceding claim, wherein obtaining the first dataset comprises receiving, by a server from a first laboratory instrument at a location remote from the server, an indication of the first dataset, wherein obtaining the second dataset comprises receiving, by the server from a second laboratory instrument at the location remote from the server, an indication of tlie second dataset, wherein merging the first dataset and second dataset to generate the third dataset is performed by the server, and wherein the method further comprises:transmitting, by the server to a computing device at the location remote from the server, an indication of the third dataset.
21. A method for improved simultaneous visualization of large, high-dimensional datasets, the method comprising:determining, for a plurality of samples, a reduced-dimensional representation thereof comprising at least two dimensions such that each sample of the plurality of samples is represented by a respective two-dimensional location with respect to the at least two dimensions of the reduced-dimensional representation, wherein each sample of the plurality of samples is represented by a categorical variable and a plurality of measured variables;displaying, in a first pane of a user interface, a colored two-dimensional plot of the plurality of samples such that a given sample of the plurality of samples is represented by a spot whose location corresponds, within the first pane, to the two-dimensional location of the given sample and whose color is indicative of the value of the categorical variable of the given sample; anddisplaying, in respective additional panes of a user interface, respective colored two-dimensional plots of the plurality of samples such that, for a given one of the additional panes, a given sample of the plurality of samples is represented by a spot whose location corresponds,within the given pane, to the two-dimensional location of the given sample and whose color is indicative of the value of a corresponding one of the plurality of variables of the given sample.
22. The method of claim 21, wherein determining the reduced-dimensional representation comprises determining the reduced-dimensional representation based on values of the plurality’ of measured variables of the plurality of samples.
23. The method of any of claims 21-22, further comprising:clustering the measured variables based on values of the plurality of measured variables of the plurality of samples, wherein displaying the respective colored two-dimensional plots of the plurality of samples in respective additional panes of a user interface comprises displaying the respective colored two-dimensional plots in an ordering based on tire clustering of the measured variables.
24. The method of any of claims 21-23, w herein each sample of the plurality of samples is represented by an additional categorical variable, wherein displaying a colored two- dimensional plot of the plurality of samples in the first pane of a user interface comprises displaying a colored two-dimensional plot of a first subset of tlie plurality of samples having additional categorical values matching a first value, wherein displaying respective colored t w o- dimensional plots of the plurality of samples in respective additional panes of a user interface comprises displaying respective colored two-dimensional plots of the first subset of the plurality of samples, and wherein the method further comprises:displaying, in a second pane of the user interface, a colored two-dimensional plot of a second subset of the plurality’ of samples having additional categorical values matching a second value such that a given sample of the second subset of the plurality of samples isrepresented by a spot whose location corresponds, within tire second pane, to the two-dimensional location of the given sample of the second subset and whose color is indicative of the value of the categorical variable of the given sample of the second subset; and displaying, in respective additional panes of a user interface, respective colored two- dimensional plots of the second subset of the plurality of samples such that, for a given one of the additional panes, a given sample of the second subset plurality of samples is represented by a spot whose location corresponds, within the given pane, to the two-dimensional location of the given sample of the second subset and whose color is indicative of the value of a corresponding one of the plurality of variables of the given sample of the second subset.
25. The method of claim 24, wherein the additional categorical variable represents a timing of the samples of the plurality of samples.
26. The method of any of claims 21-25, wherein the categorical variable represents at least one of a treatment applied to a biological sample or a cell type of cells in a biological sample.
27. The method of any of claims 21-26, wherein the plurality of samples is the third plurality of entries of the third dataset generated as in any of claims 1-20.
28. A method for improved simultaneous visualization of large, high-dimensional datasets, the method comprising:obtaining a dataset comprising a plurality of entries, wherein each entry of the dataset includes a respective value of a first categorical variable, a respective value of a secondcategorical variable that is below the first categorical variable in a hierarchical structure, and respective values of a plurality of measured variables; anddisplaying, in a user interface, a two-dimensional plot of the values of the plurality of measured variables such that the value of a given measured variable for a given entry' of the plurality of entries is represented at a location within the two-dimensional plot having a location along the first dimension that corresponds to the identity of the given measured variable and along the second dimension that corresponds to the values of the first and second categorical variables of the given entry such that the data for the plurality of entries is arranged, across the second dimension, according to the hierarchical structure.
29. The method of claim 28, wherein the second categorical variable is a continuous variable, and wherein the data for the plurality of entries is arranged, across the second dimension, such that the entries are ordered according to their values of the second categorical variable.
30. The method of claim 29, wherein tire continuous variable represents a concentration of a substance applied to a biological sample or a timing of a measurement of a biological sample.
31. The method of claim 29, wherein the first categorical variable represents at least one of a treatment applied to a biological sample or a cell type of cells in a biological sample, and w herein the second categorical variable represents a concentration of a substance applied to a biological sample.
32. The method of claim 31, wherein each entry of the dataset includes a respective value of a third categorical variable that is below the second categorical variable in the hierarchical structure and that represents an identity’ of a biological sample.
33. The method of any of claims 28-32, further comprising:clustering the measured variables based on values of the plurality of measured variables of the plurality of entries, wherein displaying the two-dimensional plots of the values of the plurality of measured variables comprises displaying the two-dimensional plots of the values of the plurality of measured variables such that the value of a given measured variable for a given entry of the plurality of entries is represented at a location along the first dimension that corresponds to the identity of the given measured variable in an ordering based on the clustering of the measured variables.
34. The method of any of claims 28-33, clustering the plurality of entries with respect to the first categorical variable based on values of the plurality of measured variables of the plurality of entries, wherein displaying the two-dimensional plots of the values of the plurality of measured variables comprises displaying the two-dimensional plots of the values of the plurality of measured variables such that the value of a given measured variable for a given entry of the plurality of entries is represented at a location along the second dimension that corresponds to the value of the first categorical variable of the given entry in an ordering based on the clustering of the plurality of entries with respect to the first categorical variable.
35. The method of any of claims 28-34, wherein displaying the two-dimensional plot of the values of the plurality of measured variables comprises displaying a heat map whose color corresponds to the values of a plurality of measured variables.
36. The method of any of claims 28-35, wherein the plurality of entries is the third plurality of entries of the third dataset generated as in any of claims 1-20.
37. A method for corresponding cells between in-sample microscopy image data and cytometer event data comprising:generating a first image of a biological sample that contains a first cell; determining, based on the first image, at least one of a location, an optical property, an image, or a morphological property of the first cell;using a high-throughput flow cytometer, generating event data for a plurality’ of particles extracted from the biological sample, wherein the plurality of particles includes the first cell, and wherein the event data includes, for each particle of the plurality of particles, a respective at least one of an optical property or a morphological property; and identifying, based on the location, optical property, image, or morphological property of the first cell and the event data, the first cell within the plurality of particles.
38. The method of claim 37, further comprising:generating a second image of the biological sample that contains the first cell; determining, based on the second image, cell data for a plurality of cells of the biological sample that are represented in the second image, wherein the cell data includes, for each cell represented in the second image, a respective at least one of a location, an optical property, an image, or a morphological property; andidentifying, based on the location, optical property’, image, or morphological property of the first cell and the cell data, the first cell within the plurality of cells.
39. The method of any of claims 37-38. further comprising:subsequent to generating the first image of the biological sample and prior to the high- throughput flow cytometer generating the event data, adding, to the biological sample or to material extracted therefrom, a dye having a second optical property, wherein the event data includes, for each particle of the plurality of particles, the second optical property, and wherein the biological sample contains a first population of cells and a second population of cells; and based on the second optical property for the first cell within the event data, identifying the first cell as being of the first population.
40. The method of claim 39, wherein first population of cells has a first optical property that is indicative of a first substance, wherein the second population of cells has the first optical property that is indicative of a second substance that differs from the first substance, and wherein determining at least one of a location, an optical property’, an image, or a morphological property of the first cell comprises determining the first optical property of the first cell, the method further comprising:responsive to identifying the first cell as being of the first population, determining a presence or quantity of the first substance on or within the first cell based on the first optical property of the first cell.
41. The method of any of claims 37-40, wherein the first image represents the first cell when the first cell is an adherent cell within the biological sample.
42. A non -transitory’ computer readable medium having stored thereon program instructions executable by at least one processor to cause the at least one processor to perform the method of any preceding claim.
43. A system comprising:at least one processor; anda non-transitory computer-readable medium, having stored therein instructions executable by the at least one processor to cause the system to perform the method of any of claims 1-41,