Systems and methods for single molecule localization microscopy (SMLM)

Voronoi tessellation-based methods efficiently cluster and analyze SMLM data, addressing computational intensity issues, enabling real-time analysis and automation in SMLM systems.

WO2025207616A1PCT designated stage Publication Date: 2025-10-02OHIO STATE INNOVATION FOUND +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/021317
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-25
Filing Date
2025-03-25
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing single molecule localization microscopy (SMLM) methods face challenges in processing and analyzing data due to high computational intensity, leading to slow data analysis and precluding real-time automation, especially when converting point localizations into discretized images, which results in information loss or immense computational loads.

Method used

The methods utilize Voronoi tessellation to group input data points into clusters, discretize them into arrays, and measure spatial relationships, allowing for efficient clustering and analysis with nearly linear time complexity, eliminating the need for hyperparameters and reducing computational overhead.

Benefits of technology

These methods enable rapid and robust analysis of SMLM data, facilitating real-time feedback control and automation by accelerating key analysis steps, thus enabling high-throughput single molecule localization microscopy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025021317_02102025_PF_FP_ABST
    Figure US2025021317_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are methods for analysis of single molecule localization microscopy (SMLM) data comprising a plurality of input data points in the form of molecular coordinates. These methods can identify clustering within input data points at all spatial scales within the input data set. These methods can comprise (i) grouping the plurality of input data points into clusters; (ii) discretizing the clusters into a discretized array; and (iii) measuring a spatial relationship between two species of localizations captured by the discretized array.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SYSTEMS AND METHODS FOR SINGLE MOLECULE LOCALIZATION MICROSCOPY (SMLM)

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims benefit of priority of U.S. Provisional Application No. 63 / 569,373, filed March 25, 2024, which is incorporated by reference in its entirety.

[0004] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0005] This invention was made with government support under grant / contract nos. R01HL148736 and R01HL165751 awarded by the National Institutes of Health. The government has certain rights in the invention.

[0006] BACKGROUND

[0007] Fluorescence microscopy methods that overcome the diffraction limit of resolution (about 200 to 300 nm) allow imaging of biological or non-biological structures with molecular specificity and optical resolution close to the molecular scale. Among superresolution microscopy approaches, those based on single molecule localization, such as PALM (photo-activated localization microscopy), STORM (stochastic optical reconstruction microscopy), or PAINT (point accumulation for imaging in nanoscale topography) are particularly attractive owing to their extreme spatial resolution and ease of implementation.

[0008] SMLM image acquisition and processing is described, in particular, in Betzig, E. et al., “Imaging intracellular fluorescent proteins at nanometer resolution”, Science, 313, 1642-1645 (2006). In these methods of single molecule localization microscopy (SMLM), small random subsets of fluorophores are imaged in many consecutive diffraction-limited images, computationally detected and localized with high precision, and the combined fluorophore localizations are used to generate a dense super-resolution image (or dense image) defined for example as a 2D histogram of independent localizations (xi, yi).

[0009] In practice, the number of diffraction-limited images (also referred to as raw images) underlying a single dense super-resolution image, denoted K, is typically comprised between 1,000 and 100,000. This constraint follows from two conditions that must be met simultaneously to ensure high spatial resolution: (1) a low number of activated fluorophores per diffraction- limited image, denoted p, to avoid overlaps between diffraction limited spots and allow the precise localization of individual molecules. The value of p is typically between 10 and 100 (this number depends on the size and resolution of the sensor, it needs to be small enough so that diffraction limited spots from activated fluorophores do not overlap or overlap only rarely); and (2) a large number of independent fhiorophore localizations, i.e. localizations corresponding to distinct fluorophores, denoted N (with N=Kxp), to ensure sufficiently dense sampling of the underlying biological structures.

[0010] For the sake of illustration, the mean number n of fluorophores appearing as diffractionlimited fluorescent spots per raw image may be approximately equal to 15 and the number K of raw images used to reconstruct one image may be equal to 80,000.

[0011] The resolution limit due to sampling was defined in one early study as twice the mean distance between closest independent localizations. To achieve a given resolution denoted R, the Nyquist criterion yields a minimum number of fluorophore localizations denoted NNyq(R). However, reanalysis and comparison with electron microscopy data has since concluded that at least five times more molecular localizations (N5xNyq(R)=5xNNyq(R)) are in fact needed. In practice, since prior knowledge of ground truth is not available, it is common practice to continue collecting images until blinking fluorophores are exhausted through photobleaching.

[0012] Additionally, SMLM, unlike most other microscopy methods, yields point coordinates in continuous space for localized molecules, rather than digitized images comprised of pixels or voxels. This has presented a challenge an analytical challenge. Conversion of point localizations into discretized images using methods such as point splatting lead to information loss, while analysis in continuous space using methods such as point pattern analysis entails immense computational loads. Thus, much work has focused on unsupervised clustering methods to identify clusters of localizations densely grouped in space enabling further analysis in continuous space or lossless discretization to enable analysis in discrete space. However, these approaches still remain computationally intensive, and thus, much slower than data collection. This precludes real-time analysis and thus, exploitation of real-time analysis for automation of SMLM experiments.

[0013] Accordingly, there is a need to improve the processing and analysis of data obtained from SMLM. The systems and methods described herein address these and other needs.

[0014] SUMMARY

[0015] Provided herein are methods for analysis of single molecule localization microscopy (SMLM) data comprising a plurality of input data points in the form of molecular coordinates.

[0016] 9 These methods can identify clustering within input data points at all spatial scales within the input data set. These methods can comprise (i) grouping the plurality of input data points into clusters; (ii) discretizing the clusters into a discretized array; and (iii) measuring a spatial relationship between two species of localizations captured by the discretized array. In some aspects, measuring the spatial relationship between two species of localizations captured by the discretized array comprises point pattern analysis, image processing, or another suitable data analysis method.

[0017] In some aspects, step (i) comprises (a) processing the plurality of input data points to define a study area; (b) computing a tessellation cell for each of the plurality of input data points; (c) removing any tessellation cells that are unbounded or lie outside of the study area; (d) computing an instantaneous spatial density at each of the plurality of input data points using the tessellation cell; and (e) identifying clustered data points as those whose instantaneous spatial density exceeds an analytically defined density threshold. In some aspects, step (i) can further comprise (f) grouping clustered data points into individual clusters (clustered points with adjacent tessellation cells).

[0018] In some aspects, step (ii) comprises processing the instantaneous spatial density at each of the plurality of input data points using the tessellation cell to produce the discretized array.

[0019] In some aspects, step (ii) comprises generating a map of instantaneous spatial density, cluster membership, or a combination thereof. In certain aspects, step (ii) further comprises binarizing the map using the density threshold to produce a binary image mask that indicates regions of clustering for the plurality of input data points.

[0020] In some aspects, measuring a spatial relationship between two species of localizations captured by the discretized array comprises evaluating the map. In certain aspects, evaluating the map comprises point pattern analysis, image processing, or another suitable data analysis method. In certain aspects, evaluating the map comprises performing discrete point pattern analysis on a pair of cluster species, such as discrete nearest neighbor point pattern analysis.

[0021] In some aspects, the discrete point pattern analysis can comprise (i) providing the pair of cluster species as a first binary discretized array and a second binary discretized array; (ii) calculating a Euclidean distance transformation for the first binary discretized array and the second binary discretized array; (iii) calculating an observed cumulative distribution function (CDF) of interpoint distances for regions of clustered points in the first binary discretized array and the second binary discretized array; (iv) calculating a random cumulative distribution function (CDF) of interpoint distances for the first binary discretized array and the second binary discretized array; and (v) comparing the observed CDF and the random CDF to identify a degree of non-random spatial association between the pair of cluster species.

[0022] In certain aspects, the discrete point pattern analysis can comprise (i) providing the pair of cluster species as a first binary image mask and a second binary image mask; (ii) calculating a Euclidean distance transformation for the first binary image mask and the second binary image mask; (iii) calculating an observed cumulative distribution function (CDF) of nearest neighbor distances for regions of clustered points in the first binary image mask and the second binary image mask; (iv) calculating a random cumulative distribution function (CDF) of nearest neighbor distances for the first binary image mask and the second binary image mask; and (v) comparing the observed CDF and the random CDF to identify a degree of non-random spatial association between the pair of cluster species.

[0023] In some aspects, comparing the observed CDF and the random CDF comprises a statistical comparison between the observed CDF and the random CDF. In some aspects, comparing the observed CDF and the random CDF comprises performing a 2-sided Kolmogorov-Smirnov (KS) test to compare the observed CDF and the random CDF.

[0024] In some aspects, computing a tessellation cell for each of the plurality of input data points comprises computing a Voronoi tessellation cell for each of the plurality of input data points.

[0025] In some aspects, computing the instantaneous spatial density at each of the plurality of input data points using the tessellation cell comprises calculating a cell volume of each tessellation cell; and computing a tessellation cell density for each tessellation cell as the inverse of its cell size.

[0026] In some aspects, the density threshold is analytically defined by distribution theory.

[0027] In some aspects, identifying clustered data points further comprises identifying one or more ensemble clusters. In some aspects, the one or more ensemble clusters are defined as connected components of a graph data structure composed of the input data points and their tessellation cell vertices.

[0028] In some aspects, binarizing the map using the density threshold produces a binary image mask that indicates non-random regions of clustered localization.

[0029] In some aspects, computing the instantaneous spatial density at each of the plurality of input data points using the tessellation cell comprises interpolating pixel values over a hyperplane defined by coordinates at a localization position and a cell density value for each clustered data point. In some aspects, the molecular coordinates comprise two or more dimensional coordinates.

[0030] The methods described herein can allow for the analysis of complex data sets with improved computational efficiency. For example, the methods described herein can reduce computational overhead, runtimes, or a combination thereof, as compared to other methods of data analysis, such as methods that utilize Monte Carlo simulations to derive a random distribution for purposes of defining a density threshold. The methods described herein can also avoid / eliminate the need to define hyperparameters for clustering, eliminating a potential source of bias in data analysis.

[0031] In some aspects, the methods for analysis described herein are characterized by less than quadratic time complexity, such as nearly linear time complexity.

[0032] In some aspects, the methods for analysis described herein are more computationally efficient than methods which require explicit evaluation of all inter-point distances between input points.

[0033] In some aspects, the method for analysis described herein can identify clustering within the plurality of input data points at all spatial scales contained or represented within the input data points.

[0034] Also provided herein are systems for analysis of a data set comprising a plurality of input data points in the form of molecular coordinates, the system comprising a computer, wherein the computer receives the plurality of input data points, and a processor of the computer executes computer-executable instructions stored in a memory of the computer to perform the methods described herein. In some embodiments, the system can further comprise a microscope configured to perform single molecule localization microscopy (SMLM). The microscope can be operatively coupled to the computer and processor, such that the SMLM collects data set comprising a plurality of input data points in the form of molecular coordinates, and transmits them to the computer and / or processor.

[0035] Also provided herein are methods for performing high-throughput single molecule localization microscopy (SMLM). These methods can comprise providing a plurality of samples; imaging each of the plurality of samples using SMLM to provide input data points in the form of molecular coordinates for the sample; and analyzing the input data points using the methods described herein.

[0036] BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a flowchart illustrating aspects of the methods and systems described herein. Figure 2 is a flowchart illustrating aspects of the methods and systems described herein. Figure 3 is a flowchart illustrating aspects of the methods and systems described herein. Figure 4 is a flowchart illustrating aspects of the methods and systems described herein. Figure 5 is a flowchart illustrating aspects of the methods and systems described herein. Figure 6 illustrates an example computer or computing device that can be used for some, a portion of, or all of the methods described herein.

[0038] Figure 7 shows a Voronoi clustering algorithm workflow diagram. (A) Input (observed) localizations simulated as a 2-dimensional point pattern with 1500 points within the unit square, of which 50% are the result of a Poisson point process and the remaining are the result of a Thomas point process with 12 clusters. (B) Voronoi tessellation of the observed point pattern. (C) Voronoi tessellation of the observed point pattern with cells linearly colored according to their local density, with larger, low-density cells colored black to red and smaller, high-density cells colored yellow to white. (D) Theoretical randomized localizations simulated as a 2- dimensional Poisson point pattern with 1500 points within the unit square. (E) Voronoi tessellation of the random point pattern. (F) Voronoi tessellation of the random point pattern with cells colored according to their density using the same scheme as panel C. (G) Empirical probability density functions of the tessellation cell volume for the observed (solid blue line) and random (dashed blue line) point pattens. The intersection of these distributions is used as the clustering threshold and is marked with a black dot. Points from the observed pattern with tessellation cells smaller than the threshold are considered clustered (red region) and cells larger than this are non-clustered (green region). (H) The observed point pattern with clustered points uniquely colored according to the ensemble cluster they belong to and non-clustered points colored gray. (I) Clustered tessellation cells of the observed point pattern colored according to their cluster identity using the same color scheme as panel H. (J) Raster image of cluster identity using the same color scheme as panels H and I. Each dimension was discretized into 200 pixels. (K) Raster image of the observed point pattern tessellation cell density using the same color mapping as panels C and F and the same discretization scheme as panel J. (L) The observed point pattern made 70% transparent and overlayed onto the discretized cluster mask. Clustered points and pixels are respectively colored blue and white and non-clustered points and pixels are respectively colored gray and black.

[0039] Figure 8 includes isometric views of Voronoi cluster analysis results for 3-dimensional (3D) data. (A) Simulated input data of a 3D point pattern with 1500 points within the unit cube, of which 40% are the result of a Poisson point process and the remaining are the result of a Thomas point process with 12 clusters. (B) Voronoi tessellation with cells linearly colored according to their local density, with larger, low density cells colored black to red and smaller, high density cells colored yellow to white. Cell transparency is logarithmically scaled according to their size, with larger cells made more transparent and smaller cells made more opaque. (C) Clustered tessellation cells uniquely colored according to the ensemble cluster they belong to. Non-clustered cells are not rendered. (D) The input point pattern with non-clustered points colored gray with 50% transparency and clustered points colored according to their cluster identity using the same coloring scheme as panel C. (E) Raster image of tessellation cell density using the same color mapping as panel B. Each dimension was discretized into 200 pixels. (F) Raster image of ensemble cluster identity using the same color scheme as panels C and D and the same discretization scheme as panel E.

[0040] Figure 9. (A-I) Voronoi cluster analysis and (J-M) SPACE point pattern analysis of two different point patterns. (A-C) Input synthetic localization data of 2 colabeled biomolecules simulated as combinations of Poisson and Thomas point processes. (D-F) Voronoi cluster analysis results with clustered points darkly shaded and outlined in black and non clustered points as 75% transparent. (G-I) Discretized cluster results as cluster masks with colored pixels indicating regions of clustered localizations. (J-K) Empirical cumulative distribution functions (CDFs) of nearest neighbor distances of (J) probe 1 clustered pixels relative to probe 2 clustered pixels and (K) probe 2 clustered pixels relative to probe 1 clustered pixels. Solid colored lines represent the distribution of distances in the observed data while dashed black lines represent the distribution of theoretical distances that would occur if the clustered pixels were randomly distributed in the image. (L) Delta CDFs representing the difference between observed and random CDFs from panels J and K. (M) Spatial association scatter plot with the spatial association index (SAI) of probe 1 clusters relative to probe 2 on the x-axis and the SAI of probe 2 clusters relative to probe 1 on the y-axis.

[0041] Figure 10. Voronoi cluster analysis results for standard reference 3D SMLM data of labeled (A-D) nuclear pore complex and (E-H) tubulin with zoomed 2pm-square regions of interest marked by white dotted lines and magnifications below. (A,E) Scatter plots of raw localizations colored according to their position along the z-dimension. (B,F) Cluster mask images with 0.05pm pixel size. Clustered pixels are colored according to their position along the z-dimension using the same color mapping as panels A and E. (C,G) Scatter plots of clustered localizations colored according to the cluster each point belongs to. (D,H) Cluster identity images with pixels colored according to the cluster they belong to using the same color mapping as panels C and G. Figure 11. NPC experimental data with added reference grid of uniformly spaced points to improve segmentation of individual complexes.

[0042] Figure 12. Characterizing the effect of uniform noise on cluster separability. (A-D) Representative Voronoi tessellations generated from point patterns simulated in the unit square and composed of varying levels of uniform noise (black dots with white outline) and two adjacent rectangular Poisson clusters with an intensity of 2e5 points per unit cluster volume. Tessellation cells are outlined in black and clustered cells are colored according to the unique cluster they belong to. (E-H) Rectangular clusters were iteratively moved closer to one another while the number of large clusters identified by the cluster analysis are counted at each iteration. (I) Minimum distance the two clusters were brought together before merging into one (Dmin), as indicated by the cluster analysis, as a function of uniform noise.

[0043] Figure 13. Characterizing the effect of noise on cluster fidelity. (A-G) Representative Voronoi tessellations generated from point patterns simulated in the unit square and composed of varying levels of (B-D) uniform or (E-G) Poisson noise with a circular Poisson cluster at the center of the study area with a diameter of 0.25 and intensity of le5. Tessellation cells are outlined in black and clustered cells are colored according to the unique cluster they belong to. (H-J) Circular clusters with varying intensity were iteratively exposed to increasing intensities of uniform (magenta) and Poisson (cyan) noise and cluster fidelity measured, defined as the ratio of the union of the discretized cluster to the size of a perfect circular mask of the same resolution with diameter of 0.25. (K) Value of cluster to noise intensity ratio where cluster fidelity drops below 90%, measured over a range of cluster intensities. (L) Histogram of 90% cluster fidelity measurements from panel K.

[0044] Figure 14. Linkage error effects on cluster analysis of point patterns with mixed Poisson (20%) and Thomas (80%) process components. Patterns are composed of 500,000 total points simulated in a 20pm cube. Thomas pattern components are simulated with either (A-D) small 50nm diameter (red) or (E-H) large 500nm diameter (blue) Gaussian clusters of constant intensity. 5 different patterns are simulated for each cluster size with varying degrees of applied Gaussian linkage error with a mean between 10 and 50nm. (A-H) Representative clusters are plotted with points colored if they’re considered clustered by the analysis and non-clustered points shaded black. Full patterns are displayed in Supplemental Figure X. (I- J) Plots summarizing the effects of linage error on cluster analysis results. (I) Percent change in the cluster tessellation cell volume threshold from the analysis of point patterns with no linkage error. (J) Global clustering similarity between analysis results of patterns with linkage error compared to those with none. Inset plot reduces the y-axis scale to zoom into the data. (K) Difference between mean cluster volume between patterns with linkage error and those without. (L) Difference between mean number of points composing each cluster between patterns with linkage error and those without.

[0045] Figure 15. The effect of point pattern sample size and dimensionality on clustering runtime and approximation of the distribution of Poisson Voronoi cell size. Poisson point patterns of dimensionality 2-6 were simulated in the unit cube with sample sizes between 2000 and 20000 points with 5 linear steps. (A-E) The sample size depicted in the x-axis are the number of viable Voronoi tessellation cells produced by each data set. Data sets with fewer than 100 viable tessellation cells were omitted from the left plots to eliminate impact of insufficient sample size. Y-axes measure the mean absolute error (MAE) between the empirical and approximated probability distribution functions (PDFs). The results for individual point patterns are depicted with black dots and the mean and standard deviation for each sample size are respectively depicted by horizontal red bars and capped error bars. (F-K) The effect of point pattern dimensionality on approximation fit was isolated by sampling 999 tessellation cell volumes from 3 of the largest point patterns at each dimensionality before generating and comparing both empirical and approximated PDFs. (F) Results for individual point patterns are depicted with black dots. (G-K) The tessellation cell volume PDFs for the worst approximation of each dimensionality were graphed, with the empirical PDF and its uncertainty respectively plotted with black dots and error bars and the approximation depicted with a red line. (L-P) Clustering runtime was graphed against point pattern sample size for each dimensionality, with the results for individual data sets depicted with black dots and linear regression fits plotted with a red line. (Q) Slope estimates from the linear regression plotted as red bars depicting the runtime seconds per 1000 points and error bars indicated the standard error of the slope estimation. P-values for slope estimates of dimensions 2-6 were respectively 1.718e-05, 4.900e- 17, 2.638e-18, 1.720e-15, and 4.274e-10.

[0046] Figure 16. Voronoi tessellation clustering for 50 different 2- and 3-dimensional (2D, 3D) Poisson point patterns simulated in the unit cube with a sample sizes of le6. (A-C) Empirical PDFs were generated using the first 945546 cell volumes for each data set to ensure equal sample size across all datasets. (A) Histograms depicting the mean absolute error (MAE) between the empirical and approximated distribution functions for 2D (blue) and 3D (green) point patterns. Overall MAE for 2D was 9.8671e-6 + / - 2.0778e-6 and overall MAE for 3D was 1.6974e-4 + / - 5.6695e-6. (B-C) Tessellation cell volume distribution functions for the worst approximation of each dimensionality were graphed, with the empirical distribution and its 95% confidence interval uncertainty respectively plotted with black dots and error bars and the approximation depicted with a red line. 2D worst fit MAE was 1.4791e-5 and 3D MAE is 1.8254e-4. (D) Histograms depicting the clustering runtime for 2D (blue) and 3D (green) point patterns. Mean 2D runtime was 66.1903s + / - 5.9428s and mean 3D runtime was 146.4457s + / - 2.0491s. Bin counts for all histograms and PDFs were determined using the Rice Rule.

[0047] Figure 17. Effect of study area size on clustering runtime and tessellation cell volume distribution approximation. 2-dimensional (2D) Poisson point patterns with a sample size of le4 were simulated in a cubic study area with side lengths between le-2 and le9 with 20 logarithmic steps. 10 different point patterns were simulated for each study area size. (A-B) No significant correlation was found between study area volume and runtime or cell volume distribution approximation fit. (C) Tessellation cell volume distribution functions with the worst approximation fit (MAE of 1.6496e-4) was graphed with the empirical distribution and its 95% confidence interval uncertainty respectively plotted with black dots and error bars and the approximation depicted with a red line.

[0048] Figure 18. Observed PDF uncertainty. (A-B) Poisson point patterns of dimensionality 2-6 were simulated in a unit cube study area with sample sizes between 2000 and 20000 points with 5 linear steps. 3 different patterns were created for each combination of pattern parameters. (A) Mean uncertainty for each point pattern’s empirical PDF of tessellation cell volumes was plotted as a function of the pattern’s sample size, measured as the number of viable tessellation cells. Dots are colored according to the pattern’s dimensionality (2D blue, 3D cyan, 4D green, 5D yellow, 6D orange). (B) 999 tessellation cell volumes were sampled from the 3 point patterns at each dimensionality with the largest sample size before computing their tessellation cell volume PDF and associated uncertainty. No statistically significant correlation was measured between pattern dimensionality and mean PDF uncertainty. (C-D) 2-dimensional Poisson (random, red) and Thomas (clustered, cyan) point processes were simulated in the unit square study area with samples sizes between 2,000 and 200,000 with 10 linear steps and 5 different patterns generated for each combination of pattern parameters. (C) Mean uncertainty for each random (red) and clustered (cyan) point pattern’s empirical PDF of tessellation cell volumes was plotted as a function of the pattern’s sample size, measured as the number of viable tessellation cells. (D) Voronoi clustering runtime for each pattern plotted as a function of sample size, measured as the number of points composing the pattern. The linear fit for the random (red) and clustered (blue) datasets are plotted with solid lines.

[0049] Figure 19. Measuring nuclear pore complex (NPC) cluster morphology. (A) Randomly chosen NPC clusters with localizations shown as black dots and fitted hollow cylinder geometry in red. Red crosshair indicates the cylinder center with pore and total radii marked by the inner and outer dashed red lines. The shaded red region indicates the body of the NPC. The leftmost NPC had an intensity of 1.7e-4 localizations per nm3, total diameter of 144nm, and pore diameter of 68nm. The middle NPC had an intensity of 1.0e-4 localizations per nm3, total diameter of 141nm, and pore diameter of 45nm. And the rightmost NPC had an intensity of l.Oe- 4 localizations per nm3, total diameter of 141nm, and pore diameter of 56nm (B-D) Histograms of NPC (B) cluster intensity, (C) total diameter, and (D) pore diameter. Of the 558 viable NPCs assessed, their mean and standard deviation of intensity, total diameter, and pore diameter were respectively 1.0e-4±3e-5 localizations per cm3, 143+llnm, and 50±14nm.

[0050] Figure 20. Linkage error effects on cluster analysis of mixed Poisson (20%) and Thomas (80%) point patterns with 500,000 total points simulated in a 20pm cube. Thomas pattern components are simulated with either (A-D) small 50nm diameter (red) or (E-H) large 500nm diameter (blue) Gaussian clusters of constant intensity. 5 different patterns are simulated for each cluster size with varying degrees of Gaussian linkage error with a mean between 10 and 50nm. (A-H) Representative whole point patterns with colored clustered points and non-clustered points shaded black. (I-O) Plots summarizing the effects of linage error on cluster analysis results. (I) Raw cluster tessellation cell volume threshold from the analysis of all point patterns and (J) the difference in threshold values between patterns with linkage error and those without. (K) Global clustering similarity between analysis results of patterns with linkage error compared to those with none. The y-axis is scaled highlight the trend for large clusters. (L) Mean cluster volume and (M) percent change of mean cluster volume between patterns with linkage error and those without. (N) Mean number of points composing each cluster and (O) percent change of mean cluster point count between patterns with linkage error and those without.

[0051] Figure 21. Embedding a Voronoi tessellation’s clustered points and cell indices into a graph data structure. (A) Given an input point cloud, (B) their Voronoi tessellation is computed. Input points are represented with red dots, tessellation cell vertices are represented with black dots, and tessellation cell edges are represented with black lines. (C) The clustering algorithm is used to identify clustered points and their associated tessellation cells, shaded in gray. Points and cell vertices are indexed for easy downstream identification and separation. (D) The clustered tessellation is embedded into a graph data structure with nodes as clustered points (red dots) and cell vertices (black dots), and edges (black lines) connecting those points to the vertices that compose their tessellation cells. Panels C and D utilize the same indexing and coloring scheme. Ensemble clusters are identified as the graph’s connected components which are individual disconnected regions of the graph, of which there are 2 in this example. Figure 22. Discretization of tessellation data into digital images. (A) Each tessellation cell is assigned a unique index (represented here with a unique color assigned to each cell interior) and the tessellation is subdivided into a uniform grid of query points (black circles) which represent pixel centroids. (B) Nearest neighbor interpolation is used to determine which tessellation cell each pixel falls within. Panels A and B use the same coloring scheme to indicate tessellation cell identity. The resulting cell index image can then be used as an indexing variable into the input point or tessellation datasets to quickly and efficiently discretize other metrics such as (C-D) ensemble cluster identity (colors indicate ensemble cluster identity), (E-F) tessellation cell volume, or (G-H) tessellation cell density. The magnitude of cell volume and density are represented with a map where blue indicates low values and yellow indicates high values.

[0052] Figure 23. Schematic illustration of a proposed autonomous high-throughput SMLM system.

[0053] DETAILED DESCRIPTION

[0054] Before the present methods and systems are disclosed and described, it is to be understood that the methods and systems are not limited to specific synthetic methods, specific components, or to particular compositions. It is also to be understood that the terminology used in this entire application is for the purpose of describing particular embodiments only and is not intended to be limiting.

[0055] As used in the specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” one particular value, to “about” another particular value, or from “about” one value to “about” another value. When such a range is expressed, another embodiment includes from the one particular value, to the other particular value, or from the one particular value to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms another embodiment. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint.

[0056] “Optional” or “optionally” means that the subsequently described event or circumstance may or may not occur, and that the description includes instances where said event or circumstance occurs and instances where it does not.

[0057] Throughout the description and claims of this specification, the word “comprise” and variations of the word, such as “comprising” and “comprises,” means “including but not limited to,” and is not intended to exclude, for example, other additives, components, integers or steps. “Exemplary” means “an example of” and is not intended to convey an indication of a preferred or ideal embodiment. “Such as” is not used in a restrictive sense, but for explanatory purposes.

[0058] Disclosed herein are components that can be used to perform the disclosed methods and systems. These and other components are disclosed herein, and it is understood that when combinations, subsets, interactions, groups, etc. of these components are disclosed that while specific reference of each various individual and collective combinations and permutation of these may not be explicitly disclosed, each is specifically contemplated and described herein, for all methods and systems. This applies to all aspects of this application including, but not limited to, steps in disclosed methods. Thus, if there are a variety of additional steps that can be performed it is understood that each of these additional steps can be performed with any specific embodiment or combination of embodiments of the disclosed methods.

[0059] As will be appreciated by one skilled in the art, the methods and systems may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the methods and systems may take the form of a computer program product on a computer-readable storage medium having computer- readable program instructions (e.g., computer software) embodied in the storage medium. More particularly, the present methods and systems may take the form of web-implemented computer software. Any suitable computer-readable storage medium may be utilized including hard disks, CD-ROMs, DVD-ROMs, optical storage devices, solid-state storage devices, or magnetic storage devices.

[0060] Embodiments of the methods and systems are described below with reference to block diagrams and flowchart illustrations of methods, systems, apparatuses and computer program products. It will be understood that each block of the block diagrams and flowchart illustrations, and combinations of blocks in the block diagrams and flowchart illustrations, respectively, can be implemented by computer program instructions. These computer program instructions may be loaded onto a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions which execute on the computer or other programmable data processing apparatus create a means for implementing the functions specified in the flowchart block or blocks.

[0061] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including computer-readable instructions for implementing the function specified in the flowchart block or blocks. The computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions that execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0062] Accordingly, blocks of the block diagrams and flowchart illustrations support combinations of means for performing the specified functions, combinations of steps for performing the specified functions and program instruction means for performing the specified functions. It will also be understood that each block of the block diagrams and flowchart illustrations, and combinations of blocks in the block diagrams and flowchart illustrations, can be implemented by special purpose hardware-based computer systems that perform the specified functions or steps, or combinations of special purpose hardware and computer instructions. The present methods and systems may be understood more readily by reference to the following detailed description of preferred embodiments and the Examples included therein and to the Figures and their previous and following description.

[0063] Provided herein are methods for analysis of single molecule localization microscopy (SMLM) data comprising a plurality of input data points in the form of molecular coordinates. These methods can identify clustering within input data points at all spatial scales within the input data set. Referring now to Figure 1 , in some examples, these methods can comprise (i) grouping the plurality of input data points into clusters; (ii) discretizing the clusters into a discretized array; and (iii) measuring a spatial relationship between two species of localizations captured by the discretized array. In some aspects, measuring the spatial relationship between two species of localizations captured by the discretized array comprises point pattern analysis, image processing, or another suitable data analysis method.

[0064] Referring still to Figure 1 , the analytical approach described herein allows a user to robustly and efficiently quantify spatial relationships within input data provided as a set of points (coordinates in Cartesian space) with a dimensionality of 2 or greater (102). All input points may fall into one species (univariate case), or into two or more categories, in which case they will be analyzed pairwise, two at a time (bivariate case). In the case of SMLM, such categories may correspond to localizations obtained from different fluorophore species used to label the sample. Clusters of points are identified within the input data of each species using unsupervised learning methods (104). Once such clusters are identified, spatial maps of identified clusters are discretized into binary image arrays with the same dimensionality as the input data (108). Spatial analysis is then performed on these discretized images to assess the strength of attraction / repulsion between signal-positive pixels within these binary image arrays (110). In the univariate case, such analysis focuses on spatial relationships between clusters of one species, whereas in the bivariate case, the analysis quantifies bidirectional spatial relationships between clusters of two categories. Thus, the methods described herein enable the robust and rapid quantification of spatial relationships within the input data. Additionally, by accelerating key bottleneck steps in the analysis, our process can be completed faster than SMLM data collection, enabling real-time analysis, and thereby, feedback control of the SMLM system for autonomous image collection and / or experimentation.

[0065] Next, unsupervised clustering can be performed. Unsupervised clustering of input points can be performed using a number of algorithms, such as DBSCAN (Density-based spatial clustering of applications with noise), OPTICS (Ordering points to identify the clustering structure), Voronoi tessellation etc. Among these Voronoi tessellation presents certain unique advantages - it can be implemented with a high degree of computational efficiency, it provides information about all length scales contained within the input data from a single algorithm pass, and it does not require any user-defined parameters.

[0066] Referring now to Figure 2, in some embodiments, methods for analysis of a data set (102) can begin by first processing the plurality of input data points to define a study area (204). The study area can be defined by calculating the smallest area / volume bounded by a right prism or in higher dimensions, a prismatic polytope that contains all of the input data points. In some embodiments, these methods can also include processing and cleaning the data set. For example, in some embodiments, these methods can also include translating all points input data points such that the minimum coordinate in every dimension is 0. In some embodiments, these methods can also include removing duplicate input data points from the data set.

[0067] Next, a tessellation cell can be computed for each of the plurality of input data points (206). In some embodiments, this can comprise computing a Voronoi tessellation cell for each of the plurality of input data points. Subsequently, any tessellation cells that are unbounded or lie outside of the study area are removed (208). In some embodiments, these methods can further comprise fitting a convex hull to each or a group of tessellation cell(s), and removing any tessellation cells that form degenerate convex hulls.

[0068] Next, an instantaneous spatial density can be computed at each of the plurality of input data points using the tessellation cell (210). In some aspects, computing the instantaneous spatial density at each of the plurality of input data points using the tessellation cell can comprise calculating a cell volume of each tessellation cell; and computing a tessellation cell density for each tessellation cell as the inverse of its cell size.

[0069] Next, clustered data points can be identified (212). Clustered data points can be identified as those whose instantaneous spatial density exceeds an analytically defined density threshold. The density threshold typically corresponds to the minimum local density of input points, which is observed more frequently in the input data than would be expected under spatial randomness. Thus, the threshold can be determined by comparing the observed distribution of tessellation cell sizes / local spatial densities with that predicted under complete spatial randomness. The distribution of tessellation cell sizes / local spatial densities under complete spatial randomness can in turn be determined by numerical simulations, wherein the input points are repeatedly randomized, or evaluated analytically by modeling it as a Gamma distribution.

[0070] In some aspects, the density threshold can be analytically defined by computing an expected or random distribution. Clusters can be defined as those groupings of points or tessellation cells whose instantaneous spatial density exceeds that calculated for the expected or random distribution. In some aspects, the density threshold can be analytically defined by distribution theory by modeling the distribution of tessellation cell volumes for randomly distributed points as a 2 or 3 -parameter Gamma distribution.

[0071] In some aspects, identifying clustered data points can further comprise identifying one or more ensemble clusters. The one or more ensemble clusters can be defined as connected components of a graph data structure composed of the input data points and their tessellation cell vertices.

[0072] In some aspects, the method can next involve processing the instantaneous spatial density at each of the plurality of input data points using the tessellation cell to produce the discretized array (214). Once clustered input points are identified, maps of identified clusters (cluster map) or local densities within clusters (density map) are generated and each clustered point is assigned to a specific cluster, enabling the generation of a cluster membership map. Conventionally, the task of assigning clustered points to specific clusters is performed using recursive search methods, which are computationally intensive (quadratic time complexity). Instead, in some aspects, a computational graph data structure can be constructed using the input data points and their Voronoi cell vertices as nodes. Clustered data points are thus assigned to one or more ensemble clusters, defined as groups or neighborhoods of input points with tessellation cells that share vertices (i.e., groups of cells that are adjacent or “touching” one another). These ensemble clusters are identified by removing the cell vertex nodes from the list of connected components. This method enables assignment of clustered points to ensemble clusters to be performed with nearly linear time complexity for computation.

[0073] In some aspects, the method can next involve generating a map of instantaneous spatial density, cluster membership, or a combination thereof (218). In certain aspects, generating the instantaneous spatial density map can comprise interpolating pixel values over a hyperplane defined by coordinates at a localization position and a cell density value for each clustered data point. In some aspects, measuring a spatial relationship between two species of localizations captured by the discretized array can comprise evaluating the map.

[0074] In certain aspects, the method can next involve binarizing the map using the density threshold to produce a binary image mask that indicates regions of clustering for the plurality of input data points. In some aspects, binarizing the map using the density threshold produces a binary image mask that indicates non-random regions of clustered localization.

[0075] Optionally, the method can further comprise evaluating the map to obtain information regarding a relationship between the plurality of input data points. For example, in some aspects, the method can further comprise evaluating the map to perform data stratification based on a spatial relationship between the plurality of input data points. In some aspects, evaluating the map can comprise identifying input data points that are proximally related to one another, identifying regions with a target density of input data points, or a combination thereof.

[0076] The map can be evaluated using a variety of suitable data analysis methods. In some aspects, evaluating the map can comprise point pattern analysis, image processing, or another suitable data analysis method. In certain aspects, evaluating the map comprises performing discrete point pattern analysis on a pair of cluster species, such as discrete nearest neighbor point pattern analysis.

[0077] In some embodiments, the cluster map, density map, and / or cluster membership map, thus generated, can then be discretized into image arrays by subdividing the study area into a uniform grid of pixels or voxels without loss of information. These methods can yield a discretized image array for each species of input data points, thus enabling univariate or bivariate analysis in discrete space. These discretized maps can be processed further using a variety of image analysis, spatial statistics and / or other methods. The discretized cluster maps are binary image masks that indicates non-random regions of clustered localization. Alternatively, discretized density maps generated in the previous step can be binarized after discretization using the density threshold assessed for non-random clustering. Alternatively, density maps can be used to perform data stratification based on one or more density criteria. By way of example, Figure 7 demonstrates the application of this process to a set of input points in 2D (Figure 7(A)). The Voronoi tessellation computed for these points (Figure 7(B)) enables the generation of a density map (Figure 7(C)). A set of randomly distributed points (Figure 7(D)) are similarly processed generating tessellation cells (Figure 7(E)) and a density map (Figure 7(F)). The distribution of observed densities, gleaned from the density map in Figure 7(D) are compiled into a cumulative distribution function (Figure 7(G), solid line). The cumulative distribution function predicted under spatial randomness (Figure 7(G), dashed line) is calculated either by averaging cumulative distribution functions for multiple sets of randomized input points (processed as in Figure 7(D)-(F)) or analytically as described above. Comparison of the observed and random cumulative distribution functions enables identification of non-random clusters (Figure 7(H)). Cluster maps, density maps, and cluster membership maps, thus generated, can be readily discretized (Figure 7(I)-(K)).

[0078] Figure 8 illustrates the processing of input points in 3D (Figure 8(A)) using Voronoi tessellation to perform cluster identification (Figure 8(B)) and cluster assignment (Figure 8(C)) and the discretization of these (Figure 8(D)-(F)).

[0079] Figure 21 illustrates an example method for performing cluster assignment, wherein the clustered input data points (Figure 21(A)) and the vertices of their tessellation cells (Figure 21(B)) are used to create a graph data structure (Figure 21(C), 21(D)). Specifically, the clustered tessellation is embedded into a graph data structure with nodes as clustered points (red dots) and cell vertices (black dots), and edges (black lines) connecting those points to the vertices that compose their tessellation cells. Ensemble clusters are identified as the graph’s connected components which are individual disconnected regions of the graph, of which there are 2 in this example.

[0080] Figure 22 illustrates an example method for discretizing tessellation outputs, which can include cluster maps, density maps, and cluster identity maps. Briefly, each tessellation cell is assigned a unique index and the tessellation then subdivided into a uniform grid of query points, which represent pixel centroids. Nearest neighbor interpolation is then used to determine which tessellation cell each pixel falls within. The resulting cell index image is then used as an indexing variable into the input point or tessellation datasets to quickly and efficiently discretize other metrics such as ensemble cluster identity, tessellation cell volume, or tessellation cell density.

[0081] In a further example, an assessment was performed of the function for estimating Voronoi tessellation cell volume PDFs for Poisson point patterns and saw good approximation for all dimensions tested, between 2D and 6D (Figure 15(A-K), 16(A-C)), though the estimated PDF’s error did increase exponentially with higher dimensionality independent of sample size. The algorithm’s run time also increased exponentially with dimensionality, making analysis of more than 6D datasets untenable on a laptop computer, even for small sample sizes of a few thousand points. This is likely due to the algorithm’s us of the Qhull library which computes the tessellations using triangulation whose time complexity is known to scale poorly with dimensionality. Study area size had no effect on runtime or PDF estimation (Figure 17).

[0082] Uncertainty of the observed pattern’s PDF generated from example Poisson patterns appeared unaffected by dimensionality but decreased with larger sample sizes (Figure 18(A-B)). Clustered patterns modeled using a Thomas Process had significantly lower PDF uncertainties across all sample sizes, though the difference diminished as sample size increased (Figure 18(C)). No difference in SPACE- VorTeCS’ run time was observed between analysis of Thomas and Poisson point patterns (Figure 18(D)).

[0083] The most similar clustering algorithm which also uses Voronoi tessellation, 3DClusterViSu, required 6 hours to analyze a 3D dataset with 783,128 points, at 27.6 seconds per 1000 points. In comparison, SPACE- VorTeCS’ run time for 3D datasets was generally 0.1 seconds per 1000 points (Figure 15(M),15(Q)) from analyses of datasets with 1 million points for an average total run time of 146.4+2.0 seconds. 2D datasets were slightly faster at about 0.05 seconds per 1000 points (Figure 15(L), 15(Q)) from analyses of datasets with 1 million points for an average total run time of 66.2+5.9 seconds (Figure 16(D)). Both algorithms were executed using computers with similar specifications.

[0084] One use of discretized, binary cluster maps is to perform spatial analysis using methods from point pattern analysis.

[0085] In the univariate case, where only one species exists among input data points, the discrete point pattern analysis can comprise (i) providing a single cluster species as a binary discretized array; (ii) calculating a Euclidean distance transformation for the first binary discretized array; (iii) calculating an observed cumulative distribution function (CDF) of interpoint distances for regions of clustered points in the binary discretized array; (iv) calculating a random cumulative distribution function (CDF) of interpoint distances by modelling it as a Poisson distribution, or by computational means; and (v) comparing the observed CDF and the random CDF to identify a degree of non-random spatial association between clusters of a single species.

[0086] Referring now to Figure 3, in some embodiments, in the bivariate case, the discrete point pattern analysis can comprise (i) providing the pair of cluster species as a first binary discretized array and a second binary discretized array (304); (ii) calculating a Euclidean distance transformation for the first binary discretized array and the second binary discretized array (306); (iii) calculating an observed cumulative distribution function (CDF) of interpoint distances for regions of clustered points in the first binary discretized array and the second binary discretized array (308); (iv) calculating a random cumulative distribution function (CDF) of interpoint distances for the first binary discretized array and the second binary discretized array from the distance transformation of each image mask relative to the study area (310): and (v) comparing the observed CDF and the random CDF to identify a degree of non-random spatial association between the pair of cluster species (312).

[0087] In some aspects, comparing the observed CDF and the random CDF comprises a statistical comparison between the observed CDF and the random CDF. In some aspects, comparing the observed CDF and the random CDF comprises performing a 2-sided Kolmogorov-Smirnov (KS) test to compare the observed CDF and the random CDF.

[0088] In certain embodiments, methods for the analysis of SMLM data comprising a plurality of input data points in the form of molecular coordinates can comprise (i) grouping the plurality of input data points into clusters; (ii) discretizing the clusters into a binary image mask; and (iii) measuring a spatial relationship between two species of localizations captured by the binary image mask.

[0089] By way of example, Figure 9 illustrates the processing of two species of input points in 2D (Figure 9(A)-(C)) using Voronoi tessellation to identify clusters (Figure 9(D)-(F)) and generate discretized binary image masks (Figure 9(G)-(I)). These, in turn, are inputs for discrete spatial analysis, enabling generation of cumulative distribution functions for observed (solid) and random (dashed) nearest-neighbor distances (Figure 9(J)-(K)), which are then subtracted (Figure 9(L)) to obtain indices of non-random spatial association (Figure 9(M)).

[0090] Figure 10 illustrates detection of biological structures (top: nuclear pore complexes, bottom: microtubules) by Voronoi tessellation-based detection of clusters from single molecule localizations. Notably cluster maps resulting from such analysis capture the geometries of underlying structures and can be discretized to enable further analysis / interpretation. In some instances, when input data contains an majority of clustered points with very few unclustered, background points, tessellation cells around points at the peripheries of clusters include spikelike projections, which can distort the geometries of smaller, clustered structures (Figure 11(B), bottom). To circumvent this issue, we inject uniformly distributed points into the input data, to improve structure segmentation. Figure 19 shows the dimensions of nuclear pore complexes estimated from SMLM data using our clustering approach

[0091] Referring now to Figure 4, in certain embodiments, steps (i) and (ii) can comprise (a) processing the plurality of input data points (402) to define a study area (404); (b) computing a tessellation cell for each of the plurality of input data points (406); (c) removing any tessellation cells that are unbounded or lie outside of the study area (408); (d) computing an instantaneous spatial density at each of the plurality of input data points using the tessellation cell (410); (e) identifying clustered data points as those whose instantaneous spatial density exceeds an analytically defined density threshold (412); (g) generating an instantaneous spatial density map (414); and (h) binarizing the map using the density threshold to produce the binary image mask that indicates regions of clustering for the plurality of input data points (416).

[0092] Subsequently, the map can be evaluated to obtain information regarding a relationship between the plurality of input data points (418). In some embodiments, evaluating the map comprises performing discrete nearest neighbor point pattern analysis on a pair of cluster species. Referring now to Figure 5, in certain embodiments, the discrete point pattern analysis can comprise (i) providing the pair of cluster species as a first binary image mask and a second binary image mask (504); (ii) calculating a Euclidean distance transformation for the first binary image mask and the second binary image mask (506); (iii) calculating an observed cumulative distribution function (CDF) of nearest neighbor distances for regions of clustered points in the first binary image mask and the second binary image mask (508); (iv) calculating a random cumulative distribution function (CDF) of nearest neighbor distances for the first binary image mask and the second binary image mask (510); and (v) comparing the observed CDF and the random CDF to identify a degree of non-random spatial association between the pair of cluster species (512).

[0093] The methods described herein can allow for the analysis of complex data sets with improved computational efficiency. For example, the methods described herein can reduce computational overhead, runtimes, or a combination thereof, as compared to other methods of data analysis, such as methods that utilize Monte Carlo simulations to derive a random distribution for purposes of defining a density threshold. The methods described herein can also avoid / eliminate the need to define hyperparameters for clustering, eliminating a potential source of bias in data analysis.

[0094] In some aspects, the methods for analysis described herein are characterized by less than quadratic time complexity, such as nearly linear time complexity.

[0095] In some aspects, the methods for analysis described herein are more computationally efficient than methods which require explicit evaluation of all inter-point distances between input points. In some aspects, the method for analysis described herein can identify clustering within the plurality of input data points at all spatial scales contained or represented within the input data points.

[0096] Also provided herein are systems for analysis of a data set comprising a plurality of input data points in the form of numbers or multidimensional coordinates, the system comprising a computer, wherein the computer receives the plurality of input data points, and a processor of the computer executes computer-executable instructions stored in a memory of the computer to perform the methods described herein. In some embodiments, the system can further comprise a microscope configured to perform single molecule localization microscopy (SMLM). The microscope can be operatively coupled to the computer and processor, such that the SMLM collects data set comprising a plurality of input data points in the form of molecular coordinates, and transmits them to the computer and / or processor.

[0097] Figure 6 illustrates an exemplary computer or computing device that can be used for some, a portion of, or all of the features and / or components of the methods and systems described herein. All or a portion of the device shown in Figure 6 may comprise all or any portion of any of the components and devices described herein that may include and / or require a processor or processing capabilities. As used herein, “computer” may include a plurality of computers. The computers may include one or more hardware components such as, for example, a processor 2021, a random-access memory (RAM) module 2022, a read-only memory (ROM) module 2023, a storage 2024, a database 2025, one or more input / output (I / O) devices 2026, and an interface 2027. Alternatively and / or additionally, the computer may include one or more software components such as, for example, a computer-readable medium including computer executable instructions for performing a method or methods associated with the exemplary embodiments. It is contemplated that one or more of the hardware components listed above may be implemented using software. For example, storage 2024 may include a software partition associated with one or more other hardware components or more general storage arrangement (e.g., Storage Area Network or “SAN”). It is understood that the components listed above are exemplary only and not intended to be limiting.

[0098] Processor 2021 may include one or more processors, each configured to execute instructions and process data to perform one or more functions associated with a computer for asset verification / validation and automated transaction processing. Processor 2021 may be communicatively coupled to RAM 2022, ROM 2023, storage 2024, database 2025, TO devices 2026, and interface 2027. Processor 2021 may be configured to execute sequences of computer program instructions to perform various processes. The computer program instructions may be loaded into RAM 2022 for execution by processor 2021.

[0099] RAM 2022 and ROM 2023 may each include one or more devices for storing information associated with operation of processor 2021. For example, ROM 2023 may include a memory device configured to access and store information associated with the computer, including information for identifying, initializing, and monitoring the operation of one or more components and subsystems. RAM 2022 may include a memory device for storing data associated with one or more operations of processor 2021. For example, ROM 2023 may load instructions into RAM 2022 for execution by processor 2021.

[0100] Storage 2024 may include any type of mass storage device configured to store information that processor 2021 may need to perform processes corresponding with the disclosed embodiments. For example, storage 2024 may include one or more magnetic and / or optical disk devices, such as hard drives, CD-ROMs, DVD-ROMs, or any other type of mass media device or system (e.g., SAN).

[0101] Database 2025 may include one or more software and / or hardware components that cooperate to store, organize, sort, filter, and / or arrange data used by the computer and / or processor 2021. For example, database 2025 may include data sets for analysis (e.g., input data points) or store information and instructions related to particular data sets. It is contemplated that database 2025 may store additional and / or different information than that listed above.

[0102] VO devices 2026 may include one or more components configured to communicate information with a user associated with computer. For example, VO devices may include a console with an integrated keyboard and mouse to allow a user to maintain a database of data, and the like. I / O devices 2026 may also include a display including a graphical user interface (GUI) for outputting information on a monitor. VO devices 2026 may also include peripheral devices such as, for example, a printer for printing information associated with the computer, a user-accessible disk drive (e.g., a USB port, a floppy, CD-ROM, or DVD-ROM drive, etc.) to allow a user to input data stored on a portable media device, a microphone, a speaker system, or any other suitable type of interface device.

[0103] Interface 2027 may include one or more components configured to transmit and receive data via a communication network, such as the Internet, a local area network, a workstation peer- to-peer network, a direct link network, a wireless network, or any other suitable communication platform. For example, interface 2027 may include one or more modulators, demodulators, multiplexers, demultiplexers, network communication devices, wireless devices, antennas, modems, and any other type of device configured to enable data communication via a communication network.

[0104] Also provided herein are methods for performing high-throughput experimentation using single molecule localization microscopy (SMLM; see Figure 23 below). These methods can comprise providing a plurality of samples: imaging each of the plurality of samples using SMLM to provide input data points in the form of molecular coordinates for the sample; and analyzing the input data points using the methods described herein. In some embodiments, the input data points can be analyzed in real-time (or near real-time), and the results of such analyses can be used to drive feedback control of the SMLM system. In some embodiments, results of such analyses can be used to automate control of the microscope (and potentially to autonomously generate statistical insights).

[0105] Described below is an example control scheme for automated SMLM image collection.

[0106] (1) Automated target detection: When approaching a new sample, which may be presented as one of many included in different wells of a multi-well plate, widefield fluorescence images of the sample can be collected and rapidly analyzed to identify signal positive locations for imaging. This step may also entail the use of generalist artificial intelligence (Al) networks, such as CellProfiler / CellPose, pre-trained to recognize cells and tissues.

[0107] (2) Collect pilot SMLM images: SMLM imaging can be performed at one or more randomly chosen sample sites, which are positive for fluorescent signal and possibly also identified as containing cells / tissues using one or more pre-defined image acquisition parameter sets (laser power, frame sampling time, number of frames).

[0108] (3) Perform clustering: Unsupervised clustering can be performed on the collected single molecule localization to identify clusters and assess localization densities within clusters. SPACE- VorTeCS can be used, as demonstrated for example in Figures 11 and 12 to determine the best effective resolution (defined as the minimum distance at which two adjacent clusters are reliably separable) attainable based on the observed localization densities. This step may optionally utilize injection of uniformly distributed points into the dataset, as illustrated in Figure 12, to enhance cluster separation. If so, calculations will be undertaken as illustrated in Figure 13 to determine the maximum usable density of uniformly distributed points for which clusters identified within experimental data will not be fragmented and their boundaries faithfully reconstructed. Overall, these calculations will help determine attainable effective resolution as a function of peak localization densities within clusters observed in single molecule localizations collected from the sample. (4) Calculate imaging parameters using measured blink rate, measured localization densities, and target resolution for cluster separation: Sample imaging parameters (excitation laser power, number of frames) can be calculated using measurements of fluorophore blink rate as a function of excitation laser power collected from the sample and the SPACE- VorTeCS- derived relationship between peak localization densities and effective resolution.

[0109] (5) Drive feedback control of microscope: In some embodiments, automated target finding can be used to select locations within each sample for SMLM imaging and continue imaging until a user-defined number of images are obtained from that sample (success) or the available imaging locations are exhausted (failure). In some embodiments, the microscope can autonomously continue to collect images from each sample until a predefined statistical confidence threshold is met for spatial analysis outputs (success) or a preset maximum number of images per sample is reached without meeting the confidence threshold for spatial analysis outputs (failure) or the available number of imaging locations within the sample is exhausted without meeting the confidence threshold for spatial analysis outputs (failure). In some embodiments, a reinforcement learning algorithm can be used to guide selection of imaging locations within a given sample towards sample regions, which generate readouts that have a greater impact on a predefined statistical hypothesis. In some embodiments, SMLM imaging parameters (excitation laser power, number of frames) can be defined by the user. In some embodiments, initial SMLM imaging parameters (excitation laser power, number of frames) can be defined by the user and reoptimized based on the SPACE- VorTeCS-derived relationship between peak localization densities and effective resolution. In some embodiments, initial SMLM imaging parameters (excitation laser power, number of frames) can be automatically defined based on previous data and / or measurements of fluorophore blink rates within the sample and reoptimized based on the SPACE- VorTeCS-derived relationship between peak localization densities and effective resolution.In some embodiments, real-time analysis of SMLM localizations can generate measurements of peak localization densities within clusters, enabling SMLM image collection at a given location to be terminated once sufficient localization densities are collected to attain the desired effective resolution.

[0110] Accelerated computing: For all analysis and automated imaging tasks described above, hardware acceleration strategies may be employed in order to ensure that computation does not become a bottleneck for automated SMLM. Such hardware-acceleration strategies include the use of (i) one or more high-performance processors, possibly featuring multicore or manycore architectures, optimized for multi-threading, (ii) graphical processing unit (GPU) computing, and (iii) encoding of key, computationally intensive algorithms (eg. SMLM localization, Voronoi tessellation) onto field-programmable gate arrays (FPGAs). In addition to these hardware strategies, a real-time operating system with optimized scheduling algorithms may be employed to achieve consistency in the time taken for tasks and prevent memory leaks.

[0111] EXAMPLES

[0112] The following examples are put forth so as to provide those of ordinary skill in the art with a complete disclosure and description of how the compounds, compositions, articles, devices and / or methods claimed herein are made and evaluated, and are intended to be purely exemplary and are not intended to limit the scope of the disclosure.

[0113] Example 1: Spatial Pattern Analysis in Continuous space Enabled by Voronoi Tessellationbased Cluster Segmentation (SPACE-VorTeCS) for Use in Single Molecule Localization Microscopy (SMLM)

[0114] Single molecule localization microscopy (SMLM) is a super-resolution microscopy technique, which surpasses the diffraction limit by individually localizing fluorophore molecules within a sample that are photochemically induced to blink at random. Since Eric Betzig and William Moemer were awarded the 2014 Nobel Prize in Chemistry for the development of this approach, the technology has developed rapidly, spawning many variants and delivering improvements in experimental throughput, resolution, multiplex capability, etc. However, one arena where SMLM has posed significant challenges is that of quantitative analysis of resulting data.

[0115] Unlike traditional microscopy techniques, which yield digital images (where the image volume is discretized into pixels or voxels), SMLM yields spatial coordinates for localized molecules in continuous space. Approaches like point splatting can convert the results into pseudo-images, which can then be processed using traditional image processing methods, but this comes with information loss and decrease in effective spatial resolution (the main motivation for using SMLM). Conversely, SMLM localizations can be analyzed using spatial statistics techniques built for continuous space, such as point process analysis. However, this approach comes with massive computational cost and the risk of capturing cursory features from the data and / or missing biologically relevant information. While these issues can be mitigated by performing spatial analysis on clusters of localizations, previous clustering approaches have been (1) computationally intensive, (2) dependent on user-defined hyperparameters, and (3) incapable of capturing information from all length scales contained within a dataset. To overcome these limitations, we have developed a computationally efficient clustering-based spatial analysis approach for SMLM data. Point pattern analysis is a powerful spatial statistics framework that assesses the degree of spatial randomness between event locations, such as the positions of different biomolecular species observed using single molecule localization microscopy (SMLM). However, the significant time complexity for measuring pairwise distances between continuous coordinates coupled with the large size of SMLM datasets makes such analysis untenable due to prohibitively long runtimes.

[0116] To address this, we developed a computationally efficient adaptation of point pattern analysis, called Spatial Pattern Analysis using Closest Events (SPACE), which gains efficiency by operating on discrete data like digital images, instead of continuous Euclidean coordinates. However, this solution would necessitate the rigorous and meaningful discretization of SMLM data. Additionally, SMLM often produces multiple localizations per fluorophore molecule, which leads to overcounting artefacts that confound analysis of nearest neighbor distances between individual localizations. Cluster analysis can address this by grouping adjacent localizations, however, most cluster analyses are prone to biases through their use of hyperparameters whose specific, used-defined values can impact clustering results. Here, we propose a solution to both problems, SMLM discretization and overcounting artefacts, using a hyperparameter-free Voronoi tessellation-based cluster analysis, which can produce the discrete inputs needed for SPACE as digital image masks labeling localization clusters. We refer to this pipeline as Spatial Pattern Analysis in Continuous space Enabled by Voronoi Tessellation-based Cluster Segmentation (SPACE- VorTeCS).

[0117] Voronoi tessellation is a kind of mathematical operation for point patterns or pointillist data (Figure 7 A) that produces a tiling of the study area such that each point is designated a surrounding area, called a cell, that specifies the space that is closest to that cell’s seed point than any other point in the pattern (Figure 7B). In 2-dimensions, cells of two adjacent points will share an edge that is the perpendicular bisector of the line connecting these points. This edge terminates where it intersects with another such edge produced by other pairs of points, thus cell edges will combine to form a polygon around their seed point, except in the case of finite point patterns where points at the pattern’s periphery will form unbounded cell edges whose cell will therefore also be unbounded. In 3 -dimensions, cell faces are formed by perpendicular bisecting planes and thus will form 3-dimensional polyhedra around seed points; and so on, these rules of Voronoi tessellation generalize to arbitrarily many dimensions. Importantly, points with more neighbors will form tessellation cells with more edges and points with closer neighbors will form smaller tessellation cells. Conversely, points with fewer neighbors will form cells with fewer edges and points with more distant neighbors will form larger cells. Thus, local point densities can be inferred from the inverse of a point’s tessellation cell volume, called the cell density (Figure 7C).

[0118] In fact, we can use this metric to determine if a cell’s seed point is clustered with its neighbors, and it is this heuristic that we leverage in our cluster analysis algorithm, SPACE- VorTeCS. First, clustered points are identified as those whose tessellation cell density exceeds a threshold that as automatically calculated as the lowest density of all the high-density cells that occur with a probability greater than expected from a Poisson point pattern of the same sample size. In practice, this threshold is the intersection point between the cell volume probability density functions (PDF) from the observed point pattern and the theoretical Poisson pattern, specifically the right-most intersection that is left of the Poisson pattern’s PDF peak (Figure 7D- 7G). This approach allows us to select a useful threshold value while avoiding the need for users to define hyperparameter values. Previous implementations of Voronoi tessellation-based techniques utilized computationally intensive Monte Carlo simulations to derive the Poisson pattern’s PDF. Though this distribution’s form is unknown and it therefore has no closed form equation, we avoid these costly simulations by analytically approximating the PDF using the generalized estimation function proposed in 2007 by Ferenc and Neda who validated it for 2 and 3D tessellations. Finally, ensemble clusters are identified as sets of clustered points with adjacent tessellation cells (Figure 7H-7I). In practice, these are efficiently computed as the connected components of a graph with nodes as the clustered points and their tessellation cell vertices, and graph edges that connect cell-vertex nodes to the node of their cell's seed point.

[0119] In addition to clustering SMEM data, this algorithm provides a quick and easy way to discretize localization clusters into a digital image with arbitrary resolution to allow for additional visualizations and utilization of image-based processing and analysis techniques such as our discrete point pattern analysis, SPACE This is done by interpolating pixel values over a hyperplane defined by coordinates at the localization positions and function values as the points’ indices. If nearest neighbor interpolation is used, the resulting image will have pixel values that indicate which tessellation cell their centroid is bounded by. This cell index image can then be used as a lookup table into the clustering results to create additional images that indicate cluster identity (Figure 7 J), cell density (Figure 7K), or any other tessellation cell metric. Both the clustering algorithm and its discretization method are defined for arbitrarily many dimensions, so it can just as easily be applied to 3-D SMEM data with no additional steps or considerations (Figure 8).

[0120] The most similar clustering algorithm which also uses Voronoi tessellation, 3DClusterViSu, required 6 hours to analyze a 3D dataset with 783,128 points, at 27.6 seconds per 1000 points. In comparison, SPACE- VorTeCS’ run time for 3D datasets was generally 0.1 seconds per 1000 points (Figure 15M,Q) from analyses of datasets with 1 million points for an average total run time of 146.4+2.0 seconds (Sup Fig 2D). 2D datasets were slightly faster at about 0.05 seconds per 1000 points (Figure 15L,Q) from analyses of datasets with 1 million points for an average total run time of 66.2+5.9 seconds (Figure 16D). Both algorithms were executed using computers with similar specifications.

[0121] Figure 12 shows Voronoi tessellation-based clustering of synthetic data modeled as two vertical stripes in the absence (Figure 12(A)) and presence (Figure 12(B)-(D)) of uniformly distributed points at different spacing / density. For each case, the distance separating these clusters (shown in normalized arbitrary units) was varied and multiple simulation runs performed for each distance value. In each case, the number of clusters identified (2, indicating successful separation of adjacent clusters, or 1, indicating failure to resolve the input into two clusters) was determined (Figure 12(E)-(I)). This, in turn, enables assessment of the effective resolution, defined as the minimum distance at which two nearby clusters can be reliably separated, as a function of the density of points within the clusters and the density of the uniformly distributed points injected to aid cluster separation (Figure 12(J)-(L)) . These results help to define a minimum density of clustered localizations that must be collected via SMLM data acquisition to achieve a desired effective resolution in the absence or presence of uniformly distributed grid points.

[0122] Figure 13 shows the results of Voronoi tessellation-based clustering of synthetic data modeled as a circle in the absence (Figure 13(A)) and presence of uniformly distributed (Figure 13(B)-(D)) or Poisson distributed (Figure 13(E)-(G)) points at different spacing / density. The impact of injecting varying amounts of uniformly or Poission -distributed points on cluster fidelity (accuracy in reproducing the precise geometry of the simulated input cluster) are evaluated (Figure 13(H)-(J)). By repeating this exercise for varying densities of points wihtin the input cluster, the influence of injecting uniformly or Poission -distributed points on cluster fidelity is evaluated as a function of both input point density and the density of injected uniformly or Poission -distributed points (Figure 13(K), (L)). These results help define upper bounds on the density of uniformly or Poission -distributed points injected to aid segmentation, such that the geometry of detected clusters is not compromised. Thus, they will be used, along with results from Figure 12, to define minimum density of clustered localizations that must be collected via SMLM data acquisition to achieve a desired effective resolution, along with, possibly, the density of uniformly or Poission -distributed points injected to aid segmentation. We also performed our own assessment of the function for estimating Voronoi tessellation cell volume PDFs for Poisson point patterns and saw good approximation for all dimensions tested, between 2D and 6D (Figure 15A-K, 16A-C), though the estimated PDF’s error did increase exponentially with higher dimensionality independent of sample size. The algorithm’s run time also increased exponentially with dimensionality, making analysis of more than 6D datasets untenable, even for small sample sizes of a few thousand points. This is likely due to the algorithm’s us of the Qhull library which computes the tessellations using triangulation whose time complexity is known to scale poorly with dimensionality. Study area size had no effect on runtime or PDF estimation (Figure 17).

[0123] Uncertainty of the observed pattern’s PDF generated from example Poisson patterns appeared unaffected by dimensionality but decreased with larger sample sizes (Figure 18A-B). Clustered patterns modeled using a Thomas Process had significantly lower PDF uncertainties across all sample sizes, though the difference diminished as sample size increased (Figure 18C). No difference in SPACE- VorTeCS’ run time was observed between analysis of Thomas and Poisson point patterns (Figure 18D).

[0124] Example Procedure for Analysis of SMLM Data

[0125] Spatial Pattern Analysis in Continuous space Enabled by Voronoi Tessellation-based Cluster Segmentation (SPACE- VorTeCS) is an algorithmic analysis pipeline that performs groups single molecule localization microscopy (SMLM) data (molecular coordinates) into clusters, discretizes the clusters into a binary image mask, and rapidly measures the spatial relationship between two species of localizations captured by the binary image masks using a form of discrete nearest neighbor point pattern analysis. As implemented in this example, SPACE- VorTeCS is performed in two stages. In the first stage (Stage A), input data points (molecular coordinates) are grouped into clusters and these clusters are discretized to produce cluster map images. In the second stage, point pattern analysis was performed at the cluster level. Example steps performed throughout the data analysis are detailed below, with inputs and outputs for each stage respectively indicated. Due to the computational efficiency achieved by this algorithm, it is a viable step for live, high-throughput screening of SMLM data. The procedures for such an experiment are detailed below as well as the steps of the spatial analysis.

[0126] An example algorithm is described below. This example algorithm is a function implemented in MATLAB that identifies clustered points from an input point cloud, groups the clustered points into ensemble clusters, and optionally discretizes the clusters into a digital image.

[0127] The function takes as inputs: 1. a set or list of a plurality of points composing a point cloud with a sample size and dimensionality greater than 2

[0128] 2. an optional definition of the study area as the bounds for each dimension to consider for analysis

[0129] 3. an optional definition of the pixel size used for discretization

[0130] The function outputs a MATLAB structure data type composed of the following fields:

[0131] 1. Points - ‘nPoints’-rows by ‘nDimensions -columns numerical matrix of input data point coordinates, where ‘nPoints’ is the number of data points and ‘nDimensions’ is the dimensionality of the coordinates.

[0132] 2. nPoints - numeric scalar representing the number of input data points. This number will be equal to the number of rows composing the ‘Points’ matrix.

[0133] 3. nDimensions - numeric scalar representing the dimensionality of the input data. This number will be equal to the number of columns composing the ‘Points’ matrix.

[0134] 4. PixelSize - numeric scalar representing the side length for the pixels which will be used to compose the cell density map (‘CellDensitylmage’) image and cluster masks (‘Clusterindeximage’, ‘ClusterMasklmage’). This number will have the same units as ‘Points’.

[0135] 5. Study Area - structure containing information regarding the size of the study area (IE design space) which bounds the input data a. Min - 1-row by ‘nDims’ -columns numeric array representing the minimum coordinate values from the input data ‘Points’ for all dimensions. b. Max - 1-row by ‘nDims ’-columns numeric array representing the maximum coordinate values from the input data ‘Points’ for all dimensions. c. Size - 1-row by ‘nDims’ -columns numeric array representing the range of coordinate values from the input data ‘Points’ for all dimensions. IE the minimum study area size needed to bound all input data points. d. Volume - numeric scalar representing the area (2-D), volume (3-D), or hypervolume (N-D) of the study area.

[0136] 6. Runtime - structure containing the runtimes for each major step of the cluster analysis. a. Process - Numeric scalar representing the number of seconds it took to parse and process the input data. b. Tessellation - Numeric scalar representing the number of seconds it took to compute the Voronoi tessellation of the input data. c. ConvexHull - Numeric scalar representing the number of seconds it took to fit convex hulls to each cell of the Voronoi tessellation. d. Threshold - Numeric scalar representing the number of seconds it took to compute the cell volume threshold for determining which input data points are clustered. e. Cluster - Numeric scalar representing the number of seconds it took to determine which unique cluster each clustered point belongs to. f. Images - Numeric scalar representing the number of seconds it took to generate the cell density map (‘CellDensityhnage’) image and cluster masks (‘Clusterindeximage’, ‘ClusterMasklmage’). g. Total - Numeric scalar representing the number of seconds it took to complete all steps of the cluster analysis. TessellationVertices - m-rows by ‘nDims’ -columns numerical matrix of the coordinates of all vertices composing the Voronoi tessellation, where ‘m’ is the number of vertices. The first row will contain all ‘Inf’ values which are used as vertex coordinates for unbounded (infinitely large) tessellation cells. CellVertexIndices - ‘nPoints’-rows by 1-column cell array where each element is a 1- row by mi-columns numeric array of row indices into ‘TessellationVertices’ which point to the vertices composing tessellation cell ‘i’, where ‘mi’ is the number of vertices that compose cell ‘i1. Unbounded cells will contain a vertex index of 1 which refers to the ‘Inf’ coordinate. CelllsBounded - ‘nPoints’-rows by 1-column logical array that indicates whether cell ‘i’ is bounded, where TRUE ( 1 ) indicates the cell is bounded and FALSE (0) indicates the cell is unbounded. Unbounded cells have an infinitely large volume and are therefore omitted from the cluster analysis. CelllsInBounds - ‘nPoints’-rows by 1-column logical array that indicates whether all vertices composing cell ‘i’ lie within the study area, where TRUE (1) indicates that all parts of that cell are inside the study area and FALSE (0) indicates that at least one vertex of that cell is outside the study area. Out of bounds cells tend to be those that belong to points on the edge of the study area. These points lack the spatial information of their unobserved neighbors, therefore their tessellation cells may not be accurate and as such are omitted from the cluster analysis. nViableCells - numeric scalar representing the number of tessellations cells that can be used for the cluster analysis. These are the cells that are bounded, lie completely within the study area, and do not form degenerate convex hulls whose vertices are co-linear (2D) or coplanar (>3D).

[0137] 12. Cellis Viable - ‘nPoints’-rows by 1-column logical array that indicates whether cell ‘i’ is used in the cluster analysis, where TRUE (1) indicates the cells that are used for cluster analysis and FALSE (0) indicates the cells that are not used because they are either unbounded, out of bounds, or degenerate.

[0138] 13. CellHullFaces - ‘nPoints’-rows by 1-column cell array where each element is a k-rows by v-columns numerical matrix of indices into ‘Tessellationvertices’, where ‘k’ is the number of edges (2D) or faces (>3D) that compose cell h’s convex hull and ‘v’ is the number of vertices that compose each edge / face. So each row of the matrix contains indices to the vertices that define an edge or face of the convex hull.

[0139] 14. CellVolume - ‘nPoints’-rows by 1-column numeric array of the volume for each tessellation cell. These number have the same units as ‘Points’ but raised to the ‘nDims’ power.

[0140] 15. CellDensity - nPoints’-rows by 1-column numeric array of the density for each tessellation cell. These values are simply the inverse of ‘cellVolume’.

[0141] 16. PointPatternlntensity - numeric scalar representing the intensity, or concentration, of the input data points if they were considered a point pattern, calculated as the sample size (number of data points) divided by the study area volume.

[0142] 17. Volume_PDF_Random_Mean - numeric scalar representing the first moment (mean) of the empirical probability density function of tessellation cell volumes for a theoretical Poisson point pattern with the same intensity and dimensionality as the observed input data points.

[0143] 18. Volume_PDF_Observed_y - row vector of the y-coordinates defining the empirical distribution of tessellation cell volumes for the observed data points stored in ‘Points’.

[0144] 19. Volume_PDF_x - row vector of the x-coordinates defining the empirical distributions of tessellation cell volumes for both the observed data points and the theoretical “random” (Poisson) point pattern.

[0145] 20. Volume_PDF_Random_y - row vector of the y-coordinates defining the empirical distribution of tessellation cell volumes for the theoretical random (Poisson) point pattern.

[0146] 21. Volume_PDF_Intersection_x - numeric scalar representing the x-coordinate of the point where the observed and random distributions intersect. 22. Volume_PDF_lntersection_y - numeric scalar representing the y-coordinate of the point where the observed and random distributions intersect.

[0147] 23. CellDensityClusterThreshold - numeric scalar used as the tessellation cell density threshold which cells must surpass for the seed point to be considered clustered. This value is simply the inverse of ‘Volume_PDF_Intersection_x’.

[0148] 24. PointlsClustered - ‘nPoints’-rows by 1 -column logical array that indicates whether input data point ‘i’ is clustered, where TRUE (1) indicates the point is clustered and FALSE (0) indicates the point is not clustered. A point is considered clustered if its tessellation cell density exceeds the threshold ‘CellDensityClusterThreshold’ .

[0149] 25. AreClustered - logical scalar that indicates whether the input data set has any clustered points, where TRUE (1) indicates that at least one point from the set is clustered and FALSE (0) indicates that none of the input data points are clustered.

[0150] 26. nClusters - numeric scalar representing the total number of ensemble point clusters. An ensemble cluster is defined as all clustered points whose tessellation cells are touching.

[0151] 27. ClusterPoints - ‘nPoints’-rows by 1-column numeric array indicating the index of the cluster each point belongs to. Non-clustered points are assigned an index of 0. The maximum value in this array will be equal to ‘nClusters’.

[0152] 28. ClusterGroups - ‘nClusters ’-rows by 1-column cell array where each element is a numeric row vector indicating the indices of the points composing cluster .

[0153] 29. ClusterSize - ‘nClusters’-rows by 1-column numeric array indicating the number of points composing each cluster.

[0154] 30. ClusterVolume - ‘nClusters ’-rows by 1-column numeric array indicating the volume of each cluster. A cluster’s volume is calculating as the sum of all tessellation cell volumes that compose the cluster.

[0155] 31. ImageSize - 1-row by ‘nDims’-columns numeric array indicating the side length of each dimension of the image, in number of pixels.

[0156] 32. Pixel2CellIndex - ‘p’-rows by 1-column numeric array indicating the index of the point whose tessellation cell bounds each image pixel, where ‘p’ is the total number of pixels composing the image. The point index for each pixel is determined using nearest neighbor interpolation with the input data points used as sample points, the points’ index used as sample values, and query points defined by the Euclidean coordinates of the pixels. 33. CellDensitylmage - numeric matrix with dimensions defined by ‘ImageSize’ indicating the density of the tessellation cell each pixel lies within. Cell density values for each pixel are determined by using ‘Pixel2CellIndex’ to index ‘CellDensity’.

[0157] 34. Clusterindeximage - numeric matrix with dimensions defined by ‘ImageSize’ indicating the index of the ensemble point cluster each pixel lies within. Cluster indices for each pixel are determined by using ‘Pixel2CellIndex’ to index ‘ClusterPoints’.

[0158] 35. ClusterMasklmage - logical matrix with dimensions defined by ‘ImageSize’ indicating whether a pixel lies within an ensemble point cluster, where TRUE (1) indicates a pixel the lies within a cluster and FALSE (0) indicates a pixel that does not lie within a cluster. These values are determined for each pixel by evaluating whether the pixels in ‘Clusterindeximage’ are greater than 0.

[0159] The algorithm is composed of the following operations:

[0160] 1. Input points are removed from the set if they are a duplicate of any other point or if any of their coordinates are non-numeric or otherwise missing.

[0161] 2. The study area is either user-defined or default as the bounding box of the input point cloud. a) If user-defined, points outside of the study area are removed from the set.

[0162] 3. The algorithm is terminated if there are fewer than 2 points remaining or if the input point cloud has less than 2 dimensions.

[0163] 4. The Voronoi tessellation of the input point cloud is computed using the built-in MATLAB function, voronoin.

[0164] 5. Input points are removed from the set if their tessellation cell is unbounded or features a vertex that resides outside of the study area.

[0165] 6. A convex hull is fit to each tessellation cell using the cells’ vertices and the built-in MATLAB function, convhulln. This function also computes the volume of each convex hull which are used as a proxy for the volume of each tessellation cell.

[0166] 7. The instantaneous density of each tessellation cell is computed as the inverse of their volume.

[0167] 8. Calculating the threshold of tessellation cell density for clustered points. . . a) Deriving the observed distribution. . .

[0168] ■ The intensity of the input point cloud is computed as the number of points divided by the volume of the study area.

[0169] ■ Random distribution mean is calculated as the inverse of the intensity of the input point cloud. ■ Observed distribution cell volumes (x-axis) is limited to 3x the mean of the random distribution to eliminate outliers.

[0170] ■ Bin count calculated using Rice Rule.

[0171] ■ Observed distribution (probability density function) calculated by binning the input points’ tessellation cell volumes.

[0172] ■ Distribution y-coordinates are normalized so the area under the curve is 1 by dividing each y-coordinate by the total sum of y-coordinates. b) Deriving the random distribution using the 2-parameter estimation function

[0173] ■ The cell volume bin values (x-coordinates) from the observed distribution are “un-normalized” by multiplying them by the intensity of the input point cloud.

[0174] ■ Parameters ‘a’ and ‘b’ are calculated using the dimensionality of the input point cloud.

[0175] ■ Random distribution y-coordinates are computed as the values of the estimation function which is calculated using the parameters ‘a’ and ‘b’ and the “un-normalized” x-coordinates.

[0176] ■ Random distribution y-coordinates are normalized so the area under the curve is 1 by dividing each y-coordinate by the total sum of the y- coordinates. c) Finding the intersection of the observed and random distributions. . .

[0177] ■ The peak of the random distribution is identified as the max y-coordinate.

[0178] ■ Distribution crossovers are identified as all positive deviations in the point-wise derivative of the logical array indicating which y-coordinate from the observed distribution is less than the corresponding y-coordinate from the random distribution

[0179] ■ The crossover point used for calculating the cell volume threshold is identified as the crossover with the smallest x-coordinate that is also left of the peak of the random distribution. If no such crossover exists, then all input points are considered “not clustered”.

[0180] ■ Because it is unlikely that both empirical distributions have y-coordinates at this exact crossover point, the crossover point is instead approximated by calculating the linear intersection of the straight lines defined by the distributions’ y-coordinates adjacent to this crossover point. d) The threshold of tessellation cell density for clustered points is calculated as the inverse of the distributions’ intersection x-coordinate (volume). . Input points are identified as “clustered” if their tessellation cell density exceeds this threshold.

[0181] 10. Identifying ensemble clusters. .. a) Using MATLAB’s built-in graph building function, a graph data structure is generated with nodes corresponding to indices which identify the clustered input points and indices which identify vertices composing the tessellation cells of the clustered input points. These nodes are indexed for organized and rapid inquiry by assigning the clustered points to nodes l:‘n’, where ‘n’ is the number of clustered points, and the associated cell vertices to nodes ‘n’+l :’n’+’m’, where ‘m’ is the total number of vertices composing all clustered points’ tessellation cells. The graph’s edges connect the nodes corresponding to clustered points to the vertices which compose their associated tessellation cell. b) The graph’s connected components are found using MATLAB’s built-in function, conncomp. The graph’s connected components are output as lists of node indices which compose each connected component of the graph. c) Ensembles of clustered input points (i.e. ensemble clusters) are defined as groups or neighborhoods of input points with tessellation cells that share vertices (i.e. groups of cells that are adjacent or “touching” one another). These ensemble clusters are identified by removing the cell vertex nodes from the list of connected components.

[0182] 11. Discretization... a) Discretize the study area into pixels by subdividing the study area into a uniform grid with a spacing equal to the user-defined pixel size. b) Each input point is assigned an identifying index that is a unique integer between l:’n’, where ‘n’ is the number of input points. c) Determine which tessellation cell each pixel’s centroid is bounded within using nearest neighbor interpolation over a hyperplane with sample coordinates as the input points’ coordinates, sample values as the input points’ identifying indices, and query points as the pixels.

[0183] ■ If the input point cloud is 2- or 3-dimensional, then interpolation is performed using MATLAB’s built-in function, scatteredlnterpolant. This function is computationally efficient but only works for low dimensional inputs.

[0184] ■ If the input point cloud is greater than 3-dimensional, then interpolation is performed using MATLAB’s built-in function, This function is computationally inefficient but works for any number of dimensions. d) The resulting digital image has pixel values equal to the index of the input point whose tessellation cell bounds the pixel’s centroid. This embedding can then be used as an indexing variable into the input points or their tessellation cells to quickly and efficiently generate other images with pixel values corresponding to any metric with values derived from the input points or their tessellation cells, such as cell density, cell volume, or the index of the ensemble cluster that input points are assigned to.

[0185] Example Procedure for High-Throughput SMLM Screening

[0186] As described above, due to the computational efficiency achieved by methods described herein, it is a viable step for live, high-throughput screening of SMLM data. Example procedures for such an experiment are detailed below.

[0187] 1. Set up samples on a large-format plate (128 or 584 well plate).

[0188] 2. Identify experimental groups and replicates. a. Define statistical comparisons to be made, specifying desired sample sizes and statistical power.

[0189] 3. Define parameters: a. Image acquisition parameters (laser wavelengths, minimum laser power, maximum laser power, frame dwell time, number of frames, number of regions of interest (ROIs) to be imaged per well). b. Localization parameters (background threshold, minimum photon yield, minimum chi-square localization confidence metric). c. Signal quality parameters (minimum blink rate, number of localizations).

[0190] 4. Initiate image collection under automated software control. a. Image one well at a time. i. Divide well area into a grid with each grid square having an area equal to the field of view of the objective. ii. Randomly select necessary number of ROIs based on parameter value from 3 a. iii. Image each ROI, sequentially collecting each spectral channel. 1. Assess signal quality parameters.

[0191] 2. If signal quality parameters fall below thresholds from 3c, adjust laser power using range defined in 3a.

[0192] 3. If max laser power is reached without achieving required signal quality, mark that ROI as a failed site and randomly select another ROI within the well.

[0193] 4. If all available ROIs within a well are exhausted without achieving successful imaging at the required number of ROIs specified signal quality levels, mark the well as a failed well and proceed to next well. iv. Perform spatial analysis as detailed below on data from each ROI as soon as a sufficient number of localizations is collected.

[0194] 1. If analysis fails, mark that ROI as a failed site and randomly select another ROI within the well.

[0195] 2. If all available ROIs within a well are exhausted without achieving successful imaging at the required number of ROIs specified signal quality levels, mark the well as a failed well and proceed to next well. b. For groups where sufficient image data is collected from the required number of ROIs and replicates (wells), perform statistical comparison of spatial analysis outputs. i. Identify these comparisons as completed. ii. Identify remaining comparisons as incomplete. c. Generate summary data and list of pending comparisons (incomplete comparisons identified in 4b-ii).

[0196] Summary

[0197] In sum, given an input set of single molecule localization points, their Voronoi tessellation is calculated where the field of view is tiled with cells corresponding to each point and that indicate the region of space closer to a cell’s seed point than any other point. When points are close together, their Voronoi cells are small, and when points are far apart, their cells are comparatively large. Thus, local point density can be inferred from the inverse of a point’s Voronoi cell size, which we call the cell density. Clustered points are then identified as those whose cell density exceeds a threshold defined as the lowest density of all the high-density cells that occur with a probability greater than expected from completely spatially random data with the same sample size. This approach allows us to select a useful threshold value while avoiding the need to define hyperparameters. Previous implementations of this technique utilized computationally intensive Monte Carlo simulations to derive the distribution of random-point cell sizes. We avoid this costly step by analytically generating this distribution with defining parameters derived a priori from the number of localizations and field of view size. A discretized cell density map is then generated by interpolating pixel values over a hyperplane defined by coordinates at the localization positions and values as the points’ cell densities. Finally, this map is binarized using the density threshold to produce an image mask where white pixels indicate regions of clustered localizations.

[0198] Additionally, we have developed a graph theory-based approach for explicitly identifying individual clusters from the tessellation results which has linear time complexity as opposed to the computationally intensive n2time complexity of traditional recursive approaches. For example, processing a dataset with 106points, which would require >6 hours of computing on a powerful desktop computer using previous density-based clustering approaches can be processed in just seconds on a laptop using our novel algorithm.

[0199] Cluster masks of different colabeled biomolecules can then be used as inputs to the SPACE framework for a fast and rigorous analysis of the biomolecules’ statistical spatial relationship. Operating on clusters of localizations reduces the impact of overcounting artefacts on SPACE results, translation to discrete space significantly reduces computational overhead and runtimes, and our method for selecting the cluster density threshold avoids the need to define hyperparameters. We have also improved upon previous tessellation-based cluster analyses by including the capability to analyze 3-dimensional data and removing the need for costly Monte Carlo simulations. Thus, SPACE- VorTeCS provides a fast, unbiased, and generalized means to easily analyze the spatial patterns in SMLM data.

[0200] The methods described and exemplified herein can be used for the analysis of SMLM molecular localization data. In addition, automated SMLM image collection was made possible using a machine learning algorithm to assess signal quality during imaging, enabling the microscope to automatically prioritize imaging at sites with good signal quality. This boosted experimental throughput from dozens of cells imaged in a day to over 10,000 cells imaged in a single day. This can enable real-time spatial analysis of SMLM data, including real-time assessments not only of signal quality but also of the sample nanostructure. This can be easily combined into the microscope’s image acquisition software to enable automated high throughput screening of samples based on single molecule -scale measurements, affording experimental capability without precedent. The spatial analysis methods described herein can also be used in non-microscopy fields. Although the development of SPACE- VorTeCS was motivated by microscopy applications, these analytical approaches can be readily applied to any type of data, which can be represented as points in N-dimensional space. Examples of potential application areas include natural language processing, large language models, data mining, medical risk prediction, etc.

[0201] While the methods and systems have been described in connection with preferred embodiments and specific examples, it is not intended that the scope be limited to the particular embodiments set forth, as the embodiments herein are intended in all respects to be illustrative rather than restrictive.

[0202] Unless otherwise expressly stated, it is in no way intended that any method set forth herein be construed as requiring that its steps be performed in a specific order. Accordingly, where a method claim does not actually recite an order to be followed by its steps or it is not otherwise specifically stated in the claims or descriptions that the steps are to be limited to a specific order, it is no way intended that an order be inferred, in any respect. This holds for any possible non-express basis for interpretation, including: matters of logic with respect to arrangement of steps or operational flow; plain meaning derived from grammatical organization or punctuation; the number or type of embodiments described in the specification.

[0203] Throughout this application, various publications may be referenced. The disclosures of these publications in their entireties are hereby incorporated by reference into this application in order to more fully describe the state of the art to which the methods and systems pertain.

[0204] It will be apparent to those skilled in the art that various modifications and variations can be made without departing from the scope or spirit. Other embodiments will be apparent to those skilled in the art from consideration of the specification and practice disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit being indicated by the following claims.

Claims

WHAT IS CLAIMED IS:

1. A method for the analysis of single molecule localization microscopy (SMLM) data comprising a plurality of input data points in the form of molecular coordinates, the method comprising:(i) grouping the plurality of input data points into clusters:(ii) discretizing the clusters into a discretized array; and(iii) measuring a spatial relationship between two species of localizations captured by the discretized array.

2. The method of claim 1 , wherein measuring the spatial relationship between two species of localizations captured by the discretized array comprises point pattern analysis, image processing, or another suitable data analysis method.

3. The method of any one of claims 1-2, wherein step (i) comprises:(a) processing the plurality of input data points to define a study area;(b) computing a tessellation cell for each of the plurality of input data points;(c) removing any tessellation cells that are unbounded or lie outside of the study area;(d) computing an instantaneous spatial density at each of the plurality of input data points using the tessellation cell; and(e) identifying clustered data points as those whose instantaneous spatial density exceeds an analytically defined density threshold.

4. The method of claim 3, wherein step (i) further comprises (f) grouping clustered data points into ensemble point clusters.

5. The method of any one of claims 1-4, wherein the method further comprises adding reference data points to the input data points to improve definition of cluster boundaries, separability of clusters, or a combination thereof.

6. The method of claim 5, wherein the reference data points possess a grid spacing that is contingent upon observed densities of input points within clusters, observed spatial separation between clusters, or a combination thereof.

7. The method of any one of claims 1-6, wherein step (i) comprises: (a) computing a plurality of tessellation cells based on the plurality of input data points; and (b) grouping the clusters based on the plurality of tessellation cells such that there is a correspondence between tessellation cells and clusters.

8. The method of any one of claims 1-7, wherein step (ii) comprises: (a) generating a plurality of grid points; and (b) assigning each of a subset of the grid points to a corresponding tessellation cell.

9. The method of any one of claims 3-8, wherein step (ii) comprises processing the instantaneous spatial density at each of the plurality of input data points using the tessellation cell to produce the discretized array.

10. The method of claim 9, wherein step (ii) comprises generating a map of instantaneous spatial density, cluster membership, or a combination thereof.

11. The method of claim 10, wherein step (ii) further comprises binarizing the map using the density threshold to produce a binary image mask that indicates regions of clustering for the plurality of input data points.

12. The method of claim 10, wherein measuring a spatial relationship between two species of localizations captured by the discretized array comprises evaluating the heatmap.

13. The method of claim 12, wherein evaluating the map comprises point pattern analysis, image processing, or another suitable data analysis method.

14. The method of claim 12, wherein evaluating the map comprises performing discrete point pattern analysis on a pair of cluster species, such as discrete nearest neighbor point pattern analysis.

15. The method of claim 14, wherein the discrete point pattern analysis comprises:(i) providing the pair of cluster species as a first binary discretized array and a second binary discretized array;(ii) calculating a Euclidean distance transformation for the first binary discretized array and the second binary discretized array;(iii) calculating an observed cumulative distribution function (CDF) of interpoint distances for regions of clustered points in the first binary discretized array and the second binary discretized array;(iv) calculating a random cumulative distribution function (CDF) of interpoint distances for the first binary discretized array and the second binary discretized array; and(v) comparing the observed CDF and the random CDF to identify a degree of nonrandom spatial association between the pair of cluster species.

16. The method of claim 14, wherein the discrete point pattern analysis comprises:(i) providing the pair of cluster species as a first binary image mask and a second binary image mask;(ii) calculating a Euclidean distance transformation for the first binary image mask and the second binary image mask;(iii) calculating an observed cumulative distribution function (CDF) of nearest neighbor distances for regions of clustered points in the first binary image mask and the second binary image mask;(iv) calculating a random cumulative distribution function (CDF) of nearest neighbor distances for the first binary image mask and the second binary image mask; and(v) comparing the observed CDF and the random CDF to identify a degree of nonrandom spatial association between the pair of cluster species.

17. The method of any one of claims 15-16, wherein comparing the observed CDF and the random CDF comprises a statistical comparison between the observed CDF and the random CDF.

18. The method of any one of claims 15-17, wherein comparing the observed CDF and the random CDF comprises performing a 2-sided Kolmogorov-Smirnov (KS) test to compare the observed CDF and the random CDF.

19. The method of any one of claims 3-18, wherein computing a tessellation cell for each of the plurality of input data points comprises computing a Voronoi tessellation cell for each of the plurality of input data points.

20. The method of any one of claims 3-19, wherein computing the instantaneous spatial density at each of the plurality of input data points using the tessellation cell comprises calculating a cell volume of each tessellation cell; and computing a tessellation cell density for each tessellation cell as the inverse of its cell size.

21. The method of any one of claims 3-20, wherein the density threshold is analytically defined by distribution theory.

22. The method of any one of claims 3-21 , wherein identifying clustered data points further comprises identifying one or more ensemble clusters.

23. The method of claim 22, wherein the one or more ensemble clusters are defined as connected components of a graph data structure composed of the input data points and their tessellation cell vertices.

24. The method of claim 11 , wherein binarizing the map using the density threshold produces a binary image mask that indicates non-random regions of clustered localization.

25. The method of any one of claims 3-24, wherein computing the instantaneous spatial density at each of the plurality of input data points using the tessellation cell comprises interpolating pixel values over a hyperplane defined by coordinates at a localization position and a cell density value for each clustered data point.

26. The method of any one of claims 1-25, wherein the molecular coordinates comprise two or more dimensional coordinates.

27. The method of any one of claims 1-26, wherein the method avoids the need to define hyperparameters for clustering.

28. The method of any one of claims 3-27, wherein the method for analysis can identify clustering within the plurality of input data points at all spatial scales contained or represented within the input data points.

29. The method of any one of claims 1-28, wherein the method for analysis reduces computational overhead, runtimes, or a combination thereof, as compared to other methods of data analysis, such as methods that utilize Monte Carlo simulations to derive a random distribution for purposes of defining a density threshold.

30. The method of any one of claims 1-29, wherein the method is characterized by less than quadratic time complexity, such as nearly linear time complexity.

31. The method of any one of claims 1-30, wherein the method is more computationally efficient than methods which require explicit evaluation of all inter-point distances between input points.

32. A method for the analysis of single molecule localization microscopy (SMLM) data comprising a plurality of input data points in the form of molecular coordinates, the method comprising:(i) computing a tessellation cell for a subset of the plurality of input data points, wherein each tessellation cell is defined by its vertices;(ii) generating a graph data structure composed of the input data points and corresponding tessellation cell vertices; and(iii) clustering the subset of the plurality of input data points according to the graph data structure.

33. A system for analysis of a data set comprising a plurality of input data points in the form of numbers or multidimensional coordinates, the system comprising a computer, wherein the computer receives the plurality of input data points in the form of molecular coordinates, and a processor of the computer executes computer-executable instructions stored in a memory of the computer to perform the method of any one of claims 1-32.

34. The system of claim 33, further comprising a microscope configured to perform single molecule localization microscopy (SMLM).

35. A method for performing high-throughput single molecule localization microscopy(SMLM), the method comprising: providing a plurality of samples;imaging each of the plurality of samples using SMLM to provide input data points in the form of molecular coordinates for the sample; and analyzing the input data points using the method of any one of claims 1-32.

36. The method of claims 35, further comprising utilizing the results of such analyses to drive feedback control of the SMLM system.

Citation Information

Patent Citations

  • High resolution dual-objective microscopy

    US20140333750A1

  • Method for detecting spatial coupling

    US20230237664A1

  • Method & apparatus for processing microscope images of cells and tissue structures

    WO2022229638A1