Systems and methods for high throughput single molecule tracking in living cells

By using variational Bayesian optimization algorithm, Gibbs sampling algorithm or adaptive hill climbing algorithm to generate possible trajectories of molecules and their related probabilities, combined with a microscope system and a high-throughput sample processing robot system, the problem of limited application scale of SMT in living cells is solved, and efficient molecular tracking and drug screening are achieved.

CN120826720APending Publication Date: 2025-10-21AIKANG THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380094538.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-22
Filing Date
2023-12-21
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing single-molecule tracking technologies (SMT) have limited application scale in living cells, making it difficult to implement high-throughput settings for system-level screening or drug discovery.

Method used

The variational Bayesian optimization algorithm, Gibbs sampling algorithm or adaptive hill climbing algorithm are used to generate possible trajectories of molecules and their associated probabilities, combined with a microscope system and a high-throughput sample processing robot system to achieve efficient image acquisition and analysis.

Benefits of technology

It provides a fast and computationally efficient molecular tracking method that can process large amounts of data with minimal human supervision, generate built-in confidence indicators and dynamic indicators, and support drug screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120826720A_ABST
    Figure CN120826720A_ABST
Patent Text Reader

Abstract

Systems and methods for high throughput single molecule tracking within living cells receive a sequence of images of visualized molecular motion. Molecules across the image are linked. Possible trajectories and associated probabilities for each molecule are generated using a variational Bayesian optimization algorithm and based on the links. Data characterizing the generated possible trajectories and their associated probabilities are provided to the consumer application or process.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Application No. 63 / 476,946, filed on December 22, 2022, and U.S. Provisional Application No. 63 / 476,941, filed on December 22, 2022, the contents of each of which are incorporated herein by reference in their entirety. Technical Field

[0003] The subject matter described herein relates to a platform for tracking single molecules within complex systems. Background Art

[0004] In the crowded environment of living cells, the movement of proteins is profoundly influenced by their interactions with the surrounding environment. Single-molecule tracking (SMT) is a method to capture protein movement as a reporter of activity. In SMT, a fluorescent protein of interest is imaged with high spatiotemporal resolution to track its movement in complex systems such as living cells. The information embedded in these trajectories has been used to study various cellular phenomena, including protein-protein interactions, such as those that mediate signal transduction, inter-organelle communication, nuclear organization, and transcriptional regulation. However, the application scale of SMT technology is limited and it has therefore been mainly used to address specific mechanistic hypotheses. For example, SMT has not yet been adapted to a throughput setting that can achieve system-level screening or drug discovery. Summary of the Invention

[0005] In a first aspect, a series of images visualizing molecular motion is received. Molecules across the images are linked. Based on the linkage, a variational Bayesian optimization algorithm is used to generate possible trajectories and associated probabilities for each molecule. Data representing the generated possible trajectories and their associated probabilities is provided to a consuming application or process.

[0006] In a related aspect, a series of images visualizing molecular motion is received. Molecules across the images are linked. Using a Gibbs sampling algorithm and based on the linkages, possible trajectories for each molecule and their associated probabilities are generated. Data representing the generated possible trajectories and their associated probabilities is provided to a consuming application or process.

[0007] In another aspect, a series of images visualizing molecular motion is received. Molecules across the images are linked. Based on the links, a possible trajectory and its associated probability are generated for each molecule using an adaptive hill climbing algorithm. Data representing the generated possible trajectories and their associated probabilities is provided to a consuming application or process.

[0008] At least a subset of the sequence of images may include at least 100 molecules per image; while in other variations, at least 1000 molecules per image; and in still other variations, at least 10,000 molecules per image.

[0009] In some variations, the molecules may have a density of at least 0.01 emitters per square micron per image, and in some variations at least 0.1 emitters per square micron per image.

[0010] Molecules within a biological sample can be labeled, the labeled biological sample can emit fluorescence, and an image sequence can be generated while the biological sample emits fluorescence.

[0011] A microscope system may be used to generate the image sequence.

[0012] The molecules can be imaged in living cells.

[0013] A probabilistic kinetic model comprising information characterizing the molecular trajectories may be inferred. The probabilistic kinetic model may include a state array, and the state array may be populated with information characterizing the molecular trajectories.

[0014] An internal confidence indicator based on the associated probability may be generated, and the provided data may include the generated internal confidence indicator. The generated internal confidence indicator may be a tracking error rate lower bound that defines a lower bound on the rate of incorrect connections caused by the link. The generated internal indicator may include calculating a confidence level for each track.

[0015] Additionally, dynamic indicators can be generated independent of a specific trajectory.

[0016] Linking may include retrieving data with multiple statistics extracted from the total number of detections or the number of detections in a cell.

[0017] Providing the data may include one or more of: visualizing at least a portion of the generated possible trajectories and their associated probabilities in a graphical user interface, storing at least a portion of the generated possible trajectories and their associated probabilities in a physical persistence, loading at least a portion of the generated possible trajectories and their associated probabilities in a memory, or transmitting at least a portion of the generated possible trajectories and their associated probabilities to a remote computing device over a network.

[0018] At least a portion of the sequence of images may comprise consecutive images from the respective film.In other variations, at least a portion of the sequence of images used for linking are non-consecutive images from the respective film.

[0019] In another aspect, a series of images visualizing molecular motion may be received. The image sequence may include a first type generated using a first imaging modality and a second type generated using a second, different imaging modality. Points may be detected in the first type of image sequence. The points detected in the first type of image sequence may be linked into tracks using a probabilistic tracking algorithm. The second type of image sequence may be segmented to generate a plurality of instance masks. Molecules within the second type of image sequence may be assigned to at least one of the plurality of instance masks. Data representing the linking and assignment may be provided to a consuming application or process.

[0020] Probabilistic pursuit algorithms can take different forms, including variational Bayesian optimization, Gibbs sampling, or adaptive hill climbing.

[0021] The first imaging modality and the second imaging modality may comprise different molecular labeling technologies.

[0022] The first type of image sequence may be a single molecule tracking (SMT) movie, and the second type of image sequence may be a non-SMT movie.

[0023] The detected spots may include subcellular components.

[0024] The molecular types within the sequence of images of the first type may be labeled with different fluorophores.

[0025] A plurality of statistical metrics associated with the at least one track or the at least one instance mask may be generated.

[0026] A hierarchy that can store instance masks.

[0027] Detection may utilize one or more of: a generalized log-likelihood point detector, a Difference of Gaussian (DoG) detector, a Laplace of Gaussian (LoG) detector, or a Determinant of Hessian (DoH) blob detector.

[0028] Sub-pixel localization can be used to associate detected points with spatiotemporal coordinates.

[0029] Sub-pixel localization may include one or more of the following: a radially symmetric localizer or maximum likelihood fitting of candidate point models using the Levenberg-Marquardt method.

[0030] In some variations, the image sequence may be generated by a fluorescence microscopy apparatus. The apparatus may include a light source, a first optical element or assembly, a second optical element or assembly, and a detector device. The light source is capable of emitting fluorescence excitation light. The light source may exhibit a power output drift of less than about 10% at an ambient temperature of 17°C + / - 5°C. The first optical element or assembly may be configured to receive the fluorescence excitation light source and shape the fluorescence excitation light source to form a light beam. The second optical element or assembly may include a water immersion objective lens configured to tilt the light beam relative to the z-axis in the xz plane. The second optical element may also be configured to focus the light beam at a sample plane located in the xy plane, thereby illuminating at least a portion of the sample plane. The detector device may be configured to receive light from the illuminated portion of the sample plane. The detector device may form one or more projected images based on the light received from the illuminated portion of the sample plane.

[0031] The apparatus may comprise a second objective lens configured to direct light emitted from the illuminated portion of the sample plane to the detector arrangement.

[0032] The detector means may comprise a semiconductor sensor.

[0033] The apparatus may include a third optical element or assembly configured to translate the light beam in an imaging plane in a direction orthogonal to the longer dimension of the light beam.

[0034] The third optical element or assembly may include a galvanometer mirror.

[0035] The detector arrangement may comprise a semiconductor sensor.The detector arrangement may support a shutter mode for synchronizing the translation of the light beam in the sample plane with the selective activation or readout of the semiconductor sensor.

[0036] A microscope system can generate a sequence of images to track the movement of molecules. The microscope system can include a stage, a light source, a water immersion objective, and a detector device. The stage can support a sample. The sample can contain molecules. The light source can emit a light beam capable of inducing a photo-based reaction from the molecules in the sample. The light source can exhibit a power output drift of less than about 10% at an ambient temperature of 17°C + / - 5°C. The water immersion objective can focus the light beam on at least a portion of a sample plane. The molecules can be positioned in the sample plane. The detector device can monitor the photo-based reaction of the molecules and can analyze it to track the movement of the molecules.

[0037] The microscope system may include a scanning optical element or assembly configured to translate the light beam in the sample plane in a direction orthogonal to the longer dimension of the light beam, thereby enabling the microscope system to have a larger total field of view in the xy plane.

[0038] The microscope system may include a z-position controller for the sample plane. The z-position controller is capable of maintaining focus in the z-direction.

[0039] The samples may be placed in open wells of a sample plate having a plurality of open wells.

[0040] The microscope system may include an xy position controller for changing the field of view of the microscope system. The changed field of view may encompass a different subset of the plurality of open wells.

[0041] The microscope system may include a temperature controlled environment configured to control the environment of the sample plate.

[0042] Samples can be placed in the open wells of the sample plate and maintained at 20%-95% humidity.

[0043] Samples can be placed in open wells of the sample plate and maintained at 5% CO2.

[0044] The microscope system may also include an automated sample handling robotic system to enable high-throughput manipulation of multiple samples on the stage. The robotic system may include a memory and a processor in communication with the memory. The robotic system may also include one or more robotic end effectors in communication with the processor. The one or more end effectors may manipulate the multiple samples on the stage based on communication with the processor.

[0045] Also described are non-transitory computer program products (i.e., physically embodied computer program products) that store instructions that, when executed by one or more data processors of one or more computing systems, cause at least one data processor to perform the operations described herein. Similarly, described are computer systems that may include one or more data processors and a memory coupled to the one or more data processors. The memory may temporarily or permanently store instructions that cause at least one processor to perform one or more operations described herein. In addition, the methods may be implemented by one or more data processors within a single computing system or distributed between two or more computing systems. Such computing systems may be connected and may exchange data and / or commands or other instructions, etc., via one or more connections, including but not limited to connections via a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), via direct connections between one or more of the plurality of computing systems, etc.

[0046] The subject matter described herein provides numerous technical advantages. For example, the current subject matter provides a fast and computationally more efficient method for tracking targets (e.g., molecules, etc.) that isolates interpretable information from background and other interfering noise. The subject matter described herein can be used to analyze complex data, which may include thousands to tens of thousands of rapidly moving targets in close proximity. Additionally, an advantage of the subject matter described herein is that it can be performed with little to no human supervision.

[0047] More specifically, the current subject matter offers numerous technical advantages related to scalability. Current platforms can generate data for over 100 molecules per frame (i.e., image) across multiple imaging systems running in series. This capability requires a tracking method that is: (1) highly scalable; and (2) provides built-in confidence / diagnostic metrics in the tracking results, as there is no human supervision of the raw data. Furthermore, the probabilistic tracking algorithm presented herein provides built-in confidence metrics for consuming applications / processes without the need for human supervision. Furthermore, the current probabilistic tracking algorithm provides dynamic metrics that can be used for drug screening independent of any specific trajectory.

[0048] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the detailed description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] This patent or application file contains at least one drawing drawn in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0050] Figure 1 Schematic diagram depicting the htSMT workflow.

[0051] Figure 2 Depicted is a schematic diagram of an exemplary image acquisition system of the present disclosure.

[0052] Figure 3 Depicted are side cross-sectional views of exemplary illumination schemes, particularly highly inclined and laminated optical sheet microscopy (HILO), which may be used in conjunction with certain aspects of the present disclosure.

[0053] Figure 4 Depicts a side cross-sectional view of an exemplary lighting scheme ( Figure 4 , top image), including HILO, highly tilted swept tile (HIST) microscopy, and single objective inclined light sheet (SOLEIL) microscopy, which can be used in conjunction with certain aspects of the present disclosure. Figure 4 The bottom image schematically depicts the camera view for each corresponding lighting scheme.

[0054] Figure 5 Depicted is a schematic diagram of an exemplary sample processing system of the present disclosure.

[0055] Figure 6 An exemplary system of a high-throughput single-molecule imaging platform for measuring protein motion within living cells is illustrated.

[0056] Figure 7 The data flow through an exemplary system of a high-throughput single-molecule imaging platform for measuring protein movement within living cells is illustrated.

[0057] Figure 8 are multiple images illustrating the difference between mask categories and instance / semantic masks.

[0058] Figure 9 An example computer-implemented environment related to the subject matter described herein is described.

[0059] Figure 10 is a diagram illustrating an example computing device architecture for implementing various aspects described herein.

[0060] Figure 11 Diagram illustrating the pipeline of scalable tracing in htSMT.

[0061] Figures 12A-12C Depicts the benchmarks of various tracking algorithms, among which Figure 12A Aspects relevant to optical dynamic simulations are described, Figure 12B The basis for comparison is stated, and Figure 12C Benchmark results on recall, precision, and F1 score are illustrated.

[0062] Figures 13A-13E Aspects of the high-throughput single-molecule tracking platform are described. Figure 13A Depicted are the diffusion state probability distributions for three cell lines expressing histone H2B-Halo, Halo-CaaX, or Halo alone, where the shaded areas represent the diffusion state unique to each cell line. Figure 13B Depicted is an example field of view, including single-molecule images of mixed H2B-Halo, Halo-CaaX, and free Halo cell lines (top) and a reference Hoechst image (bottom), along with insets showing magnifications of consecutive frames of single cells and single molecules, where image intensities are equivalently scaled across the panels. Figure 13C Describes from Figure 13B The single cell diffusion profiles extracted from Figure 13A Colored by similarity to the H2B-, CaaX-, or free Halo dynamics reported in . Figure 13DDepicted is a heatmap representation of 103,757 nuclei measured from a mixture of Halo-H2B, Halo-CaaX, or free Halo in each of 1,540 unique wells across five 384-well plates, where each horizontal line represents a nucleus. Cells were clustered using k-means clustering and were based on Figure 13A Assign labels based on the diffusion profiles determined in . Figure 13E Depicted are the collective state arrays of all trajectories recovered from a mixture of Halo-H2B, Halo-CaaX, and free Halo cells.

[0063] Figures 14A-14D A dynamic comparison of a panel of steroid hormone receptors (SHRs) is provided. Figure 14A Schematic representation of SHR function is depicted, wherein under basal conditions, SHR is sequestered in complex with HSP90 and other cofactors, and upon ligand binding, the receptor dissociates from the inactive complex, dimerizes, and binds to DNA. Figure 14B Depicts the diffusion state distribution of Halo-AR, Halo-ER, Halo-GR, and Halo-PR in U2OS cells before and after stimulation with activating ligands, where the area in the shaded region is f 结合 , where the shaded error bars represent SD. Figure 14C The selectivity of individual SHRs for their cognate ligands compared to other steroids is depicted, as shown by f 结合 Determined, where error bars represent SEM. Figure 14D The selectivity of individual SHRs for their cognate ligands compared to other steroids is depicted, as shown by D 自由 Determined, where error bars represent SEM.

[0064] Figure 15 The results show that the screening of bioactive compounds targeting estrogen receptor (ER) showed reproducibility and robustness. The reproducibility of the screening results was assessed by two biological replicates. The figure shows the f of ER for 5067 compounds. 结合 If the average f of the magenta compound in two repeated experiments 结合 Of the 30 expected positive control compounds, 26 significantly increased f in both replicates. 结合 (grey outline). The three compounds increased f 结合 , but only appeared in one replicate after filtering. Exemplary positive controls include the agonist estradiol (1) and the antagonists fulvestrant (5), 4-OHT (6), and bactroxifene (7). Linear regression fits were performed on the magenta compound to determine slope and correlation.

[0065] Figures 16A-16GWe show that selective ER modulators and degraders induce DNA binding that can be measured by SMT. Figure 16A Depicted are the diffusion state probability distributions of ER treated with 100 nM of exemplary selective ER modulators (SERMs) and selective ER degraders (SERDs), where the shaded area represents the SD. Figure 16B Depicts f 结合 Changes with 12pt dose titration of fulvestrant (5), 4-OHT (6), GDC-0810 (8), AZD9496 (9), or GDC-0927 (10), where the fitted curves and compound colors are as shown Figure 16A Shown, where error bars represent SEM of three replicates. Figure 16C Depicts the effect of f after addition of agonist or antagonist 结合 The change over time is fitted with a single exponential, and the color of the compound is as Figure 16A Shown are head and neck cancer estradiol (1, green) and DMSO for comparison, and error bars represent SEM. Figure 16D Depicts the SERM and SERD interactions 结合 The maximum effect of , where each box represents a quartile and the whiskers represent the 5th–95th percentiles of single-well measurements, measured over at least four days with at least eight wells per day for each compound. Figure 16E Fluorescence recovery after photobleaching (FRAP) of ER-Halo cells treated with DMSO alone or 100 nM SERM / D is depicted, where the curves are the mean ± SEM of 18-24 cells. Figure 16A Coloring is shown, and error bars represent SEM. Figure 16F Quantification of FRAP recovery curves is depicted to measure recovery 2 min after photobleaching, where whiskers represent the 5th–95th percentiles of single-cell measurements. Figure 16G Track length survival curves for ER-Halo cells treated with DMSO alone or with SERM / D are depicted, where track survival is plotted as the 1-CDF of the track length distribution, with faster decay indicating shorter binding times. The inset depicts Figure 16G Quantification of the middle curve, each point represents the fraction of trajectories longer than 10 s in a single biological replicate consisting of 3–10 wells per condition, and the dashed line represents the median fraction of trajectories lasting longer than 10 s in histone H2B-Halo, which represents the upper limit of measurement sensitivity.

[0066] Figures 17A-17D This demonstrates that htSMT can be used to determine chemical structure-activity relationships. Figure 17ADepicted is an example of GDC-0927-induced ER degradation in different cell lines measured by immunofluorescence against the ER, where cells were exposed to the compound for 24 hours before fixation, and image pixel intensities are equivalently scaled. Figure 17B Depicted are the correlations of potency as measured by ER degradation and cell proliferation in MCF7 cells (black) and T47d cells (magenta) for compounds in the GDC-0927 structural series. Figure 17C Depicts the f of compounds 11 to 16 during a 12pt dose titration. 结合 Points are the mean ± SEM of three biological replicates. Figure 17D Depicted for compounds in the GDC-0927 structural series, by f 结合 Figure 3. Changes in α-glucose and potency correlations of α-glucose and α-glucose in MCF7 cells (black) and T47d cells (magenta).

[0067] Figures 18A-18F It is demonstrated that bioactive molecules targeting ER-related pathways affect ER dynamics. Figure 18A Depicted are the bioactivity screening results, where selected inhibitors are grouped by pathway and uniquely colored. Figure 18B Depicted are three representative compounds targeting each of HSP90, mTOR, CDK9, and proteasome during a 12pt dose titration. 结合 , where individual molecules are represented by specific shapes and error bars represent SEM. Figure 18C Depicts the compound after addition of f 结合 Time course of estradiol treatment (black) compared to gantapib (blue circles) and HSP990 (blue squares), dots are 4 min fractions, and error bars represent SEM. Figure 18D Depicts the compound after addition of f 结合 Time course of estradiol treatment (black) compared with the HSP90 inhibitors gantapib (blue circles) and HSP990 (blue squares); and the proteasome inhibitors bortezomib (red circles) and carfilzomib (red squares). Points are sections of 7.5-min htSMT data, with mean ± SEM labeled, and the shaded area indicates the time window used during the htSMT screen. Figure 18E Track length survival curves are depicted for ER-Halo cells treated with DMSO alone, stimulated with estradiol, or treated with 100 nM HSP90 or proteasome inhibition, where track survival is plotted as the 1-CDF of the track length distribution; faster decay means shorter binding time, and where the number of tracks with a length greater than 2 is quantified inset according to treatment condition, where all conditions are normalized to the median number of tracks in DMSO. Figure 18FA diagram outlining pathway interactions based on htSMT results for ER, AR, and PR is provided.

[0068] Figures 19A-19B Provides characterization of htSMT system performance. Figure 19A Depicted are the distributions of localization errors measured in multiple independent wells using histone H2B-Halo cells, with a median localization error of 39 nm. Figure 19B Depicted are the number of measurable nuclei within a 94 × 94 μm FOV, where each boxplot represents the distribution for each well in a 384-well plate.

[0069] Figures 20A-20E It was shown that the screening of bioactive molecules produced robust data with good assay performance. Figure 20A Depicted are box plots showing the number of nuclei measured for each test compound, where the whiskers are the 1st and 99th percentiles. Figure 20B Depicted is the Z' factor analysis of a bioactivity screen, where each dot represents the Z' factor of a single 384-well plate, measuring the difference between DMSO and 25 nM estradiol treatment, and plates with very low Z' factors were removed from further analysis. Figure 20C Depicted are the estradiol dose titrations for each plate in the bioactive molecule screen, which were fitted with logistic regression to determine potency. Figure 20D Describes from Figure 20C EC50 values ​​extracted from each curve fit are shown in Figure 5, where the shaded area is the three-fold range of potency. Figure 20E The f in the control DMSO wells is depicted. 结合 The distribution of changes.

[0070] Figures 21A-21E This indicates that SERM and SERD quickly reduce free diffusion and increase 结合 . Figure 21A Plotted are the normalized occupancies of the diffusive states (diffusion coefficients 0.2–100 μm) for samples treated with 100 nM SERD or SERM compared to those treated with DMSO. 2 The histograms were normalized to an integral of 1, the shaded area represents the bin-by-bin standard deviation, and the curves are the averages of 3-4 biological replicates with 8 wells per condition. Figure 21B Depicts the Figure 17C The results of single exponential correlation fitting of the data in . Figure 21C Depicted are the changes in the diffusion state distribution over time after the addition of estradiol and fulvestrant, where each curve represents the mean of three biological replicates and the shaded area is the bin-by-bin standard deviation. Figure 21D Depicted are the f of selective AR or GR antagonists from a bioactive screening set. 结合 change. Figure 21EProvides the slow decay rate constant (k 慢 ), where the fit is performed on a set of three biological replicates and compared with Figure 17D f determined in 结合 Together, the inferred k on The upper limit of can be calculated by the equation, assuming that all binding molecules have k* off =k 慢 , and cells marked with an asterisk are cells that cannot reliably convert k 慢 Cells distinguished from photobleaching.

[0071] Figures 22A-22D GDC-0927 structural variants characterized by ER degradation or cell proliferation assays are provided. Figure 22A Western blot depicts ER expression in ER-expressing breast cancer cell lines MCF7 and T47d compared to ER-null lines SK-BR-3 or U2OS, where samples were treated with fulvestrant for 24 hours prior to lysis, and fulvestrant treatment resulted in ER degradation, even when fused to HaloTag. Figure 22B An exemplary compound dose titration is depicted, showing the change in mean nuclear intensity as a function of compound concentration. Figure 22C Depicted are examples of the effects of GDC-0927 on the proliferation of MCF7 breast cancer cells compared to staurosporine as a positive control. Figure 22D An exemplary compound dose titration is described, measuring cell proliferation of cells treated with GDC-0927 analogs (normalized to DMSO-treated cells).

[0072] Figures 23A-23D Additional studies of bioactivity screening data are provided, including non-limiting exemplary cutoffs for active molecules. Figure 23A Provides information on the effects of compounds on ERf in dose titration experiments 结合 Example of the effect of , where 92 compounds from the primary screen are ranked based on their effect size, with black compounds reproducibly showing dose-dependent changes; magenta compounds were inactive in the dose titration. Figure 23B Depicts ERf 结合 Exemplary dose titration of a compound with increasing overall change. Figure 23C Depicts Figure 15 Two replicates of the bioactivity screen in which active compounds were colored based on the structural scaffold. Figure 23D Quantification of the number of compounds associated with any given cluster is depicted, with singletons (cluster 21) representing the majority of active compounds.

[0073] Figures 24A-24D This suggests that some pathway inhibitors that regulate ER dynamics are ER-specific. Figure 24AFigure 3 depicts the f of ER after treatment with HSP90, proteasome, mTOR, and CDK inhibitors from the initial bioactivity screen. 结合 Figure 4 shows the changes in the expression of CDKs, where CDK inhibitors are further stratified into CDK4 / 6 inhibitors, CDK9 inhibitors, or inhibitors without strong selectivity for a specific family (pan-CDK), and the lines represent the median value for each target. Figure 24B The f of ER for 97 bioactive molecules was mapped 结合 Changes are colored by their pathway annotation, where compounds are plotted in their rank order based on their effect in the ER, and error bars are the SEM of three biological replicates. Figure 24C Depicts Figure 24B f of the same molecule AR 结合 Compounds are plotted in their rank order based on their effect in the ER, and error bars are the SEM of three biological replicates. Figure 24D Depicts Figure 24B f of the PR of the same molecule 结合 Compounds are plotted in their rank order based on their effect in the ER, and error bars are the SEM of three biological replicates.

[0074] Figure 25 Depicted are dose titration graphs of ER (S104A / S106A / S118A) compounds with mTOR and CDK9. Each point represents the mean and SEM of three biological replicates.

[0075] Figure 26 Depicted are the f of AR measured in the presence of 25 nM agonist, 10 μM potent antagonist, or a combination of agonist and antagonist at 25 nM and 10 μM, respectively. 结合 .

[0076] Figure 27 Depicted are changes in target A motility in response to dose titration of known target A antagonists, where each series of shapes represents a different compound and error bars represent SEM.

[0077] Figure 28 Illustrated are two representative examples of receptor tyrosine kinases (target B and target C) whose motility changes in response to dose titration of known ATP-competitive or allosteric inhibitors. Each series of shapes represents a different compound normalized to DMSO, and error bars represent SEM.

[0078] Figure 29Depicted are the changes in Q3 jump length over time after compound addition for additional representative targets. Within minutes of compound addition, treatment of the protein with on-target inhibitors increases or decreases protein motion for target B or target A, respectively. Treatment of the helicase with pathway antagonists or off-target modulators (e.g., in the case of DNA damage induction) causes changes in protein motion that take several hours or longer to reach maximal effect.

[0079] Figure 30 Depicted is the use of 10 pM JF 549 Cumulative number of trajectories per individual cell type labeled versus imaging time. Shaded error bars represent 1 standard deviation.

[0080] Figure 31 Depicted are the mRNA transcript levels of AR, ESR1, PR, and NR3C1 (expressed as log(FPKM)) in the engineered U2OS cell line compared to the parental U2OS cell line and three reference breast cancer cell lines.

[0081] Figure 32 Depicted are the reference gene sets induced after 24-hour stimulation with 25 nM estradiol. The top five induced gene sets were not significantly induced upon ectopic expression of Halo-ER in the parental U2OS line.

[0082] Figure 33 Depicted are bar graphs showing the -log(q-values) of the top 50 most significantly induced gene sets after estradiol stimulation. Gene sets characteristic of ESR1 or estrogen response are marked in dark grey.

[0083] Figure 34 Depicted is the use of 20 pM JF 549 Cumulative number of tracks of labeled DMSO-treated cells versus imaging time. Shaded error bars represent 1 standard deviation.

[0084] Figure 35 Depicts Figure 15 The f of the 30 known ER-interacting molecules circled in 结合 Changes, sorted by effect size. Error bars are SEM.

[0085] Figures 36A-36D Depicts ER( Figure 36A )、PR( Figure 36B )、AR( Figure 36C ) and GR( Figure 36D Antagonists of α-glucose agonist ... 结合 The dashed line shows the f of the vehicle-treated control. 结合 . DETAILED DESCRIPTION

[0086] The subject matter of the present disclosure relates to the development of the first industrial-scale high-throughput SMT (htSMT) technology; systems incorporating such htSMT technology; hardware and software associated with such htSMT technology; and methods of using such htSMT technology. For example, the htSMT technology described herein is capable of measuring protein movement in >1,000,000 cells per day. Additionally, using the estrogen receptor (ER) as a proof-of-concept system, the htSMT technology described herein demonstrated specific, robust, and reproducible results. The htSMT technology described herein can be used for a variety of applications, including but not limited to drug discovery activities, such as compound library screening and the elucidation of structure-activity relationships (SAR). Importantly, the htSMT technology described herein can be used to characterize the contributions of known and novel pathways to larger molecular assemblies comprising a target, such as protein signaling interaction networks.

[0087] refer to Figure 1 , various aspects of the current subject matter can be implemented using the htSMT workflow. The workflow can include various stages, as described in further detail below, such as (i) sample preparation including reagent treatment; (ii) image acquisition using sample imaging to generate a series of images and / or videos; (iii) image analysis by processing these images and videos, such as using various analyses, single emitter detection and sub-pixel localization (i.e., "super-resolution imaging"), tracking, computer vision, and machine learning algorithms; (iv) storage of information extracted from or otherwise representing or comprising the images and videos (i.e., features, original images, modified images, etc.); and (v) using the stored information to provide insights, including biological interpretations (which can be provided additionally or alternatively using various analyses, tracking, computer vision, and machine learning algorithms).

[0088] The subject matter of the present disclosure is described with reference to the accompanying drawings, in which reference numbers are used throughout to indicate similar or equivalent elements. The drawings are not drawn to scale and are provided solely for the purpose of illustrating the aspects disclosed herein. Several disclosed aspects will be described below with reference to exemplary hardware, software, and applications for illustration. It should be understood that many specific details, relationships, and methods are set forth in order to provide a more complete understanding of the subject matter disclosed herein. For the purpose of clarity of disclosure and not for limitation, the detailed description is divided into the following subsections:

[0089] 1. Definition

[0090] 2.htSMT hardware

[0091] 3.htSMT software

[0092] 4. Specific htSMT applications

[0093] 5. Exemplary Implementation

[0094] 6. Examples

[0095] 1. Definition

[0096] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those of ordinary skill in the art are generally understood. In the event of conflict, this document (including definitions) shall prevail. Preferred methods and materials are described below, although methods and materials similar or equivalent to those described herein may be used when practicing or testing the subject matter of the present disclosure. All publications, patent applications, patents and other references mentioned herein are incorporated by reference in their entirety. The materials, methods and examples disclosed herein are illustrative only and are not intended to be limiting.

[0097] As used herein, the terms "including," "comprising," "having," "has," "may," "containing," and variations thereof are intended to serve as open transitional phrases, terms, or words that do not exclude the possibility of additional acts or structures. The singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. The present disclosure also contemplates additional embodiments that "comprise," "consist of," and "consist essentially of" the embodiments or elements presented herein, whether or not explicitly stated.

[0098] For the recitation of numerical ranges herein, each intervening number within the range is expressly contemplated with equal precision. For example, for a range of 6 to 9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for a range of 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are expressly contemplated.

[0099] As used herein, the term "about" or "approximately" means that a particular value is within an acceptable error range as determined by one skilled in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within 3 or more standard deviations, as practiced in the art. Alternatively, "about" can mean a range of up to 20%, preferably up to 10%, more preferably up to 5%, and more preferably up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5-fold, and more preferably within 2-fold, of a value.

[0100] As used herein, the term "trajectory" refers to a temporally linked set of spatial coordinates corresponding to the observed positions of a fluorescent protein. In some cases, multiple trajectories can be algorithmically constructed by linking multiple fluorescent proteins whose positions are determined at consecutive time points. In some cases, when no other linking is feasible, multiple trajectories can be constructed conservatively by linking only points within a fixed search radius. In some cases, multiple trajectories can be constructed probabilistically.

[0101] As defined herein, protein motion refers to the change in position of multiple fluorescent proteins. In some cases, protein motion can be quantified by analyzing changes in spatial coordinates at consecutive time points. Motion characterized in this manner may include, but is not limited to, measurement of jump length distribution. Given a set of protein displacements between one time point and a subsequent time point, a histogram of the probability of each displacement length ("jump length") can be constructed. Quantiles of this distribution can be used to describe protein motion. In some cases, the quantile used is the median of the jump length distribution. In some cases, the quantile used is the 3rd quartile of the jump length distribution. In some cases, protein motion can be quantified by trajectory analysis. Motion characterized in this manner may include, but is not limited to, measurement of mean square displacement, which is defined by the average of the squares of all displacements in a trajectory averaged over multiple trajectories. Motion characterized in this manner may also include, but is not limited to, measurement of trajectory length or trajectory length distribution. Motion characterized in this manner may also include, but is not limited to, measurement of the mean radius of gyration, which is defined by the root mean square distance between all coordinates in a trajectory and the center of mass of the set of points contained in the trajectory, averaged over multiple trajectories. Motion characterized in this manner may also include, but is not limited to, measurement of an average bond angle, defined as the angle formed by three consecutive spatial coordinates averaged over multiple trajectories. Motion characterized in this manner may also include, but is not limited to, measurement of a maximum likelihood estimator of the diffusion coefficient, defined as an estimate of the maximum likelihood diffusion coefficient for multiple trajectories under a single-state diffusion model with constant positioning error. In some cases, protein motion may be measured by analyzing the product of a link generation algorithm. Motion characterized in this manner may include, but is not limited to, the average posterior diffusion coefficient, the average of the posterior probability distribution of the coefficients from the probabilistic link algorithm. Motion characterized in this manner may include, but is not limited to, the geometric mean posterior diffusion coefficient, the average of the logarithmic scale posterior probability distribution of the coefficients from the probabilistic link algorithm. In some cases, protein motion may be measured by performing model correlation analysis on multiple trajectories. Motion characterized in this manner may include, but is not limited to, the fraction of immobile molecules ("f") defined by a two-state model fit. 结合 ”).

[0102] As used herein, the term "motion" encompasses changes in the direction of travel of a target as well as changes in speed (increase or decrease). Thus, in some cases, tracking motion can include determining that the target has not moved, for example, when the target is in or substantially in a static binding state. Motion can be characterized in a variety of ways, including but not limited to quantifying: (a) the median of the jump length distribution (where the jump length corresponds to the observed distance traveled by the target fluorescent protein in consecutive frames); (b) the 3rd quartile of the jump length distribution; (c) the median radius of gyration; (d) the mean posterior diffusion coefficient; (e) the geometric mean posterior diffusion coefficient; (f) the mean square displacement; (g) the median bond angle; (h) the maximum likelihood estimator of the diffusion coefficient; (i) the trajectory length; and / or (j) the state occupancy by inference.

[0103] As used herein, detected motion (including but not limited to any change in motion) can occur in response to any environmental or other factor. For example, and not by way of limitation, motion or lack thereof can be induced by: (A) addition of a compound; (B) temperature change; (C) change in oxygen concentration, such as introduction of hypoxic conditions; (D) mechanical stress; (E) pH change; and / or (F) change in light (e.g., increasing or decreasing intensity).

[0104] As used herein, the term "fluorescent protein" refers to any protein that emits a fluorescent signal. In some cases, fluorescent emission occurs under irradiation with light of a specific wavelength. An example of a naturally occurring fluorescent protein is green fluorescent protein (GFP). However, in some cases, the protein of interest may be adapted to emit a fluorescent signal by introducing an encoded fluorescent tag, that is, the protein sequence is fused to the protein of interest so that it emits fluorescence. In some cases, the protein of interest may be adapted to emit a fluorescent signal by binding to a fluorescent ligand. Non-limiting examples of such encoded fluorescent tags include: Halo tags, SNAP tags, CLIP tags, TMP tags, and SunTags. Additionally or alternatively, the protein of interest may be adapted to emit a fluorescent signal by coupling to a fluorescent dye molecule (e.g., an amine or thiol-reactive dye).

[0105] As used herein, the term "compound" refers to any chemically defined entity. In some cases, the compound can be a molecule less than 1000Da, i.e., a "small molecule". In some cases, the compound can be a macromolecule, such as a nucleic acid. In some cases, the nucleic acid can have a defined sequence. In some cases, the nucleic acid includes: (A) ribonucleic acid (RNA), including, for example, modified RNA; (B) deoxyribonucleic acid (DNA), including, for example, modified DNA; and (C) a combination of (A) and (B). In some cases, the nucleic acid will be a single-stranded or double-stranded small interfering nucleic acid (e.g., double-stranded siRNA), an antisense oligonucleotide, a ribozyme, a microRNA, or an aptamer. In some cases, the compound can be a protein. For example, but not limited to, the protein compounds of the present disclosure encompass signal transduction proteins, such as protein hormones, cytokines, kinases, phosphatases, and other enzymes and transcription factors, as well as antibodies, contractile proteins, structural proteins, storage proteins, and transport proteins.

[0106] In some cases, a compound may refer to a mixture of molecules, such as a mixture having a defined composition.

[0107] Throughout the drawings and description, certain numbers are associated with certain compounds, e.g., see Figure 14B 、 15 , 16A-16C, 17C and Example 1. Specifically, the compounds and their respective numbers are: estradiol (1); 2-hydroxytestosterone (2); progesterone (3); dexamethasone (4); fulvestrant (5); 4-OHT (6); baxifen (7); GDC-0810 (8); AZD9496 (9); and GDC-0927 (10). In addition, Figure 17C A series of GDC-0927 structures are included, in which specific modifications to GDC-0927 are described and numbered from (11) to (16).

[0108] 2.htSMT hardware

[0109] Image acquisition system

[0110] refer to Figure 1 , various aspects of the current subject matter can be implemented using an htSMT workflow, wherein such workflow is combined with an image acquisition system that uses sample imaging to generate a series of images and / or videos. For example, and not by way of limitation, Figure 2Depicted is a schematic diagram of an exemplary image acquisition system of the present disclosure. An exemplary image acquisition system (2-001) includes: a light source and a single-mode optical fiber (SMF) (2-002), which is configured to emit light (2-003), which is relayed by one or more optical elements in an optical relay (2-004), which is configured to shape the light emitted from the light source to form a shaped light beam (2-005); and one or more optical elements (2-006), such as a dichroic mirror, which is configured to direct the shaped light beam to a water immersion objective (2-008), whereby a sample plane (2-010) is illuminated by an inclined light beam (2-009), so that the sample (2-012) emits light, such as fluorescence emission focused by the objective, which is focused by the water immersion objective (2-008) and one or more optical elements (2-013), such as a tube lens, and passes through an emission filter wheel (2-014) to reach an image collection system (2-015), such as a detector device.

[0111] 2.1.1. Light source

[0112] refer to Figure 2 An exemplary image acquisition system is provided, the system comprising a light source (2-002) configured to emit light (2-003). In certain implementations of the image acquisition system disclosed herein, the light source (2-002) may be configured to emit light of a single wavelength. In certain implementations of the image acquisition system disclosed herein, the light source (2-002) may be configured to emit light of two, three, four, five or more separate wavelengths. In certain implementations, the wavelength of the light emitted by the light source is predetermined. For example, but not by way of limitation, the wavelength may be predetermined so that the emitted light induces fluorescence emission when irradiating a sample (e.g., a sample comprising a fluorescent protein). In some cases, the wavelength employed in connection with the methods described herein will be in the range of 400 nm to 650 nm. In some cases, the light source (2-002) will emit light with a wavelength between 400 nm and 408 nm, between 550 nm and 565 nm, or between 638 nm and 650 nm. In certain non-limiting implementations, the light source (2-002) is configured to include three lasers having nominal center wavelengths of 405 nm, 560 nm, and 640 nm, respectively, which can be varied within the absorption band of the fluorophore used. In some cases, the 405 nm wavelength is used to excite the Hoechst dye. In some cases, the 560 nm wavelength is used to excite the dye attached to the HaloTag (e.g., JF 549 ).

[0113] In certain non-limiting implementations, the light source (2-002) is configured to catalyze a photochemical reaction. For example, and not by way of limitation, the wavelength and intensity of the illumination can cause chemical bond breakage. As an additional example, and not by way of limitation, the wavelength and intensity of the illumination can induce the adoption of a non-radiative dark state (i.e., "photobleaching" the molecule). As an additional example, and not by way of limitation, the wavelength and intensity of the illumination can induce radiative or non-radiative energy transfer between fluorophores within the sample.

[0114] In certain implementations of the image acquisition systems described herein, the light source (2-002) can be configured to deliver a predetermined amount of power to the back focal plane of the objective lens (2-007). For example, but not by way of limitation, the light source (2-002) delivers a power greater than 10 mW for certain wavelengths (e.g., 405 nm) and / or delivers a power greater than 150 mW for other wavelengths (e.g., 640 nm). Additionally or alternatively, where the light source (2-002) includes three lasers emitting at wavelengths of 405 nm, 560 nm, and 640 nm, respectively, the light source (2-002) can be configured to deliver a predetermined amount of power to the back focal plane of the objective lens (2-007). For example, but not by way of limitation, 405 nm can be configured to deliver <10 mW; 560 nm can be configured to deliver >150 mW; and 640 nm can be configured to deliver >50 mW.

[0115] In certain implementations of the image acquisition system described herein, the light source (2-002) is configured to emit pulsed light. For example, but not by way of limitation, the light source (2-002) can be configured to emit stroboscopic pulsed light. In certain non-limiting implementations, the light source (2-002) will emit 2 millisecond stroboscopic pulsed light. Additionally or alternatively, the light can be pulsed in synchronization with the start of frame acquisition, as described in detail below.

[0116] In certain implementations of the image acquisition system disclosed herein, single-mode optical fibers may be used to transmit light (2-003) from a light source (2-002) and to direct the light to an optical relay (2-004). Additionally or alternatively, multimode optical fibers having a predetermined core shape may be used to illuminate the sample.

[0117] In certain implementations of the image acquisition systems described herein, such as systems configured for high-throughput sample analysis, the light source (2-002) can be configured to exhibit low power output drift. In certain implementations, this low drift configuration improves the consistency of sample processing to facilitate high-throughput analysis. For example, but not by way of limitation, this low drift power output configuration maintains the power output within a variation of about 0% to about 15%, a variation of about 0% to about 10%, a variation of about 10%, a variation of about 9%, a variation of about 8%, a variation of about 7%, a variation of about 6%, a variation of about 5%, a variation of about 4%, a variation of about 3%, a variation of about 2%, or a variation of about 1%.

[0118] In some cases, such low-drift power output configurations maintain power output within a range of about 0% to about 15%, about 0% to about 10%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, or about 1% over a range of ambient (room) temperature (e.g., 17°C + / - 5°C). In some cases, this is achieved by using temperature sensors and / or closed-loop heaters to maintain a stable temperature within the internal light source (e.g., laser engine), thereby reducing output power drift. For example, but not by way of limitation, an insulated housing design can be used to isolate the light source from ambient temperature fluctuations. Additionally or alternatively, closed-loop heaters can be strategically placed at specific locations within the system, such as at the fiber coupler, to reduce output drift. Additionally or alternatively, a water jacket and / or chiller can be used to reduce heat buildup in the laser head. Furthermore, these thermal controls, used alone or in combination, can reduce the warm-up time to reach a stable operating state and maintain a more stable internal operating temperature when the laser is turned off and on.

[0119] 2.1.2. Optical components and sample illumination

[0120] refer to Figure 2 An exemplary image acquisition system, the system comprising a light source (2-002) configured to emit light (2-003), the light being relayed by one or more optical elements in an optical relay (2-004), the optical relay being configured to shape the light emitted from the light source to form a shaped light beam (2-005).

[0121] Although Figure 2An exemplary HILO implementation is depicted for the htSMT workflow described herein, but the htSMT workflow described herein can be combined with a variety of illumination strategies. For example, but not by way of limitation, the htSMT workflow described herein can be implemented using HILO, total internal reflection fluorescence (TIRF), HIST, or SOLEIL microscopy illumination strategies. Based on the htSMT workflow described herein, one skilled in the art will understand advantageous ways to adjust the TIRF, HIST, or SOLEIL illumination strategies for use in the present method. For example, the optical elements of any particular optical relay (2-004) can be selected and configured to produce a light beam (2-005) of appropriate shape and to provide appropriate translation of the light beam, for example, when a HIST illumination strategy is employed. Additionally or alternatively, one skilled in the art will understand how to configure the necessary optical elements to implement a TIRF illumination strategy based on the htSMT workflow described herein. For example, optical elements can be used to tilt the light beam so much that its critical angle is reached, thereby propagating an evanescent wave through the coverslip to illuminate a sample adjacent to the coverslip.

[0122] In certain non-limiting implementations of the optical relay (2-004) of the presently disclosed image acquisition system, the optical relay (2-004) will include one or more lenses. For example, but not by way of limitation, the selection and orientation of the lenses in the optical relay (2-004) will be configured to appropriately shape the light beam directed toward the sample. In certain non-limiting implementations, the optical relay (2-004) will include a lens having a predetermined focal length (e.g., 80 mm) to collimate the emitted light (2-003) from the light source (2-002). Additionally or alternatively, the optical relay (2-004) will include a lens or a series of lenses (e.g., a telescope system) to shape the light beam. The specific focal length of the lens or series of lenses will be predetermined to produce an appropriately shaped light beam.

[0123] In an implementation of the htSMT workflow described herein, wherein the image acquisition system is configured to include an illumination system based on a HIST microscope, the optical relay can be configured to include a telescope comprising two cylindrical lenses (e.g., f=400 / 250 mm and f=50 mm) to generate a tiled beam compressed 8x or 5x, which, in some implementations, is relayed by another telescope system (e.g., f=60 mm and f=150 mm) before passing through an additional lens (e.g., f=400 mm). In such an implementation of the illumination system based on a HIST microscope, the optical relay (2-004) can include one or more optical elements or components configured to translate the beam relative to the imaging plane of the sample to be analyzed, for example, in a direction orthogonal to the longer dimension of the beam. For example, but not by way of limitation, such optical elements or components configured to translate the beam relative to the imaging plane of the sample to be analyzed can include a galvanometer. Additionally or alternatively, such optical elements or components configured to translate the beam relative to the imaging plane of the sample to be analyzed can include a computer-controlled motor.

[0124] refer to Figure 2 An exemplary image acquisition system includes an optical relay (2-004) configured to shape light emitted from a light source to form a shaped beam (2-005), which is then guided by an optical element (2-006), such as a dichroic mirror, and the optical element is configured to guide the shaped beam to a water immersion objective (2-008), whereby a sample plane (2-010) is illuminated by an oblique beam (2-009). The inset is an illumination view (2-011) relative to the X and Y axes of the sample plane, including a peak intensity core (2-011A) that gradually decreases to an outer edge (2-011B) of lower intensity in a Gaussian distribution.

[0125] In certain non-limiting implementations of the image acquisition system of the present disclosure, a water immersion objective (2-008) directs an inclined light beam (2-009) onto a sample plane (2-010) to be analyzed. In certain non-limiting implementations of the image acquisition system of the present disclosure, the objective (2-008) is a water immersion objective. The use of a water immersion objective enables high-throughput sample analysis by eliminating the oil associated with the use of an oil immersion objective, thereby allowing consistent sample processing and imaging. For example, but not by way of limitation, the objective can be a 60X 1.27NA water immersion objective (Nikon). In certain implementations of the workflow described herein, the water immersion objective (2-008) will be heated by a heating element. For example, such a heating element will maintain the water immersion objective (2-008) at a temperature sufficient to avoid causing a change in the temperature of the sample contained in the sample plate (2-021).

[0126] Image acquisition

[0127] In certain non-limiting implementations of the image acquisition system of the present disclosure, the water immersion objective (2-008) is also used to focus fluorescence emitted by the sample in response to illumination provided by the tilted light beam (2-009). However, in some cases, a second objective is used to focus fluorescence emitted by the sample in response to illumination provided by the tilted light beam (2-009). In certain non-limiting implementations, the objective focuses the fluorescence emission (2-012) through an emission filter wheel (2-014), for example, a bandpass emission filter matched to the spectrum of the observed fluorophore and mounted in a high-speed filter wheel (Finger Lakes Instruments), and is collected by a detector device (2-015). In certain non-limiting implementations, the objective focuses the fluorescence emission and is directed to an optical relay before being collected by the detector device (2-015). For example, but not by way of limitation, such an optical relay may include one or more lenses and one or more additional optical elements, for example, elements configured to reject additional scattered light before being collected by the detector device (2-015). In certain non-limiting implementations, the fluorescence emission focused by the objective is directed through another dichroic mirror to split the emission across multiple areas of a detector device (2-015), where the detector device can be a CMOS camera, such as a back-illuminated CMOS camera (Prime 95b, Teledyne).

[0128] In certain non-limiting implementations, when the image acquisition system is configured in conjunction with an illumination system based on a HIST or SOLEIL microscope, the detector arrangement can be configured to synchronize detection with translation of the tilted light beam (2-009) over the sample. Figure 4 The following figure shows a schematic diagram of the method, which is relevant to the implementation of HIST and SOLEIL, where "active pixels" correspond to the aspect of the detector device actively collecting data in synchronization with the translation of the tilted beam (2-009). For example, but not by way of limitation, the detection device can be a CMOS camera, such as a back-illuminated CMOS camera (Hamamatsu Fusion BT).

[0129] In certain implementations of the image acquisition system of the present disclosure, the CMOS camera can be operated so that a series of SMT frames are collected for each field of view. For example, but not limited to, 1-100,000 SMT frames, 1-50,000 SMT frames, 1-20,000 SMT frames, 1-10,000 SMT frames, 1-1,000 SMT frames, 1-500 SMT frames, 5-250 SMT frames, 10-200 SMT frames, 100-200 SMT frames, or 200 SMT frames are collected for each field of view. In certain implementations, the CMOS camera can be configured to operate at a frame rate of 0.5 to 1000 Hz. In certain implementations, the CMOS camera can be configured to operate at a frame rate of approximately 100 Hz.

[0130] In certain non-limiting implementations of the image acquisition system of the present disclosure, the detector device is configured to transmit a signal with each frame to trigger other components of the imaging system. For example, but not by way of limitation, the detector device can trigger illumination from the light source (2-002) to collect fluorescence emissions associated with stroboscopic laser pulses. For example, but not by way of limitation, such fluorescence emission collection is associated with a 10 to 100 millisecond frame and a 2 millisecond stroboscopic laser pulse. In certain embodiments, the fluorescence emission collection is associated with a stroboscopic laser pulse of about 1 to about 4 milliseconds, such as about 1 to about 3 milliseconds or about 2 to about 3 milliseconds, wherein the duration of the stroboscopic laser pulse can be selected based on the frame rate employed (e.g., 10 to 100 millisecond frames).

[0131] In certain embodiments, the imaging acquisition system can be configured to detect a predetermined field of view. In certain embodiments, the detection field of view can have a size of about 50 μm to less than 100 μm in a first dimension and a size of about 50 μm to less than 100 μm in a second dimension. For example, but not by way of limitation, the detection field of view can have a size of about 94 μm in the first dimension and a size of about 94 μm in the second dimension.

[0132] In some implementations, the detector assembly can be used to collect fluorescence emissions at multiple wavelengths. For example, but not by way of limitation, fluorescence emissions from additional fluorophores within the same field of view can be collected at the same frame rate or at different frame rates to provide downstream registration of SMT tracks with other cellular components (e.g., the nucleus). Additional channels of the detector assembly can be used as needed to expand the number of fluorescence emissions captured simultaneously within the same field of view to provide downstream registration of SMT tracks with other cellular components (e.g., the nucleus).

[0133] Sample processing

[0134] refer to Figure 1, various aspects of the current subject matter can be implemented using htSMT workflows, where such workflows incorporate systems for sample preparation, including reagent handling. For example, and not by way of limitation, Figure 5 A schematic diagram of a sample plate (2-021) is provided, comprising a plurality of wells, such as well (2-016), in which samples may be prepared and analyzed. Figure 5 A schematic diagram of sample components, such as cells (2-018) and fluorescent proteins within the cells (2-017) is also provided. However, as described herein, Figure 5 It is not intended to convey scale, for example, each sample present in wells (2-016) may contain thousands of cells, and each cell may contain many fluorescent proteins. Figure 5 Also schematically illustrated is the ability of the sample handling system of the present disclosure to add additional reagents (2-019) to the sample in the sample plate (2-021). Such reagent addition can be handled by robotic manipulation, such as, but not limited to, translation of the robotic fluid handling system relative to the individual wells (2-016) of the sample plate (2-021), translation of the sample plate (2-021) itself, or a combination of both.

[0135] In certain implementations of the image acquisition system, the sample plate (2-021) can be maintained in a temperature-controlled environment by the environmentally controlled area (2-020). For example, but not limitation, the sample can be maintained at 22-50°C. In certain implementations of the image acquisition system, the sample plate (2-021) can be maintained in a humidity-controlled environment by the environmentally controlled area (2-020). For example, but not limitation, the sample can be maintained at a humidity of 20%-95%. In certain implementations of the image acquisition system, the sample plate (2-021) can be maintained in a defined gas environment by the environmentally controlled area (2-020). For example, but not limitation, the sample can be maintained at 5% CO2.

[0136] Cell lines and cell culture

[0137] See Figure 5 A particular advantage of the htSMT system described herein is that it can be assayed in living cells (2-018), allowing for tracking the activity, mobility, and diffusion behavior of proteins within the crowded environment of living cells. Figure 5As shown, the htSMT system of the present disclosure can be used to track fluorescently labeled proteins in a sample comprising a plurality of cells. If the sample (e.g., containing such cells) can be focused by a water immersion objective (2-008) for a sufficiently long time to direct the fluorescent emission of the fluorophore to a detector device (2-015), then consider example cells (e.g., cell lines) used in conjunction with the htSMT system described herein. For example, but not by way of limitation, cells can be directly adhered to a coverslip. As an additional example, but not by way of limitation, cells can be induced to adhere to a coverslip after the coverslip is treated with an extracellular matrix material (e.g., fibronectin, collagen, poly-D-lysine, laminin, matrigel, vitronectin, etc.). As an additional example, but not by way of limitation, cells can be induced to adhere to a coverslip after the coverslip is treated with a plasma.

[0138] Exemplary cells (e.g., cell lines) may be selected so as to minimize non-fluorophore emission reaching the detector. In certain embodiments, the cells used in the present disclosure may be mammalian, bacterial, or fungal cells. In certain embodiments, the cells are mammalian cells. In certain embodiments, the cells may be obtained from preserved tissue (e.g., fixed tissue), frozen tissue (e.g., frozen tissue sample), or fresh tissue (e.g., fresh tissue sample). In certain embodiments, cells and / or samples containing cells may be obtained from a subject. In certain embodiments, cells may be obtained from a malignant tumor of a tissue or tumor, for example, cells may be present in a tumor sample (e.g., a section of a tumor). In certain embodiments, cells may be obtained from a cell line. For example, but not as a limitation, specific cell lines that can be used in conjunction with the htSMT system described herein include: U2OS cells (ATCC catalog number HTB-96), MCF7 cells (ATCC catalog number HTB-22), T47d cells (ATCC catalog number HTB-133), and SK-BR-3 cells (ATCC catalog number HTB-30). In certain embodiments, cells may be present in a three-dimensional structure, such as an organoid or a spheroid. In certain embodiments, the cells may be present in organoids.

[0139] In certain implementations of the htSMT system disclosed herein, cells to be used are cultured as needed to provide sufficient cell numbers to achieve the desired high-throughput analysis. For example, but not limitation, cells such as U2OS cells (ATCC Catalog No. HTB-96), MCF7 cells (ATCC Catalog No. HTB-22), T47d cells (ATCC Catalog No. HTB-133), and SK-BR-3 cells (ATCC Catalog No. HTB-30) can be grown in DMEM (Catalog No. 1056601, Gibco DMEM, high glucose, GlutaMAX supplement, Thermofisher) supplemented with 10% fetal bovine serum (Catalog No. 16000044, Thermofisher) and 1% penicillin-streptomycin (Catalog No. 15140122, Thermo Fisher) and maintained in a humidified 37°C incubator with 5% CO2, with subculture performed approximately every two to three days. Additional culture strategies suitable for the cell lines and uses outlined herein are known to those skilled in the relevant art.

[0140] In certain implementations of the htSMT system disclosed herein, the cells comprise one or more fluorescent proteins. The choice of a particular protein and the manner in which it fluoresces (e.g., whether it is labeled by binding to a dye or by including an encoded fluorescent tag) may vary depending on the specifics of a particular study. For example, but not by way of limitation, one method of labeling proteins that can be used in conjunction with the htSMT system described herein is a HaloTag fusion strategy. For example, but not by way of limitation, one method of labeling is a fluorescent protein. For example, but not by way of limitation, one method of labeling is a photoconvertible fluorescent protein. For example, but not by way of limitation, one method of labeling is a photoactivatable fluorescent protein. For example, but not by way of limitation, one method of labeling is a SNAPtag fusion. For example, but not by way of limitation, one method of labeling is a CLIPtag fusion. For example, but not by way of limitation, one method of labeling is a fluorophore ligase system. For example, but not by way of limitation, one method of labeling is via a FlAsH or ReAsH tetracysteine ​​motif. For example, but not by way of limitation, one method of labeling is a strain-promoted alkyne-azide cycloaddition reaction of a fluorophore. For example, but not by way of limitation, one method of labeling is by inducing cellular uptake of a separately produced fluorescent protein. In certain implementations of the htSMT system disclosed herein, the cell comprises one or more fluorescent glycoproteins. In certain embodiments, a method of labeling a protein uses a gene editing system, such as a CRISPR-based editing system. For example, without limitation, a nucleic acid encoding a fluorescent protein (e.g., a fluorescent tag, such as HaloTag) can be inserted into a gene or inserted upstream or downstream of a gene encoding a protein to be labeled to produce a protein fluorescently labeled with HaloTag (e.g., at its C-terminus or N-terminus).

[0141] Although the HaloTag fusion method can be implemented in a variety of ways by those skilled in the art, one exemplary method is to transfect a mammalian expression vector in a cell line of interest (e.g., U2OS cells) containing a fusion gene (i.e., a protein of interest fused in frame with the HaloTag sequence) under the control of a weak L30 promoter and containing a neomycin resistance marker. In certain implementations, such transfection can be accomplished using FuGENE 6 (Catalog No. E2691, Promega) when the cell confluence reaches 70%. In certain implementations, the transfected cells can then be selected with an appropriate selection agent, such as G418 (Catalog No. 10131027, Thermo Fisher), at an appropriate concentration of, for example, 500 μg / mL. In certain implementations, the cells can then be cloned and isolated. The cells can first be isolated by rinsing with 100 nM JF 549-HTL (Cat. No. GA1110, Promega) and 50 nM Hoechst 33342 were used to stain and identify cells with the expected JF 549 The clones of the signal distribution are used to determine the clones expressing the desired fusion gene. Another exemplary method is to transfect cells with a ribonucleoprotein (RNP) complex, which includes sgRNA and Cas9 protein targeting the genomic sequence of the N-terminal or C-terminal region encoding the target protein, and combined with one or more linear dsDNA donors. In certain embodiments, each donor consists of a 200-300bp homology arm specific for each target, a codon-optimized HaloTag sequence, and a TEV linker (ENLYFQG) between the target and the HaloTag. In certain implementations, the SMT conditions can then be used to test the reaction of three to six clones to the control compound, and the most homogeneous clones can then be expanded for further testing.

[0142] While the htSMT workflow of this application is generally described with respect to the implementation of tracking compounds' effects on target fluorescent proteins, the htSMT workflow described herein is equally applicable to tracking and analyzing fluorescent target compounds. For example, but not by way of limitation, the compounds described herein can themselves be fluorescent or modified to facilitate fluorescence detection. Furthermore, changes in the motion of fluorescent compounds can be exploited to determine the SMT profile of the compound itself. Thus, all analytical strategies described herein for tracking target fluorescent proteins also apply to results obtained by tracking the compound itself.

[0143] 2.2.2. Single-molecule tracking sample preparation

[0144] refer to Figure 5 , various aspects of the current subject matter can be implemented using the htSMT workflow, wherein cells (2-018) are seeded on a sample plate (2-021), such as a 384-well tissue culture treated glass bottom plate, although other types of culture plates can also be used with the methods outlined herein, including but not limited to single chamber, 9-well glass bottom plates, 24-well glass bottom plates, 96-well glass bottom plates, 1536-well glass bottom plates, and 3456-well glass bottom plates, as well as plates made of alternative materials, such as plates made partially or entirely of plastic. In certain implementations, cells (2-018) are seeded at 1 to 20,000 cells per well (2-016), such as 50 to 10,000, 100 to 9,000, 250 to 8500, 500 to 7500, 750 to 7000, 2500 to 6500, or 6000 cells per well. The seeded cells can then be incubated under conditions suitable for adhesion, for example, overnight at 37°C and 5% CO2. To enable fluorescence emission, the cells can be incubated with a sufficient amount of label, for example, in the case of HaloTag fusions, 0.1-100 pM JF549 -HTL (Cat. No. GA1110, Promega) and 50 nM Hoechst 33342 (for labeling cell nuclei) in complete medium for one hour provides ideal results.

[0145] In certain implementations of the htSMT strategy described herein, the cells are then washed, for example, three times in DPBS and twice in imaging medium. In certain implementations, imaging medium is prepared to promote fluorescence emission, such as fluoroBrite DMEM medium (Cat. No. A1896701, Thermo Fisher), and may be supplemented with GlutaMAX (Cat. No. 35050079, Thermo Fisher) and the same serum and antibiotics as the growth medium.

[0146] Where appropriate, the compound can be added to the sample to test its effect on a specific fluorescent protein by SMT. In some implementations, the compound can be serially diluted in an Echo Qualified 384-well low dead volume source microplate (0018544, Beckman Coulter) to generate a dose titration source material. The compound can then be administered in a cell culture medium at a final dilution of, for example, 1:1000. In some implementations of the htSMT strategy described herein, each dose of compound will have at least two replicates per plate and three replicates per plate. In addition, in some implementations of the htSMT strategy described herein, 20 DMSO control wells and two dye-free control wells can be randomly distributed on each plate (2-012). In some implementations, the compound can be incubated for 0 to 48 hours before image acquisition, for example, at 37°C for 1 hour.

[0147] 3.htSMT software

[0148] 3.1.htSMT Software Overview

[0149] Figure 6An example system 600 of a high-throughput single-molecule imaging platform for measuring the movement of molecules within living cells is illustrated. An experiment 602 can be performed to collect a large amount of data from a plurality of living cells (e.g., using an imaging system 624 to identify a compound 626 and / or a target 622). The experiment 602 can include applying various identifiers to molecules of interest, such as tags that can then fluoresce or be detected in other ways (e.g., using a laser or other light source). A biological sample forming part of such an experiment 602 can be organized into a plate 604 having a plurality of wells 606. Each well 606 can have one or more associated fields of view (FOV) 610. The FOV 610 can be located within or correspond to a single well 606. A series of images can be generated for the FOV 610 to produce one or more movies 612, which can include SMT movies as well as non-SMT movies. The SMT movie can be used to track the path of a single labeled molecule (e.g., a protein), thereby generating multiple tracks. Each track can be composed of a plurality of points 614, which include the spatiotemporal coordinates of the labeled molecule at a particular time (e.g., Figure 7 610 ). Separately from tracking, and in some cases in parallel with tracking, the movie 612 can be used to identify molecules by using machine learning and / or computer vision-based image segmentation to generate masks 618. Masks 618 are spatial regions within the FOV 610 generated by segmentation. Each mask 618 can belong to a mask category, which is Figure 8 A more detailed description is given in .

[0150] Data associated with the two channels (e.g., the tracking channel and the segmentation / masking channel) can be combined to generate a plurality of metrics 620 associated with various aspects of the sample. In other words, the trajectory 616 (e.g., trajectory data) can be combined with the machine learning processed image segmentation data and further analyzed using statistical / machine learning methods. The processing of the combined data can be used to generate metrics 620, such as hit scores associated with compounds and / or targets in the biological sample, which can be stored in a database structure, such as Figure 9 Further described.

[0151] Figure 7 The data flow through an example system 700 for a high-throughput single-molecule imaging platform for measuring protein movement in living cells is illustrated. Experimental specifications 704 defining experiments 602 may be provided as data input via one or more clients 702. For example, each experiment 602 may be collected with accompanying stains (e.g., Hoechst or Potomac Red) for downstream analysis including segmentation 618. The experimental specifications 704 may define various parameters of the experiment 602, such as stains, dyes, compounds, treatments, etc. As previously described in Figure 6As described in , imaging system 706 (e.g., imaging system 624) can capture a series of images that generate one or more SMT movies 711 and / or non-SMT movies or segmented movies 708 (e.g., movie 612) that characterize molecular motion. SMT movie 711 can characterize the motion of a single fluorescent dye molecule and / or an image containing a single fluorescent dye molecule. Segmented movie 708 can include a series of images that characterize the motion of labeled molecules and / or their components. It should be understood that Hoechst staining is only one technique that can be used to label molecules, and different and / or multiple labeling techniques, such as Potomoc Red, can be utilized depending on the desired configuration. For example, MitoTracker Deep Red can be used to label mitochondria, concanavalin A-dye conjugate can be used to label the endoplasmic reticulum, SYTO 14 can be used to label nucleoli, phalloidin can be used to label actin, etc.

[0152] The SMT movie 711 can be analyzed to perform operations related to molecular tracking 710, which can include detection 712, sub-pixel localization 713, and linking 714 to identify molecular tracks 715 across the various images within the SMT movie 711. More specifically, during detection 712, one or more points within the SMT movie 711 can be detected or recovered. Each point can be assigned spatiotemporal coordinates. These spatiotemporal coordinates can be estimated using sub-pixel localization techniques 713. Linking 714 can be performed on these points to ultimately identify tracks 715.

[0153] As used herein, a link is a potential association between two points. Each link is directed, starting at one point and ending at another. A "correct link" connects two points generated by the same transmitter in different frames; otherwise, the link is "incorrect." One goal of the linking algorithm is to estimate which links are correct. Links referred to herein are of the format a:i→j. This means: link a, starts at point i and ends at point j. A link satisfies at least the following three constraints: (a) the link progresses in time, (b) the link must not connect two points that are more than a certain limit apart (referred to herein as the "search radius"), and (c) the link must not connect two points that are more than a certain limit apart in time (referred to herein as the "gap limit"). A point-link graph is a graph of the points and links of an SMT movie 711. Points are the vertices of the graph, and links are the edges of the graph. Because links progress in time, the point-link graph is a directed acyclic graph. A match is a subset of the links in the point-link graph such that no two links in the subset start or end at the same point. Trajectory 715 is used herein to refer to a continuous (end-to-end) sequence of links in the same match. A plurality of trajectories may be used to determine a dynamic index 730. Such parameters may include properties of the point that characterize the motion of the point. Such parameters may include one or more of the velocity, diffusion coefficient, or anomaly parameter of each point. The dynamic parameter of point i is referred to herein as θi The set of dynamic parameters for all points in the point-link graph is referred to here as Θ.

[0154] Separately from the processing of the SMT movie 711, and in some variations in parallel therewith, the segmentation movie 708 can be segmented, thereby generating one or more masks 720. The masks can be of various types, including but not limited to nucleus, cytoplasm, and / or irrelevant masks, which will be described in detail in the following sections. Figure 8 . The FOV 610 may contain any number of instance masks of a mask class. A semantic mask is the union of all instance masks of a mask class corresponding to a FOV (e.g., all cells, all nuclei, or all mitochondria of a FOV, etc.). Extraneous masks may contain portions of the non-SMT movie 708 that are excluded from any downstream data analysis. For example, these extraneous masks may correspond to portions of the non-SMT movie 708 that are out of focus or contain autofluorescent cell debris that prevents accurate tracking. During the segmentation process, molecules in the segmented movie 708 may be assigned to one or more masks. Image metrics 740 may be evaluated based on the masked molecules (e.g., cell health, focus quality, etc.).

[0155] Experiment information, such as dynamic metrics 730, image metrics 740, and any data derived from any of the metrics (e.g., segmentation information), can be provided to a data repository 770 for storage. Such a data repository 770 can store, for example, any results of an experiment 602, such as dynamic metrics 730, image metrics 740, and / or any data derived from any of the metrics. The data repository can include local persistence and / or dedicated servers accessed locally or via the cloud. The data repository 770 can also store metadata associated therewith and / or metadata associated with the experiment specification 704. Experiment information (e.g., results and metadata from historical experiments, etc.) can be provided to the data repository 770 via a repository application program interface (API) 750. The repository API 750 can also interface with a web-based graphical user interface front end 760 that provides such information for display on the client 702.

[0156] In some variations, the segmentation information can be used to identify subcellular compartments, such as the nucleus, nucleolus, cytoplasm, etc. The segmentation information can also be used to distinguish one cell from another. The segmentation information can be stored in a specific format (e.g., a multi-image file format such as TIFF).

[0157] The example dynamic index 730 may also include a state array. The state array is a framework for learning interpretable dynamic models from SMT trajectories and can be used to further understand the movement of the target protein and where in the cell the movement occurs. In some variations, the state array can be generated / populated using segmentation information. The output of the state array can be returned at the subcellular compartment level, allowing scientists to distinguish the dynamics of different subcellular compartments. In addition, the state array can be calculated for each individual subcellular compartment (e.g., each cell nucleus).

[0158] To facilitate application access to data (including but not limited to the state array), processed SMT data can be stored in formats that allow for the following: (a) representation of processed tracks and associated properties, such as the SNR and point shape characteristics for each SMT movie; (b) representation of mask objects, including mask categories (e.g., the associated subcellular organelle for each mask object); (c) association of tracks with mask objects (e.g., the nucleus in which each track was observed); and (d) association of all SMT movies with metadata about the original experiment, such as compound treatment, acquisition time, and imaging system name. Formats (a) and (c) can be protocol buffer schemas that define the storage format for tracks and associated mask objects. Format (b) can be a specialized image file format that includes the mask object to which each pixel in the FOV belongs. Format (d) can be a PostgreSQL database that records all captured experiments / movies. As a client of processed SMT data, the state array can leverage these data schemas to report the dynamic characteristics of tracks for each mask category or each mask object.

[0159] Figure 8 800 are multiple images illustrating the difference between mask categories and instance or semantic masks. As previously described, non-SMT movies or segmentation movies can be assigned to multiple categories. These categories may include cell nuclei (e.g., category A), cytoplasm (e.g., category B) and / or irrelevant masks (e.g., category C). Unique, individual masks can be applied to biological samples. For example, image 810 is a unique, individual instance mask applied to a cell nucleus (e.g., category A). Image 812 is a unique, individual instance mask applied to the cytoplasm (e.g., category B). Image 820 illustrates multiple instance masks applied to one or more cell nuclei, where a single color represents a different, unique individual instance mask. Image 822 illustrates multiple masks applied to one or more cytoplasms, where a single color represents a different, unique individual instance mask. Image 830 illustrates a semantic mask applied to one or more cell nuclei, which is the union of all instance masks. Image 832 illustrates a semantic mask applied to one or more cytoplasms.

[0160] Figure 9An example computer-implemented environment 900 is illustrated in which an imaging system 910 can interact with a computing architecture to execute the various algorithms described herein. Figure 9 As shown, the imaging system 910 can interface with one or more clients 950 (e.g., client 702 via a web application with a graphical user interface). The one or more clients 950 can interface with one or more servers 920 accessible via a network 930. The one or more clients 950 can host a frame grabber that captures images (e.g., movie 612) from a camera. These images can be temporarily stored on the one or more clients 950 and periodically transmitted to the one or more servers 920 via the network 930 for remote storage. The one or more servers 920 can also contain or have access to one or more data stores 940 for storing data collected and / or extracted from the sample by the imaging system 910. In some variations, the network 930 can include or be interfaced with one or more network storage arrays 960 for storing data such as captured images (e.g., movie 612).

[0161] Figure 10 1000 is a diagram illustrating an example computing device architecture for implementing various aspects described herein. In some variations, the sample computing device architecture may be the architecture of client 950 and / or server 920, and some components described with respect to diagram 1000 may be optional for client 950 and / or server 920. Bus 1004 may serve as an information highway interconnecting the other illustrated components of the hardware. Processing system 1008, labeled CPU (central processing unit) (e.g., one or more computer processors / data processors on a given computer or multiple computers), may perform the computations and logical operations required to execute a program. Optionally or in addition, processing system 1012, labeled GPU (graphics processing unit) (e.g., one or more computer processors / data processors on a given computer or multiple computers), may perform the computations and logical operations required to execute a program. Non-transitory processor-readable storage media (e.g., read-only memory (ROM) 1016 and random access memory (RAM) 1020) may communicate with processing system 1008 and / or processing system 1012 and may include one or more programming instructions for the operations specified herein. Optionally, the program instructions may be stored on a non-transitory computer-readable storage medium, such as a magnetic disk, an optical disk, a recordable memory device, a flash memory, a solid-state drive, or other physical storage medium.

[0162] In one example, the disk controller 1048 can interface one or more optional removable storage 1056 or local storage 1052 with the system bus 1004. The removable storage 1056 can be an external or internal disk drive, a solid-state drive, or an external hard drive. The local storage 1052 can be an internal hard drive and / or memory. As previously described, the various examples of removable storage 1056, local storage 1052, and disk controller 1048 are all optional devices. The system bus 1004 can also include at least one communication interface 1024 to allow communication with external devices (such as cloud storage and remote services) that are physically connected to the computing system or obtained externally via a wired or wireless network. In some cases, at least one communication interface 1024 includes or additionally comprises a network interface.

[0163] In some variations, such as for client 950, to provide for user interaction, the subject matter described herein may be implemented on a computing device having a display device 1044 (e.g., an LCD (liquid crystal display) or LED (light emitting diode) monitor) for displaying information obtained from bus 1004 to the user via display interface 1040, and an input device 1032 (e.g., a keyboard and / or pointing device (e.g., a mouse or trackball) and / or a touch screen) through which the user can provide input to the computer. Other types of input devices 1032 may also be used to provide for user interaction; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback via microphone 1036, or tactile feedback); and input from the user may be received in any form, including sound, voice, or tactile input. Input device 1032 and microphone 1036 may be coupled to bus 1004 via input device interface 1028 and communicate information via the bus. For example, input device 1032 may be imaging system 910 configured with the capability to capture the series of images described herein. The frame grabber 1058 may capture or grab a single frame from the analog or digital data that encapsulates a series of images obtained from the bus 1004. The frame grabber 1058 may include a memory capable of storing a single frame or multiple frames. The frame grabber 1058 may also provide the single frame or multiple frames to the bus 1004 for further storage, for example, on the local memory 1052 and / or the removable memory 1056. Other computing devices (e.g., dedicated servers) may omit the incorporation of Figure 10 Describes one or more components.

[0164] One or more aspects or features of the subject matter described herein may be implemented in digital electronic circuits, integrated circuits, specially designed application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system comprising at least one programmable processor, which may be dedicated or general purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to the storage system, at least one input device, and at least one output device. A programmable system or computing system may include a client and a server. The client and server are typically remote from each other and typically interact via a communication network. The relationship of client and server arises through computer programs running on respective computers that cause each other to have a client-server relationship.

[0165] These computer programs, which may also be referred to as programs, software, software applications, applications, components or codes, comprise machine instructions for a programmable processor and may be implemented in high-level procedural languages, object-oriented programming languages, functional programming languages, logic programming languages ​​and / or assembly / machine languages. As used herein, the term "machine-readable medium" refers to any computer program product, device and / or apparatus for providing machine instructions and / or data to a programmable processor, such as a disk, an optical disk, a memory and a programmable logic device (PLD), including a machine-readable medium that receives machine instructions in the form of a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor. A machine-readable medium may store such machine instructions non-temporarily, such as a non-transitory solid-state memory or a magnetic hard disk or any equivalent storage medium. A machine-readable medium may alternatively or additionally store such machine instructions in a transient manner, such as a processor cache or other random access memory associated with one or more physical processor cores.

[0166] Probabilistic methods for dense molecular tracking in living cells

[0167] One aspect of the htSMT workflow 1-100 includes recovering tracks 715 or paths of identifiers (eg, individual fluorescent emitters, etc.) from a recorded sequence of images (eg, SMT movie 711). This recovery is referred to herein as "tracking."

[0168] Tracking 710 in htSMT presents several challenges. First, each emitter is dim, contributing as few as a hundred photons per frame, necessitating sensitive detection methods. Second, the absolute intensity and noise characteristics of each movie depend on its originating imaging system; these differences can arise from variations in laser power or camera gain and offset. Therefore, it is desirable that htSMT tracking methods be insensitive to variations in the absolute intensity of the movie. Third, protein motion in cells can be rapid, causing molecules to rapidly enter and leave focus. Consequently, average track lengths can be as short as three or four frames, severely limiting the information available for predicting the molecule's future motion. Finally, tracking at high labeling densities becomes challenging due to ambiguities in associating detections with tracks. For example, tracking methods could adjust their parameters in a density-dependent manner to enable accurate tracking at a wide range of densities.

[0169] As previously described, tracking 710 of htSMT may include detection 712, sub-pixel localization 713, and linking 714. First, during detection 712, the portion of each movie frame of the SMT movie 711 containing emitters is identified. Second, sub-pixel localization 713 infers the position of the emitter to sub-pixel resolution, thereby producing the spatiotemporal coordinates of each detected emitter. For example, sub-pixel localization 713 may fit the observed light distribution around the emitter to an approximation of the imaging system's point spread function (PSF). Third, a linking algorithm associates the detected emitters with tracks. While all steps should be high-performance to meet the needs of htSMT data processing, linking 714 in particular can become very expensive in the case of high label density and / or number of detections per frame. This may be due to the combinatorial explosion of possible tracks that can be constructed from a given set of detections.

[0170] Figure 11 Diagram 1100 illustrating the extensible tracing pipeline 1110 in htSMT. Diagram 1100 is a modular structure that allows for custom composition at runtime. Figure 7 710 (e.g., detection 712, sub-pixel localization 713, and linking 714) as specified in an experiment-specific configuration file 1122. Each tracing pipeline 1110 may have an associated runtime configuration 1120 specified by the experiment-specific configuration file 1122, which may define specific types for the detector 1112, sub-pixel localizer 1113, and linker 114.

[0171] The tracking pipeline 1110 may receive as input an image sequence (e.g., an SMT movie) 1111. A detector 1112 may be applied to the image sequence 1111 to detect or recover one or more points within the image sequence 1111 using any of the following detector types: a generalized log-likelihood ratio point detector, a difference of Gaussian (DoG) detector, a Laplace of Gaussian (LoG) detector, a determinant of Hessian (DoH) blob detector, or any combination thereof. Other types of detectors may be used depending on the implementation.

[0172] The spatiotemporal coordinates associated with the detected points may use sub-pixel localization 1113. This sub-pixel localization 1113 may include any of the following types of localizers: radially symmetric localizers, maximum likelihood fitting of candidate point models using the Levenberg-Marquardt method, etc. Other types of localization techniques may be used depending on the implementation.

[0173] A linker 1114 may be used to link two points. In some variations, the linker 1114 may rely on heuristics (e.g., nearest neighbor methods, etc.) or an exact solution to the assignment problem (e.g., Hungarian algorithm, etc.). This type of linker may utilize separate steps for inferring trajectories and inferring dynamic parameters from the trajectories. In other variations, the linker 1114 may infer a joint probability distribution of possible trajectories and dynamic parameters in a scalable manner. This distribution can be used to make more informed point estimates of the "correct" trajectory, estimate confidence in any particular set of trajectories, or derive dynamic results completely independent of the trajectories. This approach, referred to herein as "probabilistic linking," can be used to estimate the trajectory distribution of thousands to tens of thousands of fast-moving objects that are in close proximity in a movie.

[0174] The trace pipeline 1110 may output an object trace 1115 that represents the possible trajectories in a graphical format. This output may depend on the type of linker used by the linker 1114.

[0175] 3.2.1. Probabilistic methods for dense molecular tracking in living cells

[0176] Two example probabilistic linking types of linker 1114 may utilize different approaches, including variational Bayesian inference (referred to herein as "vtrack") or Gibbs sampling (referred to herein as "gibbstrack").

[0177] The dynamic parameters of each point can be viewed as parameters of a motion model that defines the probability distribution of its future motion. It can be assumed that the probability of any given vector displacement between points i and j depends only on the dynamic parameters of points i and j, and not on the rest of the point link graph. i→j Equation 1 expresses this probability and defines a "motion model":

[0178] f r|θ (r i→j |θ i ,θ j ). (1)

[0179] One option for Equation 1 is to use a likelihood function for the scaled Brownian motion, which can be characterized by a single kinetic parameter (the diffusion coefficient) at each point. A simplified Bayesian model for the scaled Brownian motion is described in detail in Section 3.2.6 below.

[0180] The objective function of the linking algorithm can be defined as follows. Let M be the number of links in the point-link graph. Let E∈{0,1} M is a vector of 1s and 0s, representing a match such that if link a participates in the match, then E a =1, otherwise E a = 0. Let w∈R M is a real-valued weight vector. Let w0 be the likelihood of starting or ending the trajectory. Then the linking algorithm is to find the optimal match that satisfies Equation 2.

[0181]

[0182] Because each matching cannot contain two links that start or end at the same point, Equation 2 (eg, the criterion for the solution of the linking problem) is an instance of an unbalanced allocation problem.

[0183] Because each element of the matching vector E corresponds to a link a:i→j, each element here can be conveniently indexed by its link index a (i.e., E a ) or its point index (i.e. E i→j ) to index. Similarly, vector displacement can be expressed as r corresponding to the link a:i→j a or r i→j .

[0184] If the weight vector is constant, then Equation 2 can be solved using classical solutions to the assignment problem. These include exact solutions (such as the Hungarian algorithm) and heuristics (such as the nearest neighbor method).

[0185] However, in general, the weight vector can be a function of the dynamic parameters Θ, which can be estimated from the trajectory defined by the matching vector E. Since both E and Θ are unknown a priori, they can be estimated jointly.

[0186] The goal of the probabilistic linking algorithm is to evaluate the conditional distribution p(E,Θ|R). Here, R=(r1,…,r M) represents the vector displacement corresponding to each of the M links in the point link graph. Alternatively, these terms can be generalized to contain any additional information relevant to the link problem (e.g., spatial location, point shape characteristics, etc.). Once the distribution p(E, Θ|R) is obtained (e.g., by the vtrack or gibbstrack methods discussed in this paper), it can be used to estimate the maximum a posteriori trajectory by solving Equation 2 using the marginal link probability w = logp(E|R), or by taking the posterior average dynamic parameter to obtain the estimated values ​​of the dynamic parameters.

[0187] According to Bayes’ theorem, the conditional distribution p(E, Θ|R) can be written as Equation 3 (e.g., Bayes’ theorem for the linking problem):

[0188]

[0189] In Equation 3, the term p(R|E,Θ) is the likelihood of the observed displacement given the dynamic model Θ and the matching vector E. Since f r|θ (r i→j |θ i ,θ j ) is a given dynamic parameter θ i and θ j A single displacement r i→j The likelihood function of (Equation 1) is given by the product of the likelihood functions of each link (Equation 4). In Equation 4 (i.e., the likelihood function of the link problem), Pa(i) is the set of parents of point i, which contains the set of all points j such that j→i is an allowed link:

[0190]

[0191] In Equation 3, the term p(E, Θ) is a prior on E and Θ. Here we can assume that p(E, Θ) = p(E)p(Θ), where p(E) = constant for all allowed E, and where p(θ i ) is the prior for the dynamic parameters of point i and is chosen to be consistent with the likelihood f r|θ (r i→j |θ i ,θ j ) conjugation.

[0192] In Equation 3, the term p(R) is the so-called "evidence" and can be analytically difficult to handle. Therefore, the left side of Equation 3 cannot be evaluated in closed form. The methods described in this paper, based on Gibbs sampling (gibbstrack) and variational tracking (vtrack), avoid this problem. gibbstrack approximates the left side of Equation 3 by drawing a fixed number of random samples, while vtrack approximates the left side of Equation 3 using a mean-field approximation. Gibbstrack and vtrack are discussed in detail below.

[0193] 3.2.2. Probabilistic Linking via Variational Bayesian Optimization (vtrack)

[0194] Using vtrack, the posterior p(E, Θ|R) can be approximated using variational Bayesian optimization. In this approach, the posterior can be approximated by assuming it is a factor of E and Θ: p(E, Θ|R) ≈ q(E)q(Θ). vtrack achieves progressively better approximations to the posterior by first refining q(E) while holding q(Θ) constant, then refining q(Θ) while holding q(E) constant, and iterating between these two steps until convergence.

[0195] By manipulating the model introduced in Equation 4, a simplified version of vtrack can be derived. This derivation leaves room for the choice of motion model, as vtrack can be used with a variety of motion models. As described below, vtrack can be derived from the Brownian motion model or a more general model.

[0196] Consider the joint probability function of all variables in the model expressed in Equation 4. Based on the a priori choices discussed in Section 3.2.1, the joint probability distribution of all parameters can be decomposed as shown in Equation 5 (i.e., the joint probability density of the link problem):

[0197] p(R,E,Θ)=p(R|E,Θ(p(E)p(Θ)). (5)

[0198] Substituting Equation 4 and the prior p(R, Θ) into Equation 5 and taking the logarithm yields Equation 6 (i.e., the logarithmic joint probability density of the linking problem):

[0199]

[0200] In variational pursuit methods, we seek an analytical approximation to the true posterior q(E,Θ)≈p(E,Θ|R) that satisfies two criteria. First, q is a factor of E and Θ (Equation 7):

[0201] q(E, Θ)=q(E)q(θ). (7)

[0202] Second, q maximizes the evidence lower bound (Equations 8 and 9):

[0203] and (8)

[0204] q(E,0)=argmax q L[q]. (9)

[0205] Any distribution q(E, Θ) that satisfies Equation 7 and Equation 9 must also satisfy Equation 10 and Equation 11. In Equation 10 (i.e., the recursive equation for the factor q(E)) and Equation 11 (i.e., the recursive equation for the factor q(Θ)), logp(R, E, Θ) is given by Equation 6, and the expected value is and is taken according to the corresponding factor in q(E,Θ), and the constant explains the normalization of the corresponding factor:

[0206] and (10)

[0207]

[0208] Equations 10 and 11 can be solved sequentially to obtain a progressively better approximation to q(E, Θ). This scheme is called expectation maximization. This algorithm converges because the evidence lower bound (Equation 8) is convex for each factor in q. This method is combined with the tracking model defined in Equation 6, referred to here as vtrack.

[0209] Once q(E,Θ) is obtained, the maximum a posteriori matching vector can be evaluated using a sparse hill climbing algorithm by setting the weight vector to the marginal log probability of each link

[0210] Equations 10 and 11 can be solved for any motion model for which a conjugate prior exists (i.e., any choice of Equation 1), along with the log probability density expressed by Equation 6. This includes any motion model with a probability density belonging to the exponential family of distributions.

[0211] 3.2.3.Specific Form of Brownian Motion vtrack

[0212] As a limited demonstration of vtrack, we can next solve Equations 10 and 11 for the Brownian motion model represented by Equation 19 using the prior represented by Equation 20. Under these conditions, the log joint probability (Equation 6) becomes Equation 12 (i.e., the joint probability density of Brownian motion). In Equation 12, m is the spatial dimension, θ i is the diffusion coefficient at point i, r i→jis the spatial displacement of link i→j, Δt is the frame interval, and α0 and β0 are prior parameters.

[0213]

[0214] Substituting Equation 12 into Equation 10 and Equation 11 and solving for the factors q(E) and q(Θ), respectively, yields Equation 13 and Equation 14. In these equations, GraphSoftmax is the GraphSoftmax operator as described below, T is the temperature, w0 is the trajectory start likelihood, and is the log-likelihood vector for each link. Then, the vtrack algorithm calculates β and β for a given α via Equation 13. Then the given α and β are calculated via Equation 14 Repeat this process until convergence. Then, given the marginal link probability In the case of , the maximum a posteriori matching is estimated by the sparse hill climbing algorithm described below. Equation 13 is the approximate posterior of the dynamic parameters of the Brownian motion and can be expressed as follows:

[0215]

[0216] Equation 14 is the approximate posterior of the marginal link probability of the Brownian motion and can be expressed as follows:

[0217]

[0218] 3.2.4. Extended Variational Pursuit Algorithm

[0219] The vtrack algorithm described above can be extended to exploit additional information hidden in the point-link graph as follows. Given a point-link graph with N points and M links, let Represents the m-dimensional coordinates of each point. As before, use the vector E∈{0,1} M To indicate whether each link participates in the matching, each point i is compared with the dynamic parameter θ i Associated, and let Θ=(θ1,…,θ N ) is the set of dynamic parameters for all points. We can assume the stochastic model expressed in Equation 15. In Equation 15, Pa(i) is the "parent point" of point i, or the set of points starting from the link that ends at point i. The term f r|θ (r i→j |θ i ) represents the kinetic parameter θ i Lower Sports X j →X i The probability density of . i) is the prior of the dynamic parameters of point i, and it is selected to be r|θ (r i→j |θ i ) conjugate. Note that the term p(X i |θ u ,E) is defined recursively in terms of the parents of point i, which makes Equation 15 a generalization of Bayesian networks that allow for uncertainty in the links E.

[0220]

[0221] As with the simple vtrack algorithm, we seek an approximate posterior q(E,Θ)=q(E)q(Θ)≈p(E,Θ|X) to maximize the evidence lower bound (Equation 8). The algorithm proceeds by alternately solving Equations 10 and 11.

[0222] As an example of the solution, we can use Brownian motion to prove the result. To do this, we can assume that f r|θ (r i→j |θ i ) is a gamma distribution of the form specified in Equation 19, and p(θ i ) is the inverse gamma distribution of the form specified in Equation 19. The posterior probability is then given by Equations 16 and 17. In Equation 16, the term is the marginal link probability, and an additional mean-field approximation is made For the Brownian model, vtrack is performed as follows: (a) Evaluate the given Each α i and β i , and (b) evaluate given all α i and β i of The operator GraphSoftmax is a graph softmax operator as described below. Equation 16 is the approximate posterior of the point potential parameter and can be expressed as follows:

[0223]

[0224] Equation 17 is the approximate posterior of the link and can be expressed as follows:

[0225]

[0226] As in the case of the simple vtrack algorithm, once the approximate posterior q(E,Θ) is obtained, it can be obtained by applying the sparse hill climbing algorithm (described below) to the log-marginal link probability to estimate the maximum a posteriori trajectory.

[0227] 3.2.5. Gibbs sampling method for probabilistic links (gibbstrack)

[0228] Another way to evaluate p(E, Θ|R) is to draw random samples from this distribution and then take the mean of these samples to approximate the mean of the posterior distribution. A simple and flexible way to implement this sampling scheme is to alternately draw from the conditional distributions of E and Θ (Equation 18). Equation 18 defines Gibbs tracking and can be expressed as follows:

[0229] E~p(E|R,Θ) (18)

[0230] Θ~p(Θ|R,E).

[0231] The sampling scheme expressed in Equation 18 can be accomplished as follows. Start by estimating the dynamic parameters Θ and setting E to all zeros (this is always a valid matching vector for any point-link graph). At each iteration, the log-likelihood of each link is evaluated according to the current dynamic model (the motion model is selected by Equation 1). Let is the vector of these log-likelihoods for all links. Propose a "pivot" as follows, which corresponds to changing up to 4 elements of E and is associated with a change in the log-likelihood Δw. Draw a random number u ~ uniform (0,1), if u ≤ e Δw / T Then accept the pivot. Next, draw a sample θ i ~p(θ i |E), if Equation 1 is consistent with the prior θ i conjugate, then the sample can be found analytically.

[0232] If E1, E2, …, E n and Θ1,Θ2,…,Θ n is a sample generated in this way by Equation 18, then the marginal link probability Can be approximated as As in the case of vtrack, the maximum a posteriori estimate can be estimated by applying the sparse hill climbing algorithm (Section 3.2.7) to the marginal link probabilities

[0233] 3.2.6. Bayesian Model of Brownian Motion

[0234] One motion model for vtrack and gibbstrack is scaled Brownian motion, which uses a single kinetic parameter (diffusion coefficient) to characterize the motion of each point. Under scaled Brownian motion, the likelihood function Equation 1 becomes a gamma distribution represented by Equation 19, where m represents the dimension of the space in which the motion is observed, and θ iis the diffusion coefficient at point i, and r i→j is the vector displacement corresponding to link i→j. Equation 19 is the likelihood function of the m-dimensional Brownian motion and can be expressed as follows:

[0235]

[0236] A useful prior conjugate of the Brownian likelihood expressed in Equation 19 is the inverse gamma prior expressed in Equation 20. In Equation 20, α0 and β0 are prior hyperparameters, Δt is the frame interval, and θ is the diffusion coefficient of a single point. Equation 20 is the prior for Brownian motion and is expressed as follows:

[0237]

[0238] Given a series of observed displacements r1,r2,...,r n , the posterior distribution of the diffusion coefficient is given by Equation 21, which forms the basis for Bayesian inference of the diffusion coefficient. In Equation 21, f θ represents the inverse gamma distribution of the form given by Equation 20. Equation 21 is the posterior distribution of the Brownian motion and can be expressed as follows:

[0239]

[0240] 3.2.7. Sparse Hill Climbing Algorithm

[0241] If the weights are constant and the weight vector If and the trajectory starting weight w0 are constants, then Equation 3 can be approximately solved using a sparse hill climbing algorithm. This section introduces this algorithm, which is related to several issues discussed throughout Section 3.2. The algorithm described can be understood as a modification of one of the best performing methods in the public SMT competition.

[0242] In the following algorithm, the “pivot” can be defined as the matching vector E∈{0,1} M Each pivot is associated with a weight change Δw=W a -W b -W c +W d The weight change determines whether the pivot is accepted. Each pivot is constructed by selecting a link a:i→j and setting W a =w a To start. If there is a link b:i→k such that E b =1, then let W b =w b Otherwise, let W b =w0. If there exists a link c:l→j such that Ec =1, then let W c =w c Otherwise, let W c =w0. If the previous two conditions are true, and d:k→j is a link, then let W d =w d Otherwise, let W d =w0. If Δw = W a -W b -W c +W d >0, then accept the pivot. If the pivot is accepted, you can set E a = 1 to perform the pivot, and set E if b is a link b = 0, if c is a link, set E c = 0, and if d is a link then set E d =1.

[0243] By maintaining a record in memory of the currently assigned forward and backward links for each point, the required lookups in each pivot (eg, determining whether links b, c, and d exist) can be accomplished quickly.

[0244] For a point-link graph with N points and M links, the sparse hill climbing algorithm can be described as follows. Starting from the initial valid matching vector E∈{0,1} M and a known weight vector Start. Let E be zero and the M vector is always valid. In each iteration, select a link a:i→j. If E a =1 and 2w0-w a >0, then set E a = 0; otherwise, continue to the next iteration. On the other hand, if E a = 0, the weight change of the corresponding pivot is evaluated, and if the weight change is positive, the pivot is executed. The algorithm proceeds in this way until convergence.

[0245] One way to determine convergence is to check that the matching vector E has not changed after iterating through all M links in a random order.

[0246] Sparse hill climbing algorithms can be used for a variety of purposes, including estimating the nearest neighbor solution to a pursuit problem (by setting w for each link a:i→j) a =w0-r i→j , where r i→j is the Euclidean length of the link, and w0 is the track start probability), estimate the maximum a posteriori matching given the posterior distribution produced by gibbstrack or vtrack, or other maximization problems.

[0247] 3.2.8. Graphical softmax

[0248] One aspect of the vtrack algorithm is to normalize the incoming and outgoing links in a point-link graph given some link log-likelihood. This normalization accounts for dependencies between links caused by the topological constraints of the matching (e.g., two links in the same matching cannot start or end at the same point).

[0249] One approach to address normalization is to generalize the softmax operator (e.g., Boltzmann distribution) to doubly stochastic matrices.

[0250] Another way to solve the normalization problem is to use the "graph softmax" operator. The input of the graph softmax operator is a point-link graph containing N points and M links, and the log-likelihood vector of each link Temperature (T) and trajectory initial weight (w0). The output is a probability vector Make is the marginal probability of the link a:i→j. This operation can be expressed as

[0251] Instantiate five vectors: and will hold the link probabilities, v,w will hold the start and end probabilities of each point, and x,y are auxiliary buffers. To initialize, set all links a to And all points In each iteration, the following operations are performed: (i) set x = v and y = w, (ii) for each link a:i→j, set and (iii) For each link a:i→j, set (iv) For each point i = 1, 2, ..., N, set v i =v i / x i and w i =w i / y i Repeat these steps until convergence, then return the link probability

[0252] By subtracting the probabilities of all links entering or leaving each point, we can get The starting probability v and the ending probability w are obtained from .

[0253] The convergence rate of this algorithm depends more on the sparsity of the problem than on its size. For htSMT, convergence may take around 20 iterations.

[0254] 3.2.9. Confidence Metrics for Tracking Solutions

[0255] Ensuring high-quality data increases confidence in the results produced by the htSMT system. However, because htSMT may be generated continuously at a high rate across multiple imaging systems, the ability of human supervision to detect data issues may be limited. Therefore, a feature of the tracking pipeline described in this paper is that it provides a built-in confidence measure for the tracking solution, which can be used as an alternative to direct human supervision for diagnostic purposes.

[0256] Equation 22 is the normalized entropy of the posterior distribution of the trajectory, which can be defined as follows:

[0257]

[0258] Where N is the total number of points, is the marginal probability of link j→i under the inferred posterior distribution, and is the probability that point i starts its own trajectory.

[0259] Equation 22 defines the tracking algorithm's confidence in its solution, with values ​​closer to 0 indicating higher confidence. Values ​​below 0.4 reflect high confidence in the tracking solution. However, tracking very fast particles can be more error-prone (especially for Brownian motion), and a higher threshold may be tolerated for such use cases.

[0260] The Tracking Error Rate Lower Bound (ERLB) is a secondary diagnostic for evaluating the quality of hSMT results. This metric is designed to provide a lower bound on the fraction of incorrect links generated by the tracking algorithm. For example, a value of 0.1 indicates that at least 10% of the links generated by the tracking algorithm are likely incorrect. Since it may not be possible to know in advance which links are correct or incorrect, the ERLB is designed for the case where a subset of links is known to be incorrect.

[0261] The ERLB is calculated once for each SMT movie. The detection set from the second half of the movie is superimposed on the detection set from the first half. The tracking algorithm is then rerun on this superimposed detection set, regardless of which detections come from which half of the movie. Under these conditions, any link generated by the tracking algorithm between two detections from different halves of the SMT movie is incorrect. The score of such a link serves as a lower bound on the error rate, since it is unknown whether a link between two detections from the same half of the movie is incorrect.

[0262] The process of superimposing the two halves of the movie is done in the following way. Each point is always associated with a frame index, or an image index in the original image sequence from which the point was derived. Let S1 be the set of points in the first half of the SMT movie, and let S2 be the set of points in the second half of the movie. Let T be the total number of frames in the movie. For each point in S2, subtract floor(T / 2) from its original frame index. S1 and S2 are then concatenated to form a new set of detections This set of detections is then fed into the linking algorithm. ERLB is then calculated as the number of links generated by the linking algorithm that connect a point in S1 to a point in S2 (or vice versa), divided by the total number of links generated by the linking algorithm. The links "generated by the linking algorithm" are the E in the solution E to the problem expressed in Eq. a =1 link a.

[0263] 3.2.10. Trajectory-independent dynamics estimation using probabilistic tracking algorithms

[0264] The goal of htSMT is to infer the dynamic parameters of a target protein to identify experimental conditions that alter these dynamics. For example, the diffusion coefficient of a target protein can be estimated for each of several compound treatments to identify compounds that perturb the protein's diffusion state (e.g., by creating or disrupting protein-protein interactions).

[0265] These dynamic parameters are usually estimated from the trajectory. However, since probabilistic tracking provides the posterior distribution p(E, Θ|R) of the dynamic parameters and the trajectory, the dynamic parameters can be estimated by marginalizing over the trajectory. In the following equation, is the dynamic parameter θ of point i i The marginal posterior mean of . This estimate is θ for all possible trajectories including point i. i The weighted average, weighted by the posterior probability of each trajectory, is expressed as follows:

[0266] p(Θ|R)=∑ E p(E,Θ|R) (23)

[0267]

[0268] As an example, we describe a procedure to estimate the marginal posterior diffusion coefficient for each point over all possible pasts and all possible futures. Let Marginal link probabilities estimated for the probabilistic pursuit algorithm

[0269] There are n SMT movies with M possible links and N points. Let is the square displacement of link a. Let d be the dimension of the space, and let γ∈(0,1) be the damping constant. Set all points i to Then for each link a:i→j, set and Then, the posterior distribution of the diffusion coefficient of point i is where Δt is the frame interval, α0 and β0 are prior values, and InvGamma is the inverse gamma distribution with scale parameterization. Then the posterior mean diffusion coefficient of point i is in and

[0270] 3.2.11. Tracking Algorithm Benchmark

[0271] Figures 12A to 12C Depicts benchmarks of various tracking algorithms. Figure 12A , optical dynamics simulations are used to test the accuracy of multiple linking algorithms, including vtrack, gibbstrack and adaptive hill climbing algorithms, as well as other linking algorithms: random, conservative and nearest neighbor. Using the random linking algorithm as a control, each detection is randomly linked to another detection within its range gate (i.e., the set of detections within its search radius and gap constraints). The conservative linking algorithm is configured so that each detection is only linked to another detection if there are no other possibilities within the applicable range gate. The nearest neighbor linking algorithm stipulates that each detection is linked to the nearest neighbor in the applicable range gate. There are three types of experiments (benchmark 1, benchmark 2 and benchmark 3) with increasing difficulty. The outputs of these experiments are link recall, link precision and F1 score (i.e., the harmonic mean of recall and precision). Figure 12B The metrics are shown in this context. To ensure a fair comparison, the search radius (i.e., the maximum distance to consider a link) and the gap limit (i.e., the maximum number of gap frames to consider a link) are kept constant for all algorithms at 1.25 μm and 2 gaps, respectively. The exception is the conservative linking algorithm, which essentially requires 0 gaps (so all links are between consecutive frames). The results of the benchmark are shown in Figure 12C middle.

[0272] 4. Specific htSMT applications

[0273] Many, perhaps most, pathways that regulate fundamental cellular biochemistry rely on the transient interaction of protein sensors with protein effectors that trigger changes in cellular physiology. Although the fundamental principles of this process have long been recognized, biochemical studies of these protein interactions typically require in vitro reconstitution or interrogation via pull-down assays after cell permeabilization. The htSMT workflow described here provides a method for visualizing protein movements in large numbers of living cells, while allowing for quantitative assessment of the effects of added compounds, such as small molecule inhibitors.

[0274] refer to Figure 1 , various aspects of the htSMT workflow disclosed herein include, but are not limited to: (i) sample preparation (including reagent handling), (ii) image acquisition using sample imaging to generate a series of images and / or videos, (iii) image analysis by processing these images and videos, (iv) information storage, and (v) providing insights using the stored information (including biological interpretation). With respect to biological interpretation, the htSMT workflow described herein can provide specific insights, as described below, depending on the specific workflow employed, such as (i) htSMT screening; (ii) htSMT binding; and / or (iii) KineticSMT.

[0275] In certain embodiments, the workflow of the present disclosure may include illuminating a detection field of view disposed in a sample plane within a sample with a light beam to cause a subset of fluorescent target proteins in living cells to fluoresce, wherein the detection field of view has a size of about 50 μm to less than 100 μm in a first dimension and a size of about 50 μm to less than 100 μm in a second dimension. For example, but not by way of limitation, the detection FOV may have a size of about 94 μm in the first dimension and a size of about 94 μm in the second dimension.

[0276] In certain embodiments, the workflow of the present disclosure includes illuminating a field of view disposed in a sample plane within a sample with a light beam to cause a subset of fluorescent target proteins in living cells to fluoresce, thereby imaging a plurality of tracks. In certain embodiments, the number of tracks imaged in the detection field of view can be from about 10,000 to about 100,000, such as from about 20,000 to about 40,000. For example, but not by way of limitation, the number of tracks imaged in the detection field of view can be from about 10,000 to about 50,000, from about 10,000 to about 40,000, from about 20,000 to about 50,000, or from about 20,000 to about 40,000. In some embodiments, the number of tracks imaged in the detection field of view can be at most about 100,000, e.g., at most about 95,000, at most about 90,000, at most about 85,000, at most about 80,000, at most about 75,000, at most about 70,000, at most about 65,000, at most about 60,000, at most about 55,000, at most about 50,000, at most about 45,000, at most about 40,000, at most about 35,000, or at most about 30,000.

[0277] In certain embodiments, the detection field of view may include a plurality of cells. In certain embodiments, the number of cells imaged in the detection field of view is related to the size of the imaged cells. For example, but not limited to, the smaller the size of the cell, the larger the number of cells that can be imaged in the detection field of view. In certain embodiments, according to the size of the imaged cells, the detection field of view may include about 1 to about 50 living cells, for example, may include about 1 to about 45 cells, about 1 to about 40 cells, about 1 to about 35 cells, about 1 to about 30 cells, about 1 to about 25 cells, about 1 to about 20 cells, about 1 to about 15 cells, about 1 to about 10 cells, about 1 to about 5 cells, about 5 to about 40 cells, about 10 to about 40 cells, about 15 to about 40 cells, about 20 to about 40 cells, about 10 to about 35 cells, about 15 to about 35 cells or about 20 to about 30 cells. In certain embodiments, according to the size of the imaged cells, the detection field of view may include about 1 to about 40 living cells, such as mammalian cells. In certain embodiments, the detection field of view may include about 1 to about 30 living cells, such as mammalian cells, depending on the size of the cells being imaged. In certain embodiments, the detection field of view may include about 1 to about 20 living cells, such as mammalian cells, depending on the size of the cells being imaged. In certain embodiments, the detection field of view may include up to about 50 living cells, such as up to about 45 living cells, up to about 40 living cells, up to about 35 living cells, up to about 30 living cells, up to about 25 living cells, or up to about 20 living cells. In certain embodiments, the detection field of view may include up to about 40 cells. In certain embodiments, the detection field of view may include up to about 30 cells.

[0278] In certain embodiments, the workflows of the present disclosure may comprise detecting fluorescence of an individual fluorescent target protein from a plurality of fluorescent target proteins from a detection field of view of a sample plane at a rate of >100 detection FOVs per day, >10,000 detection FOVs per day, or >100,000 detection FOVs per day, wherein the detection field of view is about 50 μm to less than 100 μm in a first dimension and about 50 μm to less than 100 μm in a second dimension. In certain embodiments, the workflows of the present disclosure may comprise detecting fluorescence of an individual fluorescent target protein from a plurality of fluorescent target proteins from a detection field of view of a sample plane at a rate of >100 detection FOVs per day, wherein the detection field of view is about 50 μm to less than 100 μm in a first dimension and about 50 μm to less than 100 μm in a second dimension. In certain embodiments, the workflows of the present disclosure may comprise detecting fluorescence of an individual fluorescent target protein from a plurality of fluorescent target proteins from a detection field of view of a sample plane at a rate of >10,000 detection FOVs per day, wherein the detection field of view is about 50 μm to less than 100 μm in a first dimension and about 50 μm to less than 100 μm in a second dimension. In certain embodiments, the workflows of the present disclosure may comprise detecting fluorescence of an individual fluorescent target protein from a plurality of fluorescent target proteins from a detection field of view of a sample plane at a rate of >100,000 detection FOVs per day, wherein the detection field of view is about 50 μm to less than 100 μm in a first dimension and about 50 μm to less than 100 μm in a second dimension.

[0279] In certain embodiments, the workflow of the present disclosure may comprise illuminating a detection field of view disposed in a sample plane within a sample with a light beam to cause a subset of fluorescent target proteins in living cells to fluoresce, wherein the detection field of view has a size of about 50 μm to less than 100 μm in a first dimension and a size of about 50 μm to less than 100 μm in a second dimension, and wherein at most 70% of the detection field of view achieves sufficient laser illumination to track protein motion. In certain embodiments, the detection field of view has a size of about 50 μm to less than 100 μm in a first dimension and a size of about 50 μm to less than 100 μm in a second dimension, and wherein at most 60% of the detection field of view achieves sufficient laser illumination to track protein motion. In certain embodiments, the detection field of view has a size of about 50 μm to less than 100 μm in a first dimension and a size of about 50 μm to less than 100 μm in a second dimension, and wherein at most 50% of the detection field of view achieves sufficient laser illumination to track protein motion. In certain embodiments, the detection field of view has a size of about 50 μm to less than 100 μm in the first dimension and about 50 μm to less than 100 μm in the second dimension, and wherein at most 40% of the detection field of view achieves sufficient laser illumination to track protein motion. In certain embodiments, the detection field of view has a size of about 50 μm to less than 100 μm in the first dimension and about 50 μm to less than 100 μm in the second dimension, and wherein at most 30% of the detection field of view achieves sufficient laser illumination to track protein motion. In certain embodiments, the detection field of view has a size of about 50 μm to less than 100 μm in the first dimension and about 50 μm to less than 100 μm in the second dimension, and wherein at most 20% of the detection field of view achieves sufficient laser illumination to track protein motion. In certain embodiments, the detection field of view has a size of about 50 μm to less than 100 μm in a first dimension and about 50 μm to less than 100 μm in a second dimension, and wherein no more than 10% of the detection field of view achieves sufficient laser illumination to track protein motion.

[0280] In certain embodiments, the workflow of the present disclosure may include determining a change in the motion of the fluorescently labeled target protein in the presence of the compound. For example, and not by way of limitation, the average change in motion of the fluorescent target protein in the presence of the compound is at least 1%, at least 5%, or at least 10%, relative to the change observed in the absence of the compound. In certain embodiments, the average change in motion of the fluorescent target protein in the presence of the compound is about 1% to about 5%. In certain embodiments, the average change in motion of the fluorescent target protein in the presence of the compound is about 1% to about 10%.

[0281] In certain embodiments, exemplary htSMT workflows include individual strategies described above as well as combinations of these strategies where two or more strategic requirements are combined.

[0282] 4.1.htSMT screening

[0283] In certain implementations of the htSMT workflows described herein, the systems and methods are adapted to interrogate one or more compositions (e.g., "test" compounds) for their ability to affect the SMT profile associated with a labeled protein. For example, such an htSMT workflow would screen for changes in the SMT profile, e.g., an increase or decrease in the movement of a protein of interest in the presence of the composition relative to the SMT profile in the absence of the composition. It will be understood that higher order comparisons can also be made where compounds are multiplexed, including where multiple proteins fluoresce. In addition, as described above, the htSMT screening strategies described herein are equally adapted to screen for SMT profiles associated with fluorescent compounds, e.g., compounds that are naturally fluorescent or compounds that have been modified to fluoresce or are linked to a fluorophore.

[0284] Fundamental to such htSMT screening strategies is the ability of the htSMT workflow described here to extract accurate motion data at scale. Figures 13A-13E The exemplary results presented in demonstrate this capability. Figures 13A-13E In the reported experiments, 384-well plates were used, in which free Halo, Halo-CaaX, and H2B-Halo cell lines were mixed in equal proportions in each well. Imaging was performed with a 94 μm × 94 μm field of view (FOV), and an average of 10 cell nuclei ( Figure 13B 、 Figure 19B ), sufficient to ensure that most FOVs contain cells from each cell line. To limit ambiguity in cell assignment, only trajectories that lie within the nuclear segmentation region are considered. The probability distribution of motion clearly distinguishes the three cell types ( Figure 13C More importantly, by observing the single-cell state distribution profiles of 103,757 cells from five separate 384-well plates, grouped according to their distribution profiles, we recovered highly consistent estimates of protein movement at the single-cell level ( Figure 13D Furthermore, by combining data from multiple cells, the expected distribution can be provided ( Figure 13E ). In addition, combining data from many cells makes it possible to perform SpotOn or State Array analysis, where as few as 10 3Only a few seconds of imaging per FOV can generate enough trajectories (>10,000) to accurately estimate protein motion, bringing the platform's overall throughput to >90,000 FOV per day, a data acquisition rate that enables drug screening within a feasible timeframe.

[0285] Equipped with the htSMT system capable of measuring protein motion, the following disclosure establishes that measurement of protein motion can be used to functionally characterize proteins. For example, using the steroid hormone receptor (SHR), which switches between an inactive and an active state upon ligand binding ( Figure 14A ), the htSMT workflow described here can capture functionally relevant differences. Specifically, in the absence of hormones, the four SHRs exhibited similar motility profiles: a small immobile fraction and a large freely moving fraction, with average diffusion coefficients of 3.4–4.3 μm 2 / s( Figure 14B However, no correlation between movement and protein size was observed, highlighting the differences between cellular protein movement and purified systems. Highlighting the selectivity and sensitivity of htSMT, a sharp increase in immobile tracks was observed after addition of agonist, which was attributed to chromatin binding. Figure 14B As shown, the binding fraction (f 结合 ) is defined as a moving speed of less than 0.1 μm 2 The fraction of trajectories per second was 1.57 × 10 / s. Consistent with previous findings, some SHRs had a higher proportion of bound molecules than others, regardless of ligand presence. The ligand-induced effect was most pronounced for ER, with a binding rate of 34% under basal conditions and 87% after estradiol treatment ( Figure 14B ).

[0286] Supporting the use of SHR as a representative target family for htSMT analysis, SHR was highly selective for its cognate agonist in biochemical binding assays, as confirmed by measuring dose-dependent changes in motility as a function of agonist concentration. 结合 The maximum increase ( Figure 14C ) and the decrease in the free diffusion coefficient (D 自由 ; Figure 14D ) varied between SHRs. Dose titration curves also revealed that the potency (EC50) of each SHR / hormone pair varied, with ER-estradiol being the most potent and selective pair. Thus, screening by htSMT allows for precise and accurate differentiation of hormone / target specificity directly within the context of living cells.

[0287] In another example of htSMT screening assays, next-generation ER degraders such as GDC-0927, AZD9833, and GDC-9545 were optimized to enhance ER degradation. Compound-induced persistent changes in protein expression, such as ER degradation ( Figure 17A 、 Figure 22A Structural analogs of GDC-0927 have been reported and optimized for ER degradation, but the correlation between ER degradation and cell proliferation is poor ( Figure 17B 、 Figures 22B-22D However, by measuring protein persistence, a more accurate measure of inhibitory activity can be obtained than by assessing protein degradation. The potency and maximum effect of GDC-0927 structural analogs were determined using htSMT. Overall, these analogs exhibited a potency range of 15 pM to 12 nM and inhibited ERf 结合 Increased by 0.4 to 0.56 ( Figure 17C Small changes in chemical structure can result in measurable changes in compound potency and maximal efficacy as determined using htSMT.

[0288] The potency of GDC-0927 and analogs, as determined by ER degradation or htSMT, was compared to the ability of each of these compounds to block estrogen-induced breast cancer cell proliferation. Potency assessed by ER degradation was not a good predictor of potency in cell proliferation assays ( Figure 17C ). In contrast, f 结合 The SMT measurements were closely related to cell viability ( Figure 17D ; R of T47d 2 Interestingly, the SMT EC50 values ​​were on average 10-fold lower than those observed in the cell growth assay, suggesting that SMT is sensitive enough to allow the selection of chemical series that would not show effects in other cell-based assays. 结合 The correlation between the effect of ATP on the protein and its function (inhibition of cell proliferation), coupled with the throughput of the SMT system, makes it an attractive approach for identifying protein modulators with novel properties.

[0289] In addition to known ER activity modulators, many other compounds present in the tested bioactive library also elicited f 结合 To define the threshold for calling a screened molecule “active”, 92 f 结合 Compounds with different magnitudes of change were retested with dose titration ( Figure 23A 、 23B ). f 结合Using this approach, 239 compounds that affect ER mobility were identified in the bioactive library ( Figure 15 Among these compounds, the correlation between the two screening replicates was high (R 2 =0.92), and the activity levels were reproducible (slope of 0.94 for active molecules). Some active compounds could be clustered based on scaffold homology, but most clusters consisted of one or a few members ( Figure 23C 、 23D Structural clustering was used to identify known ER regulators where vendor-provided annotations were poorly defined. These results demonstrate that htSMT is reproducible and robust for screening large numbers of molecules.

[0290] Most of the active molecules screened were structurally unrelated to steroids ( Figure 23C 、 23D On the other hand, many compounds can be grouped according to their reported biological targets or pathways ( Figure 18A 、 Figure 24A 、 24B For example, heat shock proteins (HSPs) and proteasome inhibitors continuously increase f 结合 , while cyclin-dependent kinase (CDK) and mTOR inhibitors reduce f binding. Although many CDK inhibitors lack intrafamily specificity ( Figure 24A , pan-CDK), but CDK9-specific inhibitors were found to have a stronger effect on ER motility than CDK4 / 6-specific inhibitors. Furthermore, as with selective AR and GR antagonists, inhibitors targeting ALK, BTK, and FLT3 kinases that have not been shown to interact with the ER had no effect on ER motility when assessed using SMT ( Figure 18A ).

[0291] For identified cellular pathway inhibitors, dose titration was used to better characterize the effect of each inhibitor on ER motility. Potencies ranged from subnanomolar to low micromolar ( Figure 18B ), similar to the reported potency of these compounds against their cellular targets. In addition, these molecules were also tested against AR ( Figure 24C ) and PR( Figure 24D ) were tested. Each SHR was significantly different from the other SHRs in terms of the responses of compounds identified through ER-focused screening efforts. Similarly, the magnitude of the ER SMT effect was generally consistent within the target class ( Figure 18B 、 Figures 24A-24DThe finding that structurally distinct compounds exhibit similar effects based on their biological targets supports the idea that these biological targets themselves must interact with the ER, and therefore that these compounds indirectly affect ER motility. HSP90 is a chaperone for many proteins, including SHR. In the canonical model, hormone binding releases the SHR-HSP90 complex. Indeed, HSP90 inhibitors increase the f 结合 This is consistent with the fact that one of the functions of molecular chaperones is to regulate the balance of SHR binding to chromatin ( Figures 24B-24D ). Proteasome inhibition also resulted in ER pinning to chromatin, consistent with results obtained in htSMT screens of bioactive compounds. ER has been shown to be phosphorylated by CDKs, Src, or GSK-3 via the MAPK and PI3K / AKT signaling pathways, and therefore inhibition of these pathways would reasonably be expected to affect ER motility measured using SMT. While CDK inhibition resulted in increased ER mobility, inhibition of PI3K, AKT, or other upstream kinases showed no effect ( Figure 18A ).

[0292] Interestingly, SMT motility in the ER triple-point mutant was impaired by CDK and mTOR pathway inhibitors ( Figure 25 ), the mutant was engineered to lack previously defined phosphorylation sites (S104A / S106A / S118A) important for transcriptional activation, suggesting that additional phosphorylation sites may mediate the effects of CDK9 and PI3K / AKT signaling, or that other molecular targets of CDKs and PI3K / AKT may act indirectly to alter ER motility. For inhibitors of well-characterized pathways, such as those targeting CDKs and mTOR, changes in ER protein motility were subtle but consistent across compounds, indicating that these observations are biologically meaningful and highlighting the need for accurate and precise SMT measurements. Therefore, the htSMT technology described herein provides a way to provide comprehensive pathway interaction information.

[0293] To further demonstrate that monitoring changes in binding via changes in target motion can support the identification of pharmacologically relevant compounds, known AR agonists and antagonists were assayed ( Figure 26 ). Although AR agonists also increase f 结合 , but AR antagonists, whether used alone or in combination with AR agonists, can 结合 Therefore, f 结合 The increase and f 结合 Any reduction in activity helps identify mechanistically distinct mechanisms of pharmacological interaction with the observed fluorescent target protein.

[0294] In addition to changes in motion associated with chromatin binding, htSMT can also be used to monitor and identify pharmacological compounds that disrupt or enhance protein-protein interactions and protein conformational changes. One such example is the disruption of ubiquitination processes that rely on a series of protein-protein interactions. By using known antagonists to disrupt the underlying protein-protein interaction, during complex disruption, the motion of one of the proteins involved in the interaction (target A) will change significantly, indicating that the protein is more free to move ( Figure 27 Similarly, htSMT can be used to interrogate protein conformational changes and screen for compounds that allosterically modulate targets. ATP-competitive and allosteric inhibitors of various known receptor tyrosine kinases, such as target B and target C, cause significant changes in protein motion. Furthermore, htSMT is extremely sensitive to allosteric inhibition, resulting in motion changes that are 4-8 times greater than those of enzyme inhibitors ( Figure 28 ). Thus, the utility of htSMT can meaningfully interrogate protein-chromatin and protein-protein interactions.

[0295] 4.2.htSMT combination

[0296] In certain implementations of the htSMT workflow described herein, once a compound is identified that increases static binding of a target (e.g., static binding of ER to chromatin in the presence of the compound), this is sometimes referred to herein as increasing f 结合 , a second assay can be performed to obtain more detailed information about the static binding properties of the target. For example, the systems and methods described herein can be adapted to distinguish between recovery after exposure to a compound, where recovery is indicated by an increase in the residence time of the target with its binding partner (i.e., k* off reduction) driven.

[0297] For example, but not by way of limitation, a bioactivity screen of 5,067 molecules surprisingly revealed that all known ER modulators (agonists such as estradiol and potent antagonists such as fulvestrant) caused f 结合 A subset of selective ER modulators (SERMs) and selective ER degraders (SERDs) were subsequently evaluated in more detail. These molecules all competitively bind to the ER ligand binding domain. As with the bioactivity screening, both SERDs and SERMS increased f 结 combine( Figure 16A ) and slightly reduced the measured D 自由 ( Figure 21A ), with potencies ranging from 9 pM for GDC-0927 to 4.8 nM for GDC-0810 ( Figure 16B Despite the different physicochemical properties, the f of the five compounds increased within a few minutes after addition of the compounds. 结合Both increased ( Figure 16C 、 Figure 21B ), there is no evidence of a transient diffusion state ( Figure 21C Without being bound by theory, it appears that dissociation, dimerization, and chromatin association of the ER with the chaperone complex occur on rapid and seemingly comparable timescales. Because the individual steps in these transitions are indistinguishable, the association rate of the overall process is assumed to have an effective rate constant k* on Importantly, selective antagonists of the AR and GR did not induce significant modulation of ER motility, further highlighting the utility of the htSMT technology described here for characterizing the specificity of interactions between modulators of protein function and their cognate targets ( Figure 21D ).

[0298] Interestingly, the SERMs 4-hydroxytamoxifen (4OHT) and GDC-0810 showed lower f 结合 Maximum increase ( Figure 16D A similar effect has been previously described using fluorescence recovery after photobleaching (FRAP), and this has been confirmed using the Halo-ER cell line ( Figure 16E The delay in ER signal recovery after two minutes in FRAP is consistent with the f 结合 The changes are consistent ( Figure 16F ). Although FRAP is used to measure these f 结合 Although the technique has been shown to differ in performance, it faces challenges in scalability and relies heavily on prior assumptions about potential motility within the sample. In contrast, the htSMT technique described here allows for detailed characterization of the potency of 4OHT and GDC-0810 relative to other ER ligands in increasing ER chromatin binding.

[0299] Neither FRAP nor htSMT can distinguish the difference between off decrease) or increase in chromatin binding rate (k* on increase) and the recovery of the drive, both of which will lead to f 结合 By changing the SMT acquisition conditions to reduce illumination intensity and collect long frame exposures, only immobile proteins form spots. Under these imaging conditions, the distribution of track lengths provides a measure of relative dwell time. Both agonist and antagonist treatments resulted in longer binding times compared to DMSO, indicating that ligand binding reduces k*. off ( Figure 16G , Figure 21E ). Consistent with FRAP, estradiol, GDC-0927, and fulvestrant exhibited longer binding times compared to other ER modulators. 结合 and k* off The measured value can be inferred from k*on In all cases, the change in dissociation rate is related to f 结合 The increase in k* caused by the ligand is disproportionate. on The increase may lead to the observed changes in the chromatin-associated ER fraction ( Figure 21E Without being bound by theory, these data are consistent with a model in which ER rapidly associates with chromatin regardless of which molecule occupies the ligand-binding domain, but some ligands induce a conformation that can be further stabilized on chromatin by cofactors. Therefore, these data support that ER associates with chromatin in a mechanistically distinct manner. Potent ER inhibitors may promote rapid and transient chromatin association that is unable to effectively recruit the necessary cofactors to drive transcription.

[0300] 4.3.KineticSMT

[0301] In certain implementations of the htSMT workflow described herein, the systems and methods are adapted to identify the rate at which changes in protein motility occur. For example, in certain such implementations, htSMT can be used to distinguish between direct and indirect effects. Additionally or alternatively, identifying the rate at which changes in protein motility occur provides parameters for assessing cell permeability and / or active transport, target efflux and / or influx, target engagement rate, and / or target engagement and disengagement rate.

[0302] Given the live cell setup of SMT, a data collection mode was configured that allowed the measurement of protein movement at set time intervals after compound addition (kinetic SMT or kSMT). When measured as kSMT, both ER agonists and antagonists rapidly induced the immobilization of ER on chromatin (t for estradiol). 1 / 2 =1.6 minutes; Figure 16C On the other hand, HSP90 inhibitors such as gantapi and HSP990 showed a 5- to 7-minute delay before changes in ER motility were observed, after which f 结合 The increase of t 1 / 2 The overall effect of these compounds reached a plateau after one hour ( Figure 18C Proteasome inhibitors, such as bortezomib and carfilzomib, act more slowly, with changes in ER motility appearing only after 40 minutes and increasing slowly over the four-hour measurement window ( Figure 18D Similarly, differential kinetics of on-target, on-pathway, and off-target inhibitors that alter the motion of several other targets have been measured, including target A, target B, and helicase ( Figure 29 ). Exploration of SMT dynamics thus represents an important tool to facilitate the differentiation of on-target modulators from on-pathway modulators. For example, the kSMT technology described herein allows for rapid mechanistic characterization of active compounds in a drug discovery setting.

[0303] To further differentiate the effects of pathway inhibitors on ER protein motility, the relative ER residence time of each such molecule was characterized. Estradiol, SERMs, and SERDs all increased residence time and, therefore, may also increase the rate of ER association with chromatin ( Figures 16A-16D 、 Figure 21E ). In contrast, although HSP990 and inhibition of HSP90 by gantapi resulted in f 结合 However, the number of total length binding events was observed to decrease by two-fold and four-fold, respectively, while the binding time was similar to that observed with DMSO alone ( Figure 18E These results suggest that HSP90 inhibition primarily increases k* on , and k* off The ER-chromatin binding pathway was largely unaffected. On the other hand, proteasome inhibition resulted in an increase in both the number and duration of long binding events. These results suggest that ER-chromatin binding can be regulated by altering the rates of association or dissociation, and that inhibition of specific cellular partners can differentially affect these rates. Combined with the distinct kinetics of direct ER, HSP90, and proteasome modulators, these data suggest that each class of molecules alters ER motility through distinct mechanisms.

[0304] 5. Exemplary Implementation

[0305] A. The present disclosure provides a method comprising:

[0306] receiving a sequence of images visualizing molecular motion;

[0307] linking molecules across said images;

[0308] generating possible trajectories for each molecule and their associated probabilities based on the linkage using a variational Bayesian optimization algorithm; and

[0309] Data representing the possible trajectories generated and their associated probabilities is provided to a consuming application or process.

[0310] B. The present disclosure provides a method comprising:

[0311] receiving a sequence of images visualizing molecular motion;

[0312] linking molecules across said images;

[0313] generating possible trajectories for each molecule and their associated probabilities based on the linkage using a Gibbs sampling algorithm; and

[0314] Data representing the possible trajectories generated and their associated probabilities is provided to a consuming application or process.

[0315] C. The present disclosure provides a method comprising:

[0316] receiving a sequence of images visualizing molecular motion;

[0317] linking molecules across said images;

[0318] generating possible trajectories for each molecule and their associated probabilities based on the links using an adaptive hill climbing algorithm; and

[0319] Data representing the possible trajectories generated and their associated probabilities is provided to a consuming application or process.

[0320] C1. The method of any one of A to C, wherein at least a subset of the sequence of images comprises at least 100 molecules per image.

[0321] C2. The method of any one of A to C1, wherein at least a subset of the sequence of images comprises at least 1000 molecules per image.

[0322] C3. The method of any one of A to C2, wherein at least a subset of the sequence of images comprises at least 10,000 molecules per image.

[0323] C4. The method of any one of A to C3, wherein the molecules have a density of at least 0.01 emitters per square micron per image.

[0324] C5. The method of any one of A to C4, wherein the molecules have a density of at least 0.1 emitters per square micron per image.

[0325] C6. The method of any one of A to C5, further comprising:

[0326] labeling molecules within biological samples;

[0327] causing the biological sample to fluoresce; and

[0328] The sequence of images is generated while causing the biological sample to fluoresce.

[0329] C7. A method as described in C6, wherein the image sequence is generated using a microscope system.

[0330] C8. The method of any one of A to C7, wherein the molecule is imaged within a living cell.

[0331] C9. The method of any one of A to C8, further comprising:

[0332] A probabilistic dynamic model is inferred that contains information characterizing the molecular trajectory.

[0333] C10. The method of C9, wherein the probabilistic dynamic model comprises a state array, and the method further comprises: populating the state array with information characterizing the molecular trajectory.

[0334] C11. The method according to any one of A to C10, further comprising:

[0335] An internal confidence indicator based on the associated probability is generated, wherein the provided data includes the generated internal confidence indicator.

[0336] C12. A method as described in C11, wherein the generated internal confidence indicator is a tracking error rate lower bound, which defines a lower bound on the incorrect connection rate caused by the link.

[0337] C13. The method of C11, wherein the generated internal indicators include:

[0338] Compute the confidence score for each trajectory.

[0339] C14. The method according to any one of A to C13, further comprising:

[0340] Generate dynamic metrics independent of specific trajectories.

[0341] C15. A method as described in any one of A to C14, wherein the linking includes retrieving data containing multiple statistical data extracted from the total number of detections or the number of detections in the cell.

[0342] C16. A method as described in any one of A to C15, wherein providing data includes one or more of the following: visualizing at least a portion of the generated possible trajectories and their associated probabilities in a graphical user interface, storing at least a portion of the generated possible trajectories and their associated probabilities in physical persistence, loading at least a portion of the generated possible trajectories and their associated probabilities into a memory, or transmitting at least a portion of the generated possible trajectories and their associated probabilities to a remote computing device over a network.

[0343] C17. A method as described in any of A to C16, wherein at least a portion of the image sequence includes consecutive images from a corresponding movie.

[0344] C18. The method of any one of A to C17, wherein at least a portion of the sequence of images used for the linking are non-consecutive images from a corresponding movie.

[0345] D. The present disclosure provides a method for single molecule tracking, comprising:

[0346] receiving a sequence of images visualizing molecular motion, the sequence of images comprising a first type generated using a first imaging modality and a second type generated using a second, different imaging modality;

[0347] detecting points within a sequence of images of said first type;

[0348] linking the points detected within the first type of image sequence into trajectories using a probabilistic tracking algorithm;

[0349] segmenting the second type of image sequence to generate a plurality of instance masks;

[0350] assigning molecules within the second type of image sequence to at least one instance mask of the plurality of instance masks; and

[0351] Data representing the link and assignment is provided to a consuming application or process.

[0352] D1. The method as described in D, wherein the probabilistic tracking algorithm includes a variational Bayesian optimization algorithm.

[0353] D2. The method as described in D, wherein the probabilistic tracking algorithm includes a Gibbs sampling algorithm.

[0354] D3. The method as described in D, wherein the partial probabilistic tracking algorithm includes an adaptive hill climbing algorithm.

[0355] D4. The method of any one of D to D3, wherein the first imaging modality and the second imaging modality comprise different molecular labeling technologies.

[0356] D5. The method of any one of D to D4, wherein the first type of image sequence is a single molecule tracking (SMT) movie and the second type of image sequence is a non-SMT movie.

[0357] D6. The method of any one of D to D5, wherein the point of detection comprises a subcellular component.

[0358] D7. The method of any one of D to D6, wherein the molecular types within the sequence of images of the first type are labeled with different fluorophores.

[0359] D8. The method of any one of D to D7, wherein at least a subset of the sequence of images comprises at least 100 molecules per image.

[0360] D9. The method of any one of D to D8, wherein at least a subset of the sequence of images comprises at least 1000 molecules per image.

[0361] D10. The method of any one of D to D9, wherein at least a subset of the sequence of images comprises at least 10,000 molecules per image.

[0362] D11. The method of any one of D to D10, wherein the molecules have a density of at least 0.01 emitters per square micron per image.

[0363] D12. The method of any one of D to D11, wherein the molecules have a density of at least 0.1 emitters per square micron per image.

[0364] D13. The method according to any one of D to D12, further comprising:

[0365] labeling molecules within biological samples;

[0366] causing the biological sample to fluoresce; and

[0367] At least a portion of the sequence of images is generated while causing the biological sample to fluoresce.

[0368] D14. The method of D13, wherein the image sequence is generated using a microscope system.

[0369] D15. The method of any one of D to D14, wherein the molecule is imaged within a living cell.

[0370] D16. The method according to any one of D to D15, further comprising:

[0371] A probabilistic dynamic model is inferred that contains information characterizing the molecular trajectory.

[0372] D17. A method as described in D16, wherein the probabilistic dynamic model includes a state array, and the method further includes: filling the state array with information characterizing the molecular trajectory.

[0373] D18. The method according to any one of D to D17, further comprising:

[0374] An internal confidence indicator based on the associated probability is generated, wherein the provided data includes the generated internal confidence indicator.

[0375] D19. A method as described in D18, wherein the internal confidence indicator generated is a tracking error rate lower bound (ERLB), which defines a lower bound on the error connection rate caused by the link.

[0376] D20. The method of D18, wherein the generated internal indicators include:

[0377] Compute the confidence score for each trajectory.

[0378] D21. The method according to any one of D to D20, further comprising:

[0379] Generate dynamic metrics independent of specific trajectories.

[0380] D22. A method as described in any one of D to D21, wherein the linking includes retrieving data containing multiple statistical data extracted from the total number of detections or the number of detections in the cell.

[0381] D23. A method as described in any one of D to D22, wherein providing data includes one or more of the following: visualizing at least a portion of the generated possible trajectories and their associated probabilities in a graphical user interface, storing at least a portion of the generated possible trajectories and their associated probabilities in physical persistence, loading at least a portion of the generated possible trajectories and their associated probabilities into a memory, or transmitting at least a portion of the generated possible trajectories and their associated probabilities to a remote computing device via a network.

[0382] D24. The method according to any one of D to D23, further comprising:

[0383] A plurality of statistical metrics associated with the at least one track or the at least one instance mask are generated.

[0384] D25. The method of D24, further comprising:

[0385] A hierarchy storing instance masks.

[0386] D26. A method as described in any of D to D25, wherein at least a portion of the image sequence includes consecutive images from a corresponding movie.

[0387] D27. The method of any one of D to D26, wherein at least a portion of the sequence of images used for the linking are non-consecutive images from a corresponding movie.

[0388] D28. The method of any one of D to D27, wherein the detecting utilizes one or more of:

[0389] Generalized log-likelihood point detector, Difference of Gaussian (DoG) detector, Laplace of Gaussian (LoG) detector, or Determinant of Hessian (DoH) blob detector.

[0390] D29. The method of any one of D to D28, further comprising:

[0391] Detected points are associated with spatiotemporal coordinates using sub-pixel localization.

[0392] D30. The method of D29, wherein sub-pixel positioning comprises one or more of the following:

[0393] Radially symmetric locators or maximum likelihood fitting of candidate point models using the Levenberg-Marquardt method.

[0394] D31. A method as described in any of D to D30, wherein the field of view corresponds to at least a portion of the hole.

[0395] D32. The method of any one of A to D31, wherein the sequence of images is generated by an apparatus for fluorescence microscopy, the apparatus comprising:

[0396] a light source capable of emitting fluorescent excitation light, wherein the light source exhibits a power output drift of less than about 10% at an ambient temperature of 17°C + / - 5°C;

[0397] a first optical element or assembly configured to receive a fluorescent excitation light source and shape the fluorescent excitation light source to form a light beam;

[0398] a second optical element or assembly comprising a water immersion objective configured to tilt the light beam in an xz plane relative to the z-axis, wherein the second optical element is further configured to focus the light beam at a sample plane located in the xy plane, thereby illuminating at least a portion of the sample plane; and

[0399] A detector arrangement is configured to receive light from the illuminated portion of the sample plane, wherein the detector arrangement forms one or more projected images based on the light received from the illuminated portion of the sample plane.

[0400] D33. The method of D32, wherein the apparatus comprises a second objective lens configured to direct light emitted from the illuminated portion of the sample plane to the detector arrangement.

[0401] D34. A method as described in D32 or D33, wherein the detector device includes a semiconductor sensor.

[0402] D35. A method as described in any of D32 to D34, wherein the device includes a third optical element or component configured to translate the light beam in the imaging plane along a direction orthogonal to the longer dimension of the light beam.

[0403] D36. A method as described in D35, wherein the third optical element or component includes a galvanometer.

[0404] D37. A method as described in any one of D32 to D36, wherein the detector device includes a semiconductor sensor, wherein the detector device supports a shutter mode for synchronizing the translation of the light beam in the sample plane with the selective activation or readout of the semiconductor sensor.

[0405] D38. The method of any one of A to D31, wherein the sequence of images is generated by a microscope system for tracking molecular motion, the microscope system comprising:

[0406] a stage for supporting a sample, wherein the sample contains the molecule;

[0407] a light source for emitting a light beam capable of inducing a light-based reaction from the molecules in the sample, wherein the light source exhibits a power output drift of less than about 10% at an ambient temperature of 17°C + / - 5°C;

[0408] a water immersion objective for focusing the light beam onto at least a portion of the sample plane, wherein the molecules are disposed within the sample plane; and

[0409] A detector device is used to monitor the light-based reaction of the molecule, analyze the light-based reaction, and thus track the movement of the molecule.

[0410] D39. A method as described in D38, wherein the microscope system also includes a scanning optical element or component, which is configured to translate the light beam in the sample plane in a direction orthogonal to the longer dimension of the light beam, thereby enabling the microscope system to have a larger total field of view in the xy plane.

[0411] D40. The method of D39, wherein the microscope system further comprises a z position controller for the sample plane, wherein the z position controller is capable of maintaining focus in the z direction.

[0412] D41. The method of any one of D38 to D40, wherein the sample is disposed within an open well of the sample plate.

[0413] D42. The method of any one of D41, wherein the sample plate comprises a plurality of open wells.

[0414] D43. A method as described in D41 or D42, wherein the microscope system further includes an xy position controller for changing the field of view of the microscope system, and the changed field of view covers different subsets of the multiple open holes.

[0415] D44. The method of any one of D41 to D43, wherein the microscope system further comprises a temperature-controlled environment configured to control the environment of the sample plate.

[0416] D45. The method of D44, wherein the samples disposed within the open wells of the sample plate are maintained at a humidity of 20%-95%.

[0417] D46. The method of D44 or D45, wherein the sample disposed within the open well of the sample plate is maintained at 5% CO2.

[0418] D47. The method of any one of D38 to D46, wherein the microscope system further comprises an automated sample handling robotic system to enable high-throughput manipulation of multiple samples on the stage, wherein the robotic system comprises:

[0419] Memory;

[0420] a processor in communication with the memory; and

[0421] One or more robotic end effectors in communication with the processor, wherein the one or more end effectors manipulate the plurality of samples on the stage based on communication with the processor.

[0422] E. The present disclosure provides a system comprising:

[0423] at least one data processor; and

[0424] A memory storing instructions which, when executed by at least one data processor, result in the implementation of the operations of any one of the methods described in A to D31.

[0425] E1. The system of E, further comprising:

[0426] A fluorescence microscope device comprising:

[0427] a light source capable of emitting fluorescent excitation light, wherein the light source exhibits a power output drift of less than about 10% at an ambient temperature of 17°C + / - 5°C;

[0428] a first optical element or assembly configured to receive a fluorescent excitation light source and shape the fluorescent excitation light source to form a light beam;

[0429] a second optical element or assembly comprising a water immersion objective configured to tilt the light beam in an xz plane relative to the z-axis, wherein the second optical element is further configured to focus the light beam at a sample plane located in the xy plane, thereby illuminating at least a portion of the sample plane; and

[0430] A detector arrangement is configured to receive light from the illuminated portion of the sample plane, wherein the detector arrangement forms one or more projected images based on the light received from the illuminated portion of the sample plane.

[0431] E2. The system of E1, wherein the apparatus comprises a second objective lens configured to direct light emitted from the illuminated portion of the sample plane to the detector arrangement.

[0432] E3. A system as described in E1 or E2, wherein the detector device includes a semiconductor sensor.

[0433] E4. A system as described in any of E1 to E3, wherein the device includes a third optical element or component configured to translate the light beam in an imaging plane along a direction orthogonal to the longer dimension of the light beam.

[0434] E5. A system as described in E4, wherein the third optical element or component includes a galvanometer.

[0435] E6. A system as described in any of E1 to E5, wherein the detector device includes a semiconductor sensor, wherein the detector device supports a shutter mode for synchronizing the translation of the light beam in the sample plane with the selective activation or readout of the semiconductor sensor.

[0436] E7. The system of E, further comprising:

[0437] A microscope system for tracking molecular motion, comprising:

[0438] a stage for supporting a sample, wherein the sample contains the molecule;

[0439] a light source for emitting a light beam capable of inducing a light-based reaction from the molecules in the sample, wherein the light source exhibits a power output drift of less than about 10% at an ambient temperature of 17°C + / - 5°C;

[0440] a water immersion objective for focusing the light beam onto at least a portion of the sample plane, wherein the molecules are disposed within the sample plane; and

[0441] A detector device is used to monitor the light-based reaction of the molecule, analyze the light-based reaction, and thus track the movement of the molecule.

[0442] E8. A system as described in E7, wherein the microscope system also includes a scanning optical element or component that is configured to translate the light beam in the sample plane in a direction orthogonal to the longer dimension of the light beam, thereby enabling the microscope system to have a larger total field of view in the xy plane.

[0443] E9. The system of E8, wherein the microscope system further comprises a z position controller for the sample plane, wherein the z position controller is capable of maintaining focus in the z direction.

[0444] E10. The system of any one of E7 to E9, wherein the sample is disposed within an open well of the sample plate.

[0445] E11. The system of E10, wherein the sample plate comprises a plurality of open wells.

[0446] E12. The system of E11, wherein the microscope system further comprises an xy position controller for changing the field of view of the microscope system, wherein the changed field of view encompasses a different subset of the plurality of open holes.

[0447] E13. The system of E11 or E12, wherein the microscope system further comprises a temperature-controlled environment configured to control the environment of the sample plate.

[0448] E14. The system of E13, wherein the samples disposed within the open wells of the sample plate are maintained at a humidity of 20%-95%.

[0449] E15. The system of E13 or E14, wherein the samples disposed within the open wells of the sample plate are maintained at 5% CO2.

[0450] E16. The system of any one of E7 to E15, wherein the microscope system further comprises an automated sample handling robotic system to enable high-throughput manipulation of multiple samples on the stage, wherein the robotic system comprises:

[0451] a memory for storing instructions;

[0452] at least one data processor; and

[0453] One or more robotic end effectors in communication with the at least one data processor, wherein the one or more end effectors manipulate the plurality of samples on the stage based on communication with the at least one data processor.

[0454] F. The present disclosure provides a non-transitory computer program product storing instructions that, when executed by at least one data processor forming part of at least one computing device, implement the method described in any one of A to D31.

[0455] G. The present disclosure provides a system comprising:

[0456] a component for receiving a sequence of images visualizing molecular motion;

[0457] a means for linking molecules across said images;

[0458] means for generating possible trajectories for each molecule and their associated probabilities based on the links; and

[0459] A component that provides data representing possible generated trajectories and their associated probabilities to a consuming application or process.

[0460] H. The present disclosure provides a system comprising:

[0461] a component for receiving a sequence of images visualizing molecular motion;

[0462] a means for linking molecules across said images;

[0463] means for generating possible trajectories for each molecule and their associated probabilities based on the linkage using a variational Bayesian optimization algorithm; and

[0464] A component that provides data representing possible generated trajectories and their associated probabilities to a consuming application or process.

[0465] I. The present disclosure provides a system comprising:

[0466] a component for receiving a sequence of images visualizing molecular motion;

[0467] a means for linking molecules across said images;

[0468] means for generating possible trajectories for each molecule and their associated probabilities based on the linkage using a Gibbs sampling algorithm; and

[0469] A component that provides data representing possible generated trajectories and their associated probabilities to a consuming application or process.

[0470] J. The present disclosure provides a system comprising:

[0471] a component for receiving a sequence of images visualizing molecular motion;

[0472] a means for linking molecules across said images;

[0473] means for generating possible trajectories for each molecule and their associated probabilities based on the links using an adaptive hill climbing algorithm; and

[0474] A component that provides data representing possible generated trajectories and their associated probabilities to a consuming application or process.

[0475] K. The present disclosure provides a single molecule tracking system, comprising:

[0476] means for receiving a sequence of images visualizing molecular motion, the sequence of images comprising a first type generated using a first imaging modality and a second type generated using a second, different imaging modality;

[0477] means for detecting points within the sequence of images of said first type;

[0478] for linking points detected within the sequence of images of the first type to members in a trajectory using a probabilistic tracking algorithm;

[0479] means for segmenting the second type of image sequence to generate a plurality of instance masks;

[0480] means for assigning molecules within the sequence of images of the second type to at least one instance mask of the plurality of instance masks; and

[0481] Means for providing data representing the link and assignment to a consuming application or process.

[0482] 6. Examples

[0483] Example 1: High-throughput single-molecule tracking of steroid hormone receptors

[0484] A. Introduction

[0485] Steroid hormone receptors (SHRs) are a class of transcription factors that play crucial roles in normal human development and disease pathogenesis. For example, estrogen receptors (ESR1 and ESR2), androgen receptors (AR), and progesterone receptors (PR) are crucial for the acquisition of secondary sexual characteristics, while glucocorticoid receptors (GRs) help coordinate metabolism and inflammation. In the ligand-free state, SHRs are sequestered in a multiprotein complex by the molecular chaperone HSP9021. Typically, in the presence of hormones, they dimerize and bind to their cognate genomic response elements, thereby recruiting epigenetic modifiers and the transcriptional machinery. Furthermore, SHR-derived signals contribute significantly to the disease burden by promoting breast cancer (ER) or prostate cancer (AR) growth or by causing immune and metabolic dysfunction (GR). Therefore, given the wealth of information and reagents already available for these systems, as well as previous reports describing certain aspects of their cellular motility, SHRs provide an excellent proof-of-concept system for studying protein motility as a determinant of protein function.

[0486] This embodiment describes industrial-scale htSMT technology, systems comprising such htSMT technology, hardware and software associated with such htSMT technology, and methods of using such htSMT technology. For example, the htSMT technology described herein is capable of measuring protein movement in >1,000,000 cells per day. Additionally, using ER as a proof-of-concept system, the htSMT technology described herein demonstrated specific, robust, and reproducible results. The htSMT technology described herein can be used for a variety of applications, including but not limited to classical drug discovery activities, such as compound library screening and the elucidation of SAR. Importantly, the htSMT technology described herein can be used to characterize the contributions of known and novel pathways to interaction networks, such as protein signaling interaction networks.

[0487] B. Results

[0488] a. Creation and verification of htSMT system

[0489] A robotic system was developed that can handle reagents, collect high-quality, rapid SMT image series, process time-sequenced raw images to generate molecular trajectories, and extract biologically interesting features within defined cellular compartments. Figure 1 ). To examine the performance of the htSMT system over a wide range of diffusion coefficients, three U2OS cell lines expressing HaloTag fusion proteins with well-established behavior in cells were generated. Histone H2B-Halo is primarily integrated into chromatin and therefore cannot move efficiently over short periods of time and was used to estimate localization errors. The prenylated motif (Halo-CaaX) embedded in the plasma membrane exhibits moderate diffusion. Unfused HaloTag was chosen to represent the upper limit of "free" diffusion in the cell. Single-molecule trajectories measured in these cell lines yielded the expected diffusion coefficients ( Figure 13A Using the stationary H2B-Halo track, the positioning error of the htSMT system was found to be 39 nm ( Figure 19A ), which is comparable to other benchmark stroboscopic lighting datasets.

[0490] We tested the ability of the htSMT platform to extract accurate molecular trajectories on a large scale. A 384-well plate was used, in which free Halo, Halo-CaaX, and H2B-Halo cell lines were mixed in equal proportions in each well. Images were taken with a 94 μm × 94 μm field of view (FOV), and an average of 10 nuclei ( Figure 13B 、 Figure 20B ), sufficient to ensure that most FOVs contain cells from each cell line. To limit ambiguity in cell assignment, only trajectories located within the nuclear segmentation region were considered. The probability distribution of diffusion states clearly distinguishes the three cell types ( Figure 13CMore importantly, by observing the single-cell state distribution profiles of 103,757 cells from five separate 384-well plates, grouped according to their distribution profiles, we recovered highly consistent estimates of protein movement at the single-cell level ( Figure 13D ).

[0491] While single-cell measurements are powerful, the number of trajectories in a cell is limited, and thus estimates of the diffusion state can be wide. However, combining trajectories from multiple cells provides an expected distribution of diffusion states ( Figure 13E Furthermore, combining trajectories from many cells enables SpotOn or State Array analysis, where as few as 103 trajectories can satisfactorily infer the underlying diffusion state. Just a few seconds of imaging per FOV yields enough trajectories (>10,000) to accurately estimate protein motion, bringing the overall throughput of the platform to 13,000 individual wells (>90,000 FOVs per day; >1,000,000 cells per day), a data acquisition rate that enables drug screening within a feasible timeframe ( Figure 30 ).

[0492] b. Measuring protein movement in SHR using htSMT

[0493] Equipped with the htSMT system capable of measuring a wide range of protein motions, the following work established that measurements of protein motion can be used to functionally characterize proteins. SHR transitions between inactive and active states upon ligand binding ( Figure 14A ), and htSMT can capture these differences. To this end, HaloTag-fused ER, AR, PR, and GR cell lines were generated in the U2OS cell background to minimize the impact of comparing motility in different cell types. Clones were carefully selected so that HaloTag-fused SHR transcript abundance was comparable to each other and no higher than transcript levels in tissue-specific cell lines such as MCF7 and T47d, which are both ER and PR positive ( Figure 31 ).

[0494] In the absence of hormone, all four proteins exhibited similar motion profiles: a small immobile fraction and a large freely diffusing fraction with average diffusion coefficients of 3.4–4.3 μm 2 / s( Figure 14B No correlation between diffusion and protein size was observed, highlighting the differences between cellular protein motility and purified systems. Upon addition of agonist, a sharp increase in immobile tracks was observed, which was attributed to chromatin binding. The bound fraction (f) of each SHR was 1.37 × 10 Å. 结合 ) is defined as a diffusion velocity less than 0.1 μm 2 / s trajectory score ( Figure 14B Consistent with previous findings, some SHRs had a higher proportion of bound molecules than others, regardless of ligand presence. The ligand-induced effect was most pronounced for ER, with 34% binding under basal conditions and 87% binding after estradiol treatment ( Figure 14B ).

[0495] SHR are highly selective for their cognate agonists in biochemical binding assays, as confirmed by measuring dose-dependent changes in locomotion as a function of agonist concentration. 结合 The maximum increase ( Figure 14C ) and the decrease in the free diffusion coefficient (D 自由 ; Figure 14D ) varied between SHRs. Dose titration curves also revealed that the potency (EC50) of each SHR / hormone pair varied, with ER-estradiol being the most potent and selective pair. RNA-seq after estradiol stimulation revealed a clear induction of a signature ER-dependent gene set, confirming that the increased chromatin binding we observed with SMTs has a functional role in promoting ER-responsive gene programs, even in an ectopic expression setting ( Figure 32 、 Figure 33 ). Thus, SMT can precisely and accurately discriminate between antigen / target specificities directly within the living cell environment.

[0496] c. Screening of diverse panels of bioactive chemicals can identify known and novel regulators of ER dynamics

[0497] The characterization of ligand selectivity for AR, ER, GR, and PR collectively demonstrates that SMT can be used to interrogate the effects of compounds on protein dynamics at a throughput conducive to high-throughput screening. The specificity and sensitivity of the htSMT platform were next examined. A panel of 5,067 structurally diverse molecules with heterogeneous bioactivity against the ER was screened, and the f-values ​​of 1 μM compounds compared to DMSO were evaluated. 结合 change( Figure 15 The screen was run twice to assess reproducibility, and the results showed a high degree of consistency between replicates of the ER-active molecules ( Figure 15 Each compound was measured from 94 to 161 cells (25th to 75th percentile; Figure 20A This screen illustrates an important advantage of the htSMT platform described here over more manual, low-throughput methods.

[0498] The assay window for screening was robust from plate to plate. ( Figure 20B ; mean Z' factor = 0.79), and the measured potency of control estradiol remained within threefold of the mean in each case ( Figure 20C and20D ), and the distribution of negative control wells is tightly concentrated on zero ( Figure 20E ). The 30 compounds identified from the bioactive panel that showed promise in modulating ER (either as agonists or antagonists) significantly increased the SMT-measured f 结合 , including notable examples such as 4-hydroxytamoxifen, fulvestrant, and bactroxifene ( Figure 15 and Figure 35 ).

[0499] Somewhat counterintuitively, it has been reported that either strong stimulation or antagonism of the ER can lead to increased chromatin binding, but this does not appear to be a universal feature of SHR. Although the PR antagonist mifepristone has effects similar to those of the ER antagonist ( Figures 36A-36B ), but AR antagonists (such as enzalutamide and darolutamide) and GR antagonists (such as AL082D06) reduce chromatin binding. This reduction occurs when administered alone or when coadministered with a competing cognate agonist ( Figures 36C-36D These results suggest that the cellular context and interaction partners are crucial for understanding the effects of a compound on its intended target. To emphasize this point, in addition to binders to the ligand-binding domain of the ER, a number of active compounds were identified that target different nodes in the ER interaction network, including modulators of the proteasome, molecular chaperones, kinases, and others. Figure 18A ).

[0500] d. Cellular ER dynamics elucidating the structure-activity relationship (SAR) of ER modulators

[0501] A bioactivity screen of 5,067 molecules surprisingly revealed that all known ER modulators (agonists such as estradiol and potent antagonists such as fulvestrant) caused f 结合 A subset of selective ER modulators (SERMs) and selective ER degraders (SERDs) were subsequently evaluated in more detail. These molecules all competitively bind to the ER ligand binding domain. As with the bioactivity screening, both SERDs and SERMS increased f 结合 ( Figure 16A ) and slightly reduced the measured D 自由 ( Figure 21A ), with potencies ranging from 9 pM for GDC-0927 to 4.8 nM for GDC-0810 ( Figure 16B Despite the different physicochemical properties, the f of the five compounds increased within a few minutes after addition of the compounds. 结合 Both increased ( Figure 16C 、 Figure 21B ), there is no evidence of a transient diffusion state ( Figure 21CWithout being bound by theory, it appears that dissociation, dimerization, and chromatin association of the ER with the chaperone complex occur on rapid and seemingly comparable timescales. Because the individual steps in these transitions are indistinguishable, the association rate of the overall process is assumed to have an effective rate constant k* on Importantly, selective antagonists of the AR and GR did not induce significant modulation of ER motility, further highlighting the utility of the htSMT technology described here for characterizing the specificity of interactions between modulators of protein function and their cognate targets ( Figure 21D ).

[0502] Interestingly, the SERMs 4-hydroxytamoxifen (4OHT) and GDC-0810 showed lower f 结合 Maximum increase ( Figure 16D A similar effect has been previously described using fluorescence recovery after photobleaching (FRAP), and this has been confirmed using the Halo-ER cell line ( Figure 16E The delay in ER signal recovery after two minutes in FRAP is consistent with the f 结合 The changes are consistent ( Figure 16F ). Although FRAP is used to measure these f 结合 Although the technique has been shown to differ in performance, it faces challenges in scalability and relies heavily on prior assumptions about potential motility within the sample. In contrast, the htSMT technique described here allows for detailed characterization of the potency of 4OHT and GDC-0810 relative to other ER ligands in increasing ER chromatin binding.

[0503] Neither FRAP nor htSMT can distinguish the difference between off decrease) or increase in chromatin binding rate (k* on increase) and the recovery of the drive, both of which will lead to f 结合 By changing the SMT acquisition conditions to reduce illumination intensity and collect long frame exposures, only immobile proteins form spots. Under these imaging conditions, the distribution of track lengths provides a measure of relative dwell time. Both agonist and antagonist treatments resulted in longer binding times compared to DMSO, indicating that ligand binding reduces k*. off ( Figure 16G , Figure 21E ). Consistent with FRAP, estradiol, GDC-0927, and fulvestrant exhibited longer binding times compared to other ER modulators. 结合 and k* off The measured value can be inferred from k* on In all cases, the change in dissociation rate is related to f 结合 The increase in k* caused by the ligand is disproportionate.on The increase may lead to the observed changes in the chromatin-associated ER fraction ( Figure 21E Without being bound by theory, these data are consistent with a model in which ER rapidly associates with chromatin regardless of which molecule occupies the ligand-binding domain, but some ligands induce a conformation that can be further stabilized on chromatin by cofactors. Therefore, these data support that ER associates with chromatin in a mechanistically distinct manner. Potent ER inhibitors may promote rapid and transient chromatin association that is unable to effectively recruit the necessary cofactors to drive transcription.

[0504] e.htSMT can define relevant structure-activity relationships for ER antagonists

[0505] As the name suggests, next-generation ER degraders (such as GDC-0927, AZD9833, and GDC-9545) are optimized to enhance ER degradation. Indeed, compound-induced ER degradation was observed by immunofluorescence in established breast cancer model lines and the U2OS ectopic expression system ( Figure 17A 、 Figure 22A Structural analogs of GDC-0927 have been reported and optimized for ER degradation, but the correlation between ER degradation and cell proliferation is poor ( Figure 17B 、 Figures 22B-22D However, by measuring protein movement, a more precise measure of inhibitory activity can be obtained than by assessing protein degradation. The potency and maximum effect of GDC-0927 structural analogs were determined using htSMT. Overall, these analogs exhibited a potency range of 15 pM to 12 nM and inhibited ERf 结合 Increased by 0.4 to 0.56 ( Figure 17C Small changes in chemical structure can result in measurable changes in compound potency and maximal efficacy as determined using SMT.

[0506] The potency of GDC-0927 and analogs, as determined by ER degradation or SMT, was compared to the ability of each of these compounds to block estrogen-induced breast cancer cell proliferation. Potency assessed by ER degradation was not a good predictor of potency in cell proliferation assays ( Figure 17C ). In contrast, f 结合 The SMT measurements were closely related to cell viability ( Figure 16E ; R of T47d 2 Interestingly, the SMT EC50 values ​​were on average 10-fold lower than those observed in the cell growth assay, suggesting that SMT is sensitive enough to allow the selection of chemical series that would not show effects in other cell-based assays. 结合The correlation between the effect of ATP on the protein and its function (inhibition of cell proliferation), coupled with the throughput of the SMT system, makes it an attractive approach for identifying protein modulators with novel properties.

[0507] f. Individual pathway-interacting proteins exhibit unique phenotypes associated with their effects on ER motility

[0508] In addition to known ER activity modulators, many other compounds in our bioactive library also elicited f 结合 To define the threshold for calling a screened molecule “active”, 92 f 结合 Compounds with different magnitudes of change were retested with dose titration ( Figure 23A 、 23B ). f 结合 Using this approach, 239 compounds that affect ER mobility were identified in the bioactive library ( Figure 15 Among these compounds, the correlation between the two screening replicates was high (R 2 =0.92), and the activity levels were reproducible (slope of 0.94 for active molecules). Some active compounds could be clustered based on scaffold homology, but most clusters consisted of one or a few members ( Figure 23C 、 23D Structural clustering was used to identify known ER regulators where vendor-provided annotations were poorly defined ( Figure 22C These results demonstrate that htSMT is reproducible and robust for screening large numbers of molecules.

[0509] Most of the active molecules screened were structurally unrelated to steroids ( Figure 23C 、 23D On the other hand, many compounds can be grouped according to their reported biological targets or pathways ( Figure 18A 、 Figure 24A 、 24B For example, heat shock proteins (HSPs) and proteasome inhibitors continuously increase f 结合 , while cyclin-dependent kinase (CDK) and mTOR inhibitors reduce f binding. Although many CDK inhibitors lack intrafamily specificity ( Figure 24A , pan-CDK), but CDK9-specific inhibitors were found to have a stronger effect on ER motility than CDK4 / 6-specific inhibitors. Furthermore, as with selective AR and GR antagonists, inhibitors targeting ALK, BTK, and FLT3 kinases that have not been shown to interact with the ER had no effect on ER motility when assessed using SMT ( Figure 18A ).

[0510] For identified cellular pathway inhibitors, dose titration was used to better characterize the effect of each inhibitor on ER motility. Potencies ranged from subnanomolar to low micromolar ( Figure 18B ), similar to the reported potency of these compounds against their cellular targets. In addition, these molecules were also tested against AR ( Figure 24C ) and PR( Figure 24D ) were tested. Each SHR was significantly different from the other SHRs in terms of the responses of compounds identified through ER-focused screening efforts. Similarly, the magnitude of the ER SMT effect was generally consistent within the target class ( Figure 18B 、 Figures 24A-24D The finding that structurally distinct compounds exhibit similar effects based on their biological targets supports the idea that these biological targets themselves must interact with the ER, and therefore that these compounds indirectly affect ER motility. HSP90 is a chaperone for many proteins, including SHR. In the canonical model, hormone binding releases the SHR-HSP90 complex. Indeed, HSP90 inhibitors increase the f 结合 This is consistent with the fact that one of the functions of molecular chaperones is to regulate the balance of SHR binding to chromatin ( Figures 24B-24D ). Proteasome inhibition also resulted in ER pinning to chromatin, consistent with results obtained in htSMT screens of bioactive compounds. ER has been shown to be phosphorylated by CDKs, Src, or GSK-3 via the MAPK and PI3K / AKT signaling pathways, and therefore inhibition of these pathways would reasonably be expected to affect ER motility measured using SMT. While CDK inhibition resulted in increased ER mobility, inhibition of PI3K, AKT, or other upstream kinases showed no effect ( Figure 18A ).

[0511] Interestingly, SMT motility in the ER triple-point mutant was impaired by CDK and mTOR pathway inhibitors ( Figure 25 ), the mutant was engineered to lack previously defined phosphorylation sites (S104A / S106A / S118A) important for transcriptional activation, suggesting that additional phosphorylation sites may mediate the effects of CDK9 and PI3K / AKT signaling, or that other molecular targets of CDKs and PI3K / AKT may act indirectly to alter ER motility. For inhibitors of well-characterized pathways, such as those targeting CDKs and mTOR, changes in ER protein motility were subtle but consistent across compounds, indicating that these observations are biologically meaningful and highlighting the need for accurate and precise SMT measurements. Therefore, the htSMT technology described herein provides a way to provide comprehensive pathway interaction information.

[0512] Since SMT can identify compounds that act directly on the target or through some intermediate process, a strategy was adopted to distinguish between these alternative modes of action. For example, SMT can be used to distinguish direct from indirect effects on ER activity by studying the rate at which changes in protein motion occur. Given the live cell setup of SMT, a data collection mode was configured that allows the measurement of protein motion at set time intervals after compound addition (kinetic SMT or kSMT). When measured as kSMT, both ER agonists and antagonists rapidly induced the immobilization of ER on chromatin (t of estradiol). 1 / 2 =1.6 minutes; Figure 16C On the other hand, HSP90 inhibitors such as gantapi and HSP990 showed a 5- to 7-minute delay before changes in ER motility were observed, after which f 结合 The increase of t 1 / 2 The overall effect of these compounds reached a plateau after one hour ( Figure 18C Proteasome inhibitors, such as bortezomib and carfilzomib, act more slowly, with changes in ER motility appearing only after 40 minutes and increasing slowly over the four-hour measurement window ( Figure 18D ). Exploration of SMT dynamics thus represents an important tool to facilitate the differentiation of on-target modulators from on-pathway modulators. For example, the kSMT technology described herein allows for rapid mechanistic characterization of active compounds in a drug discovery setting.

[0513] To further differentiate the effects of pathway inhibitors on ER protein motility, the relative ER residence time of each such molecule was characterized. Estradiol, SERMs, and SERDs all increased residence time and, therefore, may also increase the rate of ER association with chromatin ( Figure 15 、 Figure 21E ). In contrast, although HSP990 and inhibition of HSP90 by gantapi resulted in f 结合 However, the number of total length binding events was observed to decrease by two-fold and four-fold, respectively, while the binding time was similar to that observed with DMSO alone ( Figure 18E These results suggest that HSP90 inhibition primarily increases k* on , and k* off The ER-chromatin binding pathway was largely unaffected. On the other hand, proteasome inhibition resulted in an increase in both the number and duration of long binding events. These results suggest that ER-chromatin binding can be regulated by altering the rates of association or dissociation, and that inhibition of specific cellular partners can differentially affect these rates. Combined with the distinct kinetics of direct ER, HSP90, and proteasome modulators, these data suggest that each class of molecules alters ER motility through distinct mechanisms.

[0514] C. Method

[0515] a. Cell lines

[0516] U2OS (ATCC catalog no. HTB-96), MCF7 (ATCC catalog no. HTB-22), T47d (ATCC catalog no. HTB-133), and SK-BR-3 (ATCC catalog no. HTB-30) were grown in DMEM (catalog no. 1056601, GibcoDMEM, high glucose, GlutaMAX supplement, Thermofisher) supplemented with 10% fetal bovine serum (catalog no. 16000044, Thermofisher) and 1% penicillin-streptomycin (catalog no. 15140122, Thermo Fisher) and maintained in a humidified 37°C incubator with 5% CO2, and subcultured approximately every two to three days.

[0517] b. HaloTag-expressing cell lines

[0518] For ER, AR, and PR-HaloTag fusions, mammalian expression vectors containing the fusion gene under the control of the weak L30 promoter and a neomycin resistance marker were transfected into 70% confluent U2OS cells using FuGENE 6 (Cat. No. E2691, Promega). Transfected cells were selected with 500 μg / mL of G418 (Cat. No. 10131027, Thermo Fisher) and then cloned. The cells were cloned by first incubating with 100 nM JF 549 -HTL (Cat. No. GA1110, Promega) and 50 nM Hoechst 33342 staining, and identification of cells with expected JF 549 The clones expressing the desired fusion gene were identified by comparing the signal distribution of clones. Three to six clones were then tested for response to control compounds using SMT conditions, and the most homogeneous clones were subsequently expanded for further testing. Unless otherwise stated, all experiments in the group used a single clone-isolated cell line. Since U2OS cells endogenously express GR, CRISPR / Cas9 was used to insert the HaloTag before the stop codon of endogenous NR3C1 by homology-directed repair. By using HTL-JF 646 HaloTag knock-in was verified by staining imaging and DNA sequencing.

[0519] c. Western blotting

[0520] Cells were grown under the same conditions as previously described. 1.5×10 6 cells per well were seeded in DMEM overnight in a 6-well plate and then treated with compound (DMSO or 100 nM fulvestrant) the next day for 24 hours. The cells were then lysed in 200 μL 1X cell lysis buffer (catalog number 9803, Cell Signaling). The BCA protein assay kit (catalog number 23225, Pierce TM Protein lysate concentrations were determined using a BCA protein assay kit. Capillary Western immunoassays were performed using Jess Protein Simple following the manufacturer's instructions (protein simple, USA). αER levels (1:100, RM-9101) were normalized to the loading control β-tubulin (1:100, NC0244815 LI-COR 92642213, Thermo Fisher). Peak values ​​were analyzed using Compass software (ProteinSimple, USA).

[0521] d. RNA sequencing

[0522] Cells were seeded into 12-well tissue culture plates at a density of 250,000 cells per well (U2OS-WT), 200,000 cells per well (U2OS-ER), or 300,000 cells per well (MCF7, SK-BR-3, T47d). 24 hours later, cells were treated with estradiol at a final concentration of 25 nM at the designated time points (0, 10, 60, or 3 hours). To process cells for total RNA, cells were washed twice with ice-cold PBS, lysed with 350 μL of buffer RLT (Qiagen 79216), scraped from plates (Fisher 08100241), frozen on dry ice, and stored at -20°C. The cell lysate was then thawed, homogenized using a QIAshredder column (Qiagen 79656), and processed using a standard protocol using a Qiagen RNeasy Micro kit (Qiagen 74004), including an optional on-column DNase digestion step (Qiagen 79254). All samples had a RIN score of 10 according to a TapeStation (Agilent 5067-5576). RNA sequencing libraries were prepared from total RNA by Novogene (CA). Briefly, mRNA was purified from total RNA using magnetic beads linked to poly-T oligonucleotides and fragmented. First-strand synthesis was performed using random hexamer primers, second-strand synthesis was performed using dTTP, and libraries were prepared after end repair, A-tailing, adapter ligation, amplification, and purification. The libraries were sequenced on an Illumina NovaSeq with paired 150-cycle reads. For data analysis, paired-end reads were aligned to the hg38 reference genome using Hisat2 v2.0.5, the number of reads mapping to each gene was counted using featureCounts v1.5.0-p3, and differential expression analysis was performed using DESeq2 (1.20.0).

[0523] e. Single-molecule tracking sample preparation

[0524] Cells were seeded at 6000 cells per well in 384-well tissue culture treated glass bottom plates. The seeded cells were then incubated overnight at 37°C and 5% CO2 to allow them to adhere. For all SMT experiments, cells were incubated with 5-100 pM JF 549-HTL (catalog number GA1110, Promega) and 50nM Hoechst 33342 were cultivated together in complete medium for one hour. The cells were then washed three times in DPBS and washed twice in imaging medium, which was supplemented with GlutaMAX (catalog number 35050079, Thermo Fisher) and the same serum and antibiotics as the growth medium, fluoroBrite DMEM culture medium (catalog number A1896701, Thermo Fisher). Where appropriate, the compound was serially diluted in EchoQualified 384-well low dead volume source microplates (0018544, Beckman Coulter) to generate dose titration source material. The compound was applied in cell culture medium at a final dilution of 1:1000. Each dose of compound had at least 2 replicates and 3 replicates per plate, with 20 DMSO control wells and 2 dye-free control wells randomly assigned on each plate. Unless otherwise stated, the compound was cultivated for one hour at 37°C before collecting the image.

[0525] f. Image acquisition

[0526] Unless otherwise stated, all images acquired using SMT were acquired on a custom HILO microscope based on a Nikon Ti2, a motorized stage, a stage-top environmental chamber (OKO Laboratories), a quad-band filter (Chroma), and custom laser emitters with wavelengths of 405 nm and 561 nm, delivering >10 mW and >150 mW to the back focal plane of the objective, respectively. Fluorescence emission was collected through a high-speed filter wheel (Finger Lakes Instruments) with a backlit CMOS camera (Prime 95b, Teledyne). Images were acquired using a 60X 1.27NA water immersion objective (Nikon). The environmental chamber was set to 37°C, 95% humidity, and 5% CO2. For each field of view, 200 SMT frames were collected at a frame rate of 100 Hz using a 2-millisecond stroboscopic laser pulse. Ten Hoechst channel frames were collected at the same frame rate for downstream registration of trajectories with the cell nucleus.

[0527] g. Image analysis

[0528] Image acquisition generates a JF for each field of view 549 Film and a Hoechst. JF 549 Movie used to track a single JF 549The motion of molecules is captured, while Hoechst movies are used for nuclear segmentation. Tracking is performed in three consecutive steps using a combination of existing methods: detection, sub-pixel localization, and linking. Briefly, points are detected using a generalized log-likelihood ratio detector. After detection, the estimated position of each emitter is refined to sub-pixel resolution using a Levenberg-Marquardt fit and an integrated 2D Gaussian point model starting from an initial guess provided by a radially symmetric method. Detected points are linked into trajectories using a custom modification of the hill climbing algorithm. All movies used in this manuscript use the same detection, sub-pixel localization, and linking settings.

[0529] For nuclear segmentation, all frames of the Hoechst movie were averaged to generate a mean projection. This mean projection was then segmented using a neural network trained on human labeled nuclei. Each point was assigned to at most one nucleus using its subpixel coordinates.

[0530] To recover motion information from the trajectories, we used state arrays, a Bayesian inference method with the “RBME” likelihood function and a 0.01 to 100.0 μm 2 s -1 A grid of 100 diffusion coefficients and 31 localization error amplitudes ranging from 0.02 to 0.08 μm was used. After inference, the localization errors were marginalized, resulting in a one-dimensional distribution of diffusion coefficients for each field of view. For single-cell analysis, SMT and nuclear segmentation were performed on a mixture of U2OS cells with H2B-HaloTag, HaloTag-CaaX or free HaloTag. The marginal likelihood of a set of 100 diffusion coefficients on the set of trajectories within each segmented nucleus was evaluated. These marginal likelihood functions were clustered using k-means (3 clusters) and the marginal likelihood functions of each cell were sorted by cluster index to produce a heat map. In order to estimate the binding fraction (f 结合 ), for less than 0.1μm 2 s-1 In order to estimate the free diffusion coefficient (D 自由 ), calculate 0.1μm 2 s -1 The mean of the posterior distribution above.

[0531] h. Data Analysis

[0532] Use KNIME or Spotfire (TIBCO) to analyze the trace results of the automated processing pipeline. 结合 or D 自由 Measurements are associated with experiment metadata and aggregated conditionally. 结合 The change in f is calculated for each well. 结合The median f of DMSO in the same plate 结合 The difference between the two. Wells with no cells in the field of view or out of focus were omitted from further analysis. The median fluorescence intensity of the tracking channel was used to assess the assay interference of the compound and was omitted if the median fluorescence intensity was more than 3 standard deviations higher than the median intensity of the DMSO wells. Similarly, if the active and negative controls could not be clearly distinguished or deviated significantly from the performance of the rest of the screen, the plate was removed from further analysis. Finally, compounds with a variance more than three standard deviations higher than the mean compound variance (41 compounds, 0.08%) were removed from downstream analysis. The Z' factor between the active control and DMSO on the plate was calculated as previously described. EC50 values ​​were calculated in Prism (GraphPad) by first logarithmically transforming the molecular concentrations and then fitting to a four-parameter logistic curve.

[0533] i. Active molecule clustering

[0534] The molecules identified as active (239 in total) were clustered based on chemical structure. The molecular frameworks were calculated as described by Murcko et al. and implemented in Pipeline Pilot. The molecular frameworks were clustered using functional class fingerprints (FCFP_4) with a similarity threshold cutoff of 0.3 Tanimoto distance. A total of 21 clusters were obtained, of which a single cluster was the main class (124 molecules). The second largest group was the flavonoids, with 27 members, followed by several different classes within the steroid class, with 14 and 20 members, respectively. Another class was the stilbenes, with 7 members, one of which was tamoxifen. The remaining active ingredients (47 molecules) were divided into a 3-member cluster, and all other active ingredients had 2 members per cluster.

[0535] j. Kinetic experiments

[0536] The day before, cells were seeded into 384-well plates, stained, and washed as described above. One well with 25 FOVs per well was used as a baseline reading. Compounds were then manually added to each well while imaging to a final concentration of 100 nM. Data were then collected for 20 wells. Pauses were included between each FOV so that the entire imaging protocol covered the detection window. The f of each well was determined relative to t = 0. 结合 change.

[0537] For assays up to 4 hours long, the plate was imaged twice, with 8 FOVs per well, and the FOV position was varied between reads to prevent photobleaching from affecting the data. All data presented were performed in three different biological replicates.

[0538] k. Dwell time imaging

[0539] Sample preparation and dwell time imaging experiments were performed in a similar manner to the single molecule tracking assays described above, with a few exceptions. Samples were treated with 1-10 pM JF 549 (Promega) and stained with 50 nM Hoechst 33342 for one hour. 400 frames were collected per field of view, with the camera integration time set to 250 milliseconds and the laser source at the objective lens reduced to 5 mW. The laser was continuously on during image acquisition. Compound incubation times ranged from 1 to 4 hours. At least eight replicate wells were collected for each condition.

[0540] 1. Residence time analysis

[0541] Image processing, including spot detection, localization, and trajectory reconnection, was performed using the same methods described above. Because dwell time imaging selectively tracks slowly diffusing molecules, individual localizations were limited to a maximum displacement of 300 nm for single-jump reconnection. The trajectory set for each field of view was divided into a 1-CDF distribution and fitted with a biexponential decay model CDF(t) = A(F ek) 快 t+(1-F)ek 慢 t)CDFt=A(Fe-k fast t+1-Fe-k slow t).

[0542] m. Fluorescence recovery after photobleaching

[0543] Images were collected on a custom HiLo microscope using a Spectra Light Engine RS-232 as described above. Stimulation was guided using a microscanner coupled to a Coherent OBIS 561 nm 100 mW laser. All imaging was performed using a 60×1.27 NA water immersion objective (Nikon). All experiments were performed at 37°C. For FRAP experiments, cells were seeded into 384-well plates the day before and stained with 50 nM HTL-JF. 549 Marked, and washed as described above. It is 100nM to add the compound to a final concentration of 100nM one hour before imaging. Then, the image before bleaching is collected by averaging 10 continuous images. 8-10 areas (2 backgrounds, 6-8 cells) are bleached, and 2 areas in the cell are not bleached. The bleached area is bleached with 10% power without scanning. In the next 30 seconds, an image is collected every 200 milliseconds, followed by an image collection every 1 second for 2 minutes. The background subtraction average intensity that varies over time in the region of interest is measured and standardized to the mean value of the fluorescence in the baseline image, then standardized to the unbleached area to explain the photobleaching of the fluorophore caused by the readout. For three biological experiments, data from 18-24 cells were collected in each experiment.

[0544] immunofluorescence

[0545] Cells were grown under the conditions described previously. Halo-ER U2OS cells were seeded at 6,000 cells per well and MCF7 and T47d cells were seeded at 8,000 cells per well in glass-bottom 384-well plates coated with 0.05 mg / ml PDL (Cat. No. A3890401, Thermofisher). Cells were grown overnight and then treated with compounds the next day at 37°C and 5% CO2 for 24 hours. Serial dilutions were made in qualified 384-well low dead volume source microplates (0018544, Beckman Coulter), starting at a concentration of 10 mM, with a 1:3 dilution to generate a 21-point dose response. Compounds were applied in cell culture medium at a final dilution of 1:1000. Based on the potency of each compound, 8 to 12-point dose responses were selected. Each concentration was repeated at least once in each plate, and at least 2 plates were repeated. Cells were fixed for 20 minutes by adding paraformaldehyde (catalog number 15710-S; Electron Microscopy Sciences) at a final concentration of 4%. The cells were then permeabilized for one hour at room temperature using 1× PBS blocking buffer containing 1% bovine serum albumin and 0.3% Triton-X100. Immunofluorescence staining of the ER was performed using an αER antibody (1:500, RM-9101) diluted in the same blocking buffer for one hour at room temperature. Wash thoroughly with PBS before performing secondary antibody staining. Secondary antibody staining was performed for one hour using Alexa fluor 488-conjugated anti-rabbit IgG (1:1000, catalog number A32731, Thermos Fisher). Nuclear staining was performed using a 1 mg / ml solution of Hoechst 33342. Immunofluorescence imaging was performed using an ImageXpress Micro (Molecular Devices) at 10× magnification and 4 fields per well. Fluorescence intensity within the cell nucleus was quantified using CellProfiler. All analyses and curve fitting were performed using Prism with DMSO as the baseline.

[0546] o. Cell proliferation

[0547] Cells were grown and seeded under the conditions described above. Cells were seeded in 384-well plates (Cat. No. 353963, Corning) at 1000 cells per well for Halo-ER U2OS, 1200 cells per well for SK-BR-3, and 1800 cells per well for MCF7 and T47d. Cells were grown overnight and then treated with compounds the next day. Compound concentrations and administration were the same as those described in the previous immunofluorescence assay. Phase contrast was used and the plates were scanned every 24 hours in an IncuCyte live cell analysis system (Sartorius) for a total of 5 days. Cell proliferation was quantified using a full-well confluence mask using the built-in analysis function. All analyses and curve fitting were performed using Prism with DMSO as the baseline.

[0548] Example 2: htSMT analysis of AR agonists and antagonists

[0549] Using the methods described in Example 1, for example, for preparing and analyzing cell lines expressing AR as a fluorescent target protein, this example provides additional evidence that changes in protein interactions, such as changes in protein binding, as measured by changes in target motility, can support the identification of pharmacologically relevant compounds. Specifically, known agonists and antagonists of AR were assayed as described in Example 1, except that the f of AR was measured. 结合 In the presence of 25 nM agonist, 10 μM potent antagonist or a combination of agonist and antagonist at 25 nM and 10 μM respectively.

[0550] Although AR agonists were observed to increase f 结合 , but it was observed that AR antagonists, whether in monotherapy or when co-administered with AR agonists, could increase f 结合 reduce( Figure 26 ). Therefore, this embodiment clearly shows that f 结合 The increase and f 结合 Any reduction in activity helps identify mechanistically distinct mechanisms of pharmacological interaction with the observed fluorescent target protein.

[0551] Example 3: htSMT analysis of target A antagonists

[0552] Using the methods described in Example 1, for example, for preparing and analyzing cell lines expressing Target A as a fluorescent target protein, this example provides additional evidence that changes in protein interactions, such as changes in protein-protein interactions in signaling pathways unrelated to ER signaling described in Example 1, can support the identification of pharmacologically relevant compounds. Specifically, Figure 27Depicts changes in target A motility in response to dose titration of a known target A antagonist. The shape of each series represents a different compound, and the error bars represent the SEM. The curve shapes demonstrate that increasing antagonist concentrations result in a measurable difference in the median third quartile (Q3) jump length relative to DMSO, indicating that target A is released from the bound state in the presence of the antagonist.

[0553] Example 4: htSMT analysis of competitive and allosteric antagonists

[0554] Using the methods described in Example 1, for example, for preparing and analyzing cell lines expressing exemplary receptor tyrosine kinases (target B and target C) and helicases as fluorescent target proteins, this example provides additional evidence that compounds that affect protein interactions, such as changes in protein-protein interactions in signaling pathways via competitive or allosteric inhibition, can support the identification of pharmacologically relevant compounds. Specifically, Figure 28 Depicts the changes in target B and target C motion in the presence of competitive or allosteric antagonists. The shape of each series represents a different compound normalized to DMSO, and the error bars represent the SEM. The curve shapes demonstrate that increasing antagonist concentrations result in measurable differences in median third quartile (Q3) jump length relative to DMSO, indicating that targets B and C are differentially affected by competitive and allosteric antagonists.

[0555] Similarly, Figure 29 The change in median Q3 jump length relative to DMSO after compound addition is depicted for target A, target B, and the helicase target over time. Depending on the data presented, treatment of the protein with an on-target inhibitor increases or decreases protein motion. For targets A and B, these changes occur within minutes of compound addition. In contrast, for helicases treated with pathway antagonists or off-target modulators (such as those induced by DNA damage), changes in protein motion take several hours or longer to reach maximum effect.

[0556] Depending on the desired configuration, the subject matter described herein may be embodied in systems, devices, methods and / or articles. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the subject matter. For example, the implementations described above may be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several further features disclosed above. In addition, the logic flows depicted in the accompanying drawings and / or described herein do not necessarily require the specific order shown or the sequential order to achieve the desired results. Other implementations may be within the scope of the claims below.

Claims

1. A method comprising: receiving a sequence of images visualizing molecular motion; linking molecules across said images; Using a variational Bayesian optimization algorithm and based on the linkage, generating possible trajectories for each molecule and their associated probabilities; as well as Data representing the possible trajectories generated and their associated probabilities is provided to a consuming application or process.

2. A method comprising: receiving a sequence of images visualizing molecular motion; linking molecules across said images; Using a Gibbs sampling algorithm and based on the linkage, generating possible trajectories for each molecule and their associated probabilities; as well as Data representing the possible trajectories generated and their associated probabilities is provided to a consuming application or process.

3. A method comprising: receiving a sequence of images visualizing molecular motion; linking molecules across said images; generating possible trajectories for each molecule and their associated probabilities based on the links using an adaptive hill climbing algorithm; as well as Data representing the possible trajectories generated and their associated probabilities is provided to a consuming application or process.

4. A method as claimed in any preceding claim, wherein at least a subset of the sequence of images comprises at least 100 molecules per image.

5. A method as claimed in any preceding claim, wherein at least a subset of the sequence of images comprises at least 1000 molecules per image.

6. A method as claimed in any preceding claim, wherein at least a subset of the sequence of images comprises at least 10,000 molecules per image.

7. A method as claimed in any preceding claim, wherein the molecules have a density of at least 0.01 emitters per square micron per image.

8. A method as claimed in any preceding claim, wherein the molecules have a density of at least 0.1 emitters per square micron per image.

9. The method of any one of the preceding claims, further comprising: labeling molecules within biological samples; causing the biological sample to emit fluorescence; as well as The sequence of images is generated while causing the biological sample to fluoresce.

10. The method of claim 9, wherein the generating of the sequence of images is performed using a microscope system.

11. The method of any preceding claim, wherein the molecule is imaged within a living cell.

12. The method of any one of the preceding claims, further comprising: A probabilistic dynamic model is inferred that contains information characterizing the molecular trajectory.

13. The method of claim 12, wherein the probabilistic dynamic model comprises a state array, and the method further comprises: The state array is populated with the information characterizing the molecular trajectory.

14. The method of any one of the preceding claims, further comprising: An internal confidence indicator based on the associated probability is generated, wherein the provided data includes the generated internal confidence indicator.

15. The method of claim 14, wherein the generated internal confidence indicator is a tracking error rate lower bound, which defines a lower bound on the rate of incorrect connections caused by the link.

16. The method of claim 14, wherein the generated internal indicators include: Compute the confidence score for each trajectory.

17. The method of any one of the preceding claims, further comprising: Generate dynamic metrics independent of specific trajectories.

18. The method of any of the preceding claims, wherein the linking comprises retrieving data comprising a plurality of statistical data extracted from the total number of detections or the number of detections in a cell.

19. The method of any of the preceding claims, wherein providing data comprises one or more of: visualizing at least a portion of the generated possible trajectories and their associated probabilities in a graphical user interface, storing at least a portion of the generated possible trajectories and their associated probabilities in physical persistence, loading at least a portion of the generated possible trajectories and their associated probabilities into a memory, or transmitting at least a portion of the generated possible trajectories and their associated probabilities to a remote computing device over a network.

20. A method as claimed in any preceding claim, wherein at least part of the sequence of images comprises consecutive images from a respective film.

21. A method as claimed in any preceding claim, wherein at least part of the sequence of images used for the linking are non-consecutive images from a respective film.

22. A method for single molecule tracking, comprising: receiving a sequence of images visualizing molecular motion, the sequence of images comprising a first type generated using a first imaging modality and a second type generated using a second, different imaging modality; detecting points within a sequence of images of said first type; linking the points detected within the first type of image sequence into trajectories using a probabilistic tracking algorithm; segmenting the second type of image sequence to generate a plurality of instance masks; assigning molecules within the image sequence of the second type to at least one instance mask of the plurality of instance masks; as well as Data representing the link and assignment is provided to a consuming application or process.

23. The method of claim 22, wherein the probabilistic tracking algorithm comprises a variational Bayesian optimization algorithm.

24. The method of claim 22, wherein the probabilistic tracking algorithm comprises a Gibbs sampling algorithm.

25. The method of claim 22, wherein the partial probabilistic tracking algorithm comprises an adaptive hill climbing algorithm.

26. The method of any one of claims 22 to 25, wherein the first imaging modality and the second imaging modality comprise different molecular labeling technologies.

27. The method of any one of claims 22 to 26, wherein the first type of image sequence is a single molecule tracking (SMT) movie and the second type of image sequence is a non-SMT movie.

28. The method of any one of claims 22 to 27, wherein the detected sites include subcellular components.

29. The method of any one of claims 22 to 28, wherein the molecular types within the sequence of images of the first type are labeled with different fluorophores.

30. The method of any one of claims 22 to 29, wherein at least a subset of the sequence of images comprises at least 100 molecules per image.

31. A method as claimed in any one of claims 22 to 30, wherein at least a subset of the sequence of images comprises at least 1000 molecules per image.

32. The method of any one of claims 22 to 31 , wherein at least a subset of the sequence of images comprises at least 10,000 molecules per image.

33. The method of any one of claims 22 to 32, wherein the molecules have a density of at least 0.01 emitters per square micron per image.

34. The method of any one of claims 22 to 33, wherein the molecules have a density of at least 0.1 emitters per square micron per image.

35. The method of any one of claims 22 to 34, further comprising: labeling molecules within biological samples; causing the biological sample to emit fluorescence; as well as At least a portion of the sequence of images is generated while causing the biological sample to fluoresce.

36. The method of claim 35, wherein generating the sequence of images is performed using a microscope system.

37. The method of any one of claims 22 to 36, wherein the molecule is imaged within a living cell.

38. The method of any one of claims 22 to 37, further comprising: A probabilistic dynamic model is inferred that contains information characterizing the molecular trajectory.

39. The method of claim 38, wherein the probabilistic dynamic model comprises an array of states, and the method further comprises: The state array is populated with information characterizing the molecular trajectory.

40. The method of any one of claims 22 to 39, further comprising: An internal confidence indicator based on the associated probability is generated, wherein the provided data includes the generated internal confidence indicator.

41. The method of claim 40, wherein the generated internal confidence metric is a tracking error rate lower bound (ERLB) that defines a lower bound on the rate of false connections caused by the link.

42. The method of claim 40, wherein the generated internal indicators include: Compute the confidence score for each trajectory.

43. The method of any one of claims 22 to 42, further comprising: Generate dynamic metrics independent of specific trajectories.

44. The method of any one of claims 22 to 43, wherein the linking comprises retrieving data comprising a plurality of statistics extracted from the total number of detections or the number of detections in a cell.

45. The method of any one of claims 22 to 44, wherein providing data comprises one or more of: visualizing at least a portion of the generated possible trajectories and their associated probabilities in a graphical user interface, storing at least a portion of the generated possible trajectories and their associated probabilities in physical persistence, loading at least a portion of the generated possible trajectories and their associated probabilities into a memory, or transmitting at least a portion of the generated possible trajectories and their associated probabilities to a remote computing device over a network.

46. ​​The method of any one of claims 22 to 45, further comprising: A plurality of statistical metrics associated with the at least one track or the at least one instance mask are generated.

47. The method of claim 46, further comprising: A hierarchy storing instance masks.

48. A method as claimed in any one of claims 22 to 47, wherein at least a portion of the sequence of images comprises consecutive images from a respective film.

49. A method as claimed in any one of claims 22 to 48, wherein at least a portion of the sequence of images used for the linking are non-consecutive images from a respective film.

50. The method of any one of claims 22 to 49, wherein the detecting utilizes one or more of: Generalized log-likelihood point detector, Difference of Gaussian (DoG) detector, Laplace of Gaussian (LoG) detector, or Determinant of Hessian (DoH) blob detector.

51. The method of any one of claims 22 to 50, further comprising: Detected points are associated with spatiotemporal coordinates using sub-pixel localization.

52. The method of claim 51 , wherein the sub-pixel positioning comprises one or more of: Radially symmetric locators or maximum likelihood fitting of candidate point models using the Levenberg-Marquardt method.

53. A method as claimed in any one of claims 22 to 52, wherein the field of view corresponds to at least a portion of a hole.

54. The method of any of the preceding claims, wherein the sequence of images is generated by an apparatus for fluorescence microscopy comprising: a light source capable of emitting fluorescent excitation light, wherein the light source exhibits a power output drift of less than about 10% at an ambient temperature of 17°C + / - 5°C; a first optical element or assembly configured to receive a fluorescent excitation light source and shape the fluorescent excitation light source to form a light beam; a second optical element or assembly comprising a water immersion objective configured to tilt the light beam in an xz plane relative to the z-axis, wherein the second optical element is further configured to focus the light beam at a sample plane located in the xy plane, thereby illuminating at least a portion of the sample plane; as well as A detector arrangement is configured to receive light from the illuminated portion of the sample plane, wherein the detector arrangement forms one or more projected images based on the light received from the illuminated portion of the sample plane.

55. The method of claim 54, wherein the apparatus comprises a second objective lens configured to direct light emitted from the illuminated portion of the sample plane to the detector arrangement.

56. A method as claimed in claim 54 or 55, wherein the detector means comprises a semiconductor sensor.

57. A method as claimed in any one of claims 54 to 56, wherein the apparatus comprises a third optical element or component configured to translate the light beam in an imaging plane in a direction orthogonal to the longer dimension of the light beam.

58. The method of claim 57, wherein the third optical element or component comprises a galvanometer mirror.

59. The method of any one of claims 54 to 58, wherein the detector arrangement comprises a semiconductor sensor, wherein the detector arrangement supports a shutter mode for synchronizing the translation of the light beam in the sample plane with the selective activation or readout of the semiconductor sensor.

60. The method of any one of claims 1 to 53, wherein the sequence of images is generated by a microscopy system for tracking molecular motion, the microscopy system comprising: a stage for supporting a sample, wherein the sample contains the molecule; a light source for emitting a light beam capable of inducing a light-based reaction from the molecules in the sample, wherein the light source exhibits a power output drift of less than about 10% at an ambient temperature of 17°C + / - 5°C; a water immersion objective for focusing the light beam onto at least a portion of the sample plane, wherein the molecules are disposed within the sample plane; and A detector device is used to monitor the light-based reaction of the molecule, analyze the light-based reaction, and thus track the movement of the molecule.

61. A method as claimed in claim 60, wherein the microscope system further includes a scanning optical element or assembly configured to translate the light beam in the sample plane in a direction orthogonal to the longer dimension of the light beam, thereby enabling the microscope system to have a larger total field of view in the xy plane.

62. The method of claim 61, wherein the microscope system further comprises a z position controller for the sample plane, wherein the z position controller is capable of maintaining focus in the z direction.

63. The method of any one of claims 60 to 62, wherein the sample is disposed within an open well of the sample plate.

64. The method of any one of claims 63, wherein the sample plate comprises a plurality of open wells.

65. The method of claim 63 or claim 64, wherein the microscope system further comprises an xy position controller for changing the field of view of the microscope system to encompass a different subset of the plurality of open apertures.

66. The method of any one of claims 63 to 65, wherein the microscope system further comprises a temperature controlled environment configured to control the environment of the sample plate.

67. The method of claim 66, wherein the samples disposed within the open wells of the sample plate are maintained at a humidity of 20% to 95%.

68. The method of claim 66 or claim 67, wherein the sample disposed within the open well of the sample plate is maintained at 5% CO2.

69. The method of any one of claims 60 to 68, wherein the microscope system further comprises an automated sample handling robotic system to enable high throughput manipulation of multiple samples on a stage, wherein the robotic system comprises: Memory; a processor in communication with the memory; as well as One or more robotic end effectors in communication with the processor, wherein the one or more end effectors manipulate the plurality of samples on the stage based on communication with the processor.

70. A system comprising: at least one data processor; and A memory storing instructions which, when executed by at least one data processor, result in the operations of the method as claimed in any one of claims 1 to 53.

71. The system of claim 70, further comprising: An apparatus for fluorescence microscopy, comprising: a light source capable of emitting fluorescent excitation light, wherein the light source exhibits a power output drift of less than about 10% at an ambient temperature of 17°C + / - 5°C; a first optical element or assembly configured to receive a fluorescent excitation light source and shape the fluorescent excitation light source to form a light beam; a second optical element or assembly comprising a water immersion objective configured to tilt the light beam in an xz plane relative to the z-axis, wherein the second optical element is further configured to focus the light beam at a sample plane located in the xy plane, thereby illuminating at least a portion of the sample plane; as well as A detector arrangement is configured to receive light from the illuminated portion of the sample plane, wherein the detector arrangement forms one or more projected images based on the light received from the illuminated portion of the sample plane.

72. The system of claim 71, wherein the apparatus comprises a second objective lens configured to direct light emitted from the illuminated portion of the sample plane to the detector arrangement.

73. The system of claim 71 or claim 72, wherein the detector device comprises a semiconductor sensor.

74. The system of any one of claims 71 to 73, wherein the device comprises a third optical element or component configured to translate the light beam in an imaging plane in a direction orthogonal to the longer dimension of the light beam.

75. The system of claim 74, wherein the third optical element or component comprises a galvanometer mirror.

76. A system as described in any one of claims 71 to 75, wherein the detector device includes a semiconductor sensor, wherein the detector device supports a shutter mode for synchronizing the translation of the light beam in the sample plane with the selective activation or readout of the semiconductor sensor.

77. The system of claim 70, further comprising: A microscope system for tracking molecular motion, comprising: a stage for supporting a sample, wherein the sample contains the molecule; a light source for emitting a light beam capable of inducing a light-based reaction from the molecules in the sample, wherein the light source exhibits a power output drift of less than about 10% at an ambient temperature of 17°C + / - 5°C; a water immersion objective for focusing the light beam onto at least a portion of the sample plane, wherein the molecules are disposed within the sample plane; and A detector device is used to monitor the light-based reaction of the molecule, analyze the light-based reaction, and thus track the movement of the molecule.

78. A system as described in claim 77, wherein the microscope system further includes a scanning optical element or assembly configured to translate the light beam in the sample plane in a direction orthogonal to the longer dimension of the light beam, thereby enabling the microscope system to have a larger total field of view in the xy plane.

79. The system of claim 78, wherein the microscope system further comprises a z position controller for the sample plane, wherein the z position controller is capable of maintaining focus in the z direction.

80. The system of any one of claims 77 to 79, wherein the sample is disposed within an open well of the sample plate.

81. The system of claim 80, wherein the sample plate comprises a plurality of open wells.

82. The system of claim 81, wherein the microscope system further comprises an xy position controller for changing the field of view of the microscope system to encompass a different subset of the plurality of open holes.

83. The system of any one of claims 81 to 82, wherein the microscope system further comprises a temperature controlled environment configured to control the environment of the sample plate.

84. The system of claim 83, wherein the samples disposed within the open wells of the sample plate are maintained at a humidity of 20%-95%.

85. The system of claim 83 or claim 84, wherein the sample disposed within the open well of the sample plate is maintained at 5% CO2.

86. The system of any one of claims 77 to 85, wherein the microscope system further comprises an automated sample handling robotic system to enable high throughput manipulation of multiple samples on a stage, wherein the robotic system comprises: a memory for storing instructions; at least one data processor; and One or more robotic end effectors in communication with the at least one data processor, wherein the one or more end effectors manipulate the plurality of samples on the stage based on communication with the at least one data processor.

87. A non-transitory computer program product storing instructions which, when executed by at least one data processor forming part of at least one computing device, implement the method of any one of claims 1 to 53.

88. A system comprising: a component for receiving a sequence of images visualizing molecular motion; a means for linking molecules across said images; means for generating possible trajectories for each molecule and their associated probabilities based on the links; as well as A component that provides data representing possible generated trajectories and their associated probabilities to a consuming application or process.

89. A system comprising: a component for receiving a sequence of images visualizing molecular motion; a means for linking molecules across said images; a means for generating possible trajectories for each molecule and their associated probabilities based on the linkages using a variational Bayesian optimization algorithm; as well as A component that provides data representing possible generated trajectories and their associated probabilities to a consuming application or process.

90. A system comprising: a component for receiving a sequence of images visualizing molecular motion; a means for linking molecules across said images; a means for generating possible trajectories of each molecule and their associated probabilities based on the linkage using a Gibbs sampling algorithm; as well as A component that provides data representing possible generated trajectories and their associated probabilities to a consuming application or process.

91. A system comprising: a component for receiving a sequence of images visualizing molecular motion; a means for linking molecules across said images; means for generating possible trajectories for each molecule and their associated probabilities based on the links using an adaptive hill climbing algorithm; as well as A component that provides data representing possible generated trajectories and their associated probabilities to a consuming application or process.

92. A single molecule tracking system comprising: means for receiving a sequence of images visualizing molecular motion, the sequence of images comprising a first type generated using a first imaging modality and a second type generated using a second, different imaging modality; means for detecting points within the sequence of images of said first type; for linking points detected within the sequence of images of the first type to members in a trajectory using a probabilistic tracking algorithm; means for segmenting the second type of image sequence to generate a plurality of instance masks; means for assigning molecules within the sequence of images of the second type to at least one instance mask of the plurality of instance masks; as well as Means for providing data representing the link and assignment to a consuming application or process.