Compositions and methods for 3-dimensional, in SITU positional mapping of nucleic acid targets

WO2026178395A1PCT designated stage Publication Date: 2026-08-27BRUKER SPATIAL GENOMICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/016084
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-07-24
Filing Date
2026-02-20
Publication Date
2026-08-27

Smart Images

  • Figure US2026016084_27082026_PF_FP_ABST
    Figure US2026016084_27082026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are methods and compositions comprising unique pools and sets of barcodes that can be used to target genomic nucleic acids for more accurate and efficient 3D mapping of nucleic acids in situ.
Need to check novelty before this filing date? Find Prior Art

Description

Atty. Dkt. No. 121892.00158COMPOSITIONS AND METHODS FOR 3-DIMENSIONAL, IN SITU POSITIONAL MAPPING OF NUCLEIC ACID TARGETSCROSS-REFRENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Appl. No. 63 / 761,025, filed February 20, 2025, U.S. Provisional Appl. No. 63 / 761,730, filed February 21, 2025, and U.S. Provisional Appl. No. 63 / 850,411, filed July 24, 2025. The content of each of the above-referenced applications is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0002] The present technology relates to methods of preparing pools and sets of barcodes, and the use of such barcode pools and sets for the in situ detection and 3-dimensional (3D) positional mapping of target nucleic acids, such as chromosomal DNA and / or transcripts thereof.BACKGROUND

[0003] In situ nucleic acid imaging typically involves cycles of nucleic acid hybridization and imaging using methods well-known in the art (e g., fluorescent in situ hybridization (FISH), multiplexed error-robust fluorescence in situ hybridization (MERFISH), seqFISH, RNA sequential probing of targets (SPOTs), high-coverage microscopy-based technology (Hi-M), optical reconstruction of chromatin architecture (ORCA)). Typically, for genomic imaging in situ, nucleic acid probes modified with fluorophores or other detectable moi eties are hybridized directly to chromosomal nucleic acid and visualized. In some methods, oligonucleotide probes are used, where each probe bears either one or more fluorescent moieties, and / or one or more sites for secondary hybridization by a fluorophore-bearing oligonucleotide. However, the multiplexity of these methods, e g. the number of distinct genomic loci able to be labeled and distinguished from other loci, is limited to either F*N using F spectrally distinct fluorescent moieties to label F genomic loci in each of N cycles of probe hybridization, or is bounded by Nxk using k combinations of fluorescent signals, each comprised of a specific number and / or combination of the F spectrally distinct fluorophores, referred to as “colorimetric” barcoding (e.g. red+blue as a distinct label, or 2* red vs 1 x red if various levels of red can be distinguished). Unfortunately, this number is still too low for accurate, higher resolution 3-dimensional studies of chromosome QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158structure and function. Accordingly, tools and methods for the multiplex detection of large numbers of genomic loci are needed.SUMMARY

[0004] The present disclosure provides methods and compositions comprising unique pools and sets of barcodes that can be used to target genomic nucleic acids for more accurate and efficient 3D mapping of nucleic acids in situ.

[0005] In some embodiments, methods for in situ positional mapping of a plurality of nucleic acid targets in a biological sample are provided. In some embodiments, the methods comprise detecting a pool of target barcodes linked to the plurality of nucleic acid targets, wherein each target barcode comprises N digits, wherein each position of each target barcode is represented by a color signal generated from a detectable label, and wherein no position of any barcode lacks a corresponding color signal, or wherein the sequence of digits within a barcode includes no more than 15% null digits; wherein detecting comprises: (a) labeling the plurality of nucleic acid targets with a first pool of first detection probes, wherein each first detection probe comprises a corresponding first detectable label, wherein different first detection probes are correlated to different targets or sets of targets of the plurality; (b) imaging the targets of (a), wherein each first detectable label provides the corresponding first color signal as digit 1 of the target barcodes; (c) removing the first color signal of (b); (d) labeling the plurality of nucleic acid targets with a second pool of second detection probes, wherein each second detection probe comprises a corresponding second detectable label, wherein different second detection probes are correlated to different targets or sets of targets of the plurality; (e) imaging the targets of (d) wherein each second detectable label provides the corresponding second color signal as digit 2 of the target barcodes; (f) removing the second color signal of (e); (g) repeating steps (d)-(f) with an Nth pool of detection probes, wherein each Nth detection probe in each Nth pool from the second pool to the Nth pool (i) comprises a corresponding detectable label, (ii) is correlated to different targets or sets of targets of the plurality; and (iii) provides a corresponding color signal as an Nth digit of the target barcode; wherein, in combination, the N pools of probes comprise at least 3 different colors.

[0006] In some embodiments, methods for in situ positional mapping of a plurality of nucleic acid targets in a biological sample are provided. In some embodiments, the methods comprise QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158detecting a pool of target barcodes linked to the plurality of nucleic acid targets, wherein each target barcode comprises N digits, wherein each position of each target barcode is represented by a color signal generated from a detectable label, and wherein no position of any barcode lacks a corresponding color signal, or wherein the sequence of digits within a barcode includes no more than 15% null digits; wherein detecting comprises: (a) labeling the barcodes linked to the plurality of nucleic acid targets with a first pool of first detection probes, wherein each first detection probe comprises a corresponding first detectable label and a barcode hybridization region, wherein different first detection probes are correlated to different targets or sets of targets of the plurality; (b) imaging the detectable labels of (a), wherein the first detection probes hybridizes to a barcode digit 1 position and provide the corresponding first color signal as digit 1 of the target barcode to which the first detection probe is hybridized; (c) removing the first color signal of (b); (d) labeling the barcodes linked to the plurality of nucleic acid targets with a second pool of second detection probes, wherein each second detection probe comprises a corresponding second detectable label and a barcode hybridization region, wherein different second detection probes are correlated to different targets or sets of targets of the plurality; (e) imaging the detectable labels of (d) wherein the second detection probes hybridize to a barcode digit 2 position and provide the corresponding second color signal as digit 2 of the target barcode to which the second detection probe is hybridized; (f) removing the second color signal of (e); (g) repeating steps (d)-(f) with an Nth pool of detection probes, wherein each Nth detection probe in each Nth pool from the second pool to the Nth pool (i) comprises a corresponding detectable label, (ii) is correlated to different targets or sets of targets of the plurality; and (iii) provides a corresponding color signal as an Nth digit of the target barcode; wherein, in combination, the N pools of probes comprise at least 3 different colors.

[0007] In some embodiments, methods for designing a set of molecular barcodes comprising colored fluorophores for detection of a plurality of nucleic acid targets in a biological sample are provided. In some embodiments, each position (digit) of each molecular barcode is represented by a color signal generated from a detectable label, and wherein no position (digit) of any barcode lacks a corresponding color signal. In some embodiments, the methods comprise: with the use of a programmable processor operably connected with tangible, non-transitory storage medium containing program code thereon, accessing data representing: (i) a color input value, wherein the color input value defines the number of different colors for use in the set of barcodes, wherein the color input value is at least 3; (ii)a barcode digit input value, wherein the barcode position QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158input value defines the number of digits of each barcode, wherein the position input value is N, and wherein each position of each barcode is represented by a color and no position of any barcode lacks a corresponding color signal or wherein the sequence of digits within a barcode includes no more than 15% null digits; (iii) a target input value, comprising the total number of targets (loci); (b) selecting, based on the accessing data of (a), the set of molecular barcodes, wherein: (i) the global hamming distance (gHD) for all barcodes within the set is N-2; and wherein the selected barcodes enable three-dimensional mapping of the plurality of nucleic acid targets.

[0008] In some embodiments, methods for in situ positional mapping of a plurality of nucleic acid targets in a biological sample are provided. In some embodiments, the methods comprise: detecting a pool of target barcodes linked to the plurality of nucleic acid targets, wherein each target barcode comprises N positions, and wherein detecting comprises: (a) labeling the plurality of nucleic acid targets with a first pool of first detection probes, wherein each first detection probe comprises a corresponding first detectable label, wherein different first detection probes are correlated to different targets or sets of targets of the plurality; (b) imaging the targets of (a), wherein each first detectable label provides a first color signal as position 1 of the target barcodes; (c) removing the first color signal of (b); (d) labeling the plurality of nucleic acid targets with a second pool of second detection probes, wherein each second detection probe comprises a corresponding second detectable label, wherein different second detection probes are corelated to different targets or sets of targets of the plurality; (e) imaging the targets of (d) wherein each second detectable label provides the corresponding second color as position 2 of the target barcodes; (f) removing the second color signal of (e); (g) repeating steps (d)-(f) with an Nth pool of detection probes, wherein each Nth detection probe in each pool from the second pool to the Nth pool (i) comprises a corresponding detectable label, (ii) is correlated to different targets or set of targets of the plurality; wherein a territory set of target barcodes comprise a trailing Hamming distance (trHD) value that is greater than a taper trail Hamming distance (ttHD) value of a subset of target barcode in the same territory set.

[0009] In embodiments, methods for in situ positional mapping of a plurality of nucleic acid targets in a biological sample are provided. In embodiments, the method comprises: detecting a pool of target barcodes linked to the plurality of nucleic acid targets, wherein each target barcode comprises N positions, and wherein detecting comprises: (a) labeling the barcodes QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158linked to the plurality of nucleic acid targets with a first pool of first detection probes, wherein each first detection probe comprises a corresponding first detectable label and a barcode hybridization region, wherein different first detection probes are correlated to different targets or sets of targets of the plurality; (b) imaging the detectable labels of (a), wherein the first detection probes hybridizes to a barcode digit 1 position and provide the corresponding first color signal as digit 1 of the target barcode to which the first detection probe is hybridized; (c) removing the first color signal of (b); (d) labeling the plurality of nucleic acid targets with a second pool of second detection probes, wherein each second detection probe comprises a corresponding second detectable label and a barcode hybridization region, wherein different second detection probes are corelated to different targets or sets of targets of the plurality; (e) imaging the detectable labels of (d) wherein the second detection probes hybridize to a barcode digit 2 position and provide the corresponding second color as position 2 of the target barcode to which the second detection probe is hybridized; (f) removing the second color signal of (e); (g) repeating steps (d)-(f) with an Nth pool of detection probes, wherein each Nth detection probe in each pool from the second pool to the Nth pool (i) comprises a corresponding detectable label, (ii) is correlated to different targets or set of targets of the plurality; wherein a territory set of target barcodes comprise a trailing Hamming distance (trHD) value that is greater than a taper trail Hamming distance (ttHD) value of a subset of target barcode in the same territory set.

[0010] In embodiments, computer program products for devising a set of barcodes configured to detect nucleic acid targets in a biological sample are provided. In some embodiments, the computer program product comprises: a computer usable tangible non-transitory storage medium having computer readable program code thereon, the computer readable program including at least one of the following: (1A) program code for forming a first specification matrix configured as a measure required for selection of a set of barcodes targeting a single chromosome panel, the first specification matrix having a first dimension m and a second dimension equal to the first dimension, wherein: (i) each main diagonal element of the first specification matrix is assigned a value of zero, (ii) each element of a first group of immediately neighboring each other elements of the matrix is assigned a value of a chosen trailing Hamming distance, wherein the size of the first group is no larger than a pre-determined number k < m of said immediately neighboring each other elements and the first group is immediately adjacent to every main diagonal element along the first QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158dimension and / or the second dimension, and (iii) each element of a second group of immediately neighboring each other elements of the first specification matrix along each of the first and second dimensions is assigned a value of a chosen minimum global Hamming distance, wherein said second group is either immediately adjacent to an end element of the first group that is the most distant to a corresponding main diagonal element along the first dimension and along a second dimension or wherein said second group is separated from the end element of the first group along the first dimension and along the second dimension by a corresponding taper group of elements that includes at least one matrix element assigned a value smaller than the value of the chosen trailing distance and larger than the value of the chosen minimum global Hamming distance, the taper group being immediately adjacent to both the first group and the second group; and (IB) program code for forming a second specification matrix configured as a measure required for selection of a set of barcodes targeting a multiple chromosome panel, the second specification matrix having a corresponding first dimension I and a corresponding second dimension equal to the first dimension, wherein: (a) each main diagonal element of the second specification matrix is assigned a value of zero, (b) each of multiple portions of the second specification matrix respectively correspond to each of the multiple chromosomes of the multiple chromosome panel, wherein a respective main diagonal of each of the multiple portions is a part of the main diagonal of the second specification matrix and wherein each of said respective main diagonals of the multiple portions is immediately adjacent to at least one other of said respective main diagonals of the multiple portions, (c) each element of multiple third groups of immediately neighboring each other elements of each of the multiple portions is assigned a value of a chosen trailing Hamming distance, wherein corresponding sizes of the multiple third groups are no larger than a dimension of the corresponding one of the multiple portions, wherein two or more of said multiple third groups are immediately adjacent to every respective element of the main diagonal of the respective one of the multiple portions along the first dimension and / or the second dimension of the second specification matrix, (d) each remaining element of the multiple third groups is assigned a value that is smaller than the value of the chosen trailing distance and larger than a value of a chosen minimum global Hamming distance, and (e) each element of a fourth group of elements that includes immediately neighboring each other elements of the second specification matrix outside of each of the multiple portions is assigned a value of a chosen minimum global Hamming distance.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158BRIEF DESCRIPTION OF THE FIGURES

[0011] FIG.l. Is a schematic showing the concept of global Hamming distance (gHD) among the barcodes within a pool of barcodes.

[0012] FIG. 2. Is a schematic showing the concept of global Hamming distance (gHD) and trailing Hamming distance (trHD) among sets of barcodes within a pool of barcodes. In this schematic, each set of barcodes has a Hamming distance of 5 relative to the other barcodes of the set. The entire pool of barcodes has a global Hamming distance of 3.

[0013] FIG. 3. Provides an example of a specification matrix (or digital filter) containing a set of Hamming-distance-related prescriptions for choosing / defining / devising / designing a set of barcodes that must satisfy such prescriptions and that lend themselves to detecting 30 different targets on a single chromosome. The minimum global Hamming distance of the pool of barcodes is 3; the trailing Hamming distance of sets of barcodes is 6. In this example, a set comprises the barcodes having trailing Hamming distance of 6 and that are correlated to the “k” neighboring targets. In this example, k = 5.

[0014] FIG. 4. Provides a related example of a specification matrix (or digital filter) containing a set of Hamming-distance-related prescriptions for choosing / defining / devising / designing a set of barcodes that must satisfy such prescriptions and that lend themselves to detecting 30 different targets on a single chromosome. The minimum global Hamming distance of the pool of barcodes is 3; the trailing Hamming distance of sets of barcodes is 6. In this example, a set comprises the barcodes having trailing Hamming distance of 6 and that are correlated to the “k” neighboring targets. In addition to the use of the trailing Hamming distance, the specification matrix includes a requirement of a taper(ed) Hamming distance: the “trailing HD” requirement tapers off, extending further to consider more neighboring targets than would be possible with the embodiments of FIG. 2 or FIG. 3, but with a decreased distance requirement that is greater than the minimum global Hamming distance.

[0015] FIG. 5. Provides yet another example of a specification matrix (or digital filter) that - in this case - contains a set of Hamming-distance-related prescriptions for choosing / defining / devising / designing a set of barcodes that must satisfy such prescriptions and that lend themselves QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158to detecting several loci on each of chromosomes 1 , 2, and 3, for a total of 30 different targets. The minimum global Hamming distance of the pool of barcodes is 3; the trailing Hamming distance for barcodes of nearest neighbors on each chromosome (k = 2) is 6. The barcodes of each chromosome have minimum trailing Hamming distance of 5. In this example, the barcodes correlated with chromosome 1 may be considered a set (e.g., set 1); the barcodes correlated with chromosome 2 may be considered a set (e.g., set 2), and the barcodes correlated with chromosome 3 may be considered a set (e.g., set 3). Within each of sets 1-3, a sub-set of barcodes may be defined (set la, set lb, set 1c), as those barcodes with Hamming distance of 6 relative to each other, and that are correlated to the “k” neighbors on each of the different chromosomes.

[0016] FIG. 6. Provides another example of a specification matrix (or digital filter) containing a set of Hamming-distance-related prescriptions for choosing / defining / devising / designing a set of barcodes that must satisfy such prescriptions and that lend themselves to detecting 15 loci on chromosomes 1 and 16 loci on chromosome 2, totaling 31 target loci. The global Hamming distance for the pool of barcodes is 3. The trailing Hamming distance for barcodes of nearest neighbors (k = 2) is 6, and the trailing Hamming distance for barcodes of each chromosome is at least 4. In this example, a “burst” region is introduced and exemplified, in which the nearest neighbors for certain targets is enlarged from k = 2 to k = 4, maintaining the trailing Hamming distance of 6. This allows for increased barcode density in a target region, while maximizing resolution and detection accuracy and efficiency.

[0017] FIG. 7A-7B. Illustrates color collisions in regions of high probe density. The expanded image exemplifies overlapping digits, and the inability to read or distinguish specific targets from one another.

[0018] FIGS.8A, 8B, 8C, 8D, and 8E. Illustrate the concept of color overlap (optical crowding) due to the wavelength of colors, and the need to employ Hamming Distance and / or blue shifting techniques as disclosed herein, within a given round of detection in order to maximize the number of targets that can be visualized per round.

[0019] FIG. 9. Provides a graphic illustration of how an increase in the number of colors (from 3-8) allows for an increase in the number of targets that can be detected (Y-axis) versus the number of rounds of imaging required (X-axis). For example, with 4 rounds of imaging using 3, 4 or 5 QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158different colors, -100, -250, and -600 targets can be detected. However, with 6 colors, over 1200 targets can be distinguished in only 4 rounds (this follows the power law of n-colors to the power of r-rounds (nAr)).

[0020] FIGS. 10A-10B. Provides a schematic to illustrate the concept of optical crowding. FIG.10A exemplifies optical crowding with a single detectable label. FIG. 10B shows that adding additional color can alleviate optical crowding since there is no longer an overlap of the same color.

[0021] FIG. 11. Provides a schematic of a barcode comprising null or blank digits. The introduction of null or blank digits into barcodes has been used to decrease optical crowding, as the blank or null digits produce no color. However, this leads to the need for longer barcodes, more detection cycles, and more controls to check for error based on a blank or null reading, thereby increasing assay costs, and time to result.

[0022] FIG. 12. Provides a schematic of a barcode comprising both null digits and an always on portion.

[0023] FIG. 13. is a schematic illustrating the size of signal generated by blue versus red signal. Due to their shorter wavelengths, detectable signals in the blue range are smaller than those in the red range.

[0024] FIG. 14. Provides an exemplary blue-shifted barcode scheme, where each round of barcodes (imaging cycles) 1-6 includes colors that are skewed toward shorter wavelengths or “blue-shifted,” such that for each round of detection (imaging cycle), the average wavelength count of the colors of digits has a wavelength shorter than the median wavelength of the emission wavelengths of the colors in that round.

[0025] FIG. 15. Provides a schematic showing exemplary component steps to construct a barcode probe sequence library.

[0026] FIG. 16. Provides a schematic showing exemplary steps to design a codebook.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0027] FTGSS. 17A, 17B, 17C, and 17D Illustrate exemplary codebooks generated by the methods described in Examples 1-4. Each codebook is designed using different parameters, including trailing Hamming Distance, k value, and taper trail requirements.

[0028] FIG. 18. Provides an exemplary data acquisition flow chart, illustrating non-limiting steps for collecting images using barcode pools and sets, according to the methods described herein.

[0029] FIGS. 19A, 19B, 19C, and 19D. Provide data illustrating the enhanced detection achieved using the methods disclosed herein. FIG. 19A-19C show target detection using barcodes designed by prior art methods. FIG. 19A-19C show that increased signal crowding results in concomitant decreased detection as target number increases in a given region; each probe set in 19A-19C had a global Hamming distance of 3. FIG. 19A demonstrates detection of a total of 30 different targets (60 total with two sets of chromosome). Given the relatively low density of targets in the specified region, 49 / 60 targets were detected (85%). FIG. 19B demonstrates detection of a total of 60 different targets (120 total with two sets of chromosome) in the same region. As target density increases, the ability to detect discreet barcodes decreases to 76 / 120 (64%). FIG. 19C demonstrates detection of a total of 86 different targets (172 total with two sets of chromosome) in the same region. As target density is increased even more, the ability to detect discreet barcodes decreases even more: only 87 / 172 target barcodes can be detected (50%). Recovery using only a global Hamming distance of 3 suffers overlapping targets / barcodes which are difficult to distinguish because there are too many overlapping digits. In contrast, FIG. 19D which uses Hamming Distance parameters in barcode design and positioning as disclosed herein, show a surprising and unexpected improvement in target detection as compared to FIG.19A-19C. FIG.19D demonstrates detection of the same 86 different targets as FIG. 19C (172 total with two sets of chromosome) in the same region as 19A-19C. By rearranging the barcodes to require a global Hamming distance 3, in an order such that there is a trailing Hamming distance of 4 neighbors with Maximal Hamming Distance (in this case 6 for the 6 digit barcode), 156 / 172 or 90% of the targets are detected (as compared to only 50% in FIG. 19C).

[0030] FIGS. 20A-20C. Provide single cell image overlays of jebFISH chromosome tracing using the present technology, employing a 578-plex assay. FIG. 20A is an image of chromosomeQB\121892.00158\101046255.1Atty. Dkt. No. 121892.001581 probes; FIG. 2B overlays the image of chromosome 5 probes; and FIG 20C further overlays the image of chromosome 7 probes.

[0031] FIGS.20D-20F. Provides images of the single cell shown in FIGs 20A-20C, overlaying additional images of chromosome probes. FIG 20D further overlays the image of chromosome 12 probes; FIG 20E adds chromosome 16 probes; and FIG 20F overlays the image of chromosome 19 probes.

[0032] FIGS.20G-20H. Provides images of the single cell shown in FIGs 20D-20F, overlaying additional images of chromosome probes. FIG 20G further overlays the image of chromosome 21 probes; FIG 20H overlays the image of chromosome X probes.

[0033] FIG. 201. Provides images of the single cell shown in FIGs 20G-20H, overlaying additional images to provide a single image including all chromosome probes.

[0034] FIG. 20 J. Provides a single cell image overlay of jebFISH chromosome tracing using the present technology, employing the same 578-plex assay as in FIGs 20A-20I, but in a different cell type.

[0035] FIGS. 21A-21B. Provides an ideogram depicting the location of the targets for an exemplary 419 probe panel (termed the “ChromoPaint panel”) which are mostly equally spaced along the chromosomes. The image shows chromosome 1-12 (21A), and chromosomes 13-22, X and Y (22B). Each chromosome is oriented with the p-arm on top, and the q-arm at the bottom. The pinched location within each chromosome indicates the location of the centromere. The image includes a representation of 419 different probes, starting with probe 0 at the top of the p-arm of chromosome 1, and ending with the 419thprobe at the end of the q-arm of chromosome Y. There are additional insertions of targets (probes) on Chi 19, chr20, Chr21, CHr22 and ChrY.

[0036] FIGS. 22A-22C. Provide graphical representations of target gains and losses measured in a population of MCF-7 cells using the jebFISH ChromoPaint panel. FIG. 22A provides a representation of the average target copy counts per cell by chromosome in the jebFISH ChromoPaint panel on MCF-7 cells. The top panel shows chromosomes 1-8, and the bottom panel shows chromosome 9-X. Copy amplification can be seen on the q-arm of Chromosome 20, and copy losses can be seen on the p-arm of chromosome 8. FIG. 22B provides presents similar QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158information in the form of an ideogram, illustrating the spatial positions of the target loci along each chromosome. Target gains are evident in chromosomes 8, 17, and 20. FIG. 22C illustrates the distribution of chromosome count per cell in the sample population

[0037] FIG. 22D. Provides an image of copy number gain and loss. The circos plot compares the copy number variation results obtained by array based comparative genomic hybridization (array CGH) and by the present 419 plex imaging approach. Array CGH is a widely used and accepted technique for identifying genomic gains caused by duplications or amplifications, as well as losses caused by deletions. In the circos representation, the outer ring shows the chromosome in ideogram form, including cytogenetic banding patterns, with centromeres indicated by radial tick marks. The middle ring shows the copy number results obtained by array CGH, and the inner ring shows the results obtained using the 419 plex panel applied to the MCF 7 cell line.

[0038] FIG. 23. Structural variations also associated with dysregulated ER signaling observed in MCF-7 cells with the jebFISH ChromoPaint panel. The figure illustrates spatial-relationship analysis results for the 419-plex panel applied to MCF-7 cells. On the left, the figure shows an all-by-all pairwise minimum-distance heat map representing the median spatial proximity between each pair of targets across the cell population. Target pairs that are closer together appear as darker regions in the heat map, while more distant pairs appear lighter. On the right side of the heat map is a distance scale bar, in which dark colors correspond to smaller inter-target distances and progressively lighter colors correspond to greater distances. Positioned in the center of the figure is an inset showing a portion of the copy-number variation graph from FIG. 22A, highlighting the targets located on chromosome 20. This inset is provided to enable direct comparison between copy-number gains and the spatial-distance signatures observable in the heatmap. On the right of the figure is an image of an individual MCF-7 cell with an overlay of chromosome 20. Within this overlay, targets 363, 364, and 365 are highlighted, corresponding to the loci emphasized in the copy-number inset. These targets exhibit elevated copy-number values and are visually identifiable within the overlaid chromosome 20 structure in the cell image.

[0039] FIGS.24A-24B. Structural variations also associated with dysregulated GPCR signaling observed in MCF-7 with the jebFISH ChromoPaint panel. As shown in the images in 24A, the present technology is able to detect the translocations in individual MCF-7 cells (e.g., the t(6;3)QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158translocation). FIG. 24A highlights L-shaped regions in the pairwise minimum-distance heat map that indicate translocations between painted targets located on chromosome 3 and chromosome 6. The left portion of FIG. 24A shows the full heat map with the translocation-associated regions marked. The center panel of FIG. 24A presents a magnified view of the highlighted region, showing the vertical components corresponding to targets 164 and 165, which are assigned to chromosome 6 but display spatial proximity consistent with relocation to chromosome 3. The right panel of FIG. 24A shows an individual cell with overlays of chromosomes 3 and 6, where targets 164 and 165 are circled. Their close spatial proximity to chromosome 3, as predicted by the heat-map pattern, confirms the inferred translocation. FIG. 24B extends this analysis by examining targets 88 and 89, which are normally located on chromosome 3. In the left portion of FIG. 24B, the heat map shows that these targets exhibit close spatial association with loci on chromosome 6. The center panel of FIG. 24B presents a zoomed-in view of the relevant region of the heat map, illustrating the characteristic L-shaped signature associated with interchromosomal translocation. In the right panel, overlays of chromosomes 3 and 6 in a representative cell show that targets 88 and 89 are present both in their expected positions on chromosome 3 and in close proximity to chromosome 6, consistent with a structural rearrangement involving relocation or duplication of these loci.

[0040] FIGS. 25A-25C. Provide graphical representations showing target count and ploidy gains measuredin a population ofMCFlOa cells in chromosomes 1 q, 5q, and 8q, using the jebFISH ChromoPaint panel. Fig. 25A shows the average target copy counts per cell by chromosome in the panel on MCFlOa cells. The top panel shows chromosomes 1-8, and the bottom panel shows chromosome 9-X.Fig. 25B shows the target gains and losses measured in the sample population. Fig. 25C shows the distribution of chromosome count per cell in the sample population.

[0041] FIG 25D. Shows that the measurements of gains and losses correlate with a CGH data. The figure shows a circos plot comparing the copy-number variation results obtained using the JEB-FISH method described herein (as shown in FIG. 22A) with results obtained from single-nucleotide-polymorphism (SNP) array data and array-based comparative genomic hybridization (array CGH). Both SNP arrays and array CGH are widely used and accepted techniques for identifying genomic copy-number gains resulting from duplications or amplifications, as well as losses resulting from deletions. In the circos plot, the outermost ring QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158displays the chromosomes in ideogram form with cytogenetic banding patterns, and the centromeres are marked with radial red tick marks. The outer data ring shows the results obtained using SNP arrays, with green bars positioned outside the circle representing amplifications and red bars positioned inside the circle indicating deletions. The magnitude of each bar corresponds to the amplitude of the copy-number change. The middle ring shows the results obtained by array CGH. The inner ring shows the copy-number variation profile generated by the jebFISH panel applied to the MCF-10A cell line. The results obtained using JEB-FISH generally match the overall trends observed in the array-based measurements, including copy-number gains identified on chromosomes 1, 5, and 8.

[0042] FIG. 26. Provides images of three different MCFlOa cells, cell A, cell B, and cell C.Each of the images has been adjusted to show only the probes for chromosome 1 in the jebFISH ChromoPaint panel. As shown in the image, partial ploidy gain (Iq) is observed in each of the three cells. Each cell shows 2 chromosome 1 copies with extra q-arms. This spatial analysis suggests that gains in Iq are likely due to separate whole / partial q-arm duplications.

[0043] FIG. 27. Provides images of three different MCFlOa cells, cell A, cell B, and cell C. Each of the images has been adjusted to show only the probes for chromosome 8 in the jebFISH ChromoPaint panel. As shown in the image, full and partial ploidy gain (8 and 8q) is observed in the cells. Cell A and Cell B each cell show 2 chromosome 8 copies with extra q-arms. Cell C shows three chromosome 8 copies with extra q-arms. This spatial analysis indicates that gains in chromosome 8 are due to partial q-arm amplification as well as whole chromosome amplification.

[0044] FIGS. 28A-28B. Provides heatmaps demonstrating chromosomal interactions in MCFlOa cell lines. FIG. 28A shows an all targets-by- all targets pairwise minimum distance heatmap of the targets of the Chromopaint panel applied to an MCFlOa cell line. The colors indicate the median value from all the cells of that target pair distance The distance scalebar is on the right. Dark targets are close together. Light targets are further apart. The dark off diagonal patches indicate interactions between Interactions Chr 3, 5 and 9, t(5;3;9) and t(9;3). FIG. 28B is a Hi-C inter-target interaction frequency heatmap for MCFlOa. The dark regions indicate high contact frequency (shorter distance). Again the dark off diagonal patches indicate interactions between Interactions Chr 3, 5 and 9, t(5;3;9) and t(9;3). The depicted data is from ENCODE Hi-QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158C (intact Hi-C) encodeproject.org / experiments / ENCSR370TFL / . For both heatmaps, the Y axis (top to bottom) is chromosome 1- chromosome X; the X axis (from left to right) is chromosome 1 - chromosome X.

[0045] FIGS. 29A-29C. Provides images of a single MCAFlOa cell analyzed with the jebFISH ChromoPaint panel. Each of the images has been adjusted to show specific chromosome overlays as follows. Fig. 29A shows only chromosome 5 probes. Two copies of chromosome 5 are shown, with gain in partial q-arm (5q) targets. Fig 29B shows chromosome 5 with an overlay of chromosome 3. The image shows a copy neutral translocation of partial 3p to 5q gain. FIG. 29C adds the overlay of chromosome 9. Adding chromosome 9 reveals a chromosome 5-3-9 translocation t(5;3;9) and a chromosome 9-3 translocation t(3;9). The insert depict metaphase spectral karyotyping of MCFlOa cells showing a Chr 3,5 and 9 interaction (Cancer Res 2009;69(14):5946-53).

[0046] FIGS.30A-30C. Provide graphical representations showing ploidy changes measured in a population ofHCC1954 cells (HER2+) in chromosomes 8 and 5, using the jebFISH ChromoPaint panel. FIG. 30A shows the average target copy counts per cell by chromosome in the panel on HCC 1954 cells. FIG. 30B shows the target gains and losses measured in the sample population. FIG. 30C shows the distribution of chromosome count per cell in the sample population.

[0047] FIG. 30D. Provides a graphical representation of targets in 8q reported to have multiple inter- and intra-chromosomal rearrangements and interactions e.g., Chr 8-Chr 5. Overall the data in FIGS. 30A-D show copy number gains in chromosome 5p-5q, 8q and llq regions containing hotspots of increased clustered breakpoints indicating higher genomic instability.

[0048] FIG. 31. Provides data showing the inter-chromosomal rearrangement of amplified targets using jebFISH ChromoPaint panel on the HCC194 (HER2+) cell line. The analysis identifies and quantifies the prevalence of chromosome 5 and chromosome 8 possible extrachromosomal DNA (ecDNA) rearrangements. The table in the lower right corner summarizes that 48% of all cells have at least one ec203-ecl 18 pair.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0049] FTG.32. Is an image showing the inter-chromosomal rearrangement of amplified targets using jebFISH ChromoPaint panel on the HCC194 (HER2+) cell line. The system identifies and quantifies the prevalence of chromosome 5-chromosome 8 possible ecDNA rearrangement.

[0050] FIG.33A-33B. FIG. 36A. Workflow showing puncta localization (1), decoding (3), and alignment (2, 4, 5) to generate genomic loci positions (6). Compulsory steps are 1, and 3-6, optional steps are 2 and 7. FIG. 36B. shows an example of an alignment model, XYZ by round by probe. The position of each puncta is translated in 3D by an offset specific to each round and probe combination.

[0051] FIG. 34. Provides an example of reduction in puncta alignment error in z before correction (top) and after (bottom). Each plot shows all the puncta of a single round separated by probe / color.

[0052] FIG. 35. Provides a graph showing different alignment models across the bottom with increasing complexity / number of parameters (dotted line). Alignment errors generally decrease as model complexity increases (colored lines).

[0053] FIG. 36. Provides images and data examples showing puncta within a cell using a percell alignment model before (top left) and after correction (top right). Puncta are colored in rainbow sequence by round. The scale parameters optimized during error minimization show that cells in this sample generally shrink over the 6 rounds of imaging (bottom), consistent with the single cell example on the left.

[0054] FIG. 37. provides graphical and tabular summaries of residual positional error following application of the described correction procedures.

[0055] FIG. 38. are images showing the inter-chromosomal rearrangement of amplified targets using jebFISH ChromoPaint panel on the HCC194 (HER2+) cell line. The image on the left shows that many of the amplified copies of target 203 are ecDNA which preferentially interacts with sites harboring breakpoint hotspots on 5p, 5q, and 1 Iq.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158DETAILED DESCRIPTION

[0056] Chromosomes, and genomes in general, are organized in three dimensions such that functionally related genomic elements, e.g. enhancers and their target genes, are located in close spatial proximity, or located such that they can directly interact. Determining how chromosomes and chromosome regions are positioned in relation to each other and accurately characterizing chromosome topology within a nucleus provides insights into gene regulation and cellular processes, both in healthy and diseased tissues and cells.

[0057] Methods for determining 3D chromosome structure include established hybridizationbased and / or ligation-based assays in conjunction with well-known imaging approaches (e.g., various FISH techniques, Hi-C, and GAM). Using such techniques, various analyses have confirmed 3D chromatin territories, compartments, and domains. While such techniques may lend themselves to broad surveys of 3D structure, to advance our understanding of the contributions of chromosome structure, positioning, and localized gene expression in the processes of development, disease, and cell growth and division, there is a need in the art for more accurate and more comprehensive 3D imaging methods and compositions.

[0058] Accordingly, the present technology provides novel methods for designing and using pools and sets of barcodes to more accurately, comprehensively, and efficiently analyze the 3D structures, positions, proximity, and movement of nucleic acids within cells and within the nucleus of cells.

[0059] I, Considerations for Barcode Design, Identification, Formation, Generation, and / or Definition.

[0060] This disclosure discusses methodologies for designing pools and sets of molecular barcodes for the detection of a plurality of nucleic acid targets in a biological sample. Also disclosed are corresponding methods for analyzing, detecting and / or visualizing target molecules. In particular, examples of applications of the pools and sets of barcodes for use in genome imaging and next-generation sequencing are discussed that are suitable to reveal the complexity and biological importance of the genome's 3D configuration.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0061] As is known in the art, there are many oligonucleotide-based platforms for in situ detection of nucleic acid targets utilizing barcodes. Such platforms include oligonucleotides having a variety of structures depending on the assay configuration, which ultimately link the barcode to a target. Non-limiting examples of such oligos include Oligopaints, multiplexed error-robust fluorescence in situ hybridization (MERFISH) oligos, seqFISH oligos, RNA sequential probing of targets (SPOTs) oligos, high-coverage microscopy -based technology (Hi-M) oligos, or optical reconstruction of chromatin architecture (ORCA) oligos or any to oligonucleotide used for FISH methods and / or any oligonucleotides that has a sequence complementary (e.g. recognition domain) to a target molecule, e.g., an oligonucleotide sequence, a portion of a DNA sequence, or a particular chromosome or sub-chromosomal region of a particular chromosome. For further details, see e.g., Cardozo et al., Mol Cell. 2019 Apr 4;74(l):212-222; Mateo et al., Nature. 2019 Apr;568(7750):49-54; Wang et al., Scientific Reports volume 8, Article number: 4847 (2018); Shah et al., Neuron, Volume 92, Issue 2, 19 October 2016, Pages 342-357; Eng et al., Nat Methods.2017 Dec;14(12): 1153-1155; each of which is incorporated herein by reference in its entirety.

[0062] In each of these platforms, a molecular barcode is linked to a nucleic acid target via a nucleic acid probe, wherein the nucleic acid probe is configured to hybridize to the target sequence. Additionally or alternatively, in the case of a signal amplification schemes, a molecular barcode may be linked to a target via secondary or tertiary molecules, wherein a first molecule specifically binds the target, a second (or secondary) molecule binds the first, and a third or tertiary molecule binds the second, etc. The barcoded oligo then hybridizes to the third or tertiary, or the “Nth” amplification (or intervening) oligonucleotide. The present technology is intended to not be limited by any specific oligonucleotide detection platform (i.e., the present technology is not limited to the method of linking a barcode to its nucleic acid target), although one or more non-limiting examples of target nucleic acid detection are presented herein which incorporate the novel barcode pools and / or sets.

[0063] The barcode pools and barcode sets of the present disclosure are designed to result in higher spectral resolution, and to detect more targets with greater accuracy and with fewer rounds of analysis, and to provide increased spatial resolution without the use of additional agents as fiducials, as compared to those employed with currently used methods of barcoded nucleic acid detection. By utilizing different Hamming-distance-related parameters, and optionally employing QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158one or more of an “always on” design strategy, and selecting for blue-shifted digits, the barcode pools and barcode sets of the present technology provide superior detection capabilities with respect to puncta (the barcode digit’s representative color or signal, e.g., the detectable label of a detection probe hybridized to the barcode digit) resolution, error correction, and the number of rounds needed to detect a given number of targets, and improved spatial resolution without the use of additional agents as fiducials.

[0064] Embodiments of the present invention achieve the higher spectral resolution and detection of higher number of targets while reducing the number of rounds of analysis at least in part by providing a methodology of devising a digital fdter or prescription (which may be interchangeably referred to herein as a specification matrix or a Hamming distance matrix) that contains a collection or set of Hamming-distance-related values. The Hamming-di stance values are prescribed according to the physical proximity, location, or likely physical proximity or location of the selected targets in multi-dimensional space (see e.g., FIG. 15, element 100, and FIG. 16), resulting in a codebook (and pool) of barcodes able to provide accurate and efficient detection of the predefined targets.

[0065] As used in this disclosure and unless expressly specified otherwise, each of the terms “digital filter” and “specification matrix” is defined as and refers to a multi-dimensional array of expressions arranged in rows and columns, which array is used to represent Hamming distance requirements for ordered barcodes to be selected based on such requirements. In at least one implementation, the specification matrix or digital filter may be formatted as a rectangular two-dimensional array of numbers.

[0066] A, Hamming Distance Considerations in Barcode Design

[0067] Hamming distance - as a metric used to compare data strings of equal length, which identifies the number of positions at which the corresponding data are different - is used in related art. However, according to the idea of the present invention, the proposed methodology of generation of pools and sets of barcodes employs specific Hamming distance considerations that are not present in prior art barcoding designs, and are thus not present in the resulting barcode pools. As illustrated in FIG. 1 and FIG.2 designing barcode pools and barcodes sets within a poolQB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158with differing Hamming distance constraints allows for improved detection efficiency, resolution, and error correction.

[0068] As used herein, the terms “barcode pool” and / or “pool of barcodes” or similar terms refer to the totality of barcodes used in a particular analysis or practical procedure. For example, if an analysis requires the detection of 180 different nucleic acid targes, the pool of barcodes would include 180 barcodes, here, 1 barcode per target. The term “set” - as used with respect to barcodes - refers to a group of barcodes within or from a barcode pool. For example, a barcode pool may contain several sets of bar codes, with each set including a plurality of barcodes satisfying one or more common requirements (e.g., the barcodes in a set are defined to target the same territory, have a specified Hamming distance relative to each other that is different from the global Hamming distance of the pool, etc.).

[0069] As used herein, the term “codebook” refers to the colorspace barcodes to be used together in an assay, and represents a description or the code used to generate the corresponding pool of barcodes. Essentially, the codebook provides the instructions for generating the physical pool of barcodes (e.g., oligonucleotide sequences encoding color signals). A more detailed discussion regarding the codebook is provided below.

[0070] As used herein, the term “molecular barcode” refers to a nucleic acid sequence comprising at least one barcode region, the barcode region comprising at least one molecular barcode digit, wherein each molecular barcode digit is configured to specifically hybridize to a barcode hybridization region of a detection probe. The signal from the detection probes provides the output digits. A molecular barcode digit can be 3-7 nucleotides in length, wherein each molecular barcode digit (i.e., sequence of nucleotides) is specific to a color barcode digit. By way of example but not by way of limitation, the molecular barcode digit sequence AAAAA is hybridized by detection probe TTTTT-[label 1], The molecular barcode digit sequence GGGGG is hybridized by detection probe CCCCC-[lable-2], where the color of label 1 and label 2 are different.

[0071] As used herein, the terms “barcode digit,” refers to the single unit of a multi -unit barcode. When used in reference to a molecular barcode digit, a digit comprises a sequence of nucleic acids that is configured to correspond to a color signal. When used in reference to colors (e.g., a readout), QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158a barcode digit refers to a single color at a given position on a barcode. In some embodiments, a color signal is produced from a detectable label, such as a fluorophore, linked to a detection probe, where the detection probe hybridizes to its corresponding molecular barcode digit sequence. For simplicity, in this disclosure barcode digits are represented by Roman numerals. For example, a barcode may be represented by the numerical sequence “254342”. The numeral “2” represents and refers to both a nucleic acid sequence and a color. The numeral “5” represents both a different nucleic acid sequence and a different color. Likewise, the numerals 3 and 4 each represent the respectively corresponding different nucleic acid sequence digits, each such nucleic acid sequence digit corresponding to a different color. By way of example, but not by way of limitation, such barcodes may be termed “colorspace barcodes.” In discussed examples of embodiments, each target being detected is assigned a sequence of numerical digits that represents is unique “barcode” (e g., 254342 is an example of a colorspace barcode). See e g., FIG. 17.

[0072] As used herein, the terms “barcodes” and / or “target barcode” refer to a set of colored digits or digits designed and used in the present disclosure to specifically identify one or more nucleic acid targets in a biological sample. As an example, when using detectable labels such as fluorescent labels to detect target molecules, the discussed embodiments of compositions and methods provide detection probes comprising different colors (e.g., fluorophores) that localize to a molecular barcode that is linked to its target, wherein the detection probes assemble on the molecular barcodes in a predetermined arrangement, i.e., a barcoded sequence of colors. Barcodes identified or generated using an embodiment of the proposed methodology may have 3 digits, 4 digits, 5 digits, 6 digits, 7 digits, 8 digits, 9 digits or 10 or more digits. According to the idea of the present invention and substantially in every embodiment, an individual barcode comprises at least three different color signals, four different color signals, five different color signals, six different color signals, seven different color signals, or more. By way of example but not by way of limitation, the barcodes in the pools of barcodes discussed below comprise 6 digits, wherein each digit corresponds to a corresponding color signal. In some embodiments, each digit corresponds to a different color, or in alternative embodiments, specific digits may correspond to the same color. In some embodiments, a pool of barcodes comprises 3 different colors, 4 different colors, 5 different colors, 6 different colors, 7 different colors, or more.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0073] As used herein, the term “detectable label” refers to a molecule that generates or is capable of generating a detectable signal. Detectable labels may include, for example, lightabsorbing dye, a fluorescent dye, or a radioactive label. Fluorophores are an example of detectable labels that can generate color signals. In embodiments, detectable labels are linked to detection probes.

[0074] As used herein, the term “color signal” refers to the color of a detectable label, such as a fluorescent dye, when such label is exposed to light at a wavelength corresponding to it.

[0075] In embodiments, detectable labels, molecules, and / or moieties can include those that can be detected by spectroscopic, photochemical, biochemical, immunochemical, electromagnetic, radiochemical, or chemical means, such as fluorescence, chemifluorescence, or chemiluminescence, or any other appropriate means. Detectable labels can include, but are not limited to radioisotopes, bioluminescent compounds, chromophores, antibodies, chemiluminescent compounds, fluorescent compounds, metal chelates, and enzymes.

[0076] In embodiments, the detectable label may include a fluorescent compound. As is known in the art, when a fluorescent compound is exposed to light of the proper wavelength, its presence can then be detected due to fluorescence. By way of example but not by way of limitation, a detectable label can comprise a fluorescent dye molecule, or fluorophore including, but not limited to fluorescein, phycoerythrin, phycocyanin, o-phthalaldehyde, fluorescamine, Cy3TM, Cy5TM, allophycocyanin, Texas Red, peridinin chlorophyll, cyanine, tandem conjugates such as phycoerythrin-Cy5TM, green fluorescent protein, rhodamine, coumarine, fluorescein isothiocyanate (FITC) and Oregon GreenTM, rhodamine and derivatives (e g., Texas red and tetrarhodimine isothiocyanate (TR1TC)), biotin, phycoerythrin, AMCA, CyDyesTM, 6-carboxyfhiorescein (commonly known by the abbreviations FAM and F), 6-carboxy-2',4',7',4,7-hexachlorofiuorescein (HEX), 6-carboxy-4',5'-dichloro-2',7'-dimethoxyfiuorescein, JOE or J), N,N,N',N'-tetramethyl-6carboxyrhodamine (TAMRA or T), 6-carboxy-X -rhodamine (ROX or R), 5-carboxyrhodamine-6G (R6G5 or G5), 6-carboxyrhodamine-6G (R6G6 or G6), and rhodamine 110. Additionally the family of dyes comprising or including: Cyanine, BODIPY, Coumarin, Oxazine, Carbopyronin, Pyrene, Rhodol, Acridine, Phenanthridine, Indole, Benzimidazole, Carbocyanine and Naphalimide, and derivatives thereof.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0077] As used herein the term “detection probe” refers to a molecule that includes a detectable label and a barcode hybridization region that specifically binds the nucleic acid sequence of a molecular digit on a molecular barcode. By way of example and not by way of limitation, a detection probe comprises a nucleic acid sequence comprising a molecular barcode hybridization region and a detectable label.

[0078] As used herein, the phrase “detecting a target barcode” comprises hybridizing detection probes to their corresponding molecular barcode sequences, and imaging the detectable label of the detection probes. As discussed herein, hybridizing and imaging occurs over multiple rounds (e.g., as many rounds as digits present in the barcode).

[0079] As noted previously, the present technology is not intended to be limited by the means of linking a barcode to its target (e.g., there may be one or more amplification oligonucleotides between the target and the barcode oligonucleotide, hybridized in series, for example). Accordingly, the phrase “labeling the barcode(s) linked to nucleic acid target(s) with detection probes,” or “labeling a plurality of nucleic acid targets,” wherein the nucleic acid target is linked to a barcode, encompasses hybridization of a detection probe to its corresponding molecular barcode sequence. The target is thereby labeled via the detection probe, which comprises a detectable label, and which is hybridized to the barcode that is linked to the target.

[0080] As used herein, the term “removing a color signal” refers to a step of eliminating the (spectral) signal generated by a detectable label. Removing a color signal may include removing the detectable label that generates the color signal and / or removing the detection probe comprising the detectable label. For example, signal removal may be accomplished by denaturation, dehybridization, washing steps, photobleaching, channel or wavelength switching, chemical cleavage, chemical blocking, bleaching, etc. Identification of barcodes linked to targets typically involves sequential rounds of visualization and removal of color signals. By way of example but not by way of limitation, if each barcode in a pool includes three digits, there would be a first round of visualization to detect the first digit of each barcode followed by a first round of color signal removal to remove the first color signal; a second round of visualization to detect the second digit of each barcode and a second round of signal removal to remove the second color signal; andQB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158a third round of visualization, optionally followed by a third round of color removal to remove the third color signal.

[0081] Depending on the specifics of a particular embodiment, a barcode is defined to include 3 digits, 4 digits, 5 digits, 6 digits, 7 digits or more. In some embodiments, a number of digits is represented by and labelled with an “N.” By way of example but not by way of limitation, four different barcodes (referred to as A, B, C, and D) each comprising six digits (N = 6) are shown in Table 1 below with corresponding color signals. In this example, each digit includes a color signal; no digit of any barcode is a blank, null, or “no signal.”Table 1:

[0082] As used herein, the term “Hamming distance” (HD) refers to a metric for comparing two data strings. Comparing two strings of equal length, Hamming distance is the number of string positions in which the two pieces of data are different. By way of example but not by way of limitation, the Hamming distance between two binary codes: (A) 01101 and (B) 11100 is two. The Hamming distance between two words: (A) kitten and (B) button is three. The Hamming distance between two words: (A) button and (B) carton is also three. The Hamming distance between two words: (A) kitten and (B) carton is four. When referring to two barcodes comprising an equal number of digits (colored digits), Hamming distance is defined as the number of differences (i.e., different colors) between corresponding digits of the two barcodes. Barcode examples of Table 1QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158are outlined in the Table 2 below, with addition of comparison of the Hamming distances of three different barcodes (B, C, and D) to barcode A. Each of barcodes A, B, C, and D has 6 digits. The color signal of each digit is presented in the digit position column.Table 2.

[0083] As used herein “global Hamming distance” (gHD) refers to the Hamming distance among all the barcodes in a pool. In some embodiments, the pool may include multiple sub-pools termed “sets” (e.g, a first, second, third ... Nth set), as was already alluded to above. The gHD refers to the Hamming distance among all the barcodes in the pool and encompasses the sets. By way of example, each individual barcode in a pool of barcodes used to positionally map the chromosome of a cell, would be compared to every other barcode in the pool in defining a gHD. Every pair of barcodes in the pool would have a global Hamming distance of a given value, or of a given range. For example, according to the idea of the invention, a global Hamming distance among a pool of barcodes, each having 6 digits could be “at least 3,” meaning that no two barcodes in the pool has a Hamming distance of less than 3, i.e., at least three differences exist between any two barcodes in the pool (some barcode pairs could have 3 differences between the barcode in such pairs, some barcode pairs could have 4 differences, some barcode pairs could have 5 differences, and some barcode pairs could have 6 differences). In this example, no barcode pairs would have only two differences, only one difference, or no differences at all.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0084] As used herein the term “trailing Hamming distance” (trHD) refers to a Hamming distance among or between barcodes in a defined multi-dimensional space, such as neighboring barcodes, or barcodes on the same chromosome, the same chromosome arm, or specific region of a chromosome. For a trHD, in some embodiments, the number of neighbors (k) provides the “defined space”. By way of example, the trHD can be defined as the Hamming distance of two nearest neighboring barcodes (k = 2), three (k = 3), four (k = 4), five (k = 5), or six (k = 6) neighboring barcodes. According to the idea of the invention, increasing the trailing Hamming distance allows for and leads to the accurate detection of more targets that are closer together without the color signal of the neighboring barcodes interfering with each other.

[0085] By way of example, according to the idea of the invention, at least in some embodiments trailing Hamming distance may be designed into or used as a requirement for identifying a set (or a group, or a pool) of barcodes defining a territory, e.g., a space (whether linear or multidimensional) within which neighboring barcodes are characterized by such trailing Hamming distance. Such a set of barcodes may be referred to as a territory set of barcodes, e.g., barcodes located in a defined territory or space of a biological sample, or occupying a defined locus or region of a biological sample (see FIG. 2). By way of example but not by way of limitation, a trailing Hamming distance may be associated with barcodes on a specific chromosome, chromosome arm, or sub-region of a chromosome (a gene or fragment thereof, an intron, promoter region, coding region, non-coding region, telomere, etc.). As another example, the defined space or territory may be the three-dimensional territory occupied by one or more chromosomes or portions of chromosomes within the nucleus of a cell. In some embodiments, a barcode pool comprises a gHD, and one or more subsets of barcodes in the pool (e.g., territory sets of barcodes) also comprise a trailing HD relative to all of the other barcodes in that territory set. Different territory sets of barcodes (that is, a set of barcodes representing or associated with a given territory) may have different trailing HD values, or the same trailing HD values. In some embodiments (as, for example, the one discussed below in reference to FIG. 5), a territory set of barcodes may have a first trailing HD (trHD l) and a subset of barcodes within that territory may have a second trailing HD (trHD_2). As used herein, a second trailing HD (trHD_2) is greater than a first trailing HD.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0086] In some embodiments, the trailing Hamming distance equals the number of rounds of an experiment / analysis. Thus, for a 6 round experiment, the trHD would be 6. In some embodiments, the number of rounds is equal to the number of digits of one or more barcodes (N).

[0087] In situations in which the barcode includes several blank or null sequential digits (e.g., where the barcode is 000000123456 or 000000123456000000), the trHD value is defined based on the maximum sequential numbered non-null (non-blank) set of 6 and NOT 12 or 18.

[0088] Examples of trHD are provided by the tables below. The Tables 3, 4, and 5 show 8 examples of barcodes, numbered 1 through 8, in numerical order. The k value (defining the barcodes encompassed by the trHD, that is defining neighboring barcodes) is also shown. In the first table, the k value is 2. The trailing Hamming distance in Table 3 is shown as “X”. This table shows that among the 8 barcodes (with trHD value of “X” where “X” in this example is equal to the number of sequential rounds used to identify a given barcode) the pair of barcodes identified as 1 and 2 (or, numbered as #1 and #2) must have or satisfy the requirement of the defined trHD, but barcodes 1 and 3 do not have to satisfy such requirement (they could, but it is not required). Likewise, barcodes #2 and #3 must have the defined trHD of X, but barcodes 2 and 4 do not have to (although they could, but it is not required). In some embodiments, the trHD equals the number of rounds of barcode analysis. For example, for a 6-round analysis, the trHD would be 6.Table 3QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0089] The following table provides an example of trHD encompassing 3 neighboring barcodes (k = 3). In this example, barcodes numbered as #1, #2, #3 must all have the defined trHD of X relative to each other, but barcoded #1 and #4 do not have to. Barcodes #2, #3, and #4 must have the defined trHD of X relative to each other, but barcodes #2 and #5 do not.Table 4

[0090] The following table 5 provides an example of trHD encompassing 4 neighboring barcodes (k = 4). In this example, barcodes #1, #2, #3, and #4 must all have the defined trHD of X relative to each other, but barcodes #1 and #5 do not have to. Barcodes #2, #3, #4, and #5 must have the defined trHD of X relative to each other, but barcodes #2 and #6 do not have to.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158Table 5

[0091] FIG.3 provides an illustration of trHD of 6 (X = 6), where k = 4.

[0092] While the examples in the tables 3, 4, and 5 above show a linear arrangement of a territory set of barcodes #1 through #8, it is understood that the barcodes in a territory set may be grouped together based on three-dimensional proximity; a linear arrangement (e g., a linear alignment along a single chromosome arm), is not a requirement. As used herein, a trHD value is greater than a gHD.

[0093] As used herein, the term “taper trailing Hamming distance (ttHD)” refers to a choice of Hamming distance(s) that allows for a tapered or at least quasi-gradual transition between a Hamming distance with a value corresponding to trHD and value corresponding to a gHD (and / or vice versa). Table 6 below provides an example of such gradual transition between the Hamming distance values, showing the choice of barcodes with Hamming distance reduced, for identified barcodes, by 1 and then by 2. If X=number of rounds or digits, and if X equals the maximum Hamming distance (for example 6), then - for k=4 - the set of barcodes would include barcodes with HD=6, then the following barcode (here, barcode #5) would have a minimum HD=5, then theQB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158next following barcode (here, barcode #6) would have a minimum HD = 4 etc before transitioning to the minimum of the global Hamming distance, e.g., gHD=3.Table 6

[0094] FIG. 4 illustrates the implementation of a specification matrix for devising or choosing or identifying a set of barcodes that are satisfying a requirement of having the trailing Hamming distance “tapered” or quasi-gradually changed between the two extreme values: one corresponding to the trailing Hamming distance, another corresponding to the minimum global Hamming distance. FIG.4 shows a readout of barcodes having a trHD = 6, with k = 4 (similar to those shown in FIG.3) However, barcodes of FIG.3 do not exhibit a taper from trHD of 6 to gHD of 3; rather the barcode immediately adjacent to the trHD=6 barcoded set can have HD = 3; this means that there can be 3 overlapping digits in the adjacent barcodes, which can be difficult to clearly distinguish or identify in practice. In contrast, the choice of the specification matrix summarized in FIG. 4 includes the use of taper trail Hamming distances, with Hamming distance values of adjacent barcodes gradually tapering to transition from HD = 6 to HD =3. As show in FIG.4, there is a “transition zone” going from HD =6 to HD = 3, stepping through HD = 6, to HD > 5 for 1-2 positions, to HD > 4 for 1-2 positions, before HD = 3. This improves the ability to decode the barcodes accurately by reducing the number of overlapping digits in adjacent barcodes.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0095] To capitalize on the advantages of the concept of taper trail Hamming distance and to increase the size of a barcode codebook, the trail can be reset at chromosome breakpoints, centromeres, or other distinguishing features. As shown in FIG. 5, for example, in at least one implementation of the proposed methodology the specification matrix for defining a codebook can be “reset” at the end of a chromosome, thereby allowing for an additional degree of freedom of more codes to be selected for the next “segment” of the overall barcode codebook. As a result -and in stark contradistinction with approaches used in related art - such resetting at chromosome breakpoints leads to the desired increase in codebook size.

[0096] According to the idea of the invention, the methodology of devising or designing a set of barcodes may also employ - in addition or in the alternative to aspect or features already alluded to above - another aspect termed herein a “burst,” which can be advantageously implemented for analysis of dense target specific regions. As illustrated in FIG.6, the Hamming distance in specific regions of the specification matrix may be tightened (increased) among neighbors to facilitate efficient and substantially error-free detection. In a “burst,” the barcodes in a specific segment or sub-region of a “territory” have a higher Hamming distance than other barcodes of the territory.

[0097] Thus, in designing a pool of barcodes to detect and localize a plurality of in situ nucleic acid targets, the present methodology leverages the concept of chromosome territories (or other 3D defined space(s) that include the targets) by minimizing color collisions in each territory, in each round, by maximizing the use of “colorspace” with judiciously identified Hamming distance constraints built into the barcode design, (or, put differently, used as filters or constraints on choosing barcodes for a codebook). Specifically, at least one trailing Hamming distance and optionally, a taper trail Hamming distance requirement are designed into neighboring barcodes and barcode sets, while a requirement of specific barcode bursts can be additionally used to derive even more detailed information by densely populating a specific sub-region. As summarized and illustrated in FIG. 7, by combining these Hamming distance parameters or filters or requirements into barcode pools and barcode sets within the pools, not only is the number of useful barcodes expanded, but the resolution, accuracy, and efficiency of the analysis is greatly improved over current barcoding designs and methods.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0098] As described previously, the term “codebook” refers to the set of colorspace barcodes to be used together, and represents a pool of barcodes. The codebooks generated by the present methods are unique: they are not simply a collection of barcodes. Rather, a codebook generated according to the idea of the present methodology include sequences of barcodes that are necessarily ordered according to the various Hamming distance requirements used as filters to select such barcodes for the codebook. The order of the barcodes in the codebook is meaningful and significant. The order of barcodes is based on the physical location of the selected targets with respect to one another, and the prescribed parameters (e.g., global Hamming distance, different territories and defined ttHDs, regions of interest for more intense labeling an analysis, bursts, tlHDs, etc.) that are used as conditions in order to build the pool (codebook).

[0099] In some embodiments, the target barcodes may further include a prefix barcode. As used herein, a “prefix barcode” (or, a prefix, for short) refers to additional digits on a barcode, independent of the target barcode digits, that define a set of barcodes within a pool. By way of example but not by way of limitation, the set of barcodes comprising a prefix may define barcodes correlated with a particular territory, such as a chromosome or chromosome arm, correlated with a particular nucleic acid sequence (e.g., a repeated gene sequence or regulatory region, e.g., present on different chromosomes or in different chromosomal regions), correlated with multiple regulatory elements of a gene in a genome, etc.

[0100] B, An “Always On” Feature

[0101] Another aspect of the proposed methodology related to devising useful barcodes includes an “always on” feature.

[0102] In barcoding techniques used in related art, different color targets are usually measured separately. This may be accomplished by sequentially stepping through each wavelength individually or by changing filter / dichroic sets by having separate detectors for different colors. And, it is typically possible to resolve two different spatial spots located very close together (sometimes within even the same pixel) if they are different colors. (See e.g., FIG. 8A, 8B, 8C, 8D, and 8E).QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0103] As the person of skill in the art will readily appreciate, the use of a larger number of colors allows for identification of a fixed number of targets in fewer rounds. As shown in FIG.9, in an example with 500 targets, 3 rounds are required when 8 fluorophores are used; 4 rounds are required when between 5 and 7 colors are used, 5 rounds are needed when 4 colors are used, and 6 rounds are needed with 3 colors are used. One clear disadvantage of using too many colors, however, manifest in substantial spectral overlap and signal crosstalk, as schematically illustrated in FIG. 10A-10B

[0104] One option commonly used in related art to create spectral space is to incorporate a null or “zero” signal into the barcodes thereby creating fewer optical “on” targets per round of imaging (see e.g., FIG. 11).

[0105] Disadvantages of adding such “blanks” include the inevitably following need of more rounds of detection to complete an analysis, at least to control for such null or “blank digit: is the reading of a null actually represents a presence of a “signal,” or is the assay failing and providing a false null reading? Along with the extra rounds of analysis are the concomitant costs in time and materials. In addition, 3D imaging requires the use of fiducials embedded in the sample; the introduction of blanks into the barcodes results in loss of certainty regarding spatial stability and relative position of targets.

[0106] To address these shortcomings, embodiments of the invention employ what is referred to as an “always on” feature. In some embodiments, “always on” is used in combination with the Hamming distance considerations / requirements detailed above, to prepare or devise barcode pools and sets.

[0107] As is known in the art, several barcoding methods use null or blank digits in combination with color (or other detectable label) signal digits to effectively increase the number of different digits in a given barcode. As used herein, the term “always on” generally refers to barcodes that contain a sequence of digits such that between the first “on” digit of the sequence and the last “on” digit in the sequence there are no more than 15% of blank, null, or zero digits. The “always on” feature, according to the idea of the invention, does not require the elimination of the null, but requires that an “always on” sequence of digits within a barcode includes no more than 15% null digits. FIG. 12 provides an example of a barcode having a total of 19 digits. Null QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158digits (no color or detectable signal) are represented by “0” whereas digits with a detectable signal are represented by numbers 1 through 6. In this exemplary barcode, digits 1 through 7 are nulls, digit 13 is null, and digits 15 through 19 are null. Even though 13 out of 19 digits of the barcode as a whole are null (-70%), the barcode still satisfies the “always on” requirement, because the percentage of nulls between the first “on” digit (represented by the number 1) and the last “on” digit (represented by the number 6) is no more than 15%, i.e., 1 null in 7 digits.

[0108] Non-limiting advantages of the always on feature include detection of higher plex in fewer rounds, using puncta to align each wavelength and each round, eliminating the need for fiducials, and evaluating error correction values with more confidence as there is no uncertainty surrounding a blank or null digit (i.e., eliminating the questions: is it blank or was the target missed?). Accordingly, the present technology provides for the use of the barcode digits as intrinsic fiducials to align images from round to round (in each round).

[0109] Positional accuracy without fiducials

[0110] The present technology enables high-throughput localization of genomic loci through the following steps:

[0111] Locus labeling and imaging: Each genomic locus is labeled and imaged as puncta multiple times across different rounds and probe colors.

[0112] Barcode decoding: The identity of each locus is determined by its unique barcode — defined by the specific pattern of rounds and probes in which its puncta appear.

[0113] Position calculation: The spatial position of each locus is computed from the coordinates of its matched puncta across rounds and probes.

[0114] Precise alignment of puncta is critical for both accurate barcode decoding and proper measurement of locus positions at subpixel and sub-diffraction resolution. Even small shifts can lead to incorrect associations between puncta and barcodes, resulting in missing or wrong locus identities. In extreme cases, misalignment can cause complete failure of barcode decoding.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0115] Classical approaches typically rely on adding fiducials to the sample. However, fiducial-based alignment poses challenges: fiducials must be detectable across all rounds and colors at intensities comparable to puncta, remain perfectly stationary relative to the biological sample, and be titrated to specific densities

[0116] The present technology eliminates the need for external fiducials by using the puncta themselves (e.g., the barcode digit’s representative color), combined with the expected barcode pattern, to achieve fine alignment. By leveraging the large number of puncta distributed across rounds and probes, the method provides dense intrinsic reference points for robust alignment. This enables iterative, model-based corrections that can measure and compensate for round-to-round drift, chromatic shifts, and sample deformation — delivering subpixel precision without additional calibration steps.

[0117] This technique can be applied because of the “always on” nature of the barcodes: There is a fluorescently labeled target present in every round for a particular barcode. This creates more certainty around the cluster of targets because there is not a point in time when there is no signal present by design.

[0118] FIG. 36 shows and exemplary workflow design, including steps 1-7. Each will be described in turn.

[0119] Step 1 : Puncta detection and localization. Genomic loci are labeled and imaged in 3D across multiple rounds, and in different colors within the same round. Each locus appears as fluorescent puncta in the image volume, which are detected and localized to measure their positions. The output of this step is a list of puncta, each annotated with its 3D coordinates, the round and probe in which it was observed, and additional metadata such as the cell to which it belongs. This information forms the basis for subsequent alignment and barcode decoding.

[0120] Step 2: (Optional) Coarse system-level corrections. Imaging systems often exhibit consistent, machine-specific offsets between colors or probes across runs. These system-level errors typically include chromatic aberration and optical distortions such as barrel or pincushion distortion. They can be characterized in advance and used for coarse correction of puncta positions, either at the image level or after localization. Calibration is typically performed using standard QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158multi-colored fluorescence beads fixed on a slide to measure these systematic distortions. Because these corrections are global and stable across runs, they provide an improved initial alignment for barcode decoding. Corrections applied at this stage can be excluded from the fine barcode-based alignment model described later, ensuring that subsequent steps focus on residual, run-specific misalignments. Coarse sample-level corrections may also optionally be performed. Large positional shifts can occur during an imaging run due to sample movement. These shifts can be corrected coarsely, either at the image level or after puncta localization. For example: images may be aligned by image content. Signals such as DAPI staining, autofluorescence, or background fluorescence can be used for coarse alignment. Additionally or alternatively, images can be aligned by puncta positions. The global pattern of puncta across image volumes can serve as a reference for coarse alignment. These corrections improve initial alignment and can be omitted from the fine barcode-based alignment model described later.

[0121] Step 3: Decoding puncta into genomic loci / barcode. Puncta across rounds and probes are clustered based on spatial proximity (e.g., using a fixed 3D distance radius or spatial clustering). Genomic loci (each described by a barcode — specific pattern of puncta round and probe) are then decoded from these clustered puncta (e.g., using a max-flow min-cost algorithm, where low-intensity or poorly aligned puncta contribute to higher costs). This approach is moderately robust to missing or spurious puncta but depends on reasonable initial alignment or a sufficiently large clustering distance threshold. As alignment improves, barcode detection becomes more accurate and sensitive, i.e. allowing for iterative improvements (see step 7).

[0122] Step 4: Genomic loci / barcode-level correction. Decoding puncta produces a list of genomic loci, where each locus position is estimated by averaging the positions of its associated puncta. The spatial spread of puncta within a locus (or barcode) directly reflects alignment errors: if there were no positional errors, all puncta for a given locus would perfectly colocalize. The fine-alignment method disclosed herein aims to minimize these errors using an appropriate alignment model with optimized parameters. False or mis-detected puncta that do not belong to any locus are naturally excluded from the alignment process. Similarly, alignment can be made more rigorous by removing low-confidence loci and puncta.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0123] Choice of alignment model. A mathematical model is required to quantify and compensate for the misalignments. This model can be based on prior knowledge (e.g., system chromatic aberrations) or developed through trial and error (e.g., polynomial fit with XY dependence). The model may combine global transformations (e.g., XY translation across the entire field of view) and local adjustments (e.g., per-cell shrinkage over rounds). Any per-puncta metadata collected can also be incorporated. Alignment precision generally improves with model complexity, which is constrained by the number of puncta available.

[0124] Fitting and error minimization. A suitable fitting algorithm should be used to minimize puncta alignment error, depending on the chosen alignment model. Since the true locus positions are unknown, the error metric is typically the spread or variance of puncta positions within each detected locus. For example, we fitted an alignment model using scipy. optimize. least_squares, which employs the Levenberg-Marquardt algorithm internally and minimizes a least-squares loss. Robust error minimization methods (e.g. outlier rejection) can be used to minimize the effects of bad puncta or incorrectly decoded loci.

[0125] Alignment correction and parameter analysis. Once alignment errors are minimized, the model and its fitted parameters can be applied to all puncta to improve overall alignment.

[0126] Step 5 Additional improvements may be achieved by refining the alignment model or improving the decoding of the source puncta data (see step 7). Depending on the chosen alignment model, examining the fitted parameters can provide experimental insights — for example, assessing cell deformation across imaging rounds.

[0127] Step 6: Re-decoding puncta into genomic loci / barcode. After alignment, all puncta are re-decoded using their updated positions. This results in a more accurate misalignment cost, which improves the performance of the max-flow min-cost algorithm — yielding more correct loci and reducing false positives. Additionally, the distance threshold for initial clustering can be decreased to further minimize false positives. Due to more accurate puncta positions, the accuracy of genomic loci positions is also improved.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0128] Step 7: (Optional) Iterative alignment refinement. Because fine-alignment benefits from more accurate puncta decoding — and vice versa — it can be advantageous to repeat these steps iteratively. Each iteration may yield progressive improvements in alignment; however, performance is ultimately constrained by the suitability of the alignment model and inherent positional noise (e g., localization precision).

[0129] C. Blue-shifted Barcodes

[0130] Furthermore, in addition to satisfying at least some of the Hamming distance considerations and an “always on” requirement, in some embodiments, the implementation of the proposed methodology also incorporates “blue-shifted” imaging cycles into the devised barcode pools and sets. That is, digits in a pool of barcodes, e.g., for a given imaging cycle or imaging round, comprise an overall spectral shift towards the shorter wavelengths. As already mentioned above, as the puncta density increases with increasing plexity (i.e., more color signals in a specific area), puncta resolution becomes compromised by crowding and spectral overlap (see FIG. 13).To decrease spectral overlap and increase resolution in high plex assays, blue shifted barcode pools and sets are designed. As illustrated in FIG 14, barcodes are devised with colors skewed toward shorter wavelengths or “blue shifted,” such that for each round of detection (imaging cycle), the average wavelength count of the colors of digits has a wavelength shorter than the median wavelength of the emission wavelengths of the colors in that round.

[0131] II, Systems and Methods for Generating Pools and Sets of Barcodes

[0132] Disclosed herein are method for designing pools and sets of barcodes that satisfy one or more of the above-discussed Hamming distance requirements, an “always-on” feature or limitation, and that contain “blue-shifted” digit(s). The barcodes configured according to the idea of the invention enable three-dimensional mapping of a plurality of nucleic acid targets with greater accuracy, higher resolution, and more cost-effectiveness than those afforded by the prior art barcode designs.

[0133] Referring now to FIG. 15, the first step in designing a colorspace codebook according to the present technology requires, as a first step, defining or identifying the genomic targets (100).As used herein, a “genomic target” refers to a nucleic acid sequence in the nucleus of a cell, or that QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158comprises the genome of a cell (e.g, a bacterial cell), or that is present within the cell or within a virus. The nucleic acid sequence can be DNA or RNA or a combination thereof, and can be chromosomal or extrachromosomal (e.g., a transcribed message, a plasmid, a viral sequence, a chromosome fragment). Typically, the cell is a eukaryotic cell, such as a human cell, however, the present technology is not limited by cell type, and detection of genomic targets in plants, fungi, bacteria, and virus is also contemplated.

[0134] Compared to prior art methods, the present technology lends itself to the accurate and efficient identification of a large number of genomic targets, and defining hundreds, thousands, or tens of thousands of genomic targets can be implemented as the starting point at step 100. While accurate and efficient identification of a large number of targets is one option, the present technology may also be used with a small number of targets (e.g., 1-5, 10, 20, 30, 50, 100).

[0135] Once the genomic targets have been defined by nucleic acid sequence and relative position in the genome, barcode candidates are selected from a barcode database (600) to design the colorspace codebook (codebook). FIG. 16 includes a summary of exemplary steps for designing a codebook, resulting in the final codebook.

[0136] III. Methods of Detecting and Visualizing Genomic Targets

[0137] Once a codebook is developed, barcoded probe sequences are synthesized, and methods well-known in the art for labeling the selected targets with the barcodes, and detecting barcoded targets can be employed. By way of example but not by way of limitation a schematic of data acquisition is presented in the flow chart at FIG. 18. Well-known methods may be employed to visualize and read the barcode targets.

[0138] Embodiments.

[0139] By way of example but not by way of limitation, barcoded oligonucleotide pools of the present technology may be employed as follows.

[0140] Embodiment 1. A method for in situ positional mapping of a plurality of nucleic acid targets in a biological sample, the method comprising: detecting a pool of target barcodes linked to the plurality of nucleic acid targets, wherein each target barcode comprises N QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158digits, wherein each position of each target barcode is represented by a color signal generated from a detectable label, and wherein no position of any barcode lacks a corresponding color signal, or wherein the sequence of digits within a barcode includes no more than 15% null digits; and wherein detecting comprises: (a) labeling the barcodes linked to the plurality of nucleic acid targets with a first pool of first detection probes, wherein each first detection probe comprises a corresponding first detectable label and a barcode hybridization region, wherein different first detection probes are correlated to different targets or sets of targets of the plurality; (b) imaging the detectable labels of (a), wherein the first detection probes hybridizes to a barcode digit 1 position and provide the corresponding first color signal as digit 1 of the target barcode to which the first detection probe is hybridized ;(c) removing the first color signal of (b); (d) labeling the barcodes linked to the plurality of nucleic acid targets with a second pool of second detection probes, wherein each second detection probe comprises a corresponding second detectable label and a barcode hybridization region, wherein different second detection probes are correlated to different targets or sets of targets of the plurality; (e) imaging the detectable labels of (d) wherein the second detection probes hybridize to a barcode digit 2 position and provide the corresponding second color signal as digit 2 of the target barcode to which the second detection probe is hybridized; (f) removing the second color signal of (e); (g) repeating steps (d)-(f) with an Nth pool of detection probes, wherein each Nth detection probe in each Nth pool from the second pool to the Nth pool (i) comprises a corresponding detectable label, (ii) is correlated to different targets or sets of targets of the plurality; and (iii) provides a corresponding color signal as an Nth digit of the target barcode; wherein, in combination, the N pools of probes comprise at least 3 different colors.

[0141] Embodiment 2. The method of embodiment 1, wherein at least 50% of the detectable labels in each pool have a color signal comprising a wavelength below the median wavelength of the color signals of that pool.

[0142] Embodiment s. The method of any of the previous embodiments, wherein the barcodes of the pool comprise a global hamming distance (gHD) of at least 2.

[0143] Embodiment 4. The method of embodiment 3, wherein the global hamming distance (gEID) is 3 or more.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0144] Embodiment 5. The method of embodiment 3, wherein the global hamming distance (gHD) is equal to N.

[0145] Embodiment 6. The method of any of the previous embodiments, wherein a territory set of target barcodes comprise a trailing hamming distance (trHD), wherein the trHD is at least one unit greater than the global hamming distance.

[0146] Embodiment 7. The method of embodiment 6, wherein a territory set of target barcodes is defined as sets of two nearest neighbors (k = 2).

[0147] Embodiment 8. The method of embodiment 6, wherein a territory set of target barcodes is defined as sets of three (k = 3), four (k = 4), five (k = 5), or six (k = 6) neighboring barcodes.

[0148] Embodiment 9. The method of any of the previous embodiments, wherein a territory set of target barcodes comprises a subset of barcodes, the subset comprising a second trailing Hamming distance (trHD_2).

[0149] Embodiment 10. The method of embodiment 9, wherein the number of neighboring target barcodes with trHD_2 comprises two barcodes (k = 2), three (k = 3), four (k = 4), five (k = 5), or six (k = 6) barcodes.

[0150] Embodiment 11. The method of embodiment 9 or 10, wherein the trITD_2 is equal to N.

[0151] Embodiment The method of any of the previous embodiment wherein the target barcodes comprise at least one gHD, at least one trHD, and at least one taper trail hamming distance (ttHD).

[0152] Embodiment 13. The method of embodiment 12, wherein the ttHD is between the trlTD and the gHD.

[0153] Embodiment 14. The method of any of the previous embodiments, wherein the gl D is 3, and ttHD is greater than or equal to N-2, wherein N is at least 6.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0154] Embodiment 15. The method of any of the previous embodiments, wherein the territory set of target barcodes encompasses target barcodes on the same chromosome.

[0155] Embodiment 16. The method of any of the previous embodiments, wherein territory set of target barcodes encompasses target barcodes on the same chromosome arm.

[0156] Embodiment 17. The method of any of the previous embodiments further comprising (h) positional mapping of the plurality of nucleic acid targets without the use of extrinsic fiducials.

[0157] Embodiment 18. The method of any of the previous embodiment, wherein each target barcode further comprises a prefix barcode, wherein the prefix barcode digits are independent of the target barcode digits.

[0158] Embodiment 19. The method of embodiment 18, wherein the prefix barcode comprises two digits.

[0159] Embodiment 20. The method of any of the previous embodiments, wherein at least one prefix barcode is specific to a chromosome.

[0160] Embodiment 21. The method of any of the previous embodiments, wherein at least one prefix barcode is specific to a chromosome territory.

[0161] Embodiment 22. The method of any of the previous embodiments, wherein N is 6.

[0162] Embodiment 23. A method for designing a set of molecular barcodes comprising colored fluorophores for detection of a plurality of nucleic acid targets in a biological sample, wherein each position of each molecular barcode is represented by a color signal generated from a detectable label, and wherein no position of any barcode lacks a corresponding color signal, or wherein the sequence of digits within a barcode includes no more than 15% null digits; the method comprising: (a) with the use of a programmable processor operably connected with tangible, non-transitory storage medium containing program code thereon, accessing data representing: (i) a color input value, wherein the color input value defines the number of differentQB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158colors for use in the set of barcodes, wherein the color input value is at least 3; (ii) a barcode digit input value, wherein the barcode position input value defines the number of digits of each barcode, wherein the position input value is N, and wherein each position of each barcode is represented by a color and no position of any barcode lacks a corresponding color signal; (iii) a target input value, comprising the total number of targets (loci); (b) selecting, based on the accessing data of (a), the set of molecular barcodes, wherein: (i) the global hamming distance (gHD) for all barcodes within the set is N-2; wherein the selected barcodes enable three-dimensional mapping of the plurality of nucleic acid targets.

[0163] Embodiment 24. The method of embodiment 23, further comprising at step (a) accessing data comprising one or more of (iv) a territory input value, comprising the total number of territories comprising targets in the biological sample; (v) territory barcode set input values (k), each value defining the number of barcodes in a territory ("territory set of barcodes").

[0164] Embodiment 25. The method of embodiment 24 (b), wherein: (ii) the territory set of barcodes comprises a trailing hamming distance (trHD) for each territory set of barcodes; and / or (iii) a taper trail Hamming distance (ttHD) for barcodes within a territory set.

[0165] Embodiment 26. The method of embodiment 23, wherein the territory input value is an integer greater than 2.

[0166] Embodiment 27. A method for in situ positional mapping of a plurality of nucleic acid targets in a biological sample, the method comprising: detecting a pool of target barcodes linked to the plurality of nucleic acid targets, wherein each target barcode comprises N positions, and wherein detecting comprises: (a) labeling the barcodes linked to the plurality of nucleic acid targets with a first pool of first detection probes, wherein each first detection probe comprises a corresponding first detectable label and a barcode hybridization region, wherein different first detection probes are correlated to different targets or sets of targets of the plurality; (b) imaging the detectable labels of (a), wherein the first detection probes hybridizes to a barcode digit 1 position and provide the corresponding first color signal as digit 1 of the target barcode to which the first detection probe is hybridized; (c) removing the first color signal of (b); (d) labeling the plurality of nucleic acid targets with a second pool of second detection probes, wherein each second detection probe comprises a corresponding second detectable label and a barcode QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158hybridization region, wherein different second detection probes are corelated to different targets or sets of targets of the plurality; (e) imaging the detectable labels of (d) wherein the second detection probes hybridize to a barcode digit 2 position and provide the corresponding second color as position 2 of the target barcode to which the second detection probe is hybridized; (f) removing the second color signal of (e); (g) repeating steps (d)-(f) with an Nth pool of detection probes, wherein each Nth detection probe in each pool from the second pool to the Nth pool (i) comprises a corresponding detectable label, (ii) is correlated to different targets or set of targets of the plurality; wherein a territory set of target barcodes comprise a trailing Hamming distance (trHD) value that is greater than a taper trail Hamming distance (ttHD) value of a subset of target barcode in the same territory set.

[0167] Embodiment 28. The method of embodiment 27, wherein the territory set of barcodes (k) is defined as sets of two nearest neighbors (k = 2).

[0168] Embodiment 29. The method of embodiments 27 or 28, wherein the territory set of target barcodes is defined as sets of three (k = 3), four (k = 4), five (k = 5), or six (k = 6) neighbors.

[0169] Embodiment 30. The method of any one of embodiments 27-29, wherein ttHD is greater than the trETD value.

[0170] Embodiment 31. The method of any one of embodiments 27-30, wherein the territory set of barcodes encompasses barcodes on the same chromosome.

[0171] Embodiment 32. The method of any one of embodiments 27-31, wherein the territory set of barcodes encompasses barcodes on the same chromosome arm.

[0172] Embodiment 33. The method of any one of embodiments 27-32, wherein the territory set of barcodes encompasses barcodes in three-dimensional chromosome territory.

[0173] Embodiment 34. A computer program product for devising a set of barcodes configured to detect nucleic acid targets in a biological sample, the computer program product comprising a computer usable tangible non-transitory storage medium having computer readable program code thereon, the computer readable program including at least one of the following: QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158(34A) program code for forming a first specification matrix configured as a measure required for selection of a set of barcodes targeting a single chromosome panel, the first specification matrix having a first dimension m and a second dimension equal to the first dimension, wherein: (i) each main diagonal element of the first specification matrix is assigned a value of zero, (ii) each element of a first group of immediately neighboring each other elements of the matrix is assigned a value of a chosen trailing Hamming distance, wherein the size of the first group is no larger than a predetermined number k m o said immediately neighboring each other elements and the first group is immediately adjacent to every main diagonal element along the first dimension and / or the second dimension, and (iii) each element of a second group of immediately neighboring each other elements of the first specification matrix along each of the first and second dimensions is assigned a value of a chosen minimum global Hamming distance, wherein said second group is either immediately adjacent to an end element of the first group that is the most distant to a corresponding main diagonal element along the first dimension and along a second dimension or wherein said second group is separated from the end element of the first group along the first dimension and along the second dimension by a corresponding taper group of elements that includes at least one matrix element assigned a value smaller than the value of the chosen trailing distance and larger than the value of the chosen minimum global Hamming distance, the taper group being immediately adjacent to both the first group and the second group; and (34B) program code for forming a second specification matrix configured as a measure required for selection of a set of barcodes targeting a multiple chromosome panel, the second specification matrix having a corresponding first dimension / and a corresponding second dimension equal to the first dimension, wherein: (a) each main diagonal element of the second specification matrix is assigned a value of zero, (b) each of multiple portions of the second specification matrix respectively correspond to each of the multiple chromosomes of the multiple chromosome panel, wherein a respective main diagonal of each of the multiple portions is a part of the main diagonal of the second specification matrix and wherein each of said respective main diagonals of the multiple portions is immediately adjacent to at least one other of said respective main diagonals of the multiple portions, (c) each element of multiple third groups of immediately neighboring each other elements of each of the multiple portions is assigned a value of a chosen trailing Hamming distance, wherein corresponding sizes of the multiple third groups are no larger than a dimension of the corresponding one of the multiple portions, wherein two or more of said multiple third groups are QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158immediately adjacent to every respective element of the main diagonal of the respective one of the multiple portions along the first dimension and / or the second dimension of the second specification matrix, (d) each remaining element of the multiple third groups is assigned a value that is smaller than the value of the chosen trailing distance and larger than a value of a chosen minimum global Hamming distance, and (e) each element of a fourth group of elements that includes immediately neighboring each other elements of the second specification matrix outside of each of the multiple portions is assigned a value of a chosen minimum global Hamming distance.

[0174] Embodiment 35. A computer program code according to embodiment 34, wherein the program code for forming a second specification matrix is configured to form the second specification matrix in which corresponding sizes of said two or more of the multiple third groups are not equal to one another.

[0175] Embodiment 36. The method of any one of embodiments 1-22, wherein the barcode digits are used as fiducials to align images round to round.

[0176] Embodiment 37. A method for in situ positional mapping of a plurality of nucleic acid targets in a biological sample, the method comprising: detecting a pool of target barcodes linked to the plurality of nucleic acid targets, wherein each target barcode comprises N digits; and wherein detecting comprises: (a) labeling the barcodes linked to the plurality of nucleic acid targets with a first pool of first detection probes, wherein each first detection probe comprises a corresponding first detectable label and a barcode hybridization region, wherein different first detection probes are correlated to different targets or sets of targets of the plurality; (b) imaging the detectable labels of (a), wherein the first detection probes hybridizes to a barcode digit 1 position and provide the corresponding first color signal as digit 1 of the target barcode to which the detection probe is hybridized; (c) removing the first color signal of (b); (d) labeling the barcodes linked to the plurality of nucleic acid targets with a second pool of second detection probes, wherein each second detection probe comprises a corresponding second detectable label and a barcode hybridization region, wherein different second detection probes are correlated to different targets or sets of targets of the plurality; (e) imaging the detectable labels of (d) wherein the second detection probes hybridize to a barcode digit 2 position and provide the corresponding second color signal as digit 2 of the target barcode to which the detection probe is hybridized; (f)QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158removing the second color signal of (e); (g)repeating steps (d)-(f) with an Nth pool of detection probes, wherein each Nth detection probe in each Nth pool from the second pool to the Nth pool (i) comprises a corresponding detectable label and barcode hybridization region, (ii) is correlated to different targets or sets of targets of the plurality; and (iii) provides a corresponding color signal as an Nth digit of the target barcode to which the Nth detection probe is hybridized; wherein, in combination, the N pools of probes comprise at least 3 different colors, and wherein the barcode digits from each round are used a fiducials to align images round to round.

[0177] Embodiment 38. The method of embodiment 37, wherein each position of each target barcode is represented by a color signal generated from a detectable label, and wherein no position of any barcode lacks a corresponding color signal, or wherein the sequence of digits within a barcode includes no more than 15% null digits.

[0178] Miscellaneous

[0179] The present disclosure is described herein using several definitions, as set forth below and throughout the application.

[0180] As used in this specification and the claims, the singular forms “a,” “an,” and “the” include plural forms unless the context clearly dictates otherwise. For example, the term “a protease” should be interpreted to mean “one or more proteases” unless the context clearly dictates otherwise. As used herein, the term “plurality” means “two or more.”

[0181] As used herein, “about”, “approximately,” “substantially,” and “significantly” will be understood by persons of ordinary skill in the art and will vary to some extent on the context in which they are used. If there are uses of the term which are not clear to persons of ordinary skill in the art given the context in which it is used, “about” and “approximately” will mean up to plus or minus 10% of the particular term and “substantially” and “significantly” will mean more than plus or minus 10% of the particular term.

[0182] As used herein, the terms “include” and “including” have the same meaning as the terms “comprise” and “comprising.” The terms “comprise” and “comprising” should be interpreted as being “open” transitional terms that permit the inclusion of additional components further to those components recited in the claims. The terms “consist” and “consisting of’ should be interpreted QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158as being “closed” transitional terms that do not permit the inclusion of additional components other than the components recited in the claims. The term “consisting essentially of’ should be interpreted to be partially closed and allowing the inclusion only of additional components that do not fundamentally alter the nature of the claimed subject matter.

[0183] EXAMPLES

[0184] The following examples are illustrative and are not intended to limit the scope of the claimed subject matter.

[0185] Example 1 - Designing a Colorspace Codebook: A Single Chromosome Panel

[0186] FIG.3 depicts an embodiment of a specification matrix (or digital filter) containing a set of Hamming distance related prescriptions for choosing / defining / devising / designing a set of barcodes satisfying such prescriptions.

[0187] In the example of FIG. 3, the resulting barcodes to be generated with the use of the specification matrix are used to target several specific genomic loci spanning the length of human chromosome 1, comprising 30 total target loci. The oligonucleotide probes that target each locus bear unique DNA barcodes (molecular barcodes) encoding colorspace barcode digits which encode unique target IDs. To assign a unique colorspace barcode to each target locus, a codebook of size 30 will be constructed. In this example, the length of each barcode to be generated using the specification matrix of FIG. 3 is designated to be 6 digits.

[0188] The pairwise Hamming distance requirements for the codebook is stored using a two-dimensional numerical matrix (see FIG. 3). In the matrix, each number on the X and Y axes represent a different locus (1-30). The numerical values within the matrix represent the minimum Hamming distance of the barcodes for the intersecting locus pairs. For example, locus 5 and locus 10 have a minimal Hamming distance of 3. Locus 5 and locus 3 have a minimal Hamming distance of 6 First, a specification matrix (which encodes codebook Hamming distance requirements) is constructed, and then barcodes are selected that satisfy matrix constraints in order to generate the final codebook, as described below.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0189] To begin, a 30 x 30 pairwise matrix with all zero values is initialized (e.g., by a brief Python program using the numpy package, known to those of skill in the art). In order to imbue the resulting codebook with error-correcting properties, a minimum global Hamming distance (min gHD) of three units is then enforced (i.e. min gHD = 3). To achieve this, every cell in the initialized matrix is updated to a value of 3. As an exception, the main diagonal of the matrix (or line of identity) stores barcode-with-self relationships and should be set to zero to avoid impossible Hamming distance constraints.

[0190] Next, to ensure that neighboring chromosomal targets have different colors in each round to avoid optical collisions within fluorescent channels within such rounds, a trailing Hamming distance (trHD) is enforced so that neighboring targets comprise different colors in each round. To matrix cells that run parallel to the diagonal but offset by up to the pre-determined k units (here, k=4) , assign the minimum Hamming distance equal to 6 (in this example, the number of rounds).

[0191] With the completed specification matrix in hand (see FIG. 3), the barcodes are selected that satisfy the constraints of Hamming distance values as listed in the matrix. In particular, to construct the colorspace codebook with the use of the matrix of FIG. 3 and, for example, a “bottom-up” greedy approach known in the art, the following algorithm (presented below in most general layman’s terms) may be implemented :Until we have completed the codebook:Pick a random barcode from the candidate poolIf this barcode could serve as the valid "next" barcode, select it, else skip it

[0192] A computer program product (for devising a set of barcodes that are configured to detect nucleic acid targets in a biological sample) that implements such specification matrix is within the scope of the invention. Such computer program product includes a computer usable tangible non-transitory storage medium that contains computer readable program code at least a portion of which is program code for forming a specification matrix configured as a measure required for selection of a set of barcodes targeting a single chromosome panel. Such first specification matrix has a first dimension m and a second dimension equal to the first dimension and satisfies the QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158following conditions: (i) each main diagonal element of the specification matrix is assigned a value of zero, (ii) each element of a first group of immediately neighboring each other elements of the matrix is assigned a value of a chosen trailing Hamming distance (here, the size of the first group is no larger than a pre-determined number k < m of the immediately neighboring each other elements and the first group is immediately adjacent to every main diagonal element along the first dimension and / or the second dimension), (iii) each element of a second group of immediately neighboring each other elements of the specification matrix along each of the first and second dimensions is assigned a value of a chosen minimum global Hamming distance. The second group of the immediately neighboring each other elements of the matrix is immediately adjacent to such end element of the first group that is the most distant with respect to a corresponding main diagonal element along the first dimension and along the second dimension.

[0193] Example 2 - Single Chromosome Panel with a Hamming distance “Taper”

[0194] FIG. 4 depicts an embodiment of a specification matrix (or digital filter) containing a set of Hamming distance related prescriptions for choosing / defining / devising / designing a set of barcodes satisfying such prescriptions. In this example, the resulting barcodes to be generated with the use of the specification matrix are used to target several specific genomic loci spanning the length of Human chromosome 1, comprising 30 total target loci. The oligonucleotide probes that target each locus bear unique DNA barcodes encoding colorspace barcode digits which encode unique target IDs. To assign a unique colorspace barcode to each target locus, a codebook of size 30 will be constructed. In this example, the length of each barcode to be generated using the specification matrix of FIG. 4 is designated to be 6 digits. Note, that - in comparison with the Example 1 - the definition of barcodes for the codebook of FIG. 4 is carried out by combining (that is, using in addition to one another) the requirement of the trailing Hamming distance trHD and the requirement of a tapering of the trailing distance (ttHD).

[0195] The pairwise Hamming distance requirements for the codebook, in this case, is again stored in a two-dimensional numerical matrix, wherein the value within the matrix of intersecting loci defines the minimum Hamming distance for that pair of loci. First, a specification matrix that encodes codebook Hamming distance requirements is constructed and then specific barcodes are selected that satisfy matrix constraints in order to generate the final codebook.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0196] To begin, initialize a 30 x 30 pairwise matrix with all zero values (e g., as discussed in Example 1) In order to imbue the resulting codebook with error-correcting properties, a minimum global Hamming distance of three units (min gHD = 3) is enforced. To achieve this, every cell in the initialized matrix to a value of the . As an exception, the diagonal of the matrix (or line of identity) is used to stores barcode-to-self relationships and should be set to zero to avoid impossible Hamming constraints.

[0197] Next, to ensure that neighboring chromosomal targets are of different colors in each round to avoid optical collisions within fluorescent channels within rounds, a trailing Hamming distance requirement is imposed when identifying barcodes so that neighboring targets are illuminated using different colors in each round. To achieve this, in matrix cells that run parallel to the diagonal but are offset from the diagonal by k units, the minimum distance is set to a number that is dependent, in a pre-determined fashion, on the value of k. (In one non-limiting example, those cells offset by k not exceeding 4 may be assigned the value of the trailing Hamming distance of 6 (in this example - number of rounds); those cells offset from the diagonal by k=5 may be assigned the trailing Hamming distance of 5; those cells offset offset by k=6 may be assigned the trailing Hamming distance of 4; and those cells of the matrix having a larger separation from the diagonal may be assigned not a particular value of the minimal trailing distance bur, instead, a value of the global Hamming distance - in the example of FIG. 4, the value of “3”). In this example, the trailing Hamming distance requirement is made to taper off, as shown, extending further to consider more neighboring targets than in Example 1, but with a decreased distance requirement. Overall, the pre-determined fashion - according to which this “tapering” of the integer values of the trailing Hamming distance assignment to the cells of the matrix is accomplished - is not generally described by a step-function but is preferably represented (at least in one specific example) by a substantially monotonic function.

[0198] With the completed specification matrix in hand (see FIG. 4), the barcodes are selected that satisfy the constraints of Hamming distance values as listed in the matrix. In particular, to construct the colorspace codebook with the use of the matrix of FIG. 4 and, for example, a “bottom-up” greedy approach known in the art, the following algorithm (presented below in most general layman terms) may be implemented:QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158Until we have completed the codebook:Pick a random barcode from the candidate poolIf this barcode could serve as the valid "next" barcode, select it, else skip it

[0199] A computer program product that implements a specification matrix with “tapering” Hamming distance is also within the scope of the invention. Such computer program product includes a computer usable tangible non-transitory storage medium that contains computer readable program code at least a portion of which is program code for forming a specification matrix configured as a measure required for selection of a set of barcodes targeting a single chromosome panel. Such specification matrix has a first dimension m and a second dimension equal to the first dimension and satisfies the following conditions: (i) each main diagonal element of the specification matrix is assigned a value of zero, (ii) each element of a first group of immediately neighboring each other elements of the matrix is assigned a value of a chosen trailing Hamming distance (here, the size of the first group is no larger than a pre-determined number k < m of the immediately neighboring each other elements and the first group is immediately adjacent to every main diagonal element along the first dimension and / or the second dimension), (iii) each element of a second group of immediately neighboring each other elements of the specification matrix along each of the first and second dimensions is assigned a value of a chosen minimum global Hamming distance. The second group of the immediately neighboring each other elements of the specification matrix is separated from the end element of the first group along the first dimension and along the second dimension by a corresponding taper group of elements. The taper group of elements includes at least one matrix element assigned a value that is smaller than the value of the chosen trailing distance and that is larger than the value of the chosen minimum global Hamming distance. The taper group is immediately adjacent to both the first group and the second group.

[0200] Example 3 - A Three Chromosome Panel

[0201] FIG.5 depicts an embodiment of a specification matrix (or digital filter) containing a set of Hamming distance related prescriptions for choosing / defining / devising / designing a set of QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158barcodes satisfying such prescriptions. Tn this example, the resulting barcodes to be identified with the use of the specification matrix of prescriptions are intended to target several specific genomic loci on each of Human chromosomes 1, 2, 3, 30 total target loci. The oligonucleotide probes that target each locus bear unique DNA barcodes encoding colorspace barcode digits which encode unique target IDs. To assign a unique colorspace barcode to each target loci, a codebook of size 30 is be constructed. In this example, a 6-digit barcode will be used, although other digit size can be employed, comprising more or fewer digits than 6.

[0202] The pairwise Hamming distance requirements imposed onto the definition of barcodes for the codebook is stored using a two-dimensional numerical matrix in which the numerical numbers identifying rows and columns indicate barcode index and where matrix cells are used to store the minimum required Hamming distance of the corresponding barcode pair. First, a specification matrix the encodes codebook Hamming distance requirements is constructed and then, with the use of this matrix, barcodes are selected that satisfy specific pre-defined constraints in order to generate the final codebook.

[0203] To begin, a 30 x 30 pairwise matrix is initialized with all zero values (for example, as stated in Example 1). In order to imbue the resulting codebook with error-correcting properties, a minimum global Hamming distance of three units be enforced (i.e., per the idea of the invention, min gHD=3). To achieve this, every cell in the initialized matrix is updated to a value of the minimum gHD (here, 3). As an exception, the matrix diagonal (or line of identity) is used to store barcode-to-self relationships and should be set to zero to avoid impossible Hamming constraints.

[0204] Next, to ensure that neighboring chromosomal targets are of different colors in each round to avoid optical collisions within fluorescent channels within rounds, a chosen trailing Hamming distance requirement is enforced so that neighboring targets are illuminated using different colors in each round. To this end, matrix cells that run parallel to the matrix diagonal but offset by up to k units (here, k=l, only the nearest neighbors), are updated to the value of the minimum distance equal to 6 (i.e. number of rounds).

[0205] Within chromosomes, the Hamming distance should be larger than the minimum global Hamming distance. Therefore, for each chromosomal subset of the global codebook (indicated inQB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158FIG. 5 as portions “chrl”, “chr2”, and “chr3” of the matrix), a minimum Hamming distance of 5 is imposed.

[0206] With the completed specification matrix in hand (see FIG. 5), the barcodes are selected that satisfy the constraints of Hamming distance values as listed in the matrix. In particular, to construct the colorspace codebook with the use of the matrix of FIG. 5 and, for example, a “bottom-up” greedy approach known in the art, the following algorithm (presented below in most general layman terms) may be implemented :Until we have completed the codebook:Pick a random barcode from the candidate poolIf this barcode could serve as the valid "next" barcode, select it, else skip it

[0207] Example 4: Densely Targeting Genomic Regions

[0208] FIG. 6 depicts the specification matrix (or digital filter) containing a set of Hamming distance related prescriptions for choosing / defining / devising / designing a set of barcodes satisfying such prescriptions. In this example, the resulting barcodes to be identified with the use of the specification matrix of prescriptions are intended to target several specific genomic loci on each of Human chromosomes 1 and 2, totaling 32 target loci. The oligonucleotide probes that target each loci bear unique DNA barcodes encoding colorspace barcode digits which encode unique target IDs. To assign a unique colorspace barcode to each target loci, a codebook of size 32 is constructed. Again, in this example a 6-digit barcode will be exemplified, although barcodes with fewer or more digits can also be employed.

[0209] The pairwise Hamming distance requirements for the codebook are stored using a two-dimensional numerical matrix where the numbers rows and columns indicate barcode index and in which cells store the minimum required Hamming distance of the corresponding barcode pair. First, a specification matrix which encodes codebook Hamming distance requirements is constructed and then barcodes are selected that satisfy matrix constraints of such specification matrix in order to generate the final codebook.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0210] To begin, a 32 x 32 pairwise matrix is initialized with all zero values, e.g. by writing a brief Python program using the numpy package. In order to imbue the resulting codebook with error-correcting properties, a global minimum Hamming distance of three units be enforced (i.e., per the idea of the invention, min gHD = 3). To achieve this, every cell in the initialized matrix is updated to a value of 3. As an exception, the diagonal or line of identity stores barcode-self relationships and should be set to zero to avoid impossible Hamming constraints.

[0211] Next, to ensure that neighboring chromosomal targets are of different colors in each round to avoid optical collisions within fluorescent channels within rounds, a “trailing” Hamming distance is enforced so that neighboring targets are illuminated using different colors in each round. In matrix cells that run parallel to the diagonal but offset by pre-determined k units, set the minimum distance equal to 6 (i.e. number of rounds).

[0212] Within chromosomes, the Hamming distance should be larger than the minimum global Hamming distance. Therefore, for each chromosomal subset of the global codebook (indicated in FIG.6 as matrix subsets “chrl” and “chr2”, a minimum Hamming distance of 5.

[0213] With the completed specification matrix in hand (see FIG. 6), the barcodes are selected that satisfy the constraints of Hamming distance values as listed in the matrix. In particular, to construct the colorspace codebook with the use of the matrix of FIG. 6 and , for example, a “bottom-up” greedy approach known in the art, the following algorithm (presented below in most general layman terms) may be implemented:Until we have completed the codebook:Pick a random barcode from the candidate poolIf this barcode could serve as the valid "next" barcode, select it, else skip it

[0214] The corresponding computer program product includes a computer usable tangible non-transitory storage medium that contains computer readable program code at least a portion of which is program code for forming a specification matrix configured as a measure required for selection of a set of barcodes specifically targeting multiple chromosome panels of Example 3 QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158and / or dense targeting of genomic regions of Example 4. The program code for forming a particular specification matrix configured as a measure required for selection of a set of barcodes targeting a multiple chromosome panel. The particular specification matrix has a corresponding first dimension I and a corresponding second dimension equal to the first dimension. Here, program code forms the particular specification matrix such that: (a) each main diagonal element of the second specification matrix is assigned a value of zero; (b) each of multiple portions of the particular specification matrix respectively corresponds to each of the multiple chromosomes of the multiple chromosome panel (where a respective main diagonal of each of the multiple portions is a part of the main diagonal of the particular specification matrix and where each of such respective main diagonals of the multiple portions is immediately adjacent to at least one other of the respective main diagonals of the multiple portions; and (c) each element of multiple groups of immediately neighboring each other elements of each of the multiple portions is assigned a value of a chosen trailing Hamming distance. (For condition (c), corresponding sizes of such multiple groups are no larger than a dimension of the corresponding one of the multiple portions, and two or more of such multiple groups are immediately adjacent to every respective element of the main diagonal of the respective one of the multiple portions of the particular specification matrix along the first dimension and / or the second dimension of the particular specification matrix. Further the particular specification matrix is formed such that: (d) each remaining element of the multiple groups identified in (c) is assigned a value that is smaller than the value of the chosen trailing distance and larger than a value of a chosen minimum global Hamming distance, and (e) each element of yet another group of elements that includes immediately neighboring each other elements of the particular specification matrix outside of each of the multiple portions is assigned a value of a chosen minimum global Hamming distance.

[0215] Example 5: Codebook Examples

[0216] FIGS. 17A-D illustrate exemplary, non-limiting codes books generated by the methods described in Examples 1-4 and throughout the specification as filed. Each codebook is designed using different parameters, including trailing Hamming Distance, k value, and taper trail requirements.

[0217] Example 6; 578-Plex Chromosome Tracing Using the Present Technology QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0218] In one example implementation, a 578-plex chromosome-tracing experiment was performed using IMR-90 cells. The cells were cultured in flasks at 37°C prior to seeding and overnight culture in glass microscope slides, or “chips”. The chips were configured with a leur connector for delivery of required solutons (IBIDI, p Slide I Leur). Once the cells were attached to glass cover slips, they were fixed in dilute paraformaldehyde (PFA) and subsequently permeabilized in phosphate buffered TX-100. Following permeabilization, chromosome-specific hybridization probes were delivered to the chip, which was then placed in a thermal cycler to enable a brief, high temperature denaturation followed by slow, overnight hybridization at 43°C. Following this hybridization step, fluorescently labeled barcode-digit probes delivered and imaged, producing a multi -round set of barcode readouts for each chromosome-associated target.

[0219] In this implementation, the barcodes employed an “always-on” configuration in which all loci remain detectable in every imaging round, there were no null digits and a trailing Hamming distance of 4 was used. This configuration enables reliable barcode detection at substantially higher spatial density and facilitates the spatial positional corrections described herein. In addition, the barcodes were designed using a blue-shifted spectral allocation to minimize spectral overlap and to improve the robustness of digit discrimination in high-plex contexts.

[0220] FIGS. 20A through 20J illustrate representative results. FIG. 20A shows an overlay of chromosome 1 targets mapped within a single IMR-90 cell. All targets associated with this chromosome were designed using an always-on, blue-shifted, Hamming-optimized barcode scheme, enabling accurate decoding at an increased density that is beneficial for this 578-plex experiment. The chromosome was traced based on the order of the targets along the chromosome. FIG. 20B shows the additional overlay of chromosome 5 within the same cell. FIG. 20C shows the combined overlay of chromosomes 1, 5, and 7. FIG. 20D shows the further addition of chromosome 12; FIG. 20E depicts chromosome 16; and FIG. 20F shows chromosome 19. FIG. 201 illustrates the final combined overlay in which all chromosomes targeted in the 578-plex panel are simultaneously represented.

[0221] This example demonstrates that the present technology enables high-plex chromosome tracing with sufficient spatial resolution, signal density, and decoding robustness to map hundreds of genomic loci within single cells across multiple imaging rounds.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0222] Example 7; Identification of Copy Gains and Structural Variants in ER+ Breast Cancer Cell Line MCF-7

[0223] In this example implementation, after multiple rounds of imaging to detect the different digits of the barcodes across multiple fields of view, nuclear segmentation was performed using an imaging round that included DAPI staining. Segmentation may be carried out using any suitable technique known in the art, including software such as CellPose (cellpose.org) Mesmer (Greenwald, N.F., Miller, G., Moen, E. el al. Whole-cell segmentation of tissue images with human-level performance using large-scale data annotation and deep learning. Nat Biotechnol 40, 555-565 (2022). doi.org / 10.1038 / s41587-021-01094-0), or via image processing techniques such as watershed segmentation. Once the cells were segmented, barcodes were constructed and spatially aligned using the error minimization techniques described earlier in this patent. After the barcodes were computed for each cell, copy number variation (CNV) for the targeted loci was determined. For each target, the total number of detected counts was measured in each cell and an average number of counts per cell was computed. In a normal healthy diploid cell, each target is expected to occur twice, corresponding to the two copies of each chromosome. For the 578 plex panel described herein, CNV values were computed for all targets based on the average number of copies detected per cell.

[0224] FIG. 22A presents a bar graph indicating the average number of copies detected for each target, spanning chromosome 1 through chromosome X. FIG. 22B presents similar information in the form of an ideogram, illustrating the spatial positions of the target loci along each chromosome. Colors in the ideogram indicate copy number variation, with red corresponding to zero copies, yellow corresponding to two copies, and colors progressing toward green for higher copy number levels, with the maximum value determined by the copy number distribution observed for the particular cell line. FIG. 22C shows a dot plot representing chromosome wide copy number values, where each dot is positioned at the measured copy number for that chromosome and the dot size corresponds to the spread of the distribution across the cell population. In a healthy diploid sample, all dots would be circles centered at copy number 2 for all chromosomes.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0225] FIG. 22D shows a circos plot comparing the copy number variation results obtained by array based comparative genomic hybridization (array CGH) and by the present 578 plex imaging approach. Array CGH is a widely used and accepted technique for identifying genomic gains due to duplications or amplifications, as well as losses due to deletions. In the circos representation, the outer ring shows the chromosomes in ideogram form with cytogenetic banding, and centromeres are indicated by radial tick marks. The middle ring shows the results obtained by array CGH, and the inner ring shows the results obtained using the 578 plex panel applied to the MCF 7 cell line. The 578 plex results generally match the array CGH profile, showing gains, for example, on chromosomes 8, 17, and 20, as well as losses on chromosomes 8, 21, and 18.

[0226] FIG. 23 shows an all targets by all targets pairwise minimum distance heat map for the targets of the 578 plex panel applied to MCF 7 cells. Each matrix entry represents the median pairwise distance between a given pair of targets across all analyzed cells. The color scale (at right) indicates relative proximity, with darker colors denoting targets that reside closer together in three dimensional space and lighter colors indicating targets that are farther apart. The heat map is prepared by computing the pairwise minimum distances across cells and then aggregating the median for each target pair. A subset of targets exhibits elevated copy number values consistent with gains observed in the CNV analysis; this correspondence is shown by the middle inset in FIG.23, which reproduces the relevant portion of the copy number variation graph from FIG. 22A. On the right of FIG. 23, an overlay of chromosome 20 within a representative cell shows three copies of chromosome 20. Targets 363, 364, and 365 — which exhibit high copy number gains — are indicated within the cell, clearly showing copy numbers greater than two. In addition, a cluster of targets 363 and 364 is observed spatially displaced from the whole chromosome copies, indicating the presence of extrachromosomal DNA (ecDNA).

[0227] FIG. 24A highlights L shaped regions on the pairwise minimum distance heat map that indicate translocations between painted targets on chromosome 3 and chromosome 6. The left portion of FIG. 24A shows the full heat map with the translocation associated regions marked. The middle of FIG. 24A shows a magnified view of a highlighted region and the vertical components corresponding to targets 165 and 164, which are assigned to chromosome 6 but exhibit spatial proximity consistent with relocation to chromosome 3. The right of FIG. 24A shows aQB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158representative cell with overlays of chromosomes 3 and 6; targets 164 and 165 are circled on or in very close proximity to chromosome 3, in agreement with the heat map indication.

[0228] FIG. 24B continues this analysis for targets 88 and 89, which reside on chromosome 3 and exhibit close proximity to chromosome 6. The left panel of FIG. 24B shows the heat map with the relevant region; the middle panel presents a magnified view highlighting the L shaped signature consistent with interchromosomal translocation; and the right panel shows overlays demonstrating that targets 88 and 89 are located both in their native positions on chromosome 3 and in close proximity to chromosome 6. Translocations of this type have been associated with dysregulation of G protein-coupled receptor (GPCR) signaling. These results demonstrate that the present multiplexed, always on barcoded imaging approach can detect copy number alterations, extrachromosomal DNA, and inter-chromosomal translocations at single cell resolution within a high plex experimental framework.

[0229] Example 8. Identification of Extrachromosomal Aberrations in Breast Cancer Cell Line HC1954

[0230] This Example demonstrates the detection of extrachromosomal DNA (ecDNA) using the disclosed technology.

[0231] FIG. 31 shows the results of applying the 419 plex panel to the HCC1954 cell line. The all by all pairwise minimum distance heat map reveals structural variants and translocations present in HCC1954. A prominent dark band associated with chromosome 8 indicates a region of high copy number variation and suggests the presence of extrachromosomal DNA (ecDNA). This elevated copy number signal is also reflected in the copy number variation graph shown immediately above the heat map, which highlights high values for chromosome 8 target 203, as well as targets T118 and T142.

[0232] The blue shaded region in FIG. 31 corresponds to a selected portion of the heat map and its associated copy number graph. In this magnified representation, rather than showing only the median spatial distance values between targets, the full distribution of distances is shown for target 203 with respect to targets located on chromosomes 3, 4, 5, and 6. The blue gray band represents the overall distribution of distances, while the dark blue line indicates the median value. QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158A pronounced dip in the distribution is observed at chromosome 5 for targets T118 and T142, where the predominant interaction distance falls below approximately one micrometer, indicating a structural relationship between these loci and target 203.

[0233] A table beneath the figure summarizes the positional behavior of target 203 across the dataset. Out of 10,187 total observations involving target 203, 8,956 observations (approximately 88%) show target 203 at a distance greater than one micrometer from its nearest neighbor on chromosome 8 (target 202), consistent with extrachromosomal localization. Further analysis of this subset shows that, among the 8,956 instances, target 203 is located within 500 nm of target 118 on 300 occasions. A further refinement of this subpopulation reveals that, in 273 of these cases, target 118 is positioned more than one micrometer away from its adjacent chromosomal target (target 119), confirming that the co localized signals from targets 203 and 118 originate outside of the chromosome territory. These results demonstrate that the method is highly sensitive for detecting ecDNA involving targets 203 and 118.

[0234] At the right of FIG. 32, an image of a representative cell shows overlays of chromosomes and chromosome 5. Targets 203 and 118 appear segregated from their respective chromosome territories and clustered together, as indicated by the circled region. This spatial configuration is consistent with the ecDNA inference derived from the heat map and distance distribution analyses. This interaction occurs on 48% of the cells.

[0235] FIG. 38 further highlights the positional relationships among targets T118, T142, and T203 in the 419 plex analysis. On the left, two highlighted regions on the pairwise minimum distance heat map show the interactions between targetsT118 and T203, and between targetsT142 and T203. The intersection of these highlighted regions identifies the loci at which these targets exhibit close spatial proximity, indicating structural associations between chromosome 5 loci (T118 and T142) and the chromosome 8 locus (T203).

[0236] On the right of FIG. 38 is an image of representative cells with overlays showing the spatial positions of targets T118 and T142, displayed in light and dark blue, respectively, together with target T203. The circled regions highlight instances in which targets T118 and T142 appear in close proximity to T203 both extrachromosomally and within the broader chromosome 8 territory. These visualizations confirm the spatial relationships inferred from the heat map analysis QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158and demonstrate that targets T118, T142, and T203 can co localize in both chromosomal and extrachromosomal contexts.

[0237] Further analysis determined that 7-8% of cells have at least one ec203 / ecll8 chromosome 4 or chromosome 2 rearrangement and that 17% and 0.5% of such cells have ec203-ecl 18 less than 1 gm and 500 nm, respectively from chromosome 2 and chromosome 4 (data not shown).

[0238] Example 9. Puncta Localization, Decoding, and Alignment.

[0239] FIGS. 33A-33B illustrate an exemplary implementation for localizing, decoding, and aligning of puncta to generate genomic loci positions. FIG. 33A provides an exemplary workflow, while FIG. 36B. shows an exemplary alignment model (XYZ by round by probe). The position of each puncta is translated in 3D by an offset specific to each round and probe combination

[0240] The top row of FIG. 34 shows uncorrected Z position distribution of errors for individual barcode digits, shown for each imaging wavelength and plotted separately for rounds 0 through 5. Each plot displays the mean Z offset of that digit relative to the global centroid defined by the full barcode, along with the standard deviation representing the variability of that offset across all instances in the FOV. These plots illustrate the raw per wavelength axial shifts and the magnitude of digit to digit dispersion prior to correction. Bottom row: Corresponding distributions after applying Z offset correction. Following correction, the digit means are aligned near 0 nm (typically within <10 nm), and the distributions become narrower, demonstrating improved axial registration and reduced digit specific variability across rounds. Similarly, there are X and Y distributions and corresponding corrections.

[0241] FIG. 35 illustrates the error reduction across progressively more complex alignment models. The X-axis lists alignment models in order of increasing complexity, while the right Y-axis indicates the number of parameters required for each model. The left Y-axis shows residual registration error (pm). Three curves (Fx, Fy, F_z) represent the improvement in X, Y, and Z alignment accuracy as model complexity increases. Models shown include: (1) uncorrected, (2) standard chromatic correction, (3) Z correction by round by probe, (4) XYZ correction by round QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158by probe, (5) XYZ by round by probe with Z shrinkage by round, (6) XYZ by round by probe with Z-shrinkage and XYZ rotation by round (7) XYZ by round by probe with Z-shrinkage and XYZ rotation by round and second order polynomial to account for cellular swelling or shrinkage. Increasing model sophistication yields progressively lower residual errors, particularly in Z.

[0242] FIG. 36 illustrates a three-dimensional projection of a cell in which all digits of a barcode are illuminated. Due to shrinkage of the cell over successive imaging rounds, the digit positions do not appear as spatially coincident points but instead form radial distributions in the XY projection. In the uncorrected projection, the first digit appears at a larger radial distance, and subsequent digits (second through sixth) appear progressively closer to the cell center, indicating cumulative shrinkage. Similar shrinkage patterns are also visible in the XZ and YZ projections.

[0243] An updated set of projections is shown to the right, depicting the XY, XZ, and YZ views after application of the described shrinkage and scaling corrections. Following correction, the previously smeared distributions collapse into substantially point-like localizations, with markedly reduced spread in the XY projection and increased compactness in the XZ and YZ projections.

[0244] To the right of the corrected projections, graphs are provided that show the distribution of per-round scaling factors for the cells within the field of view in the XY and Z dimensions. These scaling factors exhibit values greater than one in round 1 and decrease monotonically through round 6, indicating progressive cell shrinkage across rounds.

[0245] The correction results shown in FIG. 36 were achieved using a higher-complexity alignment model that includes: (i) XYZ correction by round by probe, (ii) Z-shrinkage terms estimated on a per-round basis, (iii) rotational error correction applied per round, and (iv) a second-order polynomial scaling function to model nonlinear expansion and contraction of the sample. The use of these higher-order components enables accurate compensation for round-dependent geometric distortions. Such corrections are facilitated by the use of an “always-on” barcode configuration, which provides a temporally consistent sequence of digit observations sufficient to associate all digits with a given barcode; in contrast, barcodes with multiple null digits would provide insufficient continuity to estimate these deformation parameters reliably.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0246] FIG. 37 illustrates graphical and tabular summaries of residual positional error following application of the described correction procedures. A top row of density plots presents residual error in three orthogonal projections (XY, XZ, YZ), wherein color denotes wavelength associated with a digit. In each plot, the central location represents the mean residual error, and the spread along respective axes represents the standard deviation in those axes. A middle row provides corresponding projections in which color denotes imaging round (rounds 1-6), indicating the per-round mean residual error after correction. A bottom row comprises tables reporting, for rounds 1-6, the mean and standard deviation of the residual error in XY and in Z.

[0247] References:

[0248] 1. Chen, et al., Science. 2015 April 24;348(6233); author manuscript available in PMC 2015 November 28.

[0249] 2. Su, et al., Cell. 2020 September 17; 182(6): 1641 - 1659; author manuscript available in PMC 2021 September 17.

[0250] It will be readily apparent to one skilled in the art that varying substitutions and modifications may be made to the invention disclosed herein without departing from the scope and spirit of the invention. The invention illustratively described herein suitably may be practiced in the absence of any element or elements, limitation or limitations which is not specifically disclosed herein. The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention. Thus, it should be understood that although the present invention has been illustrated by specific embodiments and optional features, modification and / or variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention. Exemplary variants include but are not limited to the following embodiments:

[0251] Embodiment 39. A computer-implemented method for fiducial-free alignment of multi-round, multi-color fluorescence imaging data, comprising detecting puncta in image data acquired over a plurality of imaging rounds and a plurality of probes, each punctum having QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158three-dimensional coordinates and round / probe identifiers; clustering the puncta based on spatial proximity to generate candidate groups; decoding barcodes for respective genomic loci from the candidate groups, wherein decoding assigns puncta to barcodes based at least on spatial proximity and puncta intensity; computing, for each barcode, a dispersion metric characterizing a spatial spread of assigned puncta; fitting parameters of an alignment model by minimizing an objective function of the dispersion metric across the barcodes; and applying the alignment model to correct the puncta coordinates.

[0252] Embodiment 40. The method of embodiment 39, further comprising re-decoding the barcodes using the corrected puncta coordinates.

[0253] Embodiment 4E The method of embodiment 39, wherein detecting puncta includes localizing puncta in three dimensions and associating each punctum with a cell identifier.

[0254] Embodiment 42. The method of embodiment 39, wherein clustering uses a three-dimensional distance threshold defining a barcode radius.

[0255] Embodiment 43. The method of embodiment 39, wherein decoding employs a max-flow min-cost algorithm with costs weighted by puncta intensity and spatial proximity.

[0256] Embodiment 44. The method of embodiment 39, wherein the dispersion metric comprises a variance or standard deviation of distances from puncta to a barcode centroid.

[0257] Embodiment 45. The method of embodiment 39, wherein the alignment model comprises round- and probe-specific translation parameters in X, Y, and Z.

[0258] Embodiment 46. The method of embodiment 45, wherein the alignment model further comprises per-cell scaling parameters modeling sample expansion or shrinkage across imaging rounds.

[0259] Embodiment 47. The method of embodiment 39, wherein the alignment model further includes parameters representing sample rotation and one or more polynomial distortion terms.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0260] Embodiment 48. The method of embodiment 39, wherein fitting comprises nonlinear least-squares minimization with outlier rejection of low-confidence barcodes.

[0261] Embodiment 49. The method of embodiment 39, further comprising iteratively repeating clustering, decoding, fitting, and correction with a progressively decreasing barcode radius.

[0262] Embodiment 50. The method of embodiment 39, wherein applying the alignment model reduces Z-axis dispersion attributable to round-dependent drift.

[0263] Embodiment 51. The method of embodiment 39, wherein applying the alignment model reduces XY-axis dispersion attributable to probe-specific chromatic offsets.

[0264] Embodiment 52. The method of embodiment 39, further comprising performing optional coarse alignment prior to decoding using DAPI-based image registration or pre-calibrated chromatic correction using fluorescence beads.

[0265] Embodiment 53. The method of embodiment 39, wherein the barcodes are always-on barcodes providing intrinsically present reference points in each imaging round.

[0266] Embodiment 54. The method of embodiment 40, wherein re-decoding improves barcode scoring and reduces false positives without re-localizing raw images.

[0267] Embodiment 55. A system comprising a microscope configured to acquire multi-round, multi-color fluorescence image data in three dimensions; one or more processors; and memory storing instructions that, when executed by the processors, cause the processors to detect puncta in three dimensions in the image data, cluster the puncta based on spatial proximity, decode barcodes from the clustered puncta, compute a dispersion metric for each barcode, fit an alignment model by minimizing the dispersion metric across barcodes, and apply the alignment model to correct puncta coordinates.

[0268] Embodiment 56. The system of embodiment 55, wherein the processors perform decoding using max-flow min-cost optimization.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0269] Embodiment 57. The system of embodiment 55, wherein the processors fit an alignment model having XYZ-by-round-by-probe parameters.

[0270] Embodiment 58. The system of embodiment 55, wherein the processors apply per-cell Z-scale correction across imaging rounds.

[0271] Embodiment 59. The system of embodiment 55, wherein the processors iteratively re-decode barcodes after applying corrections and reduce the clustering radius in successive iterations.

[0272] Embodiment 60. The system of embodiment 55, wherein coarse alignment is optionally performed using DAPI-based registration or bead-based chromatic calibration.

[0273] Embodiment 61. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the processors to detect puncta in three dimensions in multi -round, multi-color fluorescence image data, cluster the puncta based on spatial proximity, decode barcodes from the clustered puncta, compute a dispersion metric for each barcode, fit parameters of an alignment model by minimizing the dispersion metric across the barcodes, and apply the alignment model to correct puncta coordinates.

[0274] Embodiment 62. The medium of embodiment 61, wherein the alignment model includes at least one of: round-specific translation, probe-specific translation, round*probe interaction, per-cell Z-scale, XY-scale, rotation, or polynomial distortion.

[0275] Embodiment 63. The medium of embodiment 61, wherein the instructions reduce the clustering radius used for barcode decoding after each iteration.

[0276] Embodiment 64. The medium of embodiment 61, wherein decoding uses a misalignment- weighted max-flow min-cost solver.

[0277] Embodiment 65. The medium of embodiment 61, wherein fitting uses nonlinear least-squares with robust weighting and outlier removal of barcodes showing dispersion above a threshold.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158

[0278] Embodiment 66. An imaging apparatus comprising an objective lens, a detector configured to obtain multi-round, multi-color images, and an electronic processing unit programmed to decode barcodes from clustered puncta, fit and apply an alignment model reducing barcode dispersion, and re-decode using corrected puncta coordinates.

[0279] Embodiment 67. The apparatus of embodiment 66, wherein the processing unit applies XYZ-by-round-by-probe translation and per-cell Z-scale correction to compensate drift and sample deformation.

[0280] Citations to a number of patent and non-patent references may be made herein. Any cited references are incorporated by reference herein in their entireties. In the event that there is an inconsistency between a definition of a term in the specification as compared to a definition of the term in a cited reference, the term should be interpreted based on the definition in the specification.QB\121892.00158\101046255.1

Claims

Atty. Dkt. No. 121892.00158CLAIMSWe Claim:

1. A method for in situ positional mapping of a plurality of nucleic acid targets in a biological sample, the method comprising:detecting a pool of target barcodes linked to the plurality of nucleic acid targets, wherein each target barcode comprises N digits, wherein each position of each target barcode is represented by a color signal generated from a detectable label, and wherein no position of any barcode lacks a corresponding color signal;and wherein detecting comprises:(a) labeling the barcodes linked to the plurality of nucleic acid targets with a first pool of first detection probes, wherein each first detection probe comprises a corresponding first detectable label and a barcode hybridization region, wherein different first detection probes are correlated to different targets or sets of targets of the plurality;(b) imaging the detectable labels of (a), wherein the first detection probes hybridize to a barcode digit 1 position and provide the corresponding first color signal as digit 1 of the target barcode to which the first detection probe is hybridized ;(c) removing the first color signal of (b);(d) labeling the barcodes linked to the plurality of nucleic acid targets with a second pool of second detection probes, wherein each second detection probe comprises a corresponding second detectable label and a barcode hybridization region, wherein different second detection probes are correlated to different targets or sets of targets of the plurality;(e) imaging the detectable labels of (d) wherein the second detection probes hybridize to a barcode digit 2 position and provide the corresponding second color signal as digit 2 of the target barcode to which the second detection probe is hybridized;(f) removing the second color signal of (e);QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158(g) repeating steps (d)-(f) with an Nth pool of detection probes, wherein each Nth detection probe in each Nth pool from the second pool to the Nth pool (i) comprises a corresponding detectable label, (ii) is correlated to different targets or sets of targets of the plurality; and (iii) provides a corresponding color signal as an Nth digit of the target barcode;wherein, in combination, the N pools of probes comprise at least 3 different colors.

2. The method of claim 1, wherein at least 50% of the detectable labels in each pool have a color signal comprising a wavelength below the median wavelength of the color signals of that pool.

3. The method of any one of the previous claims, wherein the barcodes of the pool comprise a global hamming distance (gHD) of at least 2.

4. The method of claim 3, wherein the global hamming distance (gHD) is 3 or more.

5. The method of claim 3, wherein the global hamming distance (gHD) is equal to N.

6. The method of any one of the previous claims, wherein a territory set of target barcodes comprise a trailing hamming distance (trHD), wherein the trHD is at least one unit greater than the global hamming distance.

7. The method of claim 6, wherein a territory set of target barcodes is defined as sets of two nearest neighbors (k = 2).

8. The method of claim 6, wherein a territory set of target barcodes is defined as sets of three (k = 3), four (k = 4), five (k = 5), or six (k = 6) neighboring barcodes.

9. The method of any one of the previous claims, wherein a territory set of target barcodes comprises a subset of barcodes, the subset comprising a second trailing Hamming distance (trHD_2).

10. The method of claim 9, wherein the number of neighboring target barcodes with trHD_2 comprises two barcodes (k = 2), three (k = 3), four (k = 4), five (k = 5), or six (k = 6) barcodes.

11. The method of claim 9 or 10, wherein the trHD_2 is equal to N.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.0015812. The method of any one of the previous claims wherein the target barcodes comprise at least one gHD, at least one trHD, and at least one taper trail hamming distance (ttHD).

13. The method of claim 12, wherein the ttHD is between the trHD and the gHD.

14. The method of any of the previous claims, wherein the gHD is 3, and ttHD is greater than or equal to N-2, wherein N is at least 6.

15. The method of any one of the previous claims, wherein the territory set of target barcodes encompasses target barcodes on the same chromosome.

16. The method of any one of the previous claims, wherein territory set of target barcodes encompasses target barcodes on the same chromosome arm.

17. The method of any one of the previous claims further comprising (h) positional mapping of the plurality of nucleic acid targets without the use of extrinsic fiducials.

18. The method of any one of the previous claims, wherein each target barcode further comprises a prefix barcode, wherein the prefix barcode digits are independent of the target barcode digits.

19. The method of claim 18, wherein the prefix barcode comprises two digits.

20. The method of any one of the previous claims, wherein at least one prefix barcode is specific to a chromosome.

21. The method of any one of the previous claims, wherein at least one prefix barcode is specific to a chromosome territory.

22. The method of any one of the previous claims, wherein N is 6.

23. A method for designing a set of molecular barcodes comprising colored fluorophores for detection of a plurality of nucleic acid targets in a biological sample, wherein each position of each molecular barcode is represented by a color signal generated from a detectable label, and wherein no position of any barcode lacks a corresponding color signal, the method comprising:QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158(a) with the use of a programmable processor operably connected with tangible, non-transitory storage medium containing program code thereon, accessing data representing:(i) a color input value, wherein the color input value defines the number of different colors for use in the set of barcodes, wherein the color input value is at least 3;(ii) a barcode digit input value, wherein the barcode position input value defines the number of digits of each barcode, wherein the position input value is N, and wherein each position of each barcode is represented by a color and no position of any barcode lacks a corresponding color signal;(iii) a target input value, comprising the total number of targets (loci);(b) selecting, based on the accessing data of (a), the set of molecular barcodes, wherein:(i) the global hamming distance (gHD) for all barcodes within the set is N-2;wherein the selected barcodes enable three-dimensional mapping of the plurality of nucleic acid targets.

24. The method of claim 23, further comprising at step (a) accessing data comprising one or more of(iv) a territory input value, comprising the total number of territories comprising targets in the biological sample;(v) territory barcode set input values (k), each value defining the number of barcodes in a territory ("territory set of barcodes").

25. The method of claim 24 (b), wherein:(ii) the territory set of barcodes comprises a trailing hamming distance (trHD) for each territory set of barcodes; and / or(iii) a taper trail Hamming distance (ttHD) for barcodes within a territory set.

26. The method of claim 23, wherein the territory input value is an integer greater than 2.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.0015827. A method for in situ positional mapping of a plurality of nucleic acid targets in a biological sample, the method comprising:detecting a pool of target barcodes linked to the plurality of nucleic acid targets, wherein each target barcode comprises N positions, and wherein detecting comprises:(a) labeling the barcodes linked to the plurality of nucleic acid targets with a first pool of first detection probes, wherein each first detection probe comprises a corresponding first detectable label and a barcode hybridization region, wherein different first detection probes are correlated to different targets or sets of targets of the plurality;(b) imaging the detectable labels of (a), wherein the first detection probes hybridizes to a barcode digit 1 position and provide the corresponding first color signal as digit 1 of the target barcode to which the first detection probe is hybridized;(c) removing the first color signal of (b);(d) labeling the plurality of nucleic acid targets with a second pool of second detection probes, wherein each second detection probe comprises a corresponding second detectable label and a barcode hybridization region, wherein different second detection probes are corelated to different targets or sets of targets of the plurality;(e) imaging the detectable labels of (d) wherein the second detection probes hybridize to a barcode digit 2 position and provide the corresponding second color as position 2 of the target barcode to which the second detection probe is hybridized;(f) removing the second color signal of (e);(g) repeating steps (d)-(f) with an Nth pool of detection probes, wherein each Nth detection probe in each pool from the second pool to the Nth pool (i) comprises a corresponding detectable label, (ii) is correlated to different targets or set of targets of the plurality;wherein a territory set of target barcodes comprise a trailing Hamming distance (trHD) value that is greater than a taper trail Hamming distance (ttHD) value of a subset of target barcode in the same territory set.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.0015828. The method of claim 27, wherein the territory set of barcodes (k) is defined as sets of two nearest neighbors (k = 2).

29. The method of claim 27 or 28, wherein the territory set of target barcodes is defined as sets of three (k = 3), four (k = 4), five (k = 5), or six (k = 6) neighbors.

30. The method of any one of claims 27-29, wherein ttHD is greater than the trHD value.

31. The method of any one of claims 27-30, wherein the territory set of barcodes encompasses barcodes on the same chromosome.

32. The method of any one of claims 27-31, wherein the territory set of barcodes encompasses barcodes on the same chromosome arm.

33. The method of any one of claims 27-32, wherein the territory set of barcodes encompasses barcodes in three-dimensional chromosome territory.

34. A computer program product for devising a set of barcodes configured to detect nucleic acid targets in a biological sample, the computer program product comprising a computer usable tangible non-transitory storage medium having computer readable program code thereon, the computer readable program including at least one of the following:(34A) program code for forming a first specification matrix configured as a measure required for selection of a set of barcodes targeting a single chromosome panel, the first specification matrix having a first dimension m and a second dimension equal to the first dimension, wherein:(i) each main diagonal element of the first specification matrix is assigned a value of zero,(ii) each element of a first group of immediately neighboring each other elements of the matrix is assigned a value of a chosen trailing Hamming distance, wherein the size of the first group is no larger than a pre-determined number k < m of said immediately neighboring each other elements and the first group is immediately adjacent to every main diagonal element along the first dimension and / or the second dimension,andQB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158(iii) each element of a second group of immediately neighboring each other elements of the first specification matrix along each of the first and second dimensions is assigned a value of a chosen minimum global Hamming distance,wherein said second group is either immediately adjacent to an end element of the first group that is the most distant to a corresponding main diagonal element along the first dimension and along a second dimension orwherein said second group is separated from the end element of the first group along the first dimension and along the second dimension by a corresponding taper group of elements that includes at least one matrix element assigned a value smaller than the value of the chosen trailing distance and larger than the value of the chosen minimum global Hamming distance, the taper group being immediately adjacent to both the first group and the second group;and(34B) program code for forming a second specification matrix configured as a measure required for selection of a set of barcodes targeting a multiple chromosome panel, the second specification matrix having a corresponding first dimension I and a corresponding second dimension equal to the first dimension, wherein:(a) each main diagonal element of the second specification matrix is assigned a value of zero,(b) each of multiple portions of the second specification matrix respectively correspond to each of the multiple chromosomes of the multiple chromosome panel, wherein a respective main diagonal of each of the multiple portions is a part of the main diagonal of the second specification matrix and wherein each of said respective main diagonals of the multiple portions is immediately adjacent to at least one other of said respective main diagonals of the multiple portions,(c) each element of multiple third groups of immediately neighboring each other elements of each of the multiple portions is assigned a value of a chosen trailing Hamming distance,wherein corresponding sizes of the multiple third groups are no larger than a dimension of the corresponding one of the multiple portions,QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158wherein two or more of said multiple third groups are immediately adjacent to every respective element of the main diagonal of the respective one of the multiple portions along the first dimension and / or the second dimension of the second specification matrix,(d) each remaining element of the multiple third groups is assigned a value that is smaller than the value of the chosen trailing distance and larger than a value of a chosen minimum global Hamming distance,and(e) each element of a fourth group of elements that includes immediately neighboring each other elements of the second specification matrix outside of each of the multiple portions is assigned a value of a chosen minimum global Hamming distance.

35. A computer program code according to claim 34, wherein the program code for forming a second specification matrix is configured to form the second specification matrix in which corresponding sizes of said two or more of the multiple third groups are not equal to one another.

36. The method of any one of claims 1-22, wherein the barcode digits are used as fiducials to align images round to round.

37. A method for in situ positional mapping of a plurality of nucleic acid targets in a biological sample, the method comprising:detecting a pool of target barcodes linked to the plurality of nucleic acid targets, wherein each target barcode comprises N digits;and wherein detecting comprises:(a) labeling the target barcodes linked to the plurality of nucleic acid targets with a first pool of first detection probes, wherein each first detection probe comprises a corresponding first detectable label and a barcode hybridization region, wherein different first detection probes are correlated to different targets or sets of targets of the plurality;QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158(b) imaging the detectable labels of (a), wherein the first detection probes hybridizes to a barcode digit 1 position and provide the corresponding first color signal as digit 1 of the target barcode to which the detection probe is hybridized;(c) removing the first color signal of (b);(d) labeling the target barcodes linked to the plurality of nucleic acid targets with a second pool of second detection probes, wherein each second detection probe comprises a corresponding second detectable label and a barcode hybridization region, wherein different second detection probes are correlated to different targets or sets of targets of the plurality;(e) imaging the detectable labels of (d) wherein the second detection probes hybridize to a barcode digit 2 position and provide the corresponding second color signal as digit 2 of the target barcode to which the detection probe is hybridized;(f) removing the second color signal of (e);(g) repeating steps (d)-(f) with an Nth pool of detection probes, wherein each Nth detection probe in each Nth pool from the second pool to the Nth pool (i) comprises a corresponding detectable label and barcode hybridization region, (ii) is correlated to different targets or sets of targets of the plurality; and (iii) provides a corresponding color signal as an Nth digit of the target barcode to which the Nth detection probe is hybridized;wherein, in combination, the N pools of probes comprise at least 3 different colors,and wherein the barcode digits from each round are used a fiducials to align images round to round.

38. The method of claim 37, wherein each position of each target barcode is represented by a color signal generated from a detectable label, and wherein no position of any barcode lacks a corresponding color signal.

39. A computer-implemented method for fiducial-free alignment of multi-round, multi-color fluorescence imaging data, comprising detecting puncta in image data acquired over a plurality of imaging rounds and a plurality of probes, each punctum having three-dimensional coordinates and round / probe identifiers; clustering the puncta based on spatial proximity to generate candidate QB\121892.00158\101046255.1Atty. Dkt. No. 121892.00158groups; decoding barcodes for respective genomic loci from the candidate groups, wherein decoding assigns puncta to barcodes based at least on spatial proximity and puncta intensity; computing, for each barcode, a dispersion metric characterizing a spatial spread of assigned puncta; fitting parameters of an alignment model by minimizing an objective function of the dispersion metric across the barcodes; and applying the alignment model to correct the puncta coordinates.

40. The method of claim 39, further comprising re-decoding the barcodes using the corrected puncta coordinates.

41. The method of claim 39, wherein detecting puncta includes localizing puncta in three dimensions and associating each punctum with a cell identifier.

42. The method of claim 39, wherein clustering uses a three-dimensional distance threshold defining a barcode radius.

43. The method of claim 39, wherein decoding employs a max-flow min-cost algorithm with costs weighted by puncta intensity and spatial proximity.

44. The method of claim 39, wherein the dispersion metric comprises a variance or standard deviation of distances from puncta to a barcode centroid.

45. The method of claim 39, wherein the alignment model comprises round- and probe-specific translation parameters in X, Y, and Z.

46. The method of claim 45, wherein the alignment model further comprises per-cell scaling parameters modeling sample expansion or shrinkage across imaging rounds.

47. The method of claim 39, wherein the alignment model further includes parameters representing sample rotation and one or more polynomial distortion terms.

48. The method of claim 39, wherein fitting comprises nonlinear least-squares minimization with outlier rejection of low-confidence barcodes.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.0015849. The method of claim 39, further comprising iteratively repeating clustering, decoding, fitting, and correction with a progressively decreasing barcode radius.

50. The method of claim 39, wherein applying the alignment model reduces Z-axis dispersion attributable to round-dependent drift.

51. The method of claim 39, wherein applying the alignment model reduces XY-axis dispersion attributable to probe-specific chromatic offsets.

52. The method of claim 39, further comprising performing optional coarse alignment prior to decoding using DAPI-based image registration or pre-calibrated chromatic correction using fluorescence beads.

53. The method of claim 39, wherein the barcodes are always-on barcodes providing intrinsically present reference points in each imaging round.

54. The method of claim 40, wherein re-decoding improves barcode scoring and reduces false positives without re-localizing raw images.

55. A system comprising a microscope configured to acquire multi-round, multi-color fluorescence image data in three dimensions; one or more processors; and memory storing instructions that, when executed by the processors, cause the processors to detect puncta in three dimensions in the image data, cluster the puncta based on spatial proximity, decode barcodes from the clustered puncta, compute a dispersion metric for each barcode, fit an alignment model by minimizing the dispersion metric across barcodes, and apply the alignment model to correct puncta coordinates.

56. The system of claim 55, wherein the processors perform decoding using max -flow min-cost optimization.

57. The system of claim 55, wherein the processors fit an alignment model having XYZ-by-round-by-probe parameters.

58. The system of claim 55, wherein the processors apply per-cell Z-scale correction across imaging rounds.QB\121892.00158\101046255.1Atty. Dkt. No. 121892.0015859. The system of claim 55, wherein the processors iteratively re-decode barcodes after applying corrections and reduce the clustering radius in successive iterations.

60. The system of claim 55, wherein coarse alignment is optionally performed using DAPI-based registration or bead-based chromatic calibration.

61. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the processors to detect puncta in three dimensions in multi-round, multi-color fluorescence image data, cluster the puncta based on spatial proximity, decode barcodes from the clustered puncta, compute a dispersion metric for each barcode, fit parameters of an alignment model by minimizing the dispersion metric across the barcodes, and apply the alignment model to correct puncta coordinates.

62. The medium of claim 61, wherein the alignment model includes at least one of: round-specific translation, probe-specific translation, roundxprobe interaction, per-cell Z-scale, XY-scale, rotation, or polynomial distortion.

63. The medium of claim 61, wherein the instructions reduce the clustering radius used for barcode decoding after each iteration.

64. The medium of claim 61, wherein decoding uses a misalignment- weighted max-flow min-cost solver.

65. The medium of claim 61, wherein fitting uses nonlinear least-squares with robust weighting and outlier removal of barcodes showing dispersion above a threshold.

66. An imaging apparatus comprising an objective lens, a detector configured to obtain multi-round, multi-color images, and an electronic processing unit programmed to decode barcodes from clustered puncta, fit and apply an alignment model reducing barcode dispersion, and re-decode using corrected puncta coordinates.

67. The apparatus of claim 66, wherein the processing unit applies XYZ-by-round-by-probe translation and per-cell Z-scale correction to compensate drift and sample deformation.QB\121892.00158\101046255.1