Jitter-corrected image analysis

The image detection device corrects for stage jitter by dividing images into subregions and performing affine transformations to align fiducial locations, improving data quality and accuracy in nucleotide sequencing.

JP2026507380APending Publication Date: 2026-03-04ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024576791
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-27
Filing Date
2024-01-05
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

High-frequency motion (jitter) of the stage in line scanning systems induces nonlinear perturbations in sequencing images, affecting the accuracy of spot locations corresponding to DNA clusters in flow cells, leading to data quality issues.

Method used

An image detection device corrects for high-frequency jitter by dividing images into subregions, performing affine transformations on each subregion to align fiducial locations, and using pixel padding to accommodate larger areas of reference fiducials, thereby determining accurate spot positions despite jitter.

Benefits of technology

The solution effectively aligns fiducial and spot locations, improving data quality and reducing errors caused by stage jitter, enhancing the accuracy of nucleotide sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026507380000001_ABST
    Figure 2026507380000001_ABST
Patent Text Reader

Abstract

Systems, methods, and devices are described herein. For example, a detection device may include a memory and at least one processor. The detection device may be configured to acquire an image including at least one feature and a plurality of fiducials. The plurality of fiducials may be arranged in a pattern. The detection device may be configured to determine a plurality of subregions of the image, each subregion including a subset of the fiducials included in the image. The detection device may be configured to perform a geometric transformation on each subregion to generate a respective local transformation associated with each subregion. The detection device may be configured to align the respective locations of the fiducials included in the image based on the respective local transformation associated with each subregion. The size of each subregion may be selected such that each subregion is substantially invariant to stage jitter.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 441,606, filed January 27, 2023, which is incorporated herein by reference in its entirety. [Background technology]

[0002] In recent years, biotechnology companies and research institutions have improved hardware and software platforms for sequencing nucleotide bases (or whole genomes) and identifying variant calls for nucleotide bases that differ from reference bases in a reference genome. Many of these platforms utilize image detection systems to capture and resolve sequencing images of nucleotide bases, for example, through the use of line scanning and confocal control of imaging optics. However, high-frequency motion (i.e., jitter) of the stage on a line scanning system typically induces nonlinear perturbations of cluster locations in the sequencing image, thereby adversely affecting data quality. Stage jitter can perturb the location of spots corresponding to clusters of deoxyribonucleic acid (DNA) molecules in wells on a flow cell relative to the fiducials in the sequencing image. In addition, stage jitter can also perturb the fiducials. Thus, spot locations may be inaccurate. For example, Figure 9 is an exemplary heat map illustrating stage jitter present in both the X and Y directions of a tile 900 during a conventional imaging cycle. The top plot 904a illustrates the motion of the fiducial 902 in the X direction 908, and region 906 of the heat map illustrates the motion of the fiducial 902 caused by stage jitter. The bottom plot 904b illustrates the motion of the fiducial 902 in the Y direction 912, and region 910 of the heat map illustrates the motion of the fiducial 902 caused by stage jitter. Summary of the Invention [Means for solving the problem]

[0003] Systems and methods for image recognition and / or image processing are described that account for and correct for errors introduced due to high-frequency jitter in an image capture device. For example, the image detection device can correct for high-frequency motion (i.e., jitter) of a stage, which leads to inconsistent movement of spot locations within an image. The spots may correspond to clusters of DNA molecules within a well of a flow cell. The image detection device can acquire images of an object (e.g., a flow cell) having multiple fiducials arranged in a pattern. An example fiducial pattern may include a set of four fiducials linearly arranged adjacent to another set of four fiducials. Another exemplary fiducial pattern may include a set of nine fiducials linearly arranged adjacent to another set of nine fiducials, with two sets of fiducials centered between the first and second sets of nine fiducials. The image detection device can generate the locations of the fiducials and decompose the image into subregions, each subregion containing a subset of the fiducials. The image detection device can configure the size of each subregion so that each subregion is substantially invariant to stage jitter. The image detection device can perform a transformation (i.e., an affine transformation) on each subregion to generate an arrangement of fiducials within each subregion (e.g., a matrix of the affine transformation performed on each subregion). Based on the generated arrangement of fiducials, the image detection device can determine the locations of the spots on the flow cell. That is, the image detection device can use each determined matrix to align the respective locations of the fiducials while taking jitter into account and determining the positioning of the spot locations in the image.

[0004] In another example, the sequencing system can acquire an image including at least one feature and a plurality of fiducials. The plurality of fiducials may be arranged in a pattern. The sequencing system can determine a plurality of subregions of the image. Each subregion may include a subset of the fiducials included in the image. The sequencing system can perform a geometric transformation on each subregion to generate a respective local transformation associated with each subregion. The sequencing system can align the respective locations of the fiducials included in the image based on the respective local transformation associated with each subregion. In one example, the size of each subregion is selected such that each subregion is substantially invariant to stage jitter. In one or more cases, the sequencing system can determine the respective locations of each of the fiducials in the image based on the determined locations of the reference fiducials. Further, the sequencing system can generate a subimage based on the acquired image. In one example, the subimage includes pixel padding to accommodate a larger area of ​​the reference fiducials included in the image. After generating the subimage based on the acquired image, the sequencing system can determine a correlation between the acquired image and the subimage. Further, the sequencing system can determine whether to adjust the locations of the reference fiducials based on the determined correlation. In one or more cases, each subset of criteria included in a sub-region may include at least three criteria. In one or more cases, the sub-regions may be linearly positioned relative to one another. In one or more cases, the geometric transformation includes an affine transformation.

[0005] In another example, the sequencing method can include acquiring an image including at least one feature and a plurality of fiducials. The plurality of fiducials may be arranged in a pattern. The sequencing method can include determining a plurality of subregions of the image. Each subregion includes a subset of the fiducials included in the image. The sequencing method can include performing a geometric transformation on each subregion to generate a respective local transformation associated with each subregion. The sequencing method can include aligning the locations of each of the fiducials included in the image based on the respective local transformation associated with each subregion. In one example, the size of each subregion is selected such that each subregion is substantially invariant to stage jitter. In one or more cases, the subregions are linearly arranged relative to one another. In one or more cases, each subset of fiducials included in a subregion includes at least three fiducials. In one or more cases, reference fiducials are located in the image, and the sequencing method can include determining the respective locations of each of the fiducials in the image based on the determined locations of the reference fiducials. In one or more cases, the sequencing method can include generating subimages based on the acquired image. In one example, the sub-image includes pixel padding with a larger area of ​​the reference fiducials and includes them within the image. In one or more cases, the sequencing method includes determining a correlation between the acquired image and the sub-image. In one or more cases, the sequencing method includes determining whether to adjust the location of the reference fiducials based on the determined correlation. In one example, the size of each sub-region is selected such that each sub-region is substantially invariant to stage jitter of about 200 Hz or less. In one example, at least one fiducial is associated with two adjacent sub-regions. In one example, the pattern of the plurality of fiducials includes a first set of linearly arranged fiducials adjacent to a second set of linearly arranged fiducials. Further, the first set of linearly arranged fiducials and the second set of linearly arranged fiducials extend substantially parallel to each other. In one example, the image includes a sequencing image. In one example, the feature locations are associated with well locations of a flow cell. In one example, each sub-region includes a respective feature.Furthermore, each local transform associated with the subregion containing each feature is used to determine the location of each feature.

[0006] In another example, a computer-readable medium may include computer-readable instructions that, when executed by a processor, cause the processor to perform a sequencing method. In one or more cases, the sequencing method includes acquiring an image including at least one feature and a plurality of fiducials. In one example, the plurality of fiducials may be arranged in a pattern. In one or more cases, the sequencing method includes determining a plurality of subregions of the image. In one example, each subregion includes a subset of the fiducials included in the image. In one or more cases, the sequencing method includes performing a geometric transformation on each subregion to generate a respective local transformation associated with each subregion. In one or more cases, the sequencing method includes aligning the respective locations of the fiducials included in the image based on the respective local transformation associated with each subregion. In one example, the size of each subregion is selected such that each subregion is substantially invariant to stage jitter. In one example, the subregions are linearly arranged relative to one another. In one example, each subset of fiducials included in a subregion includes at least three fiducials. In one example, reference fiducials are located in the image, and the locations of the reference fiducials are used to determine the respective locations of each of the fiducials in the image. In one or more cases, the sequencing method includes generating a sub-image based on the acquired image. In one example, the sub-image includes pixel padding to accommodate a larger area of ​​the reference fiducials included in the image. Further, in one or more cases, the sequencing method includes determining a correlation between the acquired image and the sub-image. Additionally, in one or more cases, the sequencing method includes determining whether to adjust the location of the reference fiducials based on the determined correlation. In one example, the size of each sub-region is selected such that each sub-region is substantially invariant to stage jitter of about 200 Hz or less. In one example, at least one fiducial is associated with two adjacent sub-regions. In one example, the pattern of the plurality of fiducials includes a first set of linearly arranged fiducials adjacent to a second set of linearly arranged fiducials. Furthermore, the first set of linearly arranged fiducials and the second set of linearly arranged fiducials extend substantially parallel to one another. In one example, the image includes a sequencing image.In one example, the feature locations are associated with flow cell well locations. In one example, each subregion contains a respective feature, and a respective local transform associated with the subregion containing the respective feature is used to determine the location of the respective feature. [Brief explanation of the drawings]

[0007] [Figure 1A] 1 illustrates an exemplary reference alignment of an exemplary sequencing image. [Figure 1B] 1 illustrates an exemplary reference alignment of an exemplary sequencing image. [Figure 1C] 1 illustrates an exemplary reference alignment of an exemplary sequencing image. [Figure 2A] 1 illustrates a schematic diagram of an exemplary system environment. [Figure 2B] 1 illustrates an exemplary image scanning system. [Figure 2C] 1 illustrates an exemplary line scan imaging system. [Figure 3] 1 is a flowchart illustrating an exemplary image calibration process. [Figure 4] 1 is a flowchart illustrating the execution of an exemplary imaging process to correct for jitter. [Figure 5A] 1 illustrates exemplary transformations of multiple exemplary sub-regions of an exemplary image. [Figure 5B] 1 illustrates exemplary transformations of multiple exemplary sub-regions of an exemplary image. [Figure 6A] 1 illustrates an exemplary imaging data subset. [Figure 6B] 1 illustrates an exemplary imaging data subset. [Figure 7A] 1 illustrates performance metrics of an exemplary imaging process for correcting jitter. [Figure 7B] 1 illustrates performance metrics of an exemplary imaging process for correcting jitter. [Figure 7C] 1 illustrates performance metrics of an exemplary imaging process for correcting jitter. [Figure 7D]1 illustrates performance metrics of an exemplary imaging process for correcting jitter. [Figure 8] FIG. 1 is a block diagram of an exemplary computing device. [Figure 9] 10 is an exemplary heat map illustrating jitter present in both the X and Y directions of a tile during a conventional imaging cycle. DETAILED DESCRIPTION OF THE INVENTION

[0008] The present disclosure is directed to processing image data and correcting image distortions during imaging of patterned arrays having multiple repeating spots or fiducials. Processing the patterned array generates image data (or any other form of detection output of sites on the array) of an analysis array, such as those used for analyzing biological samples. Biological sample analysis may include, but is not limited to, nucleic acid analysis, protein analysis, cellular analysis, and the like. The array may include a repeating pattern of features resolved in the low micron or submicron resolution range. While the systems, devices, and methods described herein may be described with respect to analyzing regular patterns of features, it should be understood that they may also be used with random distributions of features.

[0009] It should be noted that, as used herein, "patterned array" can include, but is not limited to, sequencing arrays formed as patterned flow cells. Such arrays include wells into which analytes can be located for processing and analysis. The wells can be arranged in a repeating pattern, a non-repeating pattern, or a random arrangement on one or more surfaces of a substrate (e.g., a flow cell). For simplicity, all such devices are referred to as being included in the term "patterned array" or "array," and should be understood as such.

[0010] The systems, devices, and methods described herein are robust to changes in feature characteristics in well patterns or layouts. These changes can manifest as different signal characteristics detected for one or more features in different images. For example, in nucleic acid sequencing techniques, an array of nucleic acids is subjected to several cycles of biochemical processing and imaging (e.g., a sequencing run). In some examples, each cycle can result in one of four different labels being detected in each feature based on the nucleotide base biochemically processed in that cycle. In such examples, multiple (e.g., four) different images are acquired in a given cycle, and each feature can be detected in the image. In one example, alignment of images in a given cycle presents unique challenges, as a feature detected in one image may appear dim in other images. Furthermore, sequencing involves multiple cycles, and alignment of features represented in image data from successive cycles is used to determine the sequence of nucleotides in each well based on the sequence of labels detected in each well. Improper alignment of images within a cycle or across different cycles can adversely affect sequence analysis. For example, methods using regular patterns may be susceptible to walk-off errors during image analysis. In one example, walk-off errors occur when two overlaid images are offset by one or more repeat units of the pattern, resulting in the patterns appearing to overlap, but adjacent features in different patterns being improperly correlated in the overlay.

[0011] The terms "spot" and "feature," used interchangeably herein, may refer to points or regions in a pattern that can be distinguished from other points or regions according to their relative location. An individual feature can contain one or more molecules of a particular type. For example, a feature can contain a single target nucleic acid molecule having a particular sequence. In another example, a feature can contain several nucleic acid molecules having the same sequence (and / or its complementary sequence). Different molecules in different features of a pattern can be distinguished from one another according to the location of the feature within the pattern. Features can include, but are not limited to, wells in a substrate (e.g., a flow cell), clusters within a well of a substrate (e.g., a flow cell), protrusions from a substrate, ridges on a substrate, pads of gel material on a substrate, channels in a substrate, etc.

[0012] The term "fiducial" may refer to a distinguishable reference point in or on an object. The reference point may be, for example, but not limited to, a mark, a second object, a shape, an edge, a region, an irregularity, a channel, a pit, a post, etc. The reference point may be present in an image of the object or in another dataset derived from detecting the object. The reference point may be specified by an X and / or Y coordinate in the plane of the object. Alternatively or additionally, the reference point may be specified by a Z coordinate orthogonal to the XY plane. For example, the reference point may be specified by a Z coordinate defined by the relative location of the object and the detector. In one or more cases, one or more coordinates of the reference point may be specified relative to one or more other features of the object or of the image or other dataset derived from the object.

[0013] The term "cluster" can refer to a collection of DNA molecules. In a patterned flow cell, the clusters can be located within the wells of the flow cell.

[0014] The term "footprint" may refer to the perimeter of an object, fiducial, feature, or other object in a plane. For example, a footprint may be defined by coordinates in an XY plane orthogonal to a detector observing the plane. A footprint may be defined by shape (e.g., circular, square, rectangular, triangular, polyhedral, elliptical, etc.) and / or area (e.g., at least 1 μm 2 , 5 μm 2 , 10 μm 2 , 100 μm 2 , 1000 μm 2 , 1mm 2 etc.).

[0015] The term "image" may refer to a representation of all or a portion of an object (e.g., a sample). The representation may be an optically detected reproduction. For example, an image may be obtained from fluorescence, luminescence, scattering signals, absorption signals, etc. An image may include any optical data set that can be used to obtain information about an object (e.g., features and / or fiducials) within a defined area. The portion of an object present in an image may be the surface of the object or other XY plane. An image is a two-dimensional (2D) representation, although in some cases, information in an image may be derived from three dimensions (3D). An image need not include optically detected signals. Alternatively, non-optical signals may be present instead. An image may be provided in a computer-readable format or medium, such as one or more of those described herein.

[0016] The term "tile" may refer to one or more images of the same region of a sample. For example, each of the one or more images may represent a respective color channel. A tile may form an imaging data subset of an imaging dataset for one imaging cycle.

[0017] The term "optical signal" can refer to, for example, fluorescence, luminescence, scattering, or absorption signals. Optical signals can be detected in the ultraviolet (UV) range (approximately 200-390 nm), visible (VIS) range (approximately 391-770 nm), infrared (IR) range (approximately 0.771-25 micrometers), or other ranges of the electromagnetic spectrum. Optical signals can be detected in a manner that excludes all or part of one or more of these ranges, depending on the feature of interest.

[0018] The term "location data" may refer to information regarding the relative locations of two or more things. For example, this information may relate to the relative location of at least one fiducial and the object on which it occurs, the relative location of at least one fiducial on the object and at least one feature on the object, the relative location of two or more features on the object, the relative location of a detector and the object, the relative location of two or more portions of a fiducial, etc. The information may be included in any of a variety of formats indicating relative location, including, but not limited to, numerical coordinates, pixel identities, images, etc. The location data may be provided in a computer-readable format or medium, such as one or more of those described herein.

[0019] The term "repeating pattern" may refer to the relative location of a subset of features in one region of an object that is the same as the relative location of the subset of features in at least one other region of the object. In one or more cases, one region may be adjacent to another region in the pattern. The relative location of features in one region of the repeating pattern can be predicted from the relative location of features in another region of the repeating pattern. The subset of features used for measurement may include at least two features, but can include at least 3, 4, 5, 6, 10, or more features. Alternatively, or additionally, the subset features used for measurement may include 2, 3, 4, 5, 6, or 10 or fewer features. In some cases, the repeating pattern may include multiple repetitions of a subpattern.

[0020] The term "signal level" may refer to the amount or quantity of detected energy or encoded information having a desired or predetermined characteristic. For example, optical signals may be quantified by one or more of intensity, wavelength, energy, frequency, power, brightness, etc. Other signals may be quantified according to characteristics such as voltage, current, electric field strength, magnetic field strength, frequency, power, temperature, etc. The absence of a signal may refer to a signal level of zero or a signal level that is not significantly distinguishable from noise.

[0021] The term "virtual fiducial" may refer to a reference point applied to an object in an image and derived from a source other than the object or image, respectively. For example, a virtual fiducial may be derived from a first object (e.g., a mold object or a standard object) and applied to an image of a second object. Alternatively, a virtual fiducial may be derived from a design, drawing, or plan used to create the object. In one or more cases, the virtual fiducial may be represented or specified as described herein with respect to fiducials. In one or more cases, the virtual fiducial may be provided in a computer-readable format or medium such as those described herein.

[0022] The term "XY coordinates" may refer to information specifying a location, size, shape, and / or orientation within the XY plane. The information may be, for example, numerical coordinates in a Cartesian coordinate system. The coordinates may be provided relative to one or both of the X and Y axes, or may be provided relative to another location within the XY plane. For example, the coordinates of a feature of an object (e.g., a cluster corresponding to a well of a flow cell) may specify the location of the feature relative to a fiducial or other feature location on the object. The term "XY plane" may refer to a 2D region defined by linear axes X and Y. When used with respect to a detector (e.g., detector subsystem 215 of FIG. 2B) and an object observed by the detector, the region may be further specified as being orthogonal to the direction of observation between the detector and the object being detected. The term "Z coordinate" may refer to information specifying the location of a point, line, or region along an axis orthogonal to the XY plane. In one or more cases, the Z axis is orthogonal to the region of the object observed by the detector. For example, the direction of focus of an optical system may be specified along the Z axis.

[0023] An increasing number of applications are being developed for arrays with features bearing biological molecules, such as nucleic acids and polypeptides. Such arrays typically contain deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) probes, although these techniques may also be applicable to other protein or chemical analyses. For example, DNA and / or RNA probes can be used to identify nucleotide sequences present in humans and other organisms. In one or more applications, for example, individual DNA or RNA probes can be attached to individual features of the array. A test sample, such as from a known human or organism, can be exposed to the array so that target nucleic acids (e.g., gene fragments, mRNA, amplicons thereof, etc.) hybridize to complementary probes at each feature in the array. The probes can be labeled in a target-specific process (e.g., due to a label present on the target nucleic acid or due to an enzymatic label of the probe or target present in hybridized form at the feature). The array can then be examined by scanning specific frequencies of light across the features to identify which target nucleic acids are present in the sample.

[0024] Biological arrays can be used for gene sequencing and similar applications. Gene sequencing can involve determining the order of nucleotides in a length of target nucleic acid, such as a fragment of DNA or RNA. A relatively short sequence is sequenced for each feature, and the resulting sequence information can be used in various bioinformatics methods to logically connect the sequence fragments and reliably determine the sequence of much larger lengths of genetic material from which the fragments are derived. Automated computer-based systems for identifying signature fragments can be used for genome mapping, identifying genes and their functions, and the like. Because a large number of variants are present in the array, the array can be used to characterize genome content, which replaces the alternative of performing numerous experiments on individual probes and targets.

[0025] Any of a variety of analyte arrays can be used in the processes described herein. The array can contain features, each feature having an individual probe or a population of probes. In the latter case, the population of probes in each feature can be homogeneous and correspond to a single type of probe. For example, in the case of nucleic acid sequences, each feature can have multiple nucleic acid molecules, each with a common sequence. In other cases, the population in each feature of the array can be heterogeneous.

[0026] Similarly, protein arrays can feature a single protein or a population of proteins with the same amino acid sequence. Probes can be attached to the surface of the array, for example, via covalent binding of the probe to the surface or via non-covalent interactions of the probe with the surface. In some examples, probes, such as nucleic acid molecules, can be attached to the surface via a gel layer, as described, for example, in U.S. Pat. No. 9,012,022, issued April 21, 2015, and / or U.S. Patent Application Publication No. 2011 / 0059865, filed January 7, 2005, each of which is incorporated herein by reference. In other examples, arrays can include those used in nucleic acid sequencing applications. In some other examples, arrays used for nucleic acid sequencing often have random spatial patterns of nucleic acid features. For example, the HiSeq or MiSeq sequencing platforms available from Illumina Inc. (San Diego, Calif.) utilize flow cells in which nucleic acid sequences are formed by random seeding followed by bridge amplification. Furthermore, patterned arrays can also be used for nucleic acid sequencing or other analytical applications. The features of such patterned arrays can be used to capture single nucleic acid template molecules and seed the subsequent formation of homogenous colonies, for example, via bridge amplification.

[0027] In one or more cases, the size of features on an array (or other object used in the processes described herein) can be selected to suit a particular application. For example, in some cases, an array feature may be sized to accommodate a single nucleic acid molecule. A surface having multiple features in this size range can be used to construct an array of molecules for detection with single molecule resolution. Features in this size range can be used in arrays with features each containing a colony of nucleic acid molecules. In one or more examples, features such as wells in a flow cell can have a diameter ranging from about 200 nm to about 1 μm and a pitch of about 400 nm. In one or more examples, the tile height can be about 600 μm. In one or more examples, the tile width can be about 1000 μm. In one or more examples, the array features each have a diameter of about 1 mm. 2 Below, approximately 500μm 2 Below, approximately 100μm 2 Below, approximately 10μm 2 Below, approximately 1μm 2 Below, about 500nm 2 Less than or equal to about 100 nm 2 Below, approximately 10nm 2 Below, approximately 5nm 2 Less than or equal to 1 nm 2 Alternatively or additionally, the features of the array may have an area that is less than or equal to about 1 mm 2 More than approximately 500μm 2 More than approximately 100μm 2 or more, about 10μm 2 or more, approximately 1μm 2 or more, about 500nm 2 Over 100nm 2 or more, about 10nm 2 or more, about 5nm 2 More than or about 1 nm 2or more. Features may have sizes that fall within a range between upper and lower limits selected from those exemplified above. While some size ranges for surface features have been exemplified with respect to nucleic acids and with respect to the scale of nucleic acids, it will be understood that features in these size ranges may be used in applications that do not involve nucleic acids. It will further be understood that feature sizes need not necessarily be limited to the scale used in nucleic acid applications.

[0028] When including an object with multiple features, such as an array of features, the features may be distinct and separated by spaces between them. Thus, an array may have features separated by an edge-to-edge distance of at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, or 0.5 μm. Alternatively or additionally, an array may have features separated by an edge-to-edge distance of at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, or 100 μm. These ranges may apply to the average edge-to-edge spacing of the features, as well as the minimum or maximum spacing.

[0029] In one or more cases, the features of the array may not be distinct; instead, adjacent features may abut one another. Whether the features are distinct or not, the size of the features and / or the pitch of the features may vary so that the array can have a desired density. For example, the average feature pitch in a regular pattern may be at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, or 0.5 μm. Alternatively or additionally, the average feature pitch in a regular pattern may be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, or 100 μm. In some cases, these ranges may also apply to the maximum or minimum pitch of the regular pattern. For example, the maximum feature pitch in the regular pattern may be at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, or 0.5 μm, and / or the minimum feature pitch in the regular pattern may be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, or 100 μm.

[0030] The density of features in an array can also be understood in terms of the number of features present per unit area. For example, the average density of features in an array can be at least 1×10 3 Features / mm 2 , 1×10 4 Features / mm 2 , 1×10 5 Features / mm 2 , 1×10 6 Features / mm 2 , 1×10 7 Features / mm 2 , 1×10 8 Features / mm 2 , or 1×10 9 Features / mm 2 Alternatively or additionally, the average density of features in the array may be at most 1 x 10 9 Features / mm 2 , 1×10 8 Features / mm 2 , 1×10 7 Features / mm 2 , 1×10 6 Features / mm 2 , 1×10 5 Features / mm 2 , 1×10 4 Features / mm 2 , or 1×10 3 Features / mm 2 It may be the following:

[0031] The above ranges may apply to all or part of a regular pattern, including, for example, all or part of an array of features.

[0032] The features within the pattern can have any of a variety of shapes. For example, when viewed in a 2D plane, such as on the surface of an array, the features may appear round, circular, oval, rectangular, square, symmetrical, asymmetrical, triangular, polygonal, etc. The features can be arranged in a regular repeating pattern, including, for example, a hexagonal or rectilinear pattern. The pattern can be selected to achieve a desired level of packing. For example, round features can be packed in a hexagonal arrangement. In another example, the round features can be packed in a rectilinear pattern.

[0033] A pattern can be characterized in terms of the number of features present in a subset that forms the smallest geometric unit of the pattern. A subset can include, for example, at least 2, 3, 4, 5, 6, 10, or more features. Depending on the size and density of the features, a geometric unit can be as small as 1 mm. 2 , 500 μm 2 , 100 μm 2 , 50 μm 2 , 10 μm 2 , 1 μm 2 , 500nm 2 , 100 nm 2 , 50nm 2 , 10nm 2 Alternatively or additionally, the geometric unit may be 10 nm 2 , 50nm 2 , 100 nm 2 , 500nm 2 , 1 μm 2 , 10 μm 2 , 50 μm 2 , 100 μm 2 , 500 μm 2 , 1mm 2 , or larger. The characteristics of the features in the geometric units, such as shape, size, pitch, etc., can be selected for the features in the array or pattern.

[0034] An array with a regular pattern of features may be ordered with respect to the relative location of the features, but random with respect to one or more other characteristics of each feature. For example, in the case of a nucleic acid array, the nucleic acid features may be ordered with respect to their relative location, but may be random with respect to one's knowledge of the sequence of the nucleic acid species present in any particular feature. As a more specific example, a nucleic acid array formed by seeding a repeating pattern of features with a template nucleic acid and amplifying the template in each feature (e.g., via cluster amplification or bridge amplification) to form copies of the template in the feature may have a regular pattern of nucleic acid features, but may be random with respect to the distribution of the sequences of the nucleic acids across the array. Thus, detecting the presence of nucleic acid material on the array may result in a repeating pattern of features, while sequence-specific detection may result in a non-repeating distribution of signals across the array.

[0035] It will be understood that descriptions of pattern, order, randomness, etc. herein relate not only to features on an object, such as features on an array, but also to analytes in an image. Thus, the pattern, order, randomness, etc. can exist in any of a variety of formats used to store, manipulate, or communicate image data, including, but not limited to, computer-readable media or computer components such as a graphical user interface or other output device.

[0036] In one or more cases, fiducials are included on the object (or in the image) to facilitate identification and location of individual features on the object. Fiducials can be used to align spatially ordered patterns of features because they provide reference points for the relative locations of other features. Fiducials may be used in applications where an array is repeatedly detected to track changes that occur in individual features over time. For example, fiducials may allow individual nucleic acid clusters to be tracked through successive images obtained over multiple sequencing cycles so that the sequences of nucleic acid species present in each cluster can be individually determined. Fiducials may be included on or within an array (i.e., whether within an array, or any random or other layout), such as on one or more surfaces of a patterned array support or substrate, as well as within wells and molecular image data, including wells in which molecules are located, to facilitate identification and location of individual features on the array.

[0037] The fiducials may be arranged in various patterns, as illustrated in the example images 101, 103, and 105 of FIGS. 1A-1C. In some cases, as illustrated in FIGS. 1A and 1B, the fiducials may be linearly arranged in parallel or substantially parallel directions to one another. For example, as illustrated in the example tile 102 of FIG. 1A, a set of fiducials 108a may extend across the tile 102 in a direction parallel to another set of fiducials 108b. It should be noted that a set of fiducials may include any number of fiducials (e.g., without limitation, two fiducials, three fiducials, four fiducials, eight fiducials, and nine fiducials). Furthermore, for simplicity, the example images 101, 103, and 105 are illustrated with one tile (e.g., tiles 102, 104, and 122), respectively. However, it should be understood that the image may be divided into multiple imaging data subsets (e.g., tiles) corresponding to respective regions of the patterned sample.

[0038] In some cases, a set of criteria and an adjacent set of criteria may include the same number of criteria. For example, as illustrated in FIG. 1A, set of criteria 108a may include four criteria, and set of criteria 108b may include four criteria. In other cases, one set of criteria may have a greater number of criteria than an adjacent set of criteria. For example, set of criteria 116a may have eight criteria, and an adjacent set of criteria may have six criteria. In one or more cases, a set of criteria may be centered between two adjacent sets of criteria. For example, as illustrated in the example tile 104 of FIG. 1B, set of criteria 110c may include two criteria disposed between sets of criteria 110a, 110b. In one or more cases, the sets of criteria may be offset to account for another set of criteria. For example, as illustrated in example tile 122 of FIG. 1C , fiducials 118 in set 116a may be evenly spaced from one another, while fiducials in set 116b may be separated into groups (e.g., groups 120a and 120b). In some cases, the groups may be spaced apart such that the space between the groups is greater than the space between the fiducials within each group. In one or more cases, fiducials in one set of fiducials may be aligned with fiducials in another set of fiducials. For example, as illustrated in FIG. 1A , fiducials 106a in set 108a may be horizontally aligned with fiducials 106b in set 108b, and fiducials 107a in set 108a may be horizontally aligned with fiducials 107b in set 108b. While arrangements 101, 103, and 105 in FIGS. 1A-1C illustrate two and three sets of columns of fiducials, it should be understood that an arrangement may include any number of sets of fiducials.

[0039] Regarding the arrangement of fiducials in image 103 (i.e., a 9x2x9 flow cell fiducial layout), the arrangement is configured to ensure data quality against increased jitter. This arrangement includes two fiducials 112a and 112b, which are included in the center of the arrangement and are a sufficient distance from all other fiducials, such as those in fiducial sets 110a and 110b, to avoid locating erroneous fiducials due to drift. That is, when jitter induces perturbations in the image and distorts data quality, fiducials 112a and 112b can be used as reference fiducials to correct for drift and locate fiducials in fiducial sets 110a and 110b. Furthermore, to reduce the computational cost of locating additional fiducials and significantly improve the speed of the image registration process, the size of the fiducials in the arrangement illustrated in FIG. 1B is reduced compared to the size of the fiducials utilized in FIG. 1A. For example, the fiducials in tile 102 can each have a diameter of 50 μm. In another example, the fiducials in tile 104 can each have a diameter of 25 μm.

[0040] The fiducials may be configured in various shapes. For example, the fiducials may be a set of concentric circles (e.g., fiducials 106a and 106b in FIG. 1A ). That is, the fiducials may form, for example, a bull's-eye fiducial, examples of which are described in U.S. Pat. No. 9,512,422, issued December 6, 2016; U.S. Pat. No. 11,262,307, issued March 1, 2022; and U.S. Pat. No. 11,308,640, issued April 19, 2022, each of which is incorporated herein by reference. In another example, the fiducials may be circles (e.g., fiducials 114a and 114b in FIG. 1B ). In one or more cases, an array of wells or other features may include fiducials forming a pattern of multiple rings. In one or more cases, image registration may be performed by aligning (overlapping) fiducials in an image (e.g., a virtual fiducial image) with fiducials in a sub-image, as further described herein. The correlation of the match can be determined by calculating a similarity measure such as two-dimensional cross-correlation, sum of squared intensity differences, etc. The best correlation of the match can be identified as the positioning where the brightest pixel from the image overlaps with the brightest pixel on the reference image. Alternatively or additionally, the best correlation can be identified as a measure of how strong the best correlation between the virtual reference image and the sub-image is relative to the second-best correlation between the virtual reference image and the sub-image, as described further herein. Note that circular symmetry is optional with respect to the reference. Alternatively, other symmetries can be utilized. Furthermore, in general, symmetry is optional, and asymmetric references can be used instead. Furthermore, in one or more cases, the reference can have a footprint larger than the area of ​​each individual feature of the object to be aligned. In some cases, the reference can have a footprint larger than a geometric unit of a feature that is repeated in a repeating pattern. A larger footprint of the fiducials can reduce the risk of "walk-off" or integral offset (e.g., vertical or horizontal translation), where alignment may appear locally correct within a geometric unit of features in a repeating pattern, but each feature (or geometric unit of features) is mistaken for its neighbors.

[0041] In one example, known fiducials at well-defined locations can be located within image data using template matching techniques. For example, a theoretical grid or model of fiducial locations can be generated based on the known locations of the fiducials within a flow cell. An image containing data corresponding to one or more fiducials within the image region can then be captured. Although the image containing the fiducials may be perturbed and / or distorted such that the relative locations between the captured fiducials contain a certain amount of error, the template can be used to determine the approximate location of the fiducials within the image, regardless of the presence of such error. Once the relative locations of the fiducials within the image are determined using the template, image processing techniques (described in more detail below) can be used to determine the locations of other features within the image based on their known locations relative to the fiducials. For example, to account for errors in the relative locations of objects within the image due to various environmental and / or optical perturbations, one or more geometric transformations (e.g., affine transformations) may be performed on the image using the located fiducials. For example, to account for errors introduced by high-frequency jitter due to non-uniform motion within the stage, a given region can be divided into multiple subregions, and geometric transformations can be applied to each subregion separately. Such techniques are described in more detail below.

[0042] 2A illustrates a schematic diagram of a system environment (or "environment") 200 described herein. As illustrated, the environment 200 includes one or more server devices 202 connected via a network 212 to one or more of a client device 208, a database 216, and a sequencing device 214.

[0043] As shown in FIG. 2A , server device 202, client device 208, database 216, and sequencing device 214 can communicate with each other via network 212. Network 212 can include any suitable network over which computing devices can communicate. Network 212 can include wired and / or wireless communication networks. An exemplary wireless communication network can be comprised of one or more types of radio frequency (RF) communication signals using one or more wireless communication protocols, such as a cellular communication protocol, a wireless local area network (WLAN), a WIFI communication protocol, and / or another wireless communication protocol. While FIG. 2A illustrates components of environment 200 communicating via network 212, it will be understood that components of environment 200 can communicate directly with each other, for example, bypassing network 212. For example, client device 208 can communicate directly with sequencing device 214.

[0044] As shown by FIG. 2A , the sequencing device 214 may comprise a device for sequencing a biological sample. In one or more cases, the biological sample may include, but is not limited to, human and non-human DNA, for example, for determining individual nucleotide bases of a nucleic acid sequence (e.g., sequencing-by-synthesis). In one or more cases, the biological sample may include, but is not limited to, human and non-human RNA. The sequencing device 214 can analyze nucleic acid segments and / or oligonucleotides extracted from the sample to generate nucleotide reads and / or other data using the computer-implemented methods and systems described herein, either directly or indirectly on the sequencing device 214. More specifically, the sequencing device 214 can receive and analyze nucleic acid sequences extracted from the sample within a nucleotide sample slide (e.g., a flow cell). The sequencing device 214 can sequence nucleic acid segments into nucleotide reads using sequencing-by-synthesis (SBS).

[0045] In one or more cases, the sequencing device 214 can correct for high-frequency stage motion (i.e., jitter), which leads to inconsistent movement of spot locations within an image. That is, the sequencing device 214 can perform the processes described herein to reduce the introduction and / or amount of jitter for each subregion. The spots may correspond to clusters of DNA molecules and / or fiducials within a well of a flow cell. The sequencing device 214 can acquire images of an object (e.g., a flow cell) having multiple fiducials arranged in a pattern. One exemplary fiducial pattern can include a set of four fiducials arranged linearly adjacent to another set of four fiducials. Another exemplary fiducial pattern can include a set of nine fiducials arranged linearly adjacent to another set of nine fiducials, with two sets of fiducials centered between the first and second sets of nine fiducials. The sequencing device 214 can generate fiducial locations and decompose the image into subregions, each containing a subset of fiducials. The sequencing device 214 may perform a process (e.g., a template matching process) on the fiducials in the image to initially locate the fiducials. For example, the relative locations of the fiducials may be known before acquiring the image. The known locations of the fiducials can be used to generate a pattern or reference data that is compared to an image that includes a captured or recorded version of the fiducials. The captured fiducials in the image may not be located relative to each other in the exact position expected based on the known pattern, for example, due to limitations in the optics of the image capture device and / or other errors associated with acquiring the image. However, using the known fiducial pattern and a template matching process, the fiducials in the image can be located after the first reference fiducial is located and the image is compared to the known pattern.

[0046] Once the fiducials in the image have been located using the reference data, the image can be divided into multiple subregions, each of which includes multiple (e.g., at least three) fiducials at a given location within the respective subregion. Furthermore, the sequencing device 214 can perform a geometric transformation (i.e., an affine transformation) on each subregion to generate a representation of how the image data may have been perturbed or distorted compared to the actual location of the object captured in the image (e.g., due to nonlinearities or other errors introduced in the image capture process). In this manner, because the placement of the fiducials within each subregion and / or the location of one or more features within each subregion are known relative to the location of the fiducials, the transformation determined for the subregion can be used to locate the location of the feature within the image. Such a technique allows the location of the feature to be determined even in the event of errors introduced into the image during the image capture process. Furthermore, performing a geometric transformation for each subregion can correct for errors introduced by high-frequency jitter caused by non-uniform motion within the stage.

[0047] By performing a geometric transformation for each subregion, individual representations of relative perturbations / errors in the image data can be determined for subregions whose areas are small enough that high-frequency jitter has little effect and errors associated with such jitter are reduced. That is, the areas of the subregions can be selected so that the subregions are substantially invariant to high-frequency jitter and errors associated with such jitter that may affect the entire sequencing image. For example, the geometric transformation may correspond to an affine transformation, and a resulting affine matrix may be generated for each subregion. Based on the generated arrangement of fiducials within the subregion and the transformation determined for the subregion, the sequencing device 214 can determine the locations of flow cell spots within the subregion. That is, the sequencing device 214 can use each determined matrix to account for jitter and correct the positioning of spot locations within the image. Thus, while most or all of a sequencing image may be affected by jitter, the sequencing device 214 can construct local regions within the image that are small enough to be locally invariant to the same jitter that affects the entire sequencing image. For example, each sub-region may have a frequency of 400 Hz, or less than 200 Hz to compensate for jitter that is expected to affect the entire sequencing image.

[0048] As further illustrated by Figure 2A, server device 202 can generate, receive, analyze, store, and / or transmit digital data such as, but not limited to, image data, data for determining nucleotide base calls, or sequencing a nucleic acid polymer. As shown in Figure 2A, sequencing device 214 can generate and transmit (and server device 202 can receive) image data, nucleotide reads, and / or other data analyzed by server device 202.

[0049] In one or more cases, the server device 202 can communicate with the client device 208. For example, the server device 202 can send data, including image data, sequencing data, or other information, to the client device 208, and the server device 202 can receive input from a user via the client device 208.

[0050] In some cases, server device 202 may comprise a distributed collection of servers, where server device 202 includes several server devices distributed across network 212. In some examples, distributed server device 202 may be located in the same location or in different physical locations. In other cases, server device 202 may comprise a content server, an application server, a communication server, a web hosting server, or another type of server.

[0051] In one or more cases, the server device 202 and / or the sequencing device 214 may include a sequencing system 204 or a portion thereof. The sequencing system 204 can analyze other data, such as nucleotide reads, image data, and / or sequencing metrics generated by the sequencing device 214, to, for example, correct distortions in the image data. In another example, the sequencing system 204 can analyze other data, such as nucleotide reads, image data, and / or sequencing metrics generated by the sequencing device 214, to, for example, determine the nucleotide base sequence of a nucleic acid polymer. For example, the sequencing system 204 can receive raw data generated by the sequencing device 214. The sequencing system 204 can determine the nucleotide base sequence of a nucleic acid segment based on the received raw data. In one or more cases, the raw data can be received by the sequencing device 214 in a file format, such as, but not limited to, a FASTQ file, that can be recognized for processing. The FASTQ file can include a text file containing sequence data from clusters that pass through a filter on a flow cell. The FASTQ format is a text-based format for storing both biological sequences (e.g., nucleotide sequences, etc.) and the corresponding quality scores of the biological sequences. In one or more cases, the sequencing system 204 can process the sequencing data to determine the sequence of nucleotide bases in DNA and / or RNA segments or oligonucleotides. The sequencing system 204, or one or more portions thereof, resident on the sequencing device 214, can enable on-device analysis of the sequencing data, image data, etc. The sequencing system 204, or one or more portions thereof, resident on the sequencing device 214, can enable the sequencing device 214 to monitor the status of one or more applications operating to perform analysis on the sequencing data.

[0052] The client device 208 can generate, store, receive, and / or transmit digital data to enable the sequencing processes and analyses described herein. For example, the client device 208 can receive sequencing metrics from the sequencing device 214. The client device 208 can communicate with the server device 202 and / or the sequencing device 214 to receive one or more files including nucleotide base calls and / or other metrics. In one or more cases, the client device 208 can present or display information about the nucleotide base calls to a user within a graphical user interface of the client device 208. The client device 208 may comprise various types of client devices. In some examples, the client device 208 may be a non-mobile device such as a desktop computer, a server, etc. In other examples, the client device 208 may be a mobile device such as a laptop, a tablet, a mobile phone, a smartphone, etc. The client device 208 can include a sequencing application 210. The sequencing application 210 may be, for example, a web application or a native application (e.g., a mobile application, a desktop application) stored on and executed on the client device 208. In one or more cases, the sequencing application 210 may include instructions that (when executed) cause the client device 208 to receive data from the sequencing device 214 and present the data, such as, but not limited to, data from a variant call file, within a graphical user interface of the client device 208.

[0053] The environment 200 can include a database 216. The database 216 can store information such as, but not limited to, variant call files, nucleotide sequences of samples, nucleotide reads, nucleotide base calls, sequencing metrics, population data, imaging data, and / or other data described herein. The server device 202, the client device 208, and / or the sequencing device 214 can communicate with the database 216 (e.g., via the network 212) to store and / or access information such as, but not limited to, variant call files, nucleotide sequences of samples, nucleotide reads, nucleotide base calls, sequencing metrics, population data, imaging data, and / or other data described herein.

[0054] The environment 200 may be included in a local network or a local high-performance computing (HPC) system. In one or more other cases, the environment 200 may be included in a cloud computing environment comprising multiple server devices, such as server device 202, having distributed software and / or data. In one or more cases, the sequencing system 204 may be implemented to operate one or more subsystems as described herein. The sequencing system 204 may be distributed across server devices 202 with access to database 216 via network 212 within a cloud-based computing system.

[0055] 2B illustrates examples of one or more sequencing subsystems that may be implemented by sequencing device 214 and monitored and / or controlled by the computing subsystem when performing analyses on one or more flow cells. Additionally or alternatively, the subsystems may be controlled in response to a schedule / workflow determined by the computing subsystem. In one or more cases, detector subsystem 215 may be configured to perform analyses on one or more samples.

[0056] The exemplary sequencing device 214 may include a device for acquiring or generating an image of a sample. The system 215 (i.e., detector subsystem) of Figure 2B illustrates an exemplary imaging configuration for a backlight design implementation. While systems and methods may sometimes be described herein in the context of the exemplary system 215, it should be noted that these are merely examples in which implementations of the image distortion correction methods disclosed herein may be performed.

[0057] As illustrated in FIG. 2B , a subject sample is positioned on a flow cell 225, which is positioned on a stage 270 below the objective lens 242. Note that while FIG. 2B illustrates one flow cell 225, it should be understood that the system 215 may perform analysis on more than one flow cell 225. A light source 250 and associated optics direct a beam of light, such as laser light, to a selected sample location on the flow cell 225. The sample fluoresces, and the resulting light is collected by the objective lens 242 and directed to an image sensor in the camera system 240 for detecting the fluorescence. The stage 270 is moved relative to the objective lens 242 to position the next sample location on the flow cell 225 at the focal point of the objective lens 242. Movement of the stage 270 relative to the objective lens 242 can be achieved by moving the stage itself, the objective lens, some other component of the imaging system, or any combination of the foregoing. Further embodiments may also include moving the entire imaging system over a stationary flow cell.

[0058] Sequencing device 214 directs the flow of reagents (e.g., fluorescently labeled nucleotides, buffers, enzymes, cleavage reagents, etc.) to (and through) flow cell 225 and waste valve 220. Flow cell 225 may be a sample container containing one or more substrates onto which a sample is provided. For example, in the case of a system for analyzing a large number of different nucleic acid sequences, flow cell 225 includes one or more substrates to which nucleic acids to be sequenced are bound, attached, or associated. In one or more cases, the substrate may include any inert substrate or substrate to which nucleic acids may be attached, such as, for example, glass surfaces, plastic surfaces, latex, dextran, polystyrene surfaces, polypropylene surfaces, polyacrylamide gels, gold surfaces, and silicon wafers. In one or more cases, the substrates are within channels or other regions at multiple locations formed in a matrix or array throughout flow cell 225.

[0059] In one or more cases, the flow cell 225 can contain a biological sample that is imaged using one or more fluorescent dyes. For example, the flow cell 225 can be implemented as a patterned flow cell that includes a translucent cover plate, a substrate, and a liquid sandwiched therebetween, and the biological sample can be located on the inner surface of the translucent cover plate or the inner surface of the substrate. The flow cell 225 can include a large number (e.g., thousands, millions, or billions) of wells or regions patterned into a defined array (e.g., a hexagonal array, a rectangular array, etc.) within the substrate. Each region can form a cluster (e.g., a monoclonal cluster) of a biological sample, such as DNA, RNA, or another genomic material, that can be sequenced using sequencing-by-synthesis. The flow cell 225 can be further divided into a large number of spaced lanes (e.g., eight lanes), each lane containing a hexagonal array of clusters. An exemplary flow cell that can be used in the embodiments disclosed herein is described in U.S. Patent No. 8,778,848, the entire contents of which are incorporated by reference.

[0060] System 215 may also include a temperature station actuator 230 and a heater / cooler 235 to adjust the temperature of the fluid state in flow cell 225 as needed. A camera system 240 may be included to monitor and track the sequencing of flow cell 225. Camera system 240 may be implemented, for example, as a charge-coupled device (CCD) camera (e.g., a time delay integration (TDI) CCD camera), which may interact with various filters in a filter switching assembly 245, an objective lens 242, and a focused light source 250. Camera system 240 is not limited to a CCD camera; other camera and image sensor technologies may be used. In some examples, the sensor of camera system 240 may have a pixel size of about 5 to about 15 μm.

[0061] Output data from the sensors of camera system 240 may be communicated to a real-time analysis module (e.g., real-time analysis module 291 of FIG. 2C), which may be implemented as a software application that analyzes the image data (e.g., image quality scoring), reports or displays laser beam characteristics (e.g., focus, shape, intensity, power, brightness, position) in a graphical user interface (GUI), and dynamically corrects distortions in the image data, as further described below.

[0062] A light source 250 (e.g., an excitation laser, optionally in an assembly comprising multiple lasers) or other light source may be included to illuminate the fluorescent sequencing reaction in the sample via illumination through a fiber optic interface (which may optionally include one or more reimaging lenses, fiber optic attachments, etc.). System 215 may include a low-wattage lamp 265, a focused laser 261, and an inverse dichroic. In some cases, the focused laser 261 may be turned off during imaging. In other cases, an alternative focusing configuration may include a second focused camera (not shown), which may be a quadrant detector, position-sensitive detector (PSD), or similar detector, for measuring the location of the scattered beam reflected from the surface simultaneously with data collection. While illustrated as a backlit device, other examples may include light from a laser or other light source directed through objective lens 242 onto the sample on flow cell 225.

[0063] The flow cell 225 may be placed in a flow cell holder (e.g., a sample container holder), which may be placed on a movable staging area 270. The flow cell holder may, for example, securely hold the flow cell in the proper position or orientation relative to the light source 250, a prism (not shown) that directs laser illumination to an imaging plane, and the camera system 240 while sequencing occurs. Thus, the flow cell 225 may be mounted on a stage 270 to provide movement and alignment of the flow cell 225 relative to the objective lens 242. The stage 270 may have one or more actuators to enable the stage 270 to move in any of three dimensions. For example, actuators may be provided to enable the stage 270 to move in the X, Y, and Z directions relative to the objective lens in terms of a Cartesian coordinate system. This may allow one or more sample locations on the flow cell 225 to be positioned in optical alignment with the objective lens 242. For example, depending on the device design and imaging technique used, the patterned array contained in flow cell 225 may be initially positioned in the XY plane and moved within this plane during imaging, or the imaging component may be moved parallel to this plane during imaging. Flow cell 225 may extend in the XY plane, with the X direction being the long side of flow cell 225 and the Y direction being the short side (the flow cell is rectangular). However, it should be understood that this orientation may be reversed.

[0064] The focus (z-axis) component 217 can control the positioning of optical components relative to the flow cell 225 in a focus direction (typically referred to as the z-axis or z-direction). The focus component 217 may include one or more actuators physically coupled to the optical stage or the sample stage, or both, to move the flow cell 225 on the stage 270 relative to the optical components (e.g., the objective lens 242) to provide proper focusing for imaging operations. For example, the actuators may be physically coupled to the respective stages, e.g., by direct or indirect mechanical, magnetic, fluidic, or other connections or contacts with the stages. The one or more actuators may be configured to move the stage 270 in the z-direction while maintaining the stage 270 in the same plane (e.g., while maintaining a level or horizontal orientation perpendicular to the optical axis). The one or more actuators may be configured to tilt the stage 270. This may be done, for example, so that the flow cell 225 can be dynamically leveled to account for any tilt in its surface.

[0065] Focusing the system 215 may refer to aligning the focal plane of the objective lens with the sample being imaged at a selected sample location. However, focusing may also refer to adjustments to the system 215 to obtain desired characteristics for the representation of the sample, such as a desired level of sharpness or contrast in the image of the test sample. Because the usable depth of field of the focal plane of the objective lens may be small (sometimes on the order of 1 μm or less), the focusing component 217 may closely follow the surface being imaged. Because the flow cell 225 may not be perfectly flat when fixed within the instrument, the focusing component 217 may be set up to follow the profile of the surface being imaged while moving along the scanning direction (e.g., the y-axis).

[0066] Light emanating from the test sample at the sample location being imaged may be directed to one or more detectors of the camera system 240. An aperture may be included and positioned to allow only light emanating from the focal region to pass to the detector. The aperture may be included to improve image quality by filtering components of light emanating from regions outside the focal region. An absorption filter may be included in the filter switching assembly 245 that can be selected to record the determined emission wavelength and to filter out any stray laser light.

[0067] Although not illustrated, a controller can be provided to control the operation of system 215. The controller may be implemented to control aspects of system operation, such as focusing, stage movement, and imaging operations. In one or more cases, the controller may be implemented using hardware, a processor (e.g., machine-executable instructions), or a combination of the foregoing. For example, the controller may include one or more CPUs or processors with associated memory. In another example, the controller includes hardware or other circuitry for controlling operations, such as a computer processor and a non-transitory computer-readable medium having machine-readable instructions stored thereon. For example, the hardware or other circuitry may include one or more of the following: a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a programmable logic device (PLD), a complex programmable logic device (CPLD), a programmable logic array (PLA), a programmable array logic (PAL), or other similar processing device or circuitry. As yet another example, a controller may include a combination of this circuitry with one or more processors.

[0068] In one or more cases, flow cell 225 can be read from at least one side (i.e., the top and / or bottom). Thus, multiple readers or imaging systems can be used to read signals emanating from the channels of flow cell 225. Flow cell 225 may include one or more complementary metal oxide semiconductor (CMOS) sensors, which can be implemented in place of or in addition to an external camera and / or optics within camera system 240.

[0069] The sequencing device 214 may include an access subsystem configured to move a cartridge (e.g., from a receiving position to an engaging position) and / or actuate a door (e.g., to open or close) to provide access to a cartridge holder. The cartridge may contain one or more flow cells, fluids, reagents, or other materials for loading into the sequencing device 214. The sequencing device 214 may include a status subsystem including a light bar that provides a visual indication through color and / or intensity changes of the status of one or more processes occurring on the sequencing device 214.

[0070] Each of the subsystems on the sequencing device may be powered by power resources controlled by a power subsystem. The power subsystem may include a power source, such as an alternating-current (AC) power source or a direct current (DC) power source. The power resources may be controlled by one or more applications to perform different tasks in the sequencing process. The power source may generate supply voltages to power the subsystems within the sequencing device 214.

[0071] In one or more cases, system 215 (e.g., a detector subsystem) can acquire a targeted image of an object (e.g., a flow cell 225), the image including a repeating pattern of features on the object and at least one fiducial on the object as well. In one or more cases, system 215 can be configured to capture high-resolution imaging of the surface of a substrate (e.g., a flow cell). System 215 can have sufficient resolution to distinguish features by density, pitch, and / or feature size. System 215 can be configured to maintain the object and detector in a static relationship while acquiring area images. In one or more cases, camera system 240 of system 215 can include a scanning device used to acquire images. For example, a scanning device (e.g., a "step-and-shoot" detector) can acquire continuous area images. In one or more other cases, a scanning device can be configured to continuously scan points or lines on the surface of an object to accumulate data and build an image of the surface. Such a scanning device (e.g., a point-scanning detector) can be configured to scan points (i.e., small detection areas) on the surface of an object via a raster motion in the XY plane of the surface. In one or more other cases, a scanning device (e.g., a line scan detector) may employ confocal line scanning to generate progressive pixelated image data that can be analyzed to locate individual features / fiducials within the array. In some cases, the scanning device (e.g., a line scan detector) may be configured to scan a line along the Y dimension of the object's surface, with the line having its longest dimension occurring along the X dimension. Note that scanning detection can be achieved by moving the system 215, the object, or both. Systems 215, such as those used in nucleic acid sequencing applications, are described in U.S. Patent Application Publication Nos. 2012 / 0270305 (A1), 2013 / 0023422 (A1), and 2013 / 0260372 (A1), as well as U.S. Patent Nos. 8,158,926 and 8,241,573, each of which is incorporated herein by reference.It should be noted that the processes described herein may be used to analyze image data collected for multiple swaths or regions detected in the region of a sample flow cell as described herein, however, it should be understood that the processes described herein may also be used to analyze image data collected for other types of substrates containing arrays of molecules or other detectable features.

[0072] In one or more cases, a system 215 for generating image data representing individual features / fiducials on the flow cell and the spaces between the features / fiducials, and a representation of the fiducials provided in or on the flow cell. The sequencing system 204 can receive the image data and process the image data to extract meaningful value from the imaging data as described herein in accordance with the present disclosure. In one or more cases, the sequencing system 204 can process the image data in real time or near real time while one or more sets of image data of the flow cell are being acquired. Such real-time analysis is useful for nucleic acid sequencing applications in which arrays of nucleic acids are subjected to repeated cycles of fluidic and detection operations. Analysis of sequencing data can be computationally intensive, and therefore it is beneficial to perform the process in real time or near real time, or in the background, while other data acquisition or analysis is ongoing. Exemplary real-time analytical methods that can be used in the present methods are those used in the MiSeq™ and HiSeq™ sequencing devices commercially available from Illumina, Inc., and / or those described in U.S. Patent Application Publication No. 2012 / 0020537 A1, the entire contents of which are incorporated herein by reference. The terms "real-time" and "near real-time," when used in connection with processing samples and imaging them, are intended to mean that processing occurs at least partially during the time that the samples are being processed and imaged. In other examples, image data may be acquired and stored for subsequent analysis by a similar process. This may allow other equipment (e.g., powerful processing systems) to handle processing tasks at the same or a different physical site from the one where imaging is occurring. This may also allow for reprocessing, quality verification, and the like.

[0073] The sequencing system 204 can analyze the image data to determine the locations of individual features that are visible or encoded in the image data, as well as locations where features are not visible (i.e., locations where no features are present, or where no significant radiation was detected from a present feature). Image data analysis can also be used to determine the location of fiducials that aid in locating features. Still further, image data analysis may be used to locate patterned arrays within the system to provide useful information, such as for processing or referencing purposes.

[0074] In one or more cases, the sequencing system 204 may be configured to assign an intensity and / or digital value to each feature and / or criterion based on characteristics of the image data represented by the pixel at the corresponding location. That is, for example, the sequencing system 204 may be configured to recognize a particular color (e.g., black, white, etc.) or wavelength of light detected at a particular location as being represented by a group or cluster of pixels at that location. For example, in a DNA imaging application, the four common nucleotides may be represented by distinct, distinguishable colors (or more generally, wavelengths or wavelength ranges of light). The sequencing system 204 may assign each color an intensity and / or digital value corresponding to that nucleotide. The sequencing system 204 assigns corresponding values ​​to features to alleviate the need for further processing of the image data itself, which may be larger in quantity (e.g., many pixels may correspond to each feature) and may have significantly larger numerical values ​​(i.e., more bits to encode each pixel). In one or more cases, the sequencing system 204 can associate each assigned intensity and / or digital value with a location within an image index or map, which can be done by reference to known or detected locations of fiducials or any data encoded by such fiducials. The map can correspond to the known or determined locations of individual features within the array.

[0075] As discussed herein, the sequencing device 214 can be used for SBS, in which fluorescently labeled modified nucleotides can be used to sequence high-density clusters (i.e., potentially millions of clusters) of amplified DNA present on the surface of a substrate (e.g., a flow cell). The flow cell containing the nucleic acid sample for sequencing can take the form of an array of distinct, separately detectable single molecules, or an array of features (or clusters) containing a homogenous population of specific molecular species, such as amplified nucleic acids with a common sequence.

[0076] 2C illustrates an exemplary line-scan imaging system 219. For example, system 219 may be a two-channel line-scan modular optical imaging system. While system 219 is described as a two-channel line-scan modular optical imaging system, it should be understood that system 219 may be configured to utilize any number of channels (e.g., one channel, three channels, four channels, etc.).

[0077] In one or more cases, the system 219 can be used for nucleic acid sequencing. Applicable techniques include those in which nucleic acids are attached to fixed locations in an array (e.g., the wells of a flow cell) and the array is repeatedly imaged. The system 219 can acquire images in two different color channels that can be used to distinguish one nucleotide base type from another. For example, the system 219 can perform a process called "base calling," which generally refers to the process of determining a base call (e.g., adenine (A), cytosine (C), guanine (G), or thymine (T)) for a given spot location in an image during an imaging cycle. During two-channel base calling, image data extracted from two images can be used to determine the presence of one of four base types by encoding the base identity as a combination of the intensities of the two images. For a given spot or location in each of the two images, the base identity can be determined based on whether the signal identity combination is [on, on], [on, off], [off, on], or [off, off].

[0078] System 219 may include a line generation module (LGM) 272 having at least one light source, such as, but not limited to, light sources 274 and 276. Light sources 274 and 276 may be coherent light sources, such as, but not limited to, laser diodes that output laser beams. Light source 274 may emit light of a first wavelength (e.g., a red wavelength), and light source 276 may emit light of a second wavelength (e.g., a green wavelength). The light beams output from laser sources 274 and 276 may be directed through a beam-shaping lens 286. In some cases, a single light-shaping lens may be used to shape the light beams output from both light sources. In other cases, separate beam-shaping lenses may be used for each light beam. In some examples, the beam-shaping lens is, for example, but not limited to, a Powell lens, such that the light beams are shaped into a line pattern. The beam-shaping lenses of LGM 272 or other optical components of imaging system 219 may be configured to shape the light emitted by light sources 274 and 276 into a line pattern (e.g., by using one or more Powell lenses, or other beam-shaping lenses, diffractive or scattering components).

[0079] The LGM 272 may further include a mirror 280 and a semi-reflective mirror 282 configured to direct the light beam through a single interface port to an emission optics module (EOM) 275. The light beam may pass through a shutter element 284. The EOM 275 may include an objective lens 298 and a z-stage 296 that moves the objective lens 298 longitudinally toward or further from a target 292. For example, the target 292 may include a liquid layer 290 and a semi-transparent cover plate 288, and the biological sample may be located on the inner surface of the semi-transparent cover plate 288 as well as on the inner surface of a substrate layer located below the liquid layer. The z-stage can then move the objective lens 298 to focus the light beam onto a surface on either side of the flow cell (e.g., onto the biological sample). The biological sample may be, for example, but not limited to, DNA, RNA, protein, or other biological material responsive to optical sequencing.

[0080] The EOM 275 may include a semi-reflective mirror 278 for reflecting a focus tracking light beam emitted from a focus tracking module (FTM) 281 onto a target 292 and for reflecting light returned from the target 292 back into the FTM 281. The FTM 281 may include a focus tracking optical sensor for detecting characteristics of the returned focus tracking light beam and generating a feedback signal for optimizing the focus of the objective lens 298 on the target 292.

[0081] EOM 275 may also include a semi-reflective mirror 298 for directing light through objective lens 298 while allowing light returned from target 292 to pass through. In some cases, EOM 275 may include a tube lens 273. Light transmitted through tube lens 273 may pass through filter element 271 and into camera module (CAM) 285. CAM 285 may include one or more optical sensors 283 for detecting light emitted from the biological sample in response to the incident light beam (e.g., fluorescence in response to red and green light received from light sources 274 and 276).

[0082] Output data from the sensors of CAM 285 may be communicated to real-time analysis module 291. Real-time analysis module 291 may execute computer-readable instructions to analyze the image data (e.g., quality scoring, base calling, etc.). These operations may be performed in real time during the imaging cycle to minimize downstream analysis time and provide real-time feedback and troubleshooting during the imaging run. In one or more cases, real-time analysis module 291 may be a computing device (e.g., computing device 800) communicatively coupled to and controlling imaging system 219. In one or more cases, real-time analysis module 291 may execute computer-readable instructions to correct distortions in the image data.

[0083] As discussed herein, image distortion can be particularly detrimental to multi-cycle imaging of patterned arrays (e.g., patterned flow cells) because it can shift the actual location of features in a scanned image away from the feature's expected location. This distortion effect can be particularly pronounced along the edges of the field of view, potentially rendering imaging data from these features unusable. This can result in reduced data throughput and increased error rates during multi-cycle imaging runs. Furthermore, high-frequency motion (i.e., jitter) of the stage during the cycles of an imaging run exacerbates this distortion effect. Embodiments described herein are directed to dynamically correcting image distortion during an imaging run (e.g., a sequencing run), thereby improving data throughput and reducing error rates during the imaging run.

[0084] FIG. 3 is a flowchart illustrating a procedure 300 for performing an exemplary image calibration process. In one or more cases, the sequencing device 214 can perform the procedure 300 as illustrated in FIG. 3 to determine the location of a reference standard. One or more portions of the procedure 300 may be performed by one or more computing devices. For example, one or more portions of the procedure 300 may be performed by one or more sequencing devices. One or more portions of the procedure 300 may be stored in memory as computer-readable or machine-readable instructions that may be executed by a processor of one or more computing devices. Although portions of the procedure 300 may be described herein as being performed by a sequencing device, the procedure 300 or portions thereof (e.g., the detector subsystem 215 or the imaging system 219) may be performed by another computing device or distributed across multiple computing devices, such as one or more sequencing devices, one or more client devices, and / or one or more server devices.

[0085] The procedure 300 may begin at 302. As shown in FIG. 3, at 302, the sequencing device 214 may generate a virtual reference image. To generate the virtual reference image, the sequencing device 214 may capture an image of a sample (e.g., a sample positioned on a flow cell as described herein) in an initial imaging cycle. In one or more cases, the initial imaging cycle may be a calibration cycle or the first cycle of a multi-cycle imaging run (e.g., a DNA sequencing run). The captured image may include image data corresponding to at least one feature and multiple references.

[0086] In some cases, the virtual reference image may include multiple fiducials arranged in a pattern. For example, the virtual reference image (e.g., sequencing image 102 of FIG. 1A) may be a sequencing image including two sets of fiducials (e.g., set 108a and set 108b) linearly arranged in parallel or substantially parallel directions. In one example, a set of fiducials (e.g., set 108a) may include four fiducials, and another set of fiducials (e.g., set 108b) may include four fiducials. In some examples, a fiducial from one set of fiducials (e.g., fiducial 106a) may be aligned with a fiducial from the other set of fiducials (e.g., fiducial 106b). As noted herein, fiducials are included in an object (e.g., a flow cell) or an image (e.g., sequencing image 102 of FIG. 1A) to facilitate identification and location of individual features. One or more fiducials provide a reference point for the relative location of features associated with the object or image. For example, a feature located within the area of ​​a fiducial may be associated with that fiducial. In another example, a feature located near one or more fiducials (i.e., the fiducials closest to the feature) may be associated with those fiducials. Note that while the procedure 300 described herein may be performed using the exemplary arrangement 101 of the sequencing image 102, the procedure 300 may similarly be performed using the exemplary arrangements 103, 105, and other fiducial arrangements.

[0087] Sequencing device 214 may acquire a virtual reference image, for example, but not limited to, via imaging system 219 or system 215 as described herein. The virtual reference image includes a repeating pattern of features on the object and at least one fiducial on the object as well. In one or more cases, sequencing system 204 can process the image data in real time or near real time while one or more sets of image data of the flow cell are being acquired.

[0088] The sequencing device 214 may generate a virtual reference image based on the fiducial configuration parameters. The fiducial configuration parameters may be parameters that describe the shape and size of the fiducial. For example, for a bull's-eye fiducial, such as fiducial 106a in FIG. 1A, the configuration parameters of fiducial 106a may include the diameters of the outer and inner circles that form fiducial 106a.

[0089] In one or more cases, based on the image data and / or fiducial configuration parameters, the sequencing device 214 can associate one of the fiducials (e.g., fiducial 106a) as a reference fiducial (e.g., reference fiducial 404 illustrated in FIG. 4A). The sequencing device 214 can utilize the reference fiducial to locate one or more fiducials associated with the virtual reference image. In one or more cases, the reference fiducial is a global reference point that the sequencing device 214 utilizes to locate and / or align other fiducials within the image.

[0090] At 304, the sequencing device 214 generates a sub-image of the virtual reference image with larger padding (e.g., padding twice as large as the overall size / area of ​​the fiducial). To ensure that the fiducial is located during search (i.e., the fiducial does not move more than a certain number of pixels, such as, but not limited to, 20 pixels), the sequencing device 214 may generate the sub-image to have padding with a diameter larger than the diameter of the fiducial in the virtual reference image, e.g., 20 pixels. Based on the fiducial configuration parameters, the sequencing device 214 generates the sub-image such that the sub-image is centered or substantially centered on a fiducial (i.e., reference fiducial) in the virtual reference image. In one example, the reference fiducial may be the top-left fiducial in the virtual reference image. In some cases, the reference fiducial may be the reference fiducial for all tiles in the virtual reference image. In other cases, the reference fiducial may be the reference fiducial for each tile in the virtual reference image.

[0091] Upon acquiring the virtual reference image and sub-images, the sequencing device 214 determines 306 a correlation between the virtual reference image and the sub-images to locate the fiducials within various regions of the image. For example, the sequencing device 214 can perform a template matching process and / or an image matching process using the image data and the reference fiducials as points (e.g., unique keypoints) within the image to locate the fiducials within various regions of the image. For example, during an initial imaging cycle, the sequencing device 214 can perform a matching process to search for the reference fiducials within the image. When performing the template matching process, the sequencing device 214 may assume a shape of the object being searched for and incorporate the assumed shape into the search for the reference fiducials. In some cases, the sequencing device 214 may perform a template matching process, such as a cross-correlation process. The cross-correlation process may be, for example, but is not limited to, a Fast Fourier Transform (FFT)-based cross-correlation process to locate the object within a pixel. The sequencing device 214 can fit a two-dimensional Gaussian function to the correlation peaks resulting from the FFT-based cross-correlation process to match the object with sub-pixel accuracy. Thus, the sequencing device 214 can locate each fiducial in the image with sub-pixel accuracy. That is, the sequencing device 214 can apply an FFT to the location of the reference fiducial in the virtual reference image, with a portion of the sub-image containing, for example, the upper-left fiducial. In some other cases, the sequencing device 214 may locate the reference fiducial and other fiducials in the image using one or more of a keypoint detection process, a Harris detector, a scale-invariant feature transform (SIFT), or other similar processes. The template matching process can utilize a template containing information indicating the relative positions of the fiducials. The template can represent the known actual locations of the fiducials.By using the template, fiducials in the image can be localized relative to the localized reference fiducials even if there is some error introduced in the image for one or more fiducial locations.

[0092] As described herein, the sequencing device 214 may be configured to assign a digital value to each reference based on characteristics of the image data represented by the pixel at the corresponding location. The sequencing device 214 assigns corresponding values ​​to the references to alleviate the need to further process the image data itself. In one or more cases, the sequencing device 214 may associate each of the assigned values ​​with a location within an image index or map, which may be done by reference to the known or detected location of the reference or any data encoded by such reference.

[0093] In one or more cases, the sequencing device 214 can determine 308 a correlation score. The correlation score can be a measure of how strong the best correlation between the virtual reference image and the sub-image is relative to the second-best correlation between the virtual reference image and the sub-image. The correlation score can be a result of the output from a cross-correlation process. For example, in some cases, the output of the cross-correlation process can be graphically illustrated to indicate whether the fiducial is located within the image and / or the location of the fiducial. The output of the cross-correlation process for a fiducial associated with two images, such as a virtual reference image and a sub-image, can be illustrated as a single peak. The location of the peak can indicate the location of the fiducial in the real image. Furthermore, by locating the peak, it can be assumed that the fiducial in the virtual image is located at the center of the virtual image.

[0094] The sequencing device 214 can use the correlation score to determine whether the peak of the correlation process is the maximum peak. For example, the sequencing device 214 can locate the highest peak in the output of the cross-correlation process. The location of the highest peak can be recorded as the maximum peak. Additionally, the sequencing device 214 can locate the second highest peak in the output of the cross-correlation process and record the corresponding location as the second highest peak. In some cases, the sequencing device 114 can include a zone around the location of the highest peak (i.e., the maximum peak), and pixels within the zone are excluded from the area searched to locate the second highest peak. The sequencing device 114 can search for the second highest peak at a location outside the zone. In one or more cases, the sequencing device 114 can determine whether the location of the second highest peak is the same as the location of the maximum peak. If the locations of the second highest peak and the maximum peak are the same, the sequencing device 114 determines that there is no peak and outputs a correlation score of 0. If the locations of the second highest peak and the maximum peak are different, the sequencing device 114 generates a correlation score indicating the correlation between the maximum peak and the second highest peak.

[0095] In one or more cases, the sequencing device 214 determines whether the correlation score exceeds a threshold at 310. The threshold may be, for example, but is not limited to, any numerical value (e.g., a threshold score of 0.3) that can be associated with the correlation score indicating a confidence that the reference (e.g., reference image data) is located within the virtual reference image and / or sub-image. If the correlation score is less than or equal to the threshold score (310: No), the sequencing device 214 may determine and maintain the location of the reference reference at 316. Thus, the sequencing device 214 does not change the location of the reference reference indicated by the configuration parameters.

[0096] Referring to 312 of FIG. 3 , the sequencing device 214 may also determine a location of maximum correlation based on the determination of the correlation between the virtual reference image and the sub-image in 306. The location (i.e., translation) of maximum correlation between the virtual reference image and the sub-image may be based on the output result from the cross-correlation process. If the correlation score is greater than a threshold score (310: Yes), the sequencing device 214 adjusts the location of the reference standard in 314. In one or more cases, the sequencing device 214 adjusts the location of the reference standard to match the location of maximum correlation. Further, the sequencing device 214 may adjust the location of the reference standard by adjusting a configuration parameter indicating the location of the reference standard to a parameter corresponding to the location of the position of maximum correlation.

[0097] Based on the determined locations of the reference fiducials (e.g., either the reference fiducial locations adjusted in 314 or the reference fiducial locations maintained in 316), the sequencing device 214 determines and aligns the corresponding locations of the fiducials contained within the image and / or tile. For example, the sequencing device 214 locates one or more other peaks in the output of the cross-correlation process. The sequencing device 214 searches for one or more peaks at locations outside the exclusion zone of the maximum peak. The sequencing device 214 compares the located peaks with the maximum peak and generates an alignment score indicating whether the fiducials are located within the virtual reference image and / or sub-image. Furthermore, based on the located peaks and the determined locations of the reference fiducials, the sequencing device 214 associates the located peaks with the locations of the fiducials within the virtual reference image and / or sub-image. The locations of each fiducial (i.e., fiducial configuration parameters) and associated alignment scores may be stored in a database, such as database 216.

[0098] In one or more cases, the sequencing device 214 can divide the image data acquired in the initial imaging cycle into multiple imaging data subsets (e.g., tiles) corresponding to respective regions of the patterned sample. That is, an imaging data subset can include a subset of pixels of the imaging data set of one imaging cycle. As shown herein, for simplicity, exemplary images 101, 103, and 105 are illustrated in FIGS. 1A-1C from the perspective of one tile (e.g., tiles 102, 104, and 122, respectively). Dividing the image data into multiple tiles enables parallelization of image processing operations. In one or more cases, the size of the imaging data subsets can be determined using the placement of fiducials within the field of view of the imaging system, within the sample, or on the sample. The imaging data subsets can be divided so that the pixels of each imaging data subset or tile have a predetermined number of fiducials (e.g., at least three fiducials, four fiducials, six fiducials, eight fiducials, etc.). For example, the total number of pixels in the imaging data subset may be predetermined based on a predetermined pixel distance between the boundary of the imaging data subset and the fiducial. In another example, for an initial imaging cycle, the imaging data subset may be divided such that the size of the imaging data subset is less than or equal to half the distance between the two fiducials.

[0099] 4 is a flowchart illustrating a procedure 400 for performing an exemplary image distortion correction process. One or more portions of procedure 400 may be performed by one or more computing devices. For example, one or more portions of procedure 400 may be performed by one or more sequencing devices. One or more portions of procedure 400 may be stored in memory as computer-readable or machine-readable instructions that may be executed by a processor of one or more computing devices. Although portions of procedure 400 may be described herein as being performed by a sequencing device, procedure 400 or portions thereof (e.g., system 215 or imaging system 219) may be performed by another computing device or distributed across multiple computing devices, such as one or more sequencing devices, one or more client devices, and / or one or more server devices.

[0100] 4, at 402, the sequencing device 214 may acquire image data associated with a calibration process. In one or more cases, the image data may include fiducial configuration parameters (e.g., locations of fiducials determined during a calibration process, such as the process described in procedure 300). Additionally or alternatively, the image data may include subregion parameters that define one or more subsets of fiducials included in the image and / or tile.

[0101] The sequencing device 214 can determine 404 a plurality of subregions. In one or more cases, each subregion may include a subset of the image and / or criteria included in the tile. For example, as illustrated in FIG. 5A, tile 102 includes subregions 502a, 502b, and 502c. In one or more cases, the sequencing device 214 determines the plurality of subregions based on the criteria configuration parameters and the subregion parameters.

[0102] The subregion parameters may be based on one or both of the number of reference rows per subregion and the number of overlapping reference rows between consecutive subregions. For example, as illustrated in FIG. 5A, tile 101 of image 102 may include two rows per subregion and one overlapping row. In another example, as illustrated in FIG. 6A, tile 104 of image 103 may include eight subregions (subregions 602a, 602b, 602c, 602d, 602e, 602f, 602g, and 602h), with the number of rows per region being two and the number of overlapping rows being one. In another example, as illustrated in FIG. 6B, tile 104 of image 103 may include four subregions (subregions 606a, 606b, 606c, and 606d), with the number of rows per region being four and the number of overlapping rows being two. The sequencing device 214 can determine a set of local criteria based on the criteria configuration parameters and the subregion parameters and associate them with each subregion. In one or more cases, the sequencing device 214 divides the well locations of the flow cell. For example, the sequencing device 214 divides the well locations by associating the well locations with adjacent criteria. Based on the association with the adjacent criteria, the sequencing device 214 further associates the well locations with the subregions associated with each criteria. That is, the sequencing device 214 divides the well locations and assigns the well locations to their corresponding local subregions. For example, the sequencing device 214 may utilize an even horizontal division of the image across the subregions. The sequencing device 214 can assign an equal number of rows of wells to each subregion. In some cases, the sequencing device 214 can assign upper and lower subregions to include rows of wells near the top and bottom edges of the tile. In one or more cases, the sequencing device 214 can define the horizontal boundaries of the subregions as extending between the left and right sides of the tile, and can determine the vertical boundaries of the subregions by dividing the height of the tile by the number of subregions to be considered.

[0103] In one or more cases, the sequencing device 214 can construct the subregions so that they are linearly arranged on the image. In some cases, the sequencing device 214 constructs the subregions so that they are contiguous, i.e., one subregion is adjacent to another subregion (e.g., subregion 502c is adjacent to subregion 502b). In other cases, the sequencing device 214 constructs the subregions so that some subregions are contiguous, with one or more subregions overlapping each other.

[0104] In one or more cases, sequencing device 214 can structure the subregions such that a criterion can be associated with more than one subregion. For example, sequencing device 214 can structure subregion 502c to be associated with criterion 107a and criterion 107b, and can structure subregion 502b to be associated with criterion 107a and criterion 107b. In one or more other cases, sequencing device 214 can structure the subregions such that adjacent subregions do not share a criterion.

[0105] The sequencing device 214 can initialize a transformation for each subregion. The sequencing device 214 can estimate a local transformation matrix based on multiple parameters, including, but not limited to, translation parameters in both the X and Y directions, scale parameters, and shear parameters. For example, the sequencing device 214 can initialize a default affine transformation matrix, where the scale and shear coefficients of the default affine transformation matrix may be initialized to zero, and the translation coefficients may be initialized to one for both the X and Y directions. The sequencing device 214 can construct a transformation matrix (e.g., an affine transformation matrix) based on a reference location. As illustrated in FIG. 5A , the sequencing device 214 constructs an affine transformation matrix using a reference location that is local to each subregion, such as subregions 502a, 502b, and 502c of the tile 102. That is, the sequencing device 214 constructs an affine transformation matrix for each subregion using a reference location associated with each subregion.

[0106] In one or more cases, the sequencing device 214 constructs an affine transformation matrix using at least three fiducials. For example, the sequencing device 214 constructs a subregion 502c based on the locations of fiducials 106a, 106b, 107a, and 107b. In this example subregion, the sequencing device 214 can determine that fiducials 106a, 106b, 107a, and 107b are arranged in a 2x2 pattern within the subregion 502c. Note that FIG. 5A illustrates an image and fiducials decomposed into three subregions. It is also understood that the image and fiducials may be decomposed into any number of subregions. Furthermore, it is noted that any of a variety of geometric transformation models may be used, such as, but not limited to, a linear transformation or an affine transformation. The transformation may include, for example, one or more of rotation, translation, scaling, shear, etc. Once the affine transformation has been initialized for each sub-region, the sequencing device 214 proceeds to repeat the procedure 400 for each cycle, tile by tile, sub-region by sub-region.

[0107] 4, sequencing device 214 obtains a registration score for each reference location associated with the subregion at 405. For example, sequencing device 214 obtains the registration scores from a database, such as database 216. Sequencing device 214 may obtain the registration scores determined from an image calibration process, such as the process described in procedure 300. Sequencing device 214 may determine at 406 whether at least three registration scores exceed a threshold score.

[0108] If the three registration scores for each reference location are equal to or less than the threshold score (406: NO), the sequencing device 214 provides an indication of registration failure at 408. The sequencing device 214 provides an indication of registration failure by associating each subregion with an indication of image registration failure for the subregion. For example, the sequencing device 214 can set a local failure flag for each tile and provide a -9999 translation. A registration failure can occur when fewer than three references can be found accurately. For example, the sequencing device 214 can determine that registration has failed when the three scores for each reference location are equal to or less than the threshold score (i.e., the sequencing device 214 was unable to locate or accurately locate at least three references in each subregion). Thus, the sequencing device 214 can determine that a base call cannot be made for the associated tile in each cycle. A misalignment can also occur when three fiducials are collinear (i.e., the three fiducials are aligned in a single column or a single row). For example, in some cases, the sequencing device 214 may determine that the three scores for each fiducial location are greater than a threshold score, but the sequencing device 214 may further determine that the fiducial locations are collinear. In such cases, the sequencing device 214 provides an indication of a misalignment at 408.

[0109] If each of the three alignment scores is greater than the threshold score (406: Yes), the sequencing device 214 performs a transformation process at 410. The sequencing device 214 may apply a transformation process to transform the locations of the fiducials associated with each subregion. In one or more cases, the sequencing device 214 may perform the transformation process on each subregion and the fiducials having locations associated with each subregion. The sequencing device 214 performs the transformation process to generate respective local transformations associated with each subregion (i.e., transformed locations of the fiducials associated with each subregion). For example, the sequencing device 214 may apply a transformation process (e.g., an affine transformation) to subregion 502c and the locations of fiducials 106a, 106b, 107a, and 107b to generate locally transformed locations of fiducials 106a, 106b, 107a, and 107b associated with subregion 502c, as illustrated in FIG. 5B.

[0110] The sequencing device 214 can correct optical distortion of the image at 412. For example, to correct optical distortion of the image, the sequencing device 214 applies a nonlinear distortion process to each reference location in the image. The sequencing device 214 can determine intensity values ​​of the image, e.g., pixel intensity values, at 414. In one or more cases, the sequencing device 214 can determine the intensity values ​​of the image by sharpening the image using, for example, an adaptive equalizer. To sharpen the image, the sequencing device 214 can perform equalization by convolving the image with a channel-specific mask. The sequencing device 214 can apply the same mask to different subregions of the image. The sequencing device 214 can apply different masks to different subregions of the image.

[0111] To extract intensity values, e.g., pixel intensity values, of the sharpened image at each location, the sequencing device 214 can interpolate the equalized pixel intensities within a square centered at that location. The sequencing device 214 can use an interpolation method, such as, but not limited to, a bilinear method or a Hamming method, across several pixels. The sequencing device 214 can spatially normalize the image tiles for each subregion. For example, the sequencing device 214 can spatially normalize the image tiles for each subregion so that the 90th and 10th percentiles of the extracted intensities of the normalized tiles are equal. The sequencing device 214 can compress the intensities so that the 1st and 99th percentiles of the extracted pixel intensities have pixel intensity values ​​of 1 and 255, respectively. Lossy compression of the intensities can reduce memory load. In one or more cases, the sequencing device 214 can store an array of pixel intensity values ​​of the cluster in memory.

[0112] In one or more cases, sequencing device 214 is configured to repeat procedure 400 for each tile in the image. When a cycle associated with an image is completed, sequencing device 214 can begin a subsequent imaging cycle for the patterned sample and generate new image data for the patterned sample. Thus, sequencing device 214 is configured to repeat procedure 400 for the new image data, as described herein.

[0113] The sequencing device 214 can be configured to perform base calling to determine a base call (e.g., adenine (A), cytosine (C), guanine (G), or thymine (T)) for a given spot location in an image during an imaging cycle. For example, during two-channel base calling, the sequencing device 214 can use image data extracted from two images to determine the presence of one of four base types by encoding the base identity as a combination of the intensities of the two images. For example, for a given spot or location in each of the two images, the sequencing device 214 can be configured to determine the base identity based on whether the signal identity combination is [on, on], [on, off], [off, on], or [off, off].

[0114] Base calling can be performed by fitting a mathematical model to the intensity data. The mathematical model can be, for example, but not limited to, a k-means clustering algorithm, a k-means-like clustering algorithm, an expectation maximization clustering algorithm, a histogram-based method, or the like. In one example, the sequencing device 214 can fit four Gaussian distributions to a set of two-channel intensity data, such that one distribution applies to each of the four nucleotides represented in the dataset. In one example, the sequencing device 214 can apply an expectation maximization (EM) algorithm to the intensity data. As a result of the EM algorithm, for each X, Y value (i.e., each referring to each of the two channel intensities), a value can be generated that represents the likelihood that a particular X, Y intensity value belongs to one of the four Gaussian distributions to which the data is fitted. If four bases provide four distinct distributions, each X, Y intensity value can also have four associated likelihood values, one for each of the four bases. The maximum of the four likelihood values ​​indicates the base call. For example, if a cluster is "off" in both channels, the base call is G. If a cluster is "off" in one channel and "on" in another, the base call is either C or T (i.e., depending on which channel is "on"). If a cluster is "on" in both channels, the base call is A.

[0115] 7A illustrates a simulation comparing the effect of jitter on image registration using global affine transformation process 702 and process 704 (e.g., procedure 200) as described herein. A "0" on the jitter axis of the simulation indicates no jitter. A "1" on the jitter axis of the simulation indicates actual jitter. A "2" on the jitter axis of the simulation indicates twice the jitter. A "3" on the jitter axis of the simulation indicates three times the jitter. As illustrated, image registration processes 704a, 704b, and 704c (e.g., procedure 200) utilizing a 12x2x12 reference arrangement, a 9x2x9 reference arrangement, and a 4x4 reference arrangement, respectively, recover much of the performance loss demonstrated in image registration processes utilizing global affine transformation processes 702a, 702b, and 702c having the same reference arrangements. That is, the image registration processes 704a, 704b, and 704c (eg, procedure 200) can tolerate a higher level of jitter than the global affine transformation processes 702a, 702b, and 702c.

[0116] Figure 7B illustrates the results of an actual run using a 4x4 reference configuration. Figure 7B summarizes the improvement over six runs, showing improvements in error rate ranging from 2.7% to 17%, with an average improvement of 8%. Thus, the results show that the 4x4 reference configuration provides consistent improvements with realistic stage jitter.

[0117] 7C and 7D illustrate spatial representations of error rates across an exemplary sequencing image. FIG. 7C illustrates a spatial representation of the error rate of a global affine transformation process. FIG. 7D illustrates a spatial representation of the error rate of image registration using procedure 200. Using the global affine transformation process, FIG. 7C illustrates error rate bands 708, a periodic pattern caused by stage jitter. However, FIG. 7D illustrates that image registration using procedure 200 corrects these same bands of error. Furthermore, not all tiles may be affected by the error rate bands due to jitter. Some may be more sensitive to certain cycles, causing distortion correction to fail, resulting in band 710 appearing on the left side in both global affine transformations, but not present in the spatial representation of FIG. 7D.

[0118] 8 is a block diagram of an exemplary computing device 800. One or more computing devices, such as computing device 800, may implement one or more functions for generating and / or processing sequencing tasks as described herein. For example, computing device 800 may include one or more of sequencing device 214, client device 208, and / or server device 202 shown in FIG. 2A. As shown by FIG. 8, computing device 800 may include a processor 802, memory 804, a storage device 806, and an I / O interface 808 and / or a communication interface 810, which may be communicatively coupled by a communication infrastructure 812. It should be understood that computing device 800 may include fewer or more components than those shown in FIG. 8.

[0119] The processor 802 may include hardware for executing instructions, such as those comprising a computer program. In an example, to execute instructions for dynamically altering a workflow, the processor 802 may retrieve (or fetch) instructions from an internal register, an internal cache, memory 804, or storage device 806, decode, and execute the instructions. The memory 804 may be volatile or non-volatile memory used to store data, metadata, computer-readable or machine-readable instructions, and / or programs for execution by the processor to operate as described herein. The storage device 806 may include storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for performing the methods described herein.

[0120] I / O interface 808 may enable a user to provide input to, receive output from, and / or otherwise transfer data to and receive data from computing device 800. I / O interface 808 may include a mouse, a keypad or keyboard, a touchscreen, a camera, an optical scanner, a network interface, a modem, other known I / O devices, or a combination of such I / O interfaces. I / O interface 808 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. I / O interface 808 may be configured to provide graphical data to a display for presentation to a user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content.

[0121] Communications interface 810 may include hardware, software, or both. In any case, communications interface 810 may provide one or more interfaces for communications (e.g., packet-based communications, etc.) between computing device 800 and one or more other computing devices or networks. The communications may be wired or wireless. By way of example and not limitation, communications interface 810 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wired-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network such as Wi-Fi.

[0122] Additionally, communication interface 810 can facilitate communication with various types of wired or wireless networks. Communication interface 810 can also facilitate communication using various communication protocols. Communication infrastructure 812 can also include hardware, software, or both that couple components of computing device 800 to one another. For example, communication interface 810 may use one or more networks and / or protocols to enable multiple computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein. For example, a sequencing process may enable multiple devices (e.g., client device, sequencing device, and server device) to exchange information such as sequencing data and error notifications.

[0123] In addition to what is described herein, the methods and systems may also be implemented in, for example, a computer program, software, or firmware embodied in one or more computer-readable media for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted over wired or wireless connections) and tangible / non-transitory computer-readable storage media. Examples of tangible / non-transitory computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), removable disks, and optical media such as CD-ROM disks and digital versatile disks (DVDs).

[0124] As used herein, the term "about" in reference to a numerical value means the numerical value itself, or the numerical value itself plus or minus 10% of the numerical value of the number with which the term is used.

[0125] While the present disclosure has been described with respect to particular embodiments and generally associated methods, modifications and permutations of the embodiments and methods will be apparent to those skilled in the art. Accordingly, the above description of exemplary embodiments does not constrain the present disclosure. Other changes, substitutions, and alterations are also possible without departing from the spirit and scope of the present disclosure. [Explanation of symbols]

[0126] 101 images 102 tiles 103 images 104 tiles 105 images 106 Standards 107 Standards 108 sets 110 sets 112 Standards 114 Standards 116 sets 118 Standards 120 Groups 122 tiles 200 System Environment 202 Server Device 204 Sequencing System 208 client devices 210 Sequencing Applications 212 Network 214 Sequencing Device 215 Detector Subsystem 216 databases 217 Focus Components 219 Line Scanning Imaging System 220 Waste Valve 225 flow cell 230 Temperature Station Actuator 235 Cooler 240 Camera System 242 Objective Lens 245 Filter Switching Assembly 250 light source 261 Focused Laser 265 Low Wattage Lamp 270 Movable Staging Area 271 filter elements 272 Line Generation Module (LGM) 273 Tube Lens 274 Light source 275 Emission Optics Module (EOM) 276 Light source 278 Semi-reflective mirror 280 Mirror 281 Focus Tracking Module (FTM) 282 Semi-reflective mirror 283 Optical Sensor 284 shutter element 285 Camera module (CAM) 286 Beam Shaping Lens 288 Translucent Cover Plate 290 Liquid layer 291 Real-time Analysis Module 292 Target 296 z-stage 298 Objective Lens 702 Global Affine Transformation Process 704 Process 708 Error Rate Band 710 band 800 computing devices 802 processor 804 memory 806 Storage Devices 808 I / O Interface 810 Communication Interface 812 Communications Infrastructure 900 tiles 902 Standard 904 plots 906 area 908 X direction 910 area 912 Y direction

Claims

1. 1. A sequencing system comprising: an image capture device configured to capture an image including at least one feature and a plurality of fiducials, the plurality of fiducials being arranged in a pattern; a computing device comprising a processor and a memory; wherein the processor and memory determining a plurality of sub-regions of the image, each sub-region comprising a subset of the criteria contained in the image; performing a geometric transformation on each subregion to generate a respective local transformation associated with each subregion; aligning respective locations of the fiducials included in the image based on the respective local transformations associated with each subregion, the size of each subregion being selected such that each subregion is substantially invariant to stage jitter; A sequencing system configured to:

2. the processor and memory The sequencing system of claim 1 , configured to determine a respective location for each of the fiducials in the image based on the determined location of a reference fiducial.

3. the processor and memory generating a sub-image based on the acquired image, the sub-image including a pixel padding having a larger area of ​​the reference fiducial included in the image; determining a correlation between the acquired image and the sub-image; Deciding whether to adjust the location of the reference standard based on the determined correlation. The sequencing system of claim 2 , configured as follows:

4. The sequencing system of claim 1 , wherein each subset of criteria included in a subregion includes at least three criteria.

5. The sequencing system of claim 1 , wherein the subregions are linearly arranged relative to one another.

6. The sequencing system of claim 1 , wherein the geometric transformation comprises an affine transformation.

7. 1. A computer-implemented method comprising: acquiring an image including at least one feature and a plurality of fiducials, the plurality of fiducials being arranged in a pattern; determining a plurality of sub-regions of the image, each sub-region comprising a subset of the criteria contained in the image; performing a geometric transformation on each subregion to generate a respective local transformation associated with each subregion; aligning respective locations of the fiducials included in the image based on the respective local transformations associated with each subregion, the size of each subregion being selected such that each subregion is substantially invariant to stage jitter; A method comprising:

8. The computer-implemented method of claim 7 , wherein the sub-regions are linearly arranged relative to one another.

9. The computer-implemented method of claim 7 , wherein each subset of criteria included in a subregion includes at least three criteria.

10. 10. The computer-implemented method of claim 7, wherein reference fiducials are located within the image, the method of claim 9 further comprising determining a respective location of each of the fiducials within the image based on the determined locations of the reference fiducials.

11. generating a sub-image based on the acquired image, the sub-image including a pixel padding having a larger area than the reference fiducial included in the image; determining a correlation between the acquired image and the sub-image; determining whether to adjust the location of the reference standard based on the determined correlation; The computer-implemented method of claim 10 further comprising:

12. 8. The computer-implemented method of claim 7, wherein the size of each subregion is selected such that each subregion is substantially invariant to stage jitter of about 200 Hz or less.

13. The computer-implemented method of claim 7 , wherein at least one criterion is associated with two adjacent sub-regions.

14. 8. The computer-implemented method of claim 7, wherein the pattern of the plurality of fiducials includes a first set of linearly arranged fiducials adjacent to a second set of linearly arranged fiducials, the first set of linearly arranged fiducials and the second set of linearly arranged fiducials extending substantially parallel to one another.

15. The computer-implemented method of claim 7 , wherein the image comprises a sequencing image.

16. The computer-implemented method of claim 7 , wherein the feature locations are associated with flow cell well locations.

17. 8. The computer-implemented method of claim 7, wherein each subregion contains a respective feature, and the respective local transform associated with the subregion containing the respective feature is used to determine the location of the respective feature.

18. A non-transitory computer-readable medium containing computer-readable instructions that, when executed by a processor, cause the processor to: acquiring an image including at least one feature and a plurality of fiducials, the plurality of fiducials being arranged in a pattern; determining a plurality of sub-regions of the image, each sub-region containing a subset of the criteria contained in the image; performing a geometric transformation on each subregion to generate a respective local transformation associated with each subregion; aligning respective locations of the fiducials included in the image based on the respective local transformations associated with each subregion, the size of each subregion being selected such that each subregion is substantially invariant to stage jitter; A non-transitory computer-readable medium having a method implemented therein.

19. The non-transitory computer-readable medium of claim 8 , wherein the sub-regions are linearly arranged relative to one another.

20. 20. The non-transitory computer-readable medium of claim 18, wherein each subset of criteria included in a sub-region includes at least three criteria.

21. 20. The non-transitory computer-readable medium of claim 18, wherein reference fiducials are located within the image, the locations of the reference fiducials being used to determine a respective location of each of the fiducials within the image.

22. generating a sub-image based on the acquired image, the sub-image including a pixel padding having a larger area than the reference fiducial included in the image; determining a correlation between the acquired image and the sub-image; 22. The non-transitory computer-readable medium of claim 21, further comprising: determining whether to adjust a location of the reference standard based on the determined correlation.

23. 20. The non-transitory computer-readable medium of claim 18, wherein the size of each sub-region is selected such that each sub-region is substantially invariant to stage jitter of about 200 Hz or less.

24. 20. The non-transitory computer-readable medium of claim 18, wherein at least one criterion is associated with two adjacent sub-regions.

25. 20. The non-transitory computer-readable medium of claim 18, wherein the pattern of the plurality of fiducials includes a first set of linearly arranged fiducials adjacent to a second set of linearly arranged fiducials, the first set of linearly arranged fiducials and the second set of linearly arranged fiducials extending substantially parallel to one another.

26. 20. The non-transitory computer-readable medium of claim 18, wherein the image comprises a sequencing image.

27. 20. The non-transitory computer-readable medium of claim 18, wherein the feature locations are associated with flow cell well locations.

28. 20. The non-transitory computer-readable medium of claim 18, wherein each subregion contains a respective feature, and wherein the respective local transform associated with the subregion containing the respective feature is used to determine the location of the respective feature.