Image registration method, method for detecting image recognition basic group and computer equipment
By combining coarse and fine registration methods and utilizing benchmark templates and bright spot features, the problem of insufficient image registration accuracy in patterned surface microscopic imaging nucleic acid sequencing systems was solved, achieving high-precision determination of base positions and sequences.
Patent Information
- Application Number
- CN202511423061.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-31
- Filing Date
- 2025-09-29
- Publication Date
- 2025-12-12
Smart Images

Figure CN121120728A_ABST
Abstract
Description
[0001] This application claims priority to Chinese Patent Application No. 202411549366.4, filed on October 31, 2024, entitled “Method, Apparatus, Computer Equipment and Storage Medium for Base Recognition”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of image processing technology, and in particular to an image registration method, a method for detecting and recognizing bases in an image, and a computer device. Background Technology
[0003] The topics discussed in this section should not be considered prior art simply because they are mentioned here. Similarly, the technical problems mentioned in this section or related to the topics provided as background art should not be considered as having been previously recognized in the prior art. The topics in this section merely represent different methods, which themselves may correspond to specific embodiments of the technical solutions included in the claims.
[0004] In related technologies, sequencing generally refers to the determination of the primary structure or sequence of biopolymers, including nucleic acids such as DNA and RNA. This includes determining the order of nucleotide bases (adenine A, guanine G, thymine T / uracil U, and cytosine C) in a given nucleic acid fragment. Such methods typically involve base calling, which identifies bases at one or more positions in the nucleic acid to determine at least a segment or portion of the nucleic acid molecule's sequence.
[0005] The signal and / or signal intensity changes directly or indirectly generated by the binding or ligation of a nucleotide / base to a specific location on a nucleic acid molecule (template) can indicate the type of base at that location on the nucleic acid molecule. For example, different fluorescent molecules can be used to identify different incorporated or linked nucleotides / bases. The binding or ligation of a nucleotide / base to a specific location on a nucleic acid molecule is also called nucleotide / base incorporation into the nucleic acid molecule or base extension, and can be achieved, for example, through polymerization, ligation, and hybridization.
[0006] Specifically, in detection systems that use surface microscopy for nucleic acid sequencing, nucleic acid molecules are typically attached to a solid surface. Solid surfaces containing chambers capable of holding solutions are often referred to as chips or flow cells. Based on the solid surface layout design or fabrication process, or the distribution of the nucleic acid molecules to be tested on the solid surface, surfaces can be categorized into patterned surfaces (or regular arrays) and random surfaces (random arrays). During the acquisition of signals from any type of surface array, such as during imaging, it is common to move related hardware, such as moving the objective lens and / or moving the solid surface, to acquire images of the same object (nucleic acid molecules in the same surface region / field of view) after relevant biochemical reactions at different times. The multiple images acquired at different times are then processed, including signals identifying the positions of corresponding chemical features (e.g., nucleic acid molecules undergoing extension reactions) in the images, to determine at least a portion of the sequence information of the nucleic acid molecules to be tested in that field of view.
[0007] In comparison, image registration of patterned surfaces is generally considered simpler than that of random surfaces because the former introduces regular patterns during chip design or fabrication. This means that the spatial layout of relevant regions or sites is known before imaging and can serve as a reference image, template, or global coordinate system. Therefore, by identifying pattern features on the image and aligning it to the reference image, one or more images of the surface's field of view can be mapped or unified to the same coordinate system as the reference image, thus achieving registration.
[0008] However, in actual image acquisition, due to the precision required for hardware movement and the potential for force application or application causing positional shifts or shape changes in related mechanical structures, and / or the possibility of temperature increases or decreases during sequencing biochemical reactions affecting the shape, surface properties, or quantity, morphology, or position of some nucleic acid molecules on the solid surface, it is inevitable that at least some nucleic acid molecules in the same field of view will have positional shifts or inconsistencies across multiple images. This makes it difficult to accurately locate the true position (chemical characteristics) of the nucleic acid molecules undergoing the relevant biochemical reaction directly from the acquired images. Therefore, it is generally necessary to register multiple images of the same object to place them in the same coordinate system. This allows for the location of the nucleic acid molecules undergoing base extension reactions using the registered images, and then the type of incorporated or extended bases can be identified based on signal changes at these positions. Review articles or patent documents such as Zitová, B. & Flusser, J. Image registration methods: a survey. Image and Vision Computing 21(11), 977–1000 (2003), CN112285070A, EP3336797A1 and CN112288781A disclose image correction or image registration methods.
[0009] Improving image registration accuracy facilitates the precise localization of corresponding chemical feature positions on the image and the extraction of signals from those positions, thereby enabling the accurate identification of newly formed or extended bases or sequences at those positions. Therefore, improving the accuracy and / or efficiency of image registration remains a technically important issue. Summary of the Invention
[0010] This application aims to at least partially solve one of the aforementioned technical problems or provide a practical commercial solution. Specifically, embodiments of this application provide an image registration method, an image registration system, and a method for detecting and recognizing bases in an image, as well as related computer program products or systems, as detailed below:
[0011] The first aspect of this application provides an image registration method, wherein the image is from a detection system that performs sequencing based on patterned surface microscopy, comprising: (S10) performing coarse registration of the image based on a reference template, including: determining the positions of multiple reference points in the image based on the reference positions of corresponding reference points in the reference template, establishing a fitting relationship based on the positions of the multiple reference points, and updating the positions of the reference points in the image based on the fitting relationship; wherein the reference template includes a reference coordinate system and reflects the same surface region as the image; the surface region includes multiple adjacent labeled regions and reaction regions, wherein the labeled regions have multiple regularly distributed, separate reactive sites in two or more non-parallel directions, and the intersection of these directions is located in the labeled regions; the labeled regions in their The detection signal curves of the regularly distributed reactive sites in each direction and the adjacent reactive areas all exhibit specific intensity change characteristics; the reference point is located at the intersection of each direction of the marked area and is located by intensity change characteristics; at least a portion of the reactive sites appear as bright spots in the image, and the number of bright spots is greater than the number of reference points; and (S20) fine registration of the image is performed based on the coarse registration result and the bright spots, including: determining the distribution of bright spots in each direction of the marked area, and determining the fitting relationship of each direction based on the distribution of bright spots, and further updating the reference point position of the image based on the fitting relationship; and determining the coordinates of other areas or positions of the image based on the updated reference point position, so as to achieve the registration of the image with the reference template.
[0012] This application also provides an image registration system capable of implementing the image registration method or steps in any of the above embodiments. The system includes a coarse registration module and a fine registration module connected together. The coarse registration module is configured to perform coarse registration of an image based on a reference template, including: determining the positions of multiple reference points in the image based on the reference positions of corresponding reference points in the reference template; establishing a fitting relationship based on the positions of the multiple reference points; and updating the positions of the reference points in the image based on the fitting relationship. The reference template includes a reference coordinate system and reflects a surface region identical to the image. The surface region includes multiple adjacent marker regions and reaction regions. The marker regions have multiple regularly distributed, separate reactive sites in two or more non-parallel directions, and the intersection of these multiple directions is located within the marker regions. The marker regions exhibit specific intensity variation characteristics on the detection signal curves of adjacent reaction regions in each direction where the regularly distributed reactive sites are located. The reference points are located at the intersection of these directions of the marker regions and are located by the intensity variation characteristics. At least a portion of the reactive sites appear as bright spots in the image, and the number of bright spots is greater than the number of reference points. The fine registration module is configured to perform fine registration of the image based on the coarse registration result and the bright spots, including: determining the distribution of bright spots in each direction of the marked area, determining the fitting relationship in each direction based on the distribution of bright spots, and further updating the reference point position of the image based on the fitting relationship; and determining the coordinates of other areas or positions of the image based on the updated reference point position, so as to achieve the registration of the image with the reference template.
[0013] A second aspect of this application provides a method for detecting bases in an image, the image being from a detection system that performs sequencing based on patterned surface microscopy, the surface region reflected in the image including multiple adjacent labeled regions and reaction regions, the method comprising: determining the signal intensity value of a position corresponding to a chemical feature in the reaction region of the image, wherein the image is an image registered by an image registration method or system according to any embodiment of this application; and identifying the type of base introduced at that position based on the signal intensity value.
[0014] A third aspect of this application provides a computer device, comprising: a memory for storing data, including a computer-executable program; and a processor for executing the computer-executable program, wherein executing the computer-executable program includes implementing the image registration method and / or the method for detecting image recognition bases in any of the above embodiments.
[0015] This application also provides a computer-readable storage medium for storing a computer-executable program, wherein executing the computer-executable program includes performing the image registration method or base recognition method in any of the above embodiments or examples.
[0016] In any of the above embodiments, the image registration method, system, base identification method, and computer product or device use an image from a detection system that performs sequencing based on patterned surface microscopy. The surface region reflected in the image is a regular patterned surface, including multiple regularly arranged adjacent marker regions and reaction regions. Each marker region includes at least two non-parallel extension directions, and multiple reactive sites are regularly arranged in each direction. The registration method, system, or device includes coarse registration of the image based on a reference template, and fine registration of the image based on the coarse registration result and bright spots. Specifically, in the coarse registration stage, the reference template is used to determine the region range where each reference point is likely to be located, and the reference points in each region range are quickly located by the specific intensity change characteristics shown by the marker regions and their adjacent reaction regions in each direction. The positions of the reference points are then coarsely corrected to achieve pixel-level registration. In the fine registration stage, the position of the reference points is further corrected by the distribution of bright spots within the marker regions to achieve finer-grained, such as sub-pixel-level image registration. The methods, systems, or devices described in these embodiments utilize the layout information of patterned surface regions, such as shape and pattern characteristics, as well as the information or signal differences or changes reflected in the image by biochemical reaction signals, to perform coarse registration through global features such as the distribution of marker areas or reference points in the marker areas on the surface region, and then perform fine registration through local features such as the distribution of bright spots near the reference points. This enables accurate and reliable determination of the position of the reference points, and further determination of the coordinates of other areas or positions of the surface region reflected in the image, thereby achieving high-precision registration between the image and the reference template. Attached Figure Description
[0017] The above-described embodiments and / or additional technical features and advantages of this application will become apparent and readily understood in conjunction with the following description of the embodiments in conjunction with the accompanying drawings.
[0018] Figure 1 This is a schematic flowchart of the image registration method according to an embodiment of this application;
[0019] Figure 2 This is a schematic diagram of the surface pattern of an embodiment of this application;
[0020] Figure 3 This is a schematic diagram of the grid surface pattern in an embodiment of this application;
[0021] Figure 4 This is a schematic diagram of the surface pattern of an embodiment of this application;
[0022] Figure 5 This refers to the range of the intersection of non-parallel directions of the marked area in the image of this application embodiment and its signal detection curve;
[0023] Figure 6 A schematic diagram showing the W-shaped intensity change characteristics on the detection signal curve provided in the embodiments of this application;
[0024] Figure 7 A schematic diagram of the coarse registration process provided for an embodiment of this application;
[0025] Figure 8 A partial image of the marked area and each reaction area adjacent to it in each direction, provided for an embodiment of this application;
[0026] Figure 9 A schematic diagram of the fine registration process provided for embodiments of this application;
[0027] Figure 10 This is a flowchart illustrating the method for detecting and recognizing bases in an image provided in an embodiment of this application. Detailed Implementation
[0028] The present invention will be further described in detail below with reference to specific embodiments and accompanying drawings. Similar elements in different embodiments are referred to by the same or similar element designations. In the following embodiments, many details are described to enable a better understanding of this application. Those skilled in the art will recognize that some features may be omitted or replaced by other elements, materials, means, or methods in different embodiments. Furthermore, for processes or operational details not explained or described in detail in some embodiments, those skilled in the art can understand or implement them based on the context, related examples, and / or conventional technical knowledge in the field.
[0029] Furthermore, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can be rearranged or adjusted in a manner that can be foreseen by those skilled in the art. In other words, unless otherwise stated that a particular order must be followed, the various orders in the specification and drawings are for the purpose of clearly describing a particular embodiment and do not necessarily imply a fixed order of operation or implementation.
[0030] In this document, terms such as "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number or order of the indicated technical features. Unless otherwise stated, the singular forms of "a," "an," etc., introduced particularly due to linguistic translation conventions, include their plural referents, i.e., one or more; "a group" or "a plurality" refers to two or more. Furthermore, "comprising," "having," and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units but may also include steps or units not listed.
[0031] Unless otherwise stated, the terms "connected" and "linked" as used herein should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection, an electrical connection, or a connection that allows communication; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal connection of two components or the interaction between two components; they can refer to a connection through physical adsorption or other forces, or a chemical connection through chemical bonds, such as incorporation via polymerization reactions. Those skilled in the art will understand the specific meaning of these terms in the relevant examples based on the specific embodiments described, including the context and conventional understanding.
[0032] Please refer to Figure 1 This application provides an image registration method in certain embodiments, wherein the image to be registered comes from a detection system that performs sequencing based on patterned surface microscopy, and the image is accurately registered through coarse registration and fine registration.
[0033] Specifically, in some embodiments, the method includes: (S10) performing coarse registration of the image based on a reference template, including determining the positions of multiple reference points in the image based on the reference positions of corresponding reference points in the reference template, establishing a fitting relationship based on the positions of the multiple reference points, and updating the positions of the reference points in the image based on the fitting relationship; wherein the reference template includes a reference coordinate system and reflects the same surface region as the image; the surface region includes multiple adjacent marker regions and reaction regions, wherein the marker regions have multiple regularly distributed, reactive sites in two or more non-parallel directions, and the intersection of these directions is located in the marker regions; the marker regions have multiple regularly distributed reactive sites in each direction of their respective directions. The detection signal curves of adjacent reaction zones all exhibit specific intensity change characteristics; the so-called reference points are located at the intersection of various directions of the marked area and are located by the so-called specific intensity change characteristics; at least a portion of the so-called reactive sites appear as bright spots in the image, and the number of bright spots is greater than the number of reference points; and (S20) the image is finely registered based on the coarse registration results and the bright spots, including determining the distribution of bright spots in various directions of the marked area, and determining the fitting relationship in each direction based on the distribution of bright spots, and further updating the reference point position of the image based on such fitting relationship; and determining the coordinates of other areas or positions of the image based on the updated reference point position, so as to achieve the registration of the image with the reference template.
[0034] The so-called detection system based on patterned surface microscopic imaging for sequencing, also known as a sequencing system, refers to a detection system that converts relevant biochemical signals into image information through optical microscopic imaging, and determines the base sequence of the primary structure of nucleic acid molecules based on the processing and analysis of the image information. Such detection systems include microscopic imaging devices or systems that can acquire luminescence information from surfaces and image surface regions connected to the nucleic acid molecules to be tested; they can be implemented using sequencing principles such as sequencing by synthesis (SBS), ligation sequencing (SBL), or hybridization sequencing (SBH). Optional detection systems include, but are not limited to, Illumina's HiSeq. TM Miseq TM Nextseq TM and Novaseq TM The sequencing platform, IonTorrent from Thermo Fisher / Life Technologies. TM Platform, BGISEQ of BGI Genomics TM and MGISEQ TM / DNBSEQ TM The platform includes a single-molecule sequencing platform with a microscopic imaging system, as well as a self-built sequencing platform with a microscopic imaging system. Sequencing methods can be selected from single-end sequencing, paired-end sequencing, or any sequencing method supported by the constructed or selected sequencing platform.
[0035] Regarding the patterned surface—the object of microscopic imaging in the detection system, also referred to herein as a patterned array—it is a regular array, a regular solid-phase substrate surface carrying the nucleic acid molecules to be tested. This regular array can be loaded and unloaded and connected to the detection system. Thousands or more separation sites or sites for connecting the nucleic acid molecules to be tested are regularly arranged in the surface region or the corresponding imaging field of view; these separation sites or sites are also referred to as reactive sites in some embodiments of this application. It is understood that during sequencing imaging detection, at least a portion of the reactive sites connected to the nucleic acid molecules to be tested appear as bright spots or bright patches on the image, that is, spots with signal intensity (pixel value or grayscale value) greater than the image background.
[0036] The nucleic acid molecule to be tested, also known as the nucleic acid template, can be an unamplified single molecule or an amplified cluster or long chain containing multiple identical polynucleotide molecules, such as fluorescent clusters or DNA nanospheres (DNBs) formed by bridge amplification or rolling circle amplification used by current mainstream commercial sequencing platforms. The nucleic acid molecule to be tested can be single-stranded, double-stranded, and / or a complex hybridized with probes or primers.
[0037] Please refer to Figures 2-4 In some embodiments, multiple adjacent marker regions and reaction regions are regularly arranged or distributed within the surface area or the corresponding imaging field of view to form a surface pattern. The reaction region, as referred to herein, is the main location or area for performing sequencing-related biochemical reactions and acquiring corresponding signals. The marker region is primarily for marking or tracing purposes, such as facilitating positioning or alignment. In some embodiments, based on the linear shape or pattern characteristics contained within the marker region, it is also referred to as a marker line or trackline. Generally, the area of the reaction region is much larger than that of the marker region; for example, the size of the reaction region may occupy 90% or more of the surface area, while the size of the marker region may occupy 10% or less of the surface area. Moreover, the marker region and the reaction region exhibit significantly different signal characteristics during image acquisition and signal detection, such as significantly different pattern features, signal variation patterns, and / or signal intensity.
[0038] Additionally, please refer to Figure 2 or Figure 3 Adjacent reaction zones can be connected or separated, as can adjacent marker zones. In some examples, the reaction zones on the surface are separated, and there are connected marker zones between adjacent reaction zones. In other examples, the marker zones are separated, and there are connected reaction zones between adjacent marker zones. In still other examples, some reaction zones are separated, while other reaction zones are not separated or connected, with connected marker zones between adjacent separated reaction zones and separated marker zones between adjacent non-separated reaction zones.
[0039] From the perspective of the surface region or the corresponding imaging field of view as a whole, multiple marked areas distributed in the surface region form one or more regular shapes or patterns, such as regular patterns containing one or more straight lines and / or curves, specifically, grid patterns, triangular patterns, circular or wavy line patterns containing multiple intersecting straight lines or curves. These regular shapes or patterns have one or more orientations. All or most of the marked areas in the surface region or the corresponding imaging field of view are isomorphic, that is, their shape, size, and orientation are consistent. Thus, the same sites or positions in these various marked areas will also form corresponding regular patterns or curves on this surface region. The regular distribution of these multiple marked areas on the surface region or the regular distribution of the same positions of multiple marked areas on the surface region is one of the features or characteristics on which the image registration method, especially the coarse registration step, is based. This will be further illustrated in the following related embodiments in conjunction with the implementation or execution process of (S10).
[0040] Each marking region has a plurality of discrete reactive sites regularly distributed in two or more non-parallel directions, and the intersection of these non-parallel directions is located within the marking region. In other words, each marking region has at least two non-parallel extending directions, and the intersection position of these directions is located within the marking region. In some examples, each marking region appears to include a plurality of intersecting straight or curved regions (extension directions), please refer to Figure 2 and Figure 4 , for example, appears as a "V" shape, an "×" shape, a "十" shape, a "T" shape, a "↓" shape, a "木" shape, a "米" shape, etc.
[0041] Combined with the foregoing description of the reactive sites, it can be understood that at least a part of the reactive sites in the marking region appear as bright spots in the image. The regularly distributed reactive sites or corresponding bright spots in each direction of the marking region form a regular shape or pattern. The regular distribution of these reactive sites or corresponding bright spots in the intersecting directions of the marking region is one of the features or characteristics on which this image registration method, especially the fine registration step, is based, and this will be further illustrated by examples in the relevant embodiments below in combination with the implementation or process of (S20).
[0042] In some examples, the reaction region also has a plurality of discrete reactive sites regularly distributed. Similarly, it can be understood that at least a part of the reactive sites in the reaction region also appear as bright spots in the image. The reactive sites have a specific structure and / or specific surface properties to enable the nucleic acid molecule to be detected to be connected to the corresponding site position. In some examples, the reactive sites in the reaction region and the marking region independently present as concave structures, such as pores or wells with micrometer or nanometer scale sizes. Moreover, the distribution of the reactive sites in the reaction region is denser than that of the reactive sites in the marking region. And / or, the size of a single reactive site in the reaction region is smaller than that of a single reactive site in the marking region, but the number is several orders of magnitude more than that in the marking region, and the spacing between adjacent reaction sites is smaller than that in the marking region.
[0043] Based on the aforementioned description of the structure and shape of the reaction area and / or marker area, from another perspective—such as considering the detected signals reflected in the image, image features, or relative signal intensity—the difference between the marker area and the reaction area can be described as follows: each marker area, with its regularly distributed reactive sites, exhibits specific intensity variation characteristics on the detection signal curve compared to adjacent reaction areas. In other words, there are significant differences between the signals acquired from the reaction area and the marker area. For example, the signal intensity may rise or fall sharply when transitioning from a reaction area to a marker area or vice versa. Therefore, the curve formed by continuously acquired signals from a unit surface area along the regularly distributed direction of the marker area or any other direction, transitioning from one reaction area to a marker area and then to its adjacent reaction area, or in other words, the boundary or transition region between reaction area 1, marker area 1, and reaction area 2, generally exhibits specific intensity variation characteristics on the signal curve. These specific intensity variation characteristics are the presentation of the signal difference between the reaction area and the marker area on the detection signal curve, and their specific form or shape is related to the signal characteristics presented by the reaction area and the marker area in the image.
[0044] The so-called reference template is a standard reference or standard system for image registration, also known as a reference coordinate system, reference, reference datum, or reference template. The image to be registered is mapped to this reference through geometric transformation to achieve consistent spatial positioning and feature correspondence. A suitable reference template generally needs to meet requirements such as stability, unique correspondence, complete coverage, geometric consistency, and operability. Specifically, for example, it should not change with the acquisition conditions or the state of the sample under test; it should contain sufficient features to ensure that the reference point / registration point / feature point in the image to be registered uniquely corresponds to the corresponding point in the reference; it should cover the complete surface area reflected by the image to be registered to avoid reference loss during alignment; it should be able to establish a clear mapping relationship with the actual acquisition system; and it should contain identifiable feature points or geometric structures suitable for algorithm detection and matching, etc.
[0045] The reference template can be designed, selected, or constructed independently. Relevant published literature discloses methods for constructing or obtaining reference templates. In the relevant embodiments of this application, the surface region reflected by the image to be registered is a patterned surface. A layout template, such as an ideal geometric model like a CAD coordinate system, exists during the design stage or before fabrication of this patterned surface, serving as the basis for surface processing or fabrication. This layout template is a preferred reference template. In some embodiments, this layout template or reference template is also referred to as a priori reference or a priori template.
[0046] The terms coarse and fine registration refer to relative registration processes, and their distinction depends primarily on the accuracy target, the feature scale used, and the granularity of the algorithm. Coarse registration focuses on global, large-scale spatial alignment, aiming to quickly eliminate macroscopic geometric differences between images, such as overall translation, rotation, and scaling. It typically utilizes sparse, highly saliency features (such as the marker region or reference points in this application) to obtain a preliminary, pixel-level accuracy transformation model with low computational cost. Fine registration, on the other hand, focuses on local, small-scale fine adjustments, aiming to correct local deformations and subtle deviations remaining after coarse registration. It typically utilizes denser, more widely distributed features (such as the numerous bright spots within the marker region in this application) to pursue sub-pixel-level alignment accuracy through more complex algorithms. Together, they constitute a progressive registration process from global to local, from coarse to fine. In some embodiments, coarse registration is based on pixel-level operations, and the relevant registration parameters are generally integer pixel coordinates or discrete values on a larger scale, thereby quickly achieving overall alignment between the image and the reference template. Coarse registration can be considered to correspond to registration at the integer coordinate level. Fine registration, on the other hand, is further optimized based on the results of coarse registration. It achieves sub-pixel accuracy through interpolation, fitting, and optimization algorithms. Fine registration can be considered to correspond to registration at the non-integer (sub-pixel) coordinate level.
[0047] Taking a typical SBS sequencing system as an example, it includes multiple cycles of controlled base extension, imaging, and excision using reversible terminators, DNA polymerase, and related solution systems. Understandably, the nucleic acid molecules to be tested, attached to reactive sites on a patterned surface region, undergo polymerization and reflect or emit light (e.g., fluorescence), often appearing as bright spots with higher intensity than the background signal at the corresponding location in the image acquired in that cycle. Therefore, based on the information contained in these image sets corresponding to specific chemical characteristics (e.g., the nucleic acid molecules to be tested undergoing polymerization), it is possible to determine whether the nucleic acid molecules to be tested at a specified location have undergone polymerization. Combining this with a pre-defined correspondence between fluorescence signals and nucleotide types, the type of nucleotide that has biochemically linked to the nucleic acid molecule to be tested can be detected, thereby determining at least a portion of the sequence of the nucleic acid molecule to be tested, obtaining what are called reads. These reads, understandably, are sequences inferred based on the generation, acquisition, recognition, and analysis of relevant signals.
[0048] For example, a modified nucleotide can be equipped with or bind to a fluorescent label, as well as a reversible inhibitor group that prevents other nucleotides from polymerizing and attaching to the next position of the nucleic acid molecule to be tested (this modified nucleotide is also called a reversible terminator). After each polymerization or base extension reaction, the fluorescent label is excited to emit light, and these luminescence signals are acquired to obtain an image of the nucleic acid molecule to be tested in which a controlled base extension reaction has occurred on the surface region. Then, the inhibitor group and fluorescent label are removed to perform the next polymerization reaction and signal acquisition imaging. This process of polymerization-image-removal is repeated multiple times to obtain an image set of information related to the nucleotides attached to the nucleic acid molecule to be tested in each base extension reaction.
[0049] The repeated extension reactions and corresponding signal detections in the SBS in the above example are also called "sequencing rounds". A sequencing round can be defined as the determination of the base or nucleotide type at a position on any nucleic acid template.
[0050] See also Figure 1 In some embodiments, step (S10) performs coarse registration of the image to be registered based on the reference template, including: determining the positions of multiple reference points in the image based on the reference positions of corresponding reference points in the reference template, establishing a fitting relationship based on the positions of the multiple reference points, and updating the positions of the reference points in the image based on the fitting relationship.
[0051] Based on the foregoing explanations, a reference template is a standard reference or system for image registration, also known as a reference coordinate system, datum, reference reference, or reference template. The reference template includes a reference coordinate system and reflects the same surface area as the image. For example, the reference coordinate system can be a conventional Cartesian or polar coordinate system in mathematics, or it can be a pixel coordinate system, pixel coordinate system, or design layout coordinate system used in engineering practice. In some examples, a layout coordinate system is used as the reference template.
[0052] The surface region includes multiple adjacent marked areas and reactive areas. Each marked area has multiple separate reactive sites regularly distributed in two or more non-parallel directions, and the intersection of these directions is located within the marked area. Each marked area exhibits specific intensity variation characteristics on the detection signal curves of its regularly distributed reactive sites and adjacent reactive areas. A reference point is located at the intersection of the directions of the marked areas and is positioned using the intensity variation characteristics. At least a portion of the reactive sites in the marked areas appear as bright spots in the image, and the number of bright spots is greater than the number of reference points. For example, the surface region includes several marked areas, such as 8, 9, 10, or 11, regularly distributed on the surface region, forming a pattern as shown in the image. Figures 2-5The grid pattern shown consists of vertical or non-vertical, connected or disconnected straight lines or curves. A marker region may contain only one intersection, and each intersection contains at most one reference point (also called a reference point or registration point). Choosing this characteristic or feature as the reference point facilitates establishing a unique correspondence with the reference template, enabling rapid coarse registration. Each non-parallel direction of the marker region may contain, for example, 10, 15, 20, or 30 bright spots in a regular distribution. Each direction of the marker region and its adjacent reaction region exhibit specific intensity variation characteristics on the detection signal curve. These specific intensity variation characteristics represent the signal difference between the reaction region and the marker region on the detection signal curve. Their specific shape or form is related to the signal characteristics of the reaction region and the marker region on the image, and in some examples, they may appear as a W-shape or W-like line, or a broken line containing two or three W-shapes.
[0053] More specifically, multiple marked areas form different regular distributions on the surface area, for example... Figure 2 a and b display cross-shaped marked areas, and multiple marked areas are regularly distributed along the horizontal and vertical lines; moreover, the marked areas are connected to each other. Figure 2 a, or separate as Figure 2 b. Figure 2 In examples a or b, multiple marked areas are regularly distributed along two perpendicular straight lines on the surface area, forming a grid or grid-like pattern. In related embodiments, the straight or curved directions in which the marked areas are regularly distributed are also collectively referred to as baselines, for example... Figure 2 It includes horizontal and vertical lines. Figure 2 The multiple marked regions in 'a' are connected. Figure 2 In b, the multiple marked regions are separate. Figure 2 The marked area in a or b is located at the intersection of horizontal and vertical lines, displayed as a "+" shape, while the reaction area is located on the surface area of the unmarked area, for example... Figure 2 As shown in figure a, the surface areas are divided into separate regular square grids within a connected marked area. In this case, the reference point in the marked area is located at the intersection or convergence of the horizontal and vertical lines.
[0054] Figure 3 a, b, and c illustrate the regular distribution of multiple marked areas on a surface in other embodiments of this application. Examples include non-vertical grid distribution, hexagonal distribution, and circular distribution. Specifically, Figure 3 The bold black dots in the text represent marked areas. Multiple marked areas form multiple sets of parallel lines distributed at a 60-degree angle on the surface. These three sets of lines interweave to form a honeycomb network pattern. Figure 3 In b or c, multiple marked regions are regularly distributed to form connected or disconnected hexagons, and the surface areas of the unmarked regions are all reaction zones. Figure 3 c, the dashed line indicates that the marked area is not connected), or the interior of each hexagonal cell is a reaction zone (e.g. Figure 3 The gray-filled area in b (this only illustrates a separate reaction zone). The reference point is located at the intersection of multiple sets of baselines.
[0055] Figure 4 The diagram illustrates a regular distribution pattern of marked areas on a surface, as shown in some embodiments. This surface pattern comprises a set of curves arranged periodically according to a specific curvilinear pattern (such as a sine wave, parabola, etc.). Marked areas are located at the intersections of two or more regular curves, as indicated by the rectangles, and may be connected or disconnected. The surface areas without marked areas are reactive zones, or reactive zones are located within the regions defined by these curves, with the reference point located at the intersection of these regular curves.
[0056] By using a baseline template and combining prior data, such as the degree of deformation or shift caused to the surface region or labeled region by surface preparation processes, sequencing-related biochemical reactions, and signal acquisition, the range or probable range of the baseline point of each labeled region in the image is determined. This range includes the intersection of all directions of the labeled region and a portion of each reaction region adjacent to that intersection in all directions. Figure 5 The box marked "1" or Figure 8 The area of the image is indicated by the box. For details, please refer to... Figure 5 The marked area is cross-shaped, with a row of bright spots distributed in each of the two intersecting directions. The box marked "1" in the image represents the range of the determined reference point. Box 2 within box 1 represents a marked area within that range. The signal curve corresponding to the range of box 1 is shown in the upper left corner (also marked "1"), while the signal curve segment within box 2 is also marked "2" on the curve graph. According to this curve, from one side to the other (e.g., from the left side of the reaction area to the right side), the signal intensity per unit length decreases from one level to another, and then rises back to the original level. It can be understood that the decrease to another level indicates entry into the marked area, and the peaks (with relatively high signal intensity) in this decreased level reflect the unit length areas containing bright spots within the marked area. Based on the specific intensity variation characteristics of these curves, for example… Figure 5 The W-shaped curve after the steep drop in the curve can be used to locate the reference point.
[0057] In the coarse registration step (S10), the positions of multiple reference points in the image need to be determined first. For example, the reference positions of each corresponding reference point can be read from the reference template, such as... Figure 2The coordinates of the intersection of the horizontal and vertical lines in the marked area of the reference template are used, combined with prior knowledge such as the estimated surface position shift or deformation caused by surface preparation processes, biochemical reactions, and signal detection, to determine the location or region of the reference point in the image to be registered. Then, by identifying specific intensity change characteristics of the detection signal curve in that region, the initial position of the reference point in the image is determined. For example, specific intensity change characteristics exhibited on the detection signal curve appear as follows... Figure 6 The “W” or W-shaped curve shown represents the detection signal curve, where the horizontal axis represents distance and the vertical axis represents signal intensity (e.g., pixel value or grayscale value in a grayscale image). This “W” shape arises from the adjacency of the marker region and the reaction region, where the signal intensity of a unit area within the marker region is significantly weaker than that of a reaction region of the same size. For example, if the number of reactive sites emitting light in the marker region is relatively small and sparsely distributed—for instance, if only one row or column of reactive sites (at least partially appearing as bright spots) is set in the middle of one direction of the marker region—during imaging signal acquisition, the signal intensity of a unit area within the detected marker region is significantly weaker than that of the reaction region. Thus, the signal curves of adjacent or contiguous areas of the reaction and marker regions will exhibit a sharp drop (corresponding to the transition from the reaction region to the marker region) – a rise (corresponding to the unit areas within the marker region containing bright spots) – a sharp drop (corresponding to the unit areas within the marker region without bright spots) – a sharp rise (corresponding to the transition from the marker region to the adjacent reaction region), i.e., a W or W-shaped curve. Understandably, if two rows or two columns of reactive sites are set in each direction of the marked area, or if the reactive sites are distributed in a zigzag pattern, the peaks of the W shape are relatively flat, or there are 1.5 or more W-shaped zigzags with two peaks and three troughs, or it presents a distinct concave shape.
[0058] The detection signal curve of the surface range containing specific intensity variation features, such as a "W" or W-shaped curve, can be plotted as follows, and the reference point position can be located based on these specific intensity variation features: The range is, for example, a rectangular area. At the start of the scan, the probe is positioned on one side of the range, for example, the leftmost side of the left reaction zone within the range. When acquiring the imaging signal, the signal intensity of the unit area detected in the marked area is slightly lower than that in the reaction zone. Thus, the signal curves of adjacent or contiguous areas of the reaction zone and the marked area will form a sharp drop (corresponding to the transition from the reaction zone to the marked area) - a rise (corresponding to the marked area containing bright spots) - a sharp drop (corresponding to the marked area without bright spots) - a sharp rise (corresponding to the transition from the marked area to the adjacent reaction zone), i.e., a W or W-shaped curve, such as... Figure 6 As shown, this allows for the improvement of image coarse registration accuracy based on the intensity variation characteristics exhibited on the detection signal curve. Specifically, it identifies the "W"-shaped troughs of the curves in both the horizontal and vertical directions, and extracts the extreme points of these troughs, such as the minimum value of the horizontal trough. Figure 6 The "min" value is marked in the middle; it also retrieves the second smallest value, such as... Figure 6 The "second min" marker indicates the midpoint of the extreme point calculation. Figure 6 The "mid pos" shown indicates the position of the reference point in the image. For example, the minimum and second minimum values of the integrals of 128 pixels in the horizontal and vertical directions can be calculated separately, and their corresponding coordinate information can be recorded. Then, the midpoint between the coordinates of the minimum and second minimum values can be used to determine the position of the reference point.
[0059] After obtaining the positions of multiple reference points in the image, a fitting relationship is established based on the regular distribution of the marked areas or reference points. Then, the positions of the reference points in the image are updated based on this fitting relationship, thus completing the coarse registration correction process and obtaining the coarse registration result. For example, multiple marked areas or reference points within them form a regular distribution on the surface, and this regular distribution has multiple directions, such as... Figures 2-4 As shown, after obtaining the initial positions of all reference points in the image based on specific intensity features, the reference points are fitted according to their positional components in each direction of the regular distribution (e.g., the regular distribution is represented by mutually perpendicular grids, and the coordinate components of the reference points in the two mutually perpendicular directions of the regular distribution can be represented as x-coordinate components and y-coordinate components), thus obtaining the fitting relationship in each direction. For example, if multiple marked areas or reference points within them form a regular curve on the surface, such as a circle, after obtaining the positions of all reference points in the image based on specific intensity features, a circular fitting relationship is established based on the reference point positions. If the reference points are linearly distributed in the corresponding direction, a linear fitting relationship can be obtained; if the reference points are curvedly distributed in the corresponding direction, a curved fitting relationship can be obtained.
[0060] After obtaining the fitting relationship established based on the position of the reference point, the distance (such as vertical distance) from each reference point to the corresponding fitting relationship is calculated. For reference points that deviate significantly from the fitting relationship (such as distances exceeding 5 pixels), the reference point position is corrected based on the fitting relationship to obtain the updated reference point position, thus completing the coarse registration process. In this way, the reference point position in the image is corrected by the fitting relationship established through global features, which can quickly align the image to be registered with the reference template.
[0061] After completing the coarse registration of the image, performing step S20, fine registration of the image, can improve the image registration accuracy, including fine registration of the image based on the coarse registration result and the bright spots in the marked area.
[0062] In some embodiments, the distribution of bright spots in each direction of the marked area is first determined, then the fitting relationship of each direction of the marked area is determined based on the distribution of bright spots, and then the reference point position of the image is further updated based on the fitting relationship; and the coordinates of other areas or positions of the image are determined based on the updated reference point position, so as to achieve the registration of the image with the reference template.
[0063] Specifically, based on the updated reference point positions after coarse registration, and combining the bright spot distribution in the marked area of the reference template with the prior positional relationship of the reference points (e.g., the reference point is located at the intersection of the regular distribution of bright spots in each direction of the marked area in the reference template), the bright spots in each direction of the marked area are found in the image and their positions are recorded. The bright spot distribution in each direction of the marked area is fitted to obtain the fitting relationship for each direction (e.g., if the marked area includes two mutually perpendicular directions, the fitted linear relationship for bright spots in one direction, such as the horizontal or x-direction, is y=0.02x+102.2). The intersection points of the fitted relationships for bright spots in different directions are calculated (e.g., the intersection points of the fitted lines for bright spot distributions in the horizontal and vertical directions). These intersection points are the updated reference point positions, achieving fine registration of the image. Based on the updated reference point positions and the prior positional relationships of the reference template, the coordinates of the reaction area or other locations in the image can be quickly determined, achieving registration between the image and the reference template.
[0064] In this embodiment, the size of the labeled region does not exceed 10% of the surface area, which facilitates the acquisition of high-throughput sequencing data. The labeled region size refers to the proportion of the total area occupied by all labeled regions in the surface area, including the reactive site region of the labeled region itself and its directional extension regions (e.g., the elongated distribution region of the horizontally oriented labeled regions), but excluding the reactive background region.
[0065] In this embodiment, at least a portion of the reaction regions are separated, with a marker region between adjacent reaction regions. This allows for independent positioning of the reaction regions through the marker region, avoiding signal cross-interference between them. Separation of reaction regions means that there is a physical or signalal gap between them, and adjacent reaction regions are not directly connected. The marker region is located between adjacent reaction regions, forming an alternating distribution structure of "reaction region-marker region-reaction region". For example, the reaction region is a circular area with a diameter of d1, and the center-to-center distance between adjacent reaction regions is d2; the marker region is a long strip-shaped area located on the line connecting the centers of adjacent reaction regions, i.e., adjacent reaction regions are separated by the marker region. After positioning using the reference point of the marker region, the coordinates of the reaction regions on both sides can be directly determined, avoiding signal superposition (such as fluorescence signal crosstalk) caused by adjacent reaction regions, ensuring that the signal of each reaction region can be extracted independently, and improving the accuracy of subsequent base identification.
[0066] In another embodiment of this application, see [link to application]. Figure 2 b or Figure 3 c, Figure 2b or Figure 3 In configuration c, at least some of the marked regions are separated, with reaction zones existing between adjacent marked regions. This approach is suitable for scenarios with large reaction zones, enabling partitioned localization of the reaction zones through the separated marked regions. Separated marked regions mean that there are gaps between marked regions; adjacent marked regions are not directly connected, and the reaction zones are located between adjacent marked regions. This can create a "marked region-reaction zone-marked region" distribution structure, with each reaction zone corresponding to multiple marked regions. Separated marked regions can locate large reaction zones, avoiding local registration errors caused by excessively large reaction zones (such as deviations between the reaction zone's edge and center). Simultaneously, the collaborative localization of multiple marked regions improves the reliability of the reaction zone coordinates; even if one marked region is misaligned, it can be corrected by the other three marked regions, significantly improving registration robustness.
[0067] In the embodiments of this application, the reaction zone is also provided with a plurality of regularly distributed, separate reactive sites, and at least a portion of the reactive sites in the reaction zone appear as bright spots in the image.
[0068] In related embodiments, the reactive sites in the reaction region and the labeling region are each independently presented as depressions, and both are regularly distributed. These depressions are formed on the surface of the sequencing chip using photolithography and etching processes, including micropores (e.g., 1-10 μm in diameter) or nanopores (10-100 nm in diameter). The nucleic acid molecules to be tested are attached to the surface of these depressions, facilitating identification and resolution under microscopic optical imaging. In this embodiment, the depression arrays in the reaction region and the labeling region each form different pattern features or signal differences, which is beneficial for image registration.
[0069] In the embodiments of this application, the distribution of reactive sites in the reaction region is denser than that in the marker region. In some examples, the size of a single reactive site in the reaction region is smaller than that in the marker region, but the density of reactive sites in the reaction region is higher. This method of distinguishing the signal characteristics of the marker and reaction regions by density and / or size allows for rapid identification of the two regions, avoiding confusion during image registration. Simultaneously, the high-density and / or small-sized site design in the reaction region enhances its signal intensity per unit area, facilitating subsequent base identification, while the low-density and / or large-sized site design in the marker region ensures clear reference points at directional intersections, preventing excessive obscuring of intersection features and achieving synergistic optimization of function and localization.
[0070] The coarse registration process in the embodiments of this application will be described in detail below. Figure 7The following describes the coarse registration process in some embodiments of this application. Specifically, (S10) performing coarse registration of an image based on a reference template includes the following steps: (a) calculating the reference position of the corresponding reference point in the reference template; (b) determining the detection signal curve of the predetermined region where the reference position is located in the image; and (c) identifying the intensity change characteristics of the detection signal curve and determining the position of the reference point based on the position of the intensity change characteristics.
[0071] From the reference coordinate system of the reference template, the coordinates of the intersection of the directions of each marked area are read as the reference position of the reference point. For example, if the coordinates of the intersection of the horizontal and vertical directions of a marked area in the template are (200, 200), then the reference position of the reference point is (200, 200). Then, the detection signal curve of the predetermined region in the image where the reference position is located is determined. In some embodiments, the predetermined region is also called the range or area where the reference point is located, and includes the intersection of the directions of the marked areas and a portion of each reaction area adjacent to each direction, such as... Figure 5 The box area marked 1 in the middle or Figure 8 The bounding box region in the image. For example, in a registered image, the region centered on the reference position, encompassing the area where the marked area intersects with the adjacent reaction area (e.g., 128×128 pixels), is used for focused signal detection and to reduce background noise interference. Signal intensity is collected per unit length of the predetermined region along a direction (e.g., horizontally) from one side to the other. A signal detection curve for this predetermined region is formed using "distance-signal intensity" as the coordinate axis to reflect the signal variation characteristics of the predetermined region.
[0072] The coarse registration process is divided into three steps, and the operation objects and processing methods of each step are clearly defined. This facilitates engineering development. For example, different processing units can be used to execute the corresponding processing procedures. By focusing and detecting the predetermined area, the interference of background noise is reduced, making the intensity change features easier to identify, thereby improving the accuracy of the reference point positioning.
[0073] In this embodiment, multiple marking areas are regularly distributed along two or more non-parallel straight lines on the surface region, corresponding to... Figure 1 The coarse registration process shown may also include the following steps:
[0074] (d1) Establish the fitting relationship in the corresponding direction by using the position components of the reference point in each direction.
[0075] (e1) The reference points are classified into two categories according to preset conditions, and the fitting relationship in the corresponding direction is updated based on the positional components of the reference points in the specified categories after binary classification. The preset conditions are related to the distance from the reference points to the fitting relationship.
[0076] (f1) Correct the positional components of the reference points in another category based on the updated fit.
[0077] Figure 8 In the embodiments shown in this application, the reference template includes a horizontal baseline and a vertical baseline. Figure 8 The bold black areas represent the baselines in the corresponding directions. These marked areas are regularly distributed along mutually perpendicular horizontal (X-axis) and vertical (Y-axis) directions on the surface, forming a regular grid pattern. This pattern appears as horizontal and vertical baselines on the image. For any reference point P, its coordinates are (Xp, Yp). Its horizontal position component is its Y-coordinate Yp (because the Y-coordinates of all points on the horizontal line should be equal); its vertical position component is its X-coordinate Xp (because the X-coordinates of all points on the vertical line should be equal).
[0078] like Figure 8 As shown, the horizontal baseline consists of a series of reference points with the same Y-coordinate, and the vertical baseline consists of a series of reference points with the same X-coordinate. The process of establishing the fitting relationship includes: fitting the horizontal direction (Y-coordinate component): extracting the Y-coordinate values of all reference points, and performing line fitting using methods such as least squares to obtain a horizontal fitted line representing the overall trend of all horizontal template lines. Fitting the vertical direction (X-coordinate component): extracting the X-coordinate values of all reference points, and performing line fitting to obtain a vertical fitted line representing the overall trend of all vertical template lines.
[0079] Then, the vertical distance from each reference point P to the fitted line is calculated, including the distance from the Y-coordinate of reference point P to the horizontal fitted line and the distance from its X-coordinate to the vertical fitted line. A preset distance threshold D1 is set. Reference points whose maximum (or average) distance is less than or equal to D1 are classified into a specified category (e.g., the inlier set); reference points whose distance is greater than D1 are classified into another category (the outlier set, i.e., outliers). Step (d1) is repeated using only the reference points in the specified category (e.g., the inlier set), that is, new, more accurate vertical and horizontal fitted lines are fitted again using the X and Y coordinates of these inliers. Then, the positions of the reference points classified into another category (e.g., the outlier set) in (e1) need to be corrected. For an outlier reference point, it is projected onto the updated fitted line. That is, its X-coordinate is corrected to a value on the vertical fitted line, and its Y-coordinate is corrected to a value on the horizontal fitted line, resulting in corrected new coordinates.
[0080] In this embodiment, for scenarios where the marked area is linearly distributed, the offset of the reference point caused by local deformation of the chip can be eliminated by fitting and correcting the position components, so that the reference points in the same direction are more in line with the linear distribution; binary classification screening can remove abnormal reference points (such as deviation points caused by noise), avoid them from affecting the fitting relationship, and improve the reliability of coarse registration.
[0081] Furthermore, in the embodiments of this application, the preset conditions include multiple distance thresholds that converge from large to small. For example, the multiple distance thresholds include the largest distance threshold D1 and the second largest distance threshold D2. Correspondingly, in the above embodiments, (e1) performing binary classification on the reference points according to the preset conditions and updating the fitting relationship of the corresponding direction based on the position components of the reference points in the specified category after binary classification may include the following steps: (e11) performing binary classification on the reference points according to the distance threshold D1, including classifying the reference points whose distance fitting relationship is less than or equal to the distance threshold D1 into the specified category, and classifying the reference points whose distance fitting relationship is greater than the distance threshold D1 into another category that is not a specified category.
[0082] (e12) Update the fitting relationship in each direction using the positional components of the reference points in the specified category after (e11) in each direction.
[0083] (e13) Perform binary classification on the reference points in the specified category after (e11) according to the distance threshold D2, including classifying the reference points whose fitting relationship after (e12) is less than or equal to the distance threshold D2 into the specified category, and classifying the reference points whose fitting relationship after (e12) is greater than the distance threshold D2 into another category.
[0084] (e14) Update the fitting relationship in each direction using the positional components of the reference points in the specified category after (e13) in each direction.
[0085] In this embodiment, the predetermined condition is a series of gradually converging distance thresholds (e.g., D1, D2). Through multiple rounds of binary classification and fitting updates, the accuracy of the baseline point fitting relationship is further improved, making it suitable for sequencing scenarios with high registration accuracy requirements. The convergent distance threshold refers to the distance threshold used in multiple rounds of screening that gradually decreases (e.g., D1=5 pixels, D2=3 pixels, D3=1 pixel), making the screened baseline points increasingly closer to the ideal distribution, and the fitting relationship gradually converges to the optimal value. For example, firstly, binary classification is performed based on D1=5 pixels. For instance, the vertical distance from 20 horizontal baseline points to the initial fitting relationship (e.g., y=0.05x+187) is calculated. 19 baseline points with a distance ≤5 pixels are assigned to the specified category, and 1 baseline point with a distance = 6 pixels is assigned to a non-specified category. Then, based on the y-coordinates of the 19 specified category baseline points, a refit is obtained, such as y=0.049x+187.1. Next, calculate the vertical distance from the 19 designated category reference points to the updated fitted relation (e.g., y=0.049x+187.1). Assign 17 reference points with a distance ≤ 3 pixels to the new designated category, and assign the two reference points with distances of 3.5 pixels and 4 pixels to the undesignated category. Based on the y-coordinates of the 17 designated category reference points, refit to obtain y=0.048x+187.2 (e.g., the final fitted relation). If higher precision is required, set D3=1 pixel and repeat the above process to filter out 15 reference points with a distance ≤ 1 pixel, such as fitting y=0.0478x+187.22, further improving the fitting accuracy. It should be noted that the distance thresholds and related values and formulas for the fitted relation in this example are for better illustrating the fitting and updating process; the specific values can be determined based on the actual application scenario.
[0086] In this embodiment, the convergence distance threshold is used to gradually approximate the true distribution of the benchmark points through multiple rounds of screening and fitting. The decrease in the threshold in each round ensures the asymptotic nature of the screening and avoids the elimination of a large number of effective benchmark points due to the use of a small threshold at once, thus balancing the fitting accuracy and the number of benchmark points.
[0087] In one embodiment of this application, multiple marked areas are distributed along a regular curve direction in the surface area, and the corresponding coarse registration process may include the following steps: (d2) Establishing a corresponding curve fitting relationship through the position of the reference point.
[0088] (e2) Classify the benchmark points into two categories based on preset conditions, and update the corresponding curve fitting relationship based on the positions of the benchmark points in the specified categories after binary classification. The preset conditions are related to the distance from the benchmark point to the curve fitting relationship.
[0089] (f2) Correct the position of the benchmark point in another category based on the updated curve fitting relationship.
[0090] The distribution of marked areas along a regular curve on the surface can mean that the marked areas follow a circular, elliptical, sine, parabolic, or other regular curve (such as...) on the surface. Figure 4 As shown, the distribution is uniform, and the positions of the reference points conform to a curve equation. For example, when it conforms to a circular distribution, the corresponding circular equation can be: (xa)² + (yb)² = r², where (a,b) is the center and r is the radius. Then, the curve equation obtained by fitting the reference point positions is used to describe the curve distribution characteristics of the marked area.
[0091] For example, suppose the marked area in the image is distributed along a circular direction with 30 reference points, whose coordinates are (150, 200), (155, 195), (160, 190), ... conforming to a circular trajectory. Using a nonlinear fitting algorithm, the equation of the circular curve is obtained based on the coordinates of these 30 reference points: (x-150)² + (y-200)² = 100² (center (150, 200), radius 100 pixels), which is the curve fitting relationship. If the preset condition is set that the distance from the reference point to the circular curve is ≤ 4 pixels, the distance from each reference point to the circular curve is calculated, and reference points of a specified category (distance ≤ 4 pixels) are collected. The updated circular curve equation is then obtained by refitting: for example, (x-150.2)² + (y-199.8)² = 99.9². For reference points of a non-specified category, their projection points on the circle are calculated based on the updated circular curve equation, and these projection points are used as the corrected reference point positions. In this embodiment of the application, for the scenario of the curve distribution of the marked area, the optimization of the non-linear distribution reference points is achieved through curve fitting relationship, which breaks through the limitations of traditional straight line fitting and is applicable to special chip structures such as circles and waves; curve fitting and correction can eliminate the offset of reference points on the curve trajectory (such as local deviation caused by chip deformation), so that the distribution of reference points fits the preset curve better.
[0092] Furthermore, the preset conditions include multiple distance thresholds that converge progressively from large to small, including the largest distance threshold D1 and the second largest distance threshold D2, (e2) including,
[0093] (e21) Classify the reference points into two categories based on the distance threshold D1, including classifying reference points whose distance curve fitting relationship is less than or equal to the distance threshold D1 into a specified category, and classifying reference points whose distance curve fitting relationship is greater than the distance threshold D1 into another category that is not a specified category.
[0094] (e22) Update the curve fitting relationship by the position of the reference point in the specified category after (e21).
[0095] (e23) Divide the reference points in the specified category after (e21) into two categories based on the distance threshold D2, including assigning reference points whose distance to the curve fitting relationship after (e22) is less than or equal to the distance threshold D2 to the specified category, and assigning reference points whose distance to the curve fitting relationship after (e22) is greater than the distance threshold D2 to the other category.
[0096] (e24) Update the curve fitting relationship by the position of the reference point in the specified category after (e23).
[0097] In this embodiment, multiple rounds of screening and fitting updates using the convergence distance threshold provide a high-precision reference for subsequent benchmark point correction. Simultaneously, the multiple rounds of screening avoid the problem of insufficient effective benchmark points caused by using a small threshold all at once, balancing fitting accuracy and the number of benchmark points. It should be noted that the processing of multiple rounds of screening and fitting updates for curve features is similar to the screening and updating process for straight line fitting relationships in the previous embodiment, and will not be detailed here.
[0098] Correspondingly, in the embodiments of this application, the plurality of marked areas are regularly distributed in the surface area along two or more non-parallel directions, wherein at least one direction is a straight line direction and at least another direction is a regular curve direction. For the processing of the straight line direction, please refer to the description of the relevant steps d1 to f1 in the above embodiments; for the processing of the regular curve direction, please refer to the description of the relevant steps d2 to f2 in the above embodiments, which will not be detailed here.
[0099] The process of fine image registration is explained below.
[0100] See Figure 9 It illustrates a flowchart of a fine registration process provided in an embodiment of this application, which is the image fine registration process. Figure 1 The process shown in (S20) for fine registration of the image based on the coarse registration result and bright spots includes the following steps: S22, determining the distribution of bright spots in each direction of the marked area based on the position of the reference points and the reference template after coarse registration of the image. This process includes identifying bright spots in multiple regions in each direction.
[0101] S24. Determine the fitting relationship for each direction based on the distribution of bright spots in each direction, and select bright spots to optimize the fitting relationship.
[0102] S26. Based on the fitted relationship optimized by the bright spots, further update the reference point position of the image, and determine the coordinates of other regions or positions of the image according to the further updated reference point position, so as to achieve the registration of the image with the reference template.
[0103] The distribution of bright spots in each direction of the marked area refers to the set of positions of effective bright spots in each non-parallel direction (such as horizontal, vertical, sinusoidal, or circular curve directions). For example, if 10 bright spots are distributed along a straight line in the horizontal direction and 8 bright spots are distributed along a straight line in the vertical direction, their distribution characteristics reflect the actual positional deviation of the marked area. By filtering out abnormal bright spots (such as noise points or false bright spots with weak signals), the bright spot fitting relationship (such as a straight line or curve equation) is updated to ensure that the fitting relationship can accurately describe the true distribution of bright spots. Based on the updated reference point positions, the coordinates of the reaction area and other marked areas in the image are determined, so that the coordinate system of the image is completely aligned with that of the reference template, providing an accurate coordinate reference for subsequent base identification.
[0104] For example, based on the reference point position and reference template after coarse registration, the direction of the marked area (such as horizontal and vertical) and the estimated bright spot area in each direction are determined. For instance, if the length L of the marked area in the horizontal direction is approximately 50 pixels, to limit the search to within the marked area, the reference point can be set as the center, and the bright spot area can be delineated within ±(L / 2 - d) pixels along the direction axis (where d is a safety boundary, such as 5 pixels, to avoid touching the edge of the marked area). Accordingly, three search areas of size (2d)×(2d) (such as 10×10 pixels) can be delineated in the horizontal direction to the left, center, and right of the reference point, respectively, thus accurately covering the expected bright spot positions in that direction within the marked area. These correspond to the estimated positions of the three horizontal bright spots, "left, center, and right". In the vertical direction, with the reference point as the center, three bright spot areas (each area is 10×10 pixels) are delineated within ±20 pixels along the y-axis, corresponding to the estimated positions of the three vertical bright spots, "top, center, and bottom". Valid bright spots are identified within each bright spot area. For example, bright spots are filtered by ensuring the center pixel signal value is greater than that of its eight neighboring pixels, resulting in a set of bright spot distributions (locations) in both horizontal and vertical directions, providing a data foundation for subsequent fitting. Then, based on the initial fitting relationship obtained from the filtered bright spots, an updated fitting relationship is obtained. The reference point positions of the image are then updated based on this updated fitting relationship, and the coordinates of other areas or locations in the image are determined based on the further updated reference point positions, thus achieving registration between the image and the reference template. In this embodiment, from bright spot identification to fitting optimization, and then to reference point updating and global registration, this approach is applicable to fine registration scenarios with all types of marked area distributions (straight lines, curves, and composite distributions), further improving image registration accuracy.
[0105] In the embodiments of this application, the process of determining the distribution of bright spots in each direction of the marked area based on the position of the coarsely calibrated reference point and the reference template includes: determining the pixel range k1×k2 in multiple regions of each direction where the bright spots may be located by using the position component of the coarsely registered reference point in the direction and the relative position relationship between the reference zone and the bright spots contained in the reference coordinate system; where k1 and k2 are each independently odd numbers greater than 1; and identifying the bright spots in the pixel range k1×k2 where the bright spots may be located based on the difference in pixel size between the center pixel and the surrounding pixels.
[0106] After coarse calibration, the position component of the reference point in a certain direction refers to the coordinate projection of the reference point onto the marked area in a certain direction. For example, the position component in the horizontal direction of the marked area is the y-coordinate of the reference point, and the position component in the vertical direction is the x-coordinate of the reference point. This is used to determine the approximate height or horizontal position of the bright spot in that direction. The relative positional relationship between the reference point and the bright spot in the reference frame can characterize the coordinate offset between the bright spot and the reference point in the reference template. For example, if a bright spot in the horizontal direction is located 10 pixels to the right of the reference point, and their y-coordinates are the same, it can be stored in the form of (Δx, Δy), which is the core basis for defining the pixel range of the bright spot. The pixel range of k1×k2 where the bright spot may be located refers to the search area centered on the estimated position of the bright spot. k1 and k2 are odd numbers greater than 1 to ensure that the search range is symmetrical about the estimated position, covering possible offsets of the bright spot (such as coarse registration deviation and signal drift), while avoiding an excessively large range that would reduce search efficiency.
[0107] For example, a bright spot might be located within a 5x5 pixel area. The bright spot is determined based on the difference in pixel size between the center pixel and its surrounding pixels. Specifically, if the grayscale value of the center pixel is greater than that of the surrounding pixels, then the center pixel can be considered a bright spot within that pixel area. This effectively distinguishes bright spots from background noise, improving the accuracy of bright spot identification.
[0108] In the coarse registration process of this application embodiment, the reference points can be binary classified according to preset conditions, and then the fitting relationship of the corresponding direction is updated based on the positional components of the reference points in the specified categories after binary classification. When the preset condition is the first preset condition, in the fine registration process of this application embodiment, determining the fitting relationship of each direction according to the distribution of bright spots in each direction includes: filtering bright spots through a second preset condition, and updating the fitting relationship of the direction based on the bright spots retained after filtering, wherein the second preset condition is related to the distance from the bright spot to the fitting relationship. The second preset condition is the condition used to filter bright spots in the fine image registration process, for example, it can be that the distance from the bright spot to the current fitting relationship is not greater than a preset distance threshold. Then, based on the effective bright spots retained after filtering, a more accurate fitting relationship (such as a straight line or curve equation) is refitted, and this relationship is the basis for subsequent updating of the reference point position.
[0109] Furthermore, the second preset condition includes multiple distance thresholds that converge from large to small, including a maximum distance threshold D21 and a second-largest distance threshold D22. The fine registration process, which determines the fitting relationship for each direction based on the distribution of bright spots and filters bright spots to optimize the fitting relationship, includes the following steps: (S241) Filtering bright spots according to the distance threshold D21, including retaining bright spots with a distance fitting relationship less than or equal to the distance threshold D21, and optionally discarding bright spots with a distance fitting relationship greater than the distance threshold D21. (S243) Updating the fitting relationship using the bright spots retained after (S241).
[0110] (S245) Based on the distance threshold D22, the bright spots retained after (S241) are filtered, including retaining bright spots whose fitting relationship after (S243) is less than or equal to the distance threshold D22, and optionally discarding bright spots whose fitting relationship after (S243) is greater than the distance threshold D22; and
[0111] (S247) Update the fitted relationship using the bright spots retained after (S245).
[0112] The distance thresholds used in multiple rounds of screening are gradually reduced (e.g., D21=2 pixels, D22=1 pixel). After each round of screening, the retained bright spots are closer to the ideal distribution, and the accuracy of the fitting relationship is gradually improved. After each round of screening, the fitting relationship is refitted based on the retained effective bright spots, so that the fitting relationship is optimized as the threshold converges, avoiding the problem of insufficient effective bright spots caused by using a small threshold all at once. For example, suppose 10 bright spots are initially identified in the horizontal direction, and the initial fitting relationship is y=0.005x+196.1, with preset convergence thresholds D21=2 pixels and D22=1 pixel. Then, the vertical distance from these 10 bright spots to the initial fitting relationship is calculated. Based on the calculation results, 9 bright spots with a distance ≤ 2 pixels are retained, and 1 bright spot with a distance > 2 pixels is discarded. Based on the y-coordinates of the 9 retained bright spots, a new fitting relation is obtained by refitting: y = 0.0049x + 196.11. Then, the vertical distance from these 9 retained bright spots to the updated fitting relation is calculated. Based on the results, 8 bright spots with a distance ≤ 1 pixel are retained, and 1 bright spot with a distance > 1 pixel is discarded. Based on the y-coordinates of the 8 retained bright spots, a new fitting relation is obtained by refitting. The distance from these 8 bright spots to the fitting relation is calculated again. If all of them meet the distance threshold D23, the process of updating the fitting relation can be stopped. If there are bright spots that do not meet the threshold, they can be discarded, and the process of updating the fitting relation can continue.
[0113] In this embodiment, the multiple rounds of screening of the convergence threshold can gradually eliminate abnormal bright spots of different degrees, avoiding the mistaken removal of a large number of valid bright spots due to the use of a small threshold at once; the accuracy of the final fitting relationship is significantly improved, providing a highly accurate reference for subsequent benchmark updates, and ensuring that the alignment accuracy between the image and the benchmark template meets the signal extraction requirements of single-molecule sequencing.
[0114] In the image coarse and fine registration processes of this application, to improve registration efficiency, different tasks can be executed in parallel by multiple processing units. This is applicable to the massive image processing of high-throughput sequencing, shortening image registration time and meeting the real-time requirements of the sequencing process. Hardware with multi-core processing units or multi-threading capabilities, such as GPUs or CPUs, can be used to allocate different registration tasks to different processing units for simultaneous execution, reducing the overall processing time.
[0115] For example, the process of determining the positions of multiple reference points in an image during coarse registration is performed by multiple parallel processing units. Specifically, if an image contains 200 reference points, each reference point requires processing a predetermined region of 128×128 pixels; the processing task of the predetermined region for the 200 reference points is distributed to 200 thread blocks (each thread block processes one reference point); each thread block independently reads the predetermined region data, extracts the detection signal curve, identifies intensity change features, and determines the reference point position, without waiting for other thread blocks. If traditional serial processing takes 100ms (0.5ms per reference point), parallel processing only takes 0.6ms (including data read / write latency), improving efficiency by approximately 167 times.
[0116] For example, multiple parallel processing units can be used to perform the update process of each fitting relationship in each direction during the coarse registration process. If the marked areas in the image are distributed along three non-parallel directions (horizontal, vertical, and 45° tilt), the fitting relationship needs to be updated in each direction. The fitting update tasks in the three directions are assigned to three independent thread groups (each group contains 32 threads). Each thread group independently reads the reference point data in the corresponding direction, calculates the distance, selects reference points, and updates the fitting relationship. For example, the horizontal direction thread group processes the y-coordinate fitting of 200 reference points, and the vertical direction thread group processes the x-coordinate fitting, without interfering with each other. If serial processing takes 30ms, parallel processing takes 10ms, improving efficiency by 3 times.
[0117] For example, multiple parallel processing units are used to optimize the fitting relationships in the fine registration process of an image. In fine registration, the marker regions are distributed in two directions (horizontal and vertical). For each direction, a process of selecting bright spots and updating the corresponding fitting relationship based on the selected bright spots needs to be performed. The fitting optimization tasks in the two directions are assigned to two thread blocks. For example, the horizontal thread block handles the selection and fitting of 8 bright spots, and the vertical thread block handles the selection and fitting of 7 bright spots, while simultaneously outputting the updated fitting relationship.
[0118] For example, in the fine registration process of an image executed by multiple parallel processing units, the process of determining the coordinates of other regions or positions in the image based on the further updated reference point positions is performed in step (S26). For instance, if the image contains 1000 reaction zones, their coordinates need to be calculated based on the updated reference point positions; the task of calculating the coordinates of the 1000 reaction zones is distributed to 1000 threads (each thread processes one reaction zone); each thread independently reads the prior positional relationship between the reaction zone and the reference point, and substitutes the reference point coordinates to calculate the coordinates of the reaction zone.
[0119] In this embodiment, tasks that can be processed in parallel can be distributed among multiple parallel processing units, avoiding waste of hardware resources and ensuring load balancing among the processing units. Simultaneously, the parallel processing of each unit does not affect registration accuracy, thus improving registration efficiency.
[0120] This application also provides a method for detecting bases in an image, where the image comes from a detection system that performs sequencing based on patterned surface microscopy. The surface region reflected in the image includes multiple adjacent labeled regions and reaction regions. See also Figure 10 The method includes (S100) determining the signal intensity value of the position of the corresponding chemical feature in the reaction region of the image, the image being an image registered by the image registration method or system in any of the embodiments of the present application described above; and (S200) identifying the type of base introduced at the position based on the signal intensity value.
[0121] The location of the corresponding chemical feature in the reaction region refers to the location of the reactive site in the reaction region where DNA polymerization has occurred or can occur. For example, after alignment with the reference template, the coordinates of each reaction site can be directly determined using interpolation, and then the intensity at that coordinate position can be read. Based on the probability that this intensity comes from a certain base, the type of base incorporated or extended in the current round of reaction can be determined. For example, the base channel that contributes the most to the signal intensity can be identified as the target base, thus obtaining the base identification result.
[0122] Furthermore, in the relevant embodiments, the signal strength value is the corrected signal strength value. Correcting the signal strength value includes correcting it by at least one of the following methods: i) bleeding correction; ii) crosstalk correction; iii) phase correction.
[0123] Various corrections can utilize the GPU's block-based multi-threaded processing mode. For example, the bleeding correction coefficient can be used to perform bleeding correction on the current bright spot signal strength, thereby improving the accuracy of the bright spot signal strength.
[0124] Intensity or intensity value, sometimes referred to as brightness in related embodiments of this application, can be represented by the pixel value / grayscale value, sub-pixel value, or sub-pixel value of the location of the nucleic acid molecule in the image. For example, if a feature of a nucleic acid molecule at a certain location occupies one or more pixels or multiple sub-pixels in the image, the intensity information of the nucleic acid molecule at that location can be represented by the intensity value of any pixel or sub-pixel location it occupies. Alternatively, the center or centroid of the feature can be determined, and the pixel value or sub-pixel value of the center or centroid location can represent the intensity information of the nucleic acid molecule at that location in the image. The intensity information of nucleic acid molecules at one or more locations in the image can be the original signal intensity value in the image, or it can be the processed signal intensity value, such as the intensity value after background removal and / or after at least one correction for color difference, crosstalk, and phase, or it can be the intensity value after standardization or normalization. Crosstalk correction, phase correction (phasing or prephasing correction), etc., can be performed by disclosed methods such as those disclosed in EP3077943B1 and CN113012757B.
[0125] In some specific examples, the inter-channel interference coefficients are calibrated by pre-collecting and analyzing the raw signals of the sequencing images. For instance, the interference coefficient between the A channel and the T channel can be calibrated using the a-coefficient. AT This indicates that, correspondingly, the interference coefficient 'a' between other channels can be obtained. AC a AG a CA a CG a CT a GA a GC a GT a TA a TC a TG Using GPU block processing mode, one GPU processing thread can handle the 4-channel correction of a single bright spot, for example: A 校正后 = A 校正前 - C 校正前 × aCA - G 校正前 × a GA -T 校正前 × a TA The correction process for other bright spots is similar. Therefore, crosstalk correction of the four channels of each bright spot can be performed in parallel, improving the efficiency of bright spot signal strength correction.
[0126] During phase correction, correction can be performed based on the residual ratio coefficient between the current sequencing cycle and the preceding and following sequencing cycles. For example, the residual ratio coefficient between the current sequencing cycle and the preceding and following sequencing cycles can be calculated in real time. lastcyc phasing nextcyc prephasing currcyc phasing currcyc Perform correction, for example, the corrected signal strength is expressed in A currcyc校正后 express:
[0127] A currcyc校正后 =A currcyc校正前 + A currcyc校正前 ×phasing currcyc - A currcyc校正前 ×prephasing lastcyc + A currcyc校正前 ×prephasing currcyc - A currcyc校正前 ×phasing nextcyc .
[0128] Specifically, a histogram of ratio coefficients of size 1024 can be calculated using the GPU to obtain the values of the ratio coefficient 50 (AG channel) and 40 (CT channel) quantiles. Point signal correction is then performed in a multi-threaded manner across blocks.
[0129] This application also provides a method for detecting bases in an image, including base recognition based on a registered image. For example, the image registration method of any of the above embodiments is used to register the image with a reference template, and then base recognition is performed based on the registered image to obtain accurate base recognition results.
[0130] In other embodiments of this application, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the image registration method and / or the method for detecting image base identification as described in any of the above embodiments.
[0131] In another embodiment of this application, a computer device is also provided. The electronic device may include: a memory for storing an application program and data generated by the application program; and a processor for executing the application program to implement the image registration method as described in any of the above embodiments and / or the method for detecting and recognizing bases in an image as described in any of the above embodiments.
[0132] The processor of the computer device includes at least one set of processing units, each set of processing units includes multiple processing units, each processing unit can process in parallel, and the set of processing units is used to perform at least the following steps or processes: (i) determining the positions of multiple reference points in the image in (S10) of the image registration method in any of the above embodiments;
[0133] (ii) When performing the image registration method (S10), update the fitting relationships in each direction;
[0134] (iii) When performing the image registration method (S20), optimize each fitting relationship;
[0135] (iv) In (S26) of (S20) of the image registration method, the coordinates of other regions or locations of the image are determined based on the further updated reference point positions.
[0136] Here, the processing unit set refers to a computing module in the processor composed of multiple independent processing units (such as CPU cores, GPU stream processors, FPGA logic units, etc.). These processing units can independently execute different subtasks at the same time. Multiple processing units process different data or tasks simultaneously, rather than executing them sequentially, thereby reducing overall processing time and improving computational efficiency. For details, please refer to the relevant description of parallel processing in the aforementioned image registration embodiment, which will not be elaborated here.
[0137] It should be noted that the specific implementation of the processor in this embodiment can be referred to the corresponding content above, and will not be described in detail here.
[0138] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0139] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0140] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image registration method, wherein the image is derived from a detection system for sequencing based on patterned surface microscopy, characterized in that, include: (S10) Perform coarse registration on the image based on the reference template, including: The positions of multiple reference points in the image are determined based on the reference positions of corresponding reference points in the reference template; a fitting relationship is established based on the positions of the multiple reference points; and the positions of the reference points in the image are updated based on the fitting relationship; wherein, The reference template includes a reference coordinate system and reflects the same surface area as the image; The surface region includes multiple adjacent marker regions and reaction regions. The marker regions have multiple separate reactive sites regularly distributed in two or more non-parallel directions, and the intersection of the directions is located in the marker regions. The marker regions exhibit specific intensity variation characteristics on the detection signal curves of the adjacent reaction regions in each direction where the regularly distributed reactive sites are located. The reference point is located at the intersection of the directions of the marked area and is located by the intensity change feature; at least a portion of the reactive sites appear as bright spots on the image, and the number of bright spots is greater than the number of reference points; and, (S20) Based on the coarse registration result and the bright spots, perform fine registration on the image, including: The distribution of bright spots in each direction of the marked area is determined, and the fitting relationship of each direction is determined based on the distribution of bright spots. Then, the reference point position of the image is further updated based on the fitting relationship. The coordinates of other regions or positions of the image are determined based on the updated reference point position, so as to achieve the registration of the image with the reference template.
2. The method according to claim 1, characterized in that, The size of the marked area does not exceed 10% of the surface area.
3. The method according to claim 1, characterized in that, At least a portion of the reaction zones are separated, and the marking zone exists between two adjacent reaction zones.
4. The method according to claim 1, characterized in that, At least a portion of the marked regions are separated, and the reaction region exists between two adjacent marked regions.
5. The method according to any one of claims 1-4, characterized in that, The reaction zone has multiple separate reactive sites that are regularly distributed, and at least a portion of the reactive sites in the reaction zone appear as bright spots in the image.
6. The method according to claim 5, characterized in that, The reactive sites in the reactive zone and the marked zone are each independently presented as depressions.
7. The method according to claim 5, characterized in that, The distribution of reactive sites in the reaction zone is denser than that in the marker zone; optionally, the size of a single reactive site in the reaction zone is smaller than that of a single reactive site in the marker zone.
8. The method according to any one of claims 1-7, characterized in that, (S10) includes: (a) Calculate the reference position of the corresponding reference point in the reference template; (b) Determine the detection signal curve of a predetermined region in the image where the reference position is located, wherein the predetermined region includes the intersection of the directions of the marked region and a portion of the reaction region adjacent to the direction, and (c) Identify the intensity change characteristics of the detection signal curve, and determine the position of the reference point based on the position corresponding to the intensity change characteristics.
9. The method according to claim 8, characterized in that, The plurality of said marked areas are regularly distributed in the surface region along two or more non-parallel straight lines, and (S10) further includes: (d1) Establish a fitting relationship in the corresponding direction based on the position components of the reference point in each of the directions; (e1) The reference points are binary classified according to preset conditions, and the fitting relationship in the corresponding direction is updated based on the positional components of the reference points in the specified categories after binary classification, wherein the preset conditions are related to the distance from the reference points to the fitting relationship; and (f1) Correct the positional components of the reference points in another category based on the updated fit.
10. The method according to claim 9, characterized in that, The preset conditions include multiple distance thresholds that converge from large to small. These multiple distance thresholds include the largest distance threshold D1 and the second largest distance threshold D2. (e1) includes... (e11) The reference points are classified into two categories according to the distance threshold D1, including classifying the reference points whose distance from the fitting relationship is less than or equal to the distance threshold D1 into the specified category, and classifying the reference points whose distance from the fitting relationship is greater than the distance threshold D1 into another category that is not a specified category; (e12) Update the fitting relationship in the corresponding direction using the positional components of the reference points in the specified category after (e11) in each of the directions; (e13) Perform binary classification on the reference points in the specified category after (e11) according to the distance threshold D2, including classifying the reference points whose fitting relationship after (e12) is less than or equal to the distance threshold D2 into the specified category, and classifying the reference points whose fitting relationship after (e12) is greater than the distance threshold D2 into another category; and (e14) Update the fitting relationship in the respective directions by the positional components of the reference points in the specified categories after (e13) in each of the directions.
11. The method according to claim 8, characterized in that, The plurality of the marked areas are distributed along a regular curved direction in the surface region, and (S10) further includes: (d2) Establish the corresponding curve fitting relationship based on the position of the reference point; (e2) The reference points are binary classified according to preset conditions, and the corresponding curve fitting relationship is updated based on the position of the reference points in the specified categories after binary classification, wherein the preset conditions are related to the distance from the reference points to the curve fitting relationship; and (f2) Correct the position of the benchmark point in another category based on the updated curve fitting relationship.
12. The method according to claim 11, characterized in that, The preset conditions include multiple distance thresholds that converge from large to small. These multiple distance thresholds include the largest distance threshold D1 and the second largest distance threshold D2. (e2) includes... (e21) The reference points are classified into two categories according to the distance threshold D1, including classifying the reference points whose distance from the curve fitting relationship is less than or equal to the distance threshold D1 into the specified category, and classifying the reference points whose distance from the curve fitting relationship is greater than the distance threshold D1 into another category that is not a specified category; (e22) Update the curve fitting relationship by the position of the reference point in the specified category after (e21); (e23) Divide the reference points in the specified category after (e21) into two categories according to the distance threshold D2, including classifying the reference points whose distance to the curve fitting relationship after (e22) is less than or equal to the distance threshold D2 into the specified category, and classifying the reference points whose distance to the curve fitting relationship after (e22) is greater than the distance threshold D2 into another category; and (e24) Update the curve fitting relationship by the position of the reference point in the specified category after (e23).
13. The method according to any one of claims 1-12, characterized in that, (S20) includes, (S22) Based on the position of the reference point after (S10) and the reference template, determine the distribution of bright spots in each of the directions of the marked area, including identifying bright spots in multiple regions in each direction; (S24) Determine the fitting relationship for each direction based on the distribution of bright spots in each direction, and filter the bright spots to optimize the fitting relationship; (S26) Based on the fitting relationship after (S24), the reference point position of the image is further updated, and the coordinates of other regions or positions of the image are determined according to the further updated reference point position, so as to achieve the registration of the image with the reference template.
14. The method according to claim 13, characterized in that, (S22) includes, By using the position component of the reference point after (S10) in the direction and the relative positional relationship between the reference point and the bright spot contained in the reference coordinate system, the pixel range k1×k2 that the bright spot may be located in multiple regions in each direction is determined, where k1 and k2 are each independently an odd number greater than 1. and Based on the difference in pixel size between the center pixel and the surrounding pixels, the bright spot is identified within the pixel range k1×k2 where the bright spot may be located.
15. The method according to claim 13 or 14, characterized in that, The preset condition in (S10) is the first preset condition, and (S24) includes filtering the bright spots by the second preset condition and updating the fitting relationship in that direction based on the bright spots retained after filtering, wherein the second preset condition is related to the distance from the bright spot to the fitting relationship.
16. The method according to claim 15, characterized in that, The second preset condition includes multiple distance thresholds that converge from large to small, including the largest distance threshold D21 and the second largest distance threshold D22, (S24) includes, (S241) The bright spots are filtered according to the distance threshold D21, including retaining bright spots whose distance from the fitting relationship is less than or equal to the distance threshold D21, and optionally discarding bright spots whose distance from the fitting relationship is greater than the distance threshold D21; (S243) Update the fitted relationship using the bright spots retained after (S241); (S245) The bright spots retained after (S241) are filtered according to the distance threshold D22, including retaining the bright spots whose fitting relationship after (S243) is less than or equal to the distance threshold D22, and optionally discarding the bright spots whose fitting relationship after (S243) is greater than the distance threshold D22. as well as (S247) Update the fitted relationship using the bright spots retained after (S245).
17. The method according to any one of claims 8-16, characterized in that, At least one of the following steps or processes (i)-(iv) is performed in parallel: (i) The process of determining the positions of the plurality of reference points in the image in claim 8 (S10); (ii) The process of updating the fitting relationships in each direction in (S10) of claim 9 or 10; (iii) The optimization process of each fitting relationship in (S20) of claim 13; (iv) The process of determining the coordinates of other regions or locations of the image based on the further updated reference point position in (S26) of claim 13 (S20).
18. A method for detecting bases in an image, the image being derived from a detection system for sequencing based on patterned surface microscopy, wherein the surface region reflected in the image includes a plurality of adjacent labeled regions and reaction regions, characterized in that... include: Determine the signal intensity value of the location of the corresponding chemical feature in the reaction region of the image, wherein the image is an image registered by the method according to any one of claims 1-17; and The type of base introduced at this position is identified based on the signal strength value.
19. The method according to claim 18, characterized in that, The signal strength value is the corrected signal strength value.
20. A computer device, characterized in that, include: Memory, used to store applications and the data generated by the running of the applications; A processor for executing the application program to implement the image registration method as described in any one of claims 1-17; and / or the method for detecting image-identified bases as described in claim 18 or 19.
21. The computer device according to claim 20, characterized in that, The processor includes at least one set of processing units, each of which can be executed in parallel, and the set of processing units is used to perform at least one of the following steps or processes (i)-(iv): (i) When performing (S10) of the image registration method according to claim 8, the positions of a plurality of reference points in the image are determined; (ii) In performing (S10) of the image registration method according to claim 9 or 10, the respective fitting relationships in each direction are updated; (iii) When performing (S20) of the image registration method according to claim 13, optimize each fitting relationship; (iv) In performing (S26) of (S20) of the image registration method of claim 13, the coordinates of other regions or positions of the image are determined based on the further updated reference point positions.
Citation Information
Patent Citations
Method and device for detecting bright spots on image and image registration method and device
CN112285070A
Image registration method, device and computer program product
CN112288781A
Methods and systems for identifying bases in nucleic acids
CN113012757B
Methods and systems for analyzing image data
EP3077943B1
Single-molecule image correction method, device and system, and computer-readable storage medium
EP3336797A1
Cited By
Method for image registration, method for detecting and calling bases from images, and computer device
EP4738254A1