A spatial in situ sequencing method

By combining electric field-assisted mRNA directional migration and in situ reverse transcription with an image decoding platform, the problems of RNA degradation and signal drift in existing technologies have been solved, enabling high-resolution, high-throughput spatial transcriptomics research and constructing a spatial gene expression map at single-cell resolution, which is suitable for complex imaging scenarios.

CN122168727APending Publication Date: 2026-06-09XIAMEN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAMEN UNIV
Filing Date
2026-03-18
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing spatial transcriptomics technologies have limitations in terms of high throughput and high resolution. In particular, in situ sequencing methods suffer from RNA degradation, signal point drift and loss, insufficient resolution, and deficiencies in traditional registration strategies, making it difficult to achieve precise localization and analysis at the single-cell and subcellular levels.

Method used

By employing electric field-assisted directed migration of mRNA and in-situ capture with primers on the chip surface, combined with in-situ reverse transcription, multi-round coding probe hybridization, and fluorescence imaging, a self-developed image decoding platform was used to achieve efficient enrichment, signal amplification, and precise localization, thereby constructing a spatial gene expression map at single-cell resolution.

Benefits of technology

It significantly improves resolution and throughput in thick tissue sections and low signal-to-noise ratio scenarios, providing a high spatial resolution and high throughput spatial transcriptomics research tool suitable for complex imaging scenarios, and realizing gene expression detection and systematic analysis at single-cell resolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122168727A_ABST
    Figure CN122168727A_ABST
Patent Text Reader

Abstract

A spatial in situ sequencing method, belonging to the field of biology, is proposed. It utilizes an electric field-assisted directed migration of mRNA and in-situ capture with primers on a microarray surface to enrich tissue and release mRNA. In-situ reverse transcription generates covalently fixed cDNA, ensuring high positional stability during multiple rounds of hybridization and imaging. After reverse transcription, tissue is digested to remove tissue, reducing spatial hindrance and background interference, while the cDNA remains at its original coordinates due to covalent anchoring. Combining coding probe hybridization and RCA, single-molecule-level signal amplification and recognition are achieved. Through decoding and single-cell segmentation, transcripts are mapped to their respective cells, constructing a single-cell resolution spatial gene expression map. This method does not rely on multi-round DAPI mapping or other endogenous morphological marker-based multi-cycle image registration methods, making it suitable for high-throughput spatial in situ sequencing and significantly improving robustness and versatility in complex imaging scenarios such as thick tissue sections and low signal-to-noise ratios. It is applicable to high spatial resolution, high-throughput spatial transcriptome research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of molecular biology technology, and in particular to a spatial in situ sequencing method. Background Technology

[0002] Within living organisms, the spatial location of cells significantly influences gene expression, determining tissue function and disease progression. However, traditional single-cell transcriptome sequencing methods often neglect this spatial location information. Spatial transcriptomics, combining imaging, biomarkers, sequencing, and bioinformatics tools, spatially locates gene expression in tissue sections. This reveals the spatial distribution of different cell types within tissues, the interactions between different cell populations, and the mapping of gene expression atlases in different tissue regions, offering crucial applications for a deeper understanding of disease and cancer mechanisms.

[0003] Imaging-based spatial transcriptomics technology can achieve gene expression detection at single-cell or even subcellular resolution while preserving the in situ spatial coordinates of tissues, and construct multimodal spatial data that integrates molecular expression and tissue morphology.

[0004] Existing spatial transcriptomics technologies mainly include laser microdissection, in situ capture, in situ hybridization, and in situ sequencing. In situ hybridization and in situ sequencing can determine the location of RNA molecules at the subcellular level with high resolution. However, these methods have low throughput due to limitations in instruments and the number of fluorescence channels; and because the target molecules are not covalently linked to the substrate, signal point drift and loss are prone to occur during multiple rounds of hybridization and washing. At the same time, these methods are based on probe hybridization to locate RNA, so they can only detect known RNA sequences. In situ capture technologies such as ST (Nat.Methods 2016, 13(7): 597–600) and slide-seq (Science 2019, 363(6434): 1463–1467) can sequence RNA at high throughput, but their resolution often does not reach the single-cell level, and they have limitations such as low capture rate and gaps between capture points. At the same time, in situ capture methods can only obtain pixel-level spatial information and cannot provide detailed location and structural information at the intracellular or subcellular level. Existing sequencing methods such as EEL FISH (MethodsEnzymol. 2017, 592: 1–36) can capture and image RNA in situ from tissues, but signal points are easily lost due to RNA degradation, and the method is prone to signal congestion. To overcome these limitations, this invention develops a novel spatial in situ sequencing technology aimed at removing tissue background interference through nucleic acid blotting.

[0005] Furthermore, innovations in experimental techniques have rendered traditional analytical workflows for in situ sequencing inapplicable. For instance, DAPI staining maps, used as registration references in traditional in situ sequencing, are no longer collected in blotting techniques due to tissue removal. This necessitates the development of new benchmarks and algorithms to achieve accurate image registration. To better analyze in situ sequencing data from tissue-removed blotting, a matching and precise end-to-end decoding method is urgently needed. This invention provides a universal paradigm for in situ sequencing technology, particularly suitable for blotting-based techniques, and has broad applicability. Summary of the Invention

[0006] The purpose of this invention is to address the aforementioned problems in the prior art and provide a spatial in situ sequencing method, constructing a complete platform from tissue sample processing to data decoding and analysis. This invention is suitable for high spatial resolution and high-throughput spatial transcriptomics research, significantly improving resolution and throughput in complex imaging scenarios such as thick tissue sections and low signal-to-noise ratios.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] A spatial in situ sequencing method includes the following steps:

[0009] S1: Oligonucleotide capture primers are modified on the surface of a conductive glass slide to form a functionalized region. Tissue slices are attached to this functionalized region. An external electric field drives the mRNA released by tissue permeation to migrate in a directional manner and be captured in situ by primers on the chip surface.

[0010] S2: The captured mRNA is reverse transcribed in situ to generate cDNA covalently fixed on the chip surface. Then, the tissue matrix is ​​removed by enzymatic digestion to reduce steric hindrance and background interference.

[0011] S3: Perform multiple rounds of coding probe hybridization on the cDNA and combine it with in situ amplification to achieve single-molecule-level signal amplification, followed by multiple rounds of fluorescence imaging;

[0012] S4: A full-process analysis platform for image data decoding based on computer programs, which performs image registration, signal point identification and localization, decoding sequence information and cell segmentation on multiple rounds of fluorescence images;

[0013] S5: The decoded gene tags are assigned to the corresponding single-cell regions according to their spatial coordinates to construct a single-cell resolution spatial gene expression map.

[0014] The external electric field driving in step S1 includes: using copper tape containing conductive adhesive to connect wires to a conductive glass slide (e.g., an ITO conductive glass slide) carrying the sample and connecting it to the positive terminal of a DC power supply as the anode; using a clean ITO conductive glass slide as the cathode to vertically cover the sample and connect it to the negative terminal of the power supply; placing PDMS spacers on both sides of the sample; injecting electrophoresis buffer and applying DC voltage for electrophoresis to drive the mRNA to migrate in a directional manner and specifically bind to the capture primers on the chip surface.

[0015] The methods for in situ amplification and probe hybridization in step S3 include:

[0016] 1) The hybridization probe is circularized, RCA amplification is performed, and then hybridization is performed using a fluorescent probe, which is then detected by imaging;

[0017] 2) Amplification is performed using HCR, with fluorescent groups on the probe, and detection is achieved through imaging;

[0018] 3) Amplification is performed using the LAMP method, followed by hybridization using fluorescent probes, and detection is achieved through imaging.

[0019] Preferably, the in situ amplification and probe hybridization method in step S3 includes the following steps: designing specific probes with unique coding sequences for different target genes; after hybridization of the probes with cDNA, the probes are circularized by ligase and then amplified by RCA; in each sequencing cycle, adding fluorescently labeled readout probes complementary to specific sites of the coding sequence for hybridization imaging, and then eluting the probes to enter the next cycle.

[0020] The image registration in step S4 adopts a multi-cycle image registration method that does not rely on multiple rounds of DAPI maps or other endogenous morphological markers. It achieves cross-cycle image alignment by introducing an exogenous fluorescence reference, which includes a dual reference synergy mechanism of fluorescence reference patches and fluorescence microspheres.

[0021] The image registration specifically includes:

[0022] 1) Registration for fluorescent reference patches: The image to be transformed is brightness corrected to eliminate illumination differences between rounds. Then, feature detection algorithms (such as SIFT, ORB, BRISK and AKAZE) are used to extract key points and their feature orientation matrices of fluorescent reference patches in multiple rounds of imaging. The FLANN algorithm is used to perform approximate nearest neighbor search on the feature descriptors of key points to generate cross-cycle matching feature point pairs.

[0023] 2) For registration of fluorescent microspheres: using the density distribution characteristics of connected regions, the 8-neighborhood connected regions are labeled using the measure.label function of the skimage code package, the centroid coordinates are extracted, a distance distribution feature matrix is ​​constructed, and the optimal matching point pair is found using a linear assignment algorithm;

[0024] 3) Utilize the RANSAC algorithm to calculate the affine transformation matrix of the matching point pair set, achieving pixel-level alignment across rounds.

[0025] The signal point identification and localization in step S4 includes: preprocessing the image using top-hat transform and desharpening mask techniques to enhance blob contrast; performing blob detection using a signal point identification algorithm (such as OpenCV.SimpleBlobDetector or skimage.feature.peak_local_max); calculating the signal-to-noise ratio based on the gray-level difference between the blob center and its local background region; and defining the original quality score of the blob based on the comparison between the highest quality score and the scores of the other channels. The value is used for false positive filtering.

[0026] The decoded sequence information employs a local search strategy to correct positioning errors, based on a corrected quality score. The optimal spot is selected by minimization criteria, and then the fluorescent channels representing the signal are sequentially spliced ​​to generate a barcode to determine the gene identity.

[0027] In step S4, the cell segmentation uses the Cellpose algorithm based on a convolutional neural network to initially segment the cell nucleus, and then expands the boundary outward through morphological operations to reconstruct the complete cell region. When the cell boundaries are about to overlap, the expansion in the corresponding direction is immediately terminated.

[0028] The construction of a single-cell resolution spatial gene expression map in step S5 includes: assigning each transcript to the cell region where its spatial location falls based on a single-cell segmentation mask generated by the Cellpose algorithm; summarizing and quantitatively counting all assigned gene tags in each cell; and generating a cell-level spatial gene expression matrix for downstream cell type annotation, spatial differential expression analysis, cell state gradient inference, and cell-cell interaction network modeling.

[0029] Compared with the prior art, the beneficial effects achieved by the technical solution of this invention are:

[0030] This invention constructs an integrated platform encompassing tissue sample processing, in-situ molecular manipulation, image decoding, and bioinformatics analysis. It efficiently enriches tissue-released mRNA through the synergistic effect of electric field-assisted mRNA directional migration and in-situ capture using primers on the chip surface. Subsequently, in-situ reverse transcription generates covalently fixed cDNA, ensuring its highly stable position during multiple rounds of hybridization and imaging. Tissue removal after reverse transcription significantly reduces steric hindrance and background interference, while the cDNA remains strictly preserved in its original coordinates due to covalent anchoring. Further integration with coding probe hybridization and rolling circle amplification (RCA) achieves single-molecule-level signal amplification and highly specific recognition. Finally, relying on a self-developed Python decoding workflow and Cellpose single-cell segmentation, each transcript is precisely mapped to its corresponding cell, constructing a single-cell resolution spatial gene expression map. This invention provides a multi-cycle image registration method that does not rely on multiple rounds of DAPI maps or other endogenous morphological markers, suitable for high-throughput in-situ spatial sequencing. This method achieves cross-round image alignment by introducing an exogenous fluorescence benchmark, effectively overcoming the shortcomings of traditional DAPI-based nuclear staining registration strategies, which are prone to registration failure in samples with missing, blurred, or completely unlabeled nuclear signals. It significantly improves robustness and versatility in complex imaging scenarios such as thick tissue sections and low signal-to-noise ratios. This novel in situ spatial sequencing technology is suitable for high spatial resolution, high-throughput spatial transcriptomics research, providing a general technical paradigm for the systematic analysis of tissue microenvironments. Attached Figure Description

[0031] Figure 1 This is a flowchart of the present invention;

[0032] Figure 2 This is a flowchart of the data processing and decoding process of the present invention. Detailed Implementation

[0033] To make the technical problems, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0034] Imaging-based spatial transcriptomics technology can achieve gene expression detection at single-cell and even subcellular resolution while preserving the in-situ spatial coordinates of tissues, constructing multimodal spatial data that integrates molecular expression and tissue morphology. In situ sequencing cleverly transfers the high-throughput, sequencing-by-synthesis strategy of next-generation sequencing (NGS) to the microscopic imaging system of tissue sections, directly decoding the nucleic acid sequence of RNA transcripts in situ through multiple rounds of cyclic imaging. This invention effectively solves the problem of RNA degradation and instability in in-situ sequencing methods by capturing RNA in tissues in situ on primer-modified slides and reverse transcribing it into cDNA. Furthermore, the sequence is covalently bound to the slide surface, effectively solving the problems of RCA signal point shifting and detachment during imaging. In addition, tissue digestion after in-situ transfer of RNA to the slide surface can reduce background interference from the tissue during imaging. Figure 1 Finally, this invention provides a complete decoding method for in situ sequencing data after tissue removal from nucleic acid blots. Figure 2 For example, addressing the challenge of using multi-round DAPI maps as registration references in nucleic acid blot sequencing, a new benchmark reference and algorithm were invented to achieve accurate image registration. This invention enables high-resolution in situ detection of gene expression, providing an important tool for further histological research. Detailed steps are as follows:

[0035] S1: In situ capture of mRNA is achieved by using primers fixed on the chip to bind to an external electric field.

[0036] First, oligonucleotide capture primers were modified onto the surface of a conductive glass slide, and then tissue sections were attached to the functionalized region. Next, an electrophoresis apparatus was constructed: wires were connected to the conductive glass slide carrying the sample using copper tape containing conductive adhesive, and connected to the positive terminal of a DC power supply as the anode; the slide was placed in a custom fixture, with a 1.5 mm thick PDMS spacer strip placed on each side of the sample; a clean conductive glass slide was then used as the cathode, vertically covering the sample, ensuring its conductive surface faced the tissue, and connected to the negative terminal of the power supply via copper tape. Subsequently, electrophoresis buffer was injected into the gap between the two slides using a pipette, and a 3 V DC voltage was applied for electrophoresis for 20 minutes. This drove the mRNA released after tissue permeation to migrate directionally under the influence of the electric field and specifically bind to the capture primers on the chip surface.

[0037] S2: cDNA is generated and fixed on the chip surface through in situ reverse transcription to remove tissue background;

[0038] Capture primers immobilized on the chip surface can bind to mRNA released from tissue in situ and convert it into stable cDNA via reverse transcription. This cDNA is firmly anchored to the slide surface via covalent bonds, significantly enhancing its chemical stability and spatial localization fidelity during multiple rounds of hybridization and imaging, providing a reliable foundation for high-throughput in situ sequencing. Subsequently, protease digestion is used to remove tissue residues to reduce steric hindrance and suppress background fluorescence. The specific steps are as follows:

[0039] 1. Clean once using 1×RT buffer.

[0040] 2. Prepare the RT Mix on ice and mix thoroughly by blowing and whisking. See Table 1 for the recipe.

[0041] Table 1

[0042]

[0043] 3. Add the above RT mix into the window; incubate at 42°C for 2 hours.

[0044] 4. Remove the RT mix and wash with TE-SDS, TE-TW, and Tris-HCl (10 mM, pH=8.0) respectively;

[0045] 5. Prepare the tissue digestion solution at room temperature, mix thoroughly by pipetting, and refer to Table 2 for the formula;

[0046] Table 2

[0047]

[0048] 6. Add tissue digestion solution and incubate at 55°C for 25 minutes;

[0049] 7. After the tissue digestion is complete, remove the digestion solution, add 50µL of PMSF working solution (100µM PMSF isopropanol solution), and incubate at room temperature for 5 minutes.

[0050] S3: After probe hybridization and in situ amplification of cDNA, multiple rounds of fluorescence imaging are performed;

[0051] Specific probes with unique coding sequences (barcodes) are designed for different target genes. After hybridization with cDNA, the probes are circularized with a ligase and then amplified via RCA to generate amplification products. Because the RCA products are highly localized to the original cDNA sites, this method not only significantly enhances the detection signal-to-noise ratio but also effectively preserves the subcellular spatial resolution of the transcripts. In each sequencing cycle, a fluorescently labeled readout probe complementary to the specific site of the barcode is added, followed by hybridization and high-resolution fluorescence imaging; the probe is then eluted, and the next hybridization-imaging cycle begins. By combining multiple rounds of optical signals, the gene identity corresponding to each amplification site can be uniquely decoded. The specific steps include:

[0052] S3.1, Padlock probe hybridization

[0053] 1. Blocking: Add 50 μL Blocking Buffer (5% BSA) and incubate at room temperature for 1 h;

[0054] 2. Wash 3 times with 50 μL DEPC-PBST;

[0055] 3. Prepare Padlock hybridization reagent; see Table 3 for the formula.

[0056] Table 3

[0057]

[0058] 4. Add Padlock hybridization reagent and incubate at 37°C for 2 hours;

[0059] 5. Wash 3 times with 50 μL DEPC-PBST;

[0060] 6. Add 50 μL of 1×Hyb buffer2 (2× SSC, 20% formamide), incubate at 37℃ for 5 min, and repeat 3 times;

[0061] 7. Wash 3 times with 50 μL DEPC-PBST.

[0062] S3.2, Probe ring formation

[0063] 1. Prepare 50 μL of ligation reaction solution. See Table 4 for the formula.

[0064] Table 4

[0065]

[0066] 2. Add the reaction solution, incubate at 37°C for 0.5 h, then wash three times with 50 μL DEPC-PBST.

[0067] S3.3, RCA

[0068] 1. Prepare 50 μL of RCA reagent. See Table 5 for the formula.

[0069] Table 5

[0070]

[0071] 2. Add RCA reagent and incubate overnight at 30°C;

[0072] 3. Wash 3 times with 50 μL DEPC-PBST.

[0073] S3.4, ISS

[0074] 1. Prepare the ISS mix; see Table 6 for the recipe.

[0075] Table 6

[0076]

[0077] 2. Add 50 μL of reagent and incubate at 37°C for 45 min;

[0078] 3. Wash 3 times with 50 μL DEPC-PBST;

[0079] 4. Add 1 mL of DEPC-PBS to the slide, let it stand for 10 min, and then remove the coverslip;

[0080] 5. Wash 3 times with 50 μL DEPC-PBST;

[0081] 6. Add 150 μL stripping buffer, incubate at 37°C for 10 min, then add 150 μL formamide, incubate at 37°C for 10 min, repeat 2-3 times;

[0082] 7. Wash 3 times with 50 μL DEPC-PBST (see Table 7 for the formula), and observe under a microscope that there is no signal residue in each channel under normal exposure; proceed with the next imaging step.

[0083] Table 7

[0084]

[0085] S4: A full-process analysis platform for image data decoding based on computer programs;

[0086] This step aims to accurately reconstruct the identity and spatial location of each transcript from multiple rounds of in situ sequencing fluorescence images through a series of computational steps. First, all imaging cycles are aligned to a unified coordinate system through image registration; then, fluorescence signal points are detected and located in the registered images; next, signals from the same physical location in multiple imaging rounds are combined into a complete barcode to decode gene identity; simultaneously, single-cell segmentation is performed using the unique 4',6-diamidinyl-2-phenylindole nuclear staining (DAPI) images recorded before tissue digestion.

[0087] Therefore, the decoding process for in situ sequencing results essentially consists of image registration, signal point identification and localization, decoding sequence information, cell segmentation, and spatial single-cell gene map reconstruction. Image registration aims to align images from all imaging cycles to a unified coordinate system. Signal point identification and localization involves systematically detecting spots from the registered multi-round images and recording the final signal point location coordinates in a unified coordinate system. Decoding sequence information involves connecting signal points corresponding to the same spatial physical location in imaging from different cycles to form a complete barcode sequence. Cell segmentation utilizes image segmentation techniques, including those based on nucleus and cell membrane-specific markers, to identify single-cell regions. Specifically, it includes the following steps:

[0088] S4.1 Image Registration

[0089] Image registration aims to align images from all imaging cycles to a unified coordinate system. This invention develops a novel multi-cycle image registration strategy that achieves spatial alignment of fluorescence signals across cycles without requiring multiple DAPI images by introducing a dual exogenous reference collaborative mechanism. Specifically, a fluorescent reference patch and fluorescent microspheres are introduced as dual exogenous references. High-confidence matching point sets are collaboratively generated by fusing the patch keypoint pairs extracted by feature detection algorithms (e.g., SIFT, ORB, BRISK, and AKAZE) with the spatial density distribution features of the microspheres. Then, using the image from the first imaging cycle as a reference, the feature matrices of the keypoints are compared using the k-nearest neighbor algorithm (this method uses the KNN function from the FLANN algorithm). The high-confidence keypoint matching set is used to calculate the homography matrix of the projection transformation between the two imaging planes, achieving sub-pixel-level alignment across cycles. The specific method is as follows:

[0090] (1) A feature matching strategy based on rotation-invariant key points is adopted for the registration of fluorescent patches, and cross-cycle matching is achieved by modeling its geometric structure.

[0091] First, the given reference image With the image to be transformed Brightness correction is performed to eliminate lighting differences between rounds. The specific algorithm is as follows:

[0092] (1)

[0093] in This represents the mean of the image.

[0094] Subsequently, a feature detection algorithm is employed to extract key points and their characteristic orientation matrices from the fluorescent reference patch in multiple imaging rounds. For example, the SIFT algorithm is used, combining improved FAST corner detection, non-maximum suppression, key point orientation assignment, and rotation-invariant BRIEF descriptor generation, to extract scale- and orientation-invariant key points and their characteristic orientation matrices from each image.

[0095] Subsequently, the FLANN (Fast Library for Approximate Nearest Neighbors) algorithm is used to perform an approximate nearest neighbor search on the feature descriptors of key points, generating cross-period matching feature point pairs.

[0096] (2) A feature and quantization strategy based on the density distribution of connected regions is adopted for the registration of fluorescent microspheres. Cross-cycle matching is achieved by modeling its local spatial topology.

[0097] First, the `measure.label` function from the `skimage` package performs 8-connectivity labeling on the input binary image. This algorithm is based on the principle of connected component detection in graph theory: all adjacent (sharing edges or corners) foreground pixels in the image are grouped into the same region, and a unique integer label is assigned to each independent region, thus generating a label matrix of the same size as the original image. Then, the morphological and geometric properties of each labeled region are extracted using the `measure.regionprops` function from the `skimage` package. Based on this, to exclude small fragments caused by imaging noise or segmentation artifacts, an area threshold (area > 50 pixels) is set, retaining only connected regions that meet this condition. Finally, the centroid coordinates are extracted from the filtered region attributes as the spatial representation of each target. The centroid is defined as the arithmetic mean of the coordinates of all pixels in each region, using the formula:

[0098] (2)

[0099] Where A is the area of ​​the connected region, and R is the set of pixels in that connected region. The resulting list of coordinates forms an N×2 array, where each row corresponds to the center coordinates of a valid target. .

[0100] Next, a symmetric matrix is ​​constructed for the pairwise Euclidean distances between the centers of the connected regions. The off-diagonal elements are normalized and quantized into 12-bin histograms, which serve as the local spatial density distribution features of each point, and thus form the overall distance distribution feature matrix.

[0101] Image feature point matching is achieved based on the density distribution feature matrix of connected regions. First, the distance distribution feature vectors of the centers of each connected region in the two images are extracted. A similarity matrix is ​​constructed by calculating the correlation coefficient of the feature vectors. Then, a linear assignment algorithm is used to find the optimal matching point pair.

[0102] (3) Finally, the estimateAffinePartial2D function in the parameter estimation algorithm of RANSAC (Random Sample Consensus) is used to calculate the affine transformation matrix of the matching point pair set (including rotation, scaling and translation, a total of 4 degrees of freedom) to achieve pixel-level alignment across rounds.

[0103] S4.2 Signal Point Identification and Localization

[0104] Signal point identification and localization involves systematically detecting blobs from registered, multi-round images, implementing a false positive filtering mechanism, and recording the final signal point coordinates in a unified coordinate system. Specifically: First, top-hat transform and unsharp masking (USM) techniques are used to preprocess the image to enhance the contrast between the blobs and the background. Then, the Simple Blob Detector algorithm from the OpenCV package is used for blob detection. Further, the signal-to-noise ratio (SNR) is calculated based on the gray-level difference between the blob center and its local background region, and the signal quality score (qual) of the blobs at that location is calculated for each channel. Finally, the original quality score of the blobs is defined by comparing the highest quality score with the scores of the remaining channels. Value. Set by Thresholding for preliminary screening of candidate sites can effectively suppress false positive signals, significantly improving the decoding specificity and quantitative reliability of in situ sequencing while ensuring high recall. The formula is as follows:

[0105] (3)

[0106] (4)

[0107] (5)

[0108] in, The signal-to-noise ratio of the channel with the highest score. P is the sum of the signal-to-noise ratios of different channels at the same point in the same period, and P is the false positive probability value of the candidate point.

[0109] S4.3 Decoding Sequence Information

[0110] The decoded sequence information connects signal points corresponding to the same spatial physical location in imaging at different cycle times to form a complete barcode sequence. When connecting signals at each location, a local search strategy is used to correct for positioning errors, and the optimal candidate point is selected based on quality scores to generate a preliminary decoded sequence with confidence labels. Subsequently, combined with prior information such as known gene probe sequences and base quality fractions from the experimental design, the decoding results are filtered a second time to finally obtain high-confidence spatial gene expression data.

[0111] For each candidate signal point, this process traverses the signal point set within a predefined neighborhood space search range. Taking into account both signal quality and spatial consistency, and based on the corrected signal... The minimization criterion uniquely selects the optimal spot as the representative signal for a specific location in a unified coordinate system, as shown in the following formula:

[0112] (6)

[0113] Where d is the Euclidean distance from the candidate point to the specified spatial coordinates. This formula shows that the larger the distance d, the lower the corrected score.

[0114] Subsequently, the fluorescent channels representing the signal are sequentially spliced ​​to generate the corresponding barcode, which is then compared with a known gene coding table to determine the identity of the gene it represents.

[0115] S4.4, Cell Segmentation

[0116] Cell segmentation utilizes image segmentation techniques, including those based on nucleus and cell membrane-specific markers, to identify single-cell regions. The Cellpose model, a general cell segmentation model based on convolutional neural networks (CNNs), is employed to initially segment the cell nucleus. Based on this initial segmentation, morphological operations are used to expand the boundaries outwards to approximate the complete cell region. Specifically: First, the network is trained in a supervised manner using a large number of labeled cell images. Image features are automatically extracted through multi-layer convolution and pooling operations, and the network weights are optimized using a backpropagation algorithm, gradually approximating the actual annotations. After training, a predictive model suitable for inference is obtained. In the inference phase, the DAPI image to be segmented is input into the model. The network extracts features through forward propagation and outputs a probability map of each pixel belonging to the cell nucleus. Accurate segmentation of the cell nucleus is achieved based on a probability threshold. Based on the segmented cell nucleus boundaries, the expand algorithm from the OpenCV package is used for controlled expansion. Expansion in the corresponding direction is terminated immediately when cell boundaries are about to overlap, ensuring that the final obtained cell regions do not overlap.

[0117] S5: The decoded gene tags are assigned to their respective single-cell regions to construct a single-cell resolution spatial gene expression map;

[0118] The spatial gene map reconstruction step maps the decoded transcript information back to the original spatial coordinates of the tissue, thereby constructing a single-cell resolution spatial gene expression map, which specifically includes the following:

[0119] After the aforementioned process, each transcript is assigned precise (x, y) spatial coordinates and its corresponding gene identity tag. Subsequently, based on the single-cell segmentation mask generated by the Cellpose algorithm, each transcript is assigned to the cell region in which its spatial location falls. By summarizing and quantitatively counting all assigned gene tags within each cell, a cell-level spatial gene expression matrix map is finally generated. This method not only preserves the original spatial topology of the tissue microenvironment with high fidelity but also provides a solid foundation for downstream multidimensional biological analysis, including cell type annotation, spatial differential expression analysis, cell state gradient inference, and cell-cell interaction network modeling.

[0120] This invention proposes a novel in situ sequencing technology, constructing an integrated platform encompassing tissue sample processing, in situ molecular manipulation, image decoding, and bioinformatics analysis. The invention utilizes the synergistic effect of electric field-assisted mRNA directional migration and in situ capture by primers on the chip surface to efficiently enrich tissue-released mRNA. Subsequently, in situ reverse transcription generates covalently fixed cDNA, ensuring its highly stable position during multiple rounds of hybridization and imaging. Post-reverse transcription digestion removes tissue, significantly reducing steric hindrance and background interference, while the cDNA remains strictly preserved in its original coordinates due to covalent anchoring. Further integration with coding probe hybridization and rolling circle amplification (RCA) achieves single-molecule-level signal amplification and highly specific recognition. Finally, relying on a self-developed Python decoding workflow and Cellpose single-cell segmentation, each transcript is precisely mapped to its corresponding cell, constructing a single-cell resolution spatial gene expression map. This invention provides a multi-cycle image registration method that does not rely on multiple rounds of DAPI maps or other endogenous morphological markers, suitable for high-throughput in situ sequencing. This method achieves cross-round image alignment by introducing an exogenous fluorescence benchmark, effectively overcoming the shortcomings of traditional DAPI-based nuclear staining registration strategies, which are prone to registration failure in samples with missing, blurred, or completely unlabeled nuclear signals. It significantly improves robustness and versatility in complex imaging scenarios such as thick tissue sections and low signal-to-noise ratios. This novel in situ spatial sequencing technology is suitable for high spatial resolution, high-throughput spatial transcriptomics research, providing a general technical paradigm for the systematic analysis of tissue microenvironments.

[0121] The above are merely specific embodiments of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantial modifications made to the present invention using this concept shall be considered as infringing upon the protection scope of the present invention.

Claims

1. A spatial in situ sequencing method, characterized in that, Includes the following steps: S1: Oligonucleotide capture primers are modified on the surface of a conductive glass slide to form a functionalized region. Tissue slices are attached to this functionalized region. An external electric field drives the mRNA released by tissue permeation to migrate in a directional manner and be captured in situ by primers on the chip surface. S2: The captured mRNA is reverse transcribed in situ to generate cDNA covalently fixed on the chip surface. Then, the tissue matrix is ​​removed by enzymatic digestion to reduce steric hindrance and background interference. S3: Perform multiple rounds of coding probe hybridization on the cDNA and combine it with in situ amplification to achieve single-molecule-level signal amplification, followed by multiple rounds of fluorescence imaging; S4: A full-process analysis platform for image data decoding based on computer programs, which performs image registration, signal point identification and localization, decoding sequence information and cell segmentation on multiple rounds of fluorescence images; S5: The decoded gene tags are assigned to the corresponding single-cell regions according to their spatial coordinates to construct a single-cell resolution spatial gene expression map.

2. The in situ sequencing method as described in claim 1, characterized in that, The external electric field driving in step S1 includes: using copper tape containing conductive adhesive to connect wires to a conductive glass slide carrying the sample and connecting it to the positive terminal of a DC power supply as the anode; using a clean conductive glass slide as the cathode to vertically cover the sample and connect it to the negative terminal of the power supply; placing PDMS spacers on both sides of the sample; injecting electrophoresis buffer and applying DC voltage for electrophoresis to drive the mRNA to migrate in a directional manner and specifically bind to the capture primers on the chip surface.

3. The in-situ sequencing method as described in claim 1, characterized in that, The methods for in situ amplification and probe hybridization in step S3 include: 1) The hybridization probe is circularized, RCA amplification is performed, and then hybridization is performed using a fluorescent probe, which is then detected by imaging; 2) Amplification is performed using HCR, with fluorescent groups on the probe, and detection is achieved through imaging; 3) Amplification is performed using the LAMP method, followed by hybridization using fluorescent probes, and detection is achieved through imaging.

4. The spatial in situ sequencing method as described in claim 1, characterized in that, The in situ amplification and probe hybridization method in step S3 includes the following steps: designing specific probes with unique coding sequences for different target genes; after hybridization of the probes with cDNA, they are circularized by ligase and then amplified by RCA; in each sequencing cycle, fluorescently labeled readout probes complementary to specific sites of the coding sequence are added for hybridization imaging, and then the probes are eluted to enter the next cycle.

5. The in situ sequencing method as described in claim 1, characterized in that: The image registration in step S4 adopts a multi-cycle image registration method that does not rely on multiple rounds of DAPI maps or other endogenous morphological markers. It achieves cross-cycle image alignment by introducing an exogenous fluorescence reference, which includes a dual reference synergy mechanism of fluorescence reference patches and fluorescence microspheres.

6. The in situ sequencing method as described in claim 5, characterized in that, The image registration specifically includes: 1) Registration for fluorescent reference patches: The image to be transformed is brightness corrected to eliminate illumination differences between rounds. Then, a feature detection algorithm is used to extract the key points and their feature orientation matrices of the fluorescent reference patches in multiple rounds of imaging. The FLANN algorithm is used to perform an approximate nearest neighbor search on the feature descriptors of the key points to generate cross-cycle matching feature point pairs. 2) For registration of fluorescent microspheres: using the density distribution characteristics of connected regions, the 8-neighborhood connected regions are labeled using the measure.label function of the skimage code package, the centroid coordinates are extracted, a distance distribution feature matrix is ​​constructed, and the optimal matching point pair is found using a linear assignment algorithm; 3) Utilize the RANSAC algorithm to calculate the affine transformation matrix of the matching point pair set, achieving pixel-level alignment across rounds.

7. The in situ sequencing method as described in claim 1, characterized in that, The signal point identification and localization in step S4 includes: preprocessing the image using top-hat transform and desharpening mask techniques to enhance spot contrast; detecting spots using a signal point identification algorithm; calculating the signal-to-noise ratio based on the gray-level difference between the spot center and its local background region; and defining the original quality score of the spot based on the comparison between the highest quality score and the scores of the other channels. The value is used for false positive filtering.

8. The in situ sequencing method as described in claim 1, characterized in that: The decoded sequence information employs a local search strategy to correct positioning errors, based on a corrected quality score. The optimal spot is selected by minimization criteria, and then the fluorescent channels representing the signal are sequentially spliced ​​to generate a barcode to determine the gene identity.

9. The in situ sequencing method as described in claim 1, characterized in that: In step S4, the cell segmentation uses the Cellpose algorithm based on a convolutional neural network to initially segment the cell nucleus, and then expands the boundary outward through morphological operations to reconstruct the complete cell region. When the cell boundaries are about to overlap, the expansion in the corresponding direction is immediately terminated.

10. The in situ sequencing method as described in claim 1, characterized in that, The construction of a single-cell resolution spatial gene expression map in step S5 includes: assigning each transcript to the cell region where its spatial location falls based on the single-cell segmentation mask generated by the Cellpose algorithm, summarizing all assigned gene tags in each cell and performing quantitative counting to generate a cell-level spatial gene expression matrix.