A probe design method and system for multi-round spatial transcriptome imaging

By optimizing targeted and non-targeted probe design and simulated annealing algorithm to optimize gene coding books, the compatibility and optical crowding problems in existing spatial transcriptional composition imaging technologies are solved, and efficient probe capture and coding strategies are achieved, improving imaging quality.

CN120220832BActive Publication Date: 2025-08-26ARTIFICIAL INTELLIGENCE RES INST OF HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ARTIFICIAL INTELLIGENCE LAB)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510695492.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-26
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The existing spatial transcriptional composition imaging technology has problems such as weak compatibility in probe design, incomplete disclosure of coding book design, many pollution problems in experiments, and serious optical crowding, making it difficult to achieve efficient and highly compatible probe design.

Method used

A probe design method for multi-round spatial transcriptional composition imaging was designed, including optimized screening of targeted region probes and non-targeted region probes, combined with simulated annealing algorithm to optimize gene coding books, construct gene expression matrix through GEO database, and multiple rounds of imaging using targeted and non-targeted probes, optimized coding strategies to reduce optical congestion.

Benefits of technology

It improves the capture efficiency of the probe, enhances the true positive signal, reduces false positive signals, is compatible with a variety of coding strategies, optimizes the encoding book generation and iteration process, and reduces the optical crowding phenomenon in the experiment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220832B_ABST
    Figure CN120220832B_ABST
Patent Text Reader

Abstract

The present invention discloses a probe design method and system for multi-round spatial transcriptome imaging, which relates to the technical field of bioinformatics, including: constructing a retrieval database; traversing and obtaining the targeted segments of all isoforms of the same gene from known RNA sequences and preprocessing them, dividing the obtained candidate probe sequences into two segments for additional non-specific detection to obtain targeted probe sequences for all isoforms, and then screening all expression matrices in the retrieval database to obtain targeted region probes; generating multiple random sequences through random ATCG and preprocessing them to design non-targeted region probes based on them; obtaining a single-cell expression matrix from the retrieval database, and optimizing the cell type-level expression variance of a given gene set between different imaging rounds based on a simulated annealing algorithm to minimize the gene codebook; the probe design method and system improve the overall capture efficiency of the probe to a certain extent and are compatible with multiple encoding strategies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and in particular to a probe design method and system for multi-round spatial transcriptome imaging. Background Art

[0002] Single-cell transcriptomics technology can capture gene expression at the single-cell level, but the dissociation of tissues into single-cell suspensions will result in the loss of spatial and morphological information, making it difficult to study features related to spatial scale. Spatial transcriptomics technology can effectively obtain gene expression information and spatial location information of tissues, providing a solution to this problem. Spatial transcriptomics technology is divided into sequencing-based technology and imaging-based technology. Although sequencing-based spatial transcriptomics technology can detect genes at the scale of the entire transcriptome, it is difficult to achieve single-cell precision resolution; imaging-based spatial transcriptomics technology can detect the spatial distribution of the transcriptome with high precision and high resolution, and can better depict the spatial distribution of individual cells in tissues.

[0003] However, imaging-based spatial transcriptomics is highly dependent on the gene capture rate and designed probe sequence. Specifically, different technologies vary in probe design structure, target region specificity, and gene expression levels in different tissues. Furthermore, the specific probe sequences for current mainstream technologies, such as multiplexed error-corrected in situ hybridization (MERFISH) and spatial transcriptomics with amplified readout (STARmap), are not fully public, and their designs still have limitations and room for improvement, such as the impact of potential protein binding sites. Therefore, a new, convenient, practical, and highly compatible probe design system is needed to provide an optimized solution.

[0004] Imaging-based spatial transcriptomics generally increases gene throughput by simultaneously adding multiple different gene combinations over multiple rounds, followed by decoding to determine the correct position. Codebook design generally considers averaging light intensity as much as possible after removing genes with high and low expression levels, thus minimizing optical crowding and maximizing gene capacity. However, codebook design methods are currently not fully publicized and lack generalizability across different technologies. Furthermore, potential residual contamination from experiments can be optimized in codebook design. Therefore, a compatible and practical codebook optimization algorithm is needed. Summary of the Invention

[0005] Based on the technical problems existing in the background technology, the present invention proposes a probe design method and system for multi-round spatial transcriptome imaging, which can improve the overall capture efficiency of the probe and is compatible with multiple encoding strategies.

[0006] The present invention proposes a probe design method for multi-round spatial transcriptome imaging, comprising:

[0007] Download the original fastq data from the GEO database and preprocess it to obtain isoform and gene expression matrices, taking into account the single-cell expression matrix, and construct a search database for bulkRNA and scRNA;

[0008] Designing targeted region probes: traverse and preprocess the target regions of all isoforms of the same gene from known RNA sequences. Divide the resulting candidate probe sequences into two segments for additional nonspecific detection to obtain target probe sequences for all isoforms. Target region probes are then screened based on all expression matrices in the search database.

[0009] Design of non-targeted region probes: Generate multiple random sequences through random ATCG and pre-process them to design non-targeted region probes;

[0010] Gene codebook optimization: Obtain a single-cell expression matrix from a search database and optimize the gene set based on a simulated annealing algorithm to minimize the variance of cell type-level expression across different imaging cycles to obtain a gene codebook.

[0011] Gene space display: Multiple rounds of spatial transcriptome imaging are performed using targeted region probes, non-targeted region probes, and gene code books to obtain gene space display decoded according to the gene code book.

[0012] Furthermore, in designing the target region probe, the target segments of all isoforms of the same gene are traversed from the known RNA sequence and preprocessed, specifically:

[0013] Based on the input species name, required gene set and corresponding isoform name, the sequence is traversed from the fasta file that stores all known RNA sequences according to the required design length to generate a potential probe sequence set;

[0014] By setting the Tm value, GC content, removal of concatemers, RNA secondary structure detection, and overall BLAST nonspecificity check, the potential probe sequence set was screened to obtain nonspecific alignment results and candidate probe sequences.

[0015] Furthermore, in designing the targeted region probe, the candidate probe sequence obtained is divided into two segments for additional non-specific detection to obtain the targeted probe sequence of all isoforms. Then, the targeted region probe is obtained by screening all expression matrices in the search database, specifically:

[0016] The candidate probe sequence is divided into two segments and then subjected to separate BLAST nonspecific detection. The nonspecific removal of the two segment alignment results is performed according to the set unilateral nonspecific binding threshold;

[0017] Obtain protein-binding RNA clip-seq data from known public databases and filter them based on enrichment intensity, retaining potential binding sites with confidence exceeding a set threshold to obtain filtered clip-seq data;

[0018] Set the minimum overlap ratio threshold, retain the probe sequences after nonspecific removal and the sequences whose overlap ratio with the filtered clip-seq data does not exceed the set overlap ratio threshold, and obtain the targeted probe sequences of all isoforms;

[0019] According to the expression matrix of bulk RNA of the set species tissue in the retrieval database, the target segment shared by the highest expressed isoform and other isoforms in the same gene classification is preferentially selected as the target region probe, and a set number of target region probes are selected according to the expression levels of bulk RNA and scRNA in the retrieval database.

[0020] Furthermore, non-targeted region probes were designed: multiple random sequences were generated by random ATCG and pre-processed, and non-targeted region probes were designed based on them, specifically:

[0021] Generate a sufficient amount of random sequences of fixed length by random ATCG and retain the sequences that do not specifically bind to known RNA sequences by BLAST as potential sequence sets;

[0022] Calculate the Hamming distance of all lengths between any two sequences in the potential sequence set, and retain the sequences whose Hamming distance exceeds the set threshold.

[0023] The first 1 / 3 segment and the last 1 / 3 segment of each retained sequence were marked as edge segments, and the Hamming distance between each edge segment was calculated. Sequences with Hamming distance exceeding the set threshold were selected and retained, and non-target region probes were designed based on them.

[0024] Furthermore, one targeted region probe plus two non-targeted region probes constitute a rolling circle probe, one non-targeted region probe plus a fluorescent group constitutes a fluorescent probe, and the reverse complement of one non-targeted region probe constitutes a primer probe. The rolling circle probe, the fluorescent probe and the primer probe are assembled to obtain an experimental probe.

[0025] Furthermore, in the gene codebook optimization, the details are as follows:

[0026] The expression of bulk RNA in kilobases per million sequencing reads and the average raw counts of scRNA at the cell type level for a given gene set in the tissues of a given species were obtained from the search database. Genes with expression levels exceeding the upper expression threshold or below the lower expression threshold were not allowed to be included in the encoding round. The gene sets obtained by screening were used for encoding book design and subsequent encoding imaging experiments.

[0027] Set the number of encoding rounds and the number of encoding imaging genes, and iteratively optimize the distribution of the selected gene set among the round channels;

[0028] In each iteration, one of the four exchange strategies is randomly selected to process the gene set distribution after screening;

[0029] If the variance of the gene set distribution after the exchange strategy is less than the variance before the exchange, the next round will be entered, where the variance is the total number of different gene combinations highlighted in each round under the gene set distribution calculated based on the expression matrix at the scRNA cell type level, and then the average variance of this total number between each round is calculated;

[0030] If the variance after the exchange strategy is greater than the variance before the exchange, then randomly generate a value to determine whether it is less than ,in, is the variance difference before and after the exchange strategy, is the annealing temperature, which decays with the number of iterations;

[0031] If less than Otherwise, the exchange strategy is reselected until the iteration requirements are met.

[0032] Furthermore, the four exchange strategies are randomly exchanging two rows, exchanging two rows multiple times, randomly exchanging multiple rows in their entirety, and randomly exchanging the entire area in reverse order.

[0033] Furthermore, in the optimization of the genetic codebook, the generation of initial values ​​is constrained, specifically:

[0034] The encoding of the previous and next rounds of the same channel has been removed;

[0035] The number of genes in each round is calculated, and the gene set distribution with less than the preset number of genes in each round is removed, so that the number of genes in each round is within the set interval, thereby generating constraints on the initial value.

[0036] A probe design system for multi-round spatial transcriptome imaging, comprising a database construction module, a targeted region probe design module, a non-targeted region probe design module, a gene codebook optimization module, and a gene display module;

[0037] The database construction module is used to download raw fastq data from the GEO database and preprocess it to obtain isoform and gene expression matrices, taking into account the single-cell expression matrix, and then construct a retrieval database for bulkRNA and scRNA;

[0038] The targeted region probe design module is used to traverse and pre-process the target segments of all isoforms of the same gene from known RNA sequences, divide the obtained candidate probe sequences into two segments for additional non-specific detection, and obtain the targeted probe sequences of all isoforms. The targeted region probes are then screened based on all expression matrices in the search database.

[0039] The non-targeted region probe design module is used to generate multiple random sequences through random ATCG and pre-process them to design non-targeted region probes;

[0040] The gene codebook optimization module is used to obtain the single-cell expression matrix from the search database and optimize the gene set based on the simulated annealing algorithm to minimize the variance of cell type-level expression between different imaging rounds to obtain the gene codebook.

[0041] The gene display module is used to perform multiple rounds of spatial transcriptome imaging using targeted region probes, non-targeted region probes and gene code books to obtain gene spatial display after decoding according to the gene code books.

[0042] Furthermore, the non-targeted region probe design module is specifically used to:

[0043] Generate a sufficient amount of random sequences of fixed length by random ATCG and retain the sequences that do not specifically bind to known RNA sequences by BLAST as potential sequence sets;

[0044] Calculate the Hamming distance of all lengths between any two sequences in the potential sequence set, and retain the sequences whose Hamming distance exceeds the set threshold.

[0045] The first 1 / 3 segment and the last 1 / 3 segment of each retained sequence were marked as edge segments, and the Hamming distance between each edge segment was calculated. Sequences with Hamming distance exceeding the set threshold were selected and retained, and non-target region probes were designed based on them.

[0046] The advantages of the probe design method and system for multi-round spatial transcriptome imaging provided by the present invention are: the system can be optimized from the specificity of the targeted area and the non-specificity of the unbound area, and the overall capture efficiency of the probe can be improved to a certain extent, that is, the true positive signal is enhanced and the false positive signal is reduced. In addition, for multi-round high-throughput gene encoding, a simulated annealing algorithm optimization algorithm is provided. The algorithm is compatible with multiple encoding strategies and optimizes the initial value generation and iteration process of the encoding based on potential contamination phenomena in the experiment. At the same time, it supports feedback adjustment with actual experimental results to achieve an encoding strategy with minimal optical congestion in real conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic diagram of the process of the present invention;

[0048] Figure 2 Schematic diagram of an example search for scRNA / bulkRNA in a mouse kidney dataset;

[0049] Figure 3 Schematic diagram of the variance comparison of the coding design for the mouse brain dataset;

[0050] Figure 4 Schematic diagram of the probe assembly structure;

[0051] Figure 5 Output result diagram for the encoding scheme;

[0052] Figure 6 Shown diagram for the gene example. DETAILED DESCRIPTION

[0053] The technical solutions of the present invention are described in detail below through specific embodiments. Numerous specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art may make similar modifications without departing from the scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0054] like Figures 1 to 6 As shown, the present invention proposes a probe design method for multi-round spatial transcriptome imaging, comprising:

[0055] Step 1: Download the original fastq data from the GEO database and preprocess it to obtain isoforms and gene expression matrices, taking into account the single-cell expression matrix, and construct a search database for bulk RNA and scRNA;

[0056] Step 2: Design targeted region probes: traverse and preprocess the target regions of all isoforms of the same gene from the known RNA sequence. Split the resulting candidate probe sequences into two segments for additional nonspecific detection to obtain target probe sequences for all isoforms. Target region probes are then screened based on all expression matrices in the search database.

[0057] Step 3: Designing non-target region probes: Multiple random sequences are generated and pre-processed using random ATCG, based on which non-target region probes are designed. ATCG stands for adenine (A), thymine (T), guanine (G), and cytosine (C), which are the core elements that make up the DNA double helix structure.

[0058] Step 4: Gene codebook optimization: Obtain a single-cell expression matrix from the search database, and optimize the gene set based on the simulated annealing algorithm to minimize the variance of cell type-level expression between different imaging rounds to obtain the gene codebook;

[0059] Step 5: Gene space display: Multiple rounds of spatial transcriptome imaging are performed using targeted region probes, non-targeted region probes, and gene code books to obtain gene space display decoded according to the gene code books.

[0060] This embodiment provides a multi-round imaging spatial transcriptome probe design method, which can systematically optimize the specificity of the targeted area and the non-specificity of the non-targeted area, and has a certain improvement in the overall capture efficiency of the probe, that is, enhancing the true positive signal and reducing the false positive signal. In addition, for multi-round high-throughput gene encoding, this design provides a simulated annealing algorithm optimization algorithm, which is compatible with multiple encoding strategies and optimizes the initial value generation and iteration process of the encoding based on potential contamination phenomena in the experiment. At the same time, it supports feedback adjustment based on the actual experimental results to achieve the encoding strategy with the least optical congestion in the real situation.

[0061] The data collected and processed in this embodiment, the design of the targeted region probe and the non-targeted region probe, and the gene coding optimization method are all made public in the form of a web-based database to lower the coding threshold for users and improve practical usability.

[0062] In this example, the probe design method mainly includes three functional modules: 1) construction of bulk RNA and scRNA databases of human and mouse RNA isoforms and genes from various tissues; 2) optimization screening strategies for the targeted and non-targeted regions of the isoforms in the probe structure; and 3) a multi-round high-throughput gene encoding design optimization algorithm based on simulated annealing.

[0063] In one embodiment, bulk RNA and scRNA database searches of various tissues for human and mouse isoforms and genes, i.e., step 1, are specifically as follows:

[0064] Based on information from public literature, we collected data on 56 bulk RNA and scRNA types, including 37 human tissues and 19 mouse tissues. The scRNA data were derived from downstream expression matrices processed from public databases. The bulk RNA data were obtained by downloading raw fastq data from the GEO databases associated with various literature using the sra tool. The data were then systematically processed using trim_galore, tophat, and cufflinks software tools to obtain isoform and gene expression matrices, thereby constructing a retrieval database for bulk RNA and scRNA. Finally, the data was organized and the relevant screening code was written and deployed on the web. The sra tool is a software toolkit for processing high-throughput sequencing data, fastq is a text format for storing biological sequences and their quality information, and is the standard storage format for high-throughput sequencing data. trim_galore is used for data preprocessing and quality control, tophat is used for sequence alignment and splice site identification, cufflinks is used for transcript assembly and expression quantification, scRNA is single-cell RNA, and bulkRNA is sample transcriptome sequencing.

[0065] Researchers can obtain the expression of specific gene sets of specific tissues of their interest based on the constructed bulkRNA and scRNA retrieval database, which can serve as a reference for subsequent probe design and determine the appropriate number of target fragments and whether to encode them. Figure 2 As shown, the expression of genes of interest in the mouse kidney dataset is retrieved. Among them, genes with higher expression levels in scRNA or bulkRNA are marked in red. Subsequent non-coding imaging and a moderate reduction in the number of targeted fragments are adopted. For genes with lower expression levels, the number of targets can be increased. Design is not recommended for genes with particularly low expression levels. Gene sets with moderate expression levels can be used for subsequent encoding to achieve high-throughput imaging with fewer rounds.

[0066] The transcriptome probes designed in this example target specific isoforms. The varying expression levels of different isoforms significantly impact capture efficiency. Existing publicly available bulk RNA and scRNA databases do not include this information. This example, by collaborating on literature, downloading, and processing raw bulk RNA and downstream scRNA data, provides an important data reference for human and mouse transcriptome probe design within the field.

[0067] In one embodiment, an optimization screening strategy for the targeted and non-targeted regions of the isoform in the probe structure is provided. The various spatial transcriptome technologies disclosed have different probe design structures due to different design logics, but all include two parts: a targeted region that specifically binds to the isoform segment and a non-targeted region that cannot specifically bind to known RNA sequences. This embodiment has made certain optimization innovations in both aspects.

[0068] Specifically, the design of the target region probe in step 2 is as follows:

[0069] (a1) Based on the input species name, required gene set, and corresponding isoform name, the sequence is traversed from the fasta file that stores all known RNA sequences according to the required design length to generate a potential probe sequence set;

[0070] (a2) Screening the potential probe sequence set to obtain nonspecific alignment results by setting melting temperature, GC content, concatemer removal, RNA secondary structure detection (e.g., loop and dimer removal), and overall BLAST nonspecificity check to obtain candidate probe sequences;

[0071] Among them, the Tm value is the melting temperature, the GC content refers to the ratio of guanine (G) and cytosine (C) to the total bases in the nucleic acid molecule, and BLAST (Basic Local Alignment Search Tool) is a widely used bioinformatics tool for quickly aligning biological sequences (such as DNA, RNA or protein) in databases to find similar sequences and infer functional or evolutionary relationships.

[0072] (a3) Divide the candidate probe sequence into two segments and perform BLAST nonspecific detection separately. The nonspecific removal is performed on the two segment alignment results according to the set unilateral nonspecific binding threshold.

[0073] For example, if the candidate probe sequence is 60 nt, the 60 nt is divided into the first 30 nt and the last 30 nt and then non-specifically aligned using BLAST to remove probe sequences with non-specific binding of more than 20 nt on one side. This will reduce the non-specific binding of continuous and complete matching of the entire target segment and make the probe structure more stable.

[0074] (a4) Obtaining clip-seq data of protein-bound RNA from known public databases and filtering them based on enrichment intensity, retaining potential binding sites with confidence exceeding a set threshold, and obtaining filtered clip-seq data;

[0075] Clip-seq data is based on data from existing public databases of human and mouse protein-binding RNA fragments. For example, a peak enrichment intensity greater than 10 is set to filter the clip-seq data, where peaks specifically refer to regions with significant signal intensity in genomic sequencing data.

[0076] (a5) Setting a minimum overlap ratio threshold, retaining probe sequences after nonspecific removal whose overlap ratio with the filtered clip-seq data does not exceed the overlap ratio threshold, and obtaining targeted probe sequences for all isoforms;

[0077] By searching for overlapping regions (bedtools intersect), sequences with an overlap ratio threshold exceeding the threshold (for example, the overlap ratio threshold is set to 0.2) are removed, thereby removing potential protein binding sites for RNA that may remain in the experiment.

[0078] (a6) Based on the expression matrix of bulk RNA of the set species tissue in the retrieval database, the target segments shared by the isoform with the highest expression and other isoforms under the same gene classification are preferentially selected as the target region probes, and a set number of target region probes are selected based on the expression level of the gene in bulk RNA and scRNA in the retrieval database. The selection principle is to select a smaller number of probes if the expression level is too high. If the number of shared segments is insufficient for the design, the unique segment of the isoform with the highest expression level is used as a supplement.

[0079] Specifically, the design of non-target region probes in step 3 is as follows:

[0080] For non-targeted regions, different probe designs use BLAST to identify sequences that do not bind to specific known RNA sequence segments. However, non-targeted regions often have junctions with other known RNA sequence binding segments, and there are currently no effective screening methods for the non-specificity of these junctions. This example addresses this issue by proposing a Hamming distance-based edge overlap screening method for the first time. This method significantly reduces the Hamming distance between the edge regions of probes in non-targeted regions, thereby reducing potential signal crosstalk contamination.

[0081] Specifically, a sufficient number of random sequences of a fixed length (for example, 60 nt) are generated through random ATCG, and sequences that do not specifically bind to known RNA sequences are retained as a potential sequence set through BLAST. The Hamming distances of all lengths between sequences in the potential sequence set are calculated, and sequences whose Hamming distances exceed the set threshold are retained. The first 1 / 3 segment (first 20 nt) and the last 1 / 3 segment (last 20 nt) of each retained sequence are marked as edge segments, and the Hamming distances between each pair of edge segments are calculated. Sequences whose Hamming distances exceed the set threshold are retained and used to design non-targeted region probes.

[0082] In one embodiment, a multi-round high-throughput gene encoding design optimization algorithm based on simulated annealing, i.e., step 4, specifically includes:

[0083] This embodiment implements a gene encoding optimization algorithm based on the idea of ​​simulated annealing. The optimization target is the variance of the sum of the expression levels of cell types in the tissue-paired scRNA data between different rounds. The smaller the variance, the more even the light intensity between each round. Since the paired scRNA and the spatial transcriptome are known to have a high correlation, it can be used to minimize the possibility of optical crowding and improve imaging quality under a given gene set. The core idea of ​​the algorithm proposed in this embodiment is the principle of simulated annealing, which is as follows:

[0084] (b1) Obtain the expression of bulk RNA kilobases per million (FPKM) and scRNA cell type-level average raw counts for a given gene set in tissues of a given species from the search database; genes with expression levels exceeding the upper expression threshold or below the lower expression threshold are not allowed to be included in the coding round. The gene set is screened for coding book design and subsequent coding imaging experiments;

[0085] (b2) Setting the number of encoding rounds and the number of encoding imaging genes, and iteratively optimizing the distribution of the selected gene set across round channels;

[0086] (b3) In each iteration, one of the four exchange strategies is randomly selected to process the gene set distribution after screening; if the variance of the gene set distribution after adopting the exchange strategy is less than the variance before the exchange, it enters the next round, where the variance is the total number of different gene combinations highlighted in each round under the gene set distribution according to the expression matrix at the scRNA cell type level, and then the average variance of this total number between each round is calculated; if the variance after adopting the exchange strategy is greater than the variance before the exchange, it is judged by randomly generating a value whether it is less than ,in, is the difference between the before and after variances, is the annealing temperature, which decays with the number of iterations; if it is less than Otherwise, the exchange strategy is reselected until the iteration requirements are met.

[0087] That is, in the early stage of an iteration process, a certain probability is allowed to accept the solution with increased variance to avoid entering the local minimum rather than the global minimum, but as the number of iterations increases, the acceptance probability gradually decreases, and finally the optimal solution in the entire iterative process is used as the output.

[0088] The four exchange strategies are random exchange of two rows, multiple exchange of two rows, random exchange of multiple rows (random block exchange), and random region reverse order exchange (random region flip). Randomly adopting one of the four strategies can reduce the dominant influence of a few highly expressed genes in the entire iteration process. Specifically, the gene set distribution is in matrix form, which is The matrix, row Represents genes, columns Represents rounds, where each gene is marked as 1 in rounds ≥2 (generally, each row has ≥2 1s), indicating that the gene was imaged in the round marked as 1. Each gene has a unique distribution of rounds marked as 1. As the number of iterations increases, the number of 1s marked for different genes in different rounds changes, which is the gene set distribution. The four swapping strategies described above all target the gene set distribution during iteration. The variance is calculated based on the gene set distribution at the scRNA cell type level expression matrix, and the variance between rounds is calculated.

[0089] Furthermore, during gene codebook optimization, constraints were placed on the generation of initial values ​​to avoid potential contamination during the experiment. Specifically, the codes for the same channel from previous and subsequent rounds were removed; the number of genes in each round was calculated, and gene sets with fewer than the preset number of genes were removed, ensuring that the number of genes in each round remained within the set range. This constrained generation of initial values ​​was achieved.

[0090] Specifically, the encoding scheme for the same channel between previous and next rounds (for example, channel 561 in the first round / channel 561 in the second round) is removed. This means that this encoding scheme will not appear between gene set distributions during the iteration process. This encoding scheme can be affected by incomplete probe cleaning between the first and second rounds of the experiment, resulting in a large number of contaminating false signals. Furthermore, the gene counts in each round are iterated, and genes below a preset gene-round threshold are removed to ensure that the gene counts in each round are within the set range. In this case, the experimenter may misconfigure the experimental system due to the need for more genes to be expressed in other rounds. By generating constraints on the initial values ​​through these two removal processes, the gene encoding optimization of this embodiment also supports adjusting the cutoff value in the scRNA data based on real experimental results and limiting the maximum number of light spots captured according to different technical theories. Specifically, the expression of specific genes in the scRNA can be modified based on experimental results, and the total gene expression value in each round during the simulated annealing iteration is limited.

[0091] Therefore, this embodiment performs genetic codebook optimization based on the simulated annealing algorithm. This optimization is compatible with multiple coding strategies and optimizes the initial value generation and iteration process of the codebook according to potential contamination phenomena in the experiment. At the same time, it supports feedback adjustment based on actual experimental results to achieve the coding strategy with the least optical congestion in real situations.

[0092] Example 1

[0093] (c1) Download the original fastq data from the GEO database and preprocess it to obtain the isoform and gene expression matrix, taking into account the single-cell expression matrix, and construct the bulkRNA and scRNA retrieval database based on it.

[0094] (c2) Designing targeted region probes: Enter the desired tissue name (mouse brain), the desired gene set, and the corresponding isoform name. Search the database and generate a set of potential probe sequences of 60 nt at intervals of Xnt. X represents the set value, which can be set by the user as needed. Based on a Tm value of 40-80 and a GC content of 30%-75%, remove 6-mers (e.g., AAAAAA) and sequences with potential loops and dimers. Perform nonspecific detection using blastn (nucleic acid sequence BLAST). Set the expected (evalue) value to 1000, specify the alignment task type (task value) as short sequence alignment mode (blastn-short), and define the result file format (outfmt value) as 6. Remove sequences with nonspecific binding of more than 45 nt to obtain candidate probe sequences.

[0095] The obtained candidate probe sequence was divided into two segments and then BLAST alignment was performed (non-specific binding length of the first 30 nt and the last 30 nt), and the sequence with non-specific binding of more than 25 nt was removed to obtain the probe sequence after non-specific removal.

[0096] Obtain protein-binding RNA clip-seq data and filter them based on enrichment intensity (e.g., peak enrichment intensity greater than 10), retaining potential binding sites with high confidence, and obtain filtered clip-seq data; set the minimum overlap region ratio threshold (f value) to 0.2, and use bedtools intersect (overlapping region) to remove sequences with more than 20% overlap in the filtered clip-seq data to obtain targeted probe sequences for all isoforms.

[0097] The most highly expressed isoforms are then identified from the collected mouse brain bulk RNA data. Targeted probe sequences from all isoforms are then screened based on the kilobases per million (FPKM) expression of the bulk RNA gene. For FPKM expression <10, N+6 targeted probes are selected; for FPKM expression between 10-100, N targeted probes are selected; and for FPKM expression >100, N-6 targeted probes are selected. N is a specific value that can be selected by the user based on their needs. Design is prioritized based on segments shared by the most highly expressed isoform within the same gene classification. If the number of shared segments is insufficient, segments unique to the most highly expressed isoform are used to supplement the design until the desired number is reached. If the number is insufficient or design is unavailable, the corresponding result is returned and parameters need to be readjusted for that gene. Through the above process, targeted probe sequence sets corresponding to multiple genes are ultimately designed.

[0098] (c3) Design of non-target region probes: 10 million random sequences were generated using random ATCG 60nt. Non-specific detection was performed using blastn. The expected value (evalue) was set to 1000, the alignment task type (task value) was specified as short sequence alignment mode (blastn-short), and the result file format (outfmt value) was defined as 6. Sequences that specifically bind to known RNA sequences were removed. Sequences were separated into 30-nt left and 30-nt right edge segments, with a Hamming distance of >15 on both sides. These selected non-target region probes were retained as non-target region probes.

[0099] (c4) The non-target region probe and the target region probe are separated according to Figure 4The experimental probes are assembled into rolling circle probes, fluorescent probes, and primer probes in the form of a fluorescent probe, and submitted to a biotechnology company for synthesis. The non-binding segment 1 of the same gene has the same sequence; the non-binding segment 2 has the same sequence. The non-binding segment is the non-targeted region. The rolling circle probe is a targeted region probe plus two non-targeted region probes. The fluorescent probe is a non-targeted region probe plus a fluorescent group. The primer probe is the reverse complement of a non-targeted region probe.

[0100] (c5) Mouse brain tissue was searched in the retrieval database of collected and processed bulk RNA and scRNA. The constructed gene set was set with the expression level of bulk RNA in kilobases per million sequencing reads (FPKM) within 10-1000 and the expression level of average raw count at the scRNA cell type level within 200. Based on the output results, the gene set was divided into x genes for the coding round and other genes for non-coding imaging. The genes for non-coding imaging do not need to be coded.

[0101] Then, based on the paired mouse brain scRNA data, the appropriate encoding round and the x encoded genes were input, and the simulated annealing algorithm proposed in this example was used to optimize the gene set distribution to minimize the variance of expression levels between different rounds.

[0102] For example, we randomly select 10 sets of 200 genes from the mouse brain and test them 10 times, and get the following: Figure 3 The results shown show that compared with the random strategy, this embodiment can effectively reduce the inter-round variance.

[0103] For example, set the encoding round = 3, randomly adopt four exchange strategies during the iteration process, and at the same time, the encoding initial value generation removes the possible contamination coding in the experiment, and finally obtains the optimal example encoding scheme as follows Figure 5 , Figure 5 The example shows the result files of 14 gene encodings in three rounds. You can experiment in sequence according to the case marked as 1 here.

[0104] (c6) Gene space display: According to the experimental sequence in the optimized gene coding book, the fluorescent probes of the corresponding genes (assembled in step c4) are added to different channels in different rounds to make them light up in specific round channels. The experimental process is to cut a tissue slice with a thickness of 10μm from the mouse brain and attach it to a glass slide. After fixation, all the rolling circle probes are added to the buffer solution for incubation, and then the gaps on the rolling circle probes are connected with DNA ligase. Then, primer probes, DNA polymerase, substrates, etc. are added for rolling circle amplification. After fixation again, the fluorescent probes of the genes corresponding to the gene coding book are finally added to make the genes in the current round glow, and the original image data is obtained by imaging with a fluorescence microscope. The position of the bright spot of each round is obtained by point recognition of the original image data through the LOG operator, and the relative intensity value vector of the original image data of multiple round channels at the bright spot position is recorded. The gene decoding is performed by calculating the Euclidean distance between this vector and the coding vectors of different genes in the gene coding book. Finally, the decoded gene display is shown as follows: Figure 6 , the Allen reference brain map is a single gene ISH map, in which the dots are actual gene signals; the right side is the gene point distribution obtained by decoding after the actual experiment with the probe, marked in green, and the darker the color, the higher the density. Through actual experiments, the Penk gene is enriched in the striatum area of ​​the mouse brain and shows a two-layer distribution in the mouse cerebral cortex; the Sst gene is scattered in the mouse cerebral cortex. These results show that the gene distribution obtained after designing probes and coding tables for targeted genes is similar to the expression pattern of the Allen reference brain map, illustrating the technical availability of this embodiment.

[0105] Example 2

[0106] To lower the coding barrier for users and improve usability, this example also provides a public web-based database system to assist with user experience. Development details are as follows: The front-end and back-end framework is developed based on Django (an open-source Python web framework). Data is transferred via Ajax using JSON (a lightweight data exchange format) for front-end and back-end interaction. Server-side configuration and deployment are implemented using Nginx (an HTTP and reverse proxy web server) and uwsgi (an interface). The overall framework features a left-side navigation bar, with three main functions: Check, Codebook, and Merge, each with corresponding URL links. The right side contains the main function area, where users can enter parameters. Additional parameters are accessed through the Advanced button. Clicking this button allows parameters to be passed to the server. The server receives the input parameters, processes the logic code, and returns the data to the front-end for rendering or downloading. Check, Codebook, Merge, Advanced, and Button are all buttons, and the links correspond to the displayed pages.

[0107] (d1) Check function: Access the XXX website via the web interface. Users select the appropriate tissue data type through the Species & Tissue & Disease option, determine the required gene set through the GeneSet_Check option, enter a reasonable cutoff value (threshold for scRNA and bulkRNA screening) through the Advanced button, and then click the Check button to wait for the output results. The output results will feedback the genes that can be used for coding under the custom cutoff value, as well as the gene sets with higher expression levels recommended for non-coding detection. The higher expression levels will be displayed in red, and genes that cannot be designed for other reasons (target region cannot be designed, the kilobase per million sequencing (FPKM) expression level of bulkRNA genes is too low, the gene name is incorrect, and genes that are not in the scRNA and bulkRNA data). Finally, the user downloads the output results according to their needs and adjusts the classification of the gene sets used to design coding and non-coding genes.

[0108] (d2) Codebook Function: Access the XXX website via the web interface. Select the appropriate tissue data type using the Species & Tissue & Disease option, set the appropriate number of Encode Rounds, enter the encoding gene set compiled in the previous step into the GeneSet_Encode input box, and enter the non-coding gene set compiled in the previous step into the GeneSet_Single input box. Click the Advanced button to set the parameters for the simulated annealing algorithm. Click the Codebook button and wait for the output. The output includes a display of the simulated annealing process and the optimal encoding result. The optimal codebook is saved in the Codebook function link.

[0109] (d3) Merge function: Access the XXX website via the web interface. Select the appropriate tissue data type using the Species & Tissue & Disease tab. Then, use the Check function to export the Max Isoform name for each gene's highest expression level. Finally, click the Merge button and wait for the output. The output will automatically download as a zip file. The results can be directly submitted to the sequencing company for probe ordering and subsequent experiments.

[0110] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A probe design method for multi-round spatial transcriptome imaging, characterized in that: include: Download the original fastq data from the GEO database and preprocess it to obtain isoform and gene expression matrices, taking into account the single-cell expression matrix, and construct a search database for bulkRNA and scRNA; Designing targeted region probes: traverse and preprocess the target regions of all isoforms of the same gene from known RNA sequences. Divide the resulting candidate probe sequences into two segments for additional nonspecific detection to obtain target probe sequences for all isoforms. Target region probes are then screened based on all expression matrices in the search database. Design of non-targeted region probes: Generate multiple random sequences through random ATCG and pre-process them to design non-targeted region probes; Gene codebook optimization: The expression of bulk RNA in kilobases per million sequencing reads and the average raw counts of scRNA at the cell type level for a given gene set in the tissues of a given species are obtained from the search database. Genes with expression levels exceeding the upper expression threshold or below the lower expression threshold are not allowed to be included in the coding round. The gene sets obtained are used for codebook design and subsequent coding imaging experiments. The number of coding rounds and the number of coding imaging genes are set, and the distribution of the selected gene sets between round channels is iteratively optimized. In each iteration, one of four exchange strategies is randomly selected to process the distribution of the selected gene sets. If the variance of the gene set distribution after the exchange strategy is less than the variance before the exchange, the next round will be entered, where the variance is the total number of different gene combinations highlighted in each round under the gene set distribution according to the expression matrix at the scRNA cell type level, and then the average variance of the total number between each round is calculated; if the variance after the exchange strategy is greater than the variance before the exchange, a randomly generated value is used to determine whether it is less than ,in, is the variance difference before and after the exchange strategy, is the annealing temperature, which decays with the number of iterations; if it is less than If yes, it will enter the next round, otherwise it will reselect the exchange strategy until the iteration requirements are met; Gene space display: Multiple rounds of spatial transcriptome imaging are performed using targeted region probes, non-targeted region probes, and gene code books to obtain gene space display decoded according to the gene code book.

2. The probe design method for multi-round spatial transcriptome imaging according to claim 1, characterized in that In designing the targeted region probe, the targeted segments of all isoforms of the same gene are obtained from the known RNA sequence and preprocessed, specifically: Based on the input species name, required gene set and corresponding isoform name, the sequence is traversed from the fasta file that stores all known RNA sequences according to the required design length to generate a potential probe sequence set; By setting the Tm value, GC content, removal of concatemers, RNA secondary structure detection, and overall BLAST nonspecificity check, the potential probe sequence set was screened to obtain nonspecific alignment results and candidate probe sequences.

3. The probe design method for multi-round spatial transcriptome imaging according to claim 1, characterized in that In designing targeted region probes, the candidate probe sequences were divided into two segments for additional nonspecific detection to obtain targeted probe sequences for all isoforms. Targeted region probes were then obtained by screening all expression matrices in the search database. Specifically: The candidate probe sequence is divided into two segments and then subjected to separate BLAST nonspecific detection. The nonspecific removal of the two segment alignment results is performed according to the set unilateral nonspecific binding threshold; Obtain protein-binding RNA clip-seq data from known public databases and filter them based on enrichment intensity, retaining potential binding sites with confidence exceeding a set threshold to obtain filtered clip-seq data; Set the minimum overlap ratio threshold, retain the probe sequences after nonspecific removal and the sequences whose overlap ratio with the filtered clip-seq data does not exceed the set overlap ratio threshold, and obtain the targeted probe sequences of all isoforms; According to the expression matrix of bulk RNA of the set species tissue in the retrieval database, the target segment shared by the highest expressed isoform and other isoforms in the same gene classification is preferentially selected as the target region probe, and a set number of target region probes are selected according to the expression levels of bulk RNA and scRNA in the retrieval database.

4. The probe design method for multi-round spatial transcriptome imaging according to claim 1, characterized in that Design of non-targeted region probes: Generate multiple random sequences through random ATCG and pre-process them to design non-targeted region probes, specifically: Generate a sufficient amount of random sequences of fixed length by random ATCG and retain the sequences that do not specifically bind to known RNA sequences by BLAST as potential sequence sets; Calculate the Hamming distance of all lengths between any two sequences in the potential sequence set, and retain the sequences whose Hamming distance exceeds the set threshold. The first 1 / 3 segment and the last 1 / 3 segment of each retained sequence were marked as edge segments, and the Hamming distance between each edge segment was calculated. Sequences with Hamming distance exceeding the set threshold were selected and retained, and non-target region probes were designed based on them.

5. The probe design method for multi-round spatial transcriptome imaging according to claim 1, characterized in that One targeted region probe plus two non-targeted region probes constitute a rolling circle probe, one non-targeted region probe plus a fluorescent group constitutes a fluorescent probe, and the reverse complement of one non-targeted region probe constitutes a primer probe. The rolling circle probe, the fluorescent probe and the primer probe are assembled to obtain an experimental probe.

6. The probe design method for multi-round spatial transcriptome imaging according to claim 1, characterized in that The four exchange strategies are random exchange of two rows, multiple exchange of two rows, random exchange of multiple rows, and random reverse exchange of the entire area.

7. The probe design method for multi-round spatial transcriptome imaging according to claim 1, characterized in that: In the optimization of the genetic codebook, the generation of initial values ​​is constrained, specifically: The encoding of the previous and next rounds of the same channel has been removed; The number of genes in each round is calculated, and the gene set distribution with less than the preset number of genes in each round is removed, so that the number of genes in each round is within the set interval, thereby generating constraints on the initial value.

8. A probe design system for multi-round spatial transcriptome imaging, characterized in that It includes database construction module, targeted region probe design module, non-targeted region probe design module, gene codebook optimization module and gene display module; The database construction module is used to download raw fastq data from the GEO database and preprocess it to obtain isoform and gene expression matrices, taking into account the single-cell expression matrix, and then construct a retrieval database for bulkRNA and scRNA; The targeted region probe design module is used to traverse and pre-process the target segments of all isoforms of the same gene from known RNA sequences, divide the obtained candidate probe sequences into two segments for additional non-specific detection, and obtain the targeted probe sequences of all isoforms. The targeted region probes are then screened based on all expression matrices in the search database. The non-targeted region probe design module is used to generate multiple random sequences through random ATCG and pre-process them to design non-targeted region probes; The gene codebook optimization module is used to obtain the expression of bulk RNA kilobases per million sequencing reads and the average raw counts of scRNA at the cell type level for a given gene set in a set of species tissues from the search database. Genes with expression levels exceeding the upper expression threshold or below the lower expression threshold are not allowed to be included in the coding round. The gene sets obtained are used for codebook design and subsequent coding imaging experiments. The number of coding rounds and the number of coding imaging genes are set, and the distribution of the selected gene sets between round channels is iteratively optimized. In each iteration, one of four exchange strategies is randomly selected to process the distribution of the selected gene sets. If the variance of the gene set distribution after the exchange strategy is less than the variance before the exchange, the next round will be entered, where the variance is the total number of different gene combinations highlighted in each round under the gene set distribution according to the expression matrix at the scRNA cell type level, and then the average variance of the total number between each round is calculated; if the variance after the exchange strategy is greater than the variance before the exchange, a randomly generated value is used to determine whether it is less than ,in, is the variance difference before and after the exchange strategy, is the annealing temperature, which decays with the number of iterations; if it is less than If yes, it will enter the next round, otherwise it will reselect the exchange strategy until the iteration requirements are met; The gene display module is used to perform multiple rounds of spatial transcriptome imaging using targeted region probes, non-targeted region probes and gene code books to obtain gene spatial display after decoding according to the gene code books.

9. The probe design system for multi-round spatial transcriptome imaging according to claim 8, characterized in that: The non-targeted region probe design module is specifically used for: Generate a sufficient amount of random sequences of fixed length by random ATCG and retain the sequences that do not specifically bind to known RNA sequences by BLAST as potential sequence sets; Calculate the Hamming distance of all lengths between any two sequences in the potential sequence set, and retain the sequences whose Hamming distance exceeds the set threshold. The first 1 / 3 segment and the last 1 / 3 segment of each retained sequence were marked as edge segments, and the Hamming distance between each edge segment was calculated. Sequences with Hamming distance exceeding the set threshold were selected and retained, and non-target region probes were designed based on them.

Citation Information

Patent Citations

  • Primer probe sequence combination design method and device

    CN116130000A

  • Probe design of high-throughput single cell space transcriptome based on multi-round imaging coding

    CN119724347A