Probe design method and system for multi-round space transcriptional group imaging

By designing targeted and non-targeted regional probes and optimizing gene coding books using simulated annealing algorithm, the problems of probe design limitations and insufficient signal quality in the prior art are solved, and more efficient probe capture and signal quality improvement are achieved.

CN120220832AActive Publication Date: 2025-06-27ARTIFICIAL INTELLIGENCE RES INST OF HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ARTIFICIAL INTELLIGENCE LAB)
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510695492.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Existing spatial transcriptome technology based on imaging has limitations in probe design, such as incomplete disclosure of probe sequences, weak generalization of coding book design, and possible contamination problems in experiments, resulting in insufficient capture efficiency and signal quality.

Method used

A probe design method for multi-round spatial transcriptional composition imaging is proposed, including downloading and pre-processing data from GEO databases, designing targeted and non-targeted regional probes, and optimizing gene coding books through simulated annealing algorithms to improve probe capture efficiency and signal quality.

Benefits of technology

Through the system optimization of probe design, the overall capture efficiency of the probe is improved, the true positive signal is enhanced, the false positive signal is reduced, and a simulated annealing algorithm compatible with multiple encoding strategies is provided, which optimizes the generation and iteration process of the codebook.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220832A_ABST
    Figure CN120220832A_ABST
Patent Text Reader

Abstract

The invention discloses a probe design method and system for multi-round space transcriptional group imaging, and relates to the technical field of bioinformatics, and the probe design method comprises the following steps: constructing a retrieval database; the method comprises the following steps: traversing a known RNA (Ribonucleic Acid) sequence to obtain target sections of all isoforms of the same gene, preprocessing the target sections, dividing an obtained candidate probe sequence into two sections, carrying out additional non-specific detection to obtain target probe sequences of all isoforms, and screening according to all expression matrixes in a retrieval database to obtain target region probes; the method comprises the following steps: generating a plurality of random sequences through random ATCG, preprocessing, and designing a non-targeted region probe according to the random sequences; obtaining a single cell expression matrix from the retrieval database, and optimizing the cell type level expression quantity variance of a given gene set among different imaging rounds to be minimum based on a simulated annealing algorithm to obtain a gene codebook; according to the probe design method and system, the overall capture efficiency of the probe is improved to a certain degree, and meanwhile various coding strategies are compatible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and particularly relates to a probe design method and system for multi-round spatial transcriptome imaging. Background Art

[0002] Single-cell transcriptomics technology can capture gene expression at the single-cell level. However, dissociating tissues into single-cell suspensions causes loss of spatial and morphological information, making it difficult to study features related to the spatial scale. Spatial transcriptomics technology can well obtain gene expression information and spatial location information of tissues, providing a solution to this problem. Spatial transcriptomics technology is divided into sequencing-based technology and imaging-based technology. Although sequencing-based spatial transcriptomics technology can detect genes at the whole-transcriptome scale, it is difficult to achieve single-cell precision resolution; imaging-based spatial transcriptomics technology can detect the spatial distribution of the transcriptome with high precision and high resolution, and can better depict the spatial distribution of individual cells in tissues.

[0003] However, the gene capture rate of imaging-based spatial transcriptomics is extremely related to the designed probe sequences. Specifically, the probe design structures between different technologies, the specificity of the targeted regions, and the expression levels of different genes in different tissues will all have an impact. At the same time, the specific probe sequences of current mainstream technologies such as multiplexed error-robust fluorescence in situ hybridization (MERFISH) and spatially resolved transcript amplicon readout mapping (STARmap) are not fully disclosed, and there are still some limitations and room for improvement in their designs, such as the influence of potential protein binding sites. Therefore, a brand-new, convenient, practical, and highly compatible probe design system is needed to provide an optimized solution.

[0004] To improve gene throughput in imaging-based spatial transcriptomics, multiple combinations of different genes are generally added simultaneously in multiple rounds and then decoded to obtain the correct positions. Among them, the design scheme of the codebook is generally considered to be optimal when the genes with relatively high and low expression levels are removed and the light intensity is as average as possible, that is, the probability of optical crowding phenomenon is minimized and the gene capacity is as large as possible. However, the design methods of codebooks between different technologies are not fully disclosed and have weak generalization ability. At the same time, the potential contamination problems remaining in the experiments can also be optimized in the design of the codebook. Therefore, a highly compatible and practical codebook optimization algorithm is needed. Summary of the Invention

[0005] Based on the technical problems existing in the background art, the present invention proposes a probe design method and system for multi-round spatial transcriptome imaging, which can improve the overall capture efficiency of the probes and be compatible with multiple coding strategies.

[0006] A probe design method for multi-round spatial transcriptome imaging proposed by the present invention includes: Download the original fastq data from the GEO database and preprocess it to obtain the expression matrices of isoforms and genes, taking into account the single-cell expression matrix, and construct the retrieval databases of bulkRNA and scRNA based on this; Design targeted region probes: Traverse and obtain the targeted segments of all isoforms of the same gene from the known RNA sequences and preprocess them. Divide the obtained candidate probe sequences into two segments for additional non-specific detection to obtain the targeted probe sequences of all isoforms, and then screen the targeted region probes according to all the expression matrices in the retrieval database; Design non-targeted region probes: Generate multiple random sequences by random ATCG and preprocess them, and design non-targeted region probes based on this; Optimization of gene encoding book: Obtain the single-cell expression matrix from the retrieval database, and optimize the variance of the expression levels of the given gene set at the cell type level between different imaging rounds based on the simulated annealing algorithm to obtain the gene encoding book; Gene spatial display: Use the targeted region probes, non-targeted region probes and gene encoding book for multi-round spatial transcriptome imaging to obtain the gene spatial display decoded according to the gene encoding book.

[0007] Further, in the design of targeted region probes, traversing and obtaining the targeted segments of all isoforms of the same gene from the known RNA sequences and preprocessing them specifically includes: Based on the input species name, required gene set and corresponding isoform name, traverse the sequences from the fasta file storing all RNA sequences according to the required designed length to generate a set of potential probe sequences; Screen the set of potential probe sequences through setting Tm value, GC content, removing polycistrons, RNA secondary structure detection and overall BLAST non-specific inspection to obtain the non-specific alignment results and obtain the candidate probe sequences.

[0008] Further, in the design of targeted region probes, divide the obtained candidate probe sequences into two segments for additional non-specific detection to obtain the targeted probe sequences of all isoforms, and then screen the targeted region probes according to all the expression matrices in the retrieval database, specifically including: Divide the candidate probe sequences into two segments and perform BLAST non-specific detection separately. Remove non-specificity from the alignment results of the two segments respectively according to the set unilateral non-specific binding threshold; Retrieve the clip-seq data of proteins binding to RNA from a known public database and screen according to the enrichment intensity. Retain the potential binding sites with a confidence level exceeding the set threshold to obtain the screened clip-seq data. Set the threshold for the minimum overlapping region ratio. Retain the probe sequences whose overlapping ratio with the screened clip-seq data does not exceed the set overlapping ratio threshold after non-specific removal to obtain the targeting probe sequences for all isoforms. According to the expression matrix of bulkRNA in the set species tissue in the retrieval database, preferentially select the targeting segments shared by the highest-expressed isoform and other isoforms under the same gene classification as the targeting region probes, and select a set number of targeting region probes according to the expression levels of bulkRNA and scRNA in the retrieval database.

[0009] Further, design non-targeting region probes: Generate multiple random sequences by random ATCG and preprocess them to design non-targeting region probes. Specifically: Generate a sufficient number of random sequences of a fixed length by random ATCG and retain the sequences that do not have specific binding to the known RNA sequences through BLAST as the potential sequence set. Calculate the Hamming distance of all lengths between pairwise sequences in the potential sequence set, and retain the sequences with a Hamming distance exceeding the set threshold. Mark the first 1 / 3 segment and the last 1 / 3 segment of each retained sequence as the edge segments, calculate the Hamming distance between pairwise edge segments, and retain the sequences with a Hamming distance exceeding the set threshold, and design non-targeting region probes accordingly.

[0010] Further, 1 targeting region probe plus 2 non-targeting region probes form a rolling circle probe, 1 non-targeting region probe plus a fluorophore forms a fluorescent probe, and the reverse complement of 1 non-targeting region probe forms a primer probe. Assemble the rolling circle probe, the fluorescent probe, and the primer probe to obtain the experimental probe.

[0011] Further, in the optimization of the gene codebook, specifically as follows: Obtain the expression levels of reads per kilobase per million mapped reads of bulkRNA and the average raw count expression levels at the scRNA cell type level for a given gene set in the set species tissue from the retrieval database; Genes with expression levels exceeding the upper expression threshold and genes with expression levels below the lower expression threshold are not allowed to be added to the coding rounds. Screen the gene set for codebook design and subsequent coding imaging experiments. Set the number of coding rounds and the number of coding imaging genes, and iteratively optimize the distribution of the screened gene set among the round channels. In each iteration, one of the four exchange strategies is randomly selected to process the filtered gene set distribution; If the variance of the gene set distribution after adopting the exchange strategy is less than the variance before the exchange, it enters the next round. Here, the variance is the total number calculated based on the expression matrix at the scRNA cell type level for each different gene combination highlighted in each round of the gene set distribution, and then the average variance of this total number between each round is calculated; If the variance after adopting the exchange strategy is greater than the variance before the exchange, it is judged whether it is less than by randomly generating a value, where, is the difference in variance before and after the exchange strategy, is the annealing temperature, which decays with the number of iterations; If it is less than it enters the next round, otherwise it reselects the exchange strategy until the iteration requirements are met.

[0012] Further, the four exchange strategies are randomly swapping two rows, swapping two rows multiple times, randomly swapping multiple entire rows, and randomly reversing the order of an entire region.

[0013] Further, in the optimization of the gene coding book, the generation of the initial value is constrained. Specifically: The coding of the same channel in the previous and next rounds is removed; The number of genes in each round is calculated, and the gene set distributions with less than the preset number of genes in each round are removed, so that the number of genes in each round is within the set interval, thereby constraining the generation of the initial value.

[0014] A probe design system for multi-round spatial transcriptome imaging includes a database construction module, a targeted region probe design module, a non-targeted region probe design module, a gene coding book optimization module, and a gene display module; The database construction module is used to download the original fastq data from the GEO database and preprocess it to obtain the expression matrices of isoforms and genes, taking into account the single-cell expression matrix, and accordingly construct the retrieval databases of bulkRNA and scRNA; The targeted region probe design module is used to traverse and obtain the targeted segments of all isoforms of the same gene from the known RNA sequences and preprocess them. The obtained candidate probe sequences are divided into two segments for additional non-specific detection to obtain the targeted probe sequences of all isoforms, and then the targeted region probes are screened according to all the expression matrices in the retrieval database; The non-targeted region probe design module is used to generate multiple random sequences by randomly generating ATCG and preprocess them, and accordingly design non-targeted region probes; The gene encoding book optimization module is used to obtain the single-cell expression matrix from the retrieval database, and optimize the variance of the expression levels of a given gene set at the cell type level between different imaging rounds based on the simulated annealing algorithm to obtain the gene encoding book; The gene display module is used to perform multi-round spatial transcriptome imaging using the targeted region probes, non-targeted region probes, and the gene encoding book to obtain the gene spatial display decoded according to the gene encoding book.

[0015] Furthermore, the non-targeted region probe design module is specifically used for: Generating a sufficient amount of random sequences of a fixed length by randomly generating ATCG and retaining the sequences that do not specifically bind to the known RNA sequences through BLAST as the potential sequence set; Calculating the Hamming distance of all lengths between pairwise sequences in the potential sequence set, and retaining the sequences with the Hamming distance exceeding the set threshold; Marking the first 1 / 3 section and the last 1 / 3 section of each retained sequence as the edge sections, calculating the Hamming distance between the pairwise edge sections, and retaining the sequences with the Hamming distance exceeding the set threshold, and designing the non-targeted region probes accordingly.

[0016] The advantages of a probe design method and system for multi-round spatial transcriptome imaging provided by the present invention are as follows: it can systematically optimize the specificity of the targeted region and the non-specificity of the non-binding region, improve the overall capture efficiency of the probes, that is, enhance the true positive signal and reduce the false positive signal, and provide a simulated annealing algorithm optimization algorithm for the coding of multi-round high-throughput genes. This algorithm is compatible with multiple coding strategies and optimizes the generation and iteration process of the initial value of the coding book according to the potential contamination phenomenon in the experiment, and at the same time supports feedback adjustment with the actual experimental results to achieve the coding strategy with the least optical crowding in the real situation. Brief Description of the Drawings

[0017] Figure 1 It is a flow schematic diagram of the present invention; Figure 2 It is a schematic diagram of the example retrieval situation of scRNA / bulkRNA of the mouse kidney dataset; Figure 3 It is a schematic diagram of the variance comparison of the coding book design under the mouse brain dataset; Figure 4 It is a schematic diagram of the probe assembly structure; Figure 5 It is a schematic diagram of the output result of the coding scheme; Figure 6 It is a schematic diagram of the gene example display. Detailed Embodiments

[0018] Next, the technical solution of the present invention will be described in detail through specific embodiments. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0019] As Figures 1 to 6 shown, a probe design method for multi-round spatial transcriptome imaging proposed by the present invention includes: Step 1: Download the original fastq data from the GEO database and preprocess it to obtain the expression matrices of isoforms (homologous isomers) and genes, taking into account the single-cell expression matrix, and construct a retrieval database for bulkRNA and scRNA based on this. Step 2: Design targeted region probes: Traverse and obtain the targeted segments of all isoforms of the same gene from the known RNA sequences and preprocess them. Divide the obtained candidate probe sequences into two segments for additional non-specific detection to obtain the targeted probe sequences of all isoforms, and then screen the targeted region probes according to all the expression matrices in the retrieval database. Step 3: Design non-targeted region probes: Generate multiple random sequences by randomly combining ATCG and preprocess them to design non-targeted region probes, where ATCG represents adenine (A), thymine (T), guanine (G), and cytosine (C), which are the core elements constituting the DNA double helix structure. Step 4: Gene encoding book optimization: Obtain the single-cell expression matrix from the retrieval database, and optimize the variance of the expression levels of a given gene set at the cell type level between different imaging rounds based on the simulated annealing algorithm to obtain the gene encoding book. Step 5: Gene spatial display: Use the targeted region probes, non-targeted region probes, and gene encoding book for multi-round spatial transcriptome imaging to obtain the gene spatial display decoded according to the gene encoding book.

[0020] This embodiment provides a multi-round imaging spatial transcriptome probe design method. This probe design method can systematically optimize the specificity of the targeted region and the non-specificity of the non-targeted region, which can improve the overall capture efficiency of the probe, that is, enhance the true positive signal and reduce the false positive signal. And for the coding of multi-round high-throughput genes, this design provides a simulated annealing algorithm optimization algorithm. This algorithm is compatible with multiple coding strategies and optimizes the generation and iteration process of the initial value of the coding book according to the potential contamination phenomenon in the experiment. At the same time, it supports feedback adjustment with the actual experimental results to achieve the coding strategy with the least optical crowding in the real situation.

[0021] In this embodiment, the data collected and processed, the design of the target region probes and non-target region probes, and the gene encoding book optimization method are all publicly disclosed in the form of a Web database to reduce the code threshold for users and improve the actual usability.

[0022] In this embodiment, the probe design method mainly includes three functional modules: 1) construction of bulkRNA and scRNA databases for multiple tissues of human and mouse RNA isoforms and genes; 2) optimization and screening strategies for the target regions and non-target regions of isoforms in the probe structure; 3) multi-round high-throughput gene encoding book design optimization algorithm based on simulated annealing.

[0023] In one embodiment, the retrieval of bulkRNA and scRNA databases for human and mouse isoforms and genes, that is, step one is specifically as follows: According to the information in the published literature, 37 human tissues and 19 mouse tissues, a total of 56 kinds of bulkRNA and scRNA data were collected. Among them, the scRNA data comes from the processed downstream expression matrix of the public database, and the bulkRNA data is obtained by downloading the original fastq data from the GEO database to which different literatures belong through the sra tool and then systematically processing it through software tools such as trim_galore, tophat, and cufflinks to obtain the expression matrix of isoforms and genes, thereby constructing the retrieval databases of bulkRNA and scRNA. Finally, the data is sorted and relevant screening codes are written and deployed on the Web. Among them, the sra tool is a software toolset for processing high-throughput sequencing data, fastq is a text format for storing biological sequences and their quality information, and it is the standard storage format for high-throughput sequencing data. trim_galore is for data preprocessing and quality control, tophat is for sequence alignment and splice site identification, cufflinks is for transcript assembly and expression quantification, scRNA is single-cell RNA, and bulkRNA is sample transcriptome sequencing.

[0024] Researchers can obtain the expression of a specific gene set in a specific tissue they are interested in based on the constructed retrieval databases of bulkRNA and scRNA, so as to provide a reference basis for subsequent probe design and adopt an appropriate number of target fragments and whether to code. For example, Figure 2 As shown, retrieve the gene expression in the mouse kidney dataset. Those with higher expression in scRNA or bulkRNA are marked in red. Subsequently, non-coding imaging and a moderately reduced number of target fragments are adopted, while those with lower expression can adopt an increased number of targets. It is not recommended to design those with extremely low expression, and the gene set with moderate expression can be used for subsequent coding to achieve fewer rounds of high-throughput imaging.

[0025] The targeted region of the transcriptome probe design in this embodiment is a specific isoform segment. The different expression levels of different isoform segments will greatly affect the capture efficiency. However, the existing publicly available bulkRNA and scRNA databases do not contain information to query such relevant information. This embodiment provides important data references for designing transcriptome probes for humans and mice in the art for the first time by collecting literature, downloading and processing the original bulkRNA and downstream scRNA data.

[0026] In one of the embodiments, for the optimization and screening strategy of the targeted region and non-targeted region of the isoform in the probe structure, various existing spatial transcriptome technologies will have different probe design structures due to different design logics, but they will all include two parts: a targeted region that specifically binds to the isoform segment and a non-targeted region that cannot specifically bind to known RNA sequences; this embodiment has carried out certain optimization and innovation in both aspects.

[0027] Specifically, the design of the targeted region probe in step two is as follows: (a1) Based on the input species name, required gene set, and corresponding isoform name, traverse the sequences in the fasta file storing all RNA sequences according to the required designed length to generate a set of potential probe sequences; (a2) Screen the set of potential probe sequences by setting the melting temperature, GC content, removing polycistrons, detecting RNA secondary structure (such as removing loop rings and dimers), and overall BLAST non-specificity check to obtain non-specific alignment results and obtain candidate probe sequences; Among them, the Tm value is the melting temperature, the GC content refers to the proportion of guanine (G) and cytosine (C) in the nucleic acid molecule to the total bases, and BLAST (Basic Local Alignment Search Tool) is a widely used bioinformatics tool for quickly aligning biological sequences (such as DNA, RNA, or proteins) in a database to find similar sequences and infer functional or evolutionary relationships.

[0028] (a3) Divide the candidate probe sequences into two segments and then perform BLAST non-specific detection separately, and non-specifically remove the two-segment alignment results according to the set unilateral non-specific binding threshold; For example, if the candidate probe sequence is 60nt, then divide the 60nt into the first 30nt and the last 30nt and perform non-specific alignment through BLAST respectively, and remove the probe sequences with non-specific binding greater than 20nt on one side. This will reduce the non-specific binding situation of continuous complete matching in the entire targeted segment and make the probe structure more stable.

[0029] (a4) Obtain the clip-seq data of proteins binding to RNA from a known public database and screen according to the enrichment intensity. Retain the potential binding sites with a confidence level exceeding the set threshold to obtain the screened clip-seq data; Among them, the clip-seq data is the data of different protein-binding RNA fragments of humans and mice collected from existing public databases. For example, set the peak enrichment intensity to be greater than 10 and screen the clip-seq data, where the peak specifically refers to the region with significant signal intensity in the genomic sequencing data.

[0030] (a5)Set the minimum overlapping region ratio threshold, and retain the sequences with an overlapping ratio not exceeding the overlapping ratio threshold between the probe sequences after non-specific removal and the screened clip-seq data to obtain the targeted probe sequences of all isoforms; Remove the sequences with an overlapping ratio exceeding the overlapping ratio threshold (for example, the overlapping ratio threshold is set to 0.2) by finding the overlapping regions (bedtools intersect), so as to remove the potential binding sites of proteins to RNA that may remain in the experiment.

[0031] (a6)According to the expression matrix of bulkRNA of the set species tissue in the retrieval database, preferentially select the targeted segments shared by the highest-expressed isoform and other isoforms under the same gene classification as the targeted region probes, and select a set number of targeted region probes according to the expression levels of bulkRNA and scRNA of this gene in the retrieval database. The selection principle is that if the expression level is too high, a smaller number of probes are selected. If the number of shared segments is not enough for design, the unique segments of the isoform with the highest expression level are used for supplementation.

[0032] Specifically, the design of the non-targeted region probes in step three is as follows: For the non-targeted region, in different technologies, the probe design will screen out a set of sequences that will not bind to specific known RNA sequence segments through BLAST. However, since there are actually some connections between the non-targeted region and some other known RNA sequence binding segments, and there is currently no good screening method for the non-specificity of the connections. In this embodiment, a screening method based on the Hamming distance edge overlap degree is first proposed for this defect. This method will significantly reduce the Hamming distance between the edge regions of the non-targeted region probes, thereby reducing potential signal crosstalk pollution.

[0033] Specifically: Generate a sufficient amount of random sequences of a fixed length (e.g., 60 nt) by randomly generating ATCG, and retain the sequences that do not specifically bind to known RNA sequences through BLAST as a potential sequence set. Calculate the Hamming distance of all lengths between pairwise sequences in the potential sequence set, and retain the sequences with a Hamming distance exceeding the set threshold; Mark the first 1 / 3 segment (the first 20 nt) and the last 1 / 3 segment (the last 20 nt) of each retained sequence as marginal segments, calculate the Hamming distance between pairwise marginal segments, and retain the sequences with a Hamming distance exceeding the set threshold, based on which non-target region probes are designed.

[0034] In one embodiment, the multi-round high-throughput gene encoding book design optimization algorithm based on simulated annealing, that is, step four specifically includes: This embodiment implements the gene encoding book design optimization algorithm based on the idea of simulated annealing. The optimization goal is the variance of the sum of the expression levels at the cell type level between different rounds of tissue-paired scRNA data. The smaller the variance, the more average the light intensity between each round. Since paired scRNA and spatial transcriptomics are known to have a high correlation, they can be used to minimize the possibility of optical crowding and improve imaging quality under a given gene set. The core idea of the algorithm proposed in this embodiment is the principle of simulated annealing, specifically as follows: (b1) Obtain the fragments per kilobase of exon model per million mapped reads (FPKM) of bulkRNA in the set species tissue of the given gene set and the average raw count expression level at the scRNA cell type level from the retrieval database; Genes with an expression level exceeding the upper expression threshold and genes with an expression level below the lower expression threshold are not allowed to be added to the encoding round. The gene set is screened for encoding book design and subsequent encoding imaging experiments; (b2) Set the number of encoding rounds and the number of genes for encoding imaging, and iteratively optimize the distribution of the screened gene set among the round channels; (b3) In each iteration, randomly select one of the four exchange strategies to process the distribution of the screened gene set; If the variance after adopting the exchange strategy for the gene set distribution is less than the variance before the exchange, enter the next round, where the variance is the total number calculated according to the expression matrix at the scRNA cell type level for different gene combinations highlighted in each round under the gene set distribution, and then calculate the average variance of this total number between each round; If the variance after adopting the exchange strategy is greater than the variance before the exchange, judge whether it is less than , where, is the difference between the variances before and after, is the temperature of annealing, which decays with the number of iterations; If it is less than then enter the next round, otherwise reselect the exchange strategy until the iteration requirements are met.

[0035] That is, in the early stage of an iteration process, a certain probability is allowed to accept the solution with increased variance to avoid entering the local minimum rather than the global minimum, but as the number of iterations increases, the acceptance probability gradually decreases, and finally the optimal solution in the entire iteration process is used as the output.

[0036] The four exchange strategies are random exchange of two rows, multiple exchange of two rows, random exchange of multiple rows (random block exchange), and random region reverse order (random region flip). Randomly adopting one of the four strategies can reduce the dominant influence of a few highly expressed genes in the entire iteration process. Specifically, the gene set distribution is in matrix form, that is, The matrix, row Represents genes, columns Indicates rounds, where each gene is marked as 1 in rounds greater than or equal to 2 (generally speaking, there will be greater than or equal to 2 1s in each row), indicating that the actual imaging of the gene will be experimentally imaged in the round marked as 1, and each gene has a unique distribution of rounds marked as 1. As the iteration rounds increase, the situation of different genes marked as 1 in different rounds changes, that is, the gene set distribution. The above four exchange strategies are all for the gene set distribution in the iteration, and the variance is calculated based on the gene set distribution at the scRNA cell type level expression matrix to calculate the variance between rounds.

[0037] At the same time, in the optimization of the gene codebook, constraints are imposed on the generation of initial values ​​to avoid possible contamination in the experiment. Specifically, the codes of the previous and next rounds of the same channel are removed; the number of genes in each round is calculated, and the gene set distribution with less than the preset number of genes in each round is removed, so that the number of genes in each round is within the set range, thereby constraining the generation of initial values.

[0038] Specifically, the encoding of the previous and next rounds of the same channel (for example, the first round 561 channel / the second round 561 channel) is removed, that is, this encoding scheme will not appear between the gene set distributions during the iteration process. This encoding will be affected by the incomplete cleaning of the probes in the previous and next rounds of the experiment, resulting in a large number of contaminated error signals; and the number of genes in each round is traversed, and genes less than the preset gene round threshold are removed so that the number of genes in each round is in the set interval. In this case, the experimenter may cause the experimental system configuration to be inaccurate due to the need to express more genes in other rounds. Through the above two removal processes, the initial value is generated and constrained. The gene encoding optimization of this embodiment also supports adjusting the cutoff value in the scRNA data according to the actual experimental results, and limiting the maximum number of light spots captured according to different technical theories. Specifically, the expression of specific genes in scRNA can be modified according to the experimental results, and the sum of gene expression in each round during the simulated annealing iteration is limited.

[0039] Therefore, in this embodiment, gene encoding book optimization is performed based on the simulated annealing algorithm. This optimization is compatible with multiple encoding strategies and optimizes the generation of the initial value of the encoding book and the iterative process according to potential contamination phenomena in the experiment. At the same time, it supports feedback adjustment based on the actual experimental results to achieve the encoding strategy with the least optical crowding in the real situation.

[0040] Example 1 (c1) Download the original fastq data from the GEO database and preprocess it to obtain the expression matrices of isoforms and genes, taking into account the single-cell expression matrix, and construct the retrieval databases for bulkRNA and scRNA based on this.

[0041] (c2) Design targeted region probes: Input the tissue name - mouse brain, the gene set of requirements, and the corresponding isoform names with the set characteristics of demand. By retrieving in the retrieval database, a set of potential probe sequences of 60 nt is generated by traversing every X nt, where X is a set value and can be actually set by the user according to needs. After removing hexamers (e.g., AAAAAA) and sequences with potential loop rings and dimers according to the Tm value in the range of 40 - 80 and the GC content of 30% - 75%, non-specific detection is performed through blastn (nucleic acid sequence BLAST). Set the expected (evalue) value to 1000, specify the alignment task type (task value) as the short sequence alignment mode (blastn-short), define the format of the result file (outfmt value) as 6, and remove sequences with non-specific binding of more than 45 nt to obtain candidate probe sequences.

[0042] Divide the obtained candidate probe sequences into two segments and then perform BLAST alignment (the non-specific binding length of the first 30 nt and the last 30 nt), and remove sequences with non-specific binding of more than 25 nt in a single segment to obtain the probe sequences after non-specific removal.

[0043] Obtain the clip-seq data of protein-binding RNA and screen it according to the enrichment intensity (e.g., the peak enrichment intensity is greater than 10), retain the potential binding sites with high confidence to obtain the screened clip-seq data; set the minimum overlapping region ratio threshold (f value) to 0.2, and use bedtools intersect (overlapping region) to remove sequences with more than 20% overlap with the screened clip-seq data to obtain the targeted probe sequences of all isoforms.

[0044] Then, obtain the name of the isoform with the highest expression level from the collected bulk RNA data of mouse brains, and screen among the targeting probe sequences of all isoforms according to the fragments per kilobase of exon per million reads mapped (FPKM) of the bulk RNA gene. Screen N + 6 targeting probes with an FPKM expression level < 10, screen N targeting probes with an FPKM expression level between 10 - 100, and screen N - 6 targeting probes with an FPKM expression level > 100. N is a specific value, and the user can make an actual selection according to needs. Priority is given to designing using the fragments common to the highest-expressing isoform and other isoforms under the same gene classification. If the number of common segments is not enough for design, use the unique segments of the isoform with the highest expression level to supplement until the appropriate number. If the number is insufficient or design is not possible, return the corresponding result and re-adjust the parameters for this gene to design separately. Through the above process, a set of targeting probe sequences corresponding to multiple genes is finally designed.

[0045] (c3)Design non-target region probes: Generate 10 million random sequences by randomly generating 60 nt of ATCG, and perform non-specificity detection through blastn. Set the expected (evalue) value to 1000, specify the alignment task type (task value) as the short sequence alignment mode (blastn-short), define the format of the result file (outfmt value) as 6, and remove the sequences that specifically bind to the known RNA sequences. By judging that the Hamming distance between pairwise sequences should be > 30, and splitting them into left and right 30 nt edge segments and ensuring that the Hamming distance of both the left and right sides is > 15, retain these screened non-target region probes as non-target region probes.

[0046] (c4)Assemble the non-target region probes and the target region probes in the Figure 4 form into rolling circle probes, fluorescence probes, and primer probes, and submit them to a biological company for synthesis to obtain experimental probes. Among them, the non-binding segment 1 of the same gene is the same sequence; the non-binding segment 2 is the same sequence, the non-binding segment is the non-target region, the rolling circle probe is 1 target region probe plus 2 non-target region probes, the fluorescence probe is 1 non-target region probe plus a fluorophore, and the primer probe is the reverse complement of 1 non-target region probe.

[0047] (c5) Retrieve mouse brain tissues from the retrieval databases of bulkRNA and scRNA processed and collected. Set the expression level of Fragments Per Kilobase of exon model per Million mapped reads (FPKM) of bulkRNA within 10 - 1000, and the average raw count expression level at the scRNA cell type level within 200. Divide the gene set into x genes in the coding rounds and other genes for non-coding imaging according to the output results. The genes for non-coding imaging do not need to be encoded.

[0048] Then, according to the paired mouse brain scRNA data, input the appropriate coding rounds and the x genes to be encoded, and optimize the gene set distribution through the simulated annealing algorithm proposed in this embodiment to minimize the variance of the expression levels among different rounds as much as possible.

[0049] For example, randomly select 10 gene sets of 200 genes from the mouse brain and test them 10 times each, and obtain the results as Figure 3 shown. Compared with the random strategy, this embodiment can effectively reduce the variance among rounds.

[0050] For example, set the coding rounds = 3, randomly adopt four exchange strategies during the iteration process, and generate the initial coding value to remove the possible contaminated coding in the experiment. Finally, obtain the optimal example coding scheme as Figure 5 , Figure 5 It exemplifies the result file of 14 gene codings in three rounds. Just conduct the experiments in sequence according to the situation marked as 1 here.

[0051] (c6) Gene spatial display: According to the experimental order in the optimized gene coding book, add the fluorescence probes corresponding to the genes in different rounds and different channels (assembled in step c4) to make them light up in the specific round channels. The experimental process is to cut a tissue section of the mouse brain with a thickness of 10 μm and attach it to a glass slide. After fixation, add all the rolling circle probes to the buffer for incubation, then use DNA ligase to ligate the nicks on the rolling circle probes, and then add primer probes, DNA polymerase, substrates, etc. for rolling circle amplification. After fixation again, finally add the fluorescence probes corresponding to the genes in the gene coding book to make the genes of the current round emit light, and obtain the original image data through fluorescence microscopy imaging. After performing point recognition on the original image data through the LOG operator, obtain the positions of the bright spots in each round, and record the relative intensity value vector of the original image data of multiple round channels at the position of the bright spot. The gene decoding is performed by calculating the Euclidean distance between this vector and the coding vectors of different genes in the gene coding book. Finally, an example of the decoded gene display is as Figure 6, the Allen reference brain atlas is a single-gene ISH map, where the dots represent the actual gene signals; on the right is the gene point distribution obtained by decoding after actual experiments with probes, marked in green, and the darker the color, the higher the density. Through actual experiments, the Penk gene shows an enriched state in the striatum region of the mouse brain and a two-layer distribution in the mouse cerebral cortex; the Sst gene shows a scattered distribution in the mouse cerebral cortex. These results indicate that the gene distribution obtained by designing probes and the coding table for the target gene is similar to the expression pattern of the Allen reference brain atlas, demonstrating the technical usability of this embodiment.

[0052] Example 2 In this embodiment, to reduce the code threshold for users and improve usability, a public Web-based database system is developed and provided to offer auxiliary usage functions. The development details are as follows: It is developed based on the Django (an open-source Python Web framework) front-end and back-end frameworks. The front-end and back-end interaction uses the JSON (a lightweight data interchange format) format for data transmission via Ajax, and is configured and launched on the server side through Nginx (an HTTP and reverse proxy web server) and uwsgi (an interface). The overall framework adopts the form of a left-side navigation bar, and there are three main functions: Check, Codebook, and Merge, each with a corresponding URL link. On the right is the main function area. Users can independently input the corresponding parameters, and the additional parameters are folded through the advanced button. By clicking the button, the parameters are passed to the server. After the server receives the corresponding parameter input and processes the logical code, it returns the data to the front-end for rendering or downloading and using. Among them, Check, Codebook, Merge, advanced, and button are all buttons, and the links correspond to the displayed pages.

[0053] (d1) Check function: Access the XXX website on the web page. The user selects the appropriate tissue data type through the Species&Tissue&Disease options, determines the required gene set through the GeneSet_Check option, enters a reasonable cut value (the threshold for scRNA and bulkRNA screening) through the advanced button, and then clicks the check button and waits for the output result. The output result will feedback the genes that can be used for coding under the custom cut value, as well as the gene set with relatively high expression levels recommended for non-coding detection, where the higher part will be marked in red, and the genes that cannot be designed for other reasons (the target region cannot be designed, the fragments per kilobase of exon model per million reads mapped (FPKM) expression level of the bulkRNA gene is too low, the gene name is incorrect, or the gene is not in the scRNA and bulkRNA data). Finally, the user downloads the output result according to their needs and adjusts the gene set classification for designing coding and non-coding.

[0054] (d2)Codebook Function: Access the XXX website on the web page. The user selects the appropriate tissue data type through the Species & Tissue & Disease option, sets the appropriate number of Encode Rounds, enters the gene set for encoding sorted in the previous step into the input box of the GeneSet_Encode for encoding, enters the gene set for non-coding sorted in the previous step into the input box of the GeneSet_Single for non-coding, clicks the advanced button, and sets the relevant parameters of the simulated annealing algorithm. Click the codebook button and wait for the output result. The output result includes the process display of simulated annealing and the optimal encoding result, and saves the optimal codebook in the Codebook function link.

[0055] (d3)Merge Function: Access the XXX website on the web page. The user selects the appropriate tissue data type through the Species & Tissue & Disease option, exports the name of the Max Isoform with the highest expression level corresponding to each gene from the check function, and finally clicks the Merge button and waits for the output result. The output result will be automatically downloaded in zip format, and the result can be directly sorted and used for submitting to the sequencing company for ordering probes and directly used in experiments subsequently.

[0056] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A probe design method for multi-round spatial transcriptome imaging, characterized in that, Including: Download the original fastq data from the GEO database and preprocess it to obtain the expression matrices of isoforms and genes, taking into account the single-cell expression matrix, and construct a retrieval database for bulkRNA and scRNA accordingly; Design targeted region probes: Traverse and obtain the targeted segments of all isoforms of the same gene from the known RNA sequences and preprocess them. Divide the obtained candidate probe sequences into two segments for additional non-specific detection to obtain the targeted probe sequences of all isoforms, and then screen the targeted region probes according to all the expression matrices in the retrieval database; Design non-targeted region probes: Generate multiple random sequences by random ATCG and preprocess them to design non-targeted region probes accordingly; Gene codebook optimization: Obtain the single-cell expression matrix from the retrieval database, and optimize the variance of the expression levels of a given gene set at the cell type level between different imaging rounds based on the simulated annealing algorithm to obtain the gene codebook; Gene spatial display: Perform multiple rounds of spatial transcriptome imaging using the targeted region probes, non-targeted region probes, and gene codebook to obtain the gene spatial display decoded according to the gene codebook.

2. The probe design method for multi-round spatial transcriptome imaging according to claim 1, wherein In the design of targeted region probes, traverse and obtain the targeted segments of all isoforms of the same gene from the known RNA sequences and preprocess them. Specifically: Based on the input species name, required gene set, and corresponding isoform name, traverse the sequences in the fasta file storing all RNA sequences according to the required design length to generate a set of potential probe sequences; Screen the set of potential probe sequences through setting Tm value, GC content, removing polycistrons, RNA secondary structure detection, and overall BLAST non-specific inspection to obtain non-specific alignment results and get candidate probe sequences.

3. The probe design method for multi-round spatial transcriptome imaging according to claim 1, wherein In the design of targeted region probes, divide the obtained candidate probe sequences into two segments for additional non-specific detection to obtain the targeted probe sequences of all isoforms, and then screen the targeted region probes according to all the expression matrices in the retrieval database. Specifically: Divide the candidate probe sequences into two segments and perform BLAST non-specific detection separately, and perform non-specific removal on the alignment results of the two segments respectively according to the set unilateral non-specific binding threshold; Obtain the clip-seq data of proteins binding to RNA from the known public database and screen according to the enrichment intensity, and retain the potential binding sites with confidence exceeding the set threshold to obtain the screened clip-seq data; Set the minimum overlapping region ratio threshold, and retain the sequences with the overlapping ratio not exceeding the set overlapping ratio threshold between the probe sequences after non-specific removal and the screened clip-seq data to obtain the targeted probe sequences of all isoforms; According to the expression matrix of bulk RNA of the set species tissue in the retrieval database, preferentially select the targeting segments shared by the highest-expressed isoform and other isoforms under the same gene classification as the targeting region probes, and select a set number of targeting region probes according to the expression levels of bulk RNA and scRNA in the retrieval database.

4. The probe design method for multi-round spatial transcriptome imaging according to claim 1, characterized in that, Design non-targeting region probes: Generate multiple random sequences by random ATCG and preprocess them to design non-targeting region probes. Specifically: Generate a sufficient number of random sequences of a fixed length by random ATCG and retain the sequences that do not specifically bind to the known RNA sequences through BLAST as the potential sequence set; Calculate the Hamming distance of the full length between pairwise sequences in the potential sequence set, and retain the sequences with the Hamming distance exceeding the set threshold; Mark the first 1 / 3 segment and the last 1 / 3 segment of each retained sequence as edge segments, calculate the Hamming distance between pairwise edge segments, and retain the sequences with the Hamming distance exceeding the set threshold, and design non-targeting region probes accordingly.

5. The probe design method for multi-round spatial transcriptome imaging according to claim 1, characterized in that, One targeting region probe plus two non-targeting region probes form a rolling circle probe, one non-targeting region probe plus a fluorophore forms a fluorescent probe, and the reverse complement of one non-targeting region probe forms a primer probe. Assemble the rolling circle probe, the fluorescent probe and the primer probe to obtain the experimental probe.

6. The probe design method for multi-round spatial transcriptome imaging according to claim 1, wherein, In the gene codebook optimization, specifically as follows: Obtain the transcripts per million (TPM) expression of bulk RNA of a given gene set in the set species tissue and the average raw count expression at the scRNA cell type level from the retrieval database; Genes with expression levels exceeding the upper expression threshold and genes with expression levels below the lower expression threshold are not allowed to be added to the coding round. Screen the gene set for codebook design and subsequent coding imaging experiments; Set the number of coding rounds and the number of genes for coding imaging, and iteratively optimize the distribution of the screened gene set among the round channels; In each iteration, randomly select one of the four exchange strategies to process the distribution of the screened gene set; If the variance after adopting the exchange strategy for the gene set distribution is less than the variance before the exchange, enter the next round, where the variance is the total number calculated according to the expression matrix at the scRNA cell type level for different gene combinations highlighted in each round of the gene set distribution, and then calculate the average variance of this total number among each round; If the variance after adopting the exchange strategy is greater than the variance before the exchange, it is judged whether it is less than by randomly generating a value , where is the variance difference before and after the exchange strategy, is the annealing temperature, which decays with the number of iterations; If less than then enter the next round, otherwise reselect the exchange strategy until the iteration requirement is met.

7. The probe design method for multi-round spatial transcriptome imaging according to claim 6, wherein The four exchange strategies are randomly exchange two rows, exchange two rows multiple times, randomly exchange multiple rows as a whole, and randomly reverse the order of a region as a whole.

8. The probe design method for multi-round spatial transcriptome imaging according to claim 6, wherein In the gene codebook optimization, the generation of the initial value is constrained. Specifically: Remove the coding of the same channel in the previous and subsequent rounds; Calculate the number of genes in each round, and remove the gene set distributions with the number of genes less than the preset number in each round, so that the number of genes in each round is within the set interval, thereby constraining the generation of the initial value.

9. A probe design system for multi-round spatial transcriptome imaging, characterized in that, It includes a database construction module, a targeting region probe design module, a non-targeting region probe design module, a gene codebook optimization module, and a gene display module; The database construction module is used to download the original fastq data from the GEO database, preprocess it to obtain the expression matrices of isoforms and genes, taking into account the single-cell expression matrix, and construct the retrieval databases for bulkRNA and scRNA based on this; The targeted region probe design module is used to traverse and obtain the targeted segments of all isoforms of the same gene from the known RNA sequences and preprocess them. The obtained candidate probe sequences are divided into two segments for additional non-specific detection to obtain the targeted probe sequences of all isoforms, and then the targeted region probes are screened according to all the expression matrices in the retrieval database; The non-targeted region probe design module is used to generate multiple random sequences by random ATCG and preprocess them, and design non-targeted region probes based on this; The gene codebook optimization module is used to obtain the single-cell expression matrix from the retrieval database, and optimize the variance of the expression levels of a given gene set at the cell type level between different imaging rounds to be minimized based on the simulated annealing algorithm to obtain the gene codebook; The gene display module is used to perform multi-round spatial transcriptome imaging using the targeted region probes, non-targeted region probes and gene codebook, and obtain the gene spatial display decoded according to the gene codebook.

10. The probe design system for multi-round spatial transcriptome imaging according to claim 9, characterized in that, The non-targeted region probe design module specifically is used for: Generating a sufficient number of random sequences of a fixed length by random ATCG and retaining the sequences that do not specifically bind to the known RNA sequences through BLAST as the potential sequence set; Calculating the Hamming distance of all lengths between pairwise sequences in the potential sequence set, and retaining the sequences with the Hamming distance exceeding the set threshold; Marking the first 1 / 3 segment and the last 1 / 3 segment of each retained sequence as the edge segments, calculating the Hamming distance between pairwise edge segments, and retaining the sequences with the Hamming distance exceeding the set threshold, and designing non-targeted region probes based on this.

Citation Information

Patent Citations

  • Primer probe sequence combination design method and device

    CN116130000A

  • High-throughput single cell space transcriptome method based on multi-round imaging coding

    CN119709956A

  • Probe design of high-throughput single cell space transcriptome based on multi-round imaging coding

    CN119724347A

  • Design and Use of Sortilin Specific Molecular Imaging Ligands

    US20110166032A1

  • Expression-weighted tumor mutational burden as an oncology biomarker

    WO2023107570A1