Method and apparatus for exploring the distribution, efficacy, action, or physiological activity of genome-containing substances within tissues of interest.
The method and apparatus utilize spatially analyzed transcript information and an integrated reference transcript library to provide precise positional and therapeutic insights into cell therapy agents, addressing the limitations of existing methods by enhancing the accuracy and cost-effectiveness of cell distribution and efficacy evaluation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- PORTRAI INC
- Filing Date
- 2023-08-09
- Publication Date
- 2026-04-28
AI Technical Summary
Existing methods for evaluating cell therapy efficacy lack the resolution and accuracy to provide comprehensive insights into the distribution and molecular interactions of administered cells within tissues, particularly due to limitations in tracking and confirming the distribution of cells and understanding molecular changes.
A method and apparatus utilizing spatially analyzed transcript (ST) information and an integrated reference transcript library to classify and search for the distribution, efficacy, and physiological activity of genome-containing substances within tissues, providing precise positional information and therapeutic mechanisms using spatial transcription information.
Enables accurate and cost-effective confirmation of cell distribution and therapeutic mechanisms, offering detailed insights into pharmacokinetics and mode of action, surpassing the limitations of conventional methods like PCR.
Smart Images

Figure 2026513591000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and an apparatus for exploring the distribution, efficacy, action, or physiological activity of genomic substances in tissues of interest.
Background Art
[0002] Cell therapy involves introducing therapeutic cells into a patient and has received considerable attention as a promising treatment option for various diseases. In recent years, cell therapy has been attracting attention as a promising treatment method for various diseases such as cancer (1), autoimmune diseases (2), inflammatory diseases (3), and neurodegenerative diseases (4). Cell therapy agents are characterized by the ability to design each cell function, biocompatibility, and applicability that overcome the blood-brain barrier (BBB) (5), and these characteristics enhance the potential of cell therapy. With the emergence of advanced technologies for cell manipulation (e.g., chimeric antigen receptor-T cell or NK cell therapy (CAR-T / NK), stem cells, tumor infiltrating lymphocytes, or microbiome preparations), the potential of cell therapy continues to grow as a game-changing therapeutic approach.
[0003] Despite the potential of cell therapy as described above, there are limitations in the reproducibility of drug efficacy due to the living characteristics of cells, and it faces significant limitations due to the difficulties in preparation, delivery, and administration [6].
[0004] Traditional methods for evaluating the efficacy of cell therapies often lack the resolution and accuracy necessary to provide meaningful insights into the complex interactions and molecular changes that occur during treatment. For example, poly-chain reactions (PCR) are primarily performed on organ degradation, limiting the confirmation of the microscopic characteristics of cell therapies, particularly the confirmation of the distribution of administered cells within the tissue [7]. Fluorescently labeled cells, frequently used to obtain images of administered cells, have the disadvantage that tracers are easily separated from the cells, limiting cell tracking [8]. Most importantly, it is crucial to note that such methods fail to provide a comprehensive understanding of the detailed molecular changes and interactions that occur between administered cells and host cells within the target tissue. This has created an urgent need for innovative and comprehensive approaches to better understand the true potential of cell therapies.
[0005] Spatially analyzed transcripts have emerged as a groundbreaking tool in the field of molecular biology [9, 10]. By combining spatial information with gene expression data, this technique has enabled researchers to study the dynamic changes in gene expression patterns within tissues and cells at an unprecedented level of detail. Spatially analyzed transcripts have the potential to revolutionize the evaluation of cell therapies by providing a more comprehensive understanding of their effects at the cellular and molecular levels.
[0006] In particular, in the case of cell therapy, PK (Pharmacological Kinetic) data is an important validation parameter. For example, when treating liver fibrosis, stem cell therapy agents can be used, and in this case, the distribution of stem cells in the tissue is an important measure for evaluating the efficacy and safety of the therapy agent. This is considered to be of great importance in the pharmaceutical industry, as it is also required for IND (Investigative New Drug) applications. [Overview of the Initiative] [Problems that the invention aims to solve]
[0007] Therefore, the present invention aims to provide a search method and search apparatus that can provide accurate positional information of cell therapy agents within tissues, thereby providing diverse information such as what cell types are predominant and how they are distributed after administration of the cell therapy agent, and what therapeutic mechanism the cell therapy agent has.
[0008] The present invention aims to provide a method and apparatus useful for exploring the microscopic spatial distribution and basic therapeutic mechanism of cell therapy agents administered to tissue of interest using spatially analyzed transcript (ST) information, and for evaluating various types of properties, including the pharmacokinetics and mode of action of cell therapy agents at the tissue level. [Means for solving the problem]
[0009] In one embodiment of the present invention, In order to explore the distribution, efficacy, action, or physiological activity of genome-containing substances within tissues of interest, To provide a tissue of interest that has a disease or disease, or has been induced, with a substance containing part or all of the genome that is not present in the tissue of interest, and to obtain spatial transcription information of the tissue of interest, An integrated reference transcript library is used to classify the spatial transcript information into transcripts derived from the tissue of interest and transcripts derived from the genome, which include first organ reference transcript information to which the tissue of interest belongs and second organ reference transcript information to which the genome or a substance containing the genome belongs. A search method is provided which includes searching for the distribution, efficacy, action, or physiological activity of the genome-containing substance within a tissue of interest based on the classified information.
[0010] In another embodiment of the present invention, To provide a tissue of interest that has a disease or disease, or has been induced, with a substance containing part or all of the genome that is not present in the tissue of interest, and to obtain spatial transcription information of the tissue of interest, This includes classifying genome-derived transcription information from spatial transcription information using an integrated reference transcription library that includes first organoscopy reference transcription information to which the aforementioned tissue of interest belongs, and second organoscopy reference transcription information to which the genome or a substance containing the genome belongs. A method for classifying genome-derived transcripts from spatial transcript information is provided.
[0011] In yet another embodiment of the present invention, In order to explore the distribution, efficacy, action, or physiological activity of genome-containing substances within tissues of interest, An information receiving unit 110 provides a substance containing part or all of the genome not present in a diseased or diseased tissue of interest, acquires spatial transcription information of the tissue of interest, and receives the obtained information. An origin search unit 120 for spatial transcription information classifies the spatial transcription information into transcription information derived from the tissue of interest and transcription information derived from the genome, using an integrated reference transcription library that includes first organism reference transcription information to which the tissue of interest belongs and second organism reference transcription information to which the genome or a substance containing the genome belongs. A search device 100 is provided, which includes a molecular marker search unit 130 that searches for molecular markers related to the distribution, efficacy, action, or physiological activity of genome-containing substances in tissues of interest based on information from the origin search unit 120.
[0012] In yet another embodiment of the present invention, An information receiving unit 110 provides a substance containing part or all of the genome that is not present in a diseased or diseased tissue of interest, and receives spatial transcription information of the obtained tissue of interest. An apparatus 200 is provided for classifying genome-derived transcripts from spatial transcripts, which includes a spatial transcript origin search unit 120 that classifies the spatial transcripts into tissue-derived transcripts and genome-derived transcripts from the spatial transcripts using an integrated reference transcript library that includes first organism reference transcript information to which the tissue of interest belongs and second organism reference transcript information to which the genome or a substance containing the genome belongs. [Effects of the Invention]
[0013] The present invention utilizes an integrated reference transcript library that includes first organism reference transcript information and second organism reference transcript information to which a genome or substance containing a genome belongs. Compared to using first and second organism reference transcript information separately, the nucleotide sequence mapping is performed to the reference with the higher matching rate among the two reference transcripts, thereby improving the accuracy of mapping or classification.
[0014] The present invention provides precise positional information of a genome-containing substance (e.g., a cell therapy agent) within a tissue, thereby providing diverse information such as which cell types are predominant and how they are distributed after administration of the substance, what therapeutic mechanisms the substance has within the tissue, and possible side effects.
[0015] This invention provides the exploration of the microscopic spatial distribution and the basic therapeutic mechanism of genome-containing substances (e.g., cell therapy agents) administered to tissues of interest using spatially analyzed transcript (ST) information, and is useful for evaluating various types of properties, including the pharmacokinetics and mode of action of genome-containing substances at the tissue level.
[0016] Therefore, compared to conventional methods (e.g., PCR) that require enormous time and expense to explore the distribution and effects of substances containing genomes, this method allows for the confirmation of their distribution with high accuracy using a very simple and cost-effective method.
Brief Description of the Drawings
[0017] [Figure 1] It is a flowchart of a search method according to an aspect of the present invention. [Figure 2] It is a diagram showing a method of implementing a spatial transcript information acquisition step (S10) for sharing spatial information of a tissue of interest having a genomic substance according to an aspect of the present invention in the tissue of interest. [Figure 3] It is a diagram showing spatial transcript information for sharing spatial information of a tissue of interest in the spatial transcript information acquisition step (S10) according to an aspect of the present invention. [Figure 4] In the case of a tissue containing a genetically modified genomic substance having the same origin, an integrated reference is created, and a figure showing the result indicating the expression of a new MUC isoform gene (simply denoted as MUCX) that is not present in the original reference is shown. [Figure 5] In the search step (S30) according to an aspect of the present invention, an example of an analysis method or algorithm that can be used to search for the distribution, efficacy, action, or physiological activity of a genomic substance in a tissue of interest is shown. For example, it is a diagram showing that well-known analysis algorithms or analysis methods such as SPADE, DEG, Stopover, CellDART, MIA, RCTD, and gene ontology analysis can be used. [Figure 6] It is a schematic diagram of the process of obtaining spatial transcript information of a tissue of interest with "Nor", "Con", and "Exp" in an embodiment of the present invention. [Figure 7] In "Nor", "Con", and "Exp" of an embodiment of the present invention, it is a normalized or non-normalized histogram of "all_transcripts", "all_human_transcripts", and "%human". [Figure 8] In "Nor", "Con", and "Exp" of an embodiment of the present invention, it is a diagram showing the "spots_with_human_cell" variable, which is a binary indicator that is yellow only when %human exceeds a threshold at a spot. [Figure 9] H&E stained images of the tissues of interest in three samples according to an embodiment of the present invention, namely, "Nor", "Con", and "Exp". [Figure 10] An image showing the number of all transcripts (first column), the number of mouse transcripts (second column), and the number of human transcripts (third column) in each of the three samples according to an embodiment of the present invention, namely, "Nor", "Con", and "Exp". Here, the human transcripts are shown on a scale of the number of transcripts from 0 to 100. The human transcripts were only observable in the "Exp" sample. Also, the fourth column shows the threshold of "is_human", which is determined by adding 10 times the %human standard deviation of the "Con" sample to the average, and includes the human transcripts detectable only in "Exp" and shows a significant amount of spots. Therefore, it was confirmed that statistically significant detection of human transcripts was made on the Visium platform by hMSC cell injection. [Figure 11] A figure showing the results of clustering three samples according to an embodiment of the present invention, namely, "Nor", "Con", and "Exp", based on mouse transcripts (left side) and based on human transcripts (right side), respectively. [Figure 12] A figure showing the result of clustering by integrating each spot in three samples according to an embodiment of the present invention, namely, "Nor", "Con", and "Exp". [Figure 13]This figure shows the clustering analysis results for three samples, namely "Nor," "Con," and "Exp," according to one embodiment of the present invention. a shows the results of DimPlot (top) and SpatialDimplot (bottom) by cluster number for each sample. b shows the VlnPlot, which represents the percentage of human transcripts (%human) among all transcripts, based on cluster number (horizontal axis). Interestingly, cluster 5 showed a significantly larger number of human transcripts than the other clusters. c shows the ratio of specific cluster numbers to each sample, with cluster 5 being almost absent in the "Nor" and "Con" samples. d is a GO plot for each of the top 20 spatially abundant genes in cluster 5 that are upwardly regulated by biological processes (BP), cellular components (CC), and molecular functions (MF), respectively, the genes including Lcn2 (highest log FC), Msln, Chil3, C3, Upk3b, Spp1, Col4a1, Serpina3n, Wfdc21, Col3a1, Wfdc17, Gm13889, Chil1, Mgp, Sftpd, Slpi, Ctsc, Fmo2, Scd1, and Napsa (in order). [Figure 14]This figure shows the results of comparing transcripts of a "Con" sample according to one embodiment of the present invention with other samples, "Nor" and "Exp". a is an EnhancedVolcano plot comparing "Nor" and "Con", where "Log2 fold change" > 0 indicates a gene that is abundant in the "Con" sample. b is an EnhancedVolcano plot comparing "Con" and "Exp", where "Log2 fold change" > 0 indicates a gene that is abundant in the "Exp" sample. Here, the genes tagged with "mm10" and "GRCh38" are mouse and human-derived genes, respectively. c is a GO plot for the top 20 DEGs (adjusted p-value < 0.05, log FC order) that are abundant in "Nor" compared to "Con". d is a GO plot for the top 20 DEGs (adjusted p-value < 0.05, log FC order) that are abundant in "Con" compared to "Nor". e is a GO plot for the top 20 DEGs (adjusted p-value < 0.05, log FC order) that are abundant in "Con" compared to "Exp". f is a GO plot for the top 20 DEGs (adjusted p-value < 0.05, log FC order) that are abundant in "Exp" compared to "Con". [Figure 15] This figure shows the top 6 DEGs (adjusted p-value < 0.05; log FC order) for "Nor" compared to "Con". [Figure 16] This figure shows the top 6 DEGs (adjusted p-value < 0.05; log FC order) for "Con" compared to "Nor". [Figure 17] This figure shows the top 6 DEGs (adjusted p-value < 0.05; log FC order) for "Con" compared to "Exp". [Figure 18] This figure shows the top 6 DEGs (adjusted p-value < 0.05; log FC order) for "Exp" compared to "Nor". [Figure 19]This figure shows the percentage of human transcripts and the genes spatially associated with them. Here, a shows the top 20 genes spatially associated with %human by exponential factors, b shows the bottom 20 genes, and c shows the GO analysis for mouse genes among the top 20 spatially associated genes with %human. [Figure 20] This figure shows the CellDART results. 'a' shows the spatial mapping results for the proportions of six cell types in the "Exp" sample, along with the human transcript and %human distribution, while 'b' is a correlation coefficient matrix comparing the %human proportions and each mouse cell type proportions predicted by CellDART. Spearman correlation coefficients were used, and the results showed that endothelial cells and epithelial cells were the most and least associated cell types with %human, respectively. [Figure 21] This figure shows the Stopover results. Specifically, a is the Stopover result comparing the spatial distribution of cell types. Here, the CellDART score obtained by multiplying %human and (100-%human) was used to show the spatial distribution between human and mouse cells. J_comp and J_local in the plot represent the overall and local Jaccard index, respectively. CC indicates the concatenated components. b represents the phase of the LR pair with the highest Jaccard index for LR colocalization, where the yellow, blue, and green spots correspond to the ligand, receptor, and mixed region, respectively. [Figure 22] This figure shows the STopover results derived from a comparison of the distribution of each cell type with %human. Here, the yellow, blue, and green spots correspond to ligand, receptor, and mixed region, respectively, and the cell types are listed in order of their overall Jackard index (J_comp). [Figure 23]This figure shows the CellDART results for mouse cells, specifically aDCs, adipocytes, B cells, basophils, CD4 T cells, and CD8 T cells, in three samples, "Nor," "Con," and "Exp," which embody one example of the present invention. [Figure 24] This figure shows the CellDART results for mouse cells, specifically aDCs, adipocytes, B cells, basophils, CD4 T cells, and CD8 T cells, in three samples, "Nor," "Con," and "Exp," which embody one example of the present invention. [Figure 25] This figure shows the CellDART results for mouse cells, specifically cDCs, chondrocytes, DCs, endothelial cells, eosinophils, and erythrocytes, in three samples, "Nor," "Con," and "Exp," which embody one example of the present invention. [Figure 26] This figure shows the CellDART results for mouse cells, specifically cDCs, chondrocytes, DCs, endothelial cells, eosinophils, and erythrocytes, in three samples, "Nor," "Con," and "Exp," which embody one example of the present invention. [Figure 27] This figure shows the CellDART results for three samples, "Nor," "Con," and "Exp," representing one embodiment of the present invention, for mouse cells, specifically monocytes, mv. endothelial cells, myocytes, neutrophils, NK cells, and NKT. [Figure 28]This figure shows the CellDART results for three samples, "Nor," "Con," and "Exp," representing one embodiment of the present invention, for mouse cells, specifically monocytes, mv. endothelial cells, myocytes, neutrophils, NK cells, and NKT. [Figure 29] This figure shows the CellDART results for three samples, "Nor," "Con," and "Exp," representing one embodiment of the present invention, specifically for mouse cells including osteoblasts, pDCs, pericytes, sebocytes, skeletal muscle, and smooth muscle. [Figure 30] This figure shows the CellDART results for three samples, "Nor," "Con," and "Exp," representing one embodiment of the present invention, specifically for mouse cells including osteoblasts, pDCs, pericytes, sebocytes, skeletal muscle, and smooth muscle. [Figure 31] This figure shows the CellDART results for mouse cells, specifically Tgd cells (Tgd.cells), Th1 cells (Th1.cells), Th2 cells (Th2.cells), Tregs, ImmuneScore, and StromaScore, in three samples, "Nor," "Con," and "Exp," which embody one example of the present invention. [Figure 32] This figure shows the CellDART results for mouse cells, specifically Tgd cells (Tgd.cells), Th1 cells (Th1.cells), Th2 cells (Th2.cells), Tregs, ImmuneScore, and StromaScore, in three samples, "Nor," "Con," and "Exp," which embody one example of the present invention. [Figure 33]a is an EnhancedVolcano plot comparing "Nor" and "Exp," where "Log2 fold change" > 0 indicates a gene rich in the "Exp" sample. b is a GO plot for the top 20 DEGs (adjusted p-value < 0.05, log FC order) rich in "Exp" compared to "Con." [Figure 34] This figure shows a search device 100 according to one aspect of the present invention. [Figure 35] This is a block diagram showing the various components that make up the search device 100 shown in Figure 34. [Figure 36] Figure 34 is a block diagram showing the various components that make up the molecular marker search unit 130 of the search device 100. [Figure 37] This figure shows a classification device 200 according to one aspect of the present invention. [Modes for carrying out the invention]
[0018] The first aspect of the present invention is, In order to explore the distribution, efficacy, action, or physiological activity of genome-containing substances within tissues of interest, To provide a tissue of interest that has a disease or disease, or has been induced, with a substance containing part or all of the genome that is not present in the tissue of interest, and to obtain spatial transcription information of the tissue of interest (S10), An integrated reference transcript library is used to classify the spatial transcript information into transcripts derived from the tissue of interest and transcripts derived from the genome, which include information from a first organelle reference transcript to which the tissue of interest belongs and information from a second organelle reference transcript to which the genome or a substance containing the genome belongs (S20). A search method is provided which includes searching for the distribution, efficacy, action, or physiological activity of the genome-containing substance in the tissue of interest from the classified information (S30).
[0019] In the step of obtaining spatial transcription information of the tissue of interest (S10), any substance containing part or all of the genome that is not present in the tissue of interest can be used without restriction. For example, cells derived from one or more species selected from the group consisting of umbilical cord, bone marrow, adipose tissue, blood, umbilical cord blood, liver, skin, gastrointestinal tract, placenta, and uterus, specifically stem cells, can be used. These may be the same species or different from the origin species of the tissue of interest, and in specific cases, they may be different. If the origin species is the same as that of the tissue of interest, the genome-containing substance may have a sequence that is partially or completely different from the natural genome due to one or more of the base sequences of the natural genome of the tissue of interest being genetically altered by genetic manipulation, deformation, deletion, insertion, etc.
[0020] The genome-containing material may be derived from a mammal. The mammal may be a human, monkey, chimpanzee, dog, cat, cattle, goat, pig, mouse, or rat, and more specifically, a genome-containing material derived from a human, rat, or mouse. The genome-containing material may be a microorganism, cell, stem cell, extracellular vesicle, vector, vehicle, or virus that contains a genome. The genome includes a coding sequence and a non-coding sequence that encode proteins, and may be a target for exploring therapeutic effects or mechanisms of action, or a candidate substance for therapeutic agent screening.
[0021] The tissue of interest may be, but is not limited to, the skin, the intestines (such as the small or large intestine), the heart, lungs, kidneys, liver, spleen, muscle, or tumor tissue. The tissue of interest is one in which a major disease or disorder of social interest exists, or in which such a disease or disorder has been induced. The disease or disorder is not particularly limited and may be various cancers, brain diseases, neurological diseases, liver diseases, intestinal diseases, fibrosis, immune diseases, viral diseases, kidney diseases, inflammatory diseases, metabolic diseases, skin diseases, diabetic diseases, infectious diseases, cardiovascular diseases, neurodegenerative diseases, etc. More specifically, it may be cancer, brain diseases, diabetic diseases, inflammatory diseases, viral diseases, infectious diseases, pulmonary fibrosis, etc.
[0022] The provision of the genome-containing substance to the tissue of interest may include either directly providing the genome-containing substance into the tissue of interest, or administering it to the target body systemically (e.g., by intravenous injection) and then distributing it to the tissue of interest.
[0023] By providing the genome-containing material into the tissue of interest, the genome is expressed in the tissue of interest. The tissue of interest then contains both genome-derived transcripts and tissue-derived transcripts, thereby obtaining transcript information that shares the spatial information of the tissue of interest. Spatially resolved transcriptome information is a technique that provides hundreds to tens of thousands of gene expression data at once, obtaining full-length or partial gene expression, including spatial information (see Figure 3). The spatially resolved transcriptome information can be analyzed by fixing a frozen fragment of the tissue of interest in a mold and undergoing steps such as permeabilization, cDNA synthesis, and RNA sequencing. The fragment of the tissue of interest used to obtain the spatially resolved transcriptome information may optionally be a tissue fragment "continuous" with the fragment used to obtain a stained image of the tissue of interest. The spatially resolved transcriptome information may also include the number and type of transcript RNA reads within a spot having spatial coordinates within the tissue of interest. Examples of RNA types include whether to examine gene expression, exons, or introns; whether to include event-based alternative splicing or isoform-based alternative splicing; whether to examine mutation burden; and whether to include noncoding RNA or nontranslated RNA.
[0024] Subsequently, in the step of classifying the spatial transcript information into transcripts derived from tissues of interest and transcripts derived from genomes (S20), an integrated reference transcript library can be used, which includes first organism reference transcript information to which the tissue of interest belongs and second organism reference transcript information to which the genome or a substance containing the genome belongs. The origin information of the integrated reference transcript library can be classified, for example, using SpaceRanger.
[0025] The integrated reference transcript library can be obtained by integrating the first organelle reference transcript information to which the tissue of interest belongs with the second organelle reference transcript information to which the substance containing the genome belongs, or vice versa, or by obtaining transcript information obtained by integrating the first and second organelles. The first and second organelles may be the same or different, and in specific examples they may be different. Using the integrated reference transcript library has the advantage of significantly higher matching accuracy compared to classifying the origin of spatial transcript information by applying the first organelle reference transcript information and the second organelle reference transcript information separately to the spatial transcript information. In other words, when classifying using the integrated reference transcript library, the origin of the spatial transcript information can be confirmed with even greater accuracy than when mapping each reference separately. For example, if the tissue of interest is a mouse tissue and the genome-containing material is a human-derived cell, the first organelle reference transcript information is mm10, the second organelle reference transcript information is human reference information GRCh38, and the integrated reference transcript library is information obtained by merging mm10 and GRCh38. Alternatively, the integrated reference transcript library may be a single-cell RNA sequencing library in block-diagonal matrix form containing, for example, mouse reference transcript information (GSE124872) and human reference transcript information (GSE147287).
[0026] An example of a method for generating the aforementioned combined single-cell RNA sequencing library is:
number
[0027] Information on each organism's reference transcript can be obtained using known reference transcript information sites or known materials, based on the origin information of each organism in the tissue or genome-containing material of interest, and an integrated reference transcript library can also be obtained using known methods. An example of the above is: ncbi (e.g., E. coli). The E. coli reference can be obtained at https: / / www.ncbi.nlm.nih.gov / assembly / GCF_000005845.2 / ?shouldredirect=false, but references for other species can also be found on this website.
[0028] By matching the obtained spatial transcript information with an integrated reference transcript library, it is possible to classify the spatial transcript information into tissue-derived transcript information and genome-derived transcript information. In other words, by matching with the integrated reference transcript library, spatial transcript information tagged with origin information can be obtained. This makes it possible to easily and accurately confirm the distribution of genome-containing substances within a tissue of interest using origin information tagged with genome-derived transcript information, without complex procedures such as PCR. For example, in tissues containing cells from different origins, or in tissues containing genome-containing substances of the same origin but with genetic modifications, it is possible to confirm where and to what extent genetically modified genome-containing substances or cells from different origins penetrate or are distributed, and this can be used as biodistribution data for cells administered in cell therapy.
[0029] For tissues containing genetically modified genome-containing material from the same origin, the method for creating an integrated reference is as follows: Basically, the method for creating a custom reference is well described in the official documentation provided by 10x Genomics under the title "Build a Custom Reference (cellranger mkref)". For example, the method for creating a custom reference by inserting the (jellyfish-derived) GFP gene into the human genome reference is as follows: First, prepare a GFP.fa file containing the following content. You can use Linux commands such as cat and echo, or you can input it manually.
[0030] >GFP TACACACGAATAAAAGATAACAAAGATGAGTAAAGGAGAAGAACTTTTCACTGGAGTTGTCCCAATTCTT
[0031] GTTGAATTAGATGGCGATGTTAATGGGCAAAAATTCTCTGTCAGTGGAGAGGGTGAAGGTGATGCAACAT
[0032] ACGGAAAACTTACCCTTAAATTTATTTGCACTACTGGGAAGCTACCTGTTCCATGGCCAACACTTGTCAC
[0033] TACTTTTCCTTATGGTGTTCAATGCTTTTCAAGATACCCAGATCATATGAAACAGCATGACTTTTTCAAG
[0034] AGTGCCATGCCCGAAGGTTATGTACAGGAAAGAACTATATTTTACAAAGATGACGGGAACTACAAGACAC
[0035] GTGCTGAAGTCAAGTTTGAAGGTGATACCCTTGTTAATAGAATCGAGTTAAAAGGTATTGATTTTAAAGA
[0036] AGATGGAAACATTCTTGGACACAATTGGAATACAACTATAACTCACATAATGTATACATCATGGCAGAC
[0037] AAACCAAAGAATGGAATCAAAGTTAACTTCAAATTAGACACAACATTAAAGATGGAAGCGTTCAATTAG
[0038] CAGACCATTATCAACAAAATACTCCAATTGGCGATGGCCCTGTCCTTTTACCAGACAACCATTACCTGTC
[0039] CACACAATCTGCCCTTTCCAAAGATCCCAACGAAAAGAGAGATCACATGATCCTTCTTGAGTTTGTAACA
[0040] GCTGCTGGGATTACACATGGCATGGATGAACTATACAAATAAATGTCCAGACTTCCAATGACACTAAAG
[0041] TGTCCGAACAATTACTAAATTCTCAGGGTTCCTGGTTAAATTCAGGCTGAGACTTTATTTATATATTTAT
[0042] AGATTCATTAAAATTTTATGAATAATTTATTGATGTTATTAATAGGGGCTATTTTCTTATTAAATAGGCT
[0043] ACTGGAGTGTAT
[0044] Secondly, prepare a GFP.gtf file containing the following content. You can use Linux commands such as cat and echo, or you can enter it manually: GFP unknown exon 1 922.+.gene_id "GFP";transcript_id "GFP";gene_name "GFP";gene_biotype "protein_coding";
[0045] Thirdly, the following Linux command will insert the contents of the GFP.fa file into human_genome.fa, which stores the base sequence portion of the human reference file: cat GFP.fa>>human_genome.fa
[0046] Fourth, the following Linux command will insert the contents of the GFP.gtf file into human_genome.gtf, which stores the gene annotation portion of the human reference file: cat GFP.gtf>>human_genome.gtf
[0047] Fifth, you can generate a human+GFP custom reference through the SpaceRanger mkref command: spaceranger mkref--genome=human_plus_GFP--fasta=human_genome.fa--genes=human_genome.gtf
[0048] Subsequently, in the exploration phase (S30), the classified information (information tagged with origin information) can be used to explore the distribution, efficacy, action, or physiological activity of the genome-containing substance within the tissue of interest.
[0049] The aforementioned search may include, for example, extracting molecular markers spatially associated with the ratio of genome-derived transcripts to the total number of spatial transcripts in the tissue of interest, from spatial transcript information tagged with origin information. The molecular markers may be single molecules derived from DNA, RNA, metabolites, proteins, protein fragments, etc., or molecular information based on patterns thereof. The extraction of spatially associated molecular markers can be performed using DEG (Differentially Expressed Genes), correlation analysis by calculating correlation coefficients, or image similarity evaluation algorithms, for example, Pearson correlation coefficient, Spearman correlation coefficient, or Kendall correlation coefficient can be used for calculating the correlation coefficient. From the molecular markers, the distribution, efficacy, action, or physiological activity of genome-containing substances in the tissue of interest can be explored.
[0050] For the aforementioned search, the spatial image of the classified information or the overall spatial image can be divided into one or more clusters. These one or more clusters may be classified by the intensity of the spatial image of the tissue of interest, or by an algorithm that divides the spatial image of the tissue of interest into multiple patches and classifies them based on the similarity of the image features to each patch. For example, the spatial image of the tissue of interest may be divided into 394 × 384 patches, each patch being 5 × 5, and 512 features may be extracted from each patch, which may then be classified into one or more clusters based on each feature. The tissue images of interest may be classified into clusters such as cluster 1, cluster 2, cluster 3, etc., and the total number of clusters can be one or more, specifically 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20, and is not limited to these. The search may include comparing one or more clusters, for example, selecting a cluster in which the number of transcripts derived from genome-containing substances is significantly higher than that of other clusters, and comparing the selected cluster with the transcripts of other clusters. For example, if the number of transcripts derived from genome-containing substances is significantly higher in cluster 4 among clusters 1 to 6, and genes that correlate with cluster 4 are identified, and genes associated with hypoxia, glucose metabolism, and cell death mechanisms are dominant, then it can be inferred that cluster 4 is in a hypoxic state due to genome-containing substances, and that this is causing cell death.
[0051] For the aforementioned search, a cell typology analysis algorithm or gene ontology (GO) analysis can be performed on the spatial transcription information of the classified information or the overall spatial transcription information. As the cell typology analysis algorithm, Fisher's exact test, maximum likelihood estimation, domain adaptation classification, logistic regression analysis, or negative binomial regression analysis algorithms can be used. More specifically, known cell typology analysis algorithms such as the CellDART algorithm, spSeudoMap algorithm, xCell algorithm, RCTD algorithm, Seurat algorithm, celltypist algorithm, cell2location algorithm, Scanorama algorithm, SPOTlight algorithm, DSTG algorithm, CellTrek algorithm, sc-type algorithm, TIMER algorithm, CIBERSORT algorithm, MCP-counter algorithm, Quantiseq algorithm, EPIC algorithm, or scanpy's injest algorithm, MIA algorithm, STopover algorithm can be used.
[0052] For the aforementioned exploration, one or more of the following methods can be combined and implemented: extraction of molecular markers spatially related to the ratio of genome-derived transcripts, one or more cluster partitioning methods, image characterization methods, cell type analysis algorithms, and gene ontology analysis methods. In one example, after performing each of the above analytical methods, the results can be combined to explore the distribution, efficacy, action, or physiological activity of genome-containing substances within the tissue of interest.
[0053] In another second mode, the genome-containing material is not provided, and control transcript information of a second tissue of interest that has or has been induced with the same disease or illness as the tissue of interest is obtained and used in the search. The second mode differs from the first mode only in that it additionally uses control transcript information without the provision of genome-containing material, and the other processes can be carried out in the same way as in the modes described above. That is, steps S20 and S30 in the modes described above can be carried out in the same way for the control transcript information of the second tissue of interest.
[0054] Specifically, an integrated reference transcript library is used to classify the control transcript information into transcripts derived from the second tissue of interest and transcripts derived from the genome (S20'). This classification is then used in the step of exploring the distribution, efficacy, action, or physiological activity of the genome-containing substance within the tissue of interest (S30).
[0055] Since no genome-containing material is provided for the control transcript information, the S20' step may be omitted or may not be omitted, but if it is performed, it is expected that genome-derived transcript information will not exist. The second tissue of interest may have the same origin as the tissue of interest in the manner described above.
[0056] The specific search methods described in S30 above, such as one or more cluster partitioning methods, cell type analysis algorithms, and gene ontology analysis methods, can be applied identically to the control transcript information. In one example, after performing each of the above analysis methods, the results can be used to search for the distribution, efficacy, action, or physiological activity of genome-containing substances within tissues of interest. For example, in step S30, compared with the control transcript information, it is possible to extract transcripts that are expressed higher or lower in the spatial transcript information (e.g., spatial transcript information tagged with or untagged with origin information) within tissues of interest containing genome-containing substances. Conversely, when compared with spatial transcript information (e.g., spatial transcript information tagged with or untagged with origin information), transcripts that are expressed higher or lower in the control transcript information can be extracted. For example, in step S30, the results of cell type analysis of control transcription data can be compared with the results of cell type analysis of spatial transcription data within tissues of interest containing genome-containing substances (e.g., spatial transcription data with or without origin information tagged), and the results can be used in the search in step S30. For example, in step S30, one or more clusters for control transcription data can be compared with one or more clusters for spatial transcription data in tissues of interest containing genome-containing substances (e.g., spatial transcription data with or without origin information tagged), and the results can be used in the search in step S30.
[0057] Furthermore, in another third mode, the first and / or second modes may further include obtaining normal transcript information of a normal third tissue of interest that is free from or has not been induced by the disease or disorder, and using this information for the search. The third mode differs from the modes described above only in that it does not provide genome-containing material and additionally uses normal transcript information of a normal tissue that is free from or has not been induced by the disease or disorder, and the other processes can be carried out in the same way as in the modes described above. That is, steps S20 (or S20') and S30 in the modes described above can be carried out in the same way for normal transcript information of the third tissue of interest.
[0058] Specifically, an integrated reference transcript library is used to classify the normal transcript information into transcripts derived from the third tissue of interest and transcripts derived from the genome (S20''), which include information from a first organelle reference transcript to which the third tissue of interest belongs and information from a second organelle reference transcript to which the genome or a substance containing the genome belongs. The classified information can then be used in the step of exploring the distribution, efficacy, action, or physiological activity of the genome-containing substance within the tissue of interest (S30).
[0059] Since normal transcript information does not provide genome-containing material, step S20'' may be omitted or may not be omitted, but if performed, it is expected that genome-derived transcript information will not exist. The third tissue of interest may have the same origin as the first or second tissue of interest in the manner described above.
[0060] The specific search methods described in S30 above, such as one or more cluster partitioning methods, cell type analysis algorithms, and gene ontology analysis methods, can be applied identically to the normal transcript information. In one example, after performing each of the above analysis methods, the results can be used to search for the distribution, efficacy, action, or physiological activity of genome-containing substances within tissues of interest. For example, in step S30, compared with normal transcript information, it is possible to extract transcript information that is expressed higher or lower among spatial transcript information (e.g., spatial transcript information with or without origin information tagged) within tissues of interest containing genome-containing substances, or to extract transcript information that is expressed higher or lower among control transcript information, or to extract transcript information that is expressed higher or lower in the normal transcript information when compared with spatial transcript information (e.g., spatial transcript information with or without origin information tagged) or control transcript information. For example, in step S30, the results of cell type analysis of normal transcription information can be compared with the results of cell type analysis of spatial transcription information of genome-containing substances within tissues of interest (e.g., spatial transcription information with or without origin information tagged) and / or the results of cell type analysis of control transcription information, and the results can be used in the search in step S30. For example, in step S30, one or more clusters for normal transcription information can be compared with one or more clusters for spatial transcription information of genome-containing substances within tissues of interest (e.g., spatial transcription information with or without origin information tagged) and / or one or more clusters for control transcription information, and the results can be used in the search in step S30.
[0061] The S30 exploration may further include comparing one or more pieces of information from the spatial transcript information, control transcript information, and normal transcript information of the tissue of interest to search for enhanced or suppressed molecular markers derived from genome-containing substances.
[0062] Furthermore, in a fourth mode, one or more of the modes described above may be further included in staining one or more of the tissues of interest, second tissues of interest, and third tissues of interest to obtain stained images of each tissue. The staining method may be, but is not limited to, ALP assay (alkaline phosphate assay), Sirius red staining, Alcian blue staining, pH map, H&E staining, Trichrome staining, PAS (Periodic acid-Schiff) staining, or immunohistochemical staining. The results of the stained tissue images may be used as a complementary or supplementary means to verify the search results according to the present invention or to extract information useful when searching for molecular markers. In one example, the tissue stained images may be used as is, or each characteristic of the image may be extracted by applying an artificial intelligence-based image characteristic extraction algorithm, such as the SPADE (Spatial gene expression patterns by deep learning of tissue images) algorithm, and SPADE gene information associated with the extracted characteristics can be obtained.
[0063] In yet another fifth aspect of the present invention, To provide a tissue of interest that has a disease or disease, or has been induced, with a substance containing part or all of the genome that is not present in the tissue of interest, and to obtain spatial transcription information of the tissue of interest, This includes classifying genome-derived transcription information from spatial transcription information using an integrated reference transcription library that includes first organoscopy reference transcription information to which the aforementioned tissue of interest belongs, and second organoscopy reference transcription information to which the genome or a substance containing the genome belongs. A method is provided for classifying genome-derived transcripts from spatial transcript information of an organization of interest. This classification method can further improve the accuracy of classification compared to classifying genome-derived transcripts from spatial transcript information using first-organism or second-organism reference transcript information, respectively, without using an integrated reference transcript library.
[0064] In yet another sixth aspect of the present invention, In order to explore the distribution, efficacy, action, or physiological activity of genome-containing substances within tissues of interest, An information receiving unit 110 provides a substance containing part or all of the genome that is not present in a diseased or diseased tissue of interest, obtains spatial transcription information of the tissue of interest, and receives the obtained information. An origin search unit 120 for spatial transcription information classifies the spatial transcription information into transcription information derived from the tissue of interest and transcription information derived from the genome, using an integrated reference transcription library that includes first organism reference transcription information to which the tissue of interest belongs and second organism reference transcription information to which the genome or a substance containing the genome belongs. A search device 100 is provided, which includes a molecular marker search unit 130 that searches for molecular markers related to the distribution, efficacy, action, or physiological activity of genome-containing substances within tissues of interest based on information from the origin search unit 120.
[0065] The seventh aspect relates to a system including the search device 100 of the present invention (Figure 34).
[0066] In the sixth and seventh modes, "genome-containing substances," "tissues of interest," "spatial transcript information," "first organelle reference transcript information," "second organelle reference transcript information," "integrated reference transcript library," and "molecular markers" are the same as those described in the respective modes above, so detailed descriptions of these are omitted, and the following will describe only the non-overlapping parts.
[0067] The information receiving unit 110 receives all information on genome-derived transcripts and transcripts derived from the tissue of interest that are expressed in the tissue of interest as a result of providing the genome-containing substance into the tissue of interest. In other words, it receives transcript information that shares spatial information of the tissue of interest.
[0068] The spatial transcript information may be obtained by fixing a frozen fragment of the tissue of interest to a mold and analyzing it through steps such as permeabilization, cDNA synthesis, and RNA sequencing, or it may be the number of transcript RNA reads and RNA types within a spot having spatial coordinates within the tissue of interest. Examples of RNA types include origin information such as whether it is mouse RNA or human RNA, whether to look at gene expression, exons, or introns, whether to include event-based alternative splicing or isoform-based alternative splicing, whether to look at mutation levels, and whether to include non-coding RNA or uncoding RNA. The information receiving unit 110 receives the spatial transcript information described above.
[0069] The origin search unit 120 classifies the spatial transcript information into tissue-derived transcripts and genome-derived transcripts from an integrated reference transcript library that includes first organism reference transcript information to which the tissue of interest belongs and second organism reference transcript information to which the genome or a substance containing the genome belongs. As a result, each origin is tagged for all spatial transcript information. In other words, the origin search unit 120 can obtain spatial transcript information tagged with origin information by matching it with the integrated reference transcript library.
[0070] The molecular marker search unit 130 can search for the distribution, efficacy, action, or physiological activity of genome-containing substances within tissues of interest based on the information from the origin search unit 120.
[0071] The molecular marker search unit 130 may include, for example, a correlation analysis unit 132 that extracts molecular markers spatially related to the ratio of genome-derived transcripts to the total number of spatial transcripts in the tissue of interest from spatial transcript information tagged with origin information. The correlation analysis unit 132 may be a single molecule derived from DNA, RNA, metabolites, proteins, protein fragments, etc., or molecular information based on patterns thereof. Extracting the spatially related molecular markers can be done using DEG (Differentially Expressed Genes), correlation analysis by calculating correlation coefficients, or image similarity evaluation algorithms, and for example, Pearson correlation coefficient, Spearman correlation coefficient, or Kendall correlation coefficient calculation can be used for the correlation coefficient calculation. From the molecular markers, the distribution, efficacy, action, or physiological activity of genome-containing substances in the tissue of interest can be searched.
[0072] The molecular marker search unit 130 may include a clustering unit 135 that divides the classified spatial transfer image of the information or the overall spatial transfer image into one or more clusters. The one or more clusters may be classified by the intensity of the spatial transfer image of the tissue of interest, or by an algorithm that divides the spatial transfer image of the tissue of interest into multiple patches and classifies them by the similarity of the image features to each patch. For example, the spatial transfer image of the tissue of interest may be divided into 394 × 384 patches, each patch may be 5 × 5, and 512 features may be extracted from each patch, and each feature may be used as a criterion for classification into one or more clusters. The tissue images of interest may be classified into cluster 1, cluster 2, cluster 3, etc., and the total number of clusters can be one or more, specifically 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 clusters, and is not limited to these. The molecular marker search unit 130 may include comparing one or more clusters, for example, selecting a cluster in which the number of genome-containing substance-derived transcripts is significantly higher than that of other clusters, and comparing the selected cluster with the transcripts of other clusters.
[0073] The molecular marker search unit 130 may include a cell type analysis unit 133 and / or a gene ontology (GO) analysis unit 134 for the classified spatial transcription information or the overall spatial transcription information. The cell type analysis unit 133 can use Fisher's exact test, maximal probability assessment, domain adaptation classification, logistic regression analysis, or negative binomial regression analysis algorithms, and more specifically, it can use known cell type analysis algorithms such as the CellDART algorithm, STopover algorithm, spSeudoMap algorithm, xCell algorithm, RCTD algorithm, Seurat algorithm, celltypist algorithm, cell2location algorithm, Scanorama algorithm, SPOTlight algorithm, DSTG algorithm, CellTrek algorithm, sc-type algorithm, TIMER algorithm, CIBERSORT algorithm, MCP-counter algorithm, Quantiseq algorithm, EPIC algorithm, or scanpy's injest algorithm.
[0074] The molecular marker search unit 130 may include an image characteristic analysis unit 131 that extracts transcript information using an image (or image) characteristic extraction algorithm using artificial intelligence. The image characteristic analysis unit 131 is configured to extract transcript information using an image (or image) characteristic extraction algorithm using artificial intelligence, and the image characteristic extraction algorithm using artificial intelligence may be the SPADE (Spatial gene expression patterns by deep learning of tissue images) algorithm or the like. The SPADE algorithm is https: / / doi.org / 10.1101 / 2020.06.15.150698This is publicly known, and by reference, the entirety of it is incorporated into this application. The image feature extraction algorithm extracts features from each cluster of images of tissue staining, a transcription of tissue of interest, or the extracted features and provides SPADE gene information associated with them. The SPADE algorithm can use a pre-trained VGG16 model to extract, for example, 512 features per patch around each point, perform principal component analysis (PCA) to reduce the dimensionality of the features, and select principal components (PCs) to identify SPADE genes.
[0075] The molecular marker search unit 130 may include one or more of the above-described image characteristics analysis unit 131, correlation analysis unit 132, cell type analysis unit 133, gene ontology (GO) analysis unit 134, and clustering unit 135. In one example, after performing each of the above analysis units, the results can be combined to explore the distribution, efficacy, action, or physiological activity of genome-containing substances within the tissue of interest.
[0076] Another eighth configuration further includes a control transcript information receiving unit 111 that receives control transcript information of a second tissue of interest that has or has been diagnosed with the same disease or disorder as the tissue of interest, without providing genome-containing material to the search device or system of the sixth or seventh configuration described above. This configuration differs from the sixth configuration only in that it additionally uses a control transcript information receiving unit 111 that does not provide genome-containing material, and is otherwise identical to the configuration described above. That is, the control transcript information receiving unit 111 of the second tissue of interest can apply the origin search unit 120 and molecular marker search unit 130 of the sixth configuration, and the results can be received by the molecular marker search unit 130.
[0077] Specifically, the origin of the control transcript information is searched using an integrated reference transcript library that includes information on the first organoscopy reference transcript to which the second tissue of interest belongs, and information on the second organoscopy reference transcript to which the genome or substance containing the genome belongs (120'). The classified information can then be used in the molecular marker search unit 130 for the distribution, efficacy, action, or physiological activity of the genome-containing substance within the tissue of interest.
[0078] The control transcript information receiving unit 111 does not provide genome-containing material, so it may or may not include the origin search unit 120', but in either case, it is expected that genome-derived transcript information will not be present. The second tissue of interest may have the same origin as the tissue of interest in the manner described above.
[0079] The molecular marker search unit 130 described above can perform one or more of the following specific search methods on the control transcript information: for example, the image characteristics analysis unit 131, the correlation analysis unit 132, the cell type analysis unit 133, the gene ontology (GO) analysis unit 134, and the clustering unit 135. In one example, after performing each of the above analysis methods, the results can be used to search for the distribution, efficacy, action, or physiological activity of genome-containing substances within tissues of interest. For example, the molecular marker search unit 130 can extract higher or lower expressed transcript information from spatial transcript information (e.g., spatial transcript information tagged or untagged with origin information) within tissues of interest containing genome-containing substances, compared to the control transcript information. Conversely, it can extract higher or lower expressed transcript information from the control transcript information when compared with spatial transcript information (e.g., spatial transcript information tagged or untagged with origin information). For example, the results of cell type analysis of control transcription data can be compared with the results of cell type analysis of spatial transcription data of genome-containing substances within tissues of interest (e.g., spatial transcription data with or without origin information tagged), and the results can be used for the search of the molecular marker search unit 130. For example, one or more clusters for control transcription data can be compared with one or more clusters for spatial transcription data of genome-containing substances within tissues of interest (e.g., spatial transcription data with or without origin information tagged), and the results can be used for the search of the molecular marker search unit 130.
[0080] Furthermore, in another ninth embodiment, the sixth and / or eighth embodiment may further include a normal transcript information receiving unit 112 that acquires normal transcript information of a normal third tissue of interest that is free from or not induced by the disease or disorder and does not provide genome-containing material, and receives the acquired information. This embodiment differs from the sixth and eighth embodiments only in that it additionally uses the normal transcript information receiving unit 112, and can otherwise be implemented in the same manner as the embodiments described above. That is, the normal transcript information receiving unit 112 of the third tissue of interest can apply the origin search unit 120 and molecular marker search unit 130 of the sixth embodiment, and the results can be received by the molecular marker search unit 130.
[0081] Specifically, an integrated reference transcript library including first organoscopy reference transcript information to which the third tissue of interest belongs and second organoscopy reference transcript information to which the genome or genome-containing substance belongs is used to search for the origin of normal transcript information (120''), and the classified information can be used in the molecular marker search unit 130 for the distribution, efficacy, action, or physiological activity of the genome-containing substance within the tissue of interest.
[0082] The normal transcript information receiving unit 112 does not provide genome-containing material, so it may or may not include the origin search unit 120'', but in either case, it is expected that genome-derived transcript information will not be present. The third tissue of interest may have the same origin as the tissue of interest in the manner described above.
[0083] The molecular marker search unit 130 described above can perform one or more of the following specific search methods on the normal transcript information: for example, the image characteristics analysis unit 131, correlation analysis unit 132, cell type analysis unit 133, gene ontology (GO) analysis unit 134, and clustering unit 135. In one example, after performing each of the above analysis methods, the results can be used to search for the distribution, efficacy, action, or physiological activity of genome-containing substances within tissues of interest. For example, the molecular marker search unit 130 can extract transcription information that is expressed higher or lower in spatial transcript information (e.g., spatial transcript information tagged or untagged with origin information) within tissues of interest of genome-containing substances compared to normal transcript information, extract transcription information that is expressed higher or lower in control transcript information, or extract transcription information that is expressed higher or lower in normal transcript information when compared with spatial transcript information (e.g., spatial transcript information tagged or untagged with origin information) or control transcript information. For example, the molecular marker search unit 130 can compare the cell type analysis results of normal transcription information with the cell type analysis results of spatial transcription information of genome-containing substances (e.g., spatial transcription information with or without origin information tagged) and / or the cell type analysis results of control transcription information, examine changes in cell type, and use this for search. For example, the molecular marker search unit 130 can compare one or more clusters of normal transcription information with one or more clusters of spatial transcription information of genome-containing substances (e.g., spatial transcription information with or without origin information tagged) and / or one or more clusters of control transcription information, and use the results for search.
[0084] The molecular marker search unit 130 may further include comparing one or more pieces of information from the spatial transcript information, control transcript information, and normal transcript information of the tissue of interest to search for enhanced or suppressed molecular markers derived from the genome-containing substances.
[0085] Furthermore, in other tenth modes, An information receiving unit 110 provides a substance containing part or all of the genome not present in a diseased or diseased tissue of interest, acquires spatial transcription information of the tissue of interest, and receives the obtained information. The system includes a spatial transcription origin search unit 120 that classifies the spatial transcription information into transcription information derived from the tissue of interest and transcription information derived from the genome, using an integrated reference transcription library that includes first organoscopy reference transcription information to which the tissue of interest belongs and second organoscopy reference transcription information to which the genome or a substance containing the genome belongs. A device 200 is provided for classifying genome-derived transcripts from spatial transcript information. The device 200 has the advantage of higher classification accuracy compared to classifying genome-derived transcripts from spatial transcript information using first organism or second organism reference transcript information, respectively, without using an integrated reference transcript library.
[0086] Furthermore, the search device 100 or the classification device 200 may further include a tissue staining image receiving unit 150 that stains one or more of the tissues of interest, second tissues of interest, and third tissues of interest, and receives stained images of each tissue.
[0087] The aforementioned staining methods include, but are not limited to, ALP assay (alkaline phosphate assay), Sirius red staining, Alcian blue staining, pH map, H&E staining, trichrome staining, PAS (Periodic acid-Schiff) staining, and immunohistochemical staining. The results of the stained tissue images may be used as a complementary or supplementary means to verify the search results according to the present invention or to extract information useful when searching for molecular markers. In one example, the stained tissue images can be used as is, or each characteristic of the image can be extracted by applying an artificial intelligence-based image characteristic extraction algorithm, such as the SPADE (Spatial gene expression patterns by deep learning of tissue images) algorithm, and SPADE gene information associated with the extracted characteristics can be obtained.
[0088] The present invention will be described in more detail below with reference to examples and analytical examples. However, the following examples and analytical examples are merely illustrative of the present invention, and the scope of the present invention is not limited to these. [Examples]
[0089] method Animal experimentation and tissue acquisition This animal experiment was approved by Bundang Seoul National University Hospital under IACUC approval code BA-2211-355-002. Normal bone marrow-derived hMSCs were used in Lonza TMThe mice were manufactured using [method / technology]. Next, three groups of 9-week-old C57BL / 6 male mice were prepared: normal mice ("Nor"), mice with pulmonary fibrosis as a control group ("Con"), and experimental mice with pulmonary fibrosis that received hMSC injection via the tail vein ("Exp"). Pulmonary fibrosis was induced in both the control group ("Con") and the experimental group ("Exp") by administering bleomycin for 3 weeks prior to sacrificing. Next, optimal cutting temperature (OCT) blocks (Scigen 4586, USA) were prepared according to the Visium Spatial Protocols-Tissue Preparation Guide (document CG000240). The OCT block for "Exp" was prepared 6 hours after human stem cell injection. Subsequently, two tissue slices were prepared for each OCT mold in the group, and H&E-stained tissue slices and fresh frozen tissue slices for the Visium ST library were obtained (Figure 6).
[0090] Acquisition of ST Library Tissue fragments were fixed, stained, and permeabilized in accordance with the Visium Spatial Protocols-Spatial Gene Expression Imaging Guide (document CG000241), along with a tissue optimization (TO) step. mRNA present in the tissue was captured via the poly-A tail, and subsequent cDNA was barcoded and amplified via PCR (polymerase chain reaction) to obtain sufficient cDNA to reconstruct the library. Library quality was measured and evaluated using qPCR (quantitative PCR) and Agilent Technologies 4200 TapeStation, respectively. Finally, the library was sequenced using the Illumina HiSeq platform according to the guidelines provided in the user guide.
[0091] Furthermore, an integrated reference transcript library was obtained by merging GRCh38, a representative human reference, and mm10, a representative mouse reference, using SpaceRanger (ver.2.0.1)mkref. Then, SpaceRanger counts were performed on each of the Nor, Con, and Exp samples to generate processed ST libraries, which were then mapped to the integrated reference transcript library.
[0092] The number of raw genes at a single site at a given sequencing depth is inversely proportional to the total RNA production at that site. To address this, gene counts were adjusted using various methods, including Trimmed Mean of M-Values (TMM) and fractile normalization
[11] . Furthermore, since genes with high counts can overestimate their activity, gene counts are generally log-transformed. This normalization process can be performed in Python using the scanpy.pp.normalize_total and scanpy.pp.log1p functions. After normalizing gene counts, unreliably large count values for some RNA transcripts were corrected (Figure 7, 1st column). %human, the percentage of human transcripts relative to all transcripts at a single site, was investigated by sample and normalization. The results showed that normalized %human was even higher in the "Exp" sample compared to the unnormalized sample (Figure 7, bottom right), but showed only minimal effect in the "Nor" and "Con" samples. The results above demonstrate that normalization increases the statistical power for detecting human transcripts.
[0093] Therefore, in subsequent analyses, only human transcripts and %human were considered after the normalization process. To address misidentified human transcripts, various thresholds for %human were tested (Figure 8). Ultimately, the threshold was defined as the sum of 10 times the standard deviation of %human plus the mean in the "Con" sample for two reasons: firstly, this threshold is just above the limit of %human where no spots exceed it in the "Nor" and "Exp" samples (Figure 7, third column); and secondly, when this threshold was applied, each spot classified as a human cell in the "Exp" sample showed a histological pattern suitable for comparison in subsequent cell type matching analyses.
[0094] Figure 9 shows H&E-stained tissue slices. It can be seen that pulmonary fibrosis is present in Con and Exp, but not in Nor.
[0095] Figure 10 shows the number of transcripts per spot (all_transcipts), the number of mouse transcripts per spot (all_mouse_transcripts), and the number of human transcripts per spot (all_human_transcripts) for the three samples "Nor," "Con," and "Exp," respectively. In other words, Figure 10 is an image showing the total number of transcripts (1st column), the number of mouse transcripts (2nd column), and the number of human transcripts (3rd column). The 4th column shows each tissue in two rows, with the mean + 10σ of the number of transcripts per spot defined as the cutoff. In "is_human," yellow spots are spots that exceeded the cutoff value, and purple spots are spots that did not exceed the cutoff value. Here, human transcripts include human transcripts that can only be detected in "Exp" on a transcription scale from 0 to 100, and they represent a considerable number of spots. The distribution of %human on ST in the "Exp" sample can be considered as the distribution of administered hMSCs (4th column, Figure 10). Therefore, it was confirmed that hMSC cell injection led to the statistically significant detection of human transcripts on the Visium platform.
[0096] Clustering analysis to identify administered hMSCs in lung tissue Figure 11 shows the clustering results based on mouse transcripts (left) and human transcripts (right) for each of the three samples: "Nor," "Con," and "Exp."
[0097] Figure 12 shows the results of clustering analysis performed on the combined spots of the three samples, namely "Nor," "Con," and "Exp." Figure 13(a) shows the clustering results for each of the three samples, namely "Nor," "Con," and "Exp" (DimPlot (upper) and SpatialDimplot (lower) for the three samples). The VlnPlot shows the percentage of human transcripts (%human) among all transcripts by cluster number (horizontal axis). Interestingly, cluster 5 showed a significantly larger number of human transcripts than the other clusters (Figure 13(b)). Figure 13(c) shows the ratio of specific cluster numbers for each of the three samples, where cluster 5 was almost absent in the "Nor" and "Con" samples. DEG and GO analyses were performed on cluster 5, and the results are shown in Figure 13(d). The specific DEGs of cluster 5, i.e., the top 20 spatially abundant genes regulated upward, were Lcn2 (highest log FC), Msln, Chil3, C3, Upk3b, Spp1, Col4a1, Serpina3n, Wfdc21, Col3a1, Wfdc17, Gm13889, Chil1, Mgp, Sftpd, Slpi, Ctsc, Fmo2, Scd1, and Napsa (in order). Furthermore, GO analysis results for biological processes (BP), cellular components (CC), and molecular functions (MF) in these genes were associated with immune responses, collagen-containing extracellular matrix (ECM), and peptidase activity. This suggests that immune responses, collagen-containing extracellular matrix (ECM), and peptidase activity are generated by the distribution of hMSCs (see results in Figure 13(d)).
[0098] DEG acquisition associated with the distribution of hMSCs Data integration was performed using FindIntegrationAnchors and IntegrateData in Seurat (version 4.3.0). Subsequently, human genes were removed from the integrated Seurat individuals, and then spatial clustering analysis was performed. This removal process was considered meaningful to prevent bias towards "Exp," which contains only true human transcripts. However, this removal process was not performed in other analyses. The top 20 spatially abundant genes (adjusted p-value < 0.05, log FC order) in specific clusters were obtained from the Wilcoxon Rank Sum test using Seurat's FindAllMarkers. Gene ontology (GO) plots from clusterProfiler::enrichGO (ver. 4.6.2) were explored in R.GO plots, yielding approximately five categories based on biological processes (BP), cellular components (CC), and molecular functions (MF).
[0099] To perform comparative analysis between samples, we obtained differentially expressed genes (DEGs) using the Wilcoxon rank-sum test with Seurat's FindAllMakers. Next, we explored the top 20 DEGs (adjusted p-value < 0.05, log-FC order) using GO analysis.
[0100] To obtain genes spatially associated with %human, Spearman correlation coefficients between %human and the scaled expression of each gene were calculated only for the "Exp" sample. Subsequently, the top 20 mouse genes spatially associated with %human (adjusted p-value < 0.05, Spearman correlation order) were explored using GO analysis.
[0101] We obtained differentially expressed genes (DEGs) between "Con" and "Exp". These DEGs included all human and mouse genes, but the top six DEGs were all mouse genes (see Figures 15 to 18). To obtain genes spatially associated with hMSC distribution, we searched for genes spatially associated with %human in the "Exp" sample.
[0102] Comparative analysis was performed between each sample to obtain differentially expressed genes (DEGs) using Seurat's FindAllMakers. Subsequently, GO analysis was performed to identify the top 20 positively related DEGs (adjusted p-value < 0.05, Log FC order) and the bottom 20 negatively related DEGs (see results in Figures 19(a) and (b)). Six of the top 20 DEGs were human genes. To evaluate the molecular processes of pulmonary fibrosis in the "Exp" sample, each mouse gene with a positive spatial relationship was selected from the top 20 DEGs, and GO analysis was performed on them.
[0103] The biological pathway of each of the aforementioned top genes was the "transmembrane receptor protein serine / threonine kinase signaling pathway" (see Figure 19(c)). In addition, there were overlapping mouse genes between the top 20 spatially abundant genes in cluster 5 and each gene that was spatially associated with %human, and these were Lcn2 and Col4al (see Figures 13 and 19).
[0104] CellDART was used to identify cell type populations in pulmonary fibrosis models associated with hMSC treatment.
[0105] To confirm the distribution of %human and the spatially associated cell types, cell type inference was performed using domain adaptation (CellDART) of single-cell and spatial transcript data
[12] . Single-cell RNA sequencing (scRNA-seq) references were generated using publicly available mouse lung scRNA-seq references (GSE124872). Spearman correlation coefficients between CellDART scores for each mouse cell type and %human were calculated using only "Exp" samples.
[0106] More specifically, after preparing a single-cell RNA sequencing (scRNA-seq) reference (GSE124872) of mouse lung, cellular-level deconvolution was performed on the "Exp" sample using CellDART
[12] (Figure 20(a)). In the treated lung tissue, the Spearman correlation between %human and CellDART scores for cell types was calculated (Figure 20(b)). The results showed that endothelial cells were the most associated with %human, and epithelial cells were the least associated with %human (Figure 20(b)). This is consistent with the fact that intravenously injected cells were considered to be physically trapped in the narrow microvessels of the alveoli during blood circulation
[13] .
[0107] The variable "is_endothelial" was defined to indicate whether a spot in the "Exp" sample was composed of endothelial cells, using the mean threshold of CellDART scores for endothelial cells plus one standard deviation. Next, spots labeled "is_endothelial" (simply referred to as "endothelial spots") were compared regardless of whether they were designated "is_human," and DEGs were obtained using the same method as in previous DEG analyses. As a result, only Apoe, Col1a1, and Hnrnpab were significantly abundant in endothelial spots with low %human (adjusted p-value < 0.05). Of these, the Apoe and Col1a1 (activated fibroblast marker) genes are known to be located in fibrotic sites around blood vessels
[10] . In contrast, along with 43 human genes, mt-Co2, Lcn2, Gm42418, Chil1, Scd1, mt-Nd1, and mt-Nd2 were significantly more abundant in endothelial spots with a high percentage of human (adjusted p-value < 0.05). Interestingly, Lcn2, which was found in two previous analyses, also appeared in this analysis, supporting the overall analysis.
[0108] Stopover analysis identified cell types co-localized by %human and interspecies LR interactions.
[0109] STopover was introduced as a method for measuring the spatial colocalization of two continuous variables using ST barcodes
[16] . The algorithm calculates a Jacquard exponent that evaluates |A∩B|, which, given sets A and B, can be divided into |A∪B|. The exponent can be calculated either globally (J_comp) or locally (J_local), the latter being determined by the concatenated components (CC). The parameters of STopover were set to base. STopover was also used to explore the spatial relationships of cell types. Here, CellDART scores obtained by multiplying %human and (100-%human) were used to represent the spatial distribution of human and mouse cells, respectively.
[0110] Furthermore, we investigated ligand-receptor interactions between different species. Existing ligand-receptor databases, including CellPhoneDB
[38] and CellTalkDB
[39] , generally focus primarily on individual species such as humans or mice. Research on interspecies ligand-receptor interactions was an unexplored area. For this reason, we identified shared ligand-receptor (LR) pairs between humans and mice in CellTalkDB. Next, we transformed the shared LR pairs using a homology database of human and mouse genes derived from the Mouse Genomic Informatics (MGI) database. These pairs were replaced with two distinct types of pairs: one containing a human ligand and a mouse receptor, and the other containing a mouse ligand and a human receptor. Finally, the colocalization of such LR pairs was evaluated using STopover based on normalized gene expression using scanpy, Python.
[0111] The analysis of cell-cell interactions between different biological systems, such as mouse and human cells, can be meaningful for two reasons: firstly, because biological systems are relatively clearly distinguishable, cell-cell interactions between different systems can be more easily identified, which is particularly important in cancer-immune interactions, host-microbial community interactions, and host-pathogen interactions. Secondly, unique forms of cell-cell interactions, such as interspecies interactions, can lead to new discoveries regarding isoforms and biomarkers. Various algorithms have been developed to analyze cell-cell interactions, including CCCExplorer, CellChat, ICELLNET, and CellPhoneDB
[14] . Nevertheless, many of these algorithms have the limitation of ignoring spatial proximity. However, recent studies have attempted to solve this problem by applying machine learning algorithms to ST data
[15] . As a result of these efforts, we have introduced STopover, an algorithm that utilizes phase analysis techniques to assess the co-enrichment of two variables (e.g., cell type score, gene expression, or %human) in ST data
[16] .
[0112] When performing STopover, endothelial cells showed the highest colocalization with %human among the six cell types (Jackard index = 0.304) (Figures 21, 22). This result is consistent with previous Spearman correlation analysis performed based on CellDART (Figure 20(b)). We also explored ligand-receptor (LR) interactions between different species using STopover and CellTalkDB. CellTalkDB was modified for this study analyzing interspecies interactions. As a result, the COL1A2 (human) and Cd93 (mouse) pair showed the highest Jackard index, which means that the spatial colocalization of this LR pair was the greatest. Interestingly, Cd93 is a well-known marker gene for endothelial cells [17, 18]. This implies spatial colocalization and potential interaction between mouse endothelial cells and human stem cells. Therefore, the STopover results confirm that human stem cells are spatially aligned with mouse endothelial cells.
[0113] Cellular typological changes due to pulmonary fibrosis were observed in comparative cell typology analysis between "Nor," "Con," and "Exp."
[0114] For three samples, "Nor," "Con," and "Exp," a wide variety of mouse cell types were used, such as aDCs, adipocytes, B cells, basophils, CD4 T cells, and CD8 T cells (Figures 23 to 24); cDCs, chondrocytes, DCs, endothelial cells, eosinophils, and erythrocytes (Figures 25 to 26); monocytes, mv. endothelial cells, myocytes, neurotrophils, NK cells, and NKT cells (Figures 27 to 28); osteoblasts, pDCs, pericytes, sebocytes, and skeletal muscle cells. CellDART was performed on muscle, smooth muscle (Figures 29 to 30), Tgd cells, Th1 cells, Th2 cells, Tregs, immune score, and stroma score (Figures 31 to 32), and the results are shown in the respective figures.
[0115] From the comparison of the three samples mentioned above, the following changes were observed in "Con" compared to "Nor": [Table 1]
[0116] DEG comparison between "Nor", "Con", and "Exp". "Nor" vs "Con" DEGs in "Nor" and "Con" were examined using volcano plots. Log FC>0 indicates genes with high expression in "Con," and Log FC<0 indicates genes with high expression in "Nor." As a result, increased expression of ECM (Extracellular matrix) related genes was observed in "Con" (Figure 14(a), (c)-(d)). Furthermore, when the top 6 DEGs in "Nor" were searched compared to "Con," blood flow-related genes (e.g., Hba-a1, Hba-a2, Hbb-bs, and Hbb-bt) appeared in "Nor" (Figures 15 and 16). This is presumed to be evidence of vascular deficiency in "Con" due to fibrosis.
[0117] "Con" vs "Exp" The DEGs of "Con" and "Exp" were compared using volcano plots. Log FC>0 indicated genes with high expression in "Exp," and Log FC<0 indicated genes with high expression in "Con." As a result, blood flow activity suggesting a recovery mechanism was observed in "Exp" (Figure 14(b), (e)-(f)). In addition, blood flow-related genes (e.g., Hba-a1, Hba-a2, Hbb-bs, and Hbb-bt) reappeared in the "Exp" samples (Figure 14(b), (e)-(f), Figure 17, Figure 18). This suggests that the injected hMSCs restore blood flow or blood vessels, exhibiting a therapeutic effect against pulmonary fibrosis, at least at the molecular level.
[0118] "Nor" vs "Exp" The DEGs of "Nor" and "Exp" were compared using volcano plots. Log FC > 0 indicates genes with high expression in "Exp," and Log FC < 0 indicates genes with high expression in "Nor." As a result, increased ECM expression was observed in "Exp" (Figure 29(a), (b)).
[0119] Comparisons between the aforementioned samples confirmed that pulmonary fibrosis exhibited improved levels of extracellular moisture (ECM) and immune response.
[0120] Furthermore, Bpifa1, a common top DEG in "Con" compared to "Nor" and "Exp," is known to be associated with innate immunity in lung infection models (30) (see Figures 16 and 17). Bpifa1 is also known to be associated with the mucosal environment of lung tissue. The reason why "humoral immune response" mainly appeared in the GO term of spatially abundant genes in spatial cluster 5 (see Figure 13(d)) and the reason why certain genes were abundant in "Con" compared to "Nor" and "Exp" (see Figures 14(d) and (e)) may be related to the aforementioned Bpifa1. Therefore, a decrease in the expression of this gene may be associated with the therapeutic effect of injected human stem cells.
[0121] Furthermore, compared to "Nor," Spp1 found in "Con" is well known to be expressed in adjacent, invasive, or angiogenesis-associated macrophages (40). This also suggests a link to innate immunity in relation to the efficacy of hMSCs.
[0122] From the results mentioned above, we were able to discover the following three things.
[0123] Firstly, we were able to distinguish between RNA transcripts originating from different genomes and obtain their spatial distribution (Figure 10).
[0124] Secondly, through spatial clustering analysis (Figure 13), DEG analysis (Figure 14), and genes spatially associated with %human (Figure 19), we were able to obtain RNA transcripts of interest, i.e., genes significantly associated with hMSCs.
[0125] Finally, the injected human stem cells showed a strong correlation with endothelial cells, but a reverse correlation with epithelial cells.
[0126] The market for nucleic acid-based therapeutics is expanding with numerous drugs currently under development. Clinical trials have been conducted on stem cells
[19] , CAR-T
[20] , and nucleic acid-containing exosomes
[21] . Nevertheless, preclinical studies have reported difficulties in evaluating the molecular mechanisms of such drugs, particularly those that interact with target tissue cells. Consequently, there is a pressing need for evaluation methods for such drugs.
[0127] Previously, a method was developed to identify spatially associated molecular markers of injected drugs based on ST
[10] . This method has elucidated enhanced permeability and retention (EPR) related markers and can be applied to a variety of therapeutic agents labeled with fluorescent dyes. Nevertheless, labeling therapeutic agents with dyes has several disadvantages, including the possibility of fluorescent dye labeling failure, alteration of drug properties, and separation or degradation of the dye [22-24]. It is also applicable to methods widely used to evaluate cell therapeutic agents using tracer labeling. In this application, we introduce a label-free spatial mapping of exogenous nucleic acids using spatial transcripts. Previous studies have attempted to map RNA transcripts of other species in application areas such as xenotransplantation
[25] , host-microbiome mapping
[26] , host-virus mapping
[27] and engineered oligonucleotides
[28] . While these studies demonstrated successful analysis of mixed transcripts, they did not use the distinguished mapping technique according to the present invention to identify the spatial distribution of cell therapeutic agents in tissue or to estimate their mode of action.
[0128] Furthermore, this method can also be applied to cell therapies of host origin, such as when introducing transformation with genes not present in the host
[29] . The approach proposed in this application has the potential to extend the application of ST to spatial analysis of therapeutic agents containing exogenous nucleic acids. By using this approach, a more comprehensive understanding of the underlying mechanisms of therapeutic effects of such agents can be obtained, which may lead to important insights for optimizing cell therapy for a variety of diseases. In addition to simple distribution analysis, spatial transcription analysis of stem cell therapy in this invention can reveal the effects of transcription levels on pulmonary fibrosis tissue. Injected human stem cells showed upward regulation of many hemoglobin genes, including Hba-a1, Hba-a2, Hbb-bs, and Hbb-bt, which was upwardly regulated in "Exp" compared to the "Con" group (Figures 14 and 18). Furthermore, when compared to the "Con" group, the "Exp" group showed downward regulation of "collagen-containing extracellular matrix" and "humoral immune response" genes (e.g., Bpifa1
[30] ) (Figures 13, 14, and 17). The molecular changes observed in the "Exp" group were similar to those observed when comparing the "Nor" group to the "Con" group (e.g., Hba-a1, Hba-a2, Hbb-bs, and Hbb-bt) (Figures 15 and 16), which supports the assumption that the molecular changes are due to the effects of injected human stem cells.
[0129] Several issues must be considered when using spatial analysis of cells administered using ST. The analytical process presented here relied on less significant false-positive human transcripts in "Nor" and "Con." The occurrence of false-positive human transcripts stemmed from homology between humans and mice. False-negative human transcripts must also be addressed. In the "Exp" sample, VIM and Vim appeared in the top 20 genes spatially related to %human (Figure 19(a)). Given that VIM is a mesenchymal stem cell-related gene [31, 32], it is highly probable that the Vim values are misclassified mouse transcripts. A possible solution to the erroneous detection problem presented here is to adopt long-read sequencing to completely manipulate the differences in gene sequences. Another possible method is to develop computer algorithms that can reliably remove or correct false transcripts by manipulating spatial proximity
[33] , histological morphology [34, 35], previous datasets [12, 36, 37] or other domain knowledge.
[0130] This invention successfully demonstrated the ability of ST to map administered cell therapy agents without the use of fluorescent labeling. A key result was the spatial distribution of human stem cells injected into pulmonary fibrotic tissue, and the identification of genes associated with this distribution.
[0131] Overall, this invention demonstrates the usefulness of ST for spatial analysis of therapeutic agents containing exogenous nucleic acids.
[0132] statistics R (ver 4.0.5) and Python (ver 3.7.12) were used as programming languages. Other tools used included Seurat (ver 4.0.2), scanpy (ver 1.9.1), and SpaceRanger (ver 2.0.1). For reference, SpaceRanger was used with GRCh38 (Homo sapiens) and mm10 (Mus musculus). Differentially expressed genes (DEGs) were searched for by classifying them against the fold change (FC) value for all genes with an adjusted p-value less than 0.05. When plotting gene ontology (GO) plots, 20 host mouse genes were selected and analyzed unless otherwise specified. Unless otherwise specified in Table 2, mediating variables were set to baseline values.
[0133] [Table 2]
[0134] Abbreviation BBB: Blood-brain barrier; BP: Biological process; CAR-T / NK: Chimeric antigen receptor-T cell or NK cell therapy; CC (GO analysis): Cellular components; CC (Stopover): Linked components; CellDART: Cell type inference by domain adaptation of single-cell and spatially transcribed data; DEG: Differentiatedly expressed genes; ECM: Extracellular matrix; FC: Fold changes; GO: Gene ontology; hMSC: Human mesenchymal stem cells; LR: Ligands and receptors; MGI: Mouse genomic informatics; MOA: Mechanism of action; OCT: Optimal cutting temperature; PCR: Polymerase chain reaction; qPCR: Quantitative polymerase chain reaction; scRNA-seq: Single-cell RNA sequencing; ST: Spatially resolved transcribed; TMM: Trimmed average of M-values; TO: Tissue optimization.
[0135] 1. Saez-Ibanez AR, Upadhaya S, Partridge T, Shah M, Correa D, Campbell J. Landscape of cancer cell therapies: trends and real-world data. Nat Rev Drug Discov. 2022; 21: 631-2.
[0136] 2. Ghobadinezhad F, Ebrahimi N, Mozaffari F, Moradi N, Beiranvand S, Pournazari M, et al. The emerging role of regulatory cell-based therapy in autoimmune disease. Front Immunol. 2022; 13: 1075813.
[0137] 3. Hossein-Khannazer N, Torabi S, Hosseinzadeh R, Shahrokh S, Asadzadeh Aghdaei H, Memarnejadian A, et al. Novel cell-based therapies in inflammatory bowel diseases: the established concept, promising results. Hum Cell. 2021; 34: 1289-300.
[0138] 4. Pradhan AU, Uwishema O, Onyeaka H, Adanur I, Dost B. A review of stem cell therapy: An emerging treatment for dementia in Alzheimer’ and Parkinson’s disease. Brain Behav. 2022; 12: e2740.
[0139] 5. Xia Y, Rao L, Yao H, Wang Z, Ning P, Chen X. Engineering Macrophages for Cancer Immunotherapy and Drug Delivery. Adv Mater. 2020; 32: e2002054.
[0140] 6. Zhuang WZ, Lin YH, Su LJ, Wu MS, Jeng HY, Chang HC, et al. Mesenchymal stem / stromal cell-based therapy: mechanism, systemic safety and biodistribution for precision clinical applications. J Biomed Sci. 2021; 28: 28.
[0141] 7. Brooks A, Futrega K, Liang X, Hu X, Liu X, Crawford DHG, et al. Concise Review: Quantitative Detection and Modeling the In Vivo Kinetics of Therapeutic Mesenchymal Stem / Stromal Cells. Stem Cells Transl Med. 2018; 7: 78-86.
[0142] 8. Kreyling WG, Abdelmonem AM, Ali Z, Alves F, Geiser M, Haberl N, et al. In vivo integrity of polymer-coated gold nanoparticles. Nat Nanotechnol. 2015; 10: 619-23.
[0143] 9. Bae S, Choi H, Lee DS. Discovery of molecular features underlying the morphological landscape by integrating spatial transcriptomic data with deep features of tissue images. Nucleic Acids Res. 2021; 49: e55.
[0144] 10. Park J, Choi J, Lee JE, Choi H, Im HJ. Spatial Transcriptomics-Based Identification of Molecular Markers for Nanomedicine Distribution in Tumor Tissue. Small Methods. 2022; 6: e2201091.
[0145] 11. Robinson MD, Oshlack A. A scaling normalization method for differential expression analysis of RNA-seq data. Genome Biol. 2010; 11: R25.
[0146] 12. Bae S, Na KJ, Koh J, Lee DS, Choi H, Kim YT. CellDART: cell type inference by domain adaptation of single-cell and spatial transcriptomic data. Nucleic Acids Res. 2022; 50: e57.
[0147] 13. Jung KO, Kim TJ, Yu JH, Rhee S, Zhao W, Ha B, et al. Whole-body tracking of single cells via positron emission tomography. Nat Biomed Eng. 2020; 4: 835-44.
[0148] 14. Armingol E, Officer A, Harismendy O, Lewis NE. Deciphering cell-cell interactions and communication from gene expression. Nat Rev Genet. 2021; 22: 71-88.
[0149] 15. Fischer DS, Schaar AC, Theis FJ. Modeling intercellular communication in tissues using spatial graphs of cells. Nat Biotechnol. 2023; 41: 332-6.
[0150] 16. Bae S, Lee H, Na KJ, Lee DS, Choi H, Kim YT. STopover captures spatial colocalization and interaction in the tumor microenvironment using topological analysis in spatial transcriptomics data. bioRxiv. 2022: 2022.11. 16.516708.
[0151] 17. Ianevski A, Giri AK, Aittokallio T. Fully-automated and ultra-fast cell-type identification using specific marker combinations from single-cell transcriptomic data. Nat Commun. 2022; 13: 1246.
[0152] 18. De Falco A, Caruso F, Su XD, Iavarone A, Ceccarelli M. A variational algorithm to detect the clonal copy number substructure of tumors from scRNA-seq data. Nat Commun. 2023; 14: 1074.
[0153] 19. Ji Y, Hu C, Chen Z, Li Y, Dai J, Zhang J, et al. Clinical trials of stem cell-based therapies for pediatric diseases: a comprehensive analysis of trials registered on ClinicalTrials.gov and the ICTRP portal site. Stem Cell Res Ther. 2022; 13: 307.
[0154] 20. Ivica NA, Young CM. Tracking the CAR-T Revolution: Analysis of Clinical Trials of CAR-T and TCR-T Therapies for the Treatment of Cancer (1997-2020). Healthcare (Basel). 2021; 9.
[0155] 21. Zhang Y, Liu Q, Zhang X, Huang H, Tang S, Chai Y, et al. Recent advances in exosome-mediated nucleic acid delivery for cancer therapy. J Nanobiotechnology. 2022; 20: 279.
[0156] 22. Srivastava AK, Bulte JW. Seeing stem cells at work in vivo. Stem cell reviews and reports. 2014; 10: 127-44.
[0157] 23. Horan PK, Melnicoff MJ, Jensen BD, Slezak SE. Fluorescent cell labeling for in vivo and in vitro cell tracking. Methods Cell Biol. 1990; 33: 469-90.
[0158] 24. Choi H, Lee DS. Illuminating the physiology of extracellular vesicles. Stem Cell Res Ther. 2016; 7: 55.
[0159] 25. Ni Z, Prasad A, Chen S, Halberg RB, Arkin LM, Drolet BA, et al. SpotClean adjusts for spot swapping in spatial transcriptomics data. Nat Commun. 2022; 13: 2971.
[0160] 26. Galeano Nino JL, Wu H, LaCourse KD, Kempchinsky AG, Baryiames A, Barber B, et al. Effect of the intratumoral microbiota on spatial and cellular heterogeneity in cancer. Nature. 2022; 611: 810-7.
[0161] 27. Sounart H, Lαzαr E, Masarapu Y, Wu J, Vαrkonyi T, Glasz T, et al. Dual spatially resolved transcriptomics for SARS-CoV-2 host-pathogen colocalization studies in humans. bioRxiv. 2022: 2022.03. 14.484288.
[0162] 28. Ben-Chetrit N, Niu X, Swett AD, Sotelo J, Jiao MS, Stewart CM, et al. Integration of whole transcriptome spatial profiling with protein markers. Nat Biotechnol. 2023.
[0163] 29. 10x Genomics. Build a Custom Reference. Retrieved from https: / / support.10xgenomics.com / single-cell-gene-expression / software / pipelines / latest / using / tutorial_mr.
[0164] 30. Tsou YA, Tung MC, Alexander KA, Chang WD, Tsai MH, Chen HL, et al. The Role of BPIFA1 in Upper Airway Microbial Infections and Correlated Diseases. Biomed Res Int. 2018; 2018: 2021890.
[0165] 31. Zhang L, Wei Y, Chi Y, Liu D, Yang S, Han Z, et al. Two-step generation of mesenchymal stem / stromal cells from human pluripotent stem cells with reinforced efficacy upon osteoarthritis rabbits by HA hydrogel. Cell Biosci. 2021; 11: 6.
[0166] 32. Cooper TT, Sherman SE, Bell GI, Ma J, Kuljanin M, Jose SE, et al. Characterization of a Vimentin(high) / Nestin(high) proteome and tissue regenerative secretome generated by human pancreas-derived mesenchymal stromal cells. Stem Cells. 2020; 38: 666-82.
[0167] 33. Long Y, Ang KS, Li M, Chong KLK, Sethi R, Zhong C, et al. Spatially informed clustering, integration, and deconvolution of spatial transcriptomics with GraphST. Nat Commun. 2023; 14: 1155.
[0168] 34. Comiter C, Vaishnav ED, Ciapmricotti M, Li B, Yang Y, Rodig SJ, et al. Inference of single cell profiles from histology stains with the Single-Cell omics from Histology Analysis Framework (SCHAF). bioRxiv. 2023: 2023.03. 21.533680.
[0169] 35. Parreno-Centeno M, Malagoli Tagliazucchi G, Withnell E, Pan S, Secrier M. A deep learning and graph-based approach to characterise the immunological landscape and spatial architecture of colon cancer tissue. bioRxiv. 2022: 2022.07. 06.498984.
[0170] 36. Li H, Zhou J, Li Z, Chen S, Liao X, Zhang B, et al. A comprehensive benchmarking with practical guidelines for cellular deconvolution of spatial transcriptomics. Nat Commun. 2023; 14: 1548.
[0171] 37. Miller BF, Huang F, Atta L, Sahoo A, Fan J. Reference-free cell type deconvolution of multi-cellular pixel-resolution spatially resolved transcriptomics data. Nat Commun. 2022; 13: 2339.
[0172] 38. Efremova M, Vento-Tormo M, Teichmann SA, Vento-Tormo R. CellPhoneDB: inferring cell-cell communication from combined expression of multi-subunit ligand-receptor complexes. Nat Protoc. 2020; 15: 1484-506.
[0173] 39. Shao X, Liao J, Li C, Lu X, Cheng J, Fan X. CellTalkDB: a manually curated database of ligand-receptor interactions in humans and mice. Brief Bioinform. 2021; 22.
[0174] 40. Cheng, S. et al. A pan-cancer single-cell transcriptional atlas of tumor infiltrating myeloid cells. Cell 184, 792-809 e723 (2021)
Claims
1. In order to explore the distribution, efficacy, action, or physiological activity of genome-containing substances within tissues of interest, To provide a tissue of interest that has a disease or disease, or has been induced, with a substance containing part or all of the genome that is not present in the tissue of interest, and to obtain spatial transcription information of the tissue of interest, An integrated reference transcript library is used to classify the spatial transcript information into transcripts derived from the tissue of interest and transcripts derived from the genome, which include first organoscopy reference transcript information to which the tissue of interest belongs and second organoscopy reference transcript information to which the genome or a substance containing the genome belongs. A search method comprising exploring the distribution, mechanism, efficacy, action, or physiological activity of the genome-containing substance within a tissue of interest based on the classified information.
2. The search method according to claim 1, wherein the search includes extracting molecular markers spatially related to the ratio of the number of genome-derived transcripts to the total number of spatial transcripts in the tissue of interest.
3. The search method according to claim 1, wherein, based on the classification, each origin information is tagged to the spatial transcription information of the tissue of interest.
4. The search method according to claim 1, wherein the search includes dividing the spatial image of the classified information or the entire spatial image into one or more clusters.
5. The search method according to claim 4, comprising performing an analysis on one or more clusters using an individual cluster image characteristic extraction algorithm, a correlation analysis between cluster image intensity and gene expression level, a cell type analysis algorithm, and / or gene ontology analysis.
6. The search method according to claim 1, wherein the search includes performing a cell type analysis algorithm or gene ontology (GO) analysis on the spatial transcription information of the classified information or the overall spatial transcription information.
7. The search method according to claim 6, wherein the cell type analysis algorithm is Fisher's exact test, maximum probability assessment, domain adaptive classification, logistic regression analysis, or negative binomial regression analysis algorithm.
8. The search method according to claim 6, wherein the cell type analysis algorithm is the CellDART algorithm, spSeudoMap algorithm, xCell algorithm, STopover algorithm, RCTD algorithm, Seurat algorithm, celltypist algorithm, cell2location algorithm, Scanorama algorithm, SPOTlight algorithm, DSTG algorithm, CellTrek algorithm, sc-type algorithm, TIMER algorithm, CIBERSORT algorithm, MCP-counter algorithm, Quantiseq algorithm, EPIC algorithm, or scanpy's injest algorithm.
9. The aforementioned method, The search method according to claim 1, further comprising obtaining control transcript information of a second tissue of interest that has or has the same disease or disorder as the tissue of interest, but where no genome-containing material is provided, and using this information for the search.
10. The search method according to claim 9, further comprising performing a correlation analysis between the degree of gene expression based on the control transcription information and the degree of gene expression based on the spatial transcription information.
11. The search method according to claim 10, wherein the correlation analysis is performed using DEG (Differently Expressed Genes), correlation analysis by calculation of the correlation coefficient, or an image similarity evaluation algorithm.
12. The search method according to claim 11, wherein the correlation coefficient calculation is a Pearson correlation coefficient, Spearman correlation coefficient, or Kendall correlation coefficient calculation.
13. The aforementioned method, The search method according to claim 1, further comprising obtaining normal transcript information of a normal third tissue of interest that is free from or not induced by the aforementioned disease and does not provide genome-containing material, and using this information for the search.
14. The discovery method according to any one of claims 1 to 13, wherein the provision of the genome-containing substance to the tissue of interest includes either directly providing the substance to the tissue of interest or administering it to a target body by systemic administration and then distributing it to the tissue of interest.
15. The search method according to any one of claims 9 to 13, further comprising comparing one or more transcription information from spatial transcription information, control transcription information, and normal transcription information of the tissue of interest to search for enhanced or suppressed molecular markers derived from genome-containing substances.
16. The search method according to any one of claims 1 to 13, further comprising staining one or more of the tissue of interest, secondary tissue of interest, and tertiary tissue of interest to obtain a stained image of the tissue.
17. The search method according to claim 16, further comprising dividing the stained image into one or more clusters, and performing analysis on one or more clusters using an individual cluster image characteristic extraction algorithm, a correlation analysis between the image intensity of the cluster and the degree of gene expression of genome-containing substances, a cell type analysis algorithm, and / or gene ontology analysis.
18. The search method according to any one of claims 1 to 13, wherein the spatial transcript information is the number of transcript RNA reads and RNA types in a spot having spatial coordinates within the tissue of interest.
19. The search method according to any one of claims 1 to 13, wherein the genome-containing substance is a cell, extracellular vesicle, lymphocyte, microbiome, vector, or virus containing a genome.
20. The search method according to any one of claims 9 to 13, further comprising dividing one or more organizational images from among the organizations of interest, secondary organizations of interest, and tertiary organizations of interest into one or more clusters.
21. The search method according to claim 20, wherein the one or more clusters are classified by the intensity of the video, or by an algorithm that divides the video of an organization of interest into multiple patches and classifies them by the similarity of video features to each patch.
22. To provide a tissue of interest that has a disease or disease, or has been induced, with a substance containing part or all of the genome that is not present in the tissue of interest, and to obtain spatial transcription information of the tissue of interest, A method for classifying genome-derived transcription information from spatial transcription information, comprising classifying genome-derived transcription information from spatial transcription information using an integrated reference transcription library that includes first organoscopy reference transcription information to which the aforementioned tissue of interest belongs and second organoscopy reference transcription information to which the genome or a substance containing the genome belongs.
23. The classification method according to claim 22, wherein the method has even higher classification accuracy than the method of classifying genome-derived transcription information from spatial transcription information using first organism or second organism reference transcription information, respectively, without using an integrated reference transcription library.
24. In order to explore the distribution, efficacy, action, or physiological activity of genome-containing substances within tissues of interest, An information receiving unit (110) provides a substance containing part or all of the genome that is not present in a diseased or diseased tissue of interest, and receives spatial transcription information of the obtained tissue of interest. An origin search unit (120) for spatial transcription information classifies the spatial transcription information into transcription information derived from the tissue of interest and transcription information derived from the genome, using an integrated reference transcription library that includes first organism reference transcription information to which the tissue of interest belongs and second organism reference transcription information to which the genome or a substance containing the genome belongs. A search device (100) including a molecular marker search unit (130) that searches for molecular markers associated with the distribution, efficacy, action, or physiological activity of genome-containing substances in tissues of interest based on information from the origin search unit (120).
25. The search device (100) according to claim 24, further comprising extracting molecular markers spatially related to the ratio of genome-derived transcripts to the total number of spatial transcripts in the tissue of interest from the information of the origin search unit (120).
26. The search device (100) according to claim 24, further comprising a control transcript information receiving unit (111) that acquires control transcript information of a second tissue of interest which is not provided with genome-containing material and which has the same disease or disease as the tissue of interest, or which has been induced, and receives the acquired information.
27. The search device (100) according to claim 24, wherein the molecular marker search unit (130) performs a correlation analysis between the degree of gene expression of the control transcript information receiving unit (111) and the degree of gene expression of the spatial transcript information of the information receiving unit (110) or the origin search unit (120).
28. The search device (100) according to claim 24, further comprising a normal transcript information receiving unit (112) that acquires normal transcript information of normal third tissues of interest that are free from or not induced by the disease or disorder and that do not provide genome-containing material, and receives the acquired information.
29. The search device (100) according to any one of claims 24 to 28, wherein the molecular marker search unit (130) includes comparing spatial transcript information from one or more of the information receiving unit (110), origin search unit (120), control transcript information receiving unit (111), and normal transcript information receiving unit (112) to search for enhanced or suppressed molecular markers derived from the genome-containing substance.
30. The molecular marker search unit (130) is characterized by including at least one of the following: an image characteristic analysis unit (131) that extracts transcription information using an image characteristic extraction algorithm using artificial intelligence; a correlation analysis unit (132) that analyzes the correlation between the image intensity of tissue and the degree of gene expression and extracts transcription information; a cell type analysis unit (133) that uses a cell type analysis algorithm; a gene ontology analysis unit (134) that uses gene ontology analysis; and a clustering unit (135), as described in any one of claims 24 to 28.
31. The search device (100) according to claim 30, characterized in that the cell type analysis unit (133) uses Fisher's exact test, maximum probability assessment, domain adaptation classification, logistic regression analysis, or negative binomial regression analysis algorithm.
32. The search device (100) according to claim 30, characterized in that the cell type analysis unit (133) analyzes cell types using the CellDART algorithm, spSeudoMap algorithm, STopover algorithm, xCell algorithm, RCTD algorithm, Seurat algorithm, celltypist algorithm, cell2location algorithm, Scanorama algorithm, SPOTlight algorithm, DSTG algorithm, CellTrek algorithm, sc-type algorithm, TIMER algorithm, CIBERSORT algorithm, MCP-counter algorithm, Quantiseq algorithm, EPIC algorithm, or scanpy's injest algorithm.
33. The search device (100) according to claim 30, characterized in that the clustering unit (135) classifies the one or more clusters by the image intensity of the tissue of interest, or by an algorithm that divides the image of the tissue of interest into a plurality of patches and classifies them by the similarity of the image features to each patch.
34. An information receiving unit (110) provides a substance containing part or all of the genome that is not present in a diseased or diseased tissue of interest, and receives spatial transcription information of the obtained tissue of interest, and A device (200) for classifying genome-derived transcription information from spatial transcription information, including a spatial transcription information origin search unit (120) that classifies the spatial transcription information into tissue-derived transcription information and genome-derived transcription information using an integrated reference transcription library that includes first organism reference transcription information to which the tissue of interest belongs and second organism reference transcription information to which the genome or a substance containing the genome belongs.