A device and method for analyzing atrial fibrillation myocardial fibrosis
By using spatial transcriptome technology and Visium solutions, combined with R language and Seurat software, we analyzed myocardial fibrosis in atrial fibrillation, solving the problem that the existing technology could not clearly define the pathological mechanism of atrial fibrosis in atrial fibrillation, and achieved the screening and understanding of the transcriptional expression and molecular signaling pathways in the fibrosis area.
Patent Information
- Application Number
- CN202210106058.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2042-01-28
AI Technical Summary
Existing research methods are unable to clearly analyze and determine the pathological occurrence and development of atrial fibrosis in atrial fibrillation, especially at the cellular and molecular levels of complex heterogeneous tissues in myocardial tissue. The research methods are complex, and the spatial location information of cells, genes and mutant genes in the lesion site is unclear.
By using spatial transcriptome technology combined with the Visium spatial gene expression solution, spatial transcriptome sequencing was performed on myocardial tissue sections. Differential gene screening, dimensionality reduction, cluster analysis and differential gene enrichment analysis were performed, and data processing was performed using R language and Seurat software to obtain differentially expressed genes and signaling pathways related to atrial fibrillation and myocardial fibrosis.
It has achieved a deep understanding of the pathological mechanism of atrial fibrosis in atrial fibrillation, screened the transcriptional expression levels and molecular signaling pathways in the fibrosis area, provided research support at the genetic and molecular levels, simplified data analysis and combined with histological gene expression data.
Smart Images

Figure BDA0003493966810000141 
Figure BDA0003493966810000151 
Figure HDA0003493966820000011
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of traditional Chinese medicine internal medicine and bioinformatics technology, and relates to a method for analyzing the pathological mechanism of atrial fibrillation atrial myofibrosis using spatial transcriptome technology. Background Art
[0002] Complex, heterogeneous tissues are composed of a diverse array of cell types and spatially distributed entities. The study of these complex, heterogeneous tissues requires tools that can assess gene expression profiles and the spatial organization of multiple cell types and states. Formal analysis of tissues at the cellular and molecular levels provides biological information while preserving the organization of the tissue and cellular microenvironment.
[0003] There is no tissue dissociation bias in transcriptome analysis because the tissue is intact and cell types in heterogeneous tissues can be determined. The Visium spatial gene expression solution from 10X provides a method for analyzing the spatial whole transcriptome. By overlaying total mRNA gene expression data with morphological H&E stained images, researchers can use Visium to obtain high-throughput gene expression analysis results of the entire transcriptome from whole tissue sections and a variety of sample types. This solution enables researchers to reveal the biological structure of normal and diseased tissues and discover new tissue biomarkers; moreover, by using existing laboratory equipment and tissue analysis tools, Visium can be easily integrated into the workflow, completing the end-to-end workflow from sample slicing to sequencing library construction within one working day.
[0004] Atrial fibrillation is the most common sustained cardiac arrhythmia in adults. Its incidence and prevalence have been increasing annually in recent years. Epidemiological data show that in 2010, the global total number of patients with atrial fibrillation was approximately 33.5 million. The prevalence of atrial fibrillation in adults in my country is 0.2%, and it increases with age, reaching as high as 8.3% in those over 80 years old. Atrial fibrillation increases the risk of stroke, heart failure, and death, reduces the patient's life expectancy, and increases the hospitalization rate for arrhythmia patients, placing a significant economic burden on society.
[0005] Structural remodeling characterized by interstitial fibrosis is the substrate for the maintenance and recurrence of atrial fibrillation. Interstitial fibrosis can block or slow local atrial excitation conduction, facilitating the formation of reentrant circuits. Previous studies have shown that Qipo Shengmai Granules can effectively improve atrial fibrosis in atrial fibrillation. Further investigation of the molecular mechanisms underlying their improvement is crucial for elucidating their pharmacological effects. Qipo Shengmai Granules are composed of raw astragalus, Glehnia littoralis, Radix Ophiopogonis, Amber, Schisandra chinensis, Ziziphus jujuba seeds, Salvia miltiorrhiza, Scrophularia ningpoensis, Periostracum cicadae, and Bombyx batryticatus. Astragalus, Glehnia littoralis, and Radix Ophiopogonis serve as the main ingredients, tonifying qi and nourishing yin; Schisandra chinensis and Ziziphus jujuba seeds serve as the auxiliary ingredients, soothing the heart and calming the mind; Salvia miltiorrhiza cools the blood, clears the heart, relieves restlessness, and activates blood circulation and unblocks the meridians; Amber serves as the auxiliary ingredients, calming the heart, restoring the pulse, and calming the nerves. The entire formula synergistically nourishes qi and yin, calms the mind, and calms palpitations, and is clinically used to treat atrial fibrillation characterized by both qi and yin deficiency.
[0006] In recent years, the field of transcriptomics has developed into a more general way to study the regulatory molecular pathways of cells, even in the presence of mixed components. Although these methods provide information, they mainly focus on homogenized tissues or cells in single-cell suspensions, thus losing spatial specificity. Although in situ hybridization can show the spatial location of molecules, it has throughput limitations.
[0007] Spatial transcriptomics is the process of matching gene expression with immunohistochemical images of tissue sections, thereby locating the gene expression information of different cells in the tissue to the original spatial position of the tissue, distinguishing which genes are active in the tissue, and intuitively detecting the differences in gene expression in different parts of the tissue.
[0008] Spatial transcriptomics combines histological imaging and single-cell sequencing, preserving the location of each transcript through spatial fixation and barcoded cDNA synthesis primers. Spatially resolved mRNA data allows us to focus on specific tissue regions. Comparing transcriptome changes in the lesion region of the model group can clarify the pathological mechanisms associated with atrial fibrillation fibrosis; and comparing transcriptome changes in atrial myocardial fibrosis tissue before and after drug administration can clarify the target of drug action on myocardial tissue.
[0009] The Visium Spatial Gene Expression Solution is a solution provided by 10X for spatial transcriptome sample preparation and data analysis. Researchers use the Visium Spatial Gene Expression Slide to draw spatial gene expression maps of complex tissue samples. The Visium Spatial Gene Expression Slide (also known as a chip or visium slide) is a slide equipped with poly(A) capture and novel spatial barcoding technology. Each visium spatial gene expression slide contains four capture areas (6.5×6.5mm), each of which accommodates a separate tissue slice. Each capture area has 5,000 barcoded spots (spots), each of which contains millions of oligonucleotide chains containing poly(dT). The oligonucleotide chains can capture and identify mRNA, and each spot has a unique barcode sequence. When RNA is released from the cells in the tissue slice, the RNA that migrates to each spot is labeled with the corresponding barcode sequence, and then a library is constructed and sequenced. Next, the data is assigned according to the barcode information of the data to determine which data comes from which location. Ultimately, spatial gene expression can be visualized, that is, this spatial gene expression slide can capture and identify mRNA, and can determine the spatial location of these mRNAs.
[0010] The oligonucleotide strands have unique molecular identifiers, or UMIs, which allow counting of mRNA molecules, as well as spatial barcodes that allow identification of the mRNA's origin within the capture region. Each spot has a diameter of 55 microns, so depending on the tissue type, 1-10 cells can be detected per spot. Summary of the Invention
[0011] The purpose of the present invention is to address the problems of complex research methods at the cellular and molecular levels for complex heterogeneous tissues in myocardial tissue caused by atrial fibrillation, and unclear spatial location information of cells, genes and mutant genes in the lesion site. A method for analyzing the pathological mechanism of atrial myocardial fibrosis in atrial fibrillation using spatial transcriptome technology is provided, which solves the problem that existing research methods cannot clearly analyze and determine the mechanism of occurrence and development of atrial fibrosis in atrial fibrillation.
[0012] To achieve the purpose of the present invention, the present invention provides, on one hand, an analysis device for atrial fibrillation myocardial fibrosis, comprising:
[0013] The differentially expressed gene screening module is used to standardize the spatial transcription data of myocardial tissue slice samples from the normal group and the model group, and screen the highly variable genes in the standardized data using the VST algorithm to obtain a differentially expressed gene set with expression changes from high to low;
[0014] The dimensionality reduction processing module is used to perform linear dimensionality reduction processing on the data of the differentially expressed gene sets obtained by screening the normal group and model group samples after standardization, and obtain PCA reduced dimensionality data; then the nonlinear dimensionality reduction algorithm of T-SNE is used to perform nonlinear dimensionality reduction processing on the PCA reduced dimensionality data, and obtain T-SNE reduced dimensionality data;
[0015] Cluster analysis module, used to perform cluster analysis on the dimension-reduced data using the SNN clustering algorithm to obtain cluster subgroups of genes with differential expression changes between the normal group and the model group samples;
[0016] The cluster differential gene set screening module is used to compare the genes of each cluster subgroup of the normal group and the atrial fibrillation model group animal tissue samples with the genes of all other cluster subgroups of the tissue samples of the corresponding groups using the Wilcox rank sum test of the Seurat software, and screen the differential gene sets of each cluster (i.e., the gene sets with the largest expression differences in each cluster);
[0017] The main differential cluster screening module is used to perform differential gene enrichment analysis on the cluster differential gene sets of each cluster of the normal group and the model group samples using the ClusterProfiler package of R language, and compare the enrichment analysis results of the normal group and the model group to screen and obtain the main differential clusters;
[0018] The main differential cluster analysis module is used to perform differential gene function enrichment analysis on the cluster differential gene sets of the main differential clusters of the model group and other cluster subgroups of the model group, and obtain the relevant biological processes and signal pathways of differential genes related to atrial fibrillation myocardial fibrosis.
[0019] The differentially expressed genes set screened by the differentially expressed genes screening module is the top 3,000 differentially expressed genes from high to low.
[0020] In particular, the sctransform (SCT) algorithm of Seurat software was used to standardize the spatial transcriptome sequencing data of each sample to obtain the standard processed data of spatial transcriptome data of each sample.
[0021] In particular, it also includes a spatial transcription data expression matrix conversion module, which is used to use the 10×Genomics official software SpaceRanger to align spatial transcriptome sequencing data, quantify genes, and identify sites to obtain the spatial transcription data expression matrix.
[0022] In particular, the spatial transcriptional expression matrix is an M*N matrix where the rows are genes and the columns are spots.
[0023] Then, the expression matrix of spatial transcription data was subjected to screening of highly variable genes using the expression change differential gene screening module to obtain a set of expression change differential gene sets.
[0024] In particular, it also includes a spatial transcription module for using spatial transcriptome technology to perform spatial transcription sequencing on myocardial tissue slice samples of normal group and model group animals respectively to obtain spatial transcription data of normal group and model group animal samples.
[0025] In particular, the spatial transcription module is used to perform spatial transcription sequencing experiments according to the spatial transcriptome sequencing method of 10×Genomics visium (10x Genomics (TXG) company gene spatial transcription sequencing) to obtain spatial transcription data.
[0026] Among them, the dimensionality reduction processing module also includes using the sctransform (SCT) algorithm in the Seurat package of R language to standardize the expression change difference gene set data, and after obtaining the standardized data of highly variable genes, use the PCA method to perform linear dimensionality reduction processing.
[0027] The cluster analysis module clusters the differentially expressed genes of the normal group samples into three clusters (cluster subgroups); and clusters the differentially expressed genes of the model group samples into four clusters.
[0028] In particular, it also includes the visualization of cluster analysis processing results.
[0029] In particular, the difference comparison standard of the cluster analysis processing in the cluster difference gene set screening module is FoldChange≥2 and q-value<0.1, that is, the genes with FoldChange≥2 and q-value<0.1 are selected as genes with large changes in characteristic expression differences.
[0030] Wherein, the main difference group screening module is specifically used for:
[0031] The ClusterProfiler package in R language was used to perform biological process enrichment analysis on the cluster differential gene sets of each cluster of the normal group and model group samples, and the enrichment analysis results of the normal group and the model group were compared to screen out the main differential groups, namely the biological process analysis (BP analysis) of the GO annotation system.
[0032] In particular, the obtained biological process enrichment analysis results (Term) were visualized by drawing bar charts, bubble charts, bubble charts, etc., to compare the biological processes of each cluster of different samples and screen out the main difference groups (main difference clusters).
[0033] In particular, the biological process enrichment analysis of each cluster in the normal group and model group samples showed that compared with the normal group, cluster 2 in the model group was the main differential cluster.
[0034] In particular, the differential gene enrichment analysis is a biological process analysis (ie, BP analysis) of the GO annotation system.
[0035] Wherein, the main difference group analysis module is specifically used for:
[0036] Differential gene enrichment analysis was performed on the clustered differential gene sets of the main differential clusters of the model group and other cluster subgroups of the model group, that is, biological process enrichment analysis (i.e., biological process analysis (BP analysis) of the GO annotation system) was performed on the differential genes of the main differential clusters, and the biological processes involved in atrial fibrillation myocardial fibrosis, such as aging, collagen fiber binding, complement activation, extracellular matrix, and response to mechanical stimulation, were obtained;
[0037] KEGG pathway analysis was performed on the clustered differential gene sets of the main differential clusters in the model group and other cluster subgroups of the model group, and it was found that atrial fibrillation myocardial fibrosis was related to the complement system, ECM-receptor interaction, and AGE-RAGE signaling pathway in diabetic complications.
[0038] In particular, the main difference group analysis module also includes a cluster difference gene set screening component for the main difference groups: it is used to perform difference analysis on the genes of the main difference groups of the model group and the genes of other cluster subgroups of the model group (i.e., the collection of cluster0, cluster1, and cluster3 genes) using the Wilcox rank sum test of the Seurat software to obtain the cluster difference gene set of the main difference groups of the model group.
[0039] In particular, the clustered differential gene sets of the main differential clusters analyzed by the main differential cluster analysis module are Nppa, Apoe, Reg3b, Myl4, Bgn, Postn, Ftl1, Lbp, lgfbp4 and Fn1.
[0040] Another invention of the present invention provides a method for analyzing myocardial fibrosis caused by atrial fibrillation, comprising the following steps:
[0041] 1) Spatial transcriptome technology was used to perform spatial transcriptome sequencing on myocardial tissue sections of animals in the normal group and the atrial fibrillation model group to obtain spatial transcriptome data of the normal group and the model group samples respectively;
[0042] 2) The sctransform (SCT) algorithm in the Seurat package in the R language was used to convert the spatial transcriptome raw data of each sample into expression matrix data for standardization. The standardized data were then filtered for highly variable genes using the vst algorithm to obtain a set of differentially expressed genes.
[0043] 3) The sctransform (SCT) algorithm in the Seurat package in the R language was used to normalize the differentially expressed gene sets screened for each sample. The normalized data were then subjected to linear dimensionality reduction using the PCA method to obtain PCA-reduced data. The PCA-reduced data were then subjected to nonlinear dimensionality reduction using the T-SNE algorithm to obtain T-SNE-reduced data.
[0044] 4) Cluster analysis of the T-SNE dimensionality reduction data was performed based on the SNN clustering algorithm to obtain cluster subgroups of genes with differential expression changes in the normal group and the atrial fibrillation model group;
[0045] 5) Seurat software's Wilcox rank sum test was used to compare the genes of each cluster subgroup of the normal group and model group samples with the genes of all other cluster subgroups of the corresponding group samples, and the cluster differential gene set of each cluster was screened (i.e., the gene set with the largest expression difference in the cluster);
[0046] 6) The ClusterProfiler package in R language was used to perform differential gene enrichment analysis on the cluster differential gene sets of each cluster of the normal group and the model group samples, and the differential gene enrichment analysis results of the normal group and the model group were compared to screen for the main differential clusters;
[0047] 7) Perform differential gene function enrichment analysis on the main differential clusters of the model group samples and the differential gene sets of other clusters of the model group samples to obtain the relevant biological processes and signaling pathways of the differential genes related to atrial fibrillation myocardial fibrosis.
[0048] Wherein, the spatial transcription sequencing process in step 1) comprises the following steps:
[0049] 1A) Myocardial tissue sections from normal and atrial fibrillation model animals were permeabilized to release mRNA from cells and bind to the corresponding capture probes on a spatial transcription gene expression chip;
[0050] 1B) Reverse transcription of the mRNA bound to the capture probe of the gene expression chip is performed to synthesize a complete first-strand cDNA; second-strand cDNA is then synthesized; incubation is then performed to denature the cDNA; and finally, the denatured cDNA is recovered, amplified, and purified to obtain amplified cDNA;
[0051] 1C) The amplified cDNA was subjected to fragmentation, end-repair and A-tailing, magnetic bead paired-end fragment screening, adapter ligation, magnetic bead purification after adapter ligation, sample index PCR, and magnetic bead paired-end fragment screening after PCR to construct the Visium spatial gene expression library;
[0052] 1D) The Visium spatial gene expression library samples of the normal group and the atrial fibrillation model animal group were sequenced using the Illumina NovaSeq 6000 platform to obtain spatial transcriptome sequencing data for each sample.
[0053] In particular, the spatial transcriptome sequencing process is to conduct experiments according to the spatial transcriptome sequencing method of 10×Genomics visium (10x Genomics (TXG) company gene spatial transcriptome sequencing) to obtain spatial transcriptome data.
[0054] Particularly, the time for the tissue permeabilization treatment in step 1A) is 6-9 minutes, preferably 8 minutes.
[0055] In particular, it also includes determining the tissue permeabilization treatment time according to the following steps: first, the myocardial tissue slice is completely adhered to the Visium spatial gene tissue optimized slide (also known as the Visium spatial gene tissue optimized chip) with 8 capture areas, and according to the operating procedures of the spatial transcriptome sequencing method (10×Genomics Visium), the 6 different capture areas on the slide are respectively subjected to enzyme treatment for different times, and the signal intensity of the fluorescent-labeled cDNA in each capture area is recorded, and the enzyme treatment time corresponding to the capture area with high fluorescence intensity is selected as the tissue permeabilization treatment time.
[0056] In particular, in step 1A), tissue sections with RIN ≥ 7 and good integrity are selected for tissue permeabilization.
[0057] Among them, step 1B) also includes fragment detection of the amplified cDNA. The overall fragment size of the amplified cDNA ranges from 200 to 9000 bp, indicating that the cDNA synthesized from the sample has passed the quality control.
[0058] In particular, in step 1D), during the sequencing process, the sequencing strategy is PE150; the sequencing depth is 50,000 reads / spot.
[0059] In step 1), the animals selected are healthy male SD rats.
[0060] In particular, before performing spatial transcription sequencing processing in step 1), the method also includes: preparing a Visium spatial gene expression chip, that is, completely adhering the left atrial muscle tissue sections of the normal group and the model group animals to a Visium spatial gene expression slide (Visium spatial gene expression chip, chip, Visium Spatial Gene Expression Slide) with 4 capture areas, covering the exposed tissue with OCT and freezing it to prepare a Visium spatial gene expression chip for spatial transcription sequencing.
[0061] In step 1), the mRNA captured by the oligonucleotide chain in the capture area of the slide and marked with the unique molecular identifier UMI and the special spatial barcode in each tissue section is sequenced by spatial transcriptome technology to obtain spatial transcriptome data.
[0062] The spatial transcriptome sequencing data normalization process in step 2) includes the following steps:
[0063] The spatial transcriptome data were filtered using Fastp software to obtain sequencing data that could be directly used for subsequent analysis;
[0064] Using the barcode processing algorithm, the cell barcode information and the corresponding counts contained in the filtered statistical sequencing data are counted to determine the actual number of detected spots in the sequencing sample and obtain the true spatial transcriptome sequencing information;
[0065] Align the reads corresponding to the cellbarcode in the sequencing data to the genome of the corresponding known species, analyze the similarities and differences between the unknown sequence and the known sequence, and obtain the bam file for the comparison;
[0066] The bam file containing various information after genome alignment was converted, and the single-molecule tags aligned to the same gene in the file were merged, and the repeated UMI sequences were removed to obtain the number of UMIs corresponding to each gene. The number of genes detected in each spatial array was then counted and visualized in the spatial staining slide.
[0067] In particular, step 2A) is also included: the spatial transcriptome sequencing data of each sample is obtained using the 10×Genomics official software SpaceRanger to perform data alignment, gene quantification, and site identification to complete upstream analysis; the main function of upstream analysis is to convert the fastq files obtained by the original sequencing into the expression matrix required for downstream analysis and processing.
[0068] Then, the sctransform (SCT) algorithm in the Seurat package of R language was used to convert the spatial transcriptome sequencing data of each sample into expression matrix data for standardization.
[0069] It is used to use the 10×Genomics official software SpaceRanger to align spatial transcriptome sequencing data, quantify genes, and identify sites to obtain the spatial transcription data expression matrix, where the spatial transcription data expression matrix is an M*N matrix, where rows are genes and columns are spots.
[0070] In particular, the differentially expressed gene set in step 2) is the top 3000 differentially expressed genes with the highest to lowest expression differences.
[0071] In step 4), the genes with differential expression changes in the normal group animal samples were clustered into three clusters (cluster subgroups); the genes with differential expression changes in the model group animal samples were clustered into four clusters.
[0072] In particular, it also includes the visualization of cluster analysis processing results.
[0073] Comparing the differential gene analysis between the lesion area and the normal area can reveal the molecular mechanism of pathological changes in the lesion area.
[0074] Among them, the difference comparison standard in the screening process of the cluster differential gene set in step 5) is Fold Change ≥ 2 and q-value < 0.1, that is, the genes with Fold Change ≥ 2 and q-value < 0.1 are selected as genes with large changes in characteristic expression differences.
[0075] In particular, the screening criteria in step 5) are: genes expressed in more than 25% of the spots in at least one comparison group are subjected to differential comparative analysis.
[0076] Wherein, the differential gene enrichment analysis in step 6) is to perform biological process enrichment analysis on the clustered differential gene sets.
[0077] In particular, it also includes visualizing the terms obtained from the biological process enrichment analysis by drawing bar charts, bubble charts, etc., comparing the biological processes of each cluster of different samples, and screening to obtain the main difference groups (main difference clusters).
[0078] In particular, the biological process enrichment analysis of each cluster in the normal group and model group samples showed that compared with the normal group, cluster 2 in the model group was the main differential cluster.
[0079] In step 7), the biological enrichment analysis of the differentially expressed genes in each cluster of the normal group and model group samples was performed, i.e., the biological process analysis (BP analysis) of the GO annotation system was performed. The enrichment analysis results of the normal group and the model group were compared to screen out the main differential groups.
[0080] Among them, in step 7), the cluster differential gene set is obtained according to the following steps: the genes of the main differential groups of the model group and the genes of other cluster subgroups of the model group (i.e., the collection of cluster0, cluster1, and cluster3 genes) are differentially analyzed using the Wilcox rank sum test of Seurat software to obtain the cluster differential gene set of the main differential groups of the model group.
[0081] In particular, the clustered differentially expressed genes of the main differentially expressed groups in the model group described in step 7) are Nppa, Apoe, Reg3b, Myl4, Bgn, Postn, Ftl1, Lbp, lgfbp4 and Fn1.
[0082] In particular, the clustered differential genes of the main differential groups in the model group are involved in biological processes such as aging, collagen fiber binding, complement activation, extracellular matrix, and response to mechanical stimulation, and are related to the complement system, ECM-receptor interaction, AGE-RAGE signaling pathway in diabetic complications, etc., leading to atrial fibrosis in atrial fibrillation.
[0083] In particular, the functional enrichment analysis of the differentially expressed genes in the major differentially expressed groups in step 7) includes the following steps:
[0084] The biological process enrichment analysis (i.e., biological process analysis of the GO annotation system and BP analysis) of the clustered differential gene sets of the main differential clusters in the model group and other cluster subgroups of the model group was performed to obtain biological processes such as aging, collagen fiber binding, complement activation, extracellular matrix, and response to mechanical stimulation involved in atrial fibrillation myocardial fibrosis;
[0085] KEGG pathway analysis was performed on the clustered differential gene sets of the main differential clusters in the model group and other cluster subgroups of the model group, and it was found that atrial fibrillation myocardial fibrosis was related to the complement system, ECM-receptor interaction, and AGE-RAGE signaling pathway in diabetic complications.
[0086] The present invention overcomes the shortcomings of existing technologies by analyzing the pathogenesis of atrial fibrosis in atrial fibrillation using spatial transcriptome technology, and elucidating the mechanisms of atrial fibrillation at the cellular and molecular levels. The present invention uses spatial transcriptome technology to sequence frozen sections of pathological tissue from rats with atrial fibrillation. The resulting spatial transcriptome data is then used for subsequent analysis. Clustering and differential analysis of all detected tissue transcriptome data is then performed to identify key genes and related molecular signaling pathways regulated by Qipo Shengmai Granule in improving atrial fibrosis in rats with atrial fibrillation.
[0087] Compared with the shortcomings and deficiencies of the existing technology, the present invention has the following beneficial effects: the method of the present invention can obtain transcription information of specific spatial positions in complete tissue sections, which is helpful to screen the transcription expression levels of fibrotic areas, and further helps to screen and regulate differential genes and molecular signaling pathways in fibrotic areas of atrial fibrillation lesions, and provide technical support for the final analysis and understanding of the occurrence and development process of atrial fibrillation lesions, and provide research support at the genetic and molecular levels.
[0088] The device of the present invention obtains the entire transcriptome in the entire tissue and the complete transcriptome gene expression from the intact tissue, so there is no need to dissociate the tissue and there is no tissue dissociation bias. Combined with clinical pathological information, the pathological characteristics of the tissue are combined with gene expression. The analysis software used combines histological gene expression data to simplify data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1 This is the experimental flow chart of the spatial transcriptome sequencing method (10x Genomics visium);
[0090] Figure 2 HE staining images of rat tissue sections and spatial tissue distribution diagrams of samples, with the HE staining image on the left and the spatial tissue distribution diagram on the right. A is the normal group; B is the model group;
[0091] Figure 3A 、 3B Volcano plot of gene mutations reported for rat atrial tissue; 3A is the normal group and 3B is the model group; Figure 3A 、 3B The horizontal axis is the geometric mean of gene expression, and the vertical axis is the variation / variance statistic. The red dots represent highly variable genes, a total of 3,000, and the top 10 highly variable genes are labeled.
[0092] Figure 4 t-SNE diagram of Spots clustering of rat atrial myocardial tissue sections, where A is the normal group and B is the model group; Figure 4 Each point in the figure represents a spot, different colors represent different clusters, and the distribution areas of each category are marked in detail.
[0093] Figure 5A This is the cluster spatial distribution diagram of atrial muscle tissue sections of normal rats;
[0094] Figure 5B This is the cluster spatial distribution diagram of atrial muscle tissue sections of rats in the model group;
[0095] Figures 6A-6C These are the biological process enrichment analysis diagrams of differentially expressed genes in cluster 0, cluster 1, and cluster 2 of the atrial muscle tissue of normal rats;
[0096] Figures 6D-6G These are the biological process enrichment analysis diagrams of differentially expressed genes in cluster 0, cluster 1, cluster 2, and cluster 3 of the atrial muscle tissue of rats in the model group. DETAILED DESCRIPTION
[0097] The present invention will be further described below with reference to specific embodiments, and the advantages and features of the present invention will become clearer as the description proceeds. However, these embodiments are merely exemplary and do not constitute any limitation to the scope of the present invention. It should be understood by those skilled in the art that the details and forms of the technical solutions of the present invention may be modified or replaced without departing from the spirit and scope of the present invention, and such modifications and replacements fall within the scope of protection of the present invention.
[0098] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0099] Example 1
[0100] 1. Experimental Materials
[0101] 1-1. Experimental Animals
[0102] Healthy male SD rats (8, cleanliness grade II, Beijing Weitonglihua Laboratory Animal Technology Co., Ltd.), 4 weeks old, weighing 180-200 g, animal license number: SCXK (Beijing) 2014-0041.
[0103] The experimental animals (8) were adaptively fed for 7 days at the Animal Center of Guang'anmen Hospital, China Academy of Chinese Medical Sciences, and then the model was established. Seven days after the model was established, the animals were divided into groups for subsequent experiments.
[0104] 1-2. Experimental instruments
[0105] BL-420F multi-channel physiological signal acquisition and processing system (Chengdu Taimeng Software Co., Ltd., China); Leica CM1850 freezing microtome (Leica, Germany).
[0106] 1-3. Reagents and drugs
[0107] Acetylcholine culoride (ACh; Sigma, USA); anhydrous calcium chloride (Sigma, USA); pentobarbital sodium (Sinopharm Chemical Reagent Co., Ltd., China); sodium chloride injection (Double Crane Pharmaceuticals, China);
[0108] Qipo Shengmai Granules (QPSM, provided by the Preparation Center of Beijing Guang'anmen Hospital) are made from raw astragalus, Glehnia littoralis, Radix Ophiopogonis, Amber, Schisandrae Chinensis, Semen Ziziphi Spinosae, Salvia miltiorrhiza, Scrophulariae ningpoensis, Periostracum Cicadae, and Bombyx Batryticatus. The preparation method is as follows:
[0109] The following medicinal materials in the following weight ratio: 30 grams of raw astragalus, 12 grams of northern adenophora, 12 grams of ophiopogon, 2 grams of amber, 6 grams of schisandra, 15 grams of spiny date seeds, 15 grams of salvia miltiorrhiza, 15 grams of figwort, 10 grams of bombyx batryticatus, and 10 grams of cicada shells, were rinsed, placed in a multifunctional Chinese medicine extraction kettle and mixed evenly, and then soaked in tap water for 30 minutes, wherein the ratio of the weight of the tap water added to the weight (dry weight) of the medicinal material mixture was 9:1;
[0110] Turn on the power of the multifunctional extraction kettle for heating, heat to 100°C (not limited to 100°C, other temperatures are applicable), maintain constant temperature extraction at this temperature for 3 hours, and then filter to obtain the first extract for standby use; add tap water to the filter residue, the weight of the tap water added to the weight of the medicinal material mixture (dry weight) is 5:1, heat extraction, maintain the heating temperature at 100°C for extraction for 3 hours, filter to obtain the second extract; combine the two extracts and place them in a vacuum high-energy evaporation concentrator for reduced pressure evaporation and concentration to obtain a concentrated extract, wherein the reduced The relative pressure of the reduced pressure concentration is -0.07 MPa (i.e., with one atmospheric pressure as 0, the pressure during the reduced pressure concentration is controlled to be -0.07 MPa relative to one atmospheric pressure), the concentration temperature is 80° C., and the secondary extracts are concentrated into an extract with a relative density of 1.2; the concentrated extract is dried at 45-50° C., crushed, and passed through a 180-mesh sieve to obtain a dry extract powder; the dry extract powder is mixed with an auxiliary material, dextrin, and ethanol is used as a binder, granulated, and dried to obtain the traditional Chinese medicine composition (Qi Po Sheng Mai, QPSM) granules for atrial fibrillation of qi and yin deficiency type of the present invention.
[0111] 2. Experimental methods
[0112] 2-1. Construction of model animals
[0113] According to the previous experimental methods of our research group (Zhang Wantong, Hu Yuanhui, Shi Shuai, et al. Establishment of a vagal atrial fibrillation rat model and atrial electrophysiological study [J]. Heart Journal, 2019, 01: 8-11), the atrial fibrillation model animal was constructed. The specific methods are as follows:
[0114] After 7 days of adaptive feeding, 4 healthy male SD rats were randomly selected and injected with a mixture of acetylcholine (Ach) (99ug / mL) and calcium chloride (CaCl2, 10mg / mL) through the tail vein. The dosage volume was 0.1mL / 100g, that is, the volume of the acetylcholine-calcium chloride mixture injected was 0.1mL / 100g. After 7 consecutive days of administration, the electrocardiogram was measured immediately after the tail vein injection on the seventh day. The appearance of typical atrial fibrillation electrocardiogram AF wave on the electrocardiogram: P wave disappearance, unequal RR intervals, and replacement by f waves of varying sizes was considered a sign of successful modeling. The test results showed that all 4 rats had typical atrial fibrillation electrocardiograms, indicating that the atrial fibrillation model rats were successfully established.
[0115] 2-2. Experimental Grouping and Drug Administration
[0116] Four rats successfully modeled were randomly divided into a model group (two rats) and a Qiposhengmai group (two rats) using an Excel random number table. The model group rats were gavaged with purified water, while the Qiposhengmai group rats were gavaged with a Qiposhengmai granule solution. Administration was continued for four weeks. Each morning during these four weeks, all rats were gavaged first. Within one hour of gavage, a tail vein injection of an acetylcholine-calcium chloride mixture (0.1 mL / 100 g) was administered.
[0117] The remaining two experimental rats that were not modeled served as the normal control group. The normal group rats were first gavaged with purified water every morning. Within one hour after gavage, normal saline (injection volume 0.1 mL / 100 g) was injected into the tail vein. That is, no intravenous injection of acetylcholine-calcium chloride mixture was performed, and normal saline was given gavage for 4 weeks.
[0118] The experimental rats were divided into 3 groups: normal control group (2 rats); model group (2 rats); experimental group (QPSM group (2 rats). The rats in each group were administrated according to the following method:
[0119] 1) Normal group (Sham): purified water 1.0 ml / 100 g was administered orally once a day for 4 weeks;
[0120] 2) Model group: pure water 1.0ml / 100g gavage, qd*4w;
[0121] 3) Qipo Shengmai group (QPSM): Qipo Shengmai granule solution 1.0ml / 100g gavage, qd*4W (0.2g / ml); wherein: Qipo Shengmai granule solution is prepared by dissolving the prepared Qipo Shengmai granules in purified water and mixing, 0.2g / ml.
[0122] After four weeks of continuous drug administration, the rats in each group were subjected to the following experimental tests:
[0123] 2-3. Electrocardiogram
[0124] After 4 weeks of continuous drug administration, the SD rats were weighed in the morning on an empty stomach and the weight was recorded. The SD rats were then anesthetized with low-dose isoflurane gas, and the electrocardiograms of the anesthetized SD rats were measured, and standard lead II electrocardiograms were recorded.
[0125] Observe the recorded electrocardiogram for the appearance of typical AF waves: the disappearance of P waves and their replacement by f waves of varying sizes is a sign of atrial fibrillation; the restoration of sinus rhythm, the disappearance of f waves, and the appearance of P waves is the termination of AF waves; record the duration of AF waves.
[0126] Electrocardiogram (ECG) results showed no atrial fibrillation (AF) in normal rats. The duration of AF in the model group rats was 11.735 seconds, with typical AF waves. The duration of AF in the Qipo Shengmai granule group rats was 4.253 seconds. No AF waves were detected in the normal group rats.
[0127] 2-4. Preparation of atrial muscle tissue slice samples
[0128] After drug intervention, rat atrial myocardial samples were collected at the end of the fourth week. Rats in each group were anesthetized with an intraperitoneal injection of 3% sodium pentobarbital (30 mg / kg). After blood was drawn from the abdominal aorta, the abdominal wall was quickly cut open along the auxiliary edge below the xiphoid process. The diaphragm was opened, and the chest wall was cut open along the anterior axillary line on both sides. The heart was turned upward and cephalad, and the left atrial myocardial tissue was cut off at the root of the heart. The tissue was then placed in a culture dish and the surrounding blood was blotted dry.
[0129] Mark the freezing mold to mark the direction of the tissue and fill the mold with pre-cooled OCT embedding medium to avoid bubbles; place the atrial muscle tissue in the mold containing pre-cooled OCT embedding medium, cover the surface of any exposed tissue with OCT embedding medium, and confirm that there are no bubbles, especially near the tissue; use tweezers to place the mold in isopentane, but do not let the isopentane immerse into the mold until the tissue is frozen; the freezing time can vary according to the type and size of the tissue; after freezing, transfer the tissue to a pre-cooled cryovial and then place it on dry ice; store the OCT-embedded tissue block in a sealed container at -80℃ for long-term storage, or perform frozen sectioning immediately; freeze section at -70℃, cut 20 slices with a thickness of 10μm for each tissue, and prepare atrial muscle tissue slices of normal rats, model rats, and QPSM rats.
[0130] 2-5. Spatial transcriptome sequencing process of rat atrial muscle tissue
[0131] Spatial transcriptome sequencing of rat atrial myocardium was performed as follows: Figure 1 The experimental process of the spatial transcriptome sequencing method (10×Genomics visium) was used to perform spatial transcriptome sequencing experiments on the left atrial muscle tissue of each group of rats (one rat was randomly selected from each group). The details are as follows:
[0132] 2-5-1 Organize quality inspection
[0133] RNA was extracted from atrial myocardial tissue sections (10 sections per tissue) and quality-checked. Atrial myocardial tissue sections with a RIN value ≥7 and good integrity were completely adhered to Visium Spatial Gene Tissue Optimization slides (also known as Visium Spatial Gene Tissue Optimization Arrays) with eight capture zones for optimized tissue permeabilization time. The sections were also completely adhered to Visium Spatial Gene Expression slides (also known as Visium Spatial Gene Expression Arrays) with four capture zones. The sections were immediately frozen, the exposed tissue covered with OCT, and frozen. The sections were then stored in a sealed container at –80°C for subsequent spatial transcriptome sequencing.
[0134] 2-5-2 Optimization of tissue permeabilization time
[0135] Before constructing the library, it is necessary to optimize the optimal permeabilization time of the target tissue mRNA; the optimal permeabilization time, that is, the optimal permeabilization condition for mRNA in library construction, is determined by the time gradient of enzyme treatment and the signal intensity of fluorescently labeled cDNA.
[0136] Each tissue-optimized slide contains 8 capture areas, each measuring 8mm*8mm. Six tissue sections (time gradient) can be affixed, and the remaining two areas are used to set up a positive control and a negative control. Six tissue sections are used to determine the time gradient allocation for the optimal permeabilization time, and the different permeabilization times for the six tissue sections are 3min, 6min, 9min, 12min, 18min, and 24min.
[0137] The experimental process is as follows: first, slides are mounted according to the area, and the mounted slides are fixed, stained with HE, and imaged in the field; then time-gradient permeabilization and fluorescent-labeled cDNA synthesis are performed; finally, the tissue is removed and fluorescence imaging is performed, and the optimal permeabilization time is determined by the intensity of the fluorescence signal.
[0138] After the SD rat atrial muscle tissue slices of the present invention are subjected to different permeabilization time gradients, fluorescence scanning shows that the fluorescence intensity of the slices with permeabilization times of 6 minutes and 9 minutes is relatively strong. Therefore, the permeabilization time within the range of 6-9 minutes can be used as the optimal tissue permeabilization time of the present invention, and the longer tissue permeabilization time is preferably selected. Finally, 8 minutes is selected as the optimal time for atrial muscle tissue permeabilization; the permeabilization time is 6-9 minutes, preferably 8 minutes.
[0139] 2-5-3. Fixation, staining, and imaging
[0140] Patching: Tissue sections with a RIN value ≥7 and good integrity were completely adhered to a Visium Spatial Gene Expression Slide with four capture zones (i.e., Visium Spatial Gene Expression Chip, Chip, Visium Spatial Gene Expression Slide), and the sections were immediately frozen. The exposed tissue was covered with OCT and frozen, and stored in a sealed container at –80°C.
[0141] Fixation: Remove the slide, place the slide tissue-side up, and immediately place it on the adapter. Incubate at 37°C for 1 minute. After 1 minute, immerse the slide completely in -20°C pre-cooled methanol and incubate for 30 minutes.
[0142] HE staining: The tissue was fixed with methanol for 30 minutes and HE staining was performed. The visium spatial gene expression slides were incubated in isopropanol, hematoxylin, blue and eosin mix, washed and dried to complete HE staining.
[0143] Widefield imaging experiments: Images were captured at 20x magnification using the Metafer slide scanning platform (MetaSystems). Four fiducials per region were used on Visium spatial gene expression slides to ensure clear image quality and morphological structure. Raw images were stitched using Vslide software (MetaSystems).
[0144] Observation results such as Figure 2 As shown, Figure 2 The left side of the middle image is the HE staining of the atrial muscle tissue sections of the two groups (normal group and model group). Figure 2 It can be seen that compared with the normal group, the model group had local inflammatory cell infiltration (i.e., lesion area); compared with the model group, the inflammatory cell infiltration in the QPSM group was significantly reduced, and the inflammatory cell infiltration area was Figure 2 The model group and QPSM group have darker areas.
[0145] 2-5-4. cDNA amplification
[0146] First, tissue permeabilization was performed on the Visium spatial gene expression chip of each sample to release the mRNA in the cells and bind to the corresponding capture probes on the chip. The tissue permeabilization time was determined by the tissue optimization experiment, which was 8 minutes. The mRNA was then reverse transcribed to obtain a complete cDNA strand. The second strand of cDNA was then synthesized. At the end of the incubation, the second strand was denatured by incubating with 0.08M KOH for 10 minutes. Finally, the denatured second strand was recovered for cDNA amplification.
[0147] 2-5-5. cDNA quantification and quality control
[0148] 1 μl of the amplified cDNA from each sample was diluted to 10 μl and then fragments were detected using a HighSensitivity DNA chip (Agilent Bioanalyzer). The overall cDNA fragment size of the three samples in this experiment (normal group, model group, and QPSM group) ranged from 200 to 9000 bp, indicating that the cDNA synthesized from the three samples met quality control standards.
[0149] 2-5-6. Construction of gene expression library
[0150] 10 μL of cDNA from each of the three amplified and purified samples was taken and the library was constructed through fragmentation, end-repair and A-tailing, magnetic bead double-end fragment screening, adapter ligation, magnetic bead purification after adapter ligation, sample index PCR (i.e., barcode identification of which of the three samples the data belongs to), and magnetic bead double-end fragment screening after PCR.
[0151] 2-5-7 Library Quantification and Quality Control
[0152] Library concentration was determined using Qubit 4.0, and fragment analysis was performed using Qseq400. The typical library size ranged from 300-800 bp, with an average fragment size of 400-500 bp. The cDNA libraries of the three samples presented in this study all ranged from 300-800 bp, with an average fragment size of 400-500 bp. This indicates that the library quality control passed.
[0153] 2-5-8. Spatial transcriptome sequencing
[0154] The Illumina NovaSeq 6000 platform was used to sequence each sample in the Visium spatial gene expression library of the three samples to obtain spatial transcriptome sequencing data for each sample. The sequencing strategy was PE150 and the sequencing depth was 50,000 reads / spot.
[0155] Spatial transcriptome data can be externally input, and the external processing process can be performed according to the experimental process of the above-mentioned spatial transcriptome sequencing method (10×Genomics visium).
[0156] 2-6. Data Analysis
[0157] 2-6-A Upstream Analysis
[0158] 10×Genomics official software SpaceRanger (v1.2.2) was used to perform data alignment, gene quantification, and site identification to complete upstream analysis, that is, to convert the fastq files obtained by original sequencing into the expression matrix required for downstream analysis and processing.
[0159] 2-6-B downstream analysis
[0160] The expression matrix (upstream analysis results) was normalized, clustered, and selected for marker genes using the Seurat (v4.0.1) package in R. Functional enrichment analysis was performed using the clusterProfiler (v3) package in R. Data were filtered using default parameters for analysis. The results were visualized using the ggplot2 package in R.
[0161] 3. Results and Analysis
[0162] 3.1 Upstream analysis results
[0163] Table 1 Statistical results of raw data
[0164]
[0165]
[0166] Note: Reads: the total number of pair-end reads in Raw Data; Total_base (Mbases): the total number of bases in Raw Data (Mb); GC (%): the GC content of Raw Data, that is, the percentage of G and C bases in the total bases in Raw Data; N (%): the proportion of N in Raw Data; Q20 (%): the percentage of bases with a quality value greater than or equal to 20 in Raw Data; Q30 (%): the percentage of bases with a quality value greater than or equal to 30 in Raw Data.
[0167] Table 1 shows that the total number of spots (marked spots on the chip) for the three groups was 4458, 3791, and 3820, respectively. The median number of genes per spot was 1252, 2094, and 1530.5, indicating good capture and optimized permeabilization conditions. Sequencing saturation was 90.4, 86.2, and 82.5, respectively. The median number of UMIs per spot was 5,666, 13,020, and 9,117, respectively. The average number of reads per spot was 86,532.33, 147,189.65, and 88,098.41, respectively, indicating satisfactory sequencing depth.
[0168] 3.2 Downstream analysis results
[0169] 3-2-1. Screening of highly variable genes
[0170] The sctransform (SCT) algorithm in the Seurat package of R language was used to convert the original sequencing data of the spatial transcriptome sequencing data of each sample into expression matrix data for standardization. Then, the standardized data was filtered out by the vst algorithm to select the gene set with the largest expression change in the data as the high-variance genes (HVGs), and the top 3000 genes with the highest expression change difference from high to low (i.e., the top 3000 genes with the largest expression change) were selected to obtain the volcano map of the highly variable genes. The volcano maps of the normal group and the model group are shown in Figure 2. Figure 3A 、 3B As shown in the figure.
[0171] Gene sets with significant expression changes indicate they are meaningful for determining pathological mechanisms. Subsequent processing is based on these 3,000 genes. In the volcano plot, highly variable genes appear above less variable genes, and the 10 most variable genes in each group are labeled in their respective volcano plots.
[0172] 3-2-2. Dimensionality Reduction and Clustering
[0173] 3000 highly variable gene data were screened from each sample (i.e., normal group and model group) and further standardized using the sctransform (SCT) algorithm in the Seurat package of R language to obtain standardized data of highly variable genes.
[0174] The 3,000 highly variable gene data were normalized using the sctransform (SCT) algorithm in the Seurat package in the R language. The purpose and function of this process is to remove variations caused by differences in samples and spot sequencing depths.
[0175] The expression data quantified by 10×Genomics using SpaceRanger software is an M*N matrix (rows are genes, columns are spots). The number of spots can often reach hundreds or even thousands. Clustering such a matrix is computationally intensive. Therefore, before clustering the spots, the data must be subjected to dimensionality reduction.
[0176] First, the PCA method is used to perform linear dimensionality reduction on the expression matrix to obtain PCA linear dimensionality reduction data. Then, the T-SNE nonlinear dimensionality reduction algorithm is used to reduce the PCA linear dimensionality reduction data to obtain T-SNE nonlinear dimensionality reduction data. Finally, the T-SNE nonlinear dimensionality reduction data is clustered based on the SNN clustering algorithm to obtain the cluster subgroups of genes with differential expression changes. The t-SNE diagram of the Spots clustering (such as Figure 4 ) and cluster spatial distribution map ( Figures 5A-5B ).
[0177] in, Figure 4 A and B are the t-SNE diagrams after dimensionality reduction and clustering of gene data in the normal group and model group slices, respectively. Different colors represent different classes, which are then mapped to the slices to intuitively see the relative positions of different clusters. Figure 4 In the t-SNE diagram of the normal group, red represents cluster 0, green represents cluster 1, and blue represents cluster 2; in the t-SNE diagram of the model group, red represents cluster 0, green represents cluster 1, blue represents cluster 2, and purple represents cluster 3.
[0178] The spatial transcriptome data of atrial muscle tissue of normal rats were first subjected to dimensionality reduction processing, and then clustered into three clusters using the SNN clustering algorithm (such as Figure 4 A in ); its cluster space distribution diagram is as follows Figure 5A The spatial transcriptome data of atrial muscle tissue of model group rats were first subjected to dimensionality reduction processing, and then clustered into 4 clusters using SNN clustering algorithm (such as Figure 4 B); The cluster spatial distribution maps of the spatial transcriptome data of the atrial muscle tissue of the model group rats are Figure 5B .Note: Figures 5A-5B Each point in the figure represents a spot, which details the distribution area of each type.
[0179] Each cluster represents the gene expression in the region where the cluster is located. By comparing the gene expression differences between different clusters, we can know the gene expression differences between the two regions. Through differential gene analysis between normal areas and diseased areas, we can obtain the gene set related to pathological changes in the diseased area.
[0180] 3-3. Analysis of Differentially Expressed Genes
[0181] The Wilcox rank sum test of Seurat software was used to screen differentially expressed genes in different clusters. That is, the genes of each cluster of the three groups of rat atrial muscle tissue were compared with the genes of all other clusters. The genes expressed in more than 25% of the spots in at least one comparison group were screened for differential comparative analysis to obtain the differential gene set of each cluster.
[0182] The difference screening criteria of the present invention are that genes with Fold Change ≥ 2 and q-value < 0.1 are genes with large changes in characteristic expression differences (ie, clustered differential gene sets).
[0183] For example, the genes of cluster 2 of the model group sample were compared with all genes of other clusters (0, 1, 3) on the same slice to screen out the differential gene set of cluster 2.
[0184] 3-4. Functional enrichment analysis
[0185] With the development of high-throughput technologies, biomedical research has entered the omics era, where single-gene studies are no longer sufficient. However, such vast amounts of data present new challenges for the effective extraction and analysis of information. For example, analysis of sequencing data often yields lists of differentially expressed genes or proteins. However, for many researchers, linking this long list of genes or proteins to the biological phenomenon under investigation and its underlying mechanisms is challenging. One approach to addressing this challenge is to partition a gene or protein list into multiple components, thereby reducing the complexity of analysis. To address this issue, researchers have developed multiple annotation databases. To address this issue, researchers often perform enrichment analysis of gene function, hoping to identify biological pathways that play a key role in biological processes and thereby uncover and understand the underlying molecular mechanisms. This has led to the development of a variety of software tools.
[0186] Functional enrichment analysis can group hundreds or even thousands of genes, proteins, or other molecules into different pathways, reducing the complexity of the analysis. Activated pathways under two different experimental conditions are clearly more convincing than a simple list of genes or proteins.
[0187] 3-4-1. GO enrichment analysis of differentially expressed genes
[0188] The GO database (Gene Ontology Database) is a structured, standardized biological annotation system established by the GO Consortium in 2000. Its goal is to establish a standardized vocabulary for knowledge about genes and their products, applicable across species. The GO annotation system is a directed acyclic graph consisting of three main branches: biological process, molecular function, and cellular component.
[0189] The clustered differentially expressed gene sets for each cluster of the three groups of rat atrial muscle tissue samples were analyzed using the ClusterProfiler package in R. Biological process enrichment analysis (i.e., GO BP analysis) was performed to obtain the results of the biological process enrichment analysis. The enrichment analysis results (terms) were then visualized using bar charts and bubble charts. The biological processes of each cluster were compared across the different samples. Based on the differences in biological processes, the cluster with the greatest difference in biological processes was selected as the basis for further screening of differentially expressed clusters.
[0190] The results of biological process enrichment analysis of the three clusters of differentially expressed genes in the atrial muscle tissue of normal rats are as follows: Figures 6A-6C The results of biological process enrichment analysis of differentially expressed genes in four clusters of atrial muscle tissue in model group rats are as follows: Figures 6D-6G .
[0191] Biological process analysis of each cluster in the atrial myocardial tissue samples of the three groups of rats showed that compared with the normal group, cluster 2 in the model group and the QPSM-treated group was the main differential cluster, and was consistent with the abnormal location diagnosed by pathology (i.e., consistent with the abnormal location of the atrial myocardium of rats with atrial fibrillation shown by HE staining images). Therefore, cluster 2 in the model group and the QPSM-treated group was further determined to be the main differential group (cluster) associated with the formation of atrial myocardial fibrosis in atrial fibrillation. Through biological process analysis and combined with HE staining results, it was determined that cluster 2 in the model group and the QPSM-treated group was the main differential cluster associated with myocardial fibrosis.
[0192] Then, the Wilcox rank sum test of Seurat software was used to perform differential analysis on the genes of the main differential cluster cluster2 of the model group and the genes of other cluster subgroups of the model group (i.e., the collection of cluster0, cluster1, and cluster3 genes), and the differential gene set of the main differential cluster cluster2 of the model group was screened. It was found that the differential gene set of cluster2 of the model group mainly included genes such as Nppa, Apoe, Reg3b, Myl4, Bgn, Postn, Ftl1, Lbp, lgfbp4, and Fn1, that is, the differential gene set of atrial fibrosis (i.e., Nppa, Apoe, Reg3b, Myl4, Bgn, Postn, Ftl1, Lbp, lgfbp4, and Fn1, etc.) was obtained.
[0193] The differential analysis standard of the present invention is that the genes with Fold Change ≥ 2 and q-value < 0.1 are the differential gene set of the main differential cluster cluster 2.
[0194] For example, the genes of cluster 2 of the model group sample were compared with all genes of other clusters (0, 1, 3) on the same slice to screen out the differential gene set of cluster 2.
[0195] The biological process analysis (i.e. BP analysis in GO analysis) of the differential gene set of cluster2, the main differential cluster in the model group, was performed. It was found that the characteristic differential gene set of atrial myocardial fibrosis mainly involved biological processes such as aging, collagen fiber binding, complement activation, extracellular matrix, and response to mechanical stimulation. That is, atrial fibrillation myocardial fibrosis involves biological processes such as aging, collagen fiber binding, complement activation, extracellular matrix, and response to mechanical stimulation.
[0196] Further KEGG pathway analysis was performed on the differential gene set of cluster 2, the main differential cluster in the model group, and it was found that the differential genes in cluster 2 of the model group were related to the complement system, ECM-receptor interaction, and AGE-RAGE signaling pathway in diabetic complications, that is, atrial fibrillation myocardial fibrosis was related to the complement system, ECM-receptor interaction, and AGE-RAGE signaling pathway in diabetic complications.
Claims
1. An analysis device for atrial fibrillation myocardial fibrosis, characterized in that: include: The differentially expressed gene screening module is used to standardize the spatial transcription data of atrial myocardial tissue slice samples from the normal group and the atrial fibrillation model group, and to screen highly variable genes in the standardized data using the VST algorithm to obtain a differentially expressed gene set with expression changes from high to low; The dimensionality reduction processing module is used to perform linear dimensionality reduction processing on the data of the differentially expressed gene sets obtained by screening the normal group and the atrial fibrillation model group samples after standardization processing, and obtain PCA reduced dimensionality data; then the nonlinear dimensionality reduction algorithm of T-SNE is used to perform nonlinear dimensionality reduction processing on the PCA reduced dimensionality data, and obtain T-SNE reduced dimensionality data; The cluster analysis module is used to perform cluster analysis on the dimension-reduced data using the SNN clustering algorithm to obtain cluster subgroups of genes with differential expression changes in the normal group and the atrial fibrillation model group samples; The cluster differential gene set screening module is used to compare the genes of each cluster subgroup of the normal group and the atrial fibrillation model group with the genes of all other cluster subgroups of the tissue samples of the corresponding group using the Wilcox rank sum test of the Seurat software, and screen the differential gene sets of each cluster, that is, the gene sets with the largest expression differences in each cluster; The main differential cluster screening module is used to perform differential gene enrichment analysis on the cluster differential gene sets of each cluster of the normal group and the atrial fibrillation model group samples using the ClusterProfiler package of R language, and compare the enrichment analysis results of the normal group and the atrial fibrillation model group to screen and obtain the main differential clusters; The main differential cluster analysis module is used to perform differential gene function enrichment analysis on the cluster differential gene sets of the main differential clusters of the atrial fibrillation model group and other cluster subgroups of the atrial fibrillation model group, that is, to perform biological process enrichment analysis on the differential genes of the main differential clusters, that is, to perform biological process analysis of the GO annotation system, and obtain the biological processes that lead to atrial fibrillation myocardial fibrosis involving aging, collagen fiber binding, complement activation, extracellular matrix, and response to mechanical stimulation; KEGG pathway analysis was performed on the clustered differential gene sets of the main differential clusters of the atrial fibrillation model group and other cluster subgroups of the atrial fibrillation model group, and it was obtained that atrial fibrillation myocardial fibrosis was related to the complement system, ECM-receptor interaction, and AGE-RAGE signaling pathway in diabetic complications. The clustered differential gene sets of the main differential clusters analyzed by the main differential cluster analysis module were Nppa, Apoe, Reg3b, Myl4, Bgn, Postn, Ftl1, Lbp, lgfbp4 and Fn1.
2. The analyzing device according to claim 1, wherein: It also includes a spatial transcription module for using spatial transcriptome technology to perform spatial transcription sequencing on atrial muscle tissue slice samples of animals in the normal group and atrial fibrillation model group, respectively, to obtain spatial transcription data of the animal samples in the normal group and atrial fibrillation model group.
3. The analyzing device according to claim 1 or 2, wherein: The cluster analysis module clusters the differentially expressed genes of the normal group samples into three clusters, namely, cluster subgroups; and clusters the differentially expressed genes of the atrial fibrillation model group samples into four clusters.
4. The analyzing device according to claim 1 or 2, wherein: The difference comparison criteria in the cluster differential gene set screening module are Fold Change ≥ 2 and q-value < 0.1, that is, genes with Fold Change ≥ 2 and q-value < 0.1 are selected as genes with large changes in characteristic expression differences.
5. The analyzing device according to claim 1 or 2, wherein: The main difference cluster screening module is specifically used to: use the R language ClusterProfiler package to perform biological process enrichment analysis on the cluster difference gene sets of each cluster of normal group and model group samples, and compare the biological process enrichment analysis results of the normal group and the atrial fibrillation model group to screen and obtain the main difference clusters.
6. A method for analyzing myocardial fibrosis in atrial fibrillation, characterized by: The steps include: 1) Spatial transcriptome technology was used to perform spatial transcriptome sequencing on atrial muscle tissue sections of animals in the normal group and the atrial fibrillation model group to obtain spatial transcriptome data of samples from the normal group and the atrial fibrillation model group; 2) The spatial transcriptome data of each sample were converted into expression matrix data for standardization using the sctransform (SCT) algorithm in the Seurat package of the R language. The standardized data were then filtered for highly variable genes using the vst algorithm to obtain a set of differentially expressed genes. 3) The differentially expressed gene sets screened for each sample were normalized using the sctransform (SCT) algorithm in the Seurat package of the R language. The normalized data were then subjected to linear dimensionality reduction using the PCA method to obtain PCA-reduced data. The PCA-reduced data were then subjected to nonlinear dimensionality reduction using the T-SNE algorithm to obtain T-SNE-reduced data. 4) Cluster analysis was performed on the T-SNE dimensionality reduction data based on the SNN clustering algorithm to obtain the differentially expressed gene cluster subgroups (Clusters) between the normal group and the atrial fibrillation model group; 5) Seurat software's Wilcox rank sum test was used to compare the genes of each cluster subgroup of the normal group and the atrial fibrillation model group with the genes of all other cluster subgroups of the corresponding group samples, and the cluster differential gene set of each cluster was screened, that is, the gene set with the largest expression difference in the cluster; 6) The ClusterProfiler package in R language was used to perform differential gene enrichment analysis on the cluster differential gene sets of each cluster of the normal group and the atrial fibrillation model group samples, and the differential gene enrichment analysis results of the normal group and the atrial fibrillation model group were compared to screen for the main differential clusters; 7) Differential gene function enrichment analysis was performed on the clustered differential gene sets of the main differential clusters of the atrial fibrillation model group samples and other clusters of the atrial fibrillation model group samples, and the biological processes leading to atrial fibrillation myocardial fibrosis lesions involving aging, collagen fiber binding, complement activation, extracellular matrix, and response to mechanical stimulation were obtained; KEGG pathway analysis was performed on the clustered differential gene sets of the main differential clusters of the model group and other cluster subgroups of the atrial fibrillation model group, and the AGE-RAGE signaling pathways in atrial fibrillation myocardial fibrosis and complement system, ECM-receptor interaction, and diabetic complications were obtained. Among them, the clustered differential gene sets of the main differential clusters were Nppa, Apoe, Reg3b, Myl4, Bgn, Postn, Ftl1, Lbp, lgfbp4 and Fn1.
7. The method according to claim 6, wherein: The spatial transcriptome sequencing process in step 1) comprises the following steps: 1A) performing tissue permeabilization on atrial muscle tissue slices of animals in the normal group and the atrial fibrillation model group, respectively, to release mRNA in the cells and bind to the corresponding capture probes on the spatial transcriptome gene expression chip; 1B) Reverse transcription of the mRNA bound to the capture probe of the gene expression chip is performed to synthesize a complete first-strand cDNA; second-strand cDNA is then synthesized; incubation is then performed to denature the cDNA; and finally, the denatured cDNA is recovered, amplified, and purified to obtain amplified cDNA; 1C) The amplified cDNA was subjected to fragmentation, end-repair and A-tailing, magnetic bead paired-end fragment screening, adapter ligation, magnetic bead purification after adapter ligation, sample index PCR, and magnetic bead paired-end fragment screening after PCR to construct the Visium spatial gene expression library; 1D) The samples of the Visium spatial gene expression library of the normal group and the atrial fibrillation model animal group were sequenced using the Illumina NovaSeq 6000 platform to obtain spatial transcriptome sequencing data for each sample.
Citation Information
Patent Citations
Method for determining gene expression regulation mechanism based on unicellular transcriptome data
CN111613268A
Analysis method of spatial transcriptome sequencing data
CN112522371A
Cell subset annotation method based on single cell transcriptome sequencing
CN112700820A