Method for predicting influence of graphene quantum dots on organism based on transcriptome sequencing

By screening differentially expressed genes through transcriptome sequencing technology and monitoring the effects of graphene quantum dots on mouse bone marrow MSCs, the difficulty in detecting the effects of graphene quantum dots on organisms was solved, and the detection sensitivity and biosafety analysis capabilities were improved.

CN120666012APending Publication Date: 2025-09-19QIQIHAR MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510826677.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies have not yet been able to effectively monitor and predict the effects of graphene quantum dots on biological organisms, especially on mesenchymal stem cells, and long-term exposure to a graphene environment may produce cytotoxicity and unclear effects on the immune system.

Method used

By constructing the original gene library data of mouse bone marrow mesenchymal stem cells stimulated by graphene quantum dots and unstimulated cells, transcriptome sequencing was performed, differentially expressed genes were screened and functional annotations were performed, and reagents for genes such as Serpine1, Nlrp3, and Cxcl10 were used to monitor the effects of graphene quantum dots.

Benefits of technology

The study revealed the effects of graphene quantum dots on the biological processes of mouse bone marrow MSCs, provided highly sensitive detection indicators, analyzed their potential effects on immune function, biological development, and cell division processes, and improved the detection capabilities of graphene biosafety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120666012A_ABST
    Figure CN120666012A_ABST
Patent Text Reader

Abstract

The invention provides a method for predicting the influence of graphene quantum dots on biological organisms based on transcriptome sequencing, and belongs to the technical field of gene transcriptomics. The invention adopts a high-throughput transcriptome sequencing technology to analyze the influence of 10 [mu] g / ml GOQD on the mouse bone marrow MSC transcriptome level, and reveals the potential results of the biological process caused by the stimulation of 10 [mu] g / ml GOQD on mouse bone marrow MSC, molecular function and molecular biology signal pathway. The invention further discloses the influence of 10mu g / ml GOQD on the immune function, biological development and cell division process and cardiovascular generation and growth process of mice and an action target gene of the GOQD. The hub protein coded by the acting target gene is screened from the perspective of proteomics, and a detection index is provided for monitoring the influence of GOQD on organisms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of gene transcriptomics, and in particular relates to a method for predicting the effects of graphene quantum dots on biological organisms based on transcriptome sequencing. Background Art

[0002] Mesenchymal stem cells (MSCs) are abundant in tissues such as human bone marrow, placenta, and umbilical cord blood. They accelerate tissue repair and promote functional reconstruction in organs like the heart and nervous system. They have great potential for treating bone or cartilage defects, muscle or tendon injuries, and cardiovascular and cerebral diseases. Natural compounds or nanomaterials, such as polymers, can promote the differentiation of MSCs into osteoblasts or angiogens.

[0003] Graphene is classified into graphene oxide, redox graphene, and graphene quantum dots. Graphene exhibits excellent cytocompatibility, low toxicity, excellent antibacterial properties, and the ability to promote angiogenesis and remodel the extracellular matrix. It has applications in medicine, tissue engineering, and cell tracking. Studies have shown that graphene combined with MSCs can effectively combat bacteria and inflammation, promoting the repair of skin ulcers and making it a promising skin substitute. Electrospun polylactic-glycolic acid nanofiber membranes doped with graphene oxide promoted functional tendon-bone reattachment in a rabbit rotator cuff injury model, effectively maintaining shoulder joint stability, repairing shoulder injuries, and maintaining good shoulder kinematic function. Graphene quantum dots (GQDs), with diameters less than 10 nm, are zero-dimensional carbon-based nanomaterials characterized by high biocompatibility, low cytotoxicity, and excellent optical properties. GQDs, through π-π bonds and physical adsorption, can link antibiotics and tumor-targeting drugs, making them excellent drug carriers or potent pharmaceutical formulation materials. They have demonstrated promising results in the treatment of diseases such as Alzheimer's disease and malignant tumors. Graphene oxide quantum dots (GOQDs) enhance autophagy in human dental pulp mesenchymal stem cells (DPSCs), promoting DPSC mineralization and aiding in tooth repair and mandibular reconstruction. Therefore, ultrathin two-dimensional nanomaterials such as graphene and graphene quantum dots have broad practical applications in various tissue engineering scaffolds, including bone and blood vessel scaffolds.

[0004] With the continuous advancement of science and technology, the production and use of industrial chemicals has increased, leading to the emergence of a large number of emerging environmental pollutants with unknown toxicological properties. While plastic products have brought convenience to our lives, recent reports have found microplastic particles with a diameter of less than 10μm in the blood of many volunteers worldwide. Microplastic particles in the human body pose a serious health risk to the human nervous system, reproductive system, and locomotor system. Graphene and graphene quantum dots, which are even smaller in diameter than microplastics, are widely used in products such as batteries, heating equipment, displays, touch screens, electronic products, and water purification materials, and are closely related to people's daily lives. Although GQDs have good biocompatibility with MSCs, prolonged exposure to graphene may cause cytotoxicity, and the effects of graphene quantum dots on the human immune system and mesenchymal stem cells are still unclear.

[0005] The transcriptome is the collection of all RNA (transcripts) transcribed by a specific species, tissue, or cell type during a specific period or under certain conditions, including mRNA and non-coding RNA. Transcriptome sequencing, as it's currently known, involves sequencing and analyzing all mRNAs. Summary of the Invention

[0006] In view of this, the object of the present invention is to provide a method for predicting the effects of graphene quantum dots on biological organisms based on transcriptome sequencing.

[0007] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:

[0008] The present invention provides the use of a reagent for detecting Serpine1, Nlrp3, Cxcl10, Ccl5, Ptprc, Thbs1, and Igf1 proteins in preparing a product for monitoring the effects of graphene quantum dots on organisms.

[0009] The present invention also provides the use of a reagent for detecting Thbs1, Emilin1, Pgf, Vegfd, Ccl2, Cxcl10, Ecm1, Serpine1, Ccl5, Dab2ip, Ednra, Grem1, Igf1, Ang, Ang2, Angpt1, Cacna1c, Cd34, Csf2, Cxcl9, Lif, Prkcb, Ceacam1, Ceacam2, Fancd2, Lif, Angpt1, Ascl2, Casp6, Cd177, Gsn, Jag2, and Vegfd genes in the preparation of a product for monitoring the effects of graphene quantum dots on organisms.

[0010] The present invention also provides a method for predicting the effect of graphene quantum dots on biological organisms based on transcriptome sequencing, comprising the following steps:

[0011] 1) Construct the original gene library data of mouse bone marrow mesenchymal stem cells stimulated with 10 μg / ml graphene quantum dots and mouse bone marrow mesenchymal stem cells without graphene quantum dots stimulation, filter the original gene library data, evaluate the data quality, and clean the data to obtain valid data;

[0012] 2) Compare the valid data to the genome, perform gene expression analysis, screen the expressed genes for differentially expressed genes, and perform functional annotation based on the differentially expressed genes.

[0013] Preferably, the method for constructing the original gene library data in step 1) comprises the following steps: extracting RNA from mouse bone marrow mesenchymal stem cells stimulated with 10 μg / ml graphene quantum dots and mouse bone marrow mesenchymal stem cells not stimulated with graphene quantum dots → mRNA enrichment → double-stranded cDNA synthesis → end repair and addition of A and adapters → fragment selection and PCR amplification → library detection → Illumina sequencing;

[0014] The mRNA enrichment tool includes Oligo(dT)Beads, the first strand synthesis tool in the double-stranded cDNA synthesis includes First Strand Synthesis, and the second strand synthesis tool in the double-stranded cDNA synthesis includes Second Strand Synthesis Mix; the library detection tool includes Qseq100 DNA Analyzer, and the Illumina sequencing tool includes NovaSeq TM X Plus, PE150.

[0015] Preferably, the content of the data filtering in step 1) includes:

[0016] ①Remove reads with adapters;

[0017] ② Remove reads with N ratio greater than 0.002;

[0018] ③ When the number of low-quality bases in a single-end read exceeds 50% of the read length, the paired reads need to be removed.

[0019] Preferably, the data quality assessment in step 1) includes sequencing quality statistics, sequencing error rate and GC content analysis; the sequencing error rate tool includes Illunima Casava 1.8; and the data cleaning includes fusion gene and variant site analysis.

[0020] Preferably, the genome comparison in step 2) includes comparison statistics, genomic regional distribution, sequencing saturation analysis, transcript splicing and screening, and variable splicing analysis; the comparative statistics tools include Tophat2, Hisat2, STAR, and HISAT2, the transcript splicing and screening tools include Stringtie and gffcompare, and the variable splicing analysis tools include rMATS or JuncBASE.

[0021] Preferably, the gene expression analysis in step 2) includes expression level analysis, correlation analysis, combined with other omics analysis, regulatory network analysis and genome visualization; the correlation analysis includes principal component analysis.

[0022] Preferably, the tools for differential gene screening in step 2) include edgeR, DEseq2 and DEGseq.

[0023] Preferably, the functional annotation in step 2) includes GO, KEGG, and Reactome enrichment analysis, and the tools for the GO, KEGG, and Reactome enrichment analysis include clusterProfiler; the GO, KEGG, and Reactome enrichment analysis include GO functional enrichment, pathway enrichment, Reactome enrichment, protein interaction network analysis, GSEA analysis, mouse immune co-expression gene analysis, and structural variation analysis; the protein interaction network analysis tool includes the STRING protein interaction database, and the mouse immune co-expression gene analysis tool includes Geneontology, Venny 2.1, Metascape, TBtools-II, String database, Cytoscape3.8.0, and Cytohubba v0.1.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] The present invention uses high-throughput transcriptome sequencing technology, employing the alignment software Hisat2; the splicing software Stringtie; the quantification software HTSeq; the differential analysis software edgeR (differential gene screening threshold: |log2FC|>1 and padj<0.05); and the variable splicing analysis software rMATS. The study investigated the effects of 10 μg / ml GOQD on the transcriptome level of mouse bone marrow MSCs, revealing the biological processes induced by 10 μg / ml GOQD stimulation of mouse bone marrow MSCs, and the potential results of molecular functions and molecular biological signaling pathways. The study further revealed the effects of 10 μg / ml GOQD on mouse immune function, biological development and cell division processes, and cardiovascular growth processes, as well as its target genes. Proteomics screened the hub proteins encoded by the target genes, providing detection indicators for monitoring the effects of GOQD on organisms.

[0026] The present invention adopts transcriptome sequencing experiments, which can not only use high-throughput sequencing technology to comprehensively and quickly obtain known gene expression abundance changes, but also accurately analyze the SNP / InDel (coding sequence single nucleotide polymorphism / insertion and deletion), alternative splicing, fusion gene and other sequence and structural variations of the transcripts through the measured sequence information. In addition, by increasing the amount of sequencing data, transcriptome sequencing also has its unique advantages for detecting low-abundance transcripts and discovering new transcripts (and then predicting new genes). The difference in gene expression of mouse bone marrow mesenchymal stem cells treated with 10 μg / ml GOQD reflects the accuracy and precision of transcriptomic sequencing, and the analysis of the gene ontology and signal pathway mechanism of these differential genes reflects the advantage of high detection sensitivity, which has high detection sensitivity for predicting the effect of GOQD on mouse bone marrow mesenchymal stem cells and judging the biological safety of GOQD.

[0027] In addition to providing research data for studying the effects of graphene and its derivatives on human cells and analyzing the biosafety impact of graphene, the method of the present invention also provides a detection or diagnostic basis for the prevention and treatment of occupational diseases among producers exposed to a graphene environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is the flowchart of library construction and sequencing;

[0029] Figure 2 This is a diagram of the library construction principle (Adapter, including P5 / P7, index and Rd1SP / Rd2SP. P5 / P7 are used in the sequencer to perform bridge PCR; Index is the tag sequence used to distinguish different library information; Rd1SP / Rd2SP (sequence primer) is the sequencing primer binding region. The sequencing process theoretically starts from Rd1 / Rd2 and proceeds backward).

[0030] Figure 3 It is an information analysis flow chart;

[0031] Figure 4 The original transcriptome sequencing composition of mBMSCs in the control group (0 μg / ml GOQD) and the 10 μg / ml GOQD group (A1 to A3 are the control group (0 μg / ml GOQD), B1 to B3 are the 10 μg / ml GOQD group, and the different color ratios represent the proportions of different components, (1) Adapter-related: the proportion of reads with adapters, (2) ContainingN: the proportion of reads with N bases, (3) Low quality: the proportion of reads with low sequencing quality, (4) Clean reads: the proportion of clean reads);

[0032] Figure 5 Figure 3 is the sequencing error rate distribution of mBMSC in the control group (0 μg / ml GOQD) and the 10 μg / ml GOQD group (A1 to A3 are the control group (0 μg / ml GOQD), B1 to B3 are the 10 μg / ml GOQD group, the horizontal axis is the base position of the reads, and the vertical axis is the average error rate of the base at that position for all reads. The left side of the dotted line (0 to 150 bp) is the error rate distribution of Read1, and the right side (150 to 300 bp) is the error rate distribution of Read2. The sequencing error rate distribution check is used to detect whether there are bases at certain positions within the sequencing length range with abnormally high error rates. For example, if the sequencing error rate of the base in the middle position of Read1 or Read2 is significantly higher than that of other positions, it indicates that an abnormal base may exist).

[0033] Figure 6 The GC content distribution of the control group (0 μg / ml GOQD) and the 10 μg / ml GOQD group (A1 to A3 are the control group (0 μg / ml GOQD), B1 to B3 are the 10 μg / ml GOQD group, the horizontal axis is the base position of the reads, the vertical axis is the proportion of single bases, and different colors represent different base types);

[0034] Figure 7 The distribution of the genome in different regions of the control group (0 μg / ml GOQD) and the 10 μg / ml GOQD group (A1 to A3 are the control group (0 μg / ml GOQD), and B1 to B3 are the 10 μg / ml GOQD group);

[0035] Figure 8Saturation curves of gene expression in the control group (0 μg / ml GOQD) and the 10 μg / ml GOQD group (A1 to A3 are the control group (0 μg / ml GOQD), B1 to B3 are the 10 μg / ml GOQD group, the horizontal axis is the number of reads, and the vertical axis is the number of detected junctions (mRNA splicing))

[0036] Figure 9 It is the display of new transcript results (step 1. Transcript exon number screening: filter low-confidence single-exon transcripts in transcript splicing results and select transcripts with exon number ≥ 2; step 2. Transcript length screening: select transcripts with transcript length greater than 200bp; step 3. Transcript known annotation screening: use gffcompare software to screen out transcripts that overlap with database annotated exon regions);

[0037] Figure 10 is the FPKM information of the genes of each sample (the first column is the gene ID (ENSG, G refers to gene, the first three letters (ENS), meaning "ENSEMBL"), and each column thereafter is the FPKM value corresponding to the gene ID of each sample, A control group (0 μg / ml GOQD), B 10 μg / ml GOQD);

[0038] Figure 11 The box plots of the expression level distribution of each sample (A is the control group (0 μg / ml GOQD), B is 10 μg / ml GOQD; the horizontal axis in the figure is the sample name / group name; the vertical axis represents the log10 (FPKM) value; the box plot in each area represents four statistical quantities (maximum value, upper quartile, lower quartile and minimum value from top to bottom, respectively). This figure measures the differences between samples from the perspective of the overall discreteness of the expression level);

[0039] Figure 12 The expression level distribution density distribution diagram of each sample (A is the control group (0 μg / ml GOQD), and B is the probability density distribution diagram of the expression of all genes in the 10 μg / ml GOQD group. The horizontal axis in the figure is log10 (FPKM). The higher the value, the higher the gene expression level. The vertical axis is the density of genes, that is, the number of genes corresponding to the expression level on the horizontal axis / the total number of expressed genes detected. Each color in the figure represents a sample, and the sum of all probabilities is 1, that is, the area of ​​each region is 1. The peak of the density curve represents the area with the most concentrated gene expression in the entire sample).

[0040] Figure 13 is the Upset graph of genes identified in each sample (A is the control group (0 μg / ml GOQD), B is 10 μg / ml GOQD);

[0041] Figure 14 is a heat map of the correlation between samples (A is the control group (0 μg / ml GOQD), B is 10 μg / ml GOQD, the horizontal and vertical axes in the figure are the sample names, and the color represents the size of the Pirsh correlation coefficient R);

[0042] Figure 15 is the PCA cluster diagram of each sample (A is the control group (0 μg / ml GOQD), B is 10 μg / ml GOQD);

[0043] Figure 16 It is a list of differentially expressed genes in each sample ((1) gene_id: gene ID (2) gene_name: gene name (3) locus (chr Start End): gene location on chromosome (4) group1_FPKM: FPKM value of comparison combination 1 (5) group2_FPKM: FPKM value of comparison combination 2 (6) log2FoldChange: the difference fold value between two samples or comparison combinations, expressed as log2(group1 / gro up2), calculated by differential analysis software (7) pvalue: p value of significance test (8) padj: p value after correction for multiple hypothesis testing (9) level: level of difference change up / down);

[0044] Figure 17 The left side shows the statistical graph of differentially expressed genes in each group (the left side shows the statistical bar graph of differentially expressed genes; the horizontal axis shows different differential groups, the vertical axis shows the number of differentially expressed genes, the red column represents the number of genes with increased differential expression, and the green column represents the number of genes with decreased differential expression; the right side shows the dot graph of differentially expressed genes; the horizontal axis shows different differential groups, the vertical axis shows the gene change fold, red indicates increased differential expression, green indicates decreased differential expression, and the 10 genes with the largest difference fold are selected for each group, and the gene names are marked in the graph);

[0045] Figure 18 This is the volcano plot of differentially expressed genes in the 10 μg / ml GOQD group (the horizontal axis represents the expression fold change (log2FoldChange) of genes between different samples or comparison combinations. The larger the absolute value of the horizontal axis, the greater the expression fold change between the two comparison combinations; the vertical axis represents the significance level of the expression difference; up-regulated genes are represented by red dots, down-regulated genes are represented by green dots, and gray dots represent genes with no significant changes).

[0046] Figure 19This is a scatter plot of differentially expressed genes in the 10 μg / ml GOQD group (each point in the figure represents a gene, the horizontal axis represents the expression level of the control group samples, expressed as log10(FPKM+1) (if there are multiple samples in each group, the average is taken), and the vertical axis represents the expression level of the experimental group samples; genes with significant differential expression are represented by red and green points (red indicates increased differential expression, green indicates decreased differential expression), and genes with no significant differential expression are represented by gray points);

[0047] Figure 20 This is the hierarchical clustering heat map of differentially expressed genes in the 10 μg / ml GOQD group (the horizontal axis is the sample, the vertical axis is the differentially expressed genes, the left side clusters the genes according to the degree of expression similarity, and the expression level gradually increases from blue to red);

[0048] Figure 21 The figure shows the circle diagram of differentially expressed genes in the 10 μg / ml GOQD group (the outermost circle is the chromosome band, the innermost circle is the histogram of the log2FoldChange value of the differentially expressed genes, and the middle two circles are the log2(fpkm+1) distribution diagrams of the two groups of genes);

[0049] Figure 22 is the GO function enrichment of differentially expressed genes in the 10 μg / ml GOQD group (A is a histogram of all GO function enrichment of the top 10 differentially expressed genes in the 10 μg / ml GOQD group, B is a histogram of up-regulated GO function enrichment of the top 10 differentially expressed genes in the 10 μg / ml GOQD group, C is a histogram of down-regulated GO function enrichment of the top 10 differentially expressed genes in the 10 μg / ml GOQD group, the ordinate is GO, the abscissa is the number of differentially expressed genes enriched in the GO, asterisks are significant marks (*P<0.05, **P<0.01, ***P<0.001), colors represent -log10 (p value) (up, down, all plotted separately), CC: cellular component; BP: biological process; MF: molecular function);

[0050] Figure 23 is the KEGG pathway enrichment of differentially expressed genes in the 10 μg / ml GOQD group (A is a bar graph of all KEGG pathway enrichment of the top 20 differentially expressed genes in the 10 μg / ml GOQD group, B is a bar graph of up-regulated KEGG pathway enrichment of the top 20 differentially expressed genes in the 10 μg / ml GOQD group, C is a bar graph of down-regulated KEGG pathway enrichment of the top 20 differentially expressed genes in the 10 μg / ml GOQD group, the ordinate is GO, the abscissa is the number of differentially expressed genes enriched in the GO, asterisks are significant markers (*P<0.05, **P<0.01, ***P<0.001), and colors represent -log10 (pvalue) (up, down, all plotted separately));

[0051] Figure 24 is the Reactome enrichment of differentially expressed genes in the 10 μg / ml GOQD group (A is a histogram of all Reactome pathway enrichment of the top 20 differentially expressed genes in the 10 μg / ml GOQD group, B is a histogram of up-regulated Reactome pathway enrichment of the top 20 differentially expressed genes in the 10 μg / ml GOQD group, and C is a histogram of down-regulated Reactome pathway enrichment of the top 20 differentially expressed genes in the 10 μg / ml GOQD group. The ordinate is GO, the abscissa is the number of differentially expressed genes enriched in the GO, asterisks are significant markers (*P<0.05, **P<0.01, ***P<0.001), and colors represent -log10 (pvalue) (up, down, and all are plotted separately));

[0052] Figure 25 This is the PPI result of the 10 μg / ml GOQD group;

[0053] Figure 26 The GSEA analysis of the 10 μg / ml GOQD group ((1) ID: pathway number, (2) Description: name of the pathway classification, (3) setSize: number of genes in the expression dataset included in the pathway entry, (4) enrichmentScore: enrichment score, (5) NES: normalized ES value after correction. Since the number of gene sets in the gene database files input by different users may be different, the normalization of the enrichment score takes into account the number and size of the gene sets, (6) pvalue: statistical significance level of the enrichment score ES, which is used to characterize the credibility of the enrichment result, (7) p.adjust: calibrated pvalue, (8) qvalu e: q value for statistical test of pvalue (9) rank: pathway ranking, (10) tags indicates the proportion of core genes to the total number of genes in the gene set, while list indicates the proportion of core genes to the total number of all genes. Signal is calculated using these two indicators. When the number of core genes is the same as the total number of genes in the gene set, the signal value is the largest. When the number of genes in the gene set is close to the number of all genes, the signal value is close to 0. (11) core_enrichment: core genes, pathways with |NES|>1, p-value<0.05, and p.adjust<0.25 are significantly enriched);

[0054] Figure 27 is the Veeny plot of co-expressed genes between GOQD-DEG and mouse immunity;

[0055] Figure 28is the enrichment analysis of 155 co-expressed genes (where A is a histogram of significant biological pathways of co-expressed genes, with numbers indicating the number of biological processes included in each pathway; B is a histogram of significant molecular mechanisms of co-expressed genes, with numbers indicating the number of molecular mechanisms included in each pathway; C is a histogram of significant KEGG signaling pathways of co-expressed genes, with numbers indicating the number of KEGG signaling pathways included in each pathway);

[0056] Figure 29 is the biological development and cardiovascular biological process in which co-expressed genes and GOQD-DEGs participate (A is a bubble map of the relationship between GQD-DEGs and cardiovascular biological processes, the size of the circle is proportional to the number of enriched genes, the redder the color, the smaller the P value, B is a heat map of the relationship between co-expressed genes and GQD-DEGs participating in cardiovascular biological processes (median gene expression frequency = 1, median biological process = 5), red fonts indicate co-expressed genes, C is a bubble map of the relationship between GQD-DEGs and biological development and cell division processes, the size of the circle is proportional to the number of enriched genes, the redder the color, the smaller the P value, D is a heat map of the relationship between co-expressed genes and GQD-DEGs participating in biological development and cell division processes (median gene expression frequency = 1, median biological process = 4), red fonts indicate co-expressed genes);

[0057] Figure 30 It is the protein-protein interaction network (PPI) and hub proteins (where A is the PPI network encoded by 155 co-expressed genes, B1 to B4 are the top 20 hub proteins screened based on the BettleNeck, Degree, EPC, and MCC values, respectively, and C is the Venny diagram of hub proteins);

[0058] Figure 31It is a structural variation analysis ((1)ID: alternative splicing event number, (2)GeneID: gene number where the alternative splicing event is located, (3)geneSymbol: gene name where the alternative splicing event is located, (4)chr: chromosome where the alternative splicing event is located, (5)str and: direction of the chain where the alternative splicing event is located, (6)exonStart_0base: exon start position where the alternative splicing event occurs, (7)exonEnd: exon end position where the alternative splicing event occurs, (8)upstreamES: upstream exon start position where the alternative splicing event occurs, (9)upstreamEE: upstream exon end position where the alternative splicing event occurs, (10)downstreamES: downstream exon start position where the alternative splicing event occurs, (11)downstreamEE: downstream exon end position where the alternative splicing event occurs.(12) ID: AS event number (same as 1, can be ignored), (13) IC_SAMPLE_1: expression level of the alternative splicing event Exon Inclusion Isoform in SAMPLE_1 (comparison combination 1), (14) SC_SAMPLE_1: expression level of the alternative splicing event Exon Skipping Isoform in SAMPLE_1 (comparison combination 1), (15) IC_SAMPLE_2: expression level of the alternative splicing event Exon Inclusion Isoform in SAMPLE_2 (comparison combination 2), (16) SC_SAMPLE_2: expression level of the alternative splicing event Exon Skipping Isoform in SAMPLE_2 (comparison combination 2), (17) IncFormLen: effective length of the alternative splicing event Exo n Inclusion Isoform, (18) SkipFormLen: alternative splicing event Exon Skipping The effective length of the isoform, (19) PValue: the p-value of the significance of the expression difference of the alternative splicing event, (20) FDR: the FDR value of the significance of the expression difference of the alternative splicing event, that is, the corrected p-value, (21) IncLevel1: the ratio of the Exon Inclusion Isoform of the alternative splicing event in SAMPLE_1 to the total expression of the two isoforms, that is, the proportion of the inclusion isoform in comparison combination 1, (22) IncLevel2: the ratio of the Exon Inclusion Isoform of the alternative splicing event in SAMPLE_2 to the total expression of the two isoforms, that is, the proportion of the inclusion isoform in comparison combination 2, (23) IncLevelDifference: the difference between IncLevel1 and IncLevel2, that is, the difference in the proportion of the inclusion isoform between the two comparison combinations). DETAILED DESCRIPTION

[0059] The technical solutions provided by the present invention are described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.

[0060] Example 1

[0061] 1. Cell culture and grouping: 5.0×10 6Mouse bone marrow mesenchymal stem cells were cultured in RPMI-1640 medium supplemented with 10% fetal bovine serum at 37°C and 5% CO₂ for 12 hours. The medium was then replaced with RPMI-1640 medium supplemented with 10 μg / ml graphene oxide quantum dots (GOQDs) and lacking fetal bovine serum, and cultured for another 12 hours at 37°C and 5% CO₂ before subsequent experiments. Mouse bone marrow mesenchymal stem cells stimulated with 10 μg / ml GOQDs served as the experimental group (GOQD group), while cells cultured normally served as the control group (CON group).

[0062] 2. Genome sequencing (RNA-seq) The above two groups of cells, 3 samples in each group, were commissioned to Hangzhou Kaitai Biotechnology Co., Ltd. for genome sequencing (RNA-seq). The experimental process mainly includes: from RNA sample extraction to final data acquisition, sample detection, library construction, sequencing and other links will directly affect the quantity and quality of the data, thereby affecting the results of subsequent information analysis. In order to ensure the accuracy and reliability of sequencing data from the source, we promise to strictly control all aspects of data production to ensure the output of high-quality data from the root. The flowchart of library construction and sequencing is as follows Figure 1 shown.

[0063] 2.1 RNA sample testing: RNA is extracted from tissues or cells, and then the RNA samples are strictly quality controlled. The quality control standards mainly include the following two aspects: (1) Agarose gel electrophoresis: analyze the integrity of the sample RNA and whether there is DNA contamination; (2) Nanodrop: perform preliminary quantification and detect RNA concentration and purity (OD260 / 280 and OD260 / 230).

[0064] 2.2 Library construction: After the sample is qualified, Oligo (dT) Beads are used to isolate and obtain mRNA; a shearing reagent is added to break the mRNA into short fragments, which are used as templates to synthesize the first cDNA chain using First Strand Synthesis, and the second chain is synthesized by adding Second Strand Synthesis Mix. After the ends are blunted and "A" is added, the RNA-SeqAdapter is connected, the connection product is purified and PCR amplified, and the 300bp-600bp target product is recovered using magnetic beads for sequencing. The library construction principle is as follows Figure 2 shown.

[0065] 2.3 Library Inspection: After library construction is completed, preliminary quantification is performed using Qubit and the library is diluted to 1 ng / μl. The insert size of the library is then detected using the Qseq100 DNA Analyzer. If the insert size meets the expectation, the effective concentration of the library is accurately quantified using the Q-PCR method (library effective concentration > 2 nM) to ensure library quality.

[0066] 2.4 Sequencing: After the library is qualified, different libraries are pooled according to the effective concentration and target data volume and then NovaSeq is performed. TM X Plus, PE150 sequencing.

[0067] 3. Bioinformatics analysis process: After sequencing obtains the raw sequence (Raw data), it enters the information analysis process, which is divided into two stages:

[0068] (1) Sequencing data quality assessment: mainly through statistics on sequencing error rate, data volume, alignment rate, etc., to assess whether the library construction and sequencing have met the standards. If it meets the standards, subsequent analysis will be carried out; otherwise, the library needs to be rebuilt or additional testing is required;

[0069] (2) Information mining and analysis: The standard RNA-seq analysis process includes quality control, alignment, quantification, difference significance analysis and functional enrichment, as well as variable splicing analysis and GSEA enrichment analysis. At the same time, we launch some personalized analysis content according to the needs of different research directions, including fusion genes, molecular typing, etc. The information analysis flow chart is as follows: Figure 3 shown.

[0070] 4. Standard Analysis Results: The standard RNA-seq analysis process primarily includes quality control, alignment, quantification, differential significance analysis, and functional enrichment. The core of RNA-seq analysis is the significance of gene expression differences. Statistical methods are used to compare gene expression differences between two or more conditions, identify differentially expressed genes associated with these conditions, and then further analyze the biological significance of these differentially expressed genes.

[0071] 4.1 Sequencing Data Quality Assessment

[0072] Sequencing data is generated through multiple steps, including RNA extraction, library construction, and sequencing. These steps can produce some low-quality or invalid data, such as deviations in library length during library construction and sequencing errors during sequencing. Therefore, quality control of the raw data obtained from the machine is necessary to ensure the accuracy of subsequent analysis.

[0073] 4.1.1 Raw sequence data

[0074] The raw image files obtained from high-throughput sequencing (such as Illumina PE150 / SE50) are converted into sequenced reads (SequencedReads) through CASAVA base recognition and stored in FASTQ format.

[0075] 4.1.2 Sequencing data filtering

[0076] The raw data obtained by sequencing contains a small number of reads with sequencing adapters or low sequencing quality, such as Figure 4 In order to ensure the quality and reliability of data analysis, the original data needs to be filtered. The filtering content is as follows:

[0077] ①Remove reads with adapters;

[0078] ② Remove reads with an N ratio greater than 0.002 (N indicates that the base information cannot be determined);

[0079] ③ When the number of low-quality bases in a single-end read exceeds 50% of the read length, the paired reads need to be removed.

[0080] 4.1.3 Sequencing Error Rate Distribution: The sequencing process itself has the possibility of machine errors. The sequencing error rate distribution can reflect the quality of the sequencing data. The sequencing quality value of each base in the sequence information is stored in the FASTQ file. If the base quality value of the read is expressed as QPhred, the sequencing error rate can be calculated as e = 10(-QPhred / 10) or expressed as QPhred = -10log10(e). The concise correspondence between base calls and Phred scores in Illunima Casava version 1.8 is as follows: Figure 5 shown.

[0081] 4.1.4 GC content distribution: The ratio of guanine (G) to cytosine (C) in a nucleotide sequence is called GC content. GC content has certain species specificity, but due to the 6bp random primers used in the reverse transcription process, there will be a certain preference in the nucleotide composition of the first few bases, resulting in normal fluctuations, which then tend to stabilize. Figure 6 shown.

[0082] 4.1.5 Data Quality Summary: After raw data filtering, sequencing error rate checks, and GC content distribution checks, a clean read data summary was obtained for subsequent analysis, as shown in Table 1.

[0083] Table 1. Cleanreads data summary

[0084]

[0085] 4.2 Alignment Analysis: Aligning clean reads to the genome or transcriptome is essential for subsequent analysis. RNA-Seq data were aligned using Tophat2 (http: / / tophat.cbcb.umd.edu; Trapnell C et al., 2009) / Hisat2 (http: / / ccb.jhu.edu / software / hisat2; Kim D et al., 2015) / STAR (http: / / code.google.com / p / rna-star; Dobin A et al., 2013) software. Data were aligned to the mouse genome (GRCm39) annotated in the Ensembl database (Mus_musculus.GRCm39.109.gtf) using HISAT2 (2.2.1) software with default parameters.

[0086] Alignment output is typically a standard SAM file, developed by Sanger. This file format uses tab delimiters and is primarily used to display the results of sequencing sequence alignments to genomes. BAM files are a binary format of SAM files. Because they occupy less storage space and offer faster computational speed, many alignment software programs use the BAM format as their standard output. The specific contents of BAM files can be viewed using the samtools tool.

[0087] 4.2.1 Alignment Statistics: Coverage statistics and other statistics are calculated using valid data mapped to the reference genome. Generally, when the alignment rate of sequencing reads (Total Mapped) for human samples is higher than 80%, and the number of reads with multiple alignments to the reference sequence (Multiple Mapped Reads) is less than 20%, the results are considered reliable and suitable for subsequent analysis. This is shown in Table 2.

[0088] Table 2 Comparison statistics

[0089]

[0090]

[0091]

[0092] 4.2.2 Distribution of genomic regions: Based on the alignment results, the proportion of reads in the exon region, intron region and intergenic region of the genome is counted separately. The gene annotation of general model species is relatively complete (such as humans and mice), and the proportion of reads aligned to the exon region is the highest. Reads aligned to the intron region may come from pre-mRNA or introns retained by alternative splicing events. Reads aligned to the intergenic region may come from ncRNA or a small amount of DNA fragment contamination, or the gene annotation may not be perfect. The distribution of sequencing reads of all samples in the genomic region is as follows Figure 7 shown.

[0093] 4.2.3 Sequencing saturation analysis: Sequencing saturation refers to the relationship between the amount of sequencing and the number of genes detected. The number of genes detected by transcriptome sequencing is positively correlated with the amount of sequencing data, that is, the larger the amount of sequencing data, the more genes are detected. However, the number of genes in a species is limited, and gene transcription has time-specificity and space-specificity, so as the amount of sequencing increases, the number of genes detected will tend to a stable value. In the saturation analysis results, the number of detected junctions is used to evaluate the sequencing saturation. The sequencing saturation analysis of all samples is as follows: Figure 8 shown.

[0094] Under normal circumstances, as the amount of sequencing data increases, the number of newly detected genes should become fewer and fewer, or even zero, and the number of junctions tends to be stable, that is, the expression of transcripts tends to be stable. In this case, we can consider that the sequencing data has reached saturation, providing sufficient effective data for subsequent information analysis and laying the foundation for accurate information analysis. Figure 8 It can be seen that as the number of sequencing reads increases, the number of detected junctions gradually increases and tends to be stable, that is, the amount of sequencing data meets the analysis requirements.

[0095] 4.2.4 Alignment Visualization: We provide RNA-seq clean_reads alignment results on the genome in bam format. For some species, we also provide the corresponding reference genome and annotation files. These files can be visualized using the IGV (Integrative Genomics Viewer). The IGV viewer displays the location of single or multiple reads on the genome at different scales and the read abundance of different regions at different scales, reflecting the transcriptional level of different regions, as well as some annotation information. IGV software can be downloaded from: http: / / software.broadinstitute.org / software / igv / .

[0096] 4.3 Transcript Assembly and Screening: The Stringtie software package (https: / / ccb.jhu.edu / software / stringtie / ) is suitable for transcript assembly, quantification, and differential expression analysis in RNA-Seq. After alignment, stringtie is first used to assemble transcripts from the RNA-Seq data. Subsequently, the merge option of stringtie is used to integrate these assembled files with the reference transcriptome annotation file into a unified annotation standard. During this process, gffcompare is used to compare the assembled transcripts with known transcripts to assess the quality of the assembly.

[0097] 4.3.1 Transcript Assembly: Stringtie uses alignment results to assemble transcripts, resulting in the smallest possible set of transcripts. Specific parameters are used for strand-specific libraries, providing accurate information on transcript strand orientation.

[0098] 4.3.2 Transcript screening: We used stringtie software to merge the transcripts obtained by splicing each sample, and removed transcripts with uncertain chain direction and transcript length less than 200 nt. Next, we used gffcompare software to compare with the known database and filter out the transcripts known in the database. Finally, we predicted the coding potential of the new transcripts screened, and finally obtained Novel_transcripts. Figure 9 Only some results are shown. Detailed results: Results / 3.Assembly / Novel_trans / *.

[0099] 4.4 Quantitative Analysis: Quantification is the foundation of differential expression significance analysis. Currently, quantification can be divided into gene-level and transcript-level quantification. Gene-level quantification is robust and reliable, and algorithmically easy to implement, but it cannot accurately identify the individual transcripts of each gene. Transcript-level quantification algorithms are more difficult to implement and are less accurate than gene-level quantification. However, in recent years, new algorithms have been developed, significantly improving accuracy.

[0100] 4.4.1 Explanation of quantitative results: The gene expression value of RNA-seq is usually expressed in RPKM or FPKM. RPKM is used for single-end sequencing, and FPKM is used for double-end sequencing. The sequencing depth is corrected first, and then the length of the gene or transcript is corrected. FPKM (Fragments Per Kilobase of transcript sequence per Millions basepairssequenced) is the number of fragments per kilobase length from a gene / transcript per million fragments. It takes into account the effects of sequencing depth and gene length on fragment counts. FPKM is currently the most commonly used method for estimating gene expression levels. Quantitative results such as Figure 10 Only part of the results are shown. Detailed results can be found in: Results / 4.Quantification / genes.FPKM.xls.

[0101] 4.4.2 Expression level distribution: After calculating the expression values ​​(RPKM or FPKM) of all genes or transcripts in each sample, the distribution of gene or transcript expression levels in different samples is displayed using a box plot. Figure 11 shown.

[0102] After calculating the expression values ​​(RPKM or FPKM) of all genes or transcripts in each sample, the distribution of gene or transcript expression levels in different samples is displayed through a density distribution map. Figure 12 shown.

[0103] According to the FPKM calculation result table of all genes, the expression levels were divided into different intervals, and the number of genes in different expression intervals of each sample was counted.

[0104] After calculating the expression values ​​(RPKM or FPKM) of all genes or transcripts in each sample, the statistical results of the common and unique genes among the samples are displayed in an Upset graph. Figure 13 shown.

[0105] 4.4.3 Correlation analysis: Based on the expression values ​​(RPKM or FPKM) of all genes in each sample, the correlation coefficients of samples within and between groups are calculated and plotted into a heat map, which can intuitively show the differences between samples and the duplication of samples within a group. The higher the correlation coefficient between samples, the closer their expression patterns are. The sample correlation heat map is as follows: Figure 14 shown.

[0106] Principal component analysis (PCA) is also commonly used to evaluate inter-group differences and intra-group sample duplication. PCA uses linear algebraic calculation methods to reduce the dimensionality and extract principal components of tens of thousands of genetic variables. Ideally, in a PCA graph, inter-group samples should be dispersed, while intra-group samples should be clustered together. Figure 15 shown.

[0107] 4.5 Differential Analysis: Analysis of transcriptome differential expression profiles under different conditions is of great significance for transcriptome research. For example, it involves grouping tumor and normal, mutant and wild-type samples in N / T samples, grouping before and after treatment and clinical stage based on clinical characteristics, and molecular typing and grouping based on the characteristics of the cancer itself.

[0108] 4.5.1 List of differentially expressed genes: After the quantitative analysis is completed, the expression matrix of all samples is obtained, and then the expression difference significance analysis at the gene or transcript level can be performed to find the functional genes or transcripts of interest. Common differential analysis software includes edgeR, DEseq2 and DEGseq, among which DEseq2 and DEGseq have certain limitations in sample biological repetition (the number of samples in each comparison combination). Therefore, we use edgeR software to perform expression difference significance analysis, and use padj less than 0.05 as the difference significance standard (if padj is less than 0.05 to screen out too few differences, use pvalue less than 0.05 for difference screening. For specific difference screening conditions, see the volcano chart). As the most common RNA sequencing differential analysis software for transcriptome research, edgeR has no biological repetition restrictions on the analyzed samples, and is not limited to the number of samples in each comparison combination (1 or more), and can perform differential analysis at the gene level or transcript level. The differential analysis results are as follows: Figure 16 Only part of the results are shown. Detailed results: Results / 5.Diff / Case_vs_Control / Case_vs_Control.diffgene.xls.

[0109] 4.5.2 Statistics of the number of differentially expressed genes in each group: Statistics were performed on the differentially expressed genes between different differential schemes, and the number of differentially expressed genes between different differential schemes was analyzed and visualized. Figure 17 shown.

[0110] 4.5.3 Volcano plot of differentially expressed genes: Volcano plot can visually display the overall distribution of genes with significant expression differences. Figure 18 As shown, in the 10 μg / ml GOQD group, 695 genes were downregulated and 384 genes were upregulated.

[0111] 4.5.4 Differential gene scatter plot: Use the average FPKM of the experimental group and the control group, and use log2(FPKM+1) for processing, and draw a scatter plot of the difference results to intuitively see the expression of differentially expressed genes. Figure 19 shown.

[0112] 4.5.5 Differential gene clustering: Clustering is another way to display differentially expressed genes. Genes with similar expression patterns are grouped together. These genes may have common functions or participate in common metabolic pathways and signaling pathways. We use the mainstream hierarchical clustering method to convert the log10(FPKM+1) value (scale number) and perform clustering. Figure 20 shown.

[0113] 4.5.6 Differential gene genome circle map: We use the R language ggbio package to mark RNA on the genome based on genomic information, overall RNA expression, and differential RNA expression analysis results. Figure 21 shown.

[0114] 4.6 GO / KEGG / Reactome Enrichment Analysis: In organisms, different genes coordinate with each other to carry out their biological functions. Pathway enrichment analysis can identify the most important biochemical metabolic pathways and signal transduction pathways involved in differentially expressed genes. Enrichment analysis of differentially expressed genes can reveal which pathways are differentially expressed between sample groups from a biological pathway perspective, and is a core component of transcriptome analysis. We used clusterProfiler (http: / / www.bioconductor.org / packages / release / bioc / html / clusterProfiler.html) software to perform pathway enrichment analysis on differentially expressed genes using the widely used annotation databases GO, KEGG, and Reactome (YuG et al., 2012). GO (Gene Ontology) is a comprehensive database describing gene function, which can be divided into three components: molecular function (MF), biological process (BP), and cellular component (CC). KEGG (Kyoto Encyclopedia of Genes and Genomes) is a comprehensive database that integrates genomic, chemical, and system functional information. The Reactome database brings together various human reactions and biological pathways.

[0115] 4.6.1 GO functional enrichment: GO (Gene Ontology) is a standard classification system for gene function and a comprehensive database that describes gene function. Studying the distribution of gene sets in GO may clarify the biological function of differential gene enrichment. GO can be divided into three parts: molecular function, biological process, and cellular component. GO enrichment is considered significant when the pvalue is less than 0.05. The enrichment results are shown in Table 6.1. Only part of the results are shown. Detailed results: Results / 6.DiffEnrich / Case_vs_Control / GO / Case_vs_Control.GOenrich.xls. From the GO enrichment analysis results, the 10 most significant terms in the three categories were selected to draw bar graphs and scatter plots for display. All term information in the category was drawn. The results are shown as follows: Figure 22 shown.

[0116] 4.6.2 Pathway Enrichment Results List: Pathway enrichment results are significantly enriched with a pvalue less than 0.05. KEGG enrichment is used as an example to display the pathway enrichment list. Only part of the results are displayed. Detailed results: Results / 6.DiffEnrich / Case_vs_Control / KEGG / Case_vs_Control.KEGGenrich.xls. In the enrichment results, the 20 most significant pathways are selected to draw bar charts and scatter plots for display. If there are less than 20 pathways, all pathway information is drawn. Taking KEGG enrichment as an example, the results are as follows: Figure 23 shown.

[0117] 4.6.3 Reactome enrichment analysis: The Reactome database is a database that collects articles written by experts and reviewed by peers about various reactions and biological pathways in the human body. It brings together various human reactions and biological pathways. This database provides people with a new tool to study biological pathways at a holistic level. At the same time, it is also an improved search and data mining tool that can simplify the search and research of data related to biological pathways, making the analysis of high-throughput data sets easier. Reactome enrichment is considered significant when the pvalue is less than 0.05. The detailed enrichment results are: Results / 6.DiffEnrich / Case_vs_Control / Reactomeenrich / Case_vs_Control.Reactomeenrich.xls. From the Reactome enrichment results, the 20 most significant pathways are selected to draw bar graphs and scatter plots for display. If there are less than 20, all pathway information is drawn. The results are as follows: Figure 24 shown.

[0118] 4.6.4 Protein interaction network analysis: Using the interaction relationships in the STRING protein interaction database (http: / / string-db.org / ), for species included in the database, we directly extracted the interaction relationships of the target gene set (such as the differentially expressed gene list) from the database to construct the interaction network. For species not included in the database, we first applied Blastx to align the target gene set sequence with the protein sequences of closely related or model species included in the STRING database, and then constructed the interaction network using the protein interaction relationships of the selected closely related or model species.

[0119] We provide protein interaction network data files for genes corresponding to differentially expressed mRNAs. These files can be directly imported into Cytoscape software for visualization and editing. Customers can analyze and plot certain network topological properties. For example, the size of a node in the interaction network is proportional to its degree. The more edges connected to a node, the higher its degree and the larger the node, indicating that these nodes are likely to be central to the network. The color of a node is related to its clustering coefficient, with a color gradient from green to red corresponding to low to high clustering coefficient values. The clustering coefficient indicates the connectivity between the node's neighbors; higher clustering coefficients indicate better connectivity, and so on. Based on different research objectives and needs, customers can also adjust node position and color, annotate expression levels, and perform other operations within the network diagram. It is important to note that the results obtained through blast comparisons are not guaranteed to be completely accurate; these results are provided as a reference to assist customers in identifying potentially important transcripts. This image shows the result after importing the file into Cytoscape software according to the provided instructions. The effect after importing Cytoscape software is as follows Figure 25 As shown (only part of the results are shown, detailed results: Results / 8.PPI).

[0120] 4.6.5 GSEA analysis: GSEA (Gene Set Enrichment Analysis), also known as gene set enrichment analysis, uses a predefined gene set (usually derived from functional annotations or previous experimental results) to sort genes according to their differential expression between two types of samples, and then test whether the predefined gene set is enriched at the top or bottom of the sorting table. Gene set enrichment analysis detects expression changes of gene sets rather than individual genes, so it can include these subtle expression changes and is expected to obtain more ideal results. The results are as follows: Figure 26 As shown (detailed results: Results / 9.GSEA), the GSEA results are shown in the following table (only part of the results are shown, detailed results: Results / 9.GSEA).

[0121] 4.6.6 Analysis of Co-expressed Genes between GQD-DEGs and Mouse Immune Function: Geneontology (https: / / geneontology.org) was used to identify potential genes related to mouse immune function. Venny 2.1 (https: / / bioinfogp.cnb.csic.es / tools / venny) was used to construct a list of co-expressed genes between GQD-DEGs and mouse immune function. Metascape (https: / / metascape.org / gp) was used to enrich the Gene Ontology and Kyoto Encyclopedia of Genes and Genomes (KEGG) signaling pathways of co-expressed genes. TBtools-II used the median frequency of gene expression and biological processes as a threshold to construct heat maps of the relationships between "cardiovascular development and growth" and "biological development and cell division processes" and co-expressed genes. Protein-protein interaction (PPI) networks and hub proteins were constructed using the String database (http: / / string-db.org). The PPI network data files were imported into Cytoscape 3.8.0 for visualization and editing. Cytohubbav0.1App performed topological analysis using edge percolated component (EPC), maximum clique centrality (MCC), degree, and Bottele Neck method, with intersection proteins being the key hub proteins.

[0122] like Figure 27 As shown, there were 155 co-expressed genes between mice immunized with 10 μg / ml GOQD-DEG.

[0123] like Figure 28 As shown in Figure 2, the significant biological processes of the 155 co-expressed genes include 302, which can be summarized into 20 primary signaling pathways such as innate immune response, inflammatory response, and regulation of immune effector process. Figure 28A in . These co-expressed genes bind to cytokine receptors (cytokine receptor binding), vascular endothelial growth factor receptor binding, or glycosaminoglycan binding, etc., and exert 20 molecular mechanisms such as nuclease activity and protein kinase activator activity, covering 72 molecular action areas, see Figure 24 The co-expressed genes were mainly involved in natural killer cell mediated cytotoxicity, apoptosis, endocrine resistance, heumatoid arthritis, etc. through signaling pathways such as cytokine-cytokinereceptor interaction, AGE-RAGE signaling pathway in diabetic complications, TNF signaling pathway, Rapl signaling pathway, NF-κB signaling pathway, HIF-1 signaling pathway, Fc gamma R-mediated phagocytosis, C-type lectin receptor signaling pathway, and NOD-like receptor signaling pathway. arthritis), Chagas disease, whooping cough, etc. 16 KEGG signaling pathway reactions (70 specific signaling pathways), see Figure 28 C in.

[0124] like Figure 29As shown in the figure, the relationship between co-expressed genes and biological development and cardiovascular biological processes involved in GQD-DEG is to clarify the effect of graphene quantum dots on organisms. It was found that 1079 GQD-DEGs were involved in 35 cardiovascular-related biological processes, among which cell adhesion, cellular calcium ion homeostasis, positive regulation of angiogenesis, extracellular matrix organization, and positive regulation of endothelial cell proliferation ranked in the top 20. Figure 29 A. Thbs1, Emilin1, Pgf, Vegfd, Ccl2, Cxcl10, Ecm1, Serpine1, Ccl5, Dab2ip, Ednra, Grem1, Igf1, Ang, Ang2, Angpt1, Cacna1c, Cd34, Csf2, Cxcl9, Lif, and Prkcb co-expressed genes are expressed in the top 20 cardiovascular biological processes, see Figure 29 B in. Involving 36 biological development and cell division processes, including chondrocyte proliferation, cell-cell adhesion via plasma-membrane adhesion molecules, protein localization to the cell surface, spongiotrophoblast differentiation, regulation of cell growth, multicellular organism development, meiotic chromosome segregation, and negative regulation of the mitotic cell cycle, see Figure 29 C. Ceacam1, Ceacam2, Fancd2, Lif, Angpt1, Ascl2, Casp6, Cd177, Gsn, Jag2, Vegfd are expressed in the top 20 biological development and cell division processes, see Figure 29D in.

[0125] like Figure 30 As shown in the figure, the protein-protein interaction network (PPI) and the proteins encoded by the 155 co-expressed genes of hub proteins constitute a PPI network with 152 nodes and 461 links (p < 1.0e-16). Figure 30 A. BettleNeck, Degree, EPC and MCC values ​​are the standard top 10 hub proteins, see Figure 30 B in the figure. The key hub proteins include 7, namely Serpine1, Nlrp3, Cxcl10, Ccl5, Ptprc, Thbs1, and Igf1. Figure 30 C in.

[0126] 4.6.7 Structural variation analysis

[0127] (1) Differential alternative splicing analysis

[0128] Differential splicing, also known as alternative splicing, refers to the process by which different splicing patterns occur during the conversion of pre-mRNA to mature mRNA, allowing the same gene to produce multiple distinct mature mRNAs and ultimately different proteins. Differential splicing is a key mechanism for regulating gene expression and generating protein diversity. Differential splicing analysis involves quantification of alternative splicing events and analysis of expression significance. Differential splicing analysis can be performed using rMATS (http: / / rnaseq-mats.sourceforge.net / index.html; Shen S et al., 2014) or JuncBASE (http: / / compbio.berkeley.edu / proj / juncbase / Home.html; Brooks AN et al., 2011), with rMATS being the default. rMATS analyzes five types of alternative splicing events: SE (exon skipping), RI (intron retention), MXE (mutually exclusive exons), A5SS (differential 5' splicing site), and A3SS (differential 3' splicing site).

[0129] The principle of rMATS detecting variable splicing is roughly as follows: Each variable splicing event corresponds to two spliceosomes (Isoforms), namely Inclusion Isoform and Skipping Isoform (taking SE event as an example). The expression levels of the two isoforms are statistically analyzed and divided by their effective lengths to obtain the corrected expression level. Then, the ratio of the total expression level of Exon Inclusion Isoform in the two isoforms is calculated. Finally, the significance analysis of the difference between the two groups is performed, which is the principle of differential variable splicing analysis (only SE event results are shown). The results are as follows Figure 31 shown.

[0130] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. Application of reagents for detecting Serpine1, Nlrp3, Cxcl10, Ccl5, Ptprc, Thbs1, and Igf1 proteins in the preparation of products for monitoring the effects of graphene quantum dots on organisms.

2. Use of reagents for detecting Thbs1, Emilin1, Pgf, Vegfd, Ccl2, Cxcl10, Ecm1, Serpine1, Ccl5, Dab2ip, Ednra, Grem1, Igf1, Ang, Ang2, Angpt1, Cacna1c, Cd34, Csf2, Cxcl9, Lif, Prkcb, Ceacam1, Ceacam2, Fancd2, Lif, Angpt1, Ascl2, Casp6, Cd177, Gsn, Jag2, and Vegfd genes in the preparation of products for monitoring the effects of graphene quantum dots on organisms.

3. A method for predicting the effect of graphene quantum dots on biological organisms based on transcriptome sequencing, characterized in that: The following steps are involved: 1) Construct the original gene library data of mouse bone marrow mesenchymal stem cells stimulated with 10 μg / ml graphene quantum dots and mouse bone marrow mesenchymal stem cells without graphene quantum dots stimulation, filter the original gene library data, evaluate the data quality, and clean the data to obtain valid data; 2) Compare the valid data to the genome, perform gene expression analysis, screen the expressed genes for differentially expressed genes, and perform functional annotation based on the differentially expressed genes.

4. The method according to claim 3, characterized in that Step 1) The method for constructing the original gene library data comprises the following steps: extracting RNA from mouse bone marrow mesenchymal stem cells stimulated with 10 μg / ml graphene quantum dots and mouse bone marrow mesenchymal stem cells not stimulated with graphene quantum dots → mRNA enrichment → double-stranded cDNA synthesis → end repair and addition of A and adapters → fragment selection and PCR amplification → library detection → Illumina sequencing; The mRNA enrichment tool includes Oligo(dT)Beads, the first strand synthesis tool in the double-stranded cDNA synthesis includes First Strand Synthesis, and the second strand synthesis tool in the double-stranded cDNA synthesis includes Second Strand Synthesis Mix; the library detection tool includes Qseq100 DNA Analyzer, and the Illumina sequencing tool includes NovaSeq TM X Plus, PE150.

5. The method according to claim 3, characterized in that The content of the data filtering in step 1) includes: ①Remove reads with adapters; ② Remove reads with N ratio greater than 0.002; ③ When the number of low-quality bases in a single-end read exceeds 50% of the read length, the paired reads need to be removed.

6. The method according to claim 3, characterized in that The data quality assessment in step 1) includes sequencing quality statistics, sequencing error rate, and GC content analysis; the sequencing error rate tool includes Illunima Casava 1.8; and the cleaned data includes fusion gene and variant site analysis.

7. The method according to claim 3, characterized in that Step 2) The genome comparison includes comparison statistics, genomic regional distribution, sequencing saturation analysis, transcript splicing and screening, and variable splicing analysis; the comparative statistics tools include Tophat2, Hisat2, STAR, and HISAT2, the transcript splicing and screening tools include Stringtie and gffcompare, and the variable splicing analysis tools include rMATS or JuncBASE.

8. The method according to claim 3, characterized in that Step 2) The gene expression analysis includes expression level analysis, correlation analysis, combined with other omics analysis, regulatory network analysis and genome visualization; the correlation analysis includes principal component analysis.

9. The method according to claim 3, characterized in that The tools for differential gene screening in step 2) include edgeR, DEseq2 and DEGseq.

10. The method according to claim 3, characterized in that Step 2) The functional annotation includes GO, KEGG, and Reactome enrichment analysis, and the tools for the GO, KEGG, and Reactome enrichment analysis include clusterProfiler; the GO, KEGG, and Reactome enrichment analysis include GO functional enrichment, pathway enrichment, Reactome enrichment, protein interaction network analysis, GSEA analysis, mouse immune co-expression gene analysis, and structural variation analysis; the protein interaction network analysis tool includes the STRING protein interaction database, and the mouse immune co-expression gene analysis tool includes Geneontology, Venny 2.1, Metascape, TBtools-II, String database, Cytoscape3.8.0, and Cytohubba v0.1.