A method and system for coronavirus transcriptome identification analysis
By using a coronavirus annotator to quantify viral gene expression and align sequences, the shortcomings of transcriptome analysis for different coronaviruses were addressed. Shared and unique sgRNAs were identified, revealing differences between the virus in vivo and in vitro, and providing a deeper understanding of coronavirus biology.
Patent Information
- Application Number
- CN202111152927.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-29
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-09-29
AI Technical Summary
Current technologies lack effective tools for analyzing the transcriptomes of different coronaviruses, especially neglecting the analysis of the viral transcriptomes of different human strains, resulting in insufficient understanding of viral gene function and an inability to determine changes in viral expression profiles in response to immune responses.
The coronavirus annotator was used to quantify viral gene expression. Sequence alignment was performed using bwa software, and SNP annotation was performed using bcftools and vcf-annotator. An expression matrix was constructed, sgRNAs were identified, a phylogenetic tree was built, and the expression levels of viral strains in different groups were analyzed to identify specific sgRNAs.
It can identify real and reliable sgRNAs, determine shared and unique sgRNA sequences between SARS-CoV and SARS-CoV-2, reveal significant differences between the virus in vivo and in vitro, provide new insights into coronavirus biology, increase the prediction of coronavirus protein quantities, and reveal novel sgRNAs related to disease function.
Smart Images

Figure BDA0003287694490000041 
Figure BDA0003287694490000051 
Figure BDA0003287694490000091
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of bioinformatics, and relates to a coronavirus transcriptome identification analysis method and system. BACKGROUND
[0002] The pathogen of COVID-19 is SARS-CoV-2, a beta coronavirus similar to MERS-CoV, the only other still circulating beta coronavirus. MERS-CoV is the pathogen of Middle East Respiratory Syndrome (MERS), which is more virulent but less transmissible than SARS-CoV-2, and slightly more distant in its phylogenetic distance (less than 90% amino acid sequence homology) than SARS-CoV-2. Both viruses have a single-stranded polyadenylated RNA genome of about 30,000 bases, encoding 4 structural proteins (spike protein (S), membrane protein (M), envelope protein (E) and nucleocapsid protein (N)), which are very similar in function in both types of viruses. However, the two viruses differ in the receptors for cell invasion, accessory proteins and specific functions of 16 non-structural proteins (nsp1 to nsp16). Nsp is produced by cleavage of the ORF1a and ORF1b encoded by two large polyproteins. ORF1 is closest to the 5' end and is translated directly from the genomic RNA after entry into the host cell, and can be divided into ORF1a and ORF1b according to the ribosome skipping mechanism [1] . MERS-CoV encodes at least 5 accessory proteins (ORF3, ORF4a, ORF4b, ORF5 and ORF8b), while SARS-CoV-2 encodes at least 6 accessory proteins (ORF3a, ORF6, ORF7a, ORF7b, ORF8 and ORF10) [2] . All proteins except ORF1a and ORF1b must be translated from sgRNA [3 -4] SgRNA pairs with the 70-base position at the head of the genome through a mechanism called discontinuous extension, via a variable-length short sequence (usually 6 to 12 nucleotides (nt)) called transcription regulatory sequence (TRS) present between each gene, and then extends the negative strand to the 5' end of the positive strand, generating a short negative strand sgRNA intermediate. The RNA intermediate is then copied to generate a positive strand sgRNA encoding viral proteins [5] .
[0003] The foundation of virology is the identification and annotation of viral genes and their functions. The acquisition of coronaviral transcriptome information is itself a difficult task due to the nature of the positive-sense RNA viral sequence and the presence of subgenomic RNAs (sgRNAs). The annotation of viral transcriptomes is fundamental to understanding viral biology, which is key to stopping viral transmission, replication, and pathogenicity. Previous coronavirus outbreaks, such as the 2003 outbreak of severe acute respiratory syndrome (SARS) and the 2012 onset and still ongoing MERS outbreak [6 -7] led to an increase in research on these zoonotic coronaviruses in order to understand the origins of these viruses. By comparing the transcriptomic changes of different coronaviruses, the mechanisms behind their unique pathogenicity and infectivity can be revealed and the molecular mechanisms behind interspecies transmission can be explained. Systematic annotation of the differences in coronaviral transcriptomes hidden in the metatranscriptomic data helps to further understand the transmissibility and virulence of the viruses. However, systematic comparison of the in vitro transcriptomes of these viruses is still lacking.
[0004] For the newly emerged SARS-CoV-2 virus, sequencing plays a crucial role in the diagnosis and monitoring of strain evolution [2,8] . However, currently, the sequencing datasets of SARS-CoV-2 and MERS-CoV are limited to viral and host transcripts produced in in vitro cell line infection or infection model organisms. The analysis of viral transcriptomes from different strains of humans is overlooked due to the lack of suitable analysis tools.
[0005] Sequence homology plays a crucial role in the annotation of viral gene functions. However, sequence homology alone does not guarantee protein expression as rapidly mutating RNA viruses can contain changes in sequence, leading to the generation of new ORFs or the inability to transcribe existing ORFs. Therefore, direct analysis of viral RNA is an important step in understanding which viral genes can be expressed. In the case of SARS-CoV-2, recent studies have utilized Oxford Nanopore technology to analyze viral RNA produced in cultured cell lines, identifying the presence of canonical and non-canonical viral transcripts. These studies all used isolated viral strains to infect the Vero cell line, which is isolated from the African green monkey epithelial cell line and does not initiate an interferon (IFN) response upon infection. While these studies identified the basic characteristics of the viral transcriptome, these independent studies only describe the transcriptome of one strain and are unable to determine changes in the expression profile when the virus is responding to the most basic immune response (e.g. IFNxx) [9-11] . SUMMARY
[0006] In order to solve the problems existing in the prior art, the purpose of the present application is to provide a coronavirus transcriptome identification analysis method, which quantifies viral gene expression by using a coronavirus annotator and identifies reliable sgRNAs in a large number of publicly available metatranscriptome datasets. In addition to summarizing the changes in sgRNA profiles and their relative expression, the present application can also determine novel sgRNAs of several different coronaviruses; it can also propose core sgRNA sequences shared between SARS-CoV and SARS-CoV-2, as well as sgRNA sequences specific to MERS-CoV. In addition, among the related coronaviruses found in bats and pangolins, a new sgRNA subgroup of SARS-CoV-2 and MERS-CoV appears to be conserved evolutionarily. At the same time, the transcription of specific sgRNAs differs significantly in vivo, in vitro and between different coronaviruses.
[0007] The present application provides a coronavirus transcriptome identification analysis method, which specifically comprises the following steps:
[0008] Step one, using bwa software to sequence align the coronavirus second-generation sequencing raw fastq file obtained from the short reads archive SRA with the NCBI reference genome, to generate a BAM file;
[0009] Step two, calling and filtering single nucleotide polymorphism SNP of the BAM file obtained in step one by bcftools, and annotating the SNP by vcf-annotator; then, grouping the collected virus strains according to the presence or absence of a certain SNP, and the grouping in this step will be used for transcriptome analysis of the expression matrix constructed in step four sgRNA expression, the analysis content of which is detailed in step four. In addition, according to the base change indicated by the SNP, the specific base is replaced on the basis of the reference genome to generate a so-called synonymous genome sequence.
[0010] Step three, CIGAR string parsing and breakpoint identification operation are performed on the BAM file obtained in step one to identify sgRNA;
[0011] Step four, constructing an expression matrix of the sgRNA identification result in step three and performing transcriptome analysis, during the analysis, grouping according to the presence or absence of SNP, the source (from in vivo experiment or in vitro experiment), and the project source (according to the project identification code), and statistically comparing the sgRNA expression of different groups of virus strains.
[0012] In step one, the coronavirus includes all virus species under the family Coronaviridae.
[0013] In step one, the sequence alignment refers to comparing the nucleic acid sequence obtained by nucleic acid sequencing technology with the reference sequence to obtain the presence or absence and arrangement difference of the base of the nucleic acid sequence and the reference sequence, thereby screening the nucleic acid sequence with breakpoints.
[0014] The screening refers to searching for the read with the letter "H" or "S" in the CIGAR string in the alignment result BAM file, and recording the information thereof, where "H" refers to "hard clip", that is, the read has a completely mismatched phenomenon at one end corresponding to the reference sequence, and "S" refers to "soft clip", that is, the read has a phenomenon of inconsistency with the reference sequence at one end, but is not completely mismatched, and still has some bases consistent with the reference sequence.
[0015] In step two, the calling refers to using bcftools to read the BAM file generated in step one to describe the base difference and the position of the difference base between the aligned sequence and the reference sequence.
[0016] In step two, the filtering refers to retaining the information of the difference base with an occupancy of more than 1% at the same position.
[0017] In step two, the annotation refers to the process of determining whether the screened difference base changes the codon and amino acid at the position of the base, and recording the change from which base to which base, from which codon to which codon, and from which amino acid to which amino acid.
[0018] In step two, the replacement of the specific base refers to the process of replacing the base corresponding to the position of the difference base in the reference sequence with the difference base.
[0019] The synonymous genomic sequence generated in step two is used for multiple sequence alignment with other coronaviruses that have completed whole genome sequencing, thereby constructing an evolutionary tree of coronavirus strains, which eliminates the process of conventional genome assembly.
[0020] In step three, the string parsing refers to obtaining the CIGAR string of the alignment of the aligned sequence and the reference sequence in the BAM file, extracting information therefrom, determining the position of the aligned sequence that is completely consistent, not completely consistent or completely inconsistent with the reference sequence, and obtaining the extension length of the three alignment conditions on the aligned sequence.
[0021] In step three, the breakpoint identification refers to the process of extracting and recording the position of the junction of the region of the two alignment conditions on the aligned sequence and the reference sequence when any two of the three alignment conditions exist in the aligned sequence.
[0022] In step three, the sgRNA identification refers to a process of inferring which gene expression product the aligned sequence originally belongs to according to the base position of the sequence on both sides of the breakpoint of the aligned sequence on the reference sequence.
[0023] In step four, the expression matrix construction refers to a process of calculating the total number of each sgRNA identified in step three in the sample, calculating the proportion of each sgRNA in all sgRNAs in the sample, and listing the proportion of sgRNA in each sample with sample number as row, sgRNA type as column, or sgRNA as row, sample number as column.
[0024] The coronavirus transcriptome data set used in the present application is mainly from publicly available databases. After searching and sorting, 19 data sets containing original reads are screened from the SRA database of NCBI, which contains 588 SARS-CoV-2 and related coronavirus samples; at the same time, the data set of the recently published sample is also used
[10] .
[0025] The coronavirus transcriptome identification and analysis method described in the present application utilizes sequences generated by high-precision second-generation sequencing technology, which can identify transcriptional regulatory sequence (TRS) from a single read.
[0026] Specifically, in the method of the present application, the original read is first aligned with the respective reference viral genome sequence, including SARS-CoV-2 (GeneBank ID NC_045512.2), SARS (GeneBank ID NC_004718.3), MERS (GeneBank ID NC_019843.3) or other coronavirus reference sequence, as shown in Table 1.
[0027] Table 1 Biological project information table added in the present application
[0028]
[0029]
[0030] Through breakpoint analysis, specific sgRNA can be inferred, and reads that span the junction between the 5 leader sequence (TRS) and more distal genomic sequences can be identified ( Figure 2 A). The relative abundance of specific sgRNA is equivalent to the relative expression of genes. In order to determine how viral genotype and viral origin (e.g. in vivo and in vitro) affect viral gene expression, the present application also constructs a heatmap ( Figure 3 ).
[0031] The coronavirus annotator in this invention is designed to describe all possible breakpoints. However, to identify truly reliable sgRNAs, this invention removes rare breakpoints and breakpoints that are inconsistent between samples. A complete breakpoint contains two independent genomic locations ( Figure 2 A). Furthermore, this invention also analyzed non-sgRNA breakpoints, sequences whose 5' ends do not contain leader TRS (LTRS). Data showed that non-sgRNA breakpoints are very rare (typically less than 0.05% of total sgRNA breakpoints) and exhibit inter-sample inconsistencies. Since these breakpoints were only found in one study, this invention focuses on sgRNAs formed by conventional 5' leader transcriptional regulatory sequences (LTRS) and 3' BTRS.
[0032] Based on the methods described above, the present invention also proposes a coronavirus transcriptome identification and analysis system, the system comprising a memory and a processor; the memory stores a computer program, and when the computer program is executed by the processor, the coronavirus transcriptome identification and analysis method of the present invention is implemented.
[0033] The present invention also provides the application of the aforementioned coronavirus transcriptome identification and analysis method in the identification of novel coronavirus TRS sequences.
[0034] The present invention also provides sgRNA, the gene sequence of which is shown in SEQ ID NO.1 and SEQ ID NO.2. This sequence is the corresponding sgRNA sequence in the reference genome, and the identity of the corresponding sequence in other coronaviruses should be more than 70%.
[0035] The present invention also provides the use of the sgRNA in the preparation of products for identifying coronaviruses.
[0036] The coronaviruses include all viral species within the family Coronaviridae.
[0037] The present invention also provides a diagnostic reagent comprising sgRNA as described above.
[0038] The beneficial effects of the present application include: the coronavirus transcriptome identification analysis method provided by the present application provides a new cognitive means for coronavirus biology, and provides valuable resources for the development of future treatment methods. The results of the identification analysis by the method of the present application show that SARS-CoV, SARS-CoV-2 and MERS-CoV all have the core sgRNA expression profile generated by the classic pathway. In addition, novel sgRNAs encoding evolutionarily conserved structural polypeptides were found in several of the coronaviruses, and these sgRNAs were expressed in both in vitro and in vivo samples, which greatly increased the number of predicted coronavirus polypeptides. In addition, two newly discovered polypeptides, which can be predicted by sequence to have the ability of IFN responsiveness and blocking IL17E (IL25) signal, respectively, suggest that they may have direct functional relevance to the disease. In addition, the in vivo expression level of the sgRNAs of the S protein is significantly higher than the in vitro level, while the level of the nucleocapsid protein is the opposite, which may be related to the infectivity and transmissibility of the coronavirus. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 Flow chart of the coronavirus transcriptome identification analysis method of the present application.
[0040] Figure 2 For research overview and sgRNA map. Figure 2 A is an overview of the research process. Figure 2 B is a classic breakpoint map of SARS-CoV-2. Figure 2 C is the annotation and comparison of the open reading frames of the three coronaviruses.
[0041] Figure 3 SgRNA expression heat map of SARS-CoV-2 and SNP annotation of the corresponding strain.
[0042] Figure 4 SARS-CoV-2 and MERS newly discovered sgRNA and its corresponding expression product. Figure 4 A is the breakpoint map of the newly discovered sgRNA of SARS-CoV-2. Figure 4 B, C, D are the results of the conservation analysis of the protein products corresponding to the newly discovered sgRNA in different coronaviruses.
[0043] Figure 5 Newly discovered sgRNA breakpoint and corresponding transcriptional regulatory sequence. Figure 5 A is the two transcriptional regulatory sequences of pORF2b of SARS-CoV-2. Figure 5 B is the newly discovered M protein transcriptional regulatory sequence in SARS-CoV-2. Figure 5 C is the transcriptional regulatory sequence of the newly discovered tORF7b in SARS-CoV-2. Figure 5D is the newly discovered transcriptional regulatory sequence of pORF8c in MERS-CoV.
[0044] Figure 6 A comparison of the expression of each sgRNA in vivo and in vitro. Figure 6 A comparison of the expression of sgRNA of SARS-CoV-2 in vivo and in vitro, Figure 6 B comparison of the expression of sgRNA of MERS-CoV in vivo and in vitro, Figure 6 C expression of S and N genes of coronaviruses and evolutionary relationship of these coronaviruses.
[0045] Figure 7 A plot of reads coverage of in vivo and in vitro datasets, wherein Figure 7 A in vitro data, Figure 7 B in vivo data.
[0046] Figure 8 is a plot of breakpoints of four coronaviruses other than SARS-CoV-2, wherein Figure 8A SARS-CoV, Figure 8B MERS-CoV, Figure 8C is a coronavirus from a pangolin, and Figure 8D is HKU1.
[0047] Figure 9 A schematic of the homology of pORF2b of the coronavirus from a pangolin to the IL-17 receptor.
[0048] Figure 10 A detailed gene expression plot of SARS-CoV-2 and MERS-CoV.
[0049] Figure 11 A plot of viral load under Gleevec and IFN-β treatment. DETAILED DESCRIPTION
[0050] The present application is further described in connection with the following detailed examples and accompanying drawings. The procedures, conditions, experimental methods, etc. used in the present application are those conventional and well known in the art, unless specifically mentioned otherwise, and are not particularly limited by the present application.
[0051] All sequencing data in the examples of the present application were derived from the NCBI Short Reads Archive (SRA). Some nanopore datasets were downloaded from the online storage locations described in the respective studies
[10] BioProjects were identified by a keyword search for "coronavirus" and manual curation, only macro-transcriptomic data were kept, raw sequencing files were downloaded from SRA using a custom script containing wget, fastq files in compressed format were generated from downloaded SRA files using SRAtoolkit. After initial sequence alignment with bwa to the reference genome sequence of SARS-CoV, SARS-CoV-2 or MERS-CoV, samples with too few viral reads were filtered out. CORONATATOR only uses reads generated by second generation sequencing technology (Illumina), Nanopore data were used for comparison.
[0052] The coronavirus annotator provided by the present application includes a series of perl and bash scripts for analyzing coronavirus RNA-Seq data, and the coronavirus annotator includes three main steps, including preprocessing, breakpoint identification, sgRNA reading and analysis, and the details are as follows:
[0053] 1. Preprocessing
[0054] Aligning with the genomic sequence of SARS-CoV, SARS-CoV-2 and MERS-CoV as the reference sequence to generate the preprocessed BAM file, and the corresponding genomic sequence of bat virus and pangolin virus is taken as the reference sequence on NCBI. Call and filter SNPs by bcftools
[12] Annotate SNPs using vcf-annotator
[13] In addition, the synonymous genomic sequence is generated by screening the SNPs for further analysis.
[0055] 2. Breakpoint identification
[0056] Breakpoints are identified from soft-clip and hard-clip alignments, which are partial alignments mainly caused by reads of recombination junction positions, which are generated by the mechanism of sgRNA generated by coronavirus. In this step, a matrix of reads information, breakpoint positions, CIGAR strings and possible TRS sequences is generated.
[0057] 3. sgRNA reading and analysis
[0058] The classic sgRNA is identified and defined by two breakpoint coordinates of the reference genome sequence, which are obtained by extracting from partial alignment (one is primary alignment, the other is secondary alignment, so the result will appear when the pair of reads being aligned is aligned to two regions of the reference sequence at the same time, at this time the alignment with higher alignment degree is called primary alignment, and the secondary alignment is the secondary alignment). In order to identify the possible TRS pattern, the sequence between the breakpoint pair is extracted from the synonymous genome sequence generated before. Then, by manually comparing the distance between the known viral gene start codon and its breakpoint, the corresponding sgRNA gene is identified. The biological samples with sgRNA count more than 20 are retained for further analysis, in these samples, the sgRNA is counted according to the corresponding gene and the transcriptome matrix is obtained by normalizing the total sgRNA count.
[0059] In the implementation of the present application, the identification of new open reading frames is also carried out: using Prodigal
[14] to predict potential ORFs and output all potential genes using the -s parameter. A python script is also used to identify very short ORFs. Then, for the sgRNA supported by multiple biological projects, the distance between their breakpoints and all identified start codon sites is calculated and sorted. The ORF with the closest breakpoint upstream of the start distance is marked and manually checked for verification. By identifying new open reading frames, the reliability of the newly discovered sgRNA can be provided, and when there is a new open reading frame downstream of the breakpoint of the newly discovered sgRNA, especially when the distance is very close, the newly discovered sgRNA will be very likely to be expressed and have biological function.
[0060] In the implementation of the present application, sequence alignment and phylogenetic analysis are also carried out: the synonymous genome sequences of SARS-CoV-2, SARS-CoV, MERS-CoV, and the sequences of sgRNA greater than 20 in bat, pangolin or other human coronavirus biological samples are used for phylogenetic analysis. Multiple sequence alignment uses MAFFT
[15] , and maximum likelihood consensus tree is constructed using IQ-TREE
[16] , with 1000 bootstrap.
[0061] In the implementation of the present application, the nanopore sgRNA ratio is also converted into short read form: the dataset of Kim et al.
[10] includes nanopore data and read data. The ratio between the two is used to convert other nanopore datasets into a ratio that can be compared with other datasets in this study.
[0062] The heat map involved in the present application is Figure 3) for displaying gene expression profiles. The sgRNA expression dot plots and box plots were made with the heatmap. plus package using ggplot2 package to compare the gene expression differences between different sample sources, and t-test and wilcoxon test were used for statistical analysis.
[0063] In the functional annotation of the present application, the new peptide sequences were aligned against the UniProtKB / Swiss-Prot database using the EMBL online tool FASTA (https: / / www.ebi.ac.uk / Tools / sss / fasta / ) with default parameters. Protein domains were identified using the NCBI CD Blast online service.
[0064] The present application also verified the conservation of viral polypeptide sequences: To detect the conservation of the predicted polypeptide sequences in related virus species, the present application established a reference database containing all the predicted open codes of the genomes of related viruses. DC MegaBlast (DisContinuous MegaBlast) was used to search for inter-species homologous sequences. The parameters were set as follows: window_size 0, gapopen 0, gapextend 2, penalty-1, reward 1, num_alignments 1. Then MAFFT was used for multiple sequence alignment (MSA) of a set of homologous ORFs. Then CLUSTAO (Clustal Omega) was used to calculate the matrix of the MSA results. Nucleotide and amino acid sequences were also detected by the same method.
[0065] Example 1 demonstrates that the abundance of sgRNAs provides different levels of coronavirus ORFs
[0066] Many ORFs of SARS-CoV-2 are annotated based on sequence identity, and some of the annotation results are controversial in proteomics and sequencing studies [2,9] The present application divides the common ORFs into three categories (core, low support, and no support) Figure 2 B) Identifying true sgRNAs requires multiple studies and multiple sample analysis, because unique artificial sequences are often generated during library preparation or sequencing [17-18] In addition, many low-abundance non-canonical sgRNAs may be randomly generated abnormal transcripts without specific functions
[10] Therefore, only sgRNAs that exist in multiple studies and datasets can be considered as true sgRNA candidates. To classify each viral gene, the present embodiments consider several factors such as sgRNA relative abundance, TRS conservation, and possible influence of start-codon hijacking on leaky ribosome scanning
[19] .
[0067] The present embodiments first validated the most frequently annotated ORFs in SARS-CoV-2, SARS-CoV, and MERS-CoV by looking at the degree of support from specific sgRNA sequencing evidence. For SARS-CoV-2, after removing samples with read counts less than 20 in sgRNAs, 34 samples remained. To identify core sgRNAs in these samples and group them into the first sgRNA category of the present embodiments, all possible sgRNAs were identified according to viral species using a weighted average method and their relative abundances were recorded. At a relative abundance of 0.5%, 8 typical breakpoints were found to correspond to 8 sgRNAs, which contain 8 well-studied ORFs of SARS-CoV-2: S, E, M, N, ORF3a, ORF6, ORF7a, and ORF8 Figure 2 B-C, Table 2, Table 3, Figure 8A ).
[0068] Table 2. Annotated information of three coronaviruses
[0069] Table 2-1 SARS-CoV-2
[0070]
[0071] # a: This presents one common annotation of ORF3b in SARS-CoV-2, another annotation is shown in the bracket.
[0072] # b: ORF7b has a significantly lower relative abundance of expression in SARS-CoV-2 compared to other canonical sgRNAs and has no homologous protein in SARS-CoV.
[0073] Table 2-2 SARS-CoV
[0074]
[0075]
[0076] Table 2-3 MERS-CoV
[0077]
[0078] Table 3 sgRNA expression matrix of samples
[0079]
[0080]
[0081] The sgRNAs of these ORFs are located between 9-162 nt upstream of the start codon. N is the most abundant core sgRNA, accounting for 54% of all core sgRNAs identified in all samples. E sgRNA is the least abundant, at 1.5%, and is the only core protein not identified in a recent proteomic study [9,20] In addition, the detection rates of ORF7a, M, ORF3a, S, ORF8 and ORF6 are 10.6%, 8.4%, 6.9%, 6.1%, 5.9 and 2.7%, respectively. These 8 core sgRNAs account for 70-100% of the total sgRNAs, depending on the sample type (in vivo and in vitro), strain and read coverage Figure 3
[0082] In addition to their high relative abundance, these 8 core sgRNAs share a common TRS with the core sequence ACGAAC, which is unique to this group of sgRNAs. This core sequence is likely a necessary and sufficient condition for sgRNA formation. The same 8 core sgRNAs and the core TRS sequence are also shared by SARS-CoV Figure 8A In addition to N (with the core sequence of ACGAA), the 7 core sgRNAs of MERS-CoV (S, E, M, N, ORF3, ORF4a and ORF5) Figure 2 C) also use this core sequence.
[0083] The sgRNAs of the second class of ORFs are generally less abundant and do not use this core sequence. This group of sgRNAs includes ORF7b in SARS-CoV-2 and SARS-CoV, ORF3b in SARS, ORF4b and ORF8b in MERS-CoV. For SARS-CoV-2, the average relative abundance of E is 1.5%, the lowest among the core sgRNAs, while ORF7b is only 0.02%. The low abundance or inefficiency of this sgRNA formation is likely caused by the use of non-canonical TRS. This group of sgRNAs does not use the conserved core TRS sequence as the core sgRNAs, i.e. the sequence they rely on for recombination is always offset by a few bases from the core sequence, and often there is a mismatch between the leader TRS and the 3' TRS.
[0084] The remaining ORFs belong to a third class, which lack supporting sgRNAs. This class can be further divided into two subclasses. The first class is ORFs that are likely to be translated despite the lack of supporting sgRNAs. It has been shown that some coronaviral ORFs can be expressed through a mechanism known as ribosomal skipping
[19] . ORF9b of SARS-CoV-2 belongs to this class. In fact, several recent proteomic studies have shown that the ORF9b protein product is present in SARS-CoV-2
[20] . Its homologous gene in SARS-CoV, also named ORF9b, also belongs to this class. The ORF7b of SARS-CoV-2 and SARS-CoV have also been mentioned in previous studies to belong to this class of ORFs
[19] , and the long stretch of sequence between the start codon of ORF7b and the preceding ORF7a (362 bp in SARS-CoV-2; 365 bp in SARS-CoV) is devoid of additional start codons. However, these ORFs can still form their own sgRNAs in the present application, albeit in low amounts.
[0085] The second class is the most dubious class of ORFs, which lack supporting sgRNAs and are unlikely to be expressed due to the presence of start codons between them and the closest sgRNA breakpoint. This class includes several ORFs that are often annotated, including ORF3b, ORF9c, and ORF10 of SARS-CoV-2, and ORF8b of SARS-CoV. These ORFs are separated by multiple start codons from the preceding ORFs, and there is no corresponding sgRNA and proteomic evidence for them [10,20-21] . In fact, the existence of ORF10 has been questioned in a recent study [9,11] . The above evidence suggests that experiments of viral products in putative open reading frames without sgRNA or proteomic support can be flawed. For example, a recent study generated a synthetic product of a predicted truncated version of ORF3b in SARS-CoV-2, which was postulated to have a stronger anti-IFN activity than the SARS version of the putative truncated version
[22] .
[0086] Example 2 Identification of novel sgRNAs containing non-canonical TRSs in SARS-CoV-2, MERS-CoV, and SARS-CoV
[0087] During the formation of the core sgRNA profile, the 3’ TRS containing the minimal core sequence will pair with the leader TRS. For each particular core sgRNA, the two TRS core sequences used must have the same length and sequence, although the length of the TRS region can vary for different sgRNAs Figure 2 C). The average length of these canonical SARS-CoV-2 TRSs was found to be ~9.6 bases in this example. SARS-CoV uses the same core sequence as SARS-CoV-2, while MERS uses another class of 6-nucleotide TRSs with a different core sequence Figure 8A , Figure 8B ).
[0088] When looking for sgRNA transcripts with an abundance over 0.2%, three novel sgRNAs were found to exist in at least two independent samples and studies Figure 4 A). All three novel sgRNAs do not use the canonical TRS sequence used by the core sgRNAs. However, the breakpoints of these sgRNAs support the discontinuous extension model, as sequences from the 3’ are found in the TRS sequence of the final transcript Figure 5 ). Additionally, the existence of the negative strand versions of these sgRNAs was determined by strand-specific RNA library sequence analysis. These novel sgRNAs were not described in previous Nanopore sequencing studies [9-11] As previously described, analysis of these non-canonical breakpoint sequences found that when the current leader TRS is not complementary to the 3’ TRS core sequence, it is possible for the sequences surrounding the core sequence to be complementary, resulting in the production of a sgRNA (Figure 2b, tORF7b and 3pORF2b). Analysis in this example also confirmed that TRS sequences vary greatly between distantly related viruses, and that the length of canonical TRS sequences can exceed 30 bases in some coronaviruses (Figure 8D).
[0089] The three novel sgRNAs produced by the three novel TRSs are called pORF2b, aM and tORF7b. The longest novel sgRNA, pORF2b, is located within the S gene, with two alternative TRSs near 22501, encoding a short peptide with a domain conserved in distantly related coronaviruses Figure 4 A, Figure 4 B). The second sgRNA breakpoint is 31 bases downstream of the canonical breakpoint for M, at base 26494. This sgRNA will also express the M protein, but with a different 5’ UTR than the canonical sgRNA for M Figure 5B). The shortest of the three novel sgRNAs has a breakpoint at 27761, encoding a truncated ORF7b (tORF7b). This truncation removes 14 of the 24 amino acids that make up the peptide chain extracellular domain and transmembrane domain Figure 4 A, Figure 4 C). This sgRNA is expressed at high levels in vivo and in vitro and can have novel functions.
[0090] The pORF2b translation product is a 36 amino acid peptide chain. PSIPRED predicts
[23] It has an intracellular protein binding domain and two short alpha helices located in the transmembrane domain, one of which is partially extracellular Figure 4 A). pORF2b was found in four samples from two independent studies. In a patient sample from Washington State, USA (SRX7884411), pORF2b was expressed at the highest level, accounting for 11.1% of total sgRNAs. In another patient sample from the same project (SRX7884409), it accounted for 1.2% of the identified sgRNAs Figure 3 ). SNP analysis showed that the strains infecting these two patients only differ by one nucleotide. Patient samples from five other different strains from the same study Figure 3 ) did not produce sgRNAs for pORF2b. This suggests that the level of pORF2b transcript can have some correlation between strains, while also indicating that pORF2b is not a false positive result from sequencing randomness. SgRNA pORF2b was also identified in another independent study (PRJNA615032) that used different strains from the in vivo study Figure 3
[0091] The sequence conservation of pORF2b was determined in other Sarbecoviruses, including SARS-CoV, HKU3 (bat coronavirus), RaTG13 (bat coronavirus, believed to be directly related to SARS-CoV-2), and a coronavirus that infects pangolins (SRX7732088)
[24] All four viruses have a corresponding ORF, with RaTG13 having the highest homology at 91.85% nucleotide identity (Table 4, Figure 4 B). pORF2b, as well as the pangolin version with the C-terminal extension, has a high similarity to the ligand binding domain of human IL17RB Figure 9
[0092] Table 4 sgRNA conservation information table
[0093] Table 4-1 SARS-CoV-2
[0094]
[0095]
[0096] Table 4-2 MERS-CoV
[0097]
[0098] A third novel breakpoint is at position 27761 within ORF7b, encoding a truncated version of ORF7b (tORF7b). This transcript and its relative abundance in vivo and in vitro was identified in two independent studies including more than one viral strain. This transcript was also recently identified in a VERO cell line infected with a single strain
[10] . This novel sgRNA is expressed at relatively high levels in vivo and in vitro Figure 3 , and the homologous gene product of this sgRNA is also present in several SARS-CoV samples from two studies. This truncated version of ORF7b is missing the intracellular domain and more than half of the transmembrane domain, but retains the hydrophilic extracellular domain Figure 4 A). ORF7b, the homolog of SARS-CoV-2 ORF7b, is also present in SARS-CoV virions
[19] . The portion of ORF7b encoded by tORF7b is highly conserved in SARS (Table 4, Figure 4 C). However, SARS ORF7b in one study has a 45nt deletion in this portion, which is also present in the tORF7b deletion sequence, and results in attenuated viral response to beta interferon induction, providing a replication advantage
[25] . Future studies will reveal whether this novel sgRNA encodes a toxic peptide segment with antagonistic IFN function, while disrupting the initiation of the interferon response.
[0099] The present invention also obtained a large number of MERS-CoV in vivo and in vitro sequence datasets, enabling the identification of a large number of atypical sgRNAs (Figure 8). The sgRNA (presumably ORF8c or pORF8c) encodes a peptide segment that can be translated into 51 amino acids. This sgRNA was identified in 5 independent in vivo and in vitro studies, with abundance ranging from 0.03% to 1.0%. PSIPRED suggests that this new peptide has a transmembrane domain linked to a cytoplasmic helix structure. The conservation of this protein was confirmed in other Merbecovirus including HKU4, HKU5 and an Erinacus coronavirus. pORF8c was present in all 3 samples, with varying degrees of conservation. In Merbecovrius, the cytoplasmic N-terminus of this protein is the most conserved, while a C-terminal extended version was observed in HKU5 and Erinaceus Figure 4 D).
[0100] To exhaustively search for novel sgRNAs, the screening threshold was lowered to a relative abundance of 0.01%, while keeping the other criteria. This analysis found more novel sgRNAs of SARS-CoV-2, SARS-CoV and MERS-CoV in multiple studies (Table 5). Further sequencing and future experiments will determine the importance of pORF2b, tORF7B and aM and other numerous low-abundance novel sgRNAs.
[0101] Table 5 Information table of novel sgRNAs discovered by the present invention
[0102] Table 5-1 Novel sgRNAs of SARS-CoV-2
[0103]
[0104] #Kim_2020 is a single-sample dataset, Kim et al. 2020.
[0105] #* indicates no overlapping sequence
[0106] Table 5-2 Novel sgRNAs of SARS-CoV
[0107]
[0108]
[0109]
[0110] #a: This sgRNA has a weighted ratio below the threshold of 10^-4, but it can be translated into the tORF7b protein product. This protein peptide chain is conserved in SARS-CoV-2 and its existence is supported by multiple studies, so it is included here.
[0111] Table 5-3 Novel sgRNAs of MERS-CoV
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119] Example 3: Coronavirus detection experiment-induced changes in the relative abundance of novel pORF8c
[0120] This embodiment will verify the experimental utility of the coronavirus transcriptome identification and analysis method of the present invention, and verify the novel sgRNA's response to experimental stimulation in a similar manner to other studied genes. For this purpose, an experimental dataset was used. This embodiment tested the effects of Gleevec and IFN-β on host gene expression during in vitro treatment of MERS-CoV infection (PRJNA233943 & PRJNA233944). Figure 11 This embodiment analyzes the effects of these drugs on viral load, viral gene expression, and the expression of the newly discovered pORF8c. Initial analysis showed that a decrease in viral load led to an increase in the expression of individual viral genes. Figure 11 Importantly, even at low viral loads, the N / S ratio remained high, indicating that this ratio is not affected by viral abundance but rather by the internal and external environment. Furthermore, the effects of Gleevec and IFN-β on viral gene expression were inconsistent, varying across different viral genes. In the responses to IFN-β and Gleevec, pORF8c expression, along with core sgRNAs, particularly N and E, showed the same trend with increasing viral load. This suggests that, in this context, pORF8c exhibits similar biological responses to some core sgRNAs in terms of gene expression.
[0121] Example 4: In vivo, the relative abundance of SARS-CoV-2 S sgRNA increased.
[0122] In processing the datasets, the present invention noted that there are two different read coverage patterns on the reference genome for SARS-CoV-2, which suggests that there are two sources of reads for the virus. Upon further examination, it was found that the two patterns are due to in vivo and in vitro conditions Figure 7 ). The former consists of extracellular viral particles and infected host cells present in BALF (human), nasal wash (ferret) or lung homogenate (MERS), while the latter consists of infected cell lines, which are not affected by the antiviral responses of the organism or the innate (e.g. VERO cells do not produce IFN) that can be present. The viral sequences from samples (mainly BALF) of SARS-CoV-2 generally cover the entire viral reference genome, with little difference in the 3’ end relative to other regions. In contrast, in in vitro samples, the 3’ end of the viral genome is highly covered due to the nesting of sgRNAs formed during viral transcription.
[0123] SARS-CoV-2 and MERS-CoV are the only two virulent coronaviruses still in circulation, and both have in vivo and in vitro macrotranscriptomic datasets. The relative abundance of sgRNAs generated in vivo and in vitro for SARS-CoV-2 and MERS-CoV was analyzed in the present invention. In comparing the differences in viral sgRNAs generated in vivo versus in vitro, it was found that the S to N ratio was significantly higher in vivo than in vitro, especially for SARS-CoV-2 (0.04 in vitro vs 0.69 in vivo, p = 0.0012, Wilcoxon rank sum test) Figure 6 A-B, Figure 10 ). One general explanation for this significant difference in sgRNA levels is that the difference in environmental pressure affects the demand for these sgRNAs by viral replication. Selection pressure can promote transcription of genes that are beneficial for viral proliferation. For example, the main function of the S protein revolves around the recognition and invasion of host cells, while the function of the N protein is centered on regulating viral RNA to promote viral replication, a process mediated by binding to the 3’ end of the viral genome, the viral packaging signal and TRS [26-28] . Another possible explanation that can exist simultaneously is that the increase in the S / N ratio in vivo is caused by internal changes in the viral replication transcription complex (RTC), which makes the transcription process more favorable to TRS readthrough, preferentially generating longer sgRNAs. This global adjustment of viral transcription can involve changes in host factors, as for the infectious bronchitis virus (IBV), a gamma coronavirus, the N protein recruits the helicase DDX1 to promote TRS readthrough and thus the formation of long sgRNAs after being phosphorylated by intracellular GSK3
[29] . Under this premise, since N protein is produced from the shortest sgRNA and S protein is produced from the longest sgRNA, TRS read-through will increase the expression of S. Future electron microscopy studies of viral particles in vivo and in vitro will determine whether the abundance of S sgRNA in SARS-CoV-2 is correlated with the level of S protein on the surface of viral particles. Other examples of significant differences in sgRNA expression in vivo and in vitro include the overall increase in accessory sgRNAs levels when SARS-CoV-2 and MERS-CoV suppress immune responses through multiple pathways Figure 6 A, B)
[30] .
[0124] To better understand the relative abundance differences of SARS-CoV-2 sgRNAs in vivo and in vitro compared to other coronaviruses, and to determine whether other new sgRNAs are missed, this example analyzed other coronaviruses using CARONTATOR. The analysis included OC43, NL63, HKU1, and bat and pangolin viruses that are highly homologous to SARS-COV-2 [8,24] ( Figure 6 C, Table 1 and Figure 8). Some datasets did not have enough breakpoints to provide information. For example, the analysis of RaTG13, the bat virus with the highest homology to SARS-COV-2, only resulted in 1 breakpoint and was therefore omitted in Figure 6 C.
[0125] Among the different coronaviruses, SARS-COV-2 has the highest content of S sgRNA, especially in vivo Figure 6 C). The analysis suggests that this phenomenon is not due to strain-specificity, as high levels of S sgRNA expression were observed in in vivo samples of different strains Figure 3 . High levels of S protein expression can play a role in the virus to facilitate crossing the species barrier and thus exhibit high infectivity. Also, it is noted that the relative levels of S sgRNA positively correlate with the infectivity of the coronavirus. The viral infectivity and levels of S sgRNAs in vivo are as follows: SARS-CoV-2 > HKU1 > MERS
[31] However, S protein levels alone are not sufficient to result in high levels of transmission capacity, as factors such as the stability of the S protein, the affinity of the receptor
[32] and the stability of the viral particle
[33] will also affect the transmission of the virus.
[0126] Example 5 RTC mutation reverses the expression of N and S sgRNAs in vivo and in vitro
[0127] The observation of mutations in the RTC component of the virus in this invention also changes the expression profile of S to N. In particular, Kim et al.
[10] The virus strain used has a unique non-synonymous mutation in RTC component nsp3 (encoding the papin protease, which binds to N and M proteins) Figure 3 ). The in vitro transcriptome of this virus strain shows a strong increase in the S to N ratio, similar to the virus expression profile under in vivo conditions Figure 3 ). One in vivo experiment virus strain (SRX7852918) also has two non-synonymous mutations in nsp3, in addition to mutations in nsp6 and nsp12, and this virus has a lower S / N ratio in its transcriptome than in the in vitro conditions.
[0128] The presence of nsp3 mutations in both strains with altered gene expression is also intriguing. It has been reported that nsp3 can bind to the viral genome 3’-terminal TRS, a region that is both a global RNA packaging signal for the virus and encodes the N and M proteins [27,34-35] . In addition, phosphorylation of the N protein changes its conformation, making it more prone to bind to viral RNA, as in the aforementioned IBV, promoting TRS read-through [2,36] promoting long-chain sgRNA production. This phenomenon initially suggests that mutations within nsp3 affect the relative abundance of each sgRNA by affecting the global architecture of the virus, possibly acting through a similar mechanism as in the aforementioned IBV. In addition, the changes in the relative abundance of N sgRNA in vivo and in vitro, under the influence of the aforementioned mechanism, can feedback their interaction with nsp3 and affect the function of nsp3 mutations in vivo and in vitro Figure 6 A-B).
[0129] To date, the large amount of SARS-CoV-2 sequence data has been mainly used for the typing and tracking of emerging strains. While this is important, it has led to underutilization of valuable material. An important step has been taken in the identification of virus strains, the description of coronavirus sgRNA expression and the discovery of new unknown functions of conserved sgRNAs generated by non-canonical TRS pairing mechanisms Figure 5 ) by developing the CORONATATOR bioinformatics analysis pipeline. These newly discovered protein function predictions are still ongoing. A preliminary finding is that the pangolin virus SARS-CoV-2 pORF2b homolog has extensive similarity to the human IL17RB s ligand binding domain Figure 9 ). Coronaviruses can produce a short peptide through this sgRNA, which can theoretically disrupt IL17B and IL17E (IL25) signaling, the latter two of which are usually associated with promoting or inhibiting inflammatory responses in specific situations. Future proteomics studies and / or ribosome sequencing studies are needed to verify the presence of the protein products encoded by the novel sgRNAs identified in this application.
[0130] The analysis of the present invention also shows that different SARS-CoV-2 viral strains express different levels of each sgRNA (Table S3, Figure 3 ), especially newly discovered sgRNAs. The results of the present invention highlight that only by deep sequencing of patient samples under selective pressure of the virus, the viral pathogenesis in terms of sgRNA expression can be truly understood. This requires a thorough understanding of the samples, a thorough sequencing and analysis of different viral strains at different stages of COVID-19, which will also drive the emergence of truly personalized treatments.
[0131] While other zoonotic viruses can have extensive sequence similarity to SARS-CoV-2 at the gene or genome level, sequence similarity alone is not enough to produce a human-susceptible, pathogenic virus. When discussing the origin of zoonotic viruses, one generally does not consider the specific expression levels of viral genes, such as the S protein, but this can be important for crossing the species barrier. For example, due to the large number of unsampled zoonotic viruses, it is likely that S proteins capable of crossing the species barrier have existed for a long time, but have not been expressed at a high enough level to enable sustainable human-to-human transmission. Low levels of S protein expression can allow the virus to sporadically spread from bats to humans, but in humans the process cannot be sustained and the transmission efficiency between humans will be low due to the better sanitary conditions of human living environments. In line with this, people living near bat caves can carry virus-specific antibodies, but have never experienced severe illness
[37] .
[0132] The analysis of the present invention on the metatranscriptomic datasets identified many sources of RNA, such as host RNA and microbial RNA (although not the best way of molecular capture). At the moment, it is not clear why some people die from SARS-CoV-2 infection while others do not, and these valuable sequences should not be wasted, and these datasets could become more useful if they could carry more clinical information. In theory, most GISAID entries on SARS-CoV-2 should have supporting metatranscriptomic datasets. However, at the moment, the number of entries in GISAID describing viral genome sequences and strains far exceeds the number of raw reads data found in SRA, and sharing raw reads information would greatly help researchers study this virus and ultimately contain it.
[0133] REFERENCES
[0134] 1. Knipe, D. M., Howley, P. M., 2013. Fields virology, sixth ed. Wolters Kluwer Health / Lippincott Williams & Wilkins, Philadelphia.
[0135] 2. Wu, F., Zhao, S., Yu, B., Chen, Y. M., Wang, W., Song, Z. G., Hu, Y., Tao, Z. W., Tian, J. H., Pei, Y. Y., Yuan, M. L., Zhang, Y. L., Dai, F. H., Liu, Y., Wang, Q. M., Zheng, J. J., Xu, L., Holmes, E. C., Zhang, Y. Z., 2020. A new coronavirus associated with human respiratory disease in China. Nature 579, 265-+.
[0136] 3. Brian, D. A., Baric, R. S., 2005. Coronavirus genome structure and replication. Curr Top Microbiol 287, 1-30.
[0137] 4. Yount, B., Curtis, K. M., Fritz, E. A., Hensley, L. E., Jahrling, P. B., Prentice, E., Denison, M. R., Geisbert, T. W., Baric, R. S., 2003. Reverse genetics with a full-length infectious cDNA of severe acute respiratory syndrome coronavirus. P Natl Acad Sci USA 100, 12995-13000.
[0138] 5. Sola, I., Almazan, F., Zuniga, S., Enjuanes, L., 2015. Continuous and Discontinuous RNA Synthesis in Coronaviruses. Annu Rev Virol 2, 265-288.
[0139] 6. Assiri, A., McGeer, A., Perl, T. M., Price, C. S., Al Rabeeah, A. A., Cummings, D. A. T., Alabdullatif, Z. N., Assad, M., Almulhim, A., Makhdoom, H., Madani, H., Alhakeem, R., Al-Tawfiq, J. A., Cotten, M., Watson, S. J., Kellam, P., Zumla, A. I., Memish, Z. A., Team, K. M.-C. I., 2013. Hospital Outbreak of Middle East Respiratory Syndrome Coronavirus. New Engl J Med 369, 407-416.
[0140] 7. Peiris, J. S. M., Yuen, K. Y., Osterhaus, A. D. M. E., Stohr, K., 2003. Current concepts: The severe acute respiratory syndrome. New Engl J Med 349, 2431-2441.
[0141] 8. Zhou, P., Yang, X. L., Wang, X. G., Hu, B., Zhang, L., Zhang, W., Si, H. R., Zhu, Y., Li, B., Huang, C. L., Chen, H. D., Chen, J., Luo, Y., Guo, H., Jiang, R. D., Liu, M. Q., Chen, Y., Shen, X. R., Wang, X., Zheng, X. S., Zhao, K., Chen, Q. J., Deng, F., Liu, L. L., Yan, B., Zhan, F. X., Wang, Y. Y., Xiao, G. F., Shi, Z. L., 2020. A pneumonia outbreak associated with a new coronavirus of probable bat origin. Nature 579, 270-273.
[0142] 9. Davidson, A. D., Williamson, M. K., Lewis, S., Shoemark, D., Carroll, M. W., Heesom, K., Zambon, M., Ellis, J., Lewis, P. A., Hiscox, J. A., Matthews, D. A., 2020. Characterisation of the transcriptome and proteome of SARS-CoV-2 using direct RNA sequencing and tandem mass spectrometry reveals evidence for a cell passage induced in-frame deletion in the spike glycoprotein that removes the furin-like cleavage site. bioRxiv.
[0143] 10. Kim, D., Lee, J.-Y., Yang, J.-S., Kim, J. W., Kim, V. N., Chang, H., 2020. The Architecture of SARS-CoV-2 Transcriptome. Cell 181, 914-921. e910.
[0144] 11. Taiaroa, G., Rawlinson, D., Featherstone, L., Pitt, M., Caly, L., Druce, J., Purcell, D., Harty, L., Tran, T., Roberts, J., Catton, M., Williamson, D., Coin, L., Duchene, S., 2020. Direct RNA sequencing and early evolution of SARS-CoV-2. bioRxiv.
[0145] 12. Li, H., 2011. A statistical framework for SNP calling, mutation discovery, association mapping and population genetical parameter estimation from sequencing data. Bioinformatics 27, 2987-2993.
[0146] 13. Petit, R., 2014. vcf-annotator, https: / / github.com / rpetit3 / vcf-annotator.
[0147] 14. Hyatt, D., Chen, G. L., Locascio, P. F., Land, M. L., Larimer, F. W., Hauser, L. J., 2010. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics 11, 119.
[0148] 15. Katoh, K., Misawa, K., Kuma, K., Miyata, T., 2002. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Res 30, 3059-3066.
[0149] 16. Minh, B. Q., Schmidt, H. A., Chernomor, O., Schrempf, D., Woodhams, M. D., von Haeseler, A., Lanfear, R., 2020. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Mol Biol Evol 37, 1530-1534.
[0150] 17. Lebrigand, K., Magnone, V., Barbry, P., Waldmann, R., 2020. High throughput error corrected Nanopore single cell transcriptome sequencing. Nat Commun 11, 4025.
[0151] 18. Peng, Q., Vijaya Satya, R., Lewis, M., Randad, P., Wang, Y., 2015. Reducing amplification artifacts in high multiplex amplicon sequencing by using molecular barcodes. BMC Genomics 16, 589.
[0152] 19. Schaecher, S.R., Mackenzie, J.M., Pekosz, A., 2007. The ORF7b protein of severe acute respiratory syndrome coronavirus (SARS-CoV) is expressed in virus-infected cells and incorporated into SARS-CoV particles. J Virol 81, 718-731.
[0153] 20. Bojkova, D., Klann, K., Koch, B., Widera, M., Krause, D., Ciesek, S., Cinatl, J., Munch, C., 2020. Proteomics of SARS-CoV-2-infected host cells reveals therapy targets. Nature.
[0154] 21. Gordon, D. E., Jang, G. M., Bouhaddou, M., Xu, J., Obernier, K., White, K. M., O’Meara, M. J., Rezelj, V. V., Guo, J. Z., Swaney, D. L., Tummino, T. A., Huettenhain, R., Kaake, R. M., Richards, A. L., Tutuncuoglu, B., Foussard, H., Batra, J., Haas, K., Modak, M., Kim, M., Haas, P., Polacco, B. J., Braberg, H., Fabius, J. M., Eckhardt, M., Soucheray, M., Bennett, M. J., Cakir, M., McGregor, M. J., Li, Q., Meyer, B., Roesch, F., Vallet, T., Mac Kain, A., Miorin, L., Moreno, E., Naing, Z. Z. C., Zhou, Y., Peng, S., Shi, Y., Zhang, Z., Shen, W., Kirby, I. T., Melnyk, J. E., Chorba, J. S., Lou, K., Dai, S. A., Barrio-Hernandez, I., Memon, D., Hernandez-Armenta, C., Lyu, J., Mathy, C. J. P., Perica, T., Pilla, K. B., Ganesan, S. J., Saltzberg, D. J., Rakesh, R., Liu, X., Rosenthal, S. B., Calviello, L., Venkataramanan, S., Liboy-Lugo, J., Lin, Y., Huang, X.-P., Liu, Y., Wankowicz, S. A., Bohn, M., Safari, M., Ugur, F. S., Koh, C., Savar, N. S., Tran, Q. D., Shengjuler, D., Fletcher, S. J., O’Neal, M. C., Cai, Y., Chang, J. C. J., Broadhurst, D. J., Klippsten, S., Sharp, P. P., Wenzell, N. A., Kuzuoglu, D., Wang, H.-Y., Trenker, R., Young, J. M., Cavero, D. A., Hiatt, J., Roth, T. L.Rathore, U., Subramanian, A., Noack, J., Hubert, M., Stroud, R.M., Frankel, A.D., Rosenberg, O.S., Verba, K.A., Agard, D.A., Ott, M., Emerman, M., Jura, N., von Zastrow, M., Verdin, E., Ashworth, A., Schwartz, O., d’Enfert, C., Mukherjee, S., Jacobson, M., Malik, H.S., Fujimori, D.G., Ideker, T., Craik, C.S., Floor, S.N., Fraser, J.S., Gross, J.D., Sali, A., Roth, B.L., Ruggero, D., Taunton, J., Kortemme, T., Beltrao, P., Vignuzzi, M., Garcia-Sastre, A., Shokat, K.M., Shoichet, B.K., Krogan, N.J., 2020. A SARS-CoV-2 protein interaction map reveals targets for drug repurposing. Nature.
[0155] 22. Konno, Y., Kimura, I., Uriu, K., Fukushi, M., Irie, T., Koyanagi, Y., Nakagawa, S., Sato, K., 2020. SARS-CoV-2 ORF3b is a potent interferon antagonist whose activity is further increased by a naturally occurring elongation variant. bioRxiv.
[0156] 23. Buchan, D.W.A., Jones, D.T., 2019. The PSIPRED Protein Analysis Workbench: 20 years on. Nucleic Acids Res 47, W402-W407.
[0157] 24. Lam, T. T.-Y., Shum, M. H.-H., Zhu, H.-C., Tong, Y.-G., Ni, X.-B., Liao, Y.-S., Wei, W., Cheung, W. Y.-M., Li, W.-J., Li, L.-F., Leung, G. M., Holmes, E. C., Hu, Y.-L., Guan, Y., 2020. Identifying SARS-CoV-2 related coronaviruses in Malayan pangolins. Nature.
[0158] 25. Pfefferle, S., V., Ditt, V., Grywna, K., Muhlberger, E., Drosten, C., 2009. Reverse genetic characterization of the natural genomic deletion in SARS-Coronavirus strain Frankfurt-1 open reading frame 7b reveals an attenuating function of the 7b protein in-vitro and in-vivo. Virol J 6, 131-131.
[0159] 26. Fan, H., Ooi, A., Tan, Y. W., Wang, S., Fang, S., Liu, D. X., Lescar, J., 2005. The nucleocapsid protein of coronavirus infectious bronchitis virus: crystal structure of its N-terminal domain and multimerization properties. Structure 13, 1859-1868.
[0160] 27. Liang, Y., Wang, M.-L., Chien, C.-S., Yarmishyn, A. A., Yang, Y.-P., Lai, W.-Y., Luo, Y.-H., Lin, Y.-T., Chen, Y.-J., Chang, P.-C., Chiou, S.-H., 2020. Highlight of Immune Pathogenic Response and Hematopathologic Effect in SARS-CoV, MERS-CoV, and SARS-Cov-2 Infection. Frontiers in Immunology 11.
[0161] 28. Molenkamp, R., Spaan, W. J., 1997. Identification of a specific interaction between the coronavirus mouse hepatitis virus A59 nucleocapsid protein and packaging signal. Virology 239, 78-86.
[0162] 29. Wu, C.-H., Chen, P.-J., Yeh, S.-H., 2014. Nucleocapsid Phosphorylation and RNA Helicase DDX1 Recruitment Enables Coronavirus Transition from Discontinuous to Continuous Transcription. Cell Host & Microbe 16, 462-472.
[0163] 30. Canton, J., Fehr, A.R., Fernandez-Delgado, R., Gutierrez-Alvarez, F.J., Sanchez-Aparicio, M.T., Garcia-Sastre, A., Perlman, S., Enjuanes, L., Sola, I., 2018. MERS-CoV 4b protein interferes with the NF-kappaB-dependent innate immune response during infection. PLoS Pathog 14, e1006838.
[0164] 31. Kissler, S.M., Tedijanto, C., Goldstein, E., Grad, Y.H., Lipsitch, M., 2020. Projecting the transmission dynamics of SARS-CoV-2 through the postpandemic period. Science 368, 860-868.
[0165] 32. Wrobel, A.G., Benton, D.J., Xu, P., Roustan, C., Martin, S.R., Rosenthal, P.B., Skehel, S.J., 2020. SARS-CoV-2 and bat RaTG13 spike glycoprotein structures inform on virus evolution and furin-cleavage effects. Nat Struct Mol Biol 27, 763-767.
[0166] 33. Aboubakr, H.A., Sharafeldin, T.A., Goyal, S.M., 2020. Stability of SARS-CoV-2 and other coronaviruses in the environment and on common touch surfaces and the influence of climatic conditions: A review. Transbound Emerg Dis.
[0167] 34. Hurst, K. R., Koetzner, C. A., Masters, P. S., 2013. Characterization of a critical interaction between the coronavirus nucleocapsid protein and nonstructural protein 3 of the viral replicase-transcriptase complex. J Virol 87, 9159-9172.
[0168] 35. Lei, J., Kusov, Y., Hilgenfeld, R., 2018. Nsp3 of coronaviruses: Structures and functions of a large multi-domain protein. Antiviral Research 149, 58-74.
[0169] 36. Chang, C.-k., Hou, M.-H., Chang, C.-F., Hsiao, C.-D., Huang, T.-h., 2014. The SARS coronavirus nucleocapsid protein - Forms and functions. Antiviral Research 103, 39-50.
[0170] 37. Wang, N., Li, S.-Y., Yang, X.-L., Huang, H.-M., Zhang, Y.-J., Guo, H., Luo, C.-M., Miller, M., Zhu, G., Chmura, A. A., Hagan, E., Zhou, J.-H., Zhang, Y.-Z., Wang, L.-F., Daszak, P., Shi, Z.-L., 2018. Serological Evidence of Bat SARS-Related Coronavirus Infection in Humans, China. Virol Sin 33, 104-107.
[0171] The scope of protection of the present application is not limited to the above-mentioned embodiments. Variations and improvements of the present application, which can be thought of by those skilled in the art without departing from the spirit and scope of the present application, are included in the present application and are protected by the scope of the attached claims. SEQUENCE LISTING <110> Shanghai Institute of Immunology <120> A method and system for identifying and analyzing coronavirus transcriptome <160> 2 <170> PatentIn version 3.5 <210> 1 <211> 111 <212> RNA <213> pORF2b <400> 1 augcuuggaa caggaagaga aucagcaacu guguugcuga uuauucuguc cuauauaauu 60 ccgcaucauu uuccacuuuu aaguguuaug gagugucucc uacuaaauua a 111 <210> 2 <211> 63 <212> RNA <213> tORF7b <400> 2 augcuuauua ucuuuugguu cucacuugaa cugcaagauc auaaugaaac uugucacgcc 60 uaa 63
Claims
1. A method of coronavirus transcriptome identification analysis, characterized by, The method comprises the following steps: Step one, using bwa software to sequence align the coronavirus second-generation sequencing raw fastq file obtained from the short read library SRA with the NCBI reference genome, to generate a BAM file; In step one, the sequence alignment refers to comparing the nucleic acid sequence obtained by nucleic acid sequencing technology with the reference sequence, thereby obtaining the presence or absence and arrangement difference of the base of the nucleic acid sequence and the reference sequence, for screening the nucleic acid sequence with breakpoints; the screening refers to finding the read with letter "H" or "S" in the CIGAR string in the BAM file of the alignment result, and recording the information thereof; the letter "H" refers to "hard clip", that is, the read has a phenomenon of complete mismatch at one end with the corresponding position of the reference sequence; the letter "S" refers to "soft clip", that is, the read has a phenomenon of inconsistency with the reference sequence at one end, but is not completely mismatched, and still has some bases consistent with the reference sequence; Step two, calling and filtering single nucleotide polymorphism (SNP) on the BAM file obtained in step one by bcftools, and annotating the SNP by vcf-annotator; Subsequently, the collected virus strains are grouped according to the presence or absence of a certain SNP for subsequent transcriptome analysis; At the same time, according to the base change indicated by the SNP, the specific base is replaced on the basis of the reference genome to generate a synonymous genome sequence; Step three, CIGAR string parsing and breakpoint identification operation are performed on the BAM file obtained in step one to identify sgRNA; In step three, the string parsing refers to obtaining the CIGAR string of the alignment of the aligned sequence and the reference sequence in the BAM file, extracting information therefrom, and determining the position of the sites where the aligned sequence is completely consistent, not completely consistent or completely inconsistent with the reference sequence, and obtaining the extension length of the three kinds of alignment on the aligned sequence; the breakpoint identification refers to extracting and recording the position of the junction of the region where the two kinds of alignment exist simultaneously in the aligned sequence and the reference sequence; the sgRNA identification refers to inferring the gene expression product to which the aligned sequence originally belongs according to the base position of the sequence on both sides of the breakpoint of the aligned sequence on the reference sequence; Step four, constructing an expression matrix of the sgRNA identification result in step three and performing transcriptome analysis, and in the analysis process, grouping the virus strains according to the presence or absence of SNP, source and project source, and statistically comparing the expression amount of each sgRNA of the virus strains in different groups.
2. The method of claim 1, wherein, In step one, the coronavirus includes all virus species under the Coronaviridae family.
3. The method of claim 1, wherein, In step two, the calling refers to a process of using bcftools to read the BAM file generated in step one, describing the base difference between the aligned sequence and the reference sequence and the position of the difference base; the filtering refers to a process of retaining the information of the difference base with an occupancy of more than 1% at the same position; and the annotation refers to explicitly screening out the difference base, recording whether the codon and amino acid at the site of the difference base are changed and / or how they are changed; The replacing specific base refers to a process of replacing the base corresponding to the position of the difference base on the reference sequence with the difference base.
4. The method of claim 1, wherein, The synonymous genomic sequence generated in step two is used for multiple sequence alignment with other coronaviruses that have completed whole genome sequencing, so as to construct an evolutionary tree of coronavirus strains.
5. The method of claim 1, wherein, In step four, the constructing expression matrix refers to a process of calculating the total number of each type of sgRNA identified in step three in the sample, calculating the proportion of each type of sgRNA in all sgRNAs in the sample, and listing the proportion of sgRNA in each sample with sample number as row, sgRNA type as column, or sgRNA as row, sample number as column.
6. A coronavirus transcriptome identification analysis system, characterized by, The system comprises a storage and a processor; the storage stores a computer program, and when the computer program is executed by the processor, the method of any one of claims 1-5 is realized.
7. Use of the method of any one of claims 1-5, or the system of claim 6 in coronavirus transcriptome identification analysis.