A method for high-throughput analysis of multi-species stoichiometric transcriptomes
Patent Information
- Application Number
- CN202210159540.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-17
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2042-02-17
AI Technical Summary
目前化学计量转录组的计算方法涉及复杂的计算、统计学理论和分析方法,缺乏可视化应用
[0023]本发明具有以下优点:一是,能够快速地计算多个物种转录组化学计量分析的结果,各参数较为全面和准确,效果好,速度快;二是比较系统,效率高,自动化,能够实验多个物种转录组之间的比较分析,高通量处理数据;三是本发明将Perl语言脚本编程与几个R语言脚本编程完美流畅的结合起来,实现了软件之间的良好衔接和数据的可视化。
Smart Images

Figure CN115938474B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology, and in particular relates to a method for high-throughput analysis of chemometric transcriptomes of multiple species. Background Technology
[0002] Transcriptomics is a discipline that systematically studies gene transcription maps at the overall transcriptional level and reveals the molecular mechanisms of complex biological pathways and trait regulatory networks. In recent years, with the development of high-throughput sequencing technology, transcriptome sequencing can not only detect transcripts corresponding to existing genome sequences but also discover and quantify new transcripts, offering significant advantages in the study of alternative splicing events, new genes and transcripts, and fusion transcripts. Through existing transcriptome sequencing technologies, the system can accurately reveal complex traits in biological processes and analyze transcriptional regulatory networks, leading to its widespread application in basic research, clinical diagnosis, and drug development. The continuous generation of multi-species transcriptome data drives the development of multi-species analysis workflows. Analyzing multi-species transcriptome data will help elucidate the origin relationships and environmental adaptations among species. Convenient and rapid transcriptome data analysis methods are therefore particularly important. Based on transcriptome data, by calculating the element usage preferences in the transcriptome, we can not only understand the evolutionary patterns of element usage preferences but also use this as an assessment of the degree of regulatory responses organisms make to their surrounding environment. Chemostometric transcriptomics refers to the computational methods used in computational biology to calculate the stoichiometric characteristics of transcriptome or mRNA sequences, including the composition and abundance of elements (carbon, hydrogen, oxygen, nitrogen) and monomers (nucleotides). Chemostometric transcriptomics is an emerging interdisciplinary field encompassing chemometrics, ecology, evolutionary biology, transcriptomics, and bioinformatics. It provides a theoretical foundation for scientific research on the regulation of transcriptional levels and offers a comprehensive perspective for data mining in the post-transcriptomics era. Currently, computational methods for chemmostometric transcriptomics involve complex calculations, statistical theories, and analytical methods, lacking visualization applications. This poses a significant challenge for non-professional researchers. Furthermore, the downloaded transcriptome files require extensive preprocessing to convert them into input files with a fixed format, greatly limiting analysis by those without bioinformatics expertise or with relatively weak computer skills, ultimately hindering research in the field of biochemical composition. Summary of the Invention
[0003] The purpose of this invention is to provide a high-throughput method for analyzing the chemometric transcriptome of multiple species.
[0004] The method for high-throughput analysis of multi-species chemometric transcriptomes provided by the present invention may specifically include the following steps;
[0005] Step 1: Record the transcriptome sequence file of the first species to be tested as transcriptome data A1 (fasta or fastq format), the transcriptome sequence of the second species or transcriptome files of multiple species as transcriptome data A2 (fasta or fastq format)... Put the data A1, A2, A3, etc. of multiple species (preferably ≤4 species) into the folder in, and create a new folder out.
[0006] Step 2: Perform base and element content analysis on the transcriptome data in the folder "in". Specific operations include: traversing the transcriptome sequences in the input file, counting the number of each base in each sequence, and calculating the total amount of carbon, hydrogen, oxygen, and nitrogen elements in the sequence based on the number of atoms corresponding to each base; after calculation, output a statistical file B (output1.xls) containing base and element content data for each transcriptome sequence, and a statistical file C (output2.xls) containing average element content data for the transcriptome.
[0007] In a preferred embodiment of the present invention, step 2 above can be implemented by running the Perl script 1 "perlcount4RNA.pl", the specific code of which is as follows:
[0008]
[0009]
[0010]
[0011]
[0012]
[0013]
[0014] Step 3: Visualize the transcriptome base and element content data B. Specific operations include: reading the statistical file B, classifying the data according to species name, and using statistical plotting tools to create a comparative chart of element content distribution across different species.
[0015] In a preferred embodiment of the present invention, step 3 above can be achieved by running R script 1 (run "transcriptomics-violin plot.R" on UNIX / Linux / MacOSX systems; or directly run "transcriptomics-violin plot.R" in R or Rstudio on Windows) to obtain a comparison chart of element content distribution among different species. The specific code of R script 1 (transcriptomics-violin plot.R) is as follows:
[0016]
[0017]
[0018] Step 4: Visualize the average content data C of transcriptome elements. Specific operations include: reading the statistical file C, and using statistical plotting tools to create a Nightingale rose diagram and bar chart to visualize the average content. The Nightingale rose diagram is used to show the elemental composition ratios of different species, and the bar chart is used to show the base composition.
[0019] In a preferred embodiment of the present invention, step 4 above can be achieved by running R script 2 (run on UNIX / Linux / MacOSX systems: transcriptomics-Nightingale rose plot-bar chart.R; or directly run in Windows R or Rstudio: transcriptomics-Nightingale rose plot-bar chart.R) to obtain an average content visualization - Nightingale rose plot and bar chart. The specific code of the R script 2 (transcriptomics-Nightingale rose plot-bar chart.R) is as follows:
[0020]
[0021]
[0022] In this invention, the species to be tested in step 1 can be any species, and the transcriptome sequence can be obtained by downloading from a publicly available transcriptome database or by transcriptome sequencing.
[0023] The present invention has the following advantages: First, it can quickly calculate the results of chemometric analysis of transcriptomes of multiple species, with comprehensive and accurate parameters, good results, and fast speed; Second, it is a comparison system that is efficient, automated, and capable of comparative analysis between transcriptomes of multiple species, and high-throughput data processing; Third, the present invention perfectly and smoothly combines Perl language scripting with several R language scripting, achieving good integration between software and data visualization. Attached Figure Description
[0024] Figure 1 This is a flowchart of the high-throughput analytical chemometric transcriptomics process of the present invention;
[0025] Figure 2 This is a visualization of the transcriptome base and element content data B from R script 1 in step 3, including a comparison of element content distribution among different species.
[0026] Figure 3 The visualization of the average content data C of transcriptome elements in R script 2 in step 4 includes the average content visualization - Nightingale rose plot and bar chart. Detailed Implementation
[0027] The present invention will be described in more detail below using transcriptome data from *Halobacterium*, *Bacillus subtilis*, *Escherichia coli*, and *Arabidopsis thaliana* as examples. The data was downloaded from the NCBI database (https: / / www.ncbi.nlm.nih.gov).
[0028] The flowchart of the high-throughput analytical chemometric transcriptomics provided by this invention is shown below. Figure 1 Specifically, it includes the following steps:
[0029] Step 1: Record the transcriptome sequence file of the first species to be tested as transcriptome data A1 (fasta or fastq format), the transcriptome sequence of the second species or transcriptome files of multiple species as transcriptome data A2 (fasta or fastq format)... Put the data A1, A2, A3, etc. of multiple species (preferably ≤4 species) into the folder in, and create a new folder out.
[0030] Step 2: Operate under Linux and install the Perl software. Perform base and element content analysis on the transcriptome data in the folder "in". The core logic of this step is: the program reads the input file, identifies each sequence, and counts the bases (A, T / U, G, C) in the sequence. Then, it performs cumulative calculations according to preset stoichiometry rules (i.e., the number of C, H, O, N atoms in each base). For example, if the base A is identified, the counts of carbon, hydrogen, nitrogen, and oxygen are increased accordingly; if the base G is identified, the counts of the corresponding element are increased, and so on. Finally, the program writes the detailed data of each sequence to file B (output1.xls), with the species name taken from the input file name (excluding the extension), and writes the calculated species average data to file C (output2.xls). As a specific implementation method, run Perl script 1 "perl count4RNA.pl". The Perl script 1 is specifically: count4RNA.pl
[0031]
[0032]
[0033]
[0034]
[0035]
[0036]
[0037] Furthermore, the B(output1.xls) file format described in this invention is specifically as follows:
[0038]
[0039] Furthermore, the C(output2.xls) file format described in this invention is specifically as follows:
[0040]
[0041] Step 3: Visualize the transcriptome base and element content data B. The core logic of this step is to use plotting packages such as ggplot2 in R to read the output1.xls file generated in Step 2. The program first parses the sample ID to identify the species origin, and then plots violin plots and boxplots for elements such as carbon (C), hydrogen (H), oxygen (O), and nitrogen (N), and combines these plots to visually display the differences in element content distribution among different species. As a specific implementation method, run R script 1 (run "transcriptomics-violinplot.R" on UNIX / Linux / MacOSX systems; or run directly in R or Rstudio on Windows: transcriptomics-violinplot.R) to obtain a comparison plot of element content distribution among different species - such as... Figure 2 As shown. The specific R script 1 is: transcriptomics-violin diagram.R
[0042]
[0043]
[0044] Step 4: Visualize the average content data C of bases and elements. The core logic of this step is to read the output2.xls file generated in Step 2, reshape the data format (Melt) to adapt to the plotting requirements. The program draws a Nightingale Rose Chart to show the relative proportions of each element, and a stacked bar chart to show the composition of bases, and outputs the results as a PDF file. As a specific implementation method, run R script 2 (run "transcriptomics-Nightingale Rose-Bar Chart.R" on UNIX / Linux / MacOSX systems; or run directly in R or Rstudio on Windows: transcriptomics-Nightingale Rose-Bar Chart.R) to obtain the average content visualization - Nightingale Rose Chart and bar chart, as shown. Figure 3 As shown. The R script 2 is specifically: transcriptomics-Nightingale rose plot-bar chart.R
[0045]
Claims
1. A method for high-throughput analysis of chemometric transcriptomes of multiple species, characterized in that, The process includes the following steps: Step 1, using transcriptome sequence files of multiple species as input data (FASTA or FASTQ format), placing the transcriptome sequence files of each species in the input folder "in", and creating an output folder "out"; Step 2, iterating through each transcriptome sequence in the input folder "in", counting the number of bases in each sequence, and calculating the number of element atoms corresponding to each base according to a preset base element measurement rule to obtain the base composition data and element content data of each transcriptome sequence, and outputting the first statistical file B; simultaneously, summing and calculating the element content data of multiple transcriptome sequences of the same species to obtain the average element content data of the species, and outputting the second statistical file C; Step 3, reading statistical file B, extracting species names according to sample identifiers or input file names and classifying them by species, and drawing a distribution comparison chart for comparing the element content distribution of different species, including violin plots and box plots; Step 4, reading the second statistical file C, and drawing a Nightingale rose plot to show the element composition ratio of different species and a bar chart to show the base composition; The base element measurement rules in step 2 include counting the A, G, C, and T / U bases appearing in the sequence, and converting each base count into the cumulative value of the corresponding element based on the measurement table of each base. The first statistical file B includes at least the sequence identifier, sequence length, element content field, and base composition field. The measurement table is a pre-determined base-element atom number correspondence table, used to convert each base count into the cumulative value of the corresponding carbon, hydrogen, oxygen, and nitrogen element atom numbers.
2. The method according to claim 1, characterized in that: The species to be tested in step 1 can be any species, and the number of species is no more than 4.
3. The method according to claim 1, characterized in that: The distribution comparison diagram mentioned in step 3 includes violin plots and box plots drawn for carbon, hydrogen, oxygen, and nitrogen elements respectively, and the violin plots and box plots are output as the same visualization file.
4. The method according to claim 1, characterized in that: The Nightingale rose diagram described in step 4 uses polar coordinates to display the elemental composition ratios of different species, and the bar chart uses stacked bars to display the base composition of different species. The Nightingale rose diagram and the bar chart are then output as the same visualization file.
Citation Information
Patent Citations
Method for batch analysis of stoichiometric genomes of multiple species
CN115700884A
Calculation method for basic analysis of stoichiometric transcriptome
CN116052763A