Tomato polypeptide genome database and construction method of index file of tomato polypeptide genome database
By constructing the tomato polypeptide genome database and its index file, the research problem of non-coding small peptides in tomatoes is solved, and efficient mining of non-coding peptides is achieved, providing new ideas for tomato molecular breeding.
Patent Information
- Application Number
- CN202411949750.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-16
AI Technical Summary
The existing technology is difficult to effectively explore and study the formation mechanism and molecular functions of non-coding small peptides in tomatoes, resulting in bottlenecks in tomato molecular breeding.
By constructing the tomato polypeptide genome database and its index file, the tomato genome was translated by using fadix and EMBOSS software to merge the polypeptide genome, and the index file was constructed through linux commands and self-written python scripts.
It provides new ideas for the mining of non-coding peptides of tomatoes, reduces research costs and workload, lays the foundation for in-depth research on endogenous peptides, and is of great significance to tomato molecular breeding.
Smart Images

Figure CN120015136A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of plant genomics and bioinformatics, and more specifically, to the construction of a tomato polypeptide genome database and an index file thereof. Background Art
[0002] Tomato is an annual herbaceous plant of the Solanaceae family. It originated in South America and has a wide range of adaptability. It can be grown all over the world. Because it is rich in various nutrients and tastes delicious, it has gradually become a high-demand fruit and vegetable. It also has complete genome information and is widely studied as a model horticultural crop. However, with the development of modern elite breeding lines, the genetic background has gradually narrowed, and tomato molecular breeding has also encountered bottlenecks. Therefore, it is crucial to explore new regulatory factors that affect tomato traits.
[0003] Small peptides are compounds formed by amino acids linked together by peptide bonds, usually composed of less than 100 amino acids. They have important biological functions in organisms and also have high development and application value. Generally speaking, there are two sources of endogenous peptides in plants: one is obtained by enzymatic cleavage of proteases at corresponding cleavage sites; the other is biologically active peptides encoded by nucleic acid sequences. At present, in addition to gene-encoded small peptides, small coding frames (sORFs) for translation have been found in 5'UTR regions, 3'UTR regions, intergenic regions, lncRNAs, and mRNA precursors, suggesting that non-coding peptides exist in plants and have potential biological functions.
[0004] With the gradual deepening of research on small peptides, it has been found that many endogenous peptides have hormone-like effects and are crucial in regulating plant growth and development, abiotic and biotic stress defense, and other processes. In terms of growth and development, the Arabidopsis IDA small peptide is involved in regulating the shedding of floral organs. At the same time, knocking out two IDA genes in the floral organ abscission zone in rapeseed also produces a phenotype in which the floral organs do not shed. In terms of stress adaptation, the salt-induced AtCAPE1 small peptide negatively regulates salt tolerance by inhibiting several salt-tolerant genes that play a role in the production of osmotic regulating substances, detoxification, stomatal closure control, and cell membrane protection. In terms of immune response, the SCOOP12 small peptide in Arabidopsis can not only inhibit root growth by regulating the accumulation of H2O2, but also activate plant immunity as a damage-related molecular pattern. In 1996, a non-coding small peptide (NCP) derived from lncRNA composed of 10 amino acids was first reported in plants and named ENOD40. Subsequent studies have shown that ENOD40 can respond to auxin stimulation in flowering plants. A peptide composed of 36 amino acids, POLARIS (PLS), was also found in Arabidopsis, which is expressed in both embryos and seedling roots and can regulate root length and leaf microtubule content. Subsequent studies have found that some plant pri-miRNAs contain sORFs, which can encode small peptides with regulatory functions. For example, primiR165a of Arabidopsis can encode miPEP165a. The external application of synthetic miPEP165a to plants can specifically stimulate the accumulation of primiR165a and affect the growth of the main root. Grape vvi-miPEP171d1 is a functional small peptide encoded by primary-miR171d. There are 3 sORFs in the 500bp sequence upstream of pre-miR171d, which can increase the expression of vvi-MIR171d. In addition, it was found that the external application of vvi-miPEP171d1 can promote the development of grape adventitious roots by activating the expression of vvi-MIR171d. In tomatoes, a study characterized the primary sequence of tomato miR396a using 5′RACE, and confirmed the presence of miPEP396a in tomatoes by verifying the translation activity of the start codon. It was further found that miPEP396a mainly functions in the nucleus and regulates the expression of pri-miR396a, miR396a and their target genes. In addition, transcriptomics and metabolomics analysis showed that in vitro synthesis of miPEP396a and spraying of tomatoes significantly increased the expression of phenylpropanoid biosynthesis and hormone-related genes, and at the same time, significantly inhibited the elongation of tomato primary roots. However, due to the limitations of prediction methods and omics technologies, the current research on non-coding small peptides in tomatoes is only the tip of the iceberg, and further in-depth research is urgently needed.
[0005] Peptidomics is the study of peptide groups in biological samples. Currently, the study of peptidomics mainly involves the comprehensive analysis of endogenous peptides using high performance liquid chromatography and mass spectrometry. One of the key steps is how to identify peptides from the obtained mass spectra. The most commonly used identification method is to match the mass spectrometry data with the peptide sequences in the reference database. However, the important problem is that many peptides do not exist in a specific reference protein sequence database, so it is necessary to search a self-built sequence database to identify new small peptides. Therefore, in order to identify and further study the non-coding small peptides in tomatoes, it is very important and necessary to construct a peptide database using reference genome data. Summary of the invention
[0006] The present invention provides a method for constructing a tomato polypeptide genome and an index file thereof. By constructing the tomato polypeptide genome and the index file, a new idea is provided for mining tomato non-coding peptides, which is of great significance to tomato molecular breeding.
[0007] In order to achieve the above object, the present invention is implemented by the following technical solutions:
[0008] The first aspect of the present invention provides a method for constructing a tomato polypeptide genome database, comprising the following steps:
[0009] S1. Use fadix to split the tomato genome and obtain multiple chromosomes;
[0010] S2. Use the sixpack command under EMBOSS 6.6.0 software to perform six-frame translation on multiple chromosomes obtained in S1 to obtain multiple amino acid sequences;
[0011] S3. Use merge to combine the multiple amino acid sequences obtained in S2 to obtain the tomato polypeptide genome.
[0012] In a preferred embodiment, the six-frame translation includes three positive-strand translation frames and three negative-strand translation frames; among them, F1, F2, and F3 are positive-strand translation frames, and F5, F6, and F4 are negative-strand translation frames of the positive strands translated by F1-F3, respectively.
[0013] In a preferred embodiment, the tomato polypeptide genome is the tomato SL4.0 genome.
[0014] The invention provides a method for constructing a tomato polypeptide genome. Tomato SL4.0 is used as a reference genome and is split according to chromosome numbers using fadix. Six-frame translation is performed on each of the split chromosome genomes based on the sixpack command under EMBOSS:6.6.0 software. The amino acid sequences obtained after the six-frame translation of all chromosomes are merged to obtain the tomato polypeptide genome.
[0015] In a second aspect of the present invention, there is provided a method for constructing a tomato polypeptide genome database index file constructed by any of the above construction methods, comprising the following steps:
[0016] (1) Using Linux commands to extract the sequence length, sequence name and amino acid arrangement of multiple amino acid sequences obtained from any of the above S2;
[0017] (2) Use the paste command to merge the sequence length, sequence name, and amino acid sequence information extracted in (1), and then use the grep command to split the merged file into six reading frames (F1-F6);
[0018] (3) The positions of the nucleotides in each amino acid sequence are calculated in batches according to the length of the sequence and the bed file is output.
[0019] In a preferred embodiment, the position of the sequence on the genome is calculated based on the length of the amino acid sequence, and the calculation method is as follows:
[0020] Starting position = the end position of the previous sequence + 1
[0021] End position = end position of previous sequence + [(length of this sequence + 1) * 3]
[0022] Among them, the starting value of the first amino acid sequence of the F1 and F5 translation frames is 1, the starting value of the 4th column of the F2 and F6 translation frames is 2, and the starting value of the 4th column of the F3 and F4 translation frames is 3.
[0023] The invention provides a method for constructing a tomato polypeptide genome index file. The method comprises the following steps: using a Linux command to extract the sequence length, sequence name and amino acid arrangement of the fasta obtained after translating the six frames of each chromosome; merging the three types of information into one file using a paste command; splitting the fused file obtained above according to the six reading frames; running a self-written python script to calculate the position of the nucleotide of each fasta amino acid according to its length, and outputting a bed file, so as to obtain an index file of the tomato polypeptide genome.
[0024] Beneficial effects of the present invention:
[0025] At present, there are no reports on the construction of a tomato polypeptide genome database. Tomato non-coding small peptides play an important role in regulating its growth and development and adversity defense. The present invention uses a method for six-frame translation of tomato reference genes to construct a tomato polypeptide genome, laying the foundation for further exploration and in-depth research on the formation mechanism and molecular function of endogenous peptides. In addition, a method for constructing a tomato polypeptide genome index file has also been invented, which can narrow the scope of the research object, is low-cost, and greatly reduces the workload. Based on the existing reference genome of tomato, the present invention obtains the polypeptide genome through six-frame translation, and constructs an index file of the polypeptide genome according to a self-written script, so as to mine non-coding small peptides in tomatoes, providing a new idea for tomato molecular breeding. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 A method for constructing a tomato polypeptide genome database;
[0027] Figure 2 A method for constructing a tomato polypeptide genome index file;
[0028] Figure 3 The sequence length distribution in the tomato polypeptide genome. DETAILED DESCRIPTION
[0029] The technical solutions in the embodiments of the present invention are described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0030] Based on the existing reference genome of tomato, the present invention obtains the polypeptide genome through six-frame translation, and constructs the index file of the polypeptide genome according to the self-written script, so as to mine the non-coding small peptides in tomato, laying the foundation for in-depth study of its regulatory phenotype mechanism.
[0031] The present invention provides a method for constructing a tomato polypeptide genome:
[0032] (1) Using tomato SL4.0 as the reference genome, fadix was used to split it according to chromosome numbers;
[0033] (2) Based on the sixpack command under EMBOSS:6.6.0 software, six-frame translation was performed on the genome of each chromosome after splitting; the amino acid sequences obtained after the six-frame translation of all chromosomes were merged to obtain the tomato polypeptide genome.
[0034] The present invention also provides a method for constructing a tomato polypeptide genome index file:
[0035] (1) Using Linux commands, extract the fasta sequence length, sequence name and amino acid sequence obtained after the six-frame translation of each chromosome;
[0036] (2) Use the paste command to merge the above three types of information into one file;
[0037] (3) Split the fusion file obtained above into 6 reading frames (F1-F6) to obtain 6 files;
[0038] (4) Take the six files of each chromosome as input files, run the self-written Python script to calculate the position of the nucleotides according to the length of each fasta amino acid, and output the bed file to obtain the index file of the tomato polypeptide genome.
[0039] Example 1: Method for constructing a tomato polypeptide genome
[0040] A method for constructing a tomato polypeptide genome database, the specific method comprising: using the sixpack command under EMBOSS:6.6.0 software to translate the tomato SL4.0 genome after it is split according to chromosomes into six frames (F1-F6, wherein F1, F2, and F3 are positive chain translation frames, and F5, F6, and F4 are corresponding negative chain translation frames), and then combining the amino acid sequences after all chromosome translations ( Figure 1 ).
[0041] 1. Download of tomato reference genome: Download the tomato SL4.0 reference genome sequence and ITAG4.0 annotation file from the Sol Genomics Network (https: / / solgenomics.net / organism / solanum_lycopersicum / genome).
[0042] 2. Download and install samtools software: Download the samtools 1.18 installation package from the samtools official website (https: / / github.com / samtools / samtools / releases / download / ), and decompress and install it in the Linux system according to the instructions.
[0043] 3. Download and install EMBOSS software: Download the EMBOSS: 6.6.0 installation package from the EMBOSS official website (http: / / emboss.sourceforge.net / download.html), and unzip and install it into the Linux system according to the instructions.
[0044] 4. Splitting of the reference genome: Use the fadix-x S_lycopersicum_chromosomes.4.00.fa command to split the tomato reference genome into 13 fasta files according to different chromosome numbers.
[0045] 5. Six-frame translation: Use the sixpack SL4.0ch**.fa command to perform six-frame translation on the fasta sequence of each chromosome. Two files are obtained for each chromosome, one is the amino acid fasta file, and the other is the translation information file with the suffix sixpack.
[0046] 6. Merge all chromosome translated sequences: Use cat SL4.0ch**.fa>Sl.allchr.aa.fasta command to merge all chromosome translated amino acid fasta files together, which is the tomato polypeptide genome ( Figure 1 ), totaling 9.67 Gb, including 97,144,798 sequences, indicating that the constructed peptide genome database is of a reasonable size for subsequent peptide identification.
[0047] Example 2: Method for constructing a tomato polypeptide genome index file
[0048] A method for constructing a genome database index file constructed by a construction method in embodiment 1 comprises the following steps:
[0049] (1) The linux command was used to extract the sequence length, sequence name and amino acid arrangement of the fasta obtained after the six-frame translation of each chromosome;
[0050] (2) The above three types of information are merged using the paste command and split according to six reading frames;
[0051] (3) Use a Python script named getbed-F*.py to calculate the position of the nucleotides of each fasta amino acid according to its length and output the bed file. The specific calculation method is: the starting value of the first amino acid sequence of the F1 and F5 translation frames is 1, written in the 4th column, calculate the value of (amino acid length + 1) * 3, and write it in the 5th column; the value of the 4th column of the second amino acid sequence is the value of the 5th column of the previous row plus 1, and the value of the 5th column is the value of the 5th column of the previous row plus (amino acid length of this row + 1) * 3, and so on. After calculating the position of all amino acid sequences on the genome ( Figure 2 ). Similarly, the starting value of the 4th column of the F2 and F6 translation boxes is 2, and the starting value of the 4th column of the F3 and F4 translation boxes is 3. For the above calculation, the Python script of getbed-F*.py is as follows:
[0052]
[0053]
[0054] 1. Extract the fasta sequence length obtained after six-frame translation of each chromosome: Use the command grep">"chr**.aa.fasta|cut-f 2|cut-d,-f 4>chr**.length.txt to obtain the sequence length of all polypeptides (** represents the chromosome number).
[0055] 2. Extract the sequence names obtained after six-frame translation of each chromosome: Use the command grep">"chr**.aa.fasta|cut-d""-f 1>chr**.name.txt to obtain the sequence names of all polypeptides.
[0056] 3. Extract the amino acid arrangement obtained after six-frame translation of each chromosome: Use awk' / ^> / &&NR>1{print"";}{printf"%s", / ^> / ?$0"\n":$0}'chr**.aa.fasta|sed'$s / $ / \n / '|grep-v">">chr**.seq.txt command to obtain the amino acid arrangement of all polypeptides.
[0057] 4. Merge the above files: Use the paste chr**.length.txt chr**.name.txt chr**.seq.txt>chr**.inf.txt command to merge the three files obtained above into one file.
[0058] 5. Remove the aa characters in the first column: Use the sed -i`s / aa / / g`chr**.inf.txt command to remove the aa characters in the file.
[0059] 6. Split the sequence into 6 reading frames: Use a for loop to split each inf file into 6 reading frames. The command is as follows:
[0060]
[0061] 7. Copy the first row as the column name: Use the for loop to copy the first row of the above files as the column name. The command is as follows:
[0062]
[0063] 8. Run the python script to add nucleotide position information (* indicates reading frame number). The loop command is as follows:
[0064]
[0065] 9. Combine all the above bed files and use the awk'$1>=1{print$0}'OFS="\t" command to select the sequences with amino acid sequences greater than or equal to 1, which is the tomato polypeptide genome index file ( Figure 2 ).
[0066] 10. According to the first column of the bed file, the length distribution of peptide sequences can be counted. About 96.14% (93,392,260 sequences) of peptides are less than 50 amino acids ( Figure 3 ), indicating that the peptide fragments of the constructed peptide genome database are smaller in length and suitable for the identification of non-coding peptides.
[0067] The present invention provides a method for constructing a tomato polypeptide genome, including: using tomato SL4.0 as a reference genome, splitting it according to chromosome numbers using fadix; performing six-frame translation on each chromosome genome after splitting based on the sixpack command under EMBOSS: 6.6.0 software; merging the amino acid sequences obtained after the six-frame translation of all chromosomes to obtain the tomato polypeptide genome. A method for constructing a tomato polypeptide genome index file is provided, including: extracting the sequence length, sequence name and amino acid arrangement of fasta obtained after the six-frame translation of each chromosome using a linux command; merging the above three types of information into one file using a paste command; splitting the above fusion file according to six reading frames; running a self-written python script to calculate the position of the nucleotide according to the length of each fasta amino acid, and outputting a bed file to obtain an index file of the tomato polypeptide genome, and according to the index file, the position of the nucleotide sequence of the polypeptide translation on the genome can be found. The present invention provides a new idea for mining non-coding peptides of tomatoes by constructing a tomato polypeptide genome and an index file, which is of great significance to tomato molecular breeding.
[0068] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for constructing a tomato polypeptide genome database, characterized in that: The steps include: S1. Use fadix to split the tomato genome and obtain multiple chromosomes; S2. Use the sixpack command under EMBOSS 6.6.0 software to perform six-frame translation on multiple chromosomes obtained in S1 to obtain multiple amino acid sequences; S3. Use merge to combine the multiple amino acid sequences obtained in S2 to obtain the tomato polypeptide genome.
2. The method for constructing a tomato polypeptide genome database according to claim 1, characterized in that: The six-frame translation includes three positive-chain translation frames and three negative-chain translation frames; among them, F1, F2, and F3 are positive-chain translation frames, and F5, F6, and F4 are negative-chain translation frames of the positive chains translated by F1-F3, respectively.
3. The method for constructing a tomato polypeptide genome database according to claim 1, characterized in that: The tomato polypeptide genome is the tomato SL4.0 genome.
4. A method for constructing a tomato polypeptide genome database index file constructed by any one of the construction methods of claims 1 to 3, characterized in that: The steps include: (1) using Linux commands to extract the sequence length, sequence name and amino acid arrangement of multiple amino acid sequences obtained in S2 of any one of claims 1 to 3; (2) Use the paste command to merge the sequence length, sequence name and amino acid sequence information extracted in (1), and then use the grep command to split the merged file according to the six translation frames; The six translation frames include three positive-strand translation frames and three negative-strand translation frames, among which F1, F2, and F3 are positive-strand translation frames, and F5, F6, and F4 are negative-strand translation frames of the positive strands translated by F1-F3, respectively; (3) The positions of nucleotides are calculated in batches according to the length of each amino acid sequence and the bed file is output.
5. The method for constructing an index file of a Solanum polypeptide genome database according to claim 4, characterized in that: The position of the sequence on the genome is calculated based on the length of the amino acid sequence. The calculation method is as follows: Starting position = the end position of the previous sequence + 1 End position = end position of previous sequence + [(length of this sequence + 1) * 3] Among them, the starting value of the first amino acid sequence of the F1 and F5 translation frames is 1, the starting value of the 4th column of the F2 and F6 translation frames is 2, and the starting value of the 4th column of the F3 and F4 translation frames is 3.