A method for predicting and identifying small proteins in marine Streptomyces S187
By establishing a database for whole-genome prediction and transcriptome detection, and combining DIA mass spectrometry and de novo sequencing methods, the problems of insufficient identification accuracy and coverage of small open reading frames in bacteria were solved, and the precise identification of bacterial small proteins and improved peptide coverage were achieved, which supplemented the limitations of database searches.
Patent Information
- Application Number
- CN202310513839.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-09
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2043-05-09
AI Technical Summary
Existing technologies have problems with low accuracy and insufficient coverage when identifying small open reading frames (smORFs) in bacteria, especially in non-model organisms, where it is difficult to identify new coding polypeptides (SEPs). Traditional mass spectrometry is easily interfered by large protein peptide signals, and database retrieval methods have great limitations when gene annotations are limited.
A database of whole-genome prediction and transcriptome detection was established, combined with high-precision DIA mass spectrometry detection and de novo sequencing methods. Small proteins were identified through database search and database-independent identification methods combined with model strain data. Data analysis was performed using a variety of software tools to improve identification accuracy and coverage.
It has achieved accurate and comprehensive identification of small proteins in bacteria, improved the peptide coverage and the reliability of quantitative information, supplemented the deficiencies of database searches, discovered new SEPs and ncRNAs, and enhanced the reference value of non-model organism research.
Smart Images

Figure CN116741276B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of identification of micropeptides in bacteria, and in particular to a method for predicting and identifying small proteins in marine Streptomyces S187. Background Art
[0002] Small open reading frames (smORFs) are DNA sequences in eukaryotes and bacteria that can be translated from fewer than 100 codons. They are widely distributed in the genomes of various species, and a growing body of research suggests they have important functions. However, smORFs are often overlooked, often considered non-coding due to their small size or unusual structure.
[0003] In recent years, with the advancement of next-generation sequencing (NGS) technology, various bioinformatics-based methods have been developed to explore thousands of smORFs. Mass spectrometry (MS) is the most versatile tool for directly detecting and quantifying smORF encoded polypeptides (SEPs). However, MS-based SEP identification still has some drawbacks. For example, because SEPs are typically expressed at low levels and have small molecular weights, they can be missed during MS detection due to interference from peptide signals from highly expressed large proteins. Therefore, specific enrichment and high-precision detection methods for SEPs are needed to improve the accuracy and comprehensiveness of SEP identification. Compared to data-dependent acquisition (DDA), data-independent acquisition (DIA) is an emerging technology in proteomics that requires the generation of spectral libraries from DDA experiments. In comparison, DIA offers better data utilization and reproducibility. Notably, DIA can detect dynamic changes in SEP expression during different growth stages, with a low coefficient of variance and high peptide quantification accuracy. Furthermore, the mainstream protein profiling technique, known as the "shotgun" method, involves mass spectrometry analysis of peptide mixtures generated by enzyme digestion (most commonly trypsin) in complex protein samples. However, the resulting peptide fragments can lose protein information or reduce protein detection coverage. Due to the relatively small size of small proteins, top-down protein profiling techniques are increasingly favored by researchers.
[0004] Currently, database retrieval has become the main method for peptidomics research due to its high accuracy and simple operation. Fully annotated protein databases are already available for humans, mice, and common laboratory model organisms. However, due to the lack of available public databases, it is difficult to identify new SEPs in non-model organisms. De novo sequencing directly infers peptide sequences by comparing mass spectra with mass differences in amino acid residues, avoiding dependence on databases when discovering more new SEPs. De novo sequencing can also increase peptide sequence coverage, achieve fairly high computational speeds, and detect mutant amino acids. However, it cannot provide quantitative data on expression or the function of annotated SEPs like database searches. Summary of the Invention
[0005] To address the above-mentioned shortcomings, the present invention provides a method for predicting and identifying small proteins in marine Streptomyces S187. The present invention helps to detect polypeptides with longer amino acid lengths in peptide sequencing and improve the coverage of protein-corresponding polypeptides, providing more reliable peptide segment positioning and quantitative information for subsequent analysis.
[0006] A method for predicting and identifying small proteins in Streptomyces marineus S187 comprises the following steps:
[0007] 1) Establishment of a comprehensive database: Customize a database for non-model strains that includes genome prediction and transcriptome detection;
[0008] 2) Sample processing: Peptides from samples at different metabolic time points are extracted and enriched, and the enriched micropeptides are then detected by high-precision DIA mass spectrometry;
[0009] 3) SEP identification result analysis: The qualitative and quantitative results of the sample micropeptides were obtained through mass spectrometry data analysis, and the mass spectrometry data were used simultaneously for database search and database-independent de novo sequencing identification methods to find novel peptides and ncRNAs.
[0010] In one embodiment of the present invention, step 1) includes the following process:
[0011] 1.1) Whole-genome prediction: Two softwares were selected to predict ORFs in the whole genome of S187:
[0012] The S187 whole-genome fasta sequence was submitted to ORFFinder for six-frame translation, using four start codons: ATG, GTG, CTG, and TTG, and selecting the bacterial codon 11 translation table. Prodigal software was used to perform genome-wide ORF prediction as a supplement to the ORFFinder tool, thereby obtaining the ribosome binding motifs corresponding to the ORFs. The S187 whole-genome coding sequence was obtained using the bacterial codon 11 table as the operating parameter. The lengths of these amino acid sequences were counted, and ORFs with a length of less than 100 aa were selected as smORFs.
[0013] 1.2) Collinearity Analysis: The model strain of Streptomyces coelicolor was selected as a reference. The MCScanX tool in TBtools software was used to perform genomic collinearity analysis with S187 to explore the similarity between the two strains.
[0014] 1.3) Culture and sampling:
[0015] The preserved strain of Streptomyces xinghaiensis S187 was streaked and activated on TSB solid medium, placed in a constant temperature incubator, and the activated strain was used to prepare fermentation seed liquid;
[0016] Sampling: First, the fermentation curve and anti-complement activity curve of S187 were used to determine the activity cycle of the strain. The sampling time points were determined according to the fermentation curve and anti-complement activity curve: 36h, 48h, 72h, and 120h. Then, the samples were sampled, eluted, and stored.
[0017] 1.4) Transcriptome sequencing library construction:
[0018] Samples from the four time points were combined for RNA-seq library construction. The specific steps are as follows: After RNA extraction, purification, and library construction, the libraries were sequenced using PAIRED paired-end sequencing technology based on the Illumina HiSeq sequencing platform using NGS sequencing technology. All transcripts obtained from transcriptome sequencing were searched for open reading frames using NCBI ORFFinder, with the minimum ORF length set to 30, the Genetic code set to 11.Bacterial, and the ORF start codon selected as any start codon. The resulting six-frame translation data of the transcriptome was used as the library search database and searched during the peptide sequencing process.
[0019] Transcriptome data smRNAs prediction: FastaQC and Trimmomatic were used for quality control of transcriptome sequencing data, Bowtie2 and featureCounts were used for read alignment and quantification, and RNA with FPKM>1 and length≤303nt was screened as a candidate smRNAs library.
[0020] In one embodiment of the present invention, in step 1.3), preparing a fermentation seed solution from the activated strain includes the following steps:
[0021] Seed liquid collection: Take the S187 bacterial block activated by TSB solid medium, add it to TSB liquid medium, shake and culture at 210-230 rpm, and collect the fermentation liquid after 36 hours;
[0022] Fermentation broth collection: Add the seed solution to M33 medium and culture at 170-190 rpm with shaking. Increase the speed by 20 rpm every 48 hours until it reaches 210-230 rpm and then maintain the speed unchanged.
[0023] In one embodiment of the present invention, in step 1.3), sampling, elution, and preservation include the following steps:
[0024] 1.3.1) From the six bottles of fermentation broth from the same batch, sample three bottles of fermentation broth that are growing well. 3 mL of the sample from each bottle was added to a 50 mL centrifuge tube and mixed thoroughly. From the remaining three bottles of fermentation broth, 9 mL of the sample was transferred to a 50 mL centrifuge tube. The sample was then processed as for the first bottle to balance the sample. 10 mL of the sample was transferred to a centrifuge tube from each of the remaining two bottles for extraction and rotary evaporation. The centrifuge tubes should be kept in an ice bath throughout the entire experimental process.
[0025] 1.3.2) Removal of proteins from the culture medium: centrifuge and discard the supernatant;
[0026] 1.3.3) PBS elution: Add pre-chilled PBS solution to the centrifuged pellet and mix thoroughly. Then add PBS, vortex, centrifuge and discard the supernatant.
[0027] 1.3.4) Repeat step 1.3.3) once;
[0028] 1.3.5) Removal of CaCO3: Add pre-chilled PBS solution and mix thoroughly. Then add PBS and vortex. Centrifuge and discard the precipitate. Finally, transfer the sample to a new centrifuge tube.
[0029] 1.3.6) Repeat step 1.3.3) once more to perform elution;
[0030] 1.3.7) Place the centrifuge tube in a liquid nitrogen tank for quick freezing and store at low temperature.
[0031] In one embodiment of the present invention, step 2) includes the following process:
[0032] 2.1) Enrichment and processing of micropeptides:
[0033] Extraction of small proteins and peptides: S187 mycelial samples obtained by sampling and separation at fermentation time were crushed, and then small peptides were extracted with methanol:chloroform:water = 3:1:4 (v:v); vortexed and centrifuged; the upper aqueous supernatant was collected and centrifuged at 20,000 × g using a 10 kDa molecular weight cutoff (MWCO) and the lower solution was removed, vacuum concentrated, dried, desalted, and lyophilized.
[0034] HPLC fractionation: A portion of peptides from each sample is mixed and then separated into multiple fractions by high performance liquid chromatography;
[0035] 2.2) Peptide sequencing:
[0036] DDA mass spectrometer: Prepare mobile phases A and B; dissolve lyophilized powder in 10 μL of phase A, centrifuge at 14,000 g for 20 min at 4°C, and inject 1 μg of the supernatant for LC / MS / MS detection.
[0037] DIA mass spectrometer operation and database search: Dissolve the peptide powder in 0.1% (v / v) formic acid aqueous solution and add iRT standard peptides. Take 1 μg sample for injection and perform liquid chromatography-mass spectrometry detection.
[0038] In one embodiment of the present invention, step 3) includes the following steps:
[0039] 3.1) Database search and identification of SEPs: The quantitative values of the sample peptides were analyzed using the t-test, with a chi-square P value ≤ 0.05 and a fold change ≥ 1.5 to identify differentially expressed peptides. Based on the differences in the peptides, a voting method was used to statistically analyze the up- and down-regulation of the corresponding proteins. Small proteins with an amino acid sequence length of less than 100 aa were selected as identified SEPs.
[0040] In one embodiment of the present invention, step 3) further includes the following steps:
[0041] 3.2) De novo sequencing reanalysis: Based on the mass spectrometry data obtained from peptide group sequencing, de novo sequencing was used to reanalyze the mass spectrometry data to detect novel SEPs that may not be included in the database. The sequencing results and the novel SEP screening process in de novo sequencing are as follows:
[0042] 3.2.1) De novo sequencing main parameter settings:
[0043] PEAKS Studio peptide group sequencing mass spectrometry data were reanalyzed. Peptides were sequenced using de novo sequencing based on mass spectrometry molecular weight, and identified peptides were searched using PEAKS DB to determine peptide-protein attribution. The de novo sequencing results were combined with database search results to provide additional information on novel peptides, such as post-translational modifications and mutations. Peptides that could not be found in the database were considered de novo only peptides. The database was set to the transcriptome six-frame translation library generated by S187 transcriptome sequencing. The parameter was set to None enzyme, meaning no enzyme digestion. The search parameters included a fragment ion mass tolerance of 0.02 Da and a precursor mass tolerance of 7 ppm. Three variable modifications were allowed: Oxidation (M) 15.99, Acetylation (Protein N-term) 42.01, and Deamidation (NQ) 0.98. All peptides were quality-controlled and filtered with a -10 logP ≥ 20.
[0044] 3.2.2) De novo peptide analysis:
[0045] Three methods were used to de-redundant and genomic localize de novo peptides. First, ORFFinder was used to perform six-frame translation of the nucleotide sequences of all transcripts obtained from the whole genome and transcriptome sequencing of S187. All de novo peptides were mapped to two protein libraries using the peptide mapper software. Peptides corresponding to ORFs less than 100 aa were screened from the peptides mapped to the database and then submitted to the UniProt Peptide search tool for search. The search parameters were set as follows: the search species was limited to Actinobacteria, and leucine and isoleucine were considered equivalents, thereby filtering out peptides present in proteins reported in the UniProt database. The genomic locations of the remaining novel peptides were further determined by the ORF IDs corresponding to the peptides. The coordinates of the ORF positions and the distances between the peptides and the start and stop codons of the corresponding ORFs were calculated. SEPs were screened and peptides with an average local confidence ≥80% and no post-modification were identified as high-confidence novel peptides.
[0046] In one embodiment of the present invention, another novel peptide analysis is used in step 3.2.2). First, the actinomycete sequences in the Nr database are aligned by BLASTp, with the task set to blastp-short, the scoring matrix set to PAM30, and the E-value set to 10, thereby removing the redundant parts of the de novo peptides; subsequently, tBLASTn is used to align with the S187 genome, the PAM30 scoring matrix is set, and the best match is found; peptides with a similarity and coverage of more than 80% and an E-value ≤ 1 are screened in the match, and then a python script is used to perform ORF batch search based on the location of the peptide in the genome, and SEPs with a length of less than 100 amino acids are screened from the start codons ATG, GTG, TTG, CTG to the stop codons TGA, TAG, TAA; the online tool BLASTp and the NCBI Nr database are used to filter the homologous sequences annotated in the database; finally, peptide sequences with a length of ≥7 aa and no more than 2 amino acids in the genome alignment mismatch are selected as novel SEPs.
[0047] In one embodiment of the present invention, step 3) further includes the following steps:
[0048] 3.3) ncRNA prediction: The clean reads after quality control of S187 transcriptome sequencing were mapped back to the genome, and transcriptome assembly was performed based on the reference genome annotation file using Cufflink or Stringtie. The assembled transcripts were then compared with the genome annotation file using Cuffcompare. New transcripts were classified according to their genomic location information, and new transcripts and ncRNAs were predicted.
[0049] In one embodiment of the present invention, step 3) further includes the following steps:
[0050] 3.4) Prediction of shared SEPs between Streptomyces S187 and the model strain:
[0051] 3.4.1) Transcriptome data analysis of model strains
[0052] The transcriptome data of the model strain of Streptomyces coelicolor from the SRA public database were selected and processed using the following analysis method:
[0053] Quality control: First, use the software FastQC and Trimmomatic to perform data quality control on the raw data reads. Those that pass the test will be used as clean data for subsequent analysis.
[0054] Mapping clean data back to the genome: Transcriptome analysis was performed using two transcriptome analysis pipelines, Bowtie2 and STAR software for read mapping. First, index files for the two software were constructed based on the fasta file of the Streptomyces coelicolor reference genome. Then, clean reads in the fastq file were mapped back to the reference genome using Bowtie2 and STAR software based on the index files. The resulting sam files were converted and sorted using samtools to generate binary bam files and sorted.bam files sorted according to the genomic sequence. The sorted bam files were opened using the genome browser IGV to visualize the distribution of sequencing reads, genomes, and annotation information.
[0055] Transcript quantification and smORF screening: Reads corresponding to Bowtie2 and STAR were mapped back to the genome, and transcript quantification was performed using the featureCounts tool and RSEM software, respectively, based on the S. coelicolor genome annotation file. Read counts for corresponding gene fragments in the clean data were obtained, and the counts were normalized for sequencing depth and then for gene length to obtain FPKM (the number of reads per kilobase of gene length per million reads) to estimate gene expression levels. An FPKM greater than 1 is generally considered a necessary condition for stable transcription of a gene. Transcripts with a length of 303 nt or less and an FPKM greater than 1 in seven transcriptome data sets were screened as smORFs in S. coelicolor.
[0056] 3.4.2) Prediction of shared SEPs and S187-specific SEPs:
[0057] The Transeq tool in EMBOSS was used to translate the smORFs in the Streptomyces coelicolor transcriptome into amino acid sequences to generate SEP amino acid sequences. These sequences were then used as queries and compared with the amino acid sequences of the entire S187 genome smORF library constructed using genomic smORF prediction using BLASTp. Simultaneously, tBLASTn was used to align the sequences with the nucleotide sequence of the entire S187 genome, with an E-value of 1e-5. Successfully matched SEPs were identified as shared SEPs between the two strains and submitted to the CD-search online tool for conserved domain search and functional site prediction. SEPs possessing both characteristics were considered small proteins that were more likely to exist and function in subsequent analyses.
[0058] The SEPs shared by S. coelicolor and S187 were compared with 126 SEPs identified by DB search using BLASTp. The parameters were set to Evalue 1e-5 and output format outfmt 6. Small proteins that were not involved in the match were regarded as SEPs unique to S187.
[0059] In summary, the present invention provides a method for predicting and identifying small proteins in marine Streptomyces S187, and the beneficial effects of the present invention are:
[0060] Based on the prediction of smORFs from genomic sequence information, the present invention identified a comprehensive genomic smORF library based on complete ORFs (open reading frames), laying the foundation for subsequent analysis. Collinearity also suggests a close evolutionary relationship between S187 and the model strain. Therefore, the relatively mature research on the model strain of Streptomyces coelicolor provides important insights into S187 and smORF discovery.
[0061] Furthermore, the present invention fragments the six-frame translation theoretical enzyme digest sequence library or generates peptides, then matches the theoretical secondary spectra of the peptides with the actual spectra, and constructs a spectral library based on the secondary spectra and retention time of the peptides identified by the DDA data. Subsequently, DIA mass spectrometry is performed on the machine for detection and spectrum analysis. The spectral library generated by the DDA data is used to obtain the fragment ion characteristics corresponding to each peptide in the library. Targeted information extraction and peptide spectrum matching are performed on the DIA data with the peptide as the center. Based on the secondary spectrum matching, the spectral library is searched with the spectrum as the center, thereby completing the peptide identification of the DIA secondary spectrum, improving the sensitivity and accuracy of the identification, and faster search speed.
[0062] Furthermore, peptides identified by database search were spliced together, and statistically analyzed SEPs showed an average peptide coverage of 39.77%, with four full-length proteins identified. This suggests that the top-down, non-enzymatic small protein identification strategy may facilitate the detection of longer peptides and improve protein-specific peptide coverage in peptide sequencing, providing more reliable peptide localization and quantification information for subsequent analysis.
[0063] Furthermore, relying on database searches to identify peptides when gene annotation levels are limited still has certain limitations. However, novel peptides identified through de novo sequencing provide translational evidence to supplement gene annotation information, complementing the limitations of database searches. Using transcriptome data from model strains in public databases offers practical insights for identifying small proteins in non-model strains. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1Schematic diagram of the process for predicting and identifying small proteins in Streptomyces marineus S187.
[0065] Figure 2 A comparison chart of smORFs prediction results for the two genomes.
[0066] Figure 3 This is a diagram of the colinearity between S187 and the whole genome of Streptomyces coelicolor.
[0067] Figure 4 Functional classification diagram of 126 SEPs identified by DB search.
[0068] Figure 5 This is the coverage statistics of peptides identified by DB search after splicing; A: Peptide coverage statistics corresponding to 126 SEPs; B: Schematic diagram of peptide splicing results of four full-length proteins.
[0069] Figure 6 Figure 2 shows the de novo peptide analysis process and results.
[0070] Figure 7 The diagram shows the types of novel SEPs and the classification of new transcripts.
[0071] Figure 8 Diagram of the transcriptome analysis process and results of Streptomyces coelicolor.
[0072] Figure 9 This is a map of the shared SEPs between S187 and Streptomyces coelicolor in the DB search results. DETAILED DESCRIPTION
[0073] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention for which protection is sought, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0074] like Figure 1 As shown, the present invention aims to provide an accurate, comprehensive and highly specific identification process for studying micropeptides in non-model bacteria. The present invention is mainly divided into three modules:
[0075] (1) Establishment of a comprehensive database: A database including genome prediction and transcriptome detection was customized for the non-model strain S187 to improve the sensitivity of SEP identification.
[0076] (2) Sample processing: Peptides from samples at different metabolic time points are extracted and enriched. The enriched micropeptides are then directly subjected to high-precision DIA mass spectrometry detection instead of conventional proteomics enzyme digestion.
[0077] (3) Analysis of SEP identification results: The qualitative and quantitative results of the sample micropeptides were obtained by mass spectrometry data analysis, and the mass spectrometry data were used simultaneously for database search and database-independent de novo sequencing (de novo sequencing) identification methods to find novel peptides and ncRNAs (non-coding RNAs). In addition, by combining the transcriptome data of the model strain Streptomyces coelicolor A3 (2) in the public database, the functional micropeptides that have been overlooked were predicted and mined.
[0078] The present invention specifically adopts the following technical solutions:
[0079] 1 Database establishment
[0080] 1.1 Whole-genome prediction
[0081] Two software programs were used to predict ORFs in the entire S187 genome. First, the S187 whole-genome fasta sequence was submitted to ORFFinder (http: / / www.bioinformatics.org / sms2 / orf_find.html) for six-frame translation, using the four start codons ATG, GTG, CTG, and TTG, and a representative bacterial codon 11 translation table. Furthermore, Prodigal software was used to perform genome-wide ORF prediction as a supplement to ORFFinder, thereby obtaining the corresponding ribosome binding motifs. The S187 whole-genome coding sequence was obtained using the representative bacterial codon 11 translation table. The amino acid sequence lengths were statistically analyzed, and ORFs under 100 aa were selected as smORFs.
[0082] 1.2 Collinearity Analysis
[0083] The present invention also selected the model strain of Streptomyces coelicolor A3 (2) as a reference and used the MCScanX tool in TBtools software to perform genomic collinearity analysis with S187 to explore the similarity between the two strains. The operation steps are as follows: open TBtools software, select the One StepMCScanX tool under the Graphic item, select the genome and gff annotation files of the two strains and the path of the output folder, select 8 for CPU, 1e-10 for Evalue, and 5 for Blasthit number. Start the run, and the collinearity process and result files are generated in the output folder.
[0084] 1.3 Culture conditions and sampling methods
[0085] The preserved strain of Streptomyces marineus S187 was streaked and activated onto TSB solid medium and placed in a constant temperature incubator at 28°C. The activated strain was used to prepare the fermentation seed solution, as follows.
[0086] Seed liquid collection: Take 3-4 S187 bacterial blocks activated by TSB solid medium, add them to TSB liquid medium, culture at 220 rpm, and collect samples of fermentation liquid after 36 hours;
[0087] Fermentation broth collection: Add 10 mL of seed solution to M33 medium and culture at 180 rpm. Increase the speed by 20 rpm every 48 h until it reaches 220 rpm and then maintain the speed unchanged.
[0088] Sampling: First, the fermentation curve and anti-complement activity curve of S187 were used to determine the activity cycle of the strain. The sampling time points were determined based on the fermentation curve and anti-complement activity curve: 36h, 48h, 72h, and 120h. The sampling, elution, and storage steps are as follows:
[0089] 1) From the same batch of six bottles of fermentation broth, sample three bottles of fermentation broth that have grown well. Add 3 mL of each sample to a 50 mL centrifuge tube and mix thoroughly. From the remaining three bottles of fermentation broth, transfer 9 mL of one sample to a 50 mL centrifuge tube. This is then processed as the first bottle to balance the sample. Transfer 10 mL of each sample to a centrifuge tube for extraction and rotary evaporation. Keep the centrifuge tube in an ice bath throughout the entire experiment.
[0090] 2) Removal of proteins from the culture medium: Centrifuge at 5000 rpm for 6 min and discard the supernatant.
[0091] 3) PBS elution: Add a small amount of pre-cooled PBS solution (phosphate buffered saline) to the centrifuged pellet and mix by pipetting 40 times with a pipette. Then add PBS to 35 mL, vortex, centrifuge at 5000 g for 5 min, and discard the supernatant.
[0092] 4) Repeat step 3) once.
[0093] 5) Removal of CaCO3: Add a small amount of pre-cooled PBS solution and mix by pipetting 40 times with a pipette. Then add PBS to 35 mL, vortex, and centrifuge at 500 rpm for 3 min. Discard the precipitate and transfer the sample to a new 50 mL centrifuge tube.
[0094] 6) Repeat step 3) once more for the final elution.
[0095] 7) Place the centrifuge tube in a liquid nitrogen tank and freeze it for 10-20 minutes, then store it in a -80°C refrigerator.
[0096] 1.4 Transcriptome sequencing library construction:
[0097] The samples from the four time points were combined for RNA-seq (transcriptome sequencing technology) library construction. The specific steps are as follows: After RNA extraction, purification, and library construction, the samples were sequenced using NGS sequencing technology (Next-Generation Sequencing, next-generation sequencing), i.e., second-generation sequencing, based on the Illumina HiSeq sequencing platform, and PAIRED double-end sequencing was performed on these libraries. All transcripts obtained by transcriptome sequencing were searched for open reading frames using NCBI ORFFinder, with the minimum ORF length set to 30, the Genetic code set to 11.Bacterial, and the ORF start codon selected as any start codon. The obtained six-frame translation data of the transcriptome was used as a database and searched in the peptide group sequencing process.
[0098] Transcriptome data smRNAs prediction: FastaQC and Trimmomatic were used for quality control of transcriptome sequencing data, Bowtie2 and featureCounts were used for read alignment and quantification, and RNA with FPKM>1 and length≤303nt was screened as a candidate smRNAs library.
[0099] 2. Sample processing and mass spectrometry detection of micropeptides
[0100] The present invention sequenced the polypeptide group samples of S187 fermentation time of 48h and 120h. The strain culture, sample processing and main sequencing processes are as follows.
[0101] 2.1 Enrichment and processing of micropeptides
[0102] Since small proteins are often extracted from peptides, S187 mycelial samples, obtained at different fermentation times, were crushed (grinded in liquid nitrogen) and then small peptides were extracted using a methanol:chloroform:water ratio of 3:1:4 (v:v). The mixture was briefly vortexed and centrifuged at 4700 rpm for 10 minutes at 4°C. The supernatant was collected and centrifuged at 20,000 × g through an ultrafiltration membrane with a molecular weight cutoff of 10 kD. The lower layer was removed, concentrated in vacuo, and dried.
[0103] Desalting and lyophilization: Desalt the sample using a C18 desalting column, activate the desalting column with 100% ACN, equilibrate the column with 0.1% FA, load the sample onto the column, then wash the column with 0.1% FA to remove impurities, and finally elute with 70% ACN, collect the flow-through, and lyophilize.
[0104] HPLC fractionation: A portion of the peptides from each sample was mixed and then separated into multiple fractions by high performance liquid chromatography. Prepare mobile phase A (2% acetonitrile, 98% water, ammonia adjusted to pH = 10) and B (98% acetonitrile, 2% water, ammonia adjusted to pH = 10). Use A to dissolve the mixed freeze-dried powder and centrifuge at 12000g for 10 minutes at room temperature. Use L-3000HPLC system, the chromatographic column is Waters BEH C18 (4.6×250mm, 5μm), the column temperature is set to 50°C, and the specific elution gradient is shown in Table 1. Collect 1 tube per minute and merge into 3 fractions at intervals. The chromatographic gradient is as follows:
[0105] Table 1 HPLC chromatographic gradient
[0106]
[0107] 2.2 Peptide sequencing
[0108] DDA mass spectrometer:
[0109] The mobile phase was prepared as follows: Phase A included 100% mass spectrometry water and 0.1% formic acid. Formic acid was added to the mass spectrometry water to form a 0.1% by volume formic acid aqueous solution. Phase B included 80% acetonitrile and 0.1% formic acid. Formic acid was added to the 80% acetonitrile aqueous solution to form a 0.1% by volume formic acid acetonitrile aqueous solution.
[0110] The lyophilized powder was dissolved in 10 μL of phase A and centrifuged at 14,000 g for 20 min at 4°C. 1 μg of the supernatant was injected and detected by liquid chromatography-mass spectrometry. The elution conditions are shown in Table 2. The Orbitrap Exploris 480 mass spectrometer and NanosprayFlex TMThe (NSI) ion source was set to 1.8 kV, the ion spray voltage was set to 320°C, the ion transfer tube temperature was set to 320°C, the DDA mode was used, the mass spectrometry full scan range was m / z 350-1200, the primary mass spectrometry resolution was set to 60,000 (200 m / z), the AGC was set to Custom, and the maximum injection time was 50 ms; the parent ion was fragmented using the high-energy collision fragmentation (HCD) method, and secondary mass spectrometry detection was performed. The secondary mass spectrometry resolution was set to 15,000 (200 m / z), the AGC was set to Custom, the maximum injection time was 22 ms, and the peptide fragmentation collision energy was set to 30% to generate mass spectrometry detection raw data (.raw). The separation flow rate was 600 nl / min, and the separation gradient was as follows:
[0111] Table 2 HPLC chromatographic gradient
[0112]
[0113] DIA mass spectrometer and database search: Dissolve the peptide powder in 0.1% formic acid aqueous solution and add iRT standard peptides, take 1 μg sample and inject it for liquid chromatography-mass spectrometry detection. Use Orbitrap Exploris 480 mass spectrometer, NanosprayFlex TM The ESI ion source was set at 1.8 kV, the ion spray voltage was set at 320°C, the ion transfer tube temperature was set at 320°C, the mass spectrometer adopted data-independent acquisition (DIA) mode, the full mass spectrometer scan range was m / z 350-1500, the primary mass spectrometer resolution was set to 120,000, the C-trap maximum capacity was set to Custom, and the C-trap maximum injection time was set to 50 ms; the peptide fragmentation collision energy was set to 33%.
[0114] 3. Identification of peptide group SEPs
[0115] 3.1 Database search and identification of SEPs
[0116] The mass spectrometry analysis of DIA was processed using Spectronaut 15 software. The search parameters were set as follows:
[0117] Table 3 Spectronaut 15 parameter settings
[0118]
[0119] The quantitative values of the sample peptides were analyzed using the t-test, with a P value ≤ 0.05 and a Foldchange ≥ 1.5 to obtain differential peptides. The voting method was then used to statistically analyze the up- and down-regulation of the corresponding proteins based on the differences in the peptides, and small proteins with an amino acid sequence length of less than 100 aa were selected as the identified SEPs.
[0120] 3.2 De novo sequencing reanalysis:
[0121] In addition to using the transcriptome six-frame translation database for database search, the mass spectrometry data obtained by peptide group sequencing in the present invention were also re-analyzed using de novo sequencing to detect novel SEPs that may not be included in the database. The sequencing results and the novel SEP screening method in de novo sequencing are as follows.
[0122] 3.2.1 Main parameter settings for De novo sequencing:
[0123] Peptide sequencing mass spectrometry data were reanalyzed using PEAKS Studio (version 10.6). Peptides were identified directly by mass spectrometry molecular weight using denovo sequencing. Identified peptides were then searched against the PEAKS database to confirm peptide-protein affiliation. De novo sequencing results were combined with database search results to provide additional information on novel peptides, such as post-translational modifications and mutations. Peptides not found in the database were considered denovo-only peptides. The database was a six-frame translation library generated by S187 transcriptome sequencing. The enzyme-selected parameters were None, meaning no enzyme digestion was performed. The search parameters included a fragment ion mass tolerance of 0.02 Da and a precursor mass tolerance of 7 ppm. Three variable modifications were allowed: oxidation (M) 15.99, acetylation (protein N-term) 42.01, and deamidation (NQ) 0.98. All peptides were quality-controlled and filtered for -10 logP ≥ 20.
[0124] 3.2.2 De novo peptide analysis method:
[0125] Three methods were used to de-redundant and genomic map de novo peptides. The analysis process is shown in the figure. First, ORFFinder was used to perform a six-frame translation of the nucleotide sequences of all transcripts obtained from the S187 whole genome and S187 transcriptome sequencing to construct a library. Using the peptide mapper software PeptideMapper (v5.0.44), all de novo peptides were mapped to two protein libraries. Peptides corresponding to ORFs under 100 aa were selected from the peptides mapped to the database and then submitted to the UniProt Peptide search tool for a search. The parameters were set as follows: the search species was limited to Actinobacteria (taxonomy: 201174), and leucine and isoleucine were considered equivalents, thereby filtering out peptides present in proteins reported in the UniProt database. The genomic locations of the remaining novel peptides were further determined using the corresponding ORF IDs of the peptides. The ORF coordinates of the peptides and the distances between the peptides and the start and stop codons of the corresponding ORFs were calculated. SEPs were screened and identified. Peptides with an average local confidence (ALC) ≥ 80% and no post-modification were considered high-confidence novel peptides.
[0126] To identify a more comprehensive set of novel peptides, an alternative novel peptide analysis strategy was implemented. First, Actinobacteria sequences in the Nr database were aligned using BLASTp, using the blastp-short task, the PAM30 scoring matrix, and an E-value of 10 to remove redundant de novo peptides. Subsequently, tBLASTn was used to align with the S187 genome, using the PAM30 scoring matrix and searching for the best match. Peptides with a similarity and coverage greater than 80% and an E-value ≤ 1 were selected. Subsequently, a Python script was used to perform a batch search for ORFs based on the peptide's genomic location. Sequences with lengths less than 100 amino acids were selected, starting with the start codons ATG, GTG, TTG, and CTG and ending with the stop codons TGA, TAG, and TAA. The online BLASTp tool and the NCBI Nr database were used to filter out annotated homologous sequences in the database. Mass spectra with few impurity peaks and three pairs of b / y ions were generally considered high-quality, and peptides with high-quality spectra were selected. Finally, peptide sequences with a length of ≥7 aa and no more than 2 amino acid mismatches in genomic alignment were selected as Novel SEPs.
[0127] 3.3ncRNA Prediction
[0128] The clean reads after S187 transcriptome sequencing quality control were mapped back to the genome, and transcriptome assembly was performed based on the reference genome annotation file using Cufflink (version 2.2.1) or Stringtie (version 2.2.1). The assembled transcripts were compared with the genome annotation file using Cuffcompare (version 2.2.1), and new transcripts were classified according to genomic location information. New transcripts and ncRNAs were predicted. The database search results and de novo sequencing results were combined to predict the new transcripts.
[0129] 3.4 Prediction of shared SEPs between Streptomyces S187 and model strains:
[0130] 3.4.1 Transcriptome data analysis of model strains
[0131] The transcriptome data of the model strain of Streptomyces coelicolor from the SRA public database were selected, and the main analysis methods used are as follows.
[0132] Quality control: First, use the software FastQC and Trimmomatic to perform data quality control on the reads of the raw data. Those that pass the test will be used as clean data for subsequent analysis.
[0133] Mapping clean data back to the genome: To ensure the accuracy of the results, transcriptome analysis was performed using two transcriptome analysis processes, using Bowtie2 and STAR software for read mapping. First, it was necessary to construct index files for the two software based on the fasta file of the Streptomyces coelicolor reference genome. Then, based on the index files, Bowtie2 and STAR software were used to map the clean reads in the fastq file back to the reference genome. The resulting sam file was converted and sorted using samtools to generate a binary bam file and sorted.bam sorted according to the genomic sequence. The sorted bam file was opened using the genome browser IGV to visualize the distribution of sequencing reads, genomes, and annotation information.
[0134] Transcript quantification and smORF screening: Reads corresponding to Bowtie2 and STAR were mapped back to the genome, and transcript quantification was performed using the featureCounts tool and RSEM software, respectively, based on the S. coelicolor genome annotation file. Read counts for the corresponding gene fragments in the clean data were obtained. The counts were normalized for sequencing depth and then for gene length to obtain FPKM (The Fragments Per Kb Of CDS Per Million Mapped Reads), which is the number of reads per kilobase of gene length per million reads, to estimate gene expression levels. An FPKM greater than 1 is generally considered a prerequisite for stable transcription of a gene. Transcripts with a length of 303 nt or less and an FPKM greater than 1 in the seven transcriptome data sets were screened as potential smORFs in S. coelicolor.
[0135] 3.4.2 Prediction of shared SEPs and S187-specific SEPs
[0136] The Transeq tool in EMBOSS was used to translate the smORFs in the Streptomyces coelicolor transcriptome into amino acid sequences, generating SEP amino acid sequences. These sequences were then used as queries and aligned with the amino acid sequence of the entire S187 genome smORF library constructed using genomic smORF prediction using BLASTp. Furthermore, tBLASTn was used to align the amino acid sequence with the entire S187 genome nucleotide sequence, with an E-value of 1e-5. Successfully matched SEPs, identified as shared SEPs between the two strains, were submitted to the CD-search (https: / / www.ncbi.nlm.nih.gov / cdd / ) online tool for conserved domain search and functional site prediction. SEPs possessing both characteristics were considered for subsequent analysis as small proteins with a higher likelihood of actual existence and function.
[0137] The SEPs shared by S. coelicolor and S187 were compared with 126 SEPs identified by DB search using BLASTp. The parameters were set to Evalue 1e-5 and output format outfmt 6. Small proteins that were not involved in the match were regarded as SEPs unique to S187.
[0138] 4. Results
[0139] 4.1 Genomic smORFs prediction:
[0140] In terms of genomics, the present invention established a comprehensive S187 smORFs database. First, the present invention selected all possible ORFs from the translation database of the six coding frames of the S187 genome and screened for open reading frames with an amino acid length of ≤100aa. Figure 2 As shown, 46,629 potential smORFs were obtained as the complete genome library of the present invention. This method allows for all possible coding modes, reducing the number of missed smORF predictions. Secondly, Prodigal software, a mainstream tool for bacterial genome prediction, predicted and found 552 coding ORFs. Ribosome binding motifs corresponding to these smORFs were also obtained, with confidence levels above 0.5 for all, and 90% of smORFs with confidence levels above 0.8.
[0141] The present invention conducted a gene-level collinearity analysis on S187 and Streptomyces coelicolor based on all protein coding sequences and obtained 68 collinear blocks, such as Figure 3 The two strains' genomes contained a total of 13,792 coding genes, of which 5,671 genes, or 41.12%, exhibited collinearity. The E-values of all matching genes were less than 1e-5. Among the collinear gene pairs, 152 were less than 303 nt in length.
[0142] Based on the prediction of smORFs from genomic sequence information, the present invention identified a comprehensive genomic smORF library based on complete ORFs, laying the foundation for subsequent analysis. Collinearity also suggests a close evolutionary relationship between S187 and the model strain. Therefore, the relatively mature research on the model strain of Streptomyces coelicolor provides important insights into S187 and smORF discovery.
[0143] 4.2 Database search for SEPs identification
[0144] To detect peptides produced during secondary metabolism in S187, particularly during the production of anti-complementary substances, we first selected fermentation time points. Mycelial samples were collected at 36, 48, 72, and 120 hours. Transcriptome sequencing was performed on these mixed samples, identifying 561 transcribed smRNAs. Furthermore, all transcripts obtained from transcriptome sequencing were translated into amino acid sequences using six-frame translation to construct a library for peptide sequencing.
[0145] For peptide sequencing, samples aged 48 and 120 hours were selected. Peptide extraction and ultrafiltration through a 10 kDa filter were performed to enrich small peptides under 100 aa. Liquid chromatography separation into three fractions was then performed before DDA mass spectrometry. A six-frame translation sequence library was theoretically digested or fragmented to generate peptides. The theoretical secondary spectra of the peptides were then matched to the actual spectra. A spectral library was constructed based on the secondary spectra and retention times of the peptides identified by the DDA data. Subsequently, DIA mass spectrometry was performed on the machine for on-chip detection and spectrum analysis. Using the spectral library generated from the DDA data, the fragment ion signature corresponding to each peptide in the library was obtained. Targeted information extraction and peptide-spectrum matching were performed on the DIA data, with a peptide-centric approach. Based on the secondary spectrum matching, a spectrum-centric library search was performed to complete peptide identification from the DIA secondary spectrum. Although this library search method is limited to the identification results from DDA data and can easily lose the advantage of DIA in detecting low-abundance peptides, it can improve the sensitivity and accuracy of identification and increase the search speed.
[0146] A total of 24,155 peptides were identified through peptide group search, and the identified peptides can be mapped to 1,952 proteins in the database, including 126 small proteins. GO, KEGG and InterProScan were used to annotate the functions of 72 proteins, which are mainly divided into six categories: biosynthesis process, nucleic acid binding related, pressure stress, ribosome related, bacterial secretion system and membrane related proteins, such as Figure 4 Functional enrichment analysis showed that the identified SEPs were mainly related to ribosomes and translation processes, while protein interactions revealed that small proteins can interact with each other to regulate gene expression, participate in transmembrane transport of substances, and play a certain role in bacterial growth, development, and metabolism.
[0147] In addition, the peptides identified by database search were spliced, and the average peptide coverage in SEPs was 39.77%, and four full-length proteins were identified, such as Figure 5 This demonstrates that a top-down small protein identification strategy without enzyme cleavage may be helpful in detecting polypeptides with longer amino acid lengths and improving the coverage of protein-corresponding polypeptides in the peptide sequencing of the present invention, thereby providing more reliable peptide segment positioning and quantitative information for subsequent analysis.
[0148] 4.3 De novo sequencing SEPs identification
[0149] This study used de novo sequencing-based mass spectrometry analysis to infer peptide sequence information solely from mass spectrometry data, identifying peptides present during the S187 fermentation process. Peptide spectrum matching revealed that de novo sequencing identified 101,216 peptides that were not identified in the database search, approximately four times the 24,155 peptides identified in the database search. This demonstrates the comprehensive peptide spectrum matching capabilities of de novo sequencing in mass spectrometry analysis.
[0150] In subsequent analysis of de novo only peptides, such as Figure 6 As shown, the present invention proposes three methods to locate the identified peptides on the genome through database-dependent and database-independent methods, and filters out the peptides existing in the actinomycete database. Subsequently, smORFs screening is performed, and 15 novel SEPs with functions not included in public databases are predicted. According to their genomic location, they are divided into cryptic SEPs (CSEPs) that are in a different frame from the canonical protein and isoform SEPs (ISEPs) that are in the same frame as the canonical protein but use different start codons. Seven ISEP-corresponding peptides also appeared in peptides identified by the database search, and two of these ISEPs appeared in small proteins identified by the database search, confirming the reliability of de novo sequencing for mass spectrometry analysis. Furthermore, the fact that five ISEPs were not identified by the database search suggests that database searches can miss smORFs nested in the same frame as larger ORFs. Three ISEPs and five CSEPs were not included in the database used by the DB search, indicating that the transcriptome six-frame translation library may still be insufficient due to the transient nature or low abundance of smORF transcripts. Therefore, relying on database searches to identify peptides when gene annotation levels are limited still has certain limitations. The new peptides identified by de novo sequencing provide translational evidence to supplement gene annotation information, complementing the limitations of the DB search.
[0151] Figure 7To identify the types and classification of novel SEPs, the present invention assembled 32 ncRNAs from the S187 transcriptome data and mined a small protein encoded by an ncRNA. However, no ncRNA corresponding to novel SEPs was found in the data, indicating that the search for novel SEPs based solely on transcriptome data still has certain limitations. In addition, the abundance of CSEPs among novel SEPs in the transcriptome is low, and whether they are transcribed and subsequently translated still needs further experimental verification.
[0152] 4.4 Prediction of shared SEPs between S187 and model strains
[0153] The present invention selected 7 sets of transcriptome data of Streptomyces model strains from the SRA database and performed transcriptome analysis. Figure 8 The transcriptome analysis pipeline and results for Streptomyces coelicolor are presented. To improve the accuracy of our results and identify a more comprehensive set of SEPs, we used two complementary analysis pipelines. After data quality control and cleaning, tens of millions of high-throughput sequencing reads were mapped to the genome. Transcript quantification identified 464 smORFs (smORFs) within protein-coding genes with lengths ≤303 nt and FPKMs >1. To investigate the conservation of these smORFs across the Streptomyces genus, we aligned the S. coelicolor transcriptome smORF library with the S187 genome, identifying 248 shared SEPs. Among them, the use of Bowtie2 to quantify featureCounts obtained more smORFs, reaching 455, which was 96 more smORFs than the STAR to RSEM quantification. In addition, the Bowtie2 process alone shared 27 SEPs in the shared SEPs results with S187, which proved that this process has more comprehensive identification performance in the prediction of smORFs. At the same time, in the mining of shared SEPs, tBLASTn also has more sensitive recognition performance than BLASTp.
[0154] It is generally believed that functional SEPs are somewhat conserved. Through prediction and analysis of conserved domains and functional sites, the present invention identified 27 SEPs containing both conserved domains and functional sites, corresponding to the 17 SEPs in the S187 polypeptide group that are authentically translated. These 17 SEPs were categorized based on BLAST annotation functions, primarily into ribosomal subunits, stress-related cold shock proteins, and small proteins related to nucleic acid binding and biosynthesis, such as transport proteins associated with the phosphotransferase system.
[0155] The S187smORFs data predicted from the transcriptome of the model strain Streptomyces coelicolor were evaluated using the peptide group DBsearch identification data. Figure 9 As shown in the results of the peptide genomics database search, 79 SEPs were detected in the S. coelicolor transcriptome, accounting for 62.7% of the SEPs in the peptide database search results. Furthermore, 17 of these SEPs were shared by both strains and possessed complete structural domains and functional sites, demonstrating the practical value of using transcriptome data from model strains in public databases for identifying small proteins in non-model strains.
[0156] 4.5 Prediction of SEPs unique to S187
[0157] In addition to the shared SEPs between the two strains, S187 also possesses 47 unique SEPs not found in the transcriptome data for Streptomyces coelicolor, potentially related to its natural marine environment. Thirteen unique small proteins have been annotated. Iron and sulfur are often limiting nutrients in marine environments. Iron-sulfur cluster-binding proteins facilitate bacterial survival under low iron and sulfur conditions. Cytochrome P450s are generally associated with monooxygenases, participating in biosynthesis and metabolism and adapting to various environmental changes. Secretion of these small proteins may facilitate bacterial interactions with the environment. Stress-related proteins, such as those in the toxin-antitoxin system, maintain bacterial stability and viability. Furthermore, proteins associated with nucleic acid binding may be involved in regulating gene transcription and expression, as well as DNA repair. In summary, these unique small proteins play a crucial role in the survival and stability of S187 in marine environments, facilitating adaptive growth, development, and metabolic processes under various stresses and environmental pressures.
[0158] Table 4 Unique SEPs with annotated functions in S187
[0159]
[0160]
[0161] The above is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for predicting and identifying small proteins in Streptomyces marineus S187, characterized in that: It includes the following steps: 1) Establishment of a comprehensive database: Customize a database for non-model strains that includes genome prediction and transcriptome detection; 2) Sample processing: Peptides from samples at different metabolic time points are extracted and enriched, and the enriched micropeptides are then detected by high-precision DIA mass spectrometry; 3) SEP Identification Results Analysis: Qualitative and quantitative results of sample micropeptides are obtained through mass spectrometry data analysis. The mass spectrometry data are then used simultaneously for database search and database-independent denovo sequencing to identify novel peptides and ncRNAs. The following steps are involved: 3.1) Database search and identification of SEPs: The quantitative values of the sample peptides were analyzed using the t-test, with a chi-square P value ≤ 0.05 and a fold change ≥ 1.5 to obtain differential peptides. Based on the differences in the peptides, the voting method was used to statistically analyze the up- and down-regulation of the corresponding proteins. Small proteins with an amino acid sequence length of less than 100 aa were selected as identified SEPs. 3.2) De novo sequencing reanalysis: Based on the mass spectrometry data obtained from peptide group sequencing, de novo sequencing was used to reanalyze the mass spectrometry data to detect novel SEPs that may not be included in the database. The sequencing results and the novel SEP screening process in de novo sequencing are as follows: 3.2.1) De novo sequencing main parameter settings: PEAKS Studio peptide sequencing mass spectrometry data were reanalyzed. Peptides were sequenced using de novo sequencing based on mass spectrometry molecular weight, and identified peptides were searched using PEAKSDB to determine peptide-protein attribution. The de novo sequencing results were combined with database search results to provide additional information on novel peptides, such as post-translational modifications and mutations. Peptides not found in the database were considered de novo only peptides. The database was set to the transcriptome six-frame translation library generated by S187 transcriptome sequencing. The parameter was None enzyme, meaning no enzyme digestion. The search parameters included a fragment ion mass tolerance of 0.02 Da and a precursor mass tolerance of 7 ppm. Three variable modifications were allowed: Oxidation (M) 15.99, Acetylation (Protein N-term) 42.01, and Deamidation (NQ) 0.
98. All peptides were quality-controlled and filtered with a -10 logP ≥ 20. 3.2.2) De novo peptide analysis: Three methods were used to remove redundancy and locate de novo peptides in the genome. First, ORFFinder was used to perform six-frame translation of the nucleotide sequences of all transcripts obtained from the whole genome and transcriptome sequencing of S187. All de novo peptides were mapped to two protein libraries using the peptide mapper software. Peptides corresponding to ORFs less than 100 aa were screened from the peptides mapped to the database and then submitted to the UniProt Peptide search tool for search. The parameters were set as follows: the search species was limited to Actinobacteria, and leucine and isoleucine were considered equivalents, thereby filtering out peptides present in proteins reported in the UniProt database. The genomic locations of the remaining novel peptides were further determined by the ORFID corresponding to the peptides, and the coordinates of the ORF position and the distance between the peptides and the start and stop codons of the corresponding ORFs were calculated. SEPs were screened and peptides with an average local confidence ≥80% and no post-modification were considered as high-confidence novel peptides. 3.3) ncRNA prediction: The clean reads after S187 transcriptome sequencing quality control were mapped back to the genome, and transcriptome assembly was performed based on the reference genome annotation file using Cufflink or Stringtie. The assembled transcripts were then compared with the genome annotation file using Cuffcompare. New transcripts were classified according to their genomic location information, and new transcripts and ncRNAs were predicted. 3.4) Prediction of shared SEPs between Streptomyces S187 and the model strain: 3.4.1) Transcriptome data analysis of model strains The transcriptome data of the model strain of Streptomyces coelicolor from the SRA public database were selected and processed using the following analysis method: Quality control: First, use the software FastQC and Trimmomatic to perform data quality control on the raw data reads. Those that pass the test will be used as clean data for subsequent analysis. Mapping clean data back to the genome: Transcriptome analysis was performed using two transcriptome analysis pipelines, Bowtie2 and STAR software for read mapping. First, index files for the two software were constructed based on the fasta file of the Streptomyces coelicolor reference genome. Then, clean reads in the fastq file were mapped back to the reference genome using Bowtie2 and STAR software based on the index files. The resulting sam files were converted and sorted using samtools to generate binary bam files and sorted.bam files sorted according to the genomic sequence. The sorted bam files were opened using the genome browser IGV to visualize the distribution of sequencing reads, genomes, and annotation information. Transcript quantification and smORF screening: Reads corresponding to Bowtie2 and STAR were mapped back to the genome, and transcript quantification was performed using the featureCounts tool and RSEM software, respectively, based on the S. coelicolor genome annotation file. Read counts for corresponding gene fragments in the clean data were obtained, and the counts were normalized for sequencing depth and then for gene length to obtain FPKM (the number of reads per kilobase of gene length per million reads) to estimate gene expression levels. An FPKM greater than 1 is generally considered a necessary condition for stable transcription of a gene. Transcripts with a length of 303 nt or less and an FPKM greater than 1 in seven transcriptome data sets were screened as potential smORFs in S. coelicolor. 3.4.2) Prediction of shared SEPs and S187-specific SEPs: The Transeq tool in EMBOSS was used to translate the smORFs in the Streptomyces coelicolor transcriptome into amino acid sequences to generate SEP amino acid sequences. These sequences were then used as queries and compared with the amino acid sequences of the entire S187 genome smORF library constructed using genomic smORF prediction using BLASTp. Simultaneously, tBLASTn was used to align the sequences with the nucleotide sequence of the entire S187 genome, with an E-value of 1e-5. Successfully matched SEPs were identified as shared SEPs between the two strains and submitted to the CD-search online tool for conserved domain search and functional site prediction. SEPs possessing both characteristics were considered small proteins that were more likely to exist and function in subsequent analyses. The SEPs shared by S. coelicolor and S187 were compared with 126 SEPs identified by DB search using BLASTp. The parameters were set to Evalue 1e-5 and output format outfmt 6. Small proteins that were not involved in the match were regarded as SEPs unique to S187.
2. The method for predicting and identifying small proteins in Streptomyces marineus S187 according to claim 1, characterized in that: Step 1) includes the following process: 1.1) Whole-genome prediction: Two softwares were selected to predict ORFs in the whole genome of S187: The S187 whole-genome fasta sequence was submitted to ORFFinder for six-frame translation, using four start codons: ATG, GTG, CTG, and TTG, and selecting the bacterial codon 11 translation table. Prodigal software was used to perform genome-wide ORF prediction as a supplement to the ORFFinder tool, thereby obtaining the ribosome binding motifs corresponding to the ORFs. The S187 whole-genome coding sequence was obtained using the bacterial codon 11 table as the operating parameter. The lengths of these amino acid sequences were counted, and ORFs with a length of less than 100 aa were selected as smORFs. 1.2) Collinearity Analysis: The model strain of Streptomyces coelicolor was selected as a reference. The MCScanX tool in TBtools software was used to perform genomic collinearity analysis with S187 to explore the similarity between the two strains. 1.3) Culture and sampling: The preserved strain of Streptomyces xinghaiensis S187 was streaked and activated on TSB solid medium, placed in a constant temperature incubator, and the activated strain was used to prepare fermentation seed liquid; Sampling: First, the fermentation curve and anti-complement activity curve of S187 were used to determine the activity cycle of the strain. The sampling time points were determined according to the fermentation curve and anti-complement activity curve: 36h, 48h, 72h, and 120h. Then, the samples were sampled, eluted, and stored. 1.4) Transcriptome sequencing library construction: Samples from the four time points were combined for RNA-seq library construction. The specific steps were as follows: After RNA extraction, purification, and library construction, the libraries were sequenced using PAIRED paired-end sequencing using NGS sequencing technology based on the Illumina HiSeq sequencing platform. All transcripts obtained from transcriptome sequencing were searched for open reading frames using NCBI ORFFinder, with the minimum ORF length set to 30, the Genetic code set to 11.Bacterial, and the ORF start codon selected as any start codon. The resulting six-frame translation data of the transcriptome was used as a database and searched in the peptide sequencing process. Transcriptome data smRNAs prediction: FastaQC and Trimmomatic were used for quality control of transcriptome sequencing data, Bowtie2 and featureCounts were used for read alignment and quantification, and RNA with FPKM>1 and length≤303nt was screened as a candidate smRNAs library.
3. The method for predicting and identifying small proteins in Streptomyces marineus S187 according to claim 2, wherein in step 1.3), preparing a fermentation seed solution from the activated strain comprises the following steps: Seed liquid collection: Take the S187 bacterial block activated by TSB solid medium, add it to TSB liquid medium, shake and culture at 210-230 rpm, and collect the fermentation liquid after 36 hours; Fermentation broth collection: Add the seed solution to M33 medium and culture at 170-190 rpm with shaking. Increase the speed by 20 rpm every 48 hours until it reaches 210-230 rpm and then maintain the speed unchanged.
4. The method for predicting and identifying small proteins in marine Streptomyces according to claim 2, wherein in step 1.3), the sampling, elution, and preservation comprise the following steps: 1.3.1) From the six bottles of fermentation broth from the same batch, sample three bottles of fermentation broth that are growing well. 3 mL of the sample from each bottle was added to a 50 mL centrifuge tube and mixed thoroughly. From the remaining three bottles of fermentation broth, 9 mL of the sample was transferred to a 50 mL centrifuge tube. The sample was then processed as for the first bottle to balance the sample. 10 mL of the sample was transferred to a centrifuge tube from each of the remaining two bottles for extraction and rotary evaporation. The centrifuge tubes should be kept in an ice bath throughout the entire experimental process. 1.3.2) Removal of proteins from the culture medium: centrifuge and discard the supernatant; 1.3.3) PBS elution: Add pre-chilled PBS solution to the centrifuged pellet and mix thoroughly. Then add PBS, vortex, centrifuge and discard the supernatant. 1.3.4) Repeat step 1.3.3) once; 1.3.5) Removal of CaCO3: Add pre-chilled PBS solution and mix thoroughly. Then add PBS and vortex. Centrifuge and discard the precipitate. Finally, transfer the sample to a new centrifuge tube. 1.3.6) Repeat step 1.3.3) once more to perform elution; 1.3.7) Place the centrifuge tube in a liquid nitrogen tank for quick freezing and store at low temperature.
5. The method for predicting and identifying small proteins in Streptomyces marineus S187 according to claim 1, characterized in that: Step 2) includes the following process: 2.1) Enrichment and processing of micropeptides: Extraction of small proteins and peptides: The S187 mycelial samples obtained by sampling and separation according to the fermentation time were crushed, and then small peptides were extracted with methanol:chloroform:water = 3:1:4 (V:V); vortexed and centrifuged; The upper aqueous supernatant was collected, centrifuged at 20,000 × g using a MWCO of 10 kD, and the lower layer was removed, concentrated in vacuo, dried, desalted, and freeze-dried. HPLC fractionation: A portion of peptides from each sample is mixed and then separated into multiple fractions by high performance liquid chromatography; 2.2) Peptide sequencing: DDA mass spectrometer: prepare mobile phases A and B; Dissolve the lyophilized powder in 10 μL of phase A, centrifuge at 14,000 g for 20 min at 4°C, and inject 1 μg of the supernatant for LC / MS / MS detection. DIA mass spectrometer operation and database search: Dissolve the peptide powder in 0.1% formic acid aqueous solution and add iRT standard peptides. Take 1 μg sample for injection and perform liquid chromatography-mass spectrometry detection.
6. The method for predicting and identifying small proteins in Streptomyces marineus S187 according to claim 1, characterized in that: In step 3.2.2), another novel peptide analysis was performed. First, the actinomycete sequences in the Nr database were aligned using BLASTp, with the task set to blastp-short, the scoring matrix set to PAM30, and the E-value set to 10, thereby removing redundant de novo peptides. Subsequently, tBLASTn was used to align with the S187 genome, setting the PAM30 scoring matrix and finding the best match. Peptides with a similarity and coverage of more than 80% and an E-value ≤ 1 were screened for matches. Subsequently, a Python script was used to perform batch ORF searches based on the genomic location of the peptides. SEPs with a length of less than 100 amino acids were screened from the start codons ATG, GTG, TTG, CTG to the stop codons TGA, TAG, and TAA. The online tool BLASTp was used to filter the annotated homologous sequences in the database with the NCBI Nr database. Finally, peptide sequences with a length of ≥7 aa and no more than 2 amino acid mismatches with the genomic alignment were selected as novel SEPs.
Citation Information
Patent Citations
Bioinformatics method based on protein mass spectrum data annotation eukaryote genome
CN107103205A
Method and system for identifying protein
CN109425662A