Method for improving proteomics identification sensitivity and application thereof
By constructing a polymorphism database covering population genetic diversity, the problems of low sensitivity and false negatives caused by genetic diversity in proteomics identification have been solved, achieving higher identification accuracy and sensitivity.
Patent Information
- Application Number
- CN202511864727.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies in proteomics identification are limited by the limitations of general databases and single genome databases, which cannot effectively cover the genetic diversity of natural species, resulting in the inability to identify peptide and protein sequence variations and reduced sensitivity.
A polymorphism database of multiple sequence sets is constructed, covering population genetic diversity. Genetic variation information of biological samples is identified through nucleic acid sequencing, and protein sequences containing multiple haplotypes are generated for mass spectrometry data search and analysis.
It improves the sensitivity and accuracy of proteomics identification, enabling the identification of peptides and proteins that cannot be covered by a single reference sequence, reducing false negative results, and is suitable for sample analysis with complex genetic backgrounds.
Smart Images

Figure CN121687176A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of proteomics technology, specifically relating to a method for improving the sensitivity of proteomics identification and its application. Background Technology
[0002] High-quality mass spectrometry (HMS) is central to proteomics. It involves comparing the mass spectrometry data of peptides generated from enzymatic digestion of experimental samples with protein sequence databases—a process known as library search—to identify peptides and infer proteins. Currently, library search heavily relies on high-quality reference protein sequence databases.
[0003] The existing technology has the following main shortcomings: 1) Limitations of general-purpose databases: The most commonly used public databases, such as UniProt, typically only provide the "standard" or "consensus" sequence for that species. However, there is extensive genetic diversity in natural species, leading to differences in protein sequences between individuals. When searching using general-purpose databases, any peptides that do not match the standard sequence, such as those resulting from mutations, will be incorrectly filtered out, leading to identification failures and reduced sensitivity.
[0004] 2) Limitations of single-genome databases: Even protein sequence databases derived from genome sequencing data of a specific individual can only reflect the sequence information of that single individual. When analyzing samples from different genetic backgrounds, other haplotype sequences present in the samples cannot be covered by this database, resulting in limited improvement in sensitivity.
[0005] 3) Ignoring sequence variation: Traditional database search algorithms assume that the protein sequences of all analyzed samples are completely identical to the reference sequence, which cannot effectively deal with the complex sequence variation caused by population genetic diversity.
[0006] Existing technologies use single reference sequence data, such as the UniProt universal library or a single genome sequence, for mass spectrometry data searches. This cannot effectively cover protein sequence variations caused by genetic polymorphisms in natural species, such as single nucleotide polymorphisms (SNPs) and insertions / deletions (InDels), resulting in a large number of peptides being unidentifiable and reduced sensitivity.
[0007] Based on this, this study investigates and relates to a method for improving the sensitivity of proteomics identification and its application. Summary of the Invention
[0008] To address the aforementioned deficiencies in existing technologies, the present invention aims to provide a method and its application for improving the sensitivity of proteomics identification. Based primarily on existing methods for improving proteomics identification sensitivity, which rely on single reference sequence data such as a single genome for mass spectrometry database searches, some protein sequences, such as single nucleotide polymorphisms (SNPs) and insertions / deletions (InDels) caused by genetic diversity variations, cannot be identified. The present invention constructs a multi-sequence set representing the genetic diversity of the target species as a database search reference; that is, it builds a personalized polymorphic reference database based on population genetic diversity, instead of using a single standard sequence. The construction of this polymorphic database solves the technical problems of low sensitivity caused by missing a large number of peptides and proteins, and reduces false negative results due to sequence mismatches.
[0009] This invention is achieved through the following technical solution: A method for improving the sensitivity of proteomics identification includes the following steps: a) Sample collection and sequencing: Obtain biological samples with genetic diversity within the target species and perform nucleic acid sequencing on the biological samples; b) Construction of a polymorphism database: Identify genetic variation information in gene coding regions of the biological samples, translate and generate multiple haplotype protein sequences containing the genetic variation information, integrate the multiple haplotype protein sequences and construct a non-single sequence polymorphism database containing population-level proteins; c) Mass spectrometry data search and analysis: The proteomic mass spectrometry data of the sample to be tested is searched using the polymorphism database constructed in step b) to identify peptides and proteins.
[0010] Optionally, the nucleic acid sequencing in step a) can be whole genome resequencing, exon sequencing, or transcriptome sequencing. Optionally, the genetic variation information in step b) includes single nucleotide polymorphisms (SNPs) and insertions / deletions (InDels).
[0011] Optionally, the polymorphism database in step b) is a collection database containing all common haplotype sequences or a graphical database that integrates genetic variation information in the form of an index.
[0012] Another object of the present invention is to provide a proteomics analysis system, comprising: A memory that stores polymorphic database information constructed based on the above-mentioned methods for improving the sensitivity of proteomics identification; A processor is configured to receive mass spectrometry data of a sample to be tested and to retrieve and analyze the mass spectrometry data by calling the polymorphism database.
[0013] Another object of the present invention is to improve the application of proteomics identification sensitivity methods, including at least one of the following applications. 1) Discovery of characteristic biomarkers of diseases; 2) Identification of proteins related to desirable traits in crop or livestock breeding; 3) Qualitative and quantitative analysis of the effective components of polypeptide extracts; 4) Traceability of agricultural products or quality identification of traditional Chinese medicinal materials; 5) Proteomics studies or species evolution and adaptation studies of non-model organisms or rare species.
[0014] Compared with the prior art, the beneficial effects of the present invention are: 1) High sensitivity: By covering sequence diversity at the population level, it can identify a large number of peptides and proteins that are missed when using a single reference sequence, which is especially suitable for the analysis of samples with complex genetic backgrounds.
[0015] 2) High accuracy: Reduces false negative results caused by sequence mismatch, making protein identification results more accurately reflect the actual composition of the sample.
[0016] 3) Wide range of applications: This method can be widely applied to precision medicine, such as discovering individual-specific disease-related protein variations, animal and plant breeding, such as identifying protein variants associated with desirable traits, evolutionary biology, and other fields, and has significant scientific research and industrial value. Attached Figure Description
[0017] Figure 1 This is a schematic flowchart of the method of the present invention; Figure 2 Constructing a result graph for polymorphic data; Figure 3 The image shows the mass spectrometry data analysis results of the phenanthrene hirudin sample in Example 3. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments. The illustrative embodiments and descriptions of this invention are only used to explain this invention and are not intended to limit this invention.
[0019] Example 1: like Figure 1 As shown, a method for improving the sensitivity of proteomics identification includes the following steps: 1) Sample collection and sequencing: Obtain biological samples with genetic diversity within the target species, and perform nucleic acid sequencing on the biological samples; 2) Constructing a polymorphism database: Identify genetic variation information in gene coding regions of biological samples, translate and generate multiple haplotype protein sequences containing the genetic variation information, integrate the multiple haplotype protein sequences, and construct a non-single sequence polymorphism database containing population-level proteins; 3) Mass spectrometry data search and analysis: The proteomic mass spectrometry data of the sample to be tested is searched using the polymorphism database constructed in step 2) to identify peptides and proteins.
[0020] In step 1), obtain biological samples with genetic diversity within the target species. Systematically collect samples from different populations, subspecies, or varieties of the target species to ensure that the collected samples cover the population level. Perform nucleic acid sequencing on the biological samples. The specific nucleic acid sequencing methods are: whole genome resequencing, exon sequencing, or transcriptome sequencing.
[0021] In step 2), the genetic variation information of gene coding regions in the biological sample is identified, including single nucleotide polymorphisms (SNPs) and insertions / deletions (InDels). By splicing, aligning, and detecting variations in the sequencing data, and then annotating based on a reference genome, the genetic variation information of the coding regions is translated into corresponding amino acid variations, thereby generating a set of protein sequences containing all major haplotypes. This is then constructed into a dedicated FASTA format database or a more advanced graphical database. The polymorphism database is either a collection database containing all haplotype sequences or a graphical database that integrates genetic variation information in an indexed format.
[0022] Example 2: A proteomics analysis system, comprising: The memory stores a polymorphism database constructed according to the method for improving proteomics identification as described in Example 1; The processor is configured to receive mass spectrometry data of the sample to be tested and to call the polymorphism database to retrieve and analyze the mass spectrometry data.
[0023] The method for improving the sensitivity of proteomics identification described in Example 1 and the proteomics analysis system described in Example 2 can both be applied to: 1) the discovery of characteristic biomarkers of diseases, 2) the identification of proteins related to superior traits in crop or livestock breeding, 3) the qualitative and quantitative analysis of effective components in polypeptide extracts, 4) the traceability of agricultural products or the quality identification of traditional Chinese medicine materials, and 5) proteomics research or species evolution and adaptation research of non-model organisms or rare species.
[0024] Example 3: The following uses the deep identification of hirudin as an example to illustrate the specific implementation of the present invention.
[0025] 1) Sample collection and sequencing: More than 70 wild samples of Hirudo nipponia were collected from Guangdong, Guangxi, Hainan and Yunnan provinces of China. The samples were kept in the laboratory for more than one month and then the genome was sequenced. Head tissue of each leech was taken to enhance the genomic DNA and live tissue RNA. Whole genome sequencing and transcriptome sequencing were performed separately. De novo assembly of the genome and transcriptome was performed using MEGAHIT and TRINITY software respectively. 2) Construction of polymorphism database: Using the five published leech sequences of Hirudin as bait, the hirudin gene sequences in all genomes and transcriptomes were extracted using BLAST software and translated into protein sequences. The protein sequences corresponding to different haplotypes of each gene locus were generated using DAMBE software. These sequences were then merged with all protein sequences obtained from the Hirudin genome annotation and finally integrated into a leech population-specific polymorphic protein sequence database. 3) Proteomics analysis: Leech peptide extraction, enzymatic digestion, and mass spectrometry detection were performed according to steps 1)-2); 4) Mass spectrometry data analysis: The sequencing samples B1, B2, and B3 of Hirudin for analysis were searched using (a) the UniProt database; (b) the protein database of a single genome of this species; and (c) the polymorphism database constructed in this invention. The results confirmed that using the database (c) could identify more hirudin peptides and relative signal intensities of the spectral flow. The results of the identified proteomics analysis are shown in Table 1 below, which significantly improved the depth and breadth of the analysis.
[0026] The results of constructing a polymorphic database are as follows: Figure 2 As shown, sequences 1-5 are five hirudin polymorphic sequences identified in the genome of Hirudin hymenopus. Using the polymorphism database constructed in this invention, 36 new hirudin_Hman1 haplotypes were added, along with 47 new haplotypes of the other four hirudins.
[0027] Mass spectrometry data analysis results as follows Figure 3 As shown, many hirudin peptide fragments, such as CVLGSSTSENR, could not be found in databases (a) and (b), but two matches could be found in the polymorphism database (c). The spectral flow relative signal intensity was very high, indicating the reliability of the analysis results.
[0028] The results confirmed that using the database (c) could identify more hirudin peptides and relative signal intensities of the spectral flow. The results of the identified proteomics analysis are shown in Table 1, which significantly improved the depth and breadth of the analysis.
[0029] Table 1: Comparison of the results of the identified proteomics analysis. Using database (a), only sample B3 yielded one matching peptide; using database (b), samples B1-B3 each yielded one matching peptide; while using database (c), samples B1-B3 yielded 5, 4, and 4 matching peptides, respectively. Furthermore, the relative signal intensity of the spectral flux obtained using database (c) was significantly higher than that obtained using databases (a) and (b).
[0030] It is evident that using database (c) the polymorphism database for searching, which covers the diverse sequences of hirudin, and merging the protein sequences corresponding to different haplotypes at each gene locus based on the published leech sequences of Hirudin, ultimately forming a leech population-specific polymorphic protein sequence database, compared with database (b) the protein database in a single genome of this species, reduces the possibility of missing some peptides and proteins, improves the sensitivity of identification, reduces false negative results caused by sequence mismatch, and makes the protein identification results more realistically reflect the actual composition of the sample, thus improving the accuracy of identification.
[0031] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method of increasing the sensitivity of proteomic identification, characterized by: The method comprises the following steps: a) sample collection and sequencing: obtaining biological samples with genetic diversity in the target species, and performing nucleic acid sequencing on the biological samples; b) polymorphism database construction: identifying genetic variation information of gene coding regions in the biological samples, and translating to generate a plurality of haplotype protein sequences containing the genetic variation information, integrating the plurality of haplotype protein sequences, and constructing a polymorphism database containing non-single sequence and population level protein; c) mass spectrometry data library analysis: using the polymorphism database constructed in step b) to search the proteomic mass spectrometry data of the sample to be tested to identify peptides and proteins.
2. The method of claim 1, wherein the method is used to improve the sensitivity of proteomic identification. The nucleic acid sequencing in step a) is whole genome resequencing, exon sequencing or transcriptome sequencing.
3. The method of claim 1, wherein the method is used to improve the sensitivity of proteomic identification. The genetic variation information in step b) includes single nucleotide polymorphism (SNP) and insertion and deletion (InDel).
4. The method of claim 1, wherein the method is used to improve the sensitivity of proteomic identification. The polymorphism database in step b) is a collection database containing all haplotype sequences or a graph database integrating genetic variation information in index form.
5. A proteomic analysis system characterized by, It comprises: a memory storing a polymorphism database constructed according to any one of the methods of claims 1-5; a processor configured to receive mass spectrometry data of a sample to be tested, and call the polymorphism database to search and analyze the mass spectrometry data.
6. Use of a method according to any one of claims 1 to 5 for increasing the sensitivity of proteomic identification, characterized in that, It is applied to the discovery of disease characteristic biomarkers.
7. Use of a method according to any one of claims 1 to 5 for increasing the sensitivity of proteomic identification, characterized in that, It is applied to the identification of excellent trait related proteins in crop or livestock breeding.
8. Use of a method according to any one of claims 1 to 5 for increasing the sensitivity of proteomic identification, characterized in that, It is applied to the qualitative and quantitative analysis of effective components of polypeptide extracts.
9. Use of a method according to any one of claims 1 to 5 for increasing the sensitivity of proteomic identification, characterized in that, It is applied to the traceability of agricultural products or the quality identification of traditional Chinese medicinal materials.
10. Use of a method according to any one of claims 1 to 5 for increasing the sensitivity of proteomic identification, characterized in that, It is applied to the proteomic research of non-model organisms or rare species or the research of species evolution and adaptability. It is applied to the discovery of disease characteristic biomarkers. It is applied to the identification of excellent trait related proteins in crop or livestock breeding. It is applied to the qualitative and quantitative analysis of effective components of polypeptide extracts. It is applied to the traceability of agricultural products or the quality identification of traditional Chinese medicinal materials. It is applied to the proteomic research of non-model organisms or rare species or the research of species evolution and adaptability.
Citation Information
Patent Citations
Protein identification method
CN103177198A
A computational method for mapping peptides to proteins using sequencing data
CN103488913A
Protein tandem mass spectrometry identification method based on multiple omics abundance information
CN106404878A
Construction method of reference protein database, storage medium and electronic equipment
CN113393903A
Method for predicting and identifying small proteins in marine streptomyces S187
CN116741276A