Workflow to assign putative source to de novo peptide sequence
The workflow for assigning a putative source to de novo peptide sequences by systematically searching across multiple databases based on random hit rates addresses the reliability issues in existing methods, enhancing the accuracy of source determination.
Patent Information
- Application Number
- JP2025037981
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-03-11
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-17
AI Technical Summary
Existing methods for assigning a putative source to de novo peptide sequences lack reliability in the absence of experimental confirmation, and they struggle to efficiently determine the source among multiple databases.
A workflow that involves searching peptide sequences across multiple databases, including an extended human proteome database, in increasing order of random hit rate, to determine the putative source. This workflow uses various search types such as linear human proteome search, linear human genome search, linear mismatch search, and cis-splice search to identify the most likely source.
The proposed workflow increases the reliability of assigning a putative source to de novo peptide sequences by systematically evaluating multiple databases based on random hit rates, thereby improving the accuracy of source determination.
Smart Images

Figure 2025090717000133 
Figure 2025090717000134 
Figure 2025090717000135
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority based on U.S. Provisional Application No. 63 / 159,879, filed on March 11, 2021, and U.S. Provisional Application No. 63 / 159,880, filed on March 11, 2021, the entire contents of each of which are incorporated herein by reference.
[0002] Technical Field In some embodiments, the present invention relates to a computer method / system for optimizing search results by querying multiple databases according to false discovery rate and random hit rate.
Background Art
[0003] Background Many data sources are created and maintained around the world. It is almost guaranteed that many data sources can answer some or all of the queries using one of these data sources. The mere task of executing a query can be difficult, whether the scope of the query is within a local computer system, a private network, a local area network, or the World Wide Web. The process of querying these data sources is made even more difficult because the user has to determine which data sources are reliable enough to obtain meaningful search results. For example, the user has to consider both the relative accuracy of the source and the timeliness of the data contained within the source. These and other drawbacks are addressed herein.
[0004] In the field of bioinformatics, attempts have been made to assign putative sources (e.g., translation from RNA, synthesis from DNA, etc.) to de novo peptide sequences. By definition, putative sources are not fully experimentally confirmed and thus are imperfect.
Summary of the Invention
Means for Solving the Problem
[0005] Brief Summary Generally, embodiments of methods, systems, and devices are described herein that are directed to assigning a putative source to a de novo peptide sequence and / or creating a workflow for performing said assignment. In some embodiments, the invention includes a workflow that increases the reliability of an assigned putative source in the absence of experimental confirmation of the source assignment.
[0006] In one embodiment, the putative source of a peptide sequence can be determined based at least in part on one or more searches of peptide sequences within one or more databases, such that the one or more searches are performed in increasing order of random hit rate until the putative source is determined. The random hit rate for each individual search can be determined based at least in part on the number of random peptide sequences found by the individual search. The one or more databases can include, but are not limited to, an extended human proteome database, a human genome database, a non-endogenous proteome database, additional databases, and combinations thereof. The one or more searches can include, but are not limited to, a linear human proteome search of peptide sequences within an extended human proteome database, a linear human genome search of the translation of a human genome database, a linear mismatch search for peptides having a mismatch to peptide sequences within an extended human proteome database, a linear non-endogenous search of peptide sequences within a non-endogenous proteome database, a cis-splice search within an extended human proteome database, and a trans-splice search within an extended human proteome database. Each of the searches can indicate the possible source of the individual peptide sequences being searched if an individual source finds a match. The putative source determined for the peptide sequences being searched can be the possible source identified by the search step having the lowest random hit rate at which a match was found for the search peptide.
[0007] In one embodiment, the peptide source search steps may be ordered to create a peptide source assignment workflow. A plurality of random peptide arrays may be created, and each of the random peptide arrays may be searched by a respective peptide source search step. The random hit rate can be determined, at least in part, for each peptide source search step, based on the number of random peptide arrays found by the peptide source search step. The peptide source search steps may be ordered in the workflow from the lowest random hit rate to the highest random hit rate. The random hit rate may increase as the number of random peptide arrays found increases. Peptide source search steps may include, but are not limited to, a linear human proteome search for peptide sequences in an extended human proteome database, a linear human genome search of the translation of the human genome database, a linear mismatch search for peptides having a mismatch to a peptide sequence, a linear non-endogenous search for peptide sequences in a non-endogenous proteome database, a cis-splice search in an extended human proteome database, and a trans-splice search in an extended human proteome database.
[0008] Embodiments of the method of the present invention are described herein that include creating a plurality of simulated random queries, determining the number of matches associated with each source based on applying the plurality of simulated random queries to each source of a plurality of sources, determining the false discovery rate associated with each source based on the number of matches associated with each source, and creating a query support data structure configured to facilitate the application of new queries to the plurality of sources based on the false discovery rate.
[0009] Receiving a query, applying the query to one or more of a plurality of sources based on a query support data structure, determining an identifier associated with a source of the plurality of sources associated with the query result based on the query result, and applying the identifier to the query, method embodiments are also described.
[0010] The various steps of the methods disclosed herein, or the steps performed by the systems disclosed herein, may be performed at the same or different times, in the same or different geographical locations, e.g., countries, and / or by the same or different people.
[0011] Additional advantages of method and system embodiments are, in part, shown in the following description, in part, understood from the description, or can be learned by practicing method and system embodiments. The advantages of the disclosed embodiments of the method and system will be realized and attained by the elements and combinations particularly pointed out in the appended claims. It should be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention as claimed.
[0012] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments of the disclosed methods and systems and, together with the description, serve to explain the principles of the disclosed methods and systems. The present invention provides, for example, the following items. (Item 1) A method for determining a putative source of a peptide sequence of a peptide, comprising: receiving the peptide sequence, and determining the putative source associated with the peptide sequence based at least in part on one or more searches of the peptide sequence in one or more databases. Each individual search of the one or more searches has a random hit rate that is at least partially based on the number of random arrays found by the individual search, the one or more searches are performed in increasing order of random hit rate until the putative source is determined, Method. (Item 2) the one or more databases include an extended human proteome database, the extended human proteome database includes computer-readable representations of translations from messenger ribonucleic acid (RNA) and non-coding RNA, The method according to item 1. (Item 3) The method according to item 2, wherein the extended human proteome database includes computer-readable representations of translations from microRNA. (Item 4) The method according to item 2, wherein the extended human proteome database includes computer-readable representations of translations from long non-coding RNA. (Item 5) The method according to item 2, wherein the extended human proteome database includes computer-readable representations of translations of human endogenous retroviruses. (Item 6) the one or more searches include a linear human proteome search for the peptide sequence in the extended human proteome database, when the linear human proteome search for the peptide sequence in the extended human proteome database finds the peptide sequence in the extended human proteome database, the putative source is the linear extended human proteome source, The method according to item 2. (Item 7) identifying whether the peptide is putatively translated from messenger RNA or non-coding RNA when the source is the linear extended human proteome source, The method according to item 6, further comprising. (Item 8) The one or more databases include a human genome database, The one or more searches include a linear human genome search of the translation of the human genome database, When the linear human genome search finds a human genome sequence from which the peptide is presumably synthesized, the presumed source is a linear genome source, The method according to item 2. (Item 9) The linear human genome search excludes the portion of the human genome where the messenger RNA and the non-coding RNA of the extended human proteome database are transcribed, and includes the remaining portion of the human genome, The linear human genome search includes a search for six-frame translation of the human genome, The method according to item 8. (Item 10) The one or more searches include a linear mismatch search for peptides having a mismatch to the peptide sequence within the extended human proteome database, When the linear mismatch search finds a peptide sequence having a mismatch to the peptide sequence within the extended human proteome database, the presumed source is a linear mismatch of the extended human proteome, The method according to item 2. (Item 11) The method according to item 10, wherein the linear mismatch search is a search for peptide sequences having only a single mismatch to the peptide sequence. (Item 12) The one or more databases include a non-endogenous proteome database containing computer-readable representations of proteins translated from RNA derived from non-endogenous organisms and / or proteins synthesized by non-endogenous organisms, The one or more searches include a linear non-endogenous search for the peptide sequence within the non-endogenous proteome database, When the linear non-endogenous search finds the peptide sequence in the non-endogenous proteome database, the putative source is a linear non-endogenous proteome source. The method according to item 1. (Item 13) The method according to item 12, wherein the non-endogenous proteome database includes a Basic Local Alignment Search Tool (BLAST) database. (Item 14) One or more of the searches include a cis-splice search for peptide fragments that can be cis-spliced to match the peptide sequence in the extended human proteome database. When the cis-splice search finds a peptide fragment that can be cis-spliced to match the peptide sequence in the extended human proteome database, the source is a cis-spliced human proteome source. The method according to item 2. (Item 15) One or more of the searches include a trans-splice search for a computer-readable representation of a peptide fragment that can be trans-spliced to match the peptide sequence in the extended human proteome database. When the trans-splice search finds a computer-readable representation of a peptide fragment that can be trans-spliced to match the peptide sequence in the extended human proteome database, the source is a trans-spliced human proteome source. The method according to item 2. (Item 16) When the trans-splice search does not find a computer-readable representation of a peptide fragment that can be trans-spliced to match the peptide sequence, the putative source is determined to be unidentified, according to the method of item 15. (Item 17) One or more of the databases include a human genome database. The one or more searches are the following searches that are sequentially ordered in the workflow as follows: A linear human proteome search for the peptide sequence in the extended human proteome database, A linear human genome search of the translation of the human genome database, A linear mismatch search for peptides having a mismatch to the peptide sequence in the extended human proteome database, and A cis-splice search for peptide fragments that can be cis-spliced to match the peptide sequence in the extended human proteome database including The method according to item 2. (Item 18) The one or more databases include a non-endogenous proteome database containing computer-readable representations of proteins translated from RNA derived from non-endogenous organisms and / or proteins synthesized by non-endogenous organisms, The one or more searches further include a linear non-endogenous search for the peptide sequence in the non-endogenous proteome database, The linear non-endogenous search is sequentially ordered in the workflow after the linear mismatch search and before the cis-splice search, The method according to item 17. (Item 19) The one or more searches further include a trans-splice search for peptide fragments that can be trans-spliced to match the peptide sequence in the extended human proteome database, The trans-splice search is sequentially ordered in the workflow after the cis-splice search, The method according to item 17. (Item 20) When the putative source is determined for the peptide sequence, stopping the advancement of the workflow to subsequent searches of the one or more searches The method according to item 17, further comprising (Item 21) wherein the peptide sequence contains at least one ambiguous residue, creating a plurality of alternative peptide sequences each containing a possible residue for each of the at least one ambiguous residue, determining for each of the plurality of alternative peptide sequences an individual possible source, and determining the putative source of the peptide sequence such that the putative source is an individual possible source further comprising The method according to item 1 (Item 22) The method according to item 21, wherein the possible residues for each of the at least one ambiguous residue include leucine and isoleucine. (Item 23) determining an individual random hit rate for each of the individual possible sources such that the random hit rate increases as the number of random sequences is found by the individual searches of the one or more searches, and determining the putative source such that the individual random hit rate of the putative source is the lowest individual random hit rate for each of the possible sources further comprising the method according to item 21. (Item 24) identifying the one or more promising alternative peptide sequences of the plurality of alternative peptide sequences such that each of the one or more promising alternative peptide sequences is associated with the putative source further comprising the method according to item 21. (Item 25) The method according to item 1, wherein the peptide sequence is a de novo peptide sequence determined by mass spectrometry. (Item 26) A non-transitory computer-readable medium configured to communicate with one or more processors of a computer device, the non-transitory computer-readable medium including instructions thereon that, when executed by the processor, cause the computer device to, receive, as input, a peptide sequence, determine a putative source associated with the peptide sequence based at least in part on one or more searches of the peptide sequence in one or more databases, wherein each individual search of the one or more searches has a random hit rate based at least in part on the number of random sequences found by the individual search, the one or more searches are performed in an order of increasing random hit rate until the putative source is determined, and provide, as output, the putative source. A non-transitory computer-readable medium. (Item 27) The one or more databases include an extended human proteome database, the extended human proteome database includes computer-readable representations of translations from messenger ribonucleic acid (RNA) and non-coding RNA, The non-transitory computer-readable medium according to item 26. (Item 28) The non-transitory computer-readable medium according to item 27, wherein the extended human proteome database includes computer-readable representations of translations from microRNA. (Item 29) The non-transitory computer-readable medium according to item 27, wherein the extended human proteome database includes computer-readable representations of translations from long non-coding RNA. (Item 30) The non-transitory computer-readable medium according to item 27, wherein the extended human proteome database includes computer-readable representations of translations from human endogenous retroviruses. (Item 31) the one or more searches include a linear human proteome search for the peptide sequence in the extended human proteome database, when the linear human proteome search for the peptide sequence in the extended human proteome database finds the peptide sequence in the extended human proteome database, the putative source is a linear extended human proteome source, The non - transitory computer - readable medium according to item 27. (Item 32) when the instructions are executed by the processor, cause the computer device to, when the source is the linear extended human proteome source, identify whether the peptide is presumably translated from messenger RNA or non - coding RNA, the non - transitory computer - readable medium according to item 31. (Item 33) the one or more databases include a human genome database, the one or more searches include a linear human genome search of the translation of the human genome database, when the linear human genome search finds a human genome sequence from which the peptide is presumably synthesized, the putative source is a linear genome source, The non - transitory computer - readable medium according to item 27. (Item 34) the linear human genome search excludes the part of the human genome from which the messenger RNA and the non - coding RNA in the extended human proteome database are transcribed, and includes the remaining part of the human genome, The non - transitory computer - readable medium according to item 33. (Item 35) the one or more searches include a linear mismatch search for peptides having a mismatch to the peptide sequence in the extended human proteome database, When the linear mismatch search finds a peptide sequence having a mismatch with respect to the peptide sequence in the extended human proteome database, the putative source is a linear mismatch of the extended human proteome. The non-transitory computer-readable medium according to item 27. (Item 36) The non-transitory computer-readable medium according to item 35, wherein the linear mismatch search is a search for a peptide sequence having only a single mismatch with respect to the peptide sequence. (Item 37) The one or more databases include a non-endogenous proteome database containing computer-readable representations of proteins translated from RNA derived from non-endogenous organisms and / or proteins synthesized by non-endogenous organisms. The one or more searches include a linear non-endogenous search for the peptide sequence in the non-endogenous proteome database. When the linear non-endogenous search finds the peptide sequence in the non-endogenous proteome database, the putative source is a linear non-endogenous proteome source. The non-transitory computer-readable medium according to item 26. (Item 38) The non-transitory computer-readable medium according to item 37, wherein the non-endogenous proteome database includes a basic local alignment search tool (BLAST) database. (Item 39) The one or more searches include a cis-splicing search for peptide fragments that can be cis-spliced to match the peptide sequence in the extended human proteome database. When the cis-splicing search finds a peptide fragment that can be cis-spliced to match the peptide sequence in the extended human proteome database, the source is a cis-spliced human proteome source. The non-transitory computer-readable medium according to item 27. (Item 40) The one or more searches include a trans-splice search for a computer-readable representation of a peptide fragment that can be trans-spliced to match the peptide sequence within the extended human proteome database, wherein when the trans-splice search finds a computer-readable representation of a peptide fragment that can be trans-spliced to match the peptide sequence within the extended human proteome database, the source is a trans-spliced human proteome source, The non-transitory computer-readable medium according to item 27. (Item 41) The non-transitory computer-readable medium according to item 40, wherein when the trans-splice search does not find a computer-readable representation of a peptide fragment that can be trans-spliced to match the peptide sequence, the putative source is determined to be unidentified. (Item 42) The one or more databases include a human genome database, The one or more searches are sequentially ordered in a workflow as follows: A linear human proteome search for the peptide sequence within the extended human proteome database, A linear human genome search of the translation of the human genome database, A linear mismatch search for peptides having a mismatch to the peptide sequence within the extended human proteome database, and A cis-splice search for peptide fragments that can be cis-spliced to match the peptide sequence within the extended human proteome database including, The non-transitory computer-readable medium according to item 27. (Item 43) The one or more databases include a non-endogenous proteome database that includes computer-readable representations of proteins translated from RNA derived from non-endogenous organisms and / or proteins synthesized by non-endogenous organisms, The one or more searches further include a linear non-endogenous search for the peptide sequence within the non-endogenous proteome database, In the workflow, the linear non-endogenous search is sequentially ordered after the linear mismatch search and before the cis-splice search, The non-transitory computer-readable medium according to item 42. (Item 44) The one or more searches further include a trans-splice search for peptide fragments within the extended human proteome database that can be trans-spliced to match the peptide sequence, In the workflow, the trans-splice search is sequentially ordered after the cis-splice search, The method according to item 42. (Item 45) When the instructions are executed by the processor, cause the computer device to When the putative source is determined for the peptide sequence, stop the advancement of the workflow to subsequent searches of the one or more searches, The non-transitory computer-readable medium according to item 42. (Item 46) The peptide sequence includes at least one ambiguous residue, When the instructions are executed by the processor, cause the computer device to Create a plurality of alternative peptide sequences each containing a residue that may be possible for each of the at least one ambiguous residue, For each of the plurality of alternative peptide sequences, determine an individual possible source, and Cause the putative source of the peptide sequence to be determined such that the putative source is an individual possible source. The non-transitory computer-readable medium according to item 26. (Item 47) The non-transitory computer-readable medium according to item 46, wherein the possible residues for each of the at least one ambiguous residue include leucine and isoleucine. (Item 48) When the instructions are executed by the processor, cause the computer device to Determine an individual random hit rate for each of the individual possible sources such that the random hit rate increases as the number of random sequences is found by the individual searches of the one or more searches, and Cause the putative source to be determined such that the individual random hit rate of the putative source is the lowest of the individual random hit rates for each of the possible sources. The non-transitory computer-readable medium according to item 46. (Item 49) When the instructions are executed by the processor, cause the computer device to Identify the one or more potentially substituted peptide sequences such that each of the one or more potentially substituted peptide sequences is associated with the putative source. The non-transitory computer-readable medium according to item 46. (Item 50) The non-transitory computer-readable medium according to item 26, wherein the peptide sequence is a de novo peptide sequence determined by mass spectrometry. (Item 51) A method of ordering a peptide source assignment workflow, comprising: Creating a plurality of random peptide sequences, Determining a plurality of peptide source search steps, For each of the plurality of peptide source search steps, searching for each of the plurality of random peptide sequences; For each of the plurality of peptide source search steps, determining the random hit rate for an individual search step of the plurality of peptide source search steps, at least in part based on the number of the plurality of random peptide sequences found by the individual search step, and In the peptide source assignment workflow, ordering the peptide source search steps from the lowest random hit rate to the highest random hit rate A method comprising. (Item 52) The method according to item 51, wherein the random peptide sequence comprises a random sequence that uniformly samples all amino acids. (Item 53) The method according to item 51, wherein the random peptide sequence comprises a sequence having an amino acid frequency that matches the frequency of amino acids found in vertebrates. (Item 54) The method according to item 51, wherein each peptide of the random peptide sequence has a length of 8 to 14 amino acids. (Item 55) The method according to item 51, wherein each peptide of the random peptide sequence has a length of 9 to 14 amino acids, 10 to 14 amino acids, 11 to 14 amino acids, 12 to 14 amino acids, 13 to 14 amino acids, 8 to 13 amino acids, 8 to 12 amino acids, 8 to 11 amino acids, 8 to 10 amino acids, 8 to 9 amino acids, 9 to 13 amino acids, 9 to 12 amino acids, 9 to 11 amino acids, 9 to 10 amino acids, 10 to 13 amino acids, 10 to 12 amino acids, 10 to 11 amino acids, 11 to 13 amino acids, 11 to 12 amino acids, or 12 to 13 amino acids. (Item 56) The plurality of peptide source search steps includes a linear human proteome search for peptide sequences in an extended human proteome database, The extended human proteome database includes computer-readable representations of translation from messenger ribonucleic acid (RNA) and non-coding RNA, The method according to item 51. (Item 57) The method according to item 56, wherein the extended human proteome database includes computer-readable representations of translation from microRNA. (Item 58) The method according to item 56, wherein the extended human proteome database includes computer-readable representations of translation from long non-coding RNA. (Item 59) The method according to item 56, wherein the extended human proteome database includes computer-readable representations of translation from human endogenous retroviruses. (Item 60) The method according to item 56, wherein the plurality of peptide source search steps includes a linear human genome search of translation of the human genome database. (Item 61) The linear human genome search excludes the portion of the human genome from which the messenger RNA and the non-coding RNA of the extended human proteome database are transcribed, and includes the remaining portion of the human genome, The method according to item 60. (Item 62) The plurality of peptide source search steps includes a linear mismatch search for peptides having a mismatch to the peptide sequence in the extended human proteome database, The extended human proteome database includes computer-readable representations of translation from messenger ribonucleic acid (RNA) and non-coding RNA, The method according to item 51. (Item 63) The plurality of peptide source search steps includes a linear non-endogenous search for the peptide sequence in the non-endogenous proteome database, The non-endogenous proteome database includes a computer-readable representation of proteins translated from RNA derived from non-endogenous organisms and / or proteins synthesized by non-endogenous organisms. The method according to item 51. (Item 64) The method according to item 63, wherein the non-endogenous proteome database includes a basic local alignment search tool (BLAST) database. (Item 65) The plurality of peptide source search steps includes a cis-splice search for peptide fragments that can be cis-spliced to match the peptide sequence within an extended human proteome database. The extended human proteome database includes a computer-readable representation of translations from messenger ribonucleic acid (RNA) and non-coding RNA. The method according to item 51. (Item 66) The plurality of peptide source search steps includes a trans-splice search for peptide fragments that can be trans-spliced to match the peptide sequence within an extended human proteome database. The extended human proteome database includes a computer-readable representation of translations from messenger ribonucleic acid (RNA) and non-coding RNA. The method according to item 51. (Item 67) The method according to item 51, wherein when a peptide is not assigned a peptide source by any of the plurality of peptide source search steps, the peptide source assignment workflow ends with the unassigned peptide. (Item 68) The peptide source assignment workflow is the following searches sequentially ordered as follows: A linear human proteome search for the peptide sequence within the extended human proteome database. A linear human genome search of the translation of the human genome database. A linear mismatch search for peptides having a mismatch to the peptide sequence in the extended human proteome database, and a cis-splice search for peptide fragments that can be cis-spliced to match the peptide sequence in the extended human proteome database The method according to item 51, comprising: (Item 69) The peptide source assignment workflow includes a linear non-endogenous search for the peptide sequence in a non-endogenous proteome database, The linear non-endogenous search is sequentially ordered after the linear mismatch search and before the cis-splice search within the peptide assignment workflow, The method according to item 68. (Item 70) The peptide source assignment workflow includes a trans-splice search for peptide fragments that can be trans-spliced to match the peptide sequence in the extended human proteome database, The trans-splice search is sequentially ordered after the cis-splice search within the peptide assignment workflow, The method according to item 68. (Item 71) A non-transitory computer-readable medium configured to communicate with one or more processors of a computer device, the non-transitory computer-readable medium includes instructions thereon, and when the instructions are executed by the processor, cause the computer device to receive, as input, a plurality of peptide source search steps, create a plurality of random peptide sequences, search each of the plurality of random peptide sequences for each of the plurality of peptide source search steps, For each of the plurality of peptide source search steps, cause the random hit rate for an individual search step of the plurality of peptide source search steps to be determined, at least in part, based on the number of the plurality of random peptide sequences found by the individual search step, order the peptide source search steps in a peptide source assignment workflow from the lowest random hit rate to the highest random hit rate, and a non-transitory computer-readable medium that causes the peptide source assignment workflow to be provided as output. (Item 72) The non-transitory computer-readable medium according to item 71, wherein the random peptide sequences include random sequences that uniformly sample all amino acids. (Item 73) The non-transitory computer-readable medium according to item 71, wherein the random peptide sequences include sequences having amino acid frequencies that match the frequencies of amino acids found in vertebrates. (Item 74) The non-transitory computer-readable medium according to item 71, wherein each peptide of the random peptide sequences includes a length of 8 to 14 amino acids. (Item 75) The non-transitory computer-readable medium according to item 71, wherein each peptide of the random peptide sequences includes a length of 9 to 14 amino acids, 10 to 14 amino acids, 11 to 14 amino acids, 12 to 14 amino acids, 13 to 14 amino acids, 8 to 13 amino acids, 8 to 12 amino acids, 8 to 11 amino acids, 8 to 10 amino acids, 8 to 9 amino acids, 9 to 13 amino acids, 9 to 12 amino acids, 9 to 11 amino acids, 9 to 10 amino acids, 10 to 13 amino acids, 10 to 12 amino acids, 10 to 11 amino acids, 11 to 13 amino acids, 11 to 12 amino acids, or 12 to 13 amino acids. (Item 76) The plurality of peptide source search steps includes a linear human proteome search for peptide sequences within an extended human proteome database, The extended human proteome database includes computer-readable representations of translation from messenger ribonucleic acid (RNA) and non-coding RNA, The non-transitory computer-readable medium according to item 71. (Item 77) The non-transitory computer-readable medium according to item 76, wherein the extended human proteome database includes computer-readable representations of translation from microRNA. (Item 78) The non-transitory computer-readable medium according to item 76, wherein the extended human proteome database includes computer-readable representations of translation from long non-coding RNA. (Item 79) The non-transitory computer-readable medium according to item 76, wherein the extended human proteome database includes computer-readable representations of translation from human endogenous retroviruses. (Item 80) The non-transitory computer-readable medium according to item 76, wherein the plurality of peptide source search steps includes a linear human genome search of translation of the human genome database. (Item 81) The non-transitory computer-readable medium according to item 80, wherein the linear human genome search excludes the portion of the human genome from which the messenger RNA and the non-coding RNA of the extended human proteome database are transcribed, and includes the remaining portion of the human genome. (Item 82) The plurality of peptide source search steps includes a linear mismatch search for peptides having a mismatch to the peptide sequence in the extended human proteome database, The extended human proteome database includes computer-readable representations of translation from messenger ribonucleic acid (RNA) and non-coding RNA, The non-transitory computer-readable medium according to item 71. (Item 83) The plurality of peptide source search steps includes a linear extrinsic search for the peptide sequences in an extrinsic proteome database, wherein the extrinsic proteome database includes a computer-readable representation of proteins translated from RNA derived from an extrinsic organism and / or proteins synthesized by an extrinsic organism, The non-transitory computer-readable medium according to item 71. (Item 84) The non-transitory computer-readable medium according to item 83, wherein the extrinsic proteome database includes a Basic Local Alignment Search Tool (BLAST) database. (Item 85) The plurality of peptide source search steps includes a cis-splicing search for peptide fragments that can be cis-spliced to match the peptide sequence in an extended human proteome database, wherein the extended human proteome database includes a computer-readable representation of translations from messenger ribonucleic acid (RNA) and non-coding RNA, The non-transitory computer-readable medium according to item 71. (Item 86) The plurality of peptide source search steps includes a trans-splicing search for peptide fragments that can be trans-spliced to match the peptide sequence in an extended human proteome database, wherein the extended human proteome database includes a computer-readable representation of translations from messenger ribonucleic acid (RNA) and non-coding RNA, The non-transitory computer-readable medium according to item 71. (Item 87) The non-transitory computer-readable medium according to item 71, wherein when a peptide is not assigned a peptide source by any of the plurality of peptide source search steps, the peptide source assignment workflow ends with the unassigned peptide. (Item 88) The peptide source assignment workflow performs the following searches in sequential order, as follows: A linear human proteome search for the peptide sequence in the extended human proteome database, A linear human genome search of the translation of the human genome database, A linear mismatch search for peptides having a mismatch to the peptide sequence in the extended human proteome database, and A cis-splice search for peptide fragments that can be cis-spliced to match the peptide sequence in the extended human proteome database The non-transitory computer-readable medium according to item 71, comprising: (Item 89) The peptide source assignment workflow includes a linear non-endogenous search for the peptide sequence in a non-endogenous proteome database, The linear non-endogenous search is sequentially ordered in the peptide assignment workflow after the linear mismatch search and before the cis-splice search, The non-transitory computer-readable medium according to item 88. (Item 90) The peptide source assignment workflow includes a trans-splice search for peptide fragments that can be trans-spliced to match the peptide sequence in the extended human proteome database, The trans-splice search is sequentially ordered in the peptide assignment workflow after the cis-splice search, The non-transitory computer-readable medium according to item 88. (Item 91) Creating a plurality of simulation random queries, Determining the number of matches associated with each source based on applying the plurality of simulation random queries to each source of a plurality of sources, Determining a false discovery rate associated with each source based on the number of matches associated with each source, and Creating a query support data structure configured to facilitate application of new queries to the plurality of sources based on the false discovery rate A method comprising. (Item 92) Creating the plurality of simulation random queries is Creating a plurality of uniform random queries, or Creating a plurality of weighted random queries The method according to item 91, comprising at least one of. (Item 93) The method according to item 91, wherein the plurality of simulation random queries include a plurality of simulation random text strings. (Item 94) The method according to item 91, wherein the plurality of simulation random queries include a plurality of simulation random peptide sequences. (Item 95) Determining the false discovery rate associated with each source based on the number of matches associated with each source includes a function of the number of matches and the number of the plurality of simulation random queries, according to the method of item 91. (Item 96) Determining the false discovery rate associated with each source based on the number of matches associated with each source includes dividing the number of matches by the number of the plurality of simulation random queries, according to the method of item 91. (Item 97) One or more processors, and When executed by the one or more processors, causing the apparatus to Create a plurality of simulation random queries, Based on applying the plurality of simulation random queries to each source of a plurality of sources, determine the number of matches associated with each source, Based on the number of matches associated with each source, determine the false discovery rate associated with each source, and Create a query support data structure configured to facilitate the application of new queries to the plurality of sources based on the false discovery rate. Memory-storable processor-executable instructions An apparatus comprising. (Item 98) The processor-executable instructions that cause the apparatus to create the plurality of simulation random queries further cause the apparatus to Create a plurality of uniform random queries, or Create a plurality of weighted random queries Causing at least one of. The apparatus according to item 97. (Item 99) The apparatus according to item 97, wherein the plurality of simulation random queries include a plurality of simulation random text strings. (Item 100) The apparatus according to item 97, wherein the plurality of simulation random queries include a plurality of simulation random peptide sequences. (Item 101) The processor-executable instructions that cause the apparatus to determine the false discovery rate associated with each source based on the number of matches associated with each source further cause the apparatus to determine the false discovery rate as a function of the number of matches and the number of the plurality of simulation random queries. The apparatus according to item 97. (Item 102) The apparatus according to item 97, wherein the processor-executable instructions further cause the apparatus to determine the false discovery rate by dividing the number of matches by the number of the plurality of simulation random queries. (Item 103) When executed by a processor, cause the processor to create a plurality of simulation random queries, determine the number of matches associated with each source based on applying the plurality of simulation random queries to each source of a plurality of sources, determine a false discovery rate associated with each source based on the number of matches associated with each source, and create a query support data structure configured to facilitate the application of new queries to the plurality of sources based on the false discovery rate, One or more non-transitory computer-readable media storing processor-executable instructions therefor. (Item 104) The processor-executable instructions that cause the processor to create the plurality of simulation random queries further cause the processor to create a plurality of uniform random queries, or create a plurality of weighted random queries to cause at least one of The one or more non-transitory computer-readable media according to item 103. (Item 105) The one or more non-transitory computer-readable media according to item 103, wherein the plurality of simulation random queries include a plurality of simulation random text strings. (Item 106) The one or more non-transitory computer-readable media according to item 103, wherein the plurality of simulation random queries include a plurality of simulation random peptide sequences. (Item 107) The processor-executable instructions that cause the processor to determine the false discovery rate associated with each source based on the number of matches associated with each source, further cause the processor to determine the false discovery rate as a function of the number of matches and the number of the plurality of simulation random queries, one or more non-transitory computer-readable media according to item 103. (Item 108) The processor-executable instructions that cause the processor to determine the false discovery rate associated with each source based on the number of matches associated with each source, further cause the processor to determine the false discovery rate by dividing the number of matches by the number of the plurality of simulation random queries, one or more non-transitory computer-readable media according to item 103. (Item 109) Create a plurality of simulation random queries, Determine the number of matches associated with each source based on applying the plurality of simulation random queries to each source of the plurality of sources, Determine the false discovery rate associated with each source based on the number of matches associated with each source, and Create a query support data structure configured to facilitate the application of new queries to the plurality of sources based on the false discovery rate A computer device configured as such, and Receive the plurality of simulation random queries, Determine whether there is a number of matches for the plurality of simulation random queries, and Output a result indicating the number of matches The plurality of sources configured as such A system including. (Item 110) The computer device configured to create the plurality of simulation random queries causes the processor to create a plurality of uniform random queries, or create a plurality of weighted random queries to cause at least one of and is further configured to The system according to item 109. (Item 111) The system according to item 109, wherein the plurality of simulation random queries includes a plurality of simulation random text strings. (Item 112) The system according to item 109, wherein the plurality of simulation random queries includes a plurality of simulation random peptide sequences. (Item 113) The computer device configured to determine the false discovery rate associated with each source based on the number of matches associated with each source is further configured to determine the false discovery rate as a function of the number of matches and the number of the plurality of simulation random queries, the system according to item 109. (Item 114) The computer device configured to determine the false discovery rate associated with each source based on the number of matches associated with each source is further configured to determine the false discovery rate by dividing the number of matches by the number of the plurality of simulation random queries, the system according to item 109. (Item 115) Receiving a query, Applying the query to one or more of a plurality of sources based on a query support data structure, Determining a label associated with the source of the plurality of sources associated with the query result based on the query result, and Applying the label to the query A method comprising (Item 116) The method according to item 115, wherein the query includes a text string. (Item 117) The method according to item 115, wherein the query includes a peptide sequence. (Item 118) The method according to item 117, wherein receiving the query includes receiving the peptide sequence from a mass spectrometry system. (Item 119) The method according to item 115, further comprising determining one or more amino acids of the peptide sequence by a mass spectrometry system. (Item 120) The method according to item 115, wherein the query support data structure indicates an order of the plurality of sources for applying the query, and the order is based on a false discovery rate associated with each source of the plurality of sources. (Item 121) The method according to item 115, further comprising determining one or more permutations of the query. (Item 122) Applying the query to the one or more sources of the plurality of sources based on the query support data structure includes Applying each permutation of the one or more permutations of the query to the one or more sources of the plurality of sources, When an identical match for the one or more permutations of the query is found in a first source of the plurality of sources, stopping an additional search, applying a linear label to the one or more permutations of the query associated with the identical match, and Assigning the one or more permutations of the query associated with the identical match as the correct query The method according to item 121, including (Item 123) Applying the query to one or more sources of a plurality of sources includes Searching for identical matches to the query in the first source of the plurality of sources, and Stopping additional searches when an identical match to the query is found in the first source of the plurality of sources The method according to item 115, comprising. (Item 124) The method according to item 123, wherein the query result includes the identical match and the identifier associated with the source of the plurality of sources associated with the query result includes a linear identifier. (Item 125) Applying the query to one or more of the plurality of sources is Searching for identical matches to the query in the first source of the plurality of sources, and Stopping additional searches when an identical match to the one or more permutations of the query is found in the first source of the plurality of sources The method according to item 115, comprising. (Item 126) The method according to item 125, wherein the query result includes the identical match and the identifier associated with the source of the plurality of sources associated with the query result includes a linear identifier. (Item 127) Applying the query to one or more of the plurality of sources is Searching for identical matches to the query in any frame of the plurality of frames of the second source of the plurality of sources, and Stopping additional searches when an identical match to the query is found in any frame of the plurality of frames of the second source of the plurality of sources The method according to item 125, comprising. (Item 128) The method according to item 127, wherein the query result includes the same match, and the label associated with the source of the plurality of sources associated with the query result includes a linear label. (Item 129) Applying the query to one or more of a plurality of sources includes searching for non-identical matches for the query in a third source of the plurality of sources, and stopping additional searches when a non-identical match for the query is found in the third source of the plurality of sources. The method according to item 127, which includes the above. (Item 130) The method according to item 129, wherein the query result includes the non-identical match, and the label associated with the source of the plurality of sources associated with the query result includes a mismatch label. (Item 131) Applying the query to one or more of a plurality of sources includes searching for homologous matches for the query in a fourth source of the plurality of sources, and stopping additional searches when a homologous match for the query is found in the fourth source of the plurality of sources. The method according to item 129, which includes the above. (Item 132) The method according to item 131, wherein the query result includes the homologous match, and the label associated with the source of the plurality of sources associated with the query result includes a homologous label. (Item 133) Applying the query to one or more of a plurality of sources includes dividing the query into a set of multiple fragments, searching for each set of fragments in a fifth source of the plurality of sources, When a match for the set of fragments is found in the fifth source of the plurality of sources, stop additional searches, and When a first match for the first fragment of the set of fragments and a second match for the second fragment of the set of fragments are found in the fifth source of the plurality of sources, stop additional searches The method according to item 131, comprising: (Item 134) The method according to item 133, wherein the query result includes the match for the set of fragments, and the label associated with the source of the plurality of sources associated with the query result includes a cis-splice label. (Item 135) The method according to item 133, wherein the query result includes the first match for the first fragment of the set of fragments and the second match for the second fragment of the set of fragments, and the label associated with the source of the plurality of sources associated with the query result includes a trans-splice label. (Item 136) The method according to item 115, further comprising determining the source of the query based on the label. (Item 137) The method according to item 136, further comprising verifying the output of the mass spectrometry system based on the source of the query. (Item 138) One or more processors, and When executed by the one or more processors, cause the apparatus to Receive a query, Apply the query to one or more sources of a plurality of sources based on a query support data structure, Determine a label associated with the source of the plurality of sources associated with the query result based on the query result, and Apply the label to the query, Memory-stored processor-executable instructions An apparatus comprising (Item 139) The apparatus according to item 138, wherein the query includes a text string. (Item 140) The apparatus according to item 138, wherein the query includes a peptide sequence. (Item 141) The apparatus according to item 138, wherein the processor-executable instructions that cause the apparatus to receive the query further cause the apparatus to receive a peptide sequence from a mass spectrometry system. (Item 142) The apparatus according to item 138, wherein the processor-executable instructions further cause the apparatus to determine one or more amino acids of the peptide sequence by the mass spectrometry system. (Item 143) The apparatus according to item 138, wherein the query support data structure indicates an order of the plurality of sources for applying the query, and the order is based on a false discovery rate associated with each source of the plurality of sources. (Item 144) The apparatus according to item 138, wherein the processor-executable instructions further cause the apparatus to determine one or more permutations of the query. (Item 145) The processor-executable instructions that cause the apparatus to apply the query to the one or more sources of the plurality of sources based on the query support data structure further cause the apparatus to apply each permutation of the one or more permutations of the query to the one or more sources of the plurality of sources, stop an additional search when an identical match for the one or more permutations of the query is found in a first source of the plurality of sources, apply a linear label to the one or more permutations of the query associated with the identical match, and assign the one or more permutations of the query associated with the identical match as the correct query. The apparatus according to item 144. (Item 146) The processor-executable instructions that cause the apparatus to apply the query to one or more of a plurality of sources further cause the apparatus to search for an identical match for the query in a first source of the plurality of sources, and stop additional searches when an identical match for the query is found in the first source of the plurality of sources. The apparatus according to item 138. (Item 147) The apparatus according to item 146, wherein the query result includes the identical match, and the identifier associated with the source of the plurality of sources associated with the query result includes a linear identifier. (Item 148) The processor-executable instructions that cause the apparatus to apply the query to one or more of a plurality of sources further cause the apparatus to search for an identical match for the query in a first source of the plurality of sources, and stop additional searches when an identical match for one or more permutations of the query is found in the first source of the plurality of sources. The apparatus according to item 138. (Item 149) The apparatus according to item 148, wherein the query result includes the identical match, and the identifier associated with the source of the plurality of sources associated with the query result includes a linear identifier. (Item 150) The processor-executable instructions that cause the apparatus to apply the query to one or more of a plurality of sources further cause the apparatus to search for an identical match for the query in any frame of a plurality of frames of a second source of the plurality of sources, and When the same match for the query is found in any frame of the plurality of frames of the second source of the plurality of sources, stop additional searches. The apparatus according to item 148. (Item 151) The apparatus according to item 150, wherein the query result includes the same match and the label associated with the source of the plurality of sources associated with the query result includes a linear label. (Item 152) The processor-executable instructions that cause the apparatus to apply the query to one or more of the plurality of sources further cause the apparatus to search for a non-identical match for the query in a third source of the plurality of sources, and stop additional searches when a non-identical match for the query is found in the third source of the plurality of sources. The apparatus according to item 150. (Item 153) The apparatus according to item 152, wherein the query result includes the non-identical match and the label associated with the source of the plurality of sources associated with the query result includes a mismatch label. (Item 154) The processor-executable instructions that cause the apparatus to apply the query to one or more of the plurality of sources further cause the apparatus to search for a homologous match for the query in a fourth source of the plurality of sources, and stop additional searches when a homologous match for the query is found in the fourth source of the plurality of sources. The apparatus according to item 152. (Item 155) The apparatus according to item 154, wherein the query result includes the homologous match and the label associated with the source of the plurality of sources associated with the query result includes a homologous label. (Item 156) The processor-executable instructions that cause the device to apply the query to one or more of a plurality of sources further cause the device to split the query into a set of fragments, search for each set of fragments in a fifth source of the plurality of sources, stop additional searching if a match for the set of fragments is found in the fifth source of the plurality of sources, and stop additional searching if a first match for a first fragment of the set of fragments and a second match for a second fragment of the set of fragments are found in the fifth source of the plurality of sources. The device according to item 154. (Item 157) The device according to item 156, wherein the query result includes the match for the set of fragments, and the identifier associated with the source of the plurality of sources associated with the query result includes a cis-splice identifier. (Item 158) The device according to item 156, wherein the query result includes the first match for the first fragment of the set of fragments and the second match for the second fragment of the set of fragments, and the identifier associated with the source of the plurality of sources associated with the query result includes a trans-splice identifier. (Item 159) The device according to item 138, wherein the processor-executable instructions further cause the device to determine a source of the query based on the identifier. (Item 160) The device according to item 159, wherein the processor-executable instructions further cause the device to verify an output of a mass spectrometry system based on the source of the query. (Item 161) When executed by a processor, cause the processor to receive a query, Based on the query support data structure, apply the query to one or more of the plurality of sources, Based on the query result, determine a label associated with the source of the plurality of sources associated with the query result, and Apply the label to the query, One or more non-transitory computer-readable media storing processor-executable instructions therefor. (Item 162) The one or more non-transitory computer-readable media according to item 161, wherein the query includes a text string. (Item 163) The one or more non-transitory computer-readable media according to item 161, wherein the query includes a peptide sequence. (Item 164) The processor-executable instructions that cause the processor to receive the query further cause the processor to receive a peptide sequence from a mass spectrometry system, the one or more non-transitory computer-readable media according to item 161. (Item 165) The processor-executable instructions further cause the processor to determine one or more amino acids of the peptide sequence by the mass spectrometry system, the one or more non-transitory computer-readable media according to item 161. (Item 166) The query support data structure indicates an order of the plurality of sources for applying the query, and the order is based on a false discovery rate associated with each source of the plurality of sources, the one or more non-transitory computer-readable media according to item 161. (Item 167) The processor-executable instructions further cause the processor to determine one or more permutations of the query, the one or more non-transitory computer-readable media according to item 161. (Item 168) The processor-executable instructions that cause the processor to apply the query to the one or more sources of the plurality of sources based on the query support data structure further cause the processor to apply each of the one or more permutations of the query to the one or more sources of the plurality of sources, stop additional searching when an identical match for the one or more permutations of the query is found in a first source of the plurality of sources, apply a linear label to the one or more permutations of the query associated with the identical match, and assign the one or more permutations of the query associated with the identical match as the correct query. One or more non-transitory computer-readable media according to item 167. (Item 169) The processor-executable instructions that cause the processor to apply the query to the one or more sources of the plurality of sources further cause the processor to search for an identical match for the query in a first source of the plurality of sources, and stop additional searching when an identical match for the query is found in the first source of the plurality of sources. One or more non-transitory computer-readable media according to item 161. (Item 170) One or more non-transitory computer-readable media according to item 169, wherein the query result includes the identical match and the label associated with the source of the plurality of sources associated with the query result includes a linear label. (Item 171) The processor-executable instructions that cause the processor to apply the query to the one or more sources of the plurality of sources further cause the processor to Cause a search for the same match for the query in the first source of the plurality of sources, and Stop additional searches when the same match for the one or more permutations of the query is found in the first source of the plurality of sources. One or more non-transitory computer-readable media according to item 161. (Item 172) One or more non-transitory computer-readable media according to item 171, wherein the query result includes the same match and the identifier associated with the source of the plurality of sources associated with the query result includes a linear identifier. (Item 173) The processor-executable instructions that cause the processor to apply the query to one or more sources of a plurality of sources further cause the processor to Cause a search for the same match for the query in any frame of a plurality of frames of a second source of the plurality of sources, and Stop additional searches when the same match for the query is found in any frame of the plurality of frames of the second source of the plurality of sources. One or more non-transitory computer-readable media according to item 171. (Item 174) One or more non-transitory computer-readable media according to item 173, wherein the query result includes the same match and the identifier associated with the source of the plurality of sources associated with the query result includes a linear identifier. (Item 175) The processor-executable instructions that cause the processor to apply the query to one or more sources of a plurality of sources further cause the processor to Cause a search for a non-identical match for the query in a third source of the plurality of sources, and When a non-identical match for the query is found in the third source of the plurality of sources, stop additional searches. One or more non-transitory computer-readable media according to item 173. (Item 176) One or more non-transitory computer-readable media according to item 175, wherein the query result includes the non-identical match and the label associated with the source of the plurality of sources associated with the query result includes a mismatch label. (Item 177) The processor-executable instructions that cause the processor to apply the query to one or more of the plurality of sources further cause the processor to search for an identical match for the query in a fourth source of the plurality of sources, and stop additional searches when an identical match for the query is found in the fourth source of the plurality of sources. One or more non-transitory computer-readable media according to item 176. (Item 178) One or more non-transitory computer-readable media according to item 177, wherein the query result includes the identical match and the label associated with the source of the plurality of sources associated with the query result includes an identical label. (Item 179) The processor-executable instructions that cause the processor to apply the query to one or more of the plurality of sources further cause the processor to split the query into a set of fragments, search for each set of fragments in a fifth source of the plurality of sources, stop additional searches when a match for the set of fragments is found in the fifth source of the plurality of sources, and When a first match for the first fragment of the set of fragments and a second match for the second fragment of the set of fragments are found in the fifth source of the plurality of sources, stop additional searches. One or more non-transitory computer-readable media according to item 177. (Item 180) One or more non-transitory computer-readable media according to item 179, wherein the query result includes the match for the set of fragments, and the label associated with the source of the plurality of sources associated with the query result includes a cis-splice label. (Item 181) One or more non-transitory computer-readable media according to item 179, wherein the query result includes the first match for the first fragment of the set of fragments and the second match for the second fragment of the set of fragments, and the label associated with the source of the plurality of sources associated with the query result includes a trans-splice label. (Item 182) One or more non-transitory computer-readable media according to item 161, wherein the processor-executable instructions further cause the processor to determine the source of the query based on the label. (Item 183) One or more non-transitory computer-readable media according to item 182, wherein the processor-executable instructions further cause the processor to verify the output of a mass spectrometry system based on the source of the query. (Item 184) Receive a query, Apply the query to one or more of a plurality of sources based on a query support data structure, Determine a label associated with the source of the plurality of sources associated with the query result based on the query result, and Apply the label to the query A computer device configured to, and receives the query, and one or more of the plurality of sources configured to determine the query result A system comprising (Item 185) The system according to item 184, wherein the query includes a text string (Item 186) The system according to item 184, wherein the query includes a peptide sequence (Item 187) The system according to item 184, wherein the computer device configured to receive the query is further configured to receive a peptide sequence from a mass spectrometry system (Item 188) The system according to item 184, wherein the computer device is further configured to cause the processor to determine one or more amino acids of the peptide sequence by the mass spectrometry system (Item 189) The system according to item 184, wherein the query support data structure indicates an order of the plurality of sources for applying the query, and the order is based on a false discovery rate associated with each of the plurality of sources (Item 190) The system according to item 184, wherein the computer device is further configured to determine one or more permutations of the query (Item 191) The computer device configured to apply the query to one or more of the plurality of sources based on the query support data structure applies each of the one or more permutations of the query to one or more of the plurality of sources, If an identical match for the one or more permutations of the query is found in a first source of the plurality of sources, stop additional searches, apply a linear label to the one or more permutations of the query associated with the identical match, and assign the one or more permutations of the query associated with the identical match as the correct query The system according to item 184, further configured as described above. (Item 192) The computer device configured to cause the processor to apply the query to one or more sources of a plurality of sources, searches for an identical match for the query in a first source of the plurality of sources, and stops additional searches if an identical match for the query is found in the first source of the plurality of sources The system according to item 184, further configured as described above. (Item 193) The system according to item 192, wherein the query result includes the identical match and the label associated with the source of the plurality of sources associated with the query result includes a linear label. (Item 194) The computer device configured to apply the query to one or more sources of a plurality of sources, searches for an identical match for the query in a first source of the plurality of sources, and stops additional searches if an identical match for the one or more permutations of the query is found in the first source of the plurality of sources The system according to item 184, further configured as described above. (Item 195) The system according to item 194, wherein the query result includes the identical match and the label associated with the source of the plurality of sources associated with the query result includes a linear label. (Item 196) The computer device configured to apply the query to one or more of a plurality of sources, searches for identical matches to the query in any frame of a plurality of frames of a second source of the plurality of sources, and stops additional searches when an identical match to the query is found in any frame of the plurality of frames of the second source of the plurality of sources. The system according to item 194, further configured as such. (Item 197) The system according to item 196, wherein the query result includes the identical match, and the identifier associated with the source of the plurality of sources associated with the query result includes a linear identifier. (Item 198) The computer device configured to apply the query to one or more of a plurality of sources, searches for non-identical matches to the query in a third source of the plurality of sources, and stops additional searches when a non-identical match to the query is found in the third source of the plurality of sources. The system according to item 196, further configured as such. (Item 199) The system according to item 198, wherein the query result includes the non-identical match, and the identifier associated with the source of the plurality of sources associated with the query result includes a mismatch identifier. (Item 200) The computer device configured to apply the query to one or more of a plurality of sources, searches for homologous matches to the query in a fourth source of the plurality of sources, and stops additional searches when a homologous match to the query is found in the fourth source of the plurality of sources. The system according to item 198, further configured as follows. (Item 201) The system according to item 200, wherein the query result includes the homologous match, and the identifier associated with the source of the plurality of sources associated with the query result includes a homologous identifier. (Item 202) The computer device configured to apply the query to one or more of a plurality of sources divides the query into a set of fragments, searches for each set of fragments in a fifth source of the plurality of sources, stops additional searches when a match for the set of fragments is found in the fifth source of the plurality of sources, and stops additional searches when a first match for a first fragment of the set of fragments and a second match for a second fragment of the set of fragments are found in the fifth source of the plurality of sources The system according to item 200, further configured as follows. (Item 203) The system according to item 202, wherein the query result includes the match for the set of fragments, and the identifier associated with the source of the plurality of sources associated with the query result includes a cis-splice identifier. (Item 204) The system according to item 202, wherein the query result includes the first match for the first fragment of the set of fragments and the second match for the second fragment of the set of fragments, and the identifier associated with the source of the plurality of sources associated with the query result includes a trans-splice identifier. (Item 205) The system according to item 184, wherein the processor-executable instructions further cause the processor to determine the source of the query based on the identifier. (Item 206) The system according to item 205, wherein the computer device is further configured to verify the output of the mass spectrometry system based on the source of the query.
Brief Description of Drawings
[0013]
Figure 1
[0014]
Figure 2
[0015]
Figure 3
[0016]
Figure 4
[0017]
Figure 5
[0018]
Figure 6
[0019]
Figure 7
[0020]
Figure 8
[0021]
Figure 9
[0022]
Figure 10
[0023]
Figure 11
[0024]
Figure 12
[0025]
Figure 13
[0026]
Figure 14
[0027]
Figure 15
[0028]
Figure 16
[0029]
Figure 17
[0030]
Figure 18
[0031]
Figure 19
[0032]
Figure 20
[0033]
Figure 21
[0034]
Figure 22
[0035]
Figure 23
Figure 24
[0036] Detailed Description The disclosed method and system can be more readily understood by reference to the following detailed description of specific embodiments, examples, and drawings, and the descriptions before and after them, included herein.
[0037] It is to be understood that the methods and systems disclosed are not limited to the specific methodologies, protocols, and reagents described, as these may vary. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the invention, which is defined solely by the appended claims.
[0038] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to "a peptide" includes a plurality of such peptides, a reference to "the peptide" is a reference to one or more peptides known to those skilled in the art and their equivalents, and so forth.
[0039] The term "peptide" may be used interchangeably with "polypeptide" and refers to a polymeric form of amino acids of any length that may include genetically encoded and non-genetically encoded amino acids, chemically or biochemically modified or derivatized amino acids, and peptides having modified peptide backbones. In some embodiments, the term "peptide" refers to a string of two or more naturally occurring amino acids.
[0040] "Optional" or "optionally" means that the subsequent recited event, circumstance, or material may or may not occur, or may or may not be present, and that the description includes instances where the event, circumstance, or material occurs or is present and instances where it does not occur or is not present.
[0041] Throughout the description and claims of this specification, the word "comprise" and its variants, such as "comprising" and "comprises", mean "including but not limited to", and are not intended to exclude, for example, other additives, components, integers or steps. In particular, in a method described as including one or more steps or operations, it is specifically contemplated that each step includes what is listed (unless the step includes a limiting term such as "consisting of"), meaning that each step is not intended to exclude, for example, other additives, components, integers or steps not listed in the step.
[0042] As used herein, the term "computer-readable representation of a protein sequence" may include the protein itself, a gene sequence (e.g., DNA, RNA) from which the protein sequence may be derived through processes understood by one of ordinary skill in the relevant art (e.g., transcription, translation), and / or a sequence listing of a portion thereof. Similarly, as used herein, the term "computer-readable representation of translation from ribonucleic acid (RNA)" may include a protein or peptide that can be translated (at least theoretically) from RNA as understood by one of ordinary skill in the relevant art, the gene sequence of the RNA, the gene sequence of DNA from which the RNA can be transcribed (at least theoretically) as understood by one of ordinary skill in the relevant art, and / or a sequence listing of a portion thereof. As used herein, the term "computer-readable representation of translation from RNA" may refer to a specific type of RNA, including messenger RNA (mRNA), non-coding RNA, long non-coding RNA, microRNA, and other types of RNA understood by one of ordinary skill in the relevant art. A computer-readable representation of translation from a specific type of RNA may include a protein or peptide that can be translated (at least theoretically) from the specific type of RNA as understood by one of ordinary skill in the relevant art, the gene sequence of the specific type of RNA, the gene sequence of DNA from which the specific type of RNA can be transcribed (at least theoretically) as understood by one of ordinary skill in the relevant art, and / or a sequence listing of a portion thereof.
[0043] The terms "random hit rate" and "false discovery rate" are used interchangeably herein and are understood to mean the frequency with which randomly generated inputs are found by searching a database.
[0044] "Individual" or "subject" or "animal" refers to a human, a veterinary animal (e.g., cat, dog, cow, horse, sheep, pig, etc.) and an experimental animal model of a disease (e.g., mouse, rat). In some embodiments, the subject is a human.
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the methods and compositions, but the particularly useful methods, devices, and materials are as described. Publications cited herein, and the materials cited therein, are specifically incorporated herein by reference. The content of this specification should not be construed as an admission that the invention is not entitled to antedate such disclosure by virtue of prior invention. No admission is made that any reference constitutes prior art. The discussion of references states what their authors assert, and applicants reserve the right to challenge the accuracy and appropriateness of the cited documents. Many publications are referenced herein, and it will be clearly understood that such references do not constitute an admission that any of these documents form part of the common general knowledge in the art.
[0046] FIG. 1 shows an example of system 100. System 100 can be used to analyze one or more portions of data / information, such as query information, and to analyze complete data / information and / or receive / obtain additional data / information associated with the data / information, in order to determine / identify a data source, such as an optimal data source and / or device.
[0047] The devices and / or components of system 100 may be connected and / or communicate with each other via network 106. Network 106 may be a public network, a private network, and / or a combination thereof. Network 106 may support any wired and / or wireless communication technology and / or technique. For example, network 106 may include and / or support a cellular phone network, a data network, a content delivery network, an optical fiber network, and / or any other type of network.
[0048] System 100 may include a user device 102 (e.g., a computer device, a client device, a smart device, etc.). The user device 102 may include a communication element 103 that provides an interface to the user for interacting with the user device 102 and / or any other device / component of the system 100. The communication element 103 can be any interface for presenting and / or receiving information to / from the user, such as user feedback. The interface may include a display and / or an interactive interface (e.g., a keyboard, a touch screen, a mouse, an audio controller, etc.). The interface may include a communication interface such as a web browser (e.g., Internet Explorer®, Mozilla Firefox®, Google Chrome®, Safari®). Other software, hardware, and / or interfaces may be used to provide communication between the user and one or more of the user devices 102 and / or any other device / component of the system 100. The communication element 103 may request various files or perform queries from local and / or remote sources, such as computer devices 107-112 and / or any other device / component of the system 100. The computer devices 107-112 may be located locally or remotely with respect to the user device 102.
[0049] Communication element 103 can transmit data to local or remote devices, such as computer devices 107-112, and / or any other device / component of system 100, via wired and / or wireless communication techniques. For example, communication element 103 may utilize any suitable wired communication technique, such as Ethernet (registered trademark), coaxial cable, fiber optic, etc. Communication element 103 may utilize any suitable long-distance communication technique, such as Wi-Fi (IEEE 802.11), BLUETOOTH (registered trademark), cellular phone, artificial satellite, infrared, etc. Communication element 103 may utilize any suitable short-distance communication technique, such as BLUETOOTH (registered trademark), near-field wireless communication, infrared, etc.
[0050] User device 102 can receive and / or analyze data / information, such as query information, etc. For example, user device 102 can receive data / information, query information, etc. via communication element 103. The data / information, query information, etc. can include any type of information, such as statistical queries, analytical queries, industry-specific queries (e.g., immune peptide mix-related queries, bioinformatics-related queries, biotechnology-related queries, healthcare-related queries, business-related queries, chemistry-based queries, mathematics-based queries, etc.).
[0051] User device 102 may include a query module 105 that can analyze data / information, such as query information, etc. Query module 105 may be software, hardware, and / or a combination of software and hardware. Query module 105 may be configured for natural language processing, syntax determination / analysis, query language (coding) processing / analysis, etc.
[0052] The user device 102 (e.g., query module 105) can receive / create queries. For example, the user device can receive / create queries such as "Was the health inspection score for XYZ Restaurant the same in 2019 as in 2020?" In another example, the user device can receive / create queries such as "What was the health inspection score for XYZ Restaurant in 2020?" The query module 105 can determine / identify the parts / components of the query, for example, using natural language processing, syntax determination / analysis, query language (coding) processing / analysis, etc. The parts / components of the query may include one or more data constraints, predicates, text strings, syntax elements, semantic sectors, etc. The query module 105 may combine the parts / components of the query, for example, to determine / create fixed phrases.
[0053] Query-based set expressions can be applied to data / information sources and / or systems to determine results and / or the accuracy of the results. The results can be aggregate values / amounts of data records, such as the number / quantity of matches, hits, correspondences, etc. between the parts / components of the query, as well as indicators of one or more data records stored by and / or associated with the source and / or system. The number / quantity of matches, hits, correspondences, etc. can be evaluated and / or compared against a threshold, for example, a data discovery threshold. When the number / quantity of matches, hits, correspondences, etc. meets / exceeds the discovery threshold, the query module 105 may create data records, provide indicators, and / or assign labels to the source and / or system. The label can indicate, for example, the type and / or quantity of matches, hits, correspondences, etc. associated with the source and / or system. The label can indicate any data / information regarding the source and / or system and / or the query applied to the corresponding result.
[0054] The user device 102 may evaluate the effectiveness of any source and / or system in order to output the results of a query. For example, the user device 102 (such as the query module 105, etc.) may send the query to one or more data sources and / or process the query based thereon. For simplicity and by way of example, the computer devices 107-112 may represent one or more data sources and / or one or more search engines. Although not shown, the computer devices 107-112 may each represent a plurality of related data sources, systems, devices, repositories, etc. For example, the computer devices 107-112 may each include and / or be associated with a database (such as a data store, data repository, etc.). The database may include any type of database, such as an Internet, in-memory / centralized database, distributed database, operational database, relational database, cloud-based database, object-oriented database, query language-based database (such as NoSQL, etc.), graph database, etc. The database may contain any data / information. In some embodiments, each of the computer devices 107-112 may represent different search engines configured to search the same database (such as the Internet).
[0055] To evaluate the effectiveness of computer devices 107-112 for outputting query results, user device 102 (e.g., query module 105, etc.) can apply one or more queries to one or more of computer devices 107-112 and determine the false discovery rate (FDR) associated with computer devices 107-112. For example, user device 102 (e.g., query module 105, etc.) can determine / create a plurality of random queries. The plurality of random queries can be, for example, uniform random queries, weighted random queries, and / or any other type of query. The plurality of simulation queries can be, for example, immune peptide mix-related queries and / or bioinformatics / biotechnology-related queries, such as queries associated with a plurality of simulated random peptide sequences. The plurality of simulation random queries can be created by any known technique. For example, one that creates random numbers / characters / words can be used to create a plurality of simulation random queries and / or test queries / cases. The quantity of simulation random queries can vary based on the type of query that can have an impact, such as the combination of simulation queries and / or the number of permutations. For example, the number of simulation queries for restaurants, airfares, etc. can be different from the number of simulation queries for DNA, RNA, and / or amino acid sequences. In certain embodiments, the number of simulation queries can be suppressed by the specific length of the simulation queries. For example, the simulation queries can be limited to the number of characters and / or words. In some embodiments, the number of simulation queries can be in the range anywhere from 10 queries to 10,000,000 queries, including those values. In some embodiments, the number of simulation queries can be, although not limited to, 10 queries to 1,000 queries. In some embodiments, the number of simulation queries can be, although not limited to, 10 queries to 10,000 queries.In some embodiments, the number of simulation queries, although not limited, can be from 10 queries to 100,000 queries. In some embodiments, the number of simulation queries, although not limited, can be from 10 queries to 1,000,000 queries. In some embodiments, the number of simulation queries, although not limited, can be from 100 queries to 1,000 queries. In some embodiments, the number of simulation queries, although not limited, can be from 100 queries to 10,000 queries. In some embodiments, the number of simulation queries, although not limited, can be from 100 queries to 100,000 queries. In some embodiments, the number of simulation queries, although not limited, can be from 100 queries to 1,000,000 queries. In some embodiments, the number of simulation queries, although not limited, can be from 1,000 queries to 100,000 queries. In some embodiments, the number of simulation queries, although not limited, can be from 10,000 queries to 100,000 queries. In some embodiments, the number of simulation queries, although not limited, can be 100,000 or more queries. In some embodiments, the number of simulation queries, although not limited, can be 1,000,000 or more queries. In some embodiments, the number of simulation queries can be at least 100,000 queries. In some embodiments, the number of simulation queries can be at least 1,000,000 queries.
[0056] For example, the query module 105 may create a plurality (e.g., dozens, hundreds, millions, etc.) of simulation random queries using an application, such as MySQL, and / or may test the queries / cases in a suitable grammar / format. The suitable grammar may be any grammar, language, syntax, coding, etc. that can be understood / executed by the query module 105. The query module 105 may use a query template to create queries in any suitable grammar / format. The query template may be created according to a scripting language. The query template may be mapped to and / or corresponding to a specific test case. The query module 105 may determine the results and / or expected results for the queries determined from the query template by applying the queries to a source and / or system, such as computer devices 107 - 112.
[0057] Based on the determined queries, the query module 105 may create / determine random queries from, for example, query templates. The query module 105 may apply the random queries to each of the computer devices 107 - 112 and determine which of the computer devices 107 - 112 outputs positive results and / or expected results. The output and / or expected results may be based on, for example, the ability of any given query of the computer devices 107 - 112 to process the meaning and / or syntax of the query and read out the data / information associated with the meaning and / or syntax. The user device 102 may determine / create a false discovery rate associated with each of the computer devices 107 - 112 based on the respective output of the computer devices 107 - 112.
[0058] In one aspect, the randomly generated queries may be incorrect, meaningless, and / or illogical queries designed to evaluate the false discovery rate of any source and / or system, such as computer devices 107-112. For example, a query such as "What is the price of a ticket for a plane to Dubai?" may be created using a query template. The query module 105 may determine / create incorrect, meaningless, and / or illogical versions and / or permutations of queries such as "What is the price of an apple until noon?", "When is the price to develop a plane?", "What currency is the ticket for a plane to Dubai?", "A plane ticket when the price is low". The incorrect, meaningless, and / or illogical versions and / or permutations of the query may be determined, for example, based on synonyms, phonetic relationships, etc. of the elements (predicates, constraints, conditions, metrics, parts, etc.) of the query. The incorrect, meaningless, and / or illogical versions and / or permutations of the query may be determined by changing the positions of the elements of the query. The incorrect, meaningless, and / or illogical versions and / or permutations of the query may be determined by any method.
[0059] Query module 105 can determine at what frequency computer devices 107 - 112 output results for incorrect, meaningless, and / or illogical versions and / or permutations of queries, such as multiple random queries. It may be shown at what frequency computer devices 107 - 112 output results for incorrect, meaningless, and / or illogical versions and / or permutations of queries, and / or this may correspond to the number of matches associated with each of computer devices 107 - 112. The false discovery rate (FDR) for any given computer device 107 - 112 may be determined as a function of the number of matches and the number of multiple random queries. Determining the FDR for computer devices 107 - 112 based on the number of matches associated with each of computer devices 107 - 112 may include dividing the number of matches by the number of multiple random queries. In some embodiments, determining the FDR may take into account a relevance score associated with the matches provided by computer devices 107 - 112. For example, a search engine may identify matches and assign a relevance score to the matches indicating how well the matches are associated with the query. Each search engine may use proprietary relevance scoring techniques. Matches may be added to the FDR determination when the simulation query returns a match having a relevance score exceeding a threshold.
[0060] The user device 102 can determine / create a query support data structure configured to facilitate the application of new queries to the computer devices 107-112 based on the false discovery rates associated with each of the computer devices 107-112. For example, in one embodiment, the computer devices 107-112 may include and / or be associated with a search engine (e.g., Google®, Yahoo®, Bing®, Firefox®, etc.) and / or similar data sources, data repositories, and / or data access systems.
[0061] FIG. 2 shows an example of a data structure 200 that can be used to facilitate the application of queries to the computer devices 107-112. The query support data structure 200 can indicate the order of the computer devices 107-112 (e.g., data sources and / or search engines).
[0062] The order of data sources may be based on the false discovery rate associated with each source. The query support data structure 200 may indicate one or more search techniques for one or more of the data sources 107-112. The query support data structure 200 may, for example, in column 202, indicate multiple search techniques for a single data source (e.g., data source 107, etc.), and the query support data structure 200 may indicate a single search technique for the data sources 107-112 and combinations thereof. The query support data structure 200 may, in column 201, include identifiers of the data sources 107-112 shown in an order according to the false discovery rate. The false discovery rate may be shown, for example, in column 203 as needed. A data source associated with a lower false discovery rate may be searched before a data source having a higher false discovery rate is searched. For each data source shown in the query support data structure 200, additional data may be included. The additional data may include one or more of the location of the data source, the query syntax, one or more query parameters, combinations thereof, and the like.
[0063] Queries may be labeled based on which data source returns the query results. The label may be an indicator of the source data / information associated with the query. For example, the label may indicate one or more levels of the accuracy of the results returned by the source based on the query. As another example, the label may indicate one or more of text data, multimedia data, statistical data, historical data, private / security-protected data, public data, and / or any other label of the type of data returned by the source based on the query.
[0064] As an example shown in FIG. 3, query 300 can be applied to one or more of a plurality of data sources 307-309 (e.g., search engines, data sources 107-112, computer devices 107-112, etc.). The replacement and / or version of query 300 can also be applied to the plurality of data sources 307-309. Query 300 can be, for example, "What is the price of a ticket for an airplane going to Dubai?" The replacement and / or version of query 300 can be, for example, "What is the airfare to Dubai?", "How much is the flight to Dubai?", "Dubai airfare", etc. The order in which query 300 is applied to the plurality of data sources 307-309 can be indicated by a query support data structure based on the false discovery rate associated with each of the plurality of data sources 307-309, as described herein. In one embodiment, as shown in FIG. 3, the data sources are ordered according to FDR. For example, the FDR for data source 307 can be about 1%, the FDR for data source 308 can be about 10%, and the FDR for data source 309 can be about 68%. The data source with the lowest false discovery rate can be searched first, and the data source with the highest false discovery rate can be searched last. In one embodiment, query 300 can be stopped at any point when returning search results. In one embodiment, query 300 can be applied to each of the plurality of data sources 307-309, and after query 300 is completed, the search results can be presented together with the display of the associated data source and FDR. In this way, the user can determine from the search results that they are more reliable and whether the user wants to apply a filter based on any FDR (e.g., remove search results associated with data sources having a high FDR value).
[0065] In addition, each of the plurality of data sources 307-309 may be associated with a threshold value, e.g., a data discovery threshold value applied to a match relevance score. The data discovery threshold value may be a threshold value defined by the system and / or a threshold value defined by the user. In one embodiment, a data source associated with a low false discovery rate may be associated with a low data discovery threshold value because the data source is generally associated with "good" results and any matches from the data source must be subject to less stringent relevance requirements. A data source associated with a high false discovery rate may be associated with a high data discovery threshold value because the data source is associated with less "good" results and any matches from the data source must be subject to more stringent relevance requirements. In another embodiment, a data source associated with a low false discovery rate may be associated with a high data discovery threshold value because the data source is generally associated with "good" results and is likely to contain relevant results. A data source associated with a high false discovery rate may be associated with a low data discovery threshold value because the data source is associated with less "good" results and a low data discovery threshold value may be required to determine relevant results. In one embodiment, the data discovery threshold value may be determined and / or set by the user, e.g., via a user interface.
[0066] In one embodiment, each data source of the plurality of data sources 307-309 may be associated with the same or different data discovery thresholds. For example, when a query is applied to a first data source, the first data discovery threshold may be determined such that a match exists only if the match has a relevance score greater than the first data discovery threshold (e.g., 85%). If the match does not satisfy the first data discovery threshold, the query may be applied to a second data source associated with a second data discovery threshold that determines that a match exists only if the match has a relevance score greater than the second data discovery threshold (e.g., 90%). If the match does not satisfy the second data discovery threshold, the query may be applied to a third data source associated with a third data discovery threshold that determines that a match exists only if the match has a relevance score greater than the third data discovery threshold (e.g., 95%). If the match does not satisfy the third data discovery threshold, the result is not output.
[0067] When applying query 300 to a data source, if a match that satisfies a data discovery threshold for query 300 (e.g., a threshold determined by the system, a threshold configurable by the user, etc.) is found in data source 307 and / or through it, the result may receive a first label (a very accurate result) at 312, and all relevant and / or possible results may be included in the output. Otherwise, query 300 may be applied to the next data source 308. If a match that satisfies the data discovery threshold (e.g., a threshold determined by the system, a threshold configurable by the user, etc.) is found, the result may receive a second label (a result that seems accurate) at 313, and all relevant and / or possible results may be included in the output. Otherwise, query 300 may be applied to the next data source 309. If a match that satisfies the data discovery threshold (e.g., a threshold determined by the system, a threshold configurable by the user, etc.) is found, the result may receive a third label (an accurate result) at 314, and all relevant and / or possible results may be included in the output. If no match is determined / identified, no result may receive a fourth label (no result) at 316.
[0068] Next, when applying exemplary embodiments of the disclosed methods and systems to de novo peptide sequencing, FIG. 4 shows a schematic of how linear, cis, and trans-spliced peptides are produced. For example, a linear peptide sequence matches like its parent protein, a fragment of a cis-spliced peptide is from the same protein, and a fragment of a trans-spliced peptide is from a different protein.
[0069] Figure 5 shows an example of system 500. An example of system 500 can be configured for mass spectrometry. Mass spectrometer 504 enables the accurate determination of the molecular masses of peptides and their sequences. For example, mass spectrometer 504 can output data / information, such as mass spectrometry data, which can be used for protein identification, de novo sequencing, and identification of post-translational modifications. System 500 can be configured to assign a source for de novo sequencing peptides.
[0070] Tandem mass spectrometry (MS / MS) has become an excellent high-throughput technique for protein identification. Tandem mass spectrometer 504 can be configured to ionize a mixture of peptides in sample 502 having different peptide sequences, measure the mass / charge ratios of their individual parent masses, selectively fragment each peptide into pieces, and measure the mass / charge ratios of the fragment ions. Tandem mass spectrometer 504 can be, as a non-limiting example, a linear ion trap mass spectrometer (LTQ) combined with a Fourier transform ion cyclotron resonance mass spectrometer (LTQ-FT). Thus, a tandem mass spectrum can be regarded as the collection of fragment masses from a single peptide. This collection or set of fragment masses or fragment mass values is the "fingerprint" for identifying the peptide. Therefore, the problem of peptide sequencing is to derive the sequence of a given peptide from their MS / MS spectra. For an ideal fragmentation process and an ideal mass spectrometer, the sequence of a peptide can be easily determined by converting the mass differences of consecutive ions in the spectrum into corresponding amino acids. This ideal situation would occur if the fragmentation process could be controlled such that each peptide is cleaved between every two consecutive amino acids and a single charge is retained only on the N-terminal piece. However, in practice, the fragmentation process in a mass spectrometer is not ideal.
[0071] The problem with tandem mass spectrometry peptide sequencing is, considering the spectrum S, the ion type Δ, and the mass m, to find the peptide of mass m that has the maximum match to the spectrum S. Peptide fragmentation in a tandem mass spectrometer can be characterized by a series of numbers Δ = {δ l ,..., δ k}. The δ-ion P’ ⊂ P of a partial peptide is a modification of P that has the mass m(P’) - δ. For tandem mass spectrometry, the theoretical spectrum of a peptide P can be calculated by subtracting all possible ion types {δ l ,..., δ k} from the masses of all partial peptides of P (i.e., every partial peptide gives rise to k masses in the theoretical spectrum). The (experimental) spectrum S = {s l ,..., s m} is the mass of a series of fragment ions. The match between the spectrum S and the peptide P is the number of masses that the experimental and theoretical spectra have in common.
[0072] The LTQ-FT mass spectrometer can generate spectra at a rate of 100,000 spectra per day per instrument. Software is important and a limiting factor in mass spectrometry proteomics analysis, and typical large datasets can require computation times of days or weeks on expensive computers or grids. Most peptide identification algorithms use a database search method that matches spectra against a protein database. FIG. 5 illustrates an exemplary process for spectral matching techniques for peptide identification. Specifically, a sample 502 is provided to a mass spectrometer 504. The mass spectrometer 504 may include any number of mass spectrometers in a tandem arrangement, for example, two mass spectrometers. Although a two-step process is illustrated, single-step processes are also known. In a first mass spectrometer 504A, peptide ions are selected such that a targeted component of specific mass is separated from the remainder of the sample. The targeted component is then activated or fragmented at 504B. In the case of peptides, the result will be a mixture of the ionized parent peptide (“precursor ion”) and lower mass component peptides that are ionized into various states. Several activation methods can be used, including collision with a neutral gas (also referred to as collision-induced dissociation). The parent peptide and its fragments are then provided to a second mass spectrometer 504C, which outputs the intensity and m / z for each of the plurality of fragments in the fragment mixture. This information can be output as a fragment mass spectrum 506. In the fragment mass spectrum 506, each fragment ion is represented as a bar graph, where the value on the horizontal axis indicates the mass-to-charge ratio (m / z) and the value on the vertical axis represents the intensity. The fragment mass spectrum 506 may take the form of mass spectrometry data.
[0073] Computer device 512 may be configured to analyze mass spectrometry data (e.g., fragment mass spectrum 506) created by mass spectrometer 504 to identify one or more amino acids based on a comparison with information contained within protein sequence library 508 of information derived from the mass spectrometry data. In some implementations, a user operating computer device 512 may access mass spectrometry data analyzer 514 that executes on computer device 512. In some implementations, the user provides mass spectrometry data created by mass spectrometer 504 to mass spectrometry data analyzer 514. In other implementations, the user selects mass spectrometry data from available mass spectrometry data (e.g., data previously downloaded, transferred, or otherwise made available to computer device 512 by mass spectrometer 504). In some implementations, mass spectrometer 504 includes computer device 512. For example, computer device 512 may be implemented as one or more computer processors operating within a mass spectrometry system. Each implementation is understood to describe additional embodiments of the methods and systems described herein.
[0074] In some implementations, mass spectrometry data analyzer 514 calculates additional data from the mass spectrometry data. For example, based on experimental information contained within the mass spectrometry data, the mass-to-charge ratio of an ion (e.g., calculated as the centroid of a peak in a so-called “profile” spectrum), the relative intensity of a peak, and / or the charge.
[0075] In certain embodiments, the sub-sequences contained in the protein sequence library 508 are used as a basis for predicting a plurality of mass spectra 510. The predicted mass spectra 510 of the sub-sequences are compared to the experimentally-derived fragment spectra 506 using the mass spectrometry data analyzer 514 of the computer device 512, and one or more of the predicted mass spectra that most closely match the experimentally-derived fragment spectra 506 can be identified.
[0076] In certain embodiments, de novo peptide sequencing may be performed using, for example, a spectral graph approach, where spectra are represented as graphs having peaks as vertices connected by edges when their mass differences correspond to the mass of amino acids. The vertices of the spectral graph are further scored based on peak intensity and neutral loss, and the peptide sequence is obtained by finding the longest path in the graph. De novo peptide sequencing can be considered as a search in a database of all possible peptides. For typical spectra identified in a database search, there can be hundreds, and even thousands, of very different peptide sequences that match the spectra. As a result, de novo peptide sequencing algorithms output the reconstruction of multiple peptides rather than a single reconstruction.
[0077] In some embodiments, the protein sequence library 508 may include a spectral dictionary that can be used to create full-length peptide reconstructions with a high probability of containing the correct peptide. However, an open question is how many reconstructions need to be created to avoid loss of the correct peptide. Creating too few peptides results in false negative errors, while creating too many peptides results in false positive errors. Some de novo algorithms output a single or fixed number (determined prior to the search) of peptides. For some spectra, creating only one reconstruction may be sufficient to guarantee finding the correct peptide, while in other cases (equivalent with the same parent mass), thousands of reconstructions may be insufficient. The problem of creating different numbers of reconstructions for each spectrum becomes particularly important for long-chain peptides due to the increased complexity of the search space.
[0078] The predicted peptide sequences obtained from the comparison of the mass spectrometry data by the mass spectrometry data analyzer 514 and the protein sequence library 508 may be provided to the query module 505. The query module 505 may be configured to identify the source of the peptide sequences using a plurality of data sources 518A-518N in communication with the query module via the network 520. The plurality of data sources 518A-518N may include any number and any type of data sources. Each of the plurality of data sources 518A-518N may include, and / or be associated with, a database (e.g., a data store, a data repository, etc.). The database may include any type of database, such as an in-memory / centralized database, a distributed database, a transactional database, a relational database, a cloud-based database, an object-oriented database, a query language-based database (e.g., NoSQL, etc.), a graph database, etc. The database may include any data / information, such as data / information associated with peptides, etc.
[0079] In some embodiments, data sources 518A - 518N may include an extended human proteome database. The extended human proteome database can include a computer-readable representation of protein sequences. The extended human proteome database can include a computer-readable representation of the translation of non-coding RNA. The extended human proteome database can include long non-coding RNA (lncRNA). The extended human proteome database can include microRNA (miRNA), which is a type of non-coding RNA. The extended human proteome database can include RNA transcribed from human endogenous retrovirus (HERV). The extended human proteome database can further include messenger RNA (mRNA) that classically encodes proteins. In some embodiments, at least a portion of the computer-readable representation of the protein sequences of the extended human proteome database can be associated with a specific subject, such that the workflow can assign a subject-specific putative source to de novo peptide sequences derived from the subject.
[0080] The extended human proteome database can include peptides derived from non-classically translated regions of the human genome, i.e., peptides derived from regions annotated as non-coding. The extended human proteome database can include one or more databases that contain some or all of OpenProt and / or data similar to some or all of OpenProt, as understood by those skilled in the relevant art. OpenProt is disclosed, for example, in Brunet M. A., Brunelle M., Lucier J.-F., Delcourt V., Levesque M., Grenier F., et al. (2019). OpenProt: A More Comprehensive Guide to Explore Eukaryotic Coding Potential and Proteomes. Nucleic Acids Res. 47, D403-D410. 10.1093 / nar / gky936, which is hereby incorporated by reference in its entirety. The extended human proteome database can include a computer-readable representation of protein sequences representing the translation of non-coding RNAs by including some or all of OpenProt and / or one or more databases that contain non-coding RNA sequences and / or their translations. OpenProt is a polycistronic model of the eukaryotic genome and includes all open reading frames (ORFs) at least 30 codons in length.
[0081] The extended human proteome database can include the translation of lncRNAs, i.e., non-classically translated regions of the human genome. lncRNAs were first characterized as mRNA-like non-coding RNAs in that they undergo splicing and have features such as poly(A) signals / tails, but any criterion of "transcripts longer than 200 nucleotides" was later added to its "definition". The extended human proteome database can include one or more databases containing some or all of NONCODE and / or data similar to some or all of NONCODE, as understood by those skilled in the relevant technical fields. NONCODE is disclosed, for example, in Bu, D. et al. NONCODE v3.0: Integrative annotation of long noncoding RNAs. Nucleic Acids Res. 40, D210-5 (2012), which is hereby incorporated by reference in its entirety. The extended human proteome database can include a computer-readable representation of protein sequences representing the translation of lncRNAs by including some or all of NONCODE and / or one or more databases containing lncRNA sequences and / or their translations.
[0082] The extended human proteome database can include the translation of microRNA (miRNA), a type of non-coding RNA having a length of about 22 bases. Typically, miRNAs regulate gene expression by blocking the translation of specific mRNAs and cause their degradation. The extended human proteome database can include part or all of miRBase and / or one or more databases containing data similar to part or all of miRBase, as understood by those skilled in the relevant art. miRBase is disclosed, for example, in Kozomara, A., Birgaoanu, M. & Griffiths-Jones, S. miRBase: from microRNA sequences to function. Nucleic Acids Res. 47, D155-D162 (2019), which is hereby incorporated by reference in its entirety. The extended human proteome database can include a computer-readable representation of protein sequences representing the translation of miRNAs by including part or all of the miRNAs and / or one or more databases containing miRNA sequences and / or their translations.
[0083] The extended human proteome database can include the transcription of HERVs, which are human genomic sequences corresponding to endogenous viral elements. The extended human proteome database can include one or more databases that contain some or all of the gEVE, and / or data similar to some or all of the gEVE, as understood by those skilled in the relevant technical fields. The gEVE is disclosed, for example, in Nakagawa, S. & Takahashi, M. U. gEVE: a genome-based endogenous viral element database provides comprehensive viral protein-coding sequences in mammalian genomes. Database (Oxford). (2016) doi:10.1093 / database / baw087, which is hereby incorporated by reference in its entirety. The extended human proteome database can include a computer-readable representation of protein sequences representing the translation of HERVs by including some or all of the gEVE, and / or one or more databases that contain HERV sequences and / or their translations.
[0084] An extended human proteome database can include mRNA by including some or all of UniProt and / or one or more databases containing data similar to some or all of UniProt, as understood by those skilled in the relevant art. The extended human proteome database can include UniProt to the extent that OpenProt, in an extended human proteome database that includes OpenProt, utilizes UniProt. Additionally, or alternatively, UniProt or a portion thereof can be included separately from OpenProt within the extended human proteome database. In a preferred embodiment, the extended human proteome database includes one or more databases containing reviewed UniProt and / or data similar to some or all of the reviewed UniProt, as understood by those skilled in the relevant art. In some embodiments, the extended human proteome database includes one or more databases containing unreviewed UniProt and / or data similar to some or all of the unreviewed UniProt, as understood by those skilled in the relevant art.
[0085] The extended proteome database can be stored in a single memory or distributed across multiple memories. The extended proteome database can include multiple completely different databases that can be queried as one database by a single query of a workflow, such as, but not limited to, the workflow illustrated in FIG. 7 and modifications thereof, and other workflows disclosed herein.
[0086] In one embodiment, data sources 518A-518N can include a human genomic database that includes all or a portion of the human genome from which a computer-readable representation of a protein can be computationally synthesized. The human genome includes approximately three billion base pairs of deoxyribonucleic acid (DNA) that make up the entire set of chromosomes of a human organism. The human genome includes the translated regions of DNA that encode all of the genes (20,000-25,000) of a human organism, as well as non-coding regions of DNA that are not encoded by any gene. In some embodiments, the human genomic database can include the entirety of the human genome, including the translated and non-coding regions of DNA. In some embodiments, the human genomic database can include the non-coding portion and / or frame reads of the human genome, excluding the portions of the human genome from which the mRNA and non-coding RNA of an extended human proteome database are transcribed and / or frame reads. In some embodiments, a protein can be computationally synthesized based on one, two, three, four, five, and / or six-frame translations of all or a portion of the human genome, such that a portion of the human genome may or may not be translated using the same number of frame reads as another portion of the human genome.
[0087] In one embodiment, data sources 518A-518N can include a non-endogenous proteome database that includes computer-readable representations of proteins and / or peptides that are sourced from non-endogenous sources to humans, including but not limited to bacterial sources, viral sources, and other organisms. In one embodiment, the non-endogenous proteome database can include one or more databases that include NCBI BLAST databases and / or data similar to a part or all of NCBI BLAST, as would be understood by one of ordinary skill in the relevant art. NCBI BLAST is disclosed, for example, in Johnson, M. et al. NCBI BLAST: a better web interface. Nucleic Acids Res. 36, W5-9 (2008), which is hereby incorporated by reference in its entirety. Data sources 518A-518N can include a computer-readable representation of a protein sequence representing a translation of a source non-endogenous to a human, by including some or all of NCBI BLAST and / or one or more databases containing such sequences and / or their translations.
[0088] In certain embodiments, data sources 518A-518N can include a computer-readable representation of a protein and / or peptide specifically associated with an individual subject. These subject-specific data can be incorporated into one or more databases disclosed herein (e.g., an extended human proteome database, a human genome database, a non-endogenous proteome database, etc.) and / or can be included in separate subject-specific databases.
[0089] Query module 505 can utilize query support data structure 516 to guide the identification process. Query support data structure 516 can indicate the order of search steps for multiple data sources applied to a query. The order can be based on the random hit rate associated with each search step. Query support data structure 516 can indicate one or more search techniques for one or more of the multiple data sources 518A-518N. Query support data structure 516 can indicate multiple search techniques for a single data source, and query support data structure 516 can indicate a single search technique for multiple data sources 518A-518N and combinations thereof.
[0090] The query support data structure 516 can include a peptide source assignment workflow for assigning putative sources to peptide sequence inputs to a workflow, where the putative source indicates the origin of the most likely peptide sequence. Each search step of the query support data structure 516 can include a peptide source search step that indicates the possible sources of an individual peptide sequence if the peptide source search step finds a match. The linear extended human proteome source can be indicated by a linear human proteome search for peptide sequences within the extended human proteome database. The linear genome source can be indicated by a linear human genome search of the translation of the human genome database. The linear mismatch can be indicated by a linear mismatch search for peptides having a mismatch to peptide sequences within the extended human proteome database, a linear mismatch search for peptides having a mismatch to peptides derived from the translation of the human genome, and / or a linear mismatch search of the target-specific database. The linear non-endogenous proteome source can be indicated by a linear non-endogenous search for peptide sequences within the non-endogenous proteome database. The cis-spliced human proteome source can be identified by a cis-splice search of the extended human proteome database. The trans-spliced human proteome source can be indicated by a trans-splice search of the extended human proteome database. The putative source assigned to a peptide sequence can be the source first found in the workflow, i.e., the search step has the lowest random hit rate.
[0091] FIG. 6 shows an example of a query support data structure 600. The query support data structure 600 may include search steps for searching data sources of a plurality of data sources 518A to 518N, shown in order by random hit rate. Search steps associated with a lower random hit rate may be searched before performing search steps having a higher random hit rate. For each search step shown in the query support data structure 600, additional search steps may be included and search steps may be omitted.
[0092] In one embodiment, the query support data structure 600 may have been previously created or may be created as needed. The query support data structure 600 can be created, for example, by creating a plurality of simulation random queries, determining the number of matches associated with each search step based on applying the plurality of simulation random queries to each search step, determining the random hit rate associated with each search step based on the number of matches associated with each search step, and creating a query support data structure configured to facilitate the application of new queries to a plurality of sources based on the random hit rate. The plurality of simulation random queries may include at least one of a plurality of uniform random queries or a plurality of weighted random queries. A uniform random query (e.g., a peptide sequence) can be created by uniformly randomly sampling all amino acids. A weighted random query (e.g., a peptide sequence) can be created by randomly sampling amino acids having a frequency of amino acids that matches the frequency of amino acids found in vertebrates. Determining the random hit rate associated with each search step based on the number of matches associated with each source may include a function of the number of matches and the number of simulation random queries. As a non-limiting example, the random hit rate associated with each source may be determined by dividing the number of matches by the number of simulation random queries. The random hit rate may further depend on the size and / or complexity of the data source being searched.
[0093] In one embodiment, mass spectrometry data may be used as, or processed and then used as, a query applied to one or more of a plurality of data sources 518A - 518N according to query support data structure 600. The query may be further processed before being applied to one or more of the plurality of data sources 518A - 518N. In one embodiment, one or more permutations of the query may be determined. For example, one or more permutations of a peptide sequence may be determined, and the one or more permutations may be used as a query in addition to the original query. For example, a peptide sequence provided as a query to the workflow of query support data structure 600 may include one or more ambiguous residues. For example, leucine (L) and isoleucine (I) have the same mass and thus cannot be distinguished in de novo search sequencing. To illustrate this, for a given peptide containing I / L, all permutations of the I and L residues may be considered such that the associated permuted peptide sequences are provided as queries to the workflow of query support data structure 600. For example, for the peptide "ATTSLLHN (SEQ ID NO: 1)", there are four possible permutations: ATTSLLHN (SEQ ID NO: 1), ATTSLIHN (SEQ ID NO: 2), ATTSILHN (SEQ ID NO: 3), and ATTSIIHN (SEQ ID NO: 4). Each permuted peptide sequence may be used as a query. An individual putative source can be assigned to each permuted peptide sequence according to the peptide source assignment workflow of query module 505. Next, the assigned putative source of the permutation is a possible source for the provided peptide sequence having ambiguous residues. The possible source indicated by the peptide source step with the lowest random hit rate can be assigned as the putative source of the provided peptide sequence having ambiguous residues. Further, the permutations of the provided peptide can be filtered to remove permutations for which no putative source is assigned.
[0094] Figure 7 is a flow diagram depicting an overview of steps of an example peptide assignment workflow. Using de novo sequenced peptide sequence 701, one or more replacements 702 of the de novo sequenced peptide sequence can be created.
[0095] In the first peptide source search step 703, the queries (701 and 702) can be applied to an extended human proteome database to identify identical matches. If an identical match is found for any replacement, the peptide sequence can be labeled "linear" at 704, and all possible protein sources of the peptide can be included in the output of the workflow. Peptide sequences 701 and replacements 702 found by a linear human proteome search for peptide sequences within the extended human proteome database 703 can be assigned a linear extended human proteome source. The assigned source can be included in the output of the workflow. Replacements found by a linear human proteome search within the extended human proteome database 703 can be included in the output of the workflow.
[0096] In the second peptide source search step 705, a query (701 and 702) can be applied to the translated human genome frame using BLAT or a similar algorithm tool. BLAT is disclosed, for example, in GenomeRes. 2002 Apr; 12(4): 656-664. BLAT-The BLAST-Like AlignmentTool, which is hereby incorporated by reference in its entirety. An example of a BLAT command can be, by way of non-limiting example, "blat -t=dnax -q=prot -minScore=7 -stepSize=1 hg38.2bit Fasta_query output.psl psl2bed < output.psl > perfect_match.bed". If an identical match is found, the peptide sequence can be labeled "linear" at 706, and the possible source sequences can be included in the output. The peptide sequences 701 and its permutation 702 found by the linear human genome search 705 can assign a linear genome source. The assigned source can be included in the output of the workflow.
[0097] In the third peptide source search step 707, the peptide sequences of the queries (701 and 702) can be mapped to an extended human proteome database 703 that allows for some mismatches (non-limiting examples of mismatches include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.). In certain embodiments, the number of mismatches can be 1. An example of the BLAT command can be, for example, "blat -t=prot -q=prot -minScore=7 -stepSize=1 combined_DB.processed.fasta Fasta_query output_blat_hits.psl". If a peptide sequence with a mismatch is found (as an example, 1 mismatch), the peptide sequence can be labeled "1 mismatch" at 708. The peptide sequences 701 and its permutation 702 found by the linear mismatch search for peptides having a mismatch to the peptide sequences within the extended human proteome database 707 are assigned a linear mismatch of the extended human proteome as the source. The assigned source can be included in the output of the workflow.
[0098] In the fourth peptide source search step 709, the peptide sequences of the queries (701 and 702) can be mapped to other organisms at 709, for example, by using the BLAST NCBI tool. If any identical matches (e.g., homologous matches) are found, the results can be annotated as "linear BLAST" at 710. The peptide sequences 701 and its permutation 702 found by a linear non-endogenous search for the peptide sequences in the non-endogenous proteome database 709 are assigned a linear non-endogenous proteome source. The assigned source can be included in the output of the workflow. In some embodiments, the fourth peptide source search step 709 can be omitted, and the workflow illustrated in FIG. 7 can be modified to omit the blocks 709 and output 710 associated with this step. In such embodiments, the workflow can proceed from the third peptide source search step 707 to the fifth peptide source search step 711.
[0099] In the fifth peptide source search step 711, the peptide sequences of the queries (701 and 702) can be fragmented into two or more fragments (if each fragment is more than one amino acid). The fragments can be used as queries applied to the extended human proteome database. If there are matches for both fragments in the same protein, the peptide sequence can be labeled as "cis-splice" at 712. The peptide sequences 701 and its permutation 702 found by the cis-splice search in the extended human proteome database 711 are assigned a cis-splice human proteome source. The assigned source can be included in the output of the workflow.
[0100] In the sixth peptide source search step 713, if hits exist for both fragments in two different proteins, the peptide sequence can be labeled "trans-splice" at 714. Peptide sequences 701 found by trans-splice search of the extended human proteome database 711 and its permutation 702 are assigned a trans-splice human proteome source. The assigned source can be included in the output of the workflow. In some embodiments, the sixth peptide source search step 713 can be omitted, and the workflow illustrated in FIG. 7 can be modified to omit block 713 and output 714 associated with this step. In such embodiments, the workflow can proceed from the fifth peptide source search step 711 to block 715.
[0101] Any remaining peptide sequences can be labeled as unassigned (N / A) at 715. The workflow can stop proceeding to subsequent peptide source search steps when assigning the putative source to the peptide sequences of queries 701, 702.
[0102] Returning to FIG. 5, in one embodiment, computer device 512 can verify the data / information received from mass spectrometer 504 based on the labeling of the peptide sequences determined according to query support data structure 600.
Examples
[0103] The examples presented herein generally include a peptide source assignment workflow having search steps sequenced in increasing order of random hit rate, as well as methods and systems for using and creating a peptide source assignment workflow. The examples presented below are specific to peptide labeling, but other applications including those disclosed above can be performed according to similar methodologies. The examples presented below can reduce the false labeling of peptides as cis-splice and trans-splice compared to previous systems and methodologies.
[0104] Antigen-presenting cells use major histocompatibility (MHC) complex I or II, respectively, to present peptides to CD8+ or CD4+ T cells. The characterization of peptides presented to T cells, known as the immunopeptidome, is being studied in the fields of infectious diseases, autoimmunity, and cancer immunotherapy. Cancer-related MHC-presented peptides that induce an immune response are potentially safe and effective targets for cancer immunotherapy. The discovery and characterization of the immunopeptidome can be achieved using a number of techniques, such as whole exome sequencing, RNA sequencing, ribosome profiling, and peptide sequencing based on tandem mass spectrometry (MS / MS). Next-generation sequencing approaches can characterize the potential endogenous immunopeptidome, but only direct detection of peptides, such as by MS / MS, can provide experimental evidence for the presence of peptides presented by the MHC complex. In particular, in addition to using peptides bound to the MHC complex for the characterization of the immunopeptidome, peptides can also originate from multiple other gene- and transcription-based abnormalities. Examples of additional means for identifying abnormal peptides include cancer-specific genes and transposon overexpression (e.g., but not limited to, cancer-testis genes, transposons, and human endogenous retroviruses (HERV)), alternative splicing, stop codon read-through, or alternative open reading frame translation.
[0105] Immune peptide mixes using peptide-MHC elution followed by MS / MS have traditionally required a reference database of peptides that might be detected. Recent advances in peptide spectrum matching software have obviated the need for a reference database search for de novo sequencing, such that the software can directly identify the sequences of unknown peptides, post-translational modifications (PTMs), and amino acid substitutions from the MS / MS spectra. Using these methods, it is possible to understand the diversity of peptides that can bind to MHC complexes rather than their protein sources. For MHC I, the classical mechanism of peptide presentation begins with proteasomal cleavage of proteins in the cytoplasm, yielding fragments 8 - 12 amino acids in length. These peptides then bind to the MHC I complex prior to their movement to the cell membrane. However, some studies have suggested that, in addition to cleavage, the proteasome can catalyze reverse reactions and ligate small peptides together in a process called proteasome-catalyzed peptide splicing (PCPS). Classical cleavage yields peptides whose sequences are identical to the parent protein (referred to herein as linear), whereas pieces of spliced peptides can be from the same protein (referred to herein as cis-splicing) or, in theory, from different proteins (referred to herein as trans-splicing).
[0106] Before attempting to use de novo sequencing approaches to identify peptides of unknown origin ("hidden peptides"), many of these hidden peptides have been identified as potentially arising from post-translational splicing. However, the abundance and even the existence of spliced peptides are issues that have been debated in the art. Strategies for identifying spliced peptides in the MHC-I immunopeptidome by mass spectrometry have been previously developed. A database containing all possible cis-spliced peptides that enables querying of MS / MS spectra for cis-spliced peptides was created. It has been reported that approximately 30% of p-HLA are cis-spliced peptides at short distances. The same group also developed a pipeline for mapping the spliced immunopeptidome of MHC class I in cancer cells. This study suggested that a substantial (about 25%) portion of the peptides could be mapped to cis-spliced sequences in the HCT116 and HCC1143 cell lines derived from colon cancer and breast cancer, respectively. Trans-spliced peptides were excluded from the analysis because their existence in vivo has been controversial and their addition to the database greatly increases its complexity. Subsequently, a bioinformatics workflow, called HybridFinder, was developed to identify linear, cis, and trans-spliced peptides. HybridFinder first searches for exact matches of peptides in the UniProt human protein sequence database and then searches for all possible cis and then trans-spliced forms of that peptide in the human proteome. Using HybridFinder, MS / MS data containing peptides eluted from MHC I complexes purified from 17 HLA heterozygous cell lines were analyzed. Cis and trans-spliced peptides were found to present up to 45% of the MHC-binding peptides. A. Results 1. Expansion of the search for sources of non-classical human peptides
[0107] Disclosed herein is a strategy for determining the order of putative sources when assigning a source to a de novo array peptide. Using this strategy, a peptide source assignment workflow is developed that searches for the source of a peptide from among a plurality of sources in a particular order having an order optimized to minimize the assignment of peptides to incorrect sources. For example, assignment of de novo peptides to post-translational cis or trans splicing occurs very frequently by chance, and most peptides can be attributed to other sources that are less likely to occur by chance. As disclosed herein, a rigorous derivation of the optimal order of peptide source assignment is presented, along with the utility of the workflow in identifying most of the valid sources of de novo peptides, thus facilitating the understanding of the immunopeptidome.
[0108] Previous studies have shown that up to 45% of MHC-binding peptides are not mapped as in the UniProt human proteome. The workflows disclosed herein include a developed database of data from several other possible sources that may also be due to unmapped peptides. Peptides from non-classically translated regions of the human genome were searched, for example, peptides from regions annotated as non-coding. For this source, OpenProt was used, which includes all open reading frames (ORFs) of at least 30 codons in length, and this was supplemented with the rest of the human genome translated in six frames. Also included were translations of known transcribed elements, which include long non-coding RNAs (lncRNAs), microRNAs (miRNAs), and HERVs that can be spliced and thus may contain sequences not found via translation of genomic DNA. Below, this combination of sources is referred to as the extended human proteome database. In addition, single mismatches to sequences encoded in the human proteome can arise due to unknown SNPs, missense mutations, or repetitive errors in any of transcription, translation, or MS amino acid identification. Mismatched peptides were searched for using BLAT and the de novo peptide sequences were aligned to the extended human proteome database with an allowed single mismatch. Finally, some peptides may be of other organismal origin, particularly bacterial or viral sources. For these sources, de novo peptide sequences were searched in the BLAST database (see methods). 2. Optimal ordering of putative sources by estimation of random hit rates
[0109] For each possible source of the peptides described above (e.g., computer devices 107-112 in FIG. 1, peptide sources identified by search steps 703, 705, 707, 709, 711, 713 in FIG. 7, etc.) that include cis and trans splicing of peptides (e.g., peptides in FIG. 3, database, query 701, 703 in FIG. 7, etc.), an opportunity to randomly find a match was determined. To estimate the random hit rate associated with each possible putative source of the peptide, it was determined how many randomly created arrays would be found in each source (e.g., the number of randomly created arrays found in each search step 703, 705, 707, 709, 711, 713 in FIG. 7). The random hit rate associated with the source and / or search step may further depend on the size and / or complexity of the database. Peptide sequences of 8-12mers (1,000 per length) were created in two ways: random sequences that uniformly sample all amino acids (hereinafter referred to as uniformly random) or sequences having amino acid frequencies that match the frequencies of amino acids found in vertebrates (hereinafter referred to as weighted random); see Table 1. However, peptides of 8-14 amino acid lengths may be used (e.g., but not limited to, 8-14 amino acids, 9-14 amino acids, 10-14 amino acids, 11-14 amino acids, 12-14 amino acids, 13-14 amino acids, 8-13 amino acids, 8-12 amino acids, 8-11 amino acids, 8-10 amino acids, 8-9 amino acids, 9-13 amino acids, 9-12 amino acids, 9-11 amino acids, 9-10 amino acids, 10-13 amino acids, 10-12 amino acids, 10-11 amino acids, 11-13 amino acids, 11-12 amino acids, 12-13 amino acids).
[0110] A random array was used to estimate the random hit rate of each putative source of peptides (Figure 8). Using a uniform random array, 14 (0.28%) out of 5,000 peptides were found in the extended human proteome (e.g., but not limited to, UniProt, OpenProt, lncRNA, miRNA, and HERV). When searching the classical non-coding regions of the human genome using BLAT, 178 (3.3%) out of 5,000 random peptides could be mapped. The estimation of the random hit rate was also determined when searching for peptide mapping to the human proteome with a single mismatch, and it was found that 192 (3.9%) out of 5,000 peptides could be mapped. When searching for peptides that could be derived from non-human organisms in the BLAST database, 604 (12.1%) out of 5,000 sequences could be mapped. Finally, when searching for peptides cis-spliced in the extended human proteome, 1,936 / 5,000 (38%) could be mapped, and for trans-spliced peptides, 3,598 / 5,000 (71%) could be mapped (Figure 8). For weighted random peptides, 50% and 68% of the peptides could be assigned to cis or trans splicing, respectively (Figure 8).
[0111] An enhanced peptide mapping pipeline was designed that assigns sources for peptides in decreasing order of random hit rate. When using any set of simulation data to order peptide sources by random hit rate, the enhanced pipeline searches for peptide sources in the following order: 1) the extended human proteome database (assigned linearly), 2) the non-coding regions of the human genome using BLAT (assigned linearly), 3) single mismatch peptides in the extended human proteome (assigned linearly), 4) the BLAST database (assigned linearly), 5) cis-spliced peptides and 6) trans-spliced peptides in the extended human proteome. When ordering in sequence, 4,495 / 5,000 (90%) of the uniform random sequences and 4,847 / 5,000 (97%) of the weighted random sequences were found by this pipeline (Figure 10). Although this random hit rate is high, the researchers can select an appropriate threshold and exclude peptides mapped from sources with a high random hit rate. In fact, previous studies have excluded searches for trans-spliced peptides due to their presumed high random hit rate and assumed rarity of occurrence. 3. Identify peptides whose sequences are assigned with higher confidence during de novo sequencing in the first part of the assignment workflow
[0112] Tests were conducted to determine whether the proportion of peptides found in the actual experiment matches the final order. The peptide source assignment workflow applied to six novel immunopeptide mix data is set up from the IM9 and Raji cell lines (see methods). During de novo sequencing, the amino acid calls give a local confidence score, and the quality of the sequencing across the peptide can be quantified by the average local confidence (ALC%) score. The ALC% score is generated by MS / MS and is associated with each de novo peptide sequence 701 (Figure 7).
[0113] We hypothesized that peptides with higher ALC% are more likely to be assigned to more reliable sources with lower random hit rates, i.e., sources with a previous source in our workflow. Indeed, over six experiments, most of the peptides with the highest ALC% were found in the first source of the pipeline (linear extended human proteome source), in stark contrast to the source patterns found for randomly generated peptides (Figure 11). As the ALC% decreases, more peptides can be found in later sources in the pipeline, with the most significant increases for cis-spliced peptides in both cell lines, as well as for blast and trans-spliced peptides for samples from the IM9 and Raji cell lines, respectively (Figure 11). Depending on the true composition of a particular sample, different sources in the pipeline can be differentially enriched in the final call.
[0114] The second fragmentation in MS / MS experiments (MS2 scans) can be inherently prominent due to poor fragmentation or ionization of a particular peptide. To evaluate the proportion of de novo calls that are ambiguous as a function of ALC%, a series of MS2 scans were obtained from the IM9 cell line where both de novo identified peptides and conventional database calls were available. As hypothesized, the de novo ALC% decreased, and as a result, the proportion of peptide calls that matched between de novo and conventional database searches also decreased (Figure 12). Collectively, this indicates that de novo peptide sequences with low ALC% and their sources must be placed under additional scrutiny. 4. Reanalysis of single allele cell line data using the peptide source assignment workflow
[0115] Peptide identification by the peptide source assignment workflow was compared to the HybridFinder against peptides eluted from the MHC complex in a dataset from an immune peptide mix derived from a collection of cell lines engineered to express a single HLA allele. See Figure 9 for the HybridFinder workflow. We found that most of the peptides identified by HybridFinder as cis or trans-spliced could also be mapped to sources with a lower random hit rate. For example, for cell lines expressing HLA-A * 02:04, of the 1,075 peptides classified as spliced by HybridFinder, 215 were classified as linear from the extended human database, 120 were classified as linear with 1 mismatch, 301 were classified as linear from the BLAST database, and we found that 636 / 1,075 (60%) of the putative spliced peptides were reclassified as linear. Additionally, 133 of the peptides classified as trans-spliced could be reclassified as cis-spliced using the extended human proteome (Figure 10). Across all cell lines, 36% of the putative cis-spliced peptides could be reclassified as linear, and 45.9% of the putative trans-spliced peptides were reclassified as linear or cis-spliced (Figure 13).
[0116] The peptide source assignment workflow presented herein indicates that putative spliced peptides may be peptides resulting from mutant DNA sequences, non-classically spliced RNA sequences, non-classically translated regions of the human genome, mismatched human sequences, or bacterial proteins. In short, down from 29% using the Hybrid Finder, 20% of the peptides are assigned as spliced peptides by the workflow presented herein (Figure 13). Since this results in assigning peptides to the putative source with the lowest random hit rate, providing a workflow that results in higher peptide assignment reliability, this overall reduction in the identification of putative spliced peptides is notable. Since spliced peptides have the highest random hit rate compared to other potential sources presented herein, a significant portion of the peptides assigned as spliced by the Hybrid Finder may be inappropriately assigned. Thus, the workflow presented herein is an improvement over the Hybrid Finder due to the overall reduction in the identification of putative spliced peptides compared to the Hybrid Finder. In some embodiments, the method of the invention reduces the identification of spliced peptides by 5 - 60%. In some embodiments, the method of the invention reduces the identification of spliced peptides by 5 - 50%. In some embodiments, the method of the invention reduces the identification of spliced peptides by 5 - 40%. In some embodiments, the method of the invention reduces the identification of spliced peptides by 5 - 30%. In some embodiments, the method of the invention reduces the identification of spliced peptides by 5 - 20%. In some embodiments, the method of the invention reduces the identification of spliced peptides by 5 - 10%. In some embodiments, the method of the invention reduces the identification of spliced peptides by 10 - 60%. In some embodiments, the method of the invention reduces the identification of spliced peptides by 10 - 50%. In some embodiments, the method of the invention reduces the identification of spliced peptides by 10 - 40%.In some embodiments, the method of the present invention reduces the identification of spliced peptides by 10-30%. In some embodiments, the method of the present invention reduces the identification of spliced peptides by 10-20%.
[0117] In some embodiments, the method of the present invention reduces the identification of spliced peptides by 20-60%. In some embodiments, the method of the present invention reduces the identification of spliced peptides by 30-60%. In some embodiments, the method of the present invention reduces the identification of spliced peptides by 40-60%. In some embodiments, the method of the present invention reduces the identification of spliced peptides by 50-60%. In some embodiments, the method of the present invention reduces the identification of spliced peptides by 20-50%. In some embodiments, the method of the present invention reduces the identification of spliced peptides by 30-40%.
[0118] In some embodiments, the method of the present invention reduces the identification of spliced peptides by 5-70%. In some embodiments, the method of the present invention reduces the identification of spliced peptides by 14-60%.
[0119] At each step, the random peptide mapping results were used to estimate how many peptides might have been found by chance. When compared to weighted random peptides, more peptides were detected in the cell line assigned as linear (P<1e-308, two-sided Fisher's exact test), linear with a single mismatch (P = 3.16e-05), linear from the BLAST database (P = 0.00126), cis-splice (P = 9.9e-92), and trans-splice (P = 3.29e-73) (Figure 14). More peptides were found than expected by searching for random peptides, indicating that each source contributes to the immunopeptidome found in each cell line due to the optimal ordering of source assignments. 5. Enrich for expressed regions of peptides mapped across the entire human genome
[0120] Peptides that are within the human genome but mapped outside the UniProt proteome were examined for where they fit in terms of genome annotation. The first three steps of the pipeline can map peptides to regions of the human genome. In the first step, peptides that map exclusively in the OpenProt database fall into ORFs that do not exist in the UniProt human proteome. Analyzing the locations of these peptides in the human genome (Figure 16) and comparing them to the locations of all proteins in OpenProt, these peptides are enriched in exons, promoters, and 5’UTRs (Figure 17). The exon enrichment could be due to frameshift translation. In subsequent steps, OpenProt includes all proteins whose ORFs are longer than 30 amino acids, and although these peptides must be derived from ORFs shorter than 30 amino acids, the peptides are mapped to the six-frame translation of the human genome. The genomic annotation distribution of the peptides mapped in this step is closer to that of the human genome, i.e., most of the peptides are mapped to intergenic or intron regions (Figure 18), indicating that the assignment of these peptides is contaminated by random matching. However, peptides from the proteome dataset are more enriched for exons, promoters, and 5’UTRs than peptides from random uniform or weighted simulations (Figure 19). In the third step of the enrichment pipeline, peptides can be mapped to an extended human proteome with a single mismatch (Figure 20). The peptides mapped in this step show stronger enrichment in exons, introns, and promoters than the enrichment found for peptides in weighted or uniform simulation datasets (Figure 21). At each step, there is a consistent decrease in the intergenic region and enrichment in the transcriptional region, as found in other studies focusing on unidentified peptides in the immunopeptidome.Enrichment of the transcribed sequences supports the idea that the peptides assigned in these steps of the pipeline are accurately assigned even if they do not map to proteins in the UniProt database. Peptides identified by BLAST are not enriched for any bacterial genus
[0121] Searches within the BLAST database have the highest random hit rate for linear peptides in the peptide source assignment workflow. Peptides from cell lines had somewhat more matches in the BLAST database than expected based on uniform or weighted random data (Figure 14), but it was determined whether BLAST assignments showed enrichment of specific microorganisms that could be contaminants. To calculate enrichment, peptides that could not be uniquely mapped to a single species were removed, and then Fisher's exact test was applied to the counts of peptide mappings to each genus in each cell line and all cell lines. After correcting for multiple hypothesis testing, no genus was significantly enriched in any cell line or when all cell lines were considered together. There were three possible, non - mutually exclusive causes for the observed lack of enrichment. First, there were no contaminating organisms in the immunopeptidome preparation. Second, the presence of organisms was not represented in the BLAST database. Third, peptides arising from contaminating organisms could not be uniformly mapped to a single organism and were thus excluded from the above analysis. In the first two possibilities, the reason why cell lines have more BLAST matches than expected is not clear. Collectively, these results do not support a biological basis for the peptides assigned in the BLAST step; rather, the matches found in this step may be random and spurious. 7. Repetitive novel peptides
[0122] To identify common peptides reclassified across multiple datasets, peptides common to three or more cell lines were selected. For example, QSPVALRPL (SEQ ID NO: 5) was highly repetitive and identified as a trans-splice by the hybrid finder algorithm, but was reclassified as linear using the disclosed pipeline. The same peptide is listed in the immune epitope database as part of the un-identified peptides. In further inspection, this is an out-of-frame peptide in the FAM96A gene, which is an apoptosis-promoting tumor suppressor in gastrointestinal stromal tumors (see, for example, Schwamb et al. Int. J. Cancer (2015) Sept 15; 137(6):1318-29, which is hereby incorporated by reference in its entirety). If out-of-frame translation is specific to cancer samples, this peptide could be a target for cancer immunotherapy. B. Discussion
[0123] As the number of peptides identified in immunopeptide mix experiments using de novo sequencing increases, the need for better characterization of the immunopeptidome is more pressing than ever. Previous studies attributed PCPS as the primary source of peptides of unknown protein identity. A peptide source assignment workflow is described herein that assigns the parent protein of de novo sequenced peptides from several sources with a lower random hit rate than the set of all possible PCPS peptides. It was found that 32% of the putative PCPS peptides could be explained by known proteins, translation of probably untranslated parts of the human genome, or a single mismatch with bacterial and viral peptides. Not surprisingly, most of the peptides are encoded by known expression regions. Finally, repetitive out-of-frame peptides were identified in the tumor suppressor gene FAM96A, which could be a target for cancer immunotherapy purposes. C. Methods 1. Datasets i. Simulation Random Peptides
[0124] Two sets of peptide sequences were simulated for an estimate of the random hit rate. Using the "random" built-in python library, sets of amino acid sequences of lengths 8 - 12 were created, with 1,000 peptides for each length and a total of 5,000 random peptides for each set. For the first simulated peptide sequence set, all amino acids had an equal probability of being incorporated into the sequence, and this set was termed "uniform random". In the second set, amino acids had an incorporated probability matching their frequencies in vertebrates, and this set was termed "weighted random". The two sets of peptide sequences are included in Table 1. ii. IM9 and Raji cell line immune peptide mixes
[0125] Three replicates of the IM9 and Raji cell lines were processed by MS / MS: 210210_IM - 9_1_IFN_cl1, 210210_IM - 9_2_IFN_cl1, 210210_IM - 9_3_IFN_cl1, 180316_RAJI_NoIFN, 180323_Raji_IFN, and 180323_Raji. All replicates of the IM9 cell line were simulated with IFNγ, while only two replicates of the Raji cell line received the same treatment.
[0126] IFNγ can enhance the expression of surface major histocompatibility complex (HLA) molecules, increase the processing and presentation of tumor - specific antigens, and promote T - cell recognition and cytotoxicity. IFNγ also not only up - regulates many components of the antigen - presentation pathway but also induces a shift between constitutive immunoproteasome subunits with different catalytic activities in the proteasome, creating different populations of HLA - associated peptides. We used IFNγ treatment of the cell lines to increase the opportunity to expand the immune peptidome detectable by mass spectrometry. a. Immunoprecipitation
[0127] The HLA-Pan class I (W6 / 32) column was prepared using NHS-activated Sepharose 4 beads (GE Healthcare 17090601) and a coupling buffer of 0.2 M sodium bicarbonate and 0.5 M sodium chloride, and it was washed with 0.1 M Tris-HCl and 0.1 M acetate buffer at pH 8.5. Affinity purification was performed under gravity, and the flow-through was captured for further analysis. The bound HLA molecules were eluted under gravity using 0.1 M glycine (Sigma), pH 2.7 (Figure 1). 0.1% trifluoroacetic acid (catalog number: LC485-1 Honeywell) was added to the glycine eluate. HLA-associated peptides were eluted in a two-step elution using a Sep-Pak (catalog number: WAT054960 Waters). HLA-specific peptides were eluted using 30% acetonitrile (catalog number: LC34967 Honeywell) / 0.1% trifluoroacetic acid, and HLA molecules were eluted using 70% acetonitrile / 0.1% trifluoroacetic acid. Aliquots of the lysate, flow-through, glycine, 30% acetonitrile / 0.1% trifluoroacetic acid, and 70% acetonitrile / 0.1% trifluoroacetic acid eluate were collected throughout the process.
[0128] The peptide and HLA fractions were placed on a SpeedVac Vacuum Concentrator (Thermo) for 2 hours. Each sample was resuspended in 0.1% trifluoroacetic acid after SpeedVac. The peptide fraction was further purified using a C-18 ZipTip® (catalog number: ZTC185096 Millipore). Then, all samples were analyzed using an Orbitrap Fusion™ Lumos™ Tribrid™ mass spectrometer (Thermo) for peptide sequencing. b. Data analysis
[0129] Raw data files from an Orbitrap Fusion™ Lumos™ (Thermo) LC / MS were searched against the human Uniprot database, a custom database for the protein of interest, and de novo using PEAKS® Studio X (BSI) proteomics software. iii. HLA single allele immune peptide mix
[0130] For MS / MS data from HLA-I single allele cell lines, peptides were downloaded from the supplementary table of Faridi, P. et al. Sci. Immunol. Vol 3, issue 28, pg 3947, Oct. 12 (2018), which is incorporated herein by reference in its entirety. The data included the expression of 8 HLA-A alleles (A0101, A0203, A0204, A0207, A0301, A3101, A6802, A2402) and 9 different HLA-B alleles (B5801, B5703, B5701, B4402, B5101, B0801, B1502, B2705, B0702). In total, over 51,000 unique peptides were present. 2. Reproducibility of the Hybrid Finder
[0131] To facilitate comparison with the peptide source assignment workflow described, the workflow was compared to the hybrid finder described in Faridi, P. et al. Sci. Immunol. Vol 3, issue 28, pg 3947, Oct. 12 (2018), which is incorporated herein by reference in its entirety. was reproduced. First, each peptide was searched in the UniProt human reference proteome database. Peptides with identical matches were annotated as linear. For peptides without a linear match, all possible splits of the peptide were created where the length of the smaller piece was longer than 1 amino acid. Next, possible matches for each fragment were searched through the database. When identical matches for both fragments were detected in a single protein, the peptide was annotated as cis-splice. The matches can be ordered in reverse. Otherwise, if the matches are available in two separate proteins, the peptide was annotated as trans-splice. Peptides for which the split pairs did not match any protein sequence were annotated as not applicable (N / A). 3. Extended Human Proteome Database
[0132] FASTA files of reviewed and unreviewed human sequences from OpenProt (www.openprot.org) and UniProt (www.uniprot.org) were combined, including protein sequences from some viruses that use humans as hosts (UniProt proteome version UP0000056430, downloaded in May 2020). This database was extended to include translated protein sequences from lncRNA (NONCODE version v5.19, downloaded in May 2020), miRNA (last updated 3 / 10 / 18, downloaded in May 2020), and endogenous viral elements (gEVE database ORFs21, downloaded in May 2020). This database was used when the workflow searches for linear human peptides and single-mismatch human peptides (steps 1 and 3), as well as in the search for cis- and trans-spliced peptides. 4. Peptide Source Assignment Workflow
[0133] The inherent random hit rate was measured in each of the respective sources where the peptides in the immunopeptide mix experiment could be found using the simulated random dataset described above. The steps of the workflow were ordered in ascending order of the random hit rate to construct the workflow. The steps applied to each de novo sequenced peptide are as follows:
[0134] Step 1: Search for identical sequence matches in the extended human proteome database (described above). Leucine (L) and isoleucine (I) have the same mass and thus it is impossible to distinguish them in de novo search sequencing. To account for this, for a given peptide containing I / L, all permutations of the I and L residues are considered. For example, for the peptide "ATTSLLHN (SEQ ID NO: 1)", there are four possible permutations: ATTSLLHN (SEQ ID NO: 1), ATTSLIHN (SEQ ID NO: 2), ATTSILHN (SEQ ID NO: 3), and ATTSIIHN (SEQ ID NO: 4). If the algorithm finds an identical match (e.g., 100% identical) for any permutation, the peptide is annotated as "linear" and all possible protein sources of the peptide are included in the output. Since a match has been identified, the algorithm proceeds to the next step, for example, it does not need to continue to Step 2. Otherwise, if no match is identified, the algorithm proceeds to Step 2.
[0135] Step 2: Search for identical matches in any of the six frames of the translated human genome using BLAT32. The following command is used.
[0136] blat -t=dnax -q=prot -minScore=7 -stepSize=1 hg38.2bit Fasta_query output.psl
[0137] psl2bed < output.psl > perfect_match.bed
[0138] If the same match is found, the peptide is annotated as "linear" and the possible source sequences are included in the output. Otherwise, the peptide moves to step 3.
[0139] Step 3: Map the peptide to the extended human proteome database. At this time, use the following code: "blat -t=prot -q=prot -minScore=7 -stepSize=1 combined_DB.processed.fasta Fasta_query output_blat_hits.psl" at the genomic position of the BLAT hit analysis to allow for 1 mismatch.
[0140] If an array with a single mismatch is found, the peptide is annotated as "1 mismatch". Otherwise, the peptide moves to step 4.
[0141] Step 4: Map the array to other organisms using the BLAST NCBI tool. If any identical match is found, the result is annotated as "linear BLAST".
[0142] Step 5: For the remaining peptides, the algorithm creates all possible splits of the peptides where the length of the smaller piece is longer than 1. Then it searches for matches of both fragments in all human sequence databases. If there is a match for both chunks in the same protein, the tool annotates the peptide as "cis-splice". Otherwise, if there is a hit for both fragments in two different proteins, the tool annotates the peptide as "trans-splice". The remaining peptides without any match are assigned as not applicable (N / A). 5. Genomic position of BLAT hit analysis
[0143] The genomic locations of BLAT hits were analyzed using the annotatepeaks.pl script from the HOMER suite. Specifically:
[0144] annotatePeaks.pl ${file} hg38 -annStats ${file}.summary.txt
[0145] Only basic annotations were considered for further analysis. To calculate the enrichment of genomic locations of peptides found in the OpenProt database with either an exact match (step 1) or a single mismatch (step 3), a Fisher's exact test was performed to compare the number of peptides in each genomic annotation in the sample against the entire OpenProt database. For peptides mapped to any translated region in the human genome (step 2), the enrichment of p-values calculated by HOMER was used for over- or under-representation of each genomic annotation. 6. Tools
[0146] Python, bedops, psl2bed, BLAT, BLAST, HOMER.
[0147] Figure 22 shows a system 2200 for performing the method described herein. In certain embodiments, system 2200 can be configured to execute the workflow illustrated in FIG. 7. In certain embodiments, system 2200 can include some or all of the databases utilized by the workflow illustrated in FIG. 7. In certain embodiments, system 2200 can be configured to communicate with one or more of the databases utilized by the workflow illustrated in FIG. 7. In certain embodiments, system 2200 can include some or all of the data sources 518A-518N illustrated in FIG. 5. In certain embodiments, system 2200 can be configured to communicate with one or more of the data sources 518A-518N illustrated in FIG. 5.
[0148] The devices / components described in this specification may include a computer 2201, as shown in FIG. 22. The computer 2201 may include one or more processors 2203, a system memory 2212, and a bus 2213 that connects the various components of the computer 2201 including from the one or more processors 2203 to the system memory 2212. In the case of multiple processors 2203, the computer 2201 may utilize parallel computing.
[0149] The bus 2213 may include one or more of several possible types of bus structures, such as a memory bus, a memory controller, a peripheral bus, an accelerated graphics port, and one or more of a processor or local bus that uses any of a variety of bus configurations.
[0150] The computer 2201 may operate on and / or include various computer-readable media (e.g., non-transitory). The computer-readable media can be any available media that is accessible to the computer 2201 and includes non-transitory, volatile and / or non-volatile media, removable and non-removable media. The system memory 2212 has computer-readable media in the form of volatile memory such as random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM). The system memory 2212 can store data such as mass spectrometry data 2207, and / or program modules such as an operating system 2205 and query analysis software 2206 that is accessible to and / or operates on the one or more processors 2203. The system memory 2212 can further include some or all of a database utilized by the workflow illustrated in FIG. 7, and / or some or all of the data sources 518A - 518N illustrated in FIG. 5.
[0151] Computer 2201 may also include other removable / non-removable volatile / non-volatile computer storage media. Mass storage device 2204 may provide non-volatile storage of computer code, computer-readable instructions, data structures, program modules, and other data for computer 2201. Mass storage device 2204 may be, but is not limited to, a hard disk, a removable magnetic disk, a removable optical disk, a magnetic cassette or other magnetic storage device, a flash memory card, a CD-ROM, a digital versatile disk (DVD) or other optical storage, a random access memory (RAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0152] Any number of program modules may be stored on mass storage device 2204. Operating system 2205 and query analysis software 2206 may be stored on mass storage device 2204. One or more (or some combination thereof) of operating system 2205 and query analysis software 2206 may include program modules and query analysis software 2206. Mass spectrometry data 2207 may also be stored on mass storage device 2204. Mass spectrometry data 2207 may be stored in any one of one or more databases known in the art. The database may be centralized within network 2215 or distributed across multiple locations. Mass storage device 2204 may further include some or all of the database utilized by the workflow illustrated in FIG. 7 and / or some or all of data sources 518A-518N illustrated in FIG. 5.
[0153] The user may input commands and information into computer 2201 via an input device (not shown). Such input devices include, but are not limited to, a keyboard, a pointing device (e.g., a computer mouse, a remote control device), a microphone, a joystick, a scanner, a tactile input device such as a glove, and other body covers, motion sensors, and the like. These and other input devices may be connected to one or more processors 2203 via a human-machine interface 2202 connected to bus 2213, but may also be connected by other interfaces and bus structures, such as a parallel port, a game port, an IEEE 1394 port (also known as a Firewire (registered trademark) port), a serial port, a network adapter 2208, and / or a universal serial bus (USB).
[0154] The display device 2211 can also be connected to the bus 2213 via an interface such as the display adapter 2209. It is contemplated that the computer 2201 may have two or more display adapters 2209, and the computer 2201 may also have two or more display devices 2211. The display device 2211 may be a monitor, an LCD (liquid crystal display), a light emitting diode (LED) display, a television, a smart lens, smart glasses, and / or a projector. In addition to the display device 2211, other output peripheral devices may include components such as a speaker (not shown) and a printer (not shown) that can be connected to the computer 2201 via the input / output interface 2210. Any step and / or result of the method can be output (or caused to be output) to the output device in any form. Such output may be in any form of visual display including, but not limited to, text, graphics, animation, audio, tactile, etc. The display 2211 and the computer 2201 may be part of one device or separate devices.
[0155] Computer 2201 may operate in a network environment using logical connections to one or more remote computer devices 2214a, b, c. The remote computer devices 2214a, b, c can be personal computers, computer stations (e.g., workstations), portable computers (e.g., laptops, cell phones, tablet devices), smart devices (e.g., smartphones, smartwatches, activity trackers, smart apparel, smart accessories), security and / or monitoring devices, servers, routers, network computers, peer devices, edge devices, or other common network nodes, etc. The logical connection between computer 2201 and remote computer devices 2214a, b, c can be made via a network 2215, such as a local area network (LAN) and / or a common wide area network (WAN). Such network connections can be through a network adapter 2208. The network adapter 2208 can be implemented in both wired and wireless environments. Such network environments are conventional and common in homes, offices, enterprise-scale computer networks, intranets, and the Internet.
[0156] Application programs and other executable program components, such as operating system 2205, are shown herein as separate blocks, but it should be recognized that such programs and components may be present in different storage components of computer device 2201 at various times and are executed by one or more processors 2203 of computer 2201. The implementation of query analysis software 2206 may be stored in or transmitted in some form of computer-readable medium. Any of the disclosed methods can be performed by processor-executable instructions embodied on a computer-readable medium.
[0157] In one embodiment, the query analysis software 2206 may be configured to perform some or all of the search steps 703, 705, 707, 709, 711, 713 illustrated in FIG. 7.
[0158] In one embodiment, the query analysis software 2206 may be configured to perform the method 2300 shown in FIG. 23. The method 2300 may be performed, in whole or in part, by a single computer device, a plurality of electronic devices, etc. The method 2300 may include, at 2302, creating a plurality of simulated random queries. Creating a plurality of simulated random queries may include at least one of creating a plurality of uniform random queries or creating a plurality of weighted random queries. The plurality of simulated random queries may include a plurality of simulated random text strings. The plurality of simulated random queries may include a plurality of simulated random peptide sequences.
[0159] The method 2300 may include, at 2304, determining the number of matches associated with each source based on applying the plurality of simulated random queries to each source of the plurality of sources.
[0160] In one embodiment, method 2300 may include, at 2306, determining a false discovery rate associated with each source based on the number of matches associated with each source. In one embodiment, a function of the number of matches and the number of simulated random queries may be determined. In one embodiment, the determination may be made by dividing the number of matches by the number of simulated random queries. In one embodiment, determining a false discovery rate associated with each source based on the number of matches associated with each source may include a function of the number of matches and the number of simulated random queries. In one embodiment, determining a false discovery rate associated with each source based on the number of matches associated with each source may include dividing the number of matches by the number of simulated random queries.
[0161] Method 2300 may include, at 2308, creating a query support data structure configured to facilitate the application of new queries to a plurality of sources based on the false discovery rate.
[0162] In one embodiment, query analysis software 2206 may be configured to perform method 2400 shown in FIG. 24. Method 2400 may be performed, in whole or in part, by a single computer device, a plurality of electronic devices, etc. Method 2400 may include, at 2402, receiving a query. The query may include a text string. The query may include a peptide sequence. Receiving the query may include receiving a peptide sequence from a mass spectrometry system. Method 2400 may include determining, by the mass spectrometry system, one or more amino acids of the peptide sequence.
[0163] Method 2400 may include, at 2404, applying a query to one or more of a plurality of sources based on a query support data structure. The query support data structure may indicate an order of the plurality of sources to which the query is applied. The order may be based on a false discovery rate associated with each of the plurality of sources. Method 2400 may also include determining one or more permutations of the query. Applying the query to one or more of a plurality of sources based on the query support data structure may include applying each of the one or more permutations of the query to one or more of the plurality of sources, stopping additional searches when an identical match for the one or more permutations of the query is found in a first source of the plurality of sources, applying a linear label to the one or more permutations of the query associated with the identical match, and assigning the one or more permutations of the query associated with the identical match as the correct query.
[0164] Applying a query to one or more of a plurality of sources may include searching for an identical match for the query in a first source of the plurality of sources and stopping additional searches when an identical match for the query is found in the first source of the plurality of sources. The query result may include the identical match, and the label associated with the source of the plurality of sources associated with the query result may include a linear label.
[0165] Applying a query to one or more of a plurality of sources may include searching for an identical match for the query in a first source of the plurality of sources and stopping additional searches when an identical match for one or more permutations of the query is found in the first source of the plurality of sources. The query result may include the identical match, and the label associated with the source of the plurality of sources associated with the query result may include a linear label.
[0166] Applying a query to one or more of a plurality of sources may include searching for an identical match to the query in any frame of a plurality of frames of a second source of the plurality of sources, and stopping additional searches if an identical match to the query is found in any frame of the plurality of frames of the second source of the plurality of sources. The query result may include the identical match, and the label associated with the source of the plurality of sources associated with the query result may include a linear label.
[0167] Applying a query to one or more of a plurality of sources may include searching for a non-identical match to the query in a third source of the plurality of sources, and stopping additional searches if a non-identical match to the query is found in the third source of the plurality of sources. The query result may include the non-identical match, and the label associated with the source of the plurality of sources associated with the query result may include a mismatch label.
[0168] Applying a query to one or more of a plurality of sources may include searching for a homologous match to the query in a fourth source of the plurality of sources, and stopping additional searches if a homologous match to the query is found in the fourth source of the plurality of sources. The query result may include the homologous match, and the label associated with the source of the plurality of sources associated with the query result may include a homologous label.
[0169] Applying a query to one or more of a plurality of sources may include splitting the query into a set of fragments, searching for each set of fragments in a fifth source of the plurality of sources, stopping additional searches if a match for the set of fragments is found in the fifth source of the plurality of sources, and stopping additional searches if a first match for a first fragment of the set of fragments and a second match for a second fragment of the set of fragments are found in the fifth source of the plurality of sources. The query results may include matches for the set of fragments, and the labels associated with the sources of the plurality of sources associated with the query results may include cis-splice labels. The query results may include a first match for a first fragment of the set of fragments and a second match for a second fragment of the set of fragments, and the labels associated with the sources of the plurality of sources associated with the query results may include trans-splice labels.
[0170] Method 2400 may include, at 2406, determining a label associated with a source of a plurality of sources associated with the query results based on the query results.
[0171] Method 2400 may include, at 2408, applying the label to the query. Method 2400 may also include determining a source of the query based on the label. Method 2400 may also include verifying an output of a mass spectrometry system based on the source of the query.
[0172] Considering the described devices, systems, and methods, as well as their variants, certain more specifically described embodiments of the present invention are described below in this specification. However, these specifically enumerated embodiments should not be construed as having any limiting effect on any different claims that include different teachings or more general teachings described herein, nor should the "specific" embodiments be construed as being limited in any way other than the literal meaning of the language used herein.
[0173] Embodiment 1: A method for determining an estimated source of a peptide sequence of a peptide, comprising receiving the peptide sequence and determining an estimated source associated with the peptide sequence based at least in part on one or more searches of peptide sequences in one or more databases, wherein each individual search of the one or more searches has a random hit rate based at least in part on the number of random sequences found by the individual search, and the one or more searches are performed in increasing order of the random hit rate until the estimated source is determined.
[0174] Embodiment 2: An embodiment as in Embodiment 1, wherein the one or more databases include an extended human proteome database, and the extended human proteome database includes computer-readable representations of translations from messenger ribonucleic acid (RNA) and non-coding RNA.
[0175] Embodiment 3: An embodiment as in Embodiment 2, wherein the extended human proteome database includes computer-readable representations of translations from microRNA.
[0176] Embodiment 4: An embodiment according to any one of Embodiments 2 to 3, wherein the extended human proteome database includes computer-readable representations of translations from long non-coding RNA.
[0177] Embodiment 5: Any one of Embodiments 2 to 4, wherein the extended human proteome database includes a computer-readable representation of the translation of human endogenous retroviruses.
[0178] Embodiment 6: Any one of Embodiments 2 to 5, wherein one or more searches include a linear human proteome search for peptide sequences in the extended human proteome database, and when the linear human proteome search for peptide sequences in the extended human proteome database finds a peptide sequence, the putative source is a linear extended human proteome source.
[0179] Embodiment 7: An embodiment as in Embodiment 6, further including identifying whether a peptide is presumably translated from messenger RNA or non-coding RNA when the source is a linear extended human proteome source.
[0180] Embodiment 8: Any one of Embodiments 2 to 7, wherein one or more databases include a human genome database, one or more searches include a linear human genome search for the translation of the human genome database, and when the linear human genome search finds a human genome sequence from which a peptide is presumably synthesized, the putative source is a linear genome source.
[0181] Embodiment 9: An embodiment as in Embodiment 8, wherein the linear human genome search excludes the portions of the human genome where messenger RNA and non-coding RNA of the extended human proteome database are transcribed, includes the remaining portions of the human genome, and the linear human genome search includes a search for six-frame translation of the human genome.
[0182] Embodiment 10: One or more searches include a linear mismatch search for peptides having a mismatch to a peptide sequence in an extended human proteome database, and when the linear mismatch search finds a peptide sequence having a mismatch to a peptide sequence in the extended human proteome database, the putative source is a linear mismatch of the extended human proteome, any of Embodiments 2-9.
[0183] Embodiment 11: A linear mismatch search is a search for a peptide sequence having only a single mismatch to a peptide sequence, an embodiment such as Embodiment 10.
[0184] Embodiment 12: One or more databases include a non-endogenous proteome database containing a computer-readable representation of a protein translated from RNA derived from a non-endogenous organism and / or a protein synthesized by a non-endogenous organism, and one or more searches include a linear non-endogenous search for a peptide sequence in the non-endogenous proteome database, and when the linear non-endogenous search finds a peptide sequence in the non-endogenous proteome database, the putative source is a linear non-endogenous proteome source, any of Embodiments 1-11.
[0185] Embodiment 13: A non-endogenous proteome database includes a basic local alignment search tool (BLAST) database, an embodiment such as Embodiment 12.
[0186] Embodiment 14: One or more searches include a cis-splice search for a peptide fragment that can be cis-spliced to match a peptide sequence in an extended human proteome database, and when the cis-splice search finds a peptide fragment that can be cis-spliced to match a peptide sequence in the extended human proteome database, the source is a cis-spliced human proteome source, any of Embodiments 2-13.
[0187] Embodiment 15: One or more searches include a trans-splice search for a computer-readable representation of a peptide fragment that can be trans-spliced to match a peptide sequence within an extended human proteome database, and when the trans-splice search finds a computer-readable representation of a peptide fragment that can be trans-spliced to match the peptide sequence within the extended human proteome database, the source is a trans-spliced human proteome source, any of Embodiments 2 to 14.
[0188] Embodiment 16: When the trans-splice search does not find a computer-readable representation of a peptide fragment that can be trans-spliced to match the peptide sequence, the putative source is determined to be unidentified, an embodiment like Embodiment 15.
[0189] Embodiment 17: One or more databases include a human genome database, and one or more searches are sequentially ordered in a workflow as follows: a linear human proteome search for a peptide sequence within an extended human proteome database, a linear human genome search for a translation of the human genome database, a linear mismatch search for peptides having a mismatch to the peptide sequence within the extended human proteome database, and a cis-splice search for peptide fragments that can be cis-spliced to match the peptide sequence within the extended human proteome database, any of Embodiments 2 to 16.
[0190] Embodiment 18: One or more databases include a non-endogenous proteome database that includes computer-readable representations of proteins translated from RNAs derived from non-endogenous organisms and / or proteins synthesized by non-endogenous organisms, one or more searches further include a linear non-endogenous search for peptide sequences within the non-endogenous proteome database, and the linear non-endogenous search is sequentially ordered in the workflow after the linear mismatch search and before the cis-splice search, an embodiment such as Embodiment 17.
[0191] Embodiment 19: One or more searches further include a trans-splice search for peptide fragments that can be trans-spliced to match a peptide sequence within an extended human proteome database, and the trans-splice search is sequentially ordered in the workflow after the cis-splice search, any of the embodiments of Embodiments 17 - 18.
[0192] Embodiment 20: When the putative source is determined for a peptide sequence, further including stopping the advancement of the workflow to subsequent searches of one or more searches, any of the embodiments of Embodiments 17 - 19.
[0193] Embodiment 21: The peptide sequence includes at least one ambiguous residue, the method includes creating a plurality of substituted peptide sequences each including possible residues for each of the at least one ambiguous residue, determining an individual possible source for each of the plurality of substituted peptide sequences, and determining the putative source of the peptide sequence such that the individual possible source is the putative source, any of the embodiments of Embodiments 1 - 20.
[0194] Embodiment 22: The possible residues for each of the at least one ambiguous residue include leucine and isoleucine, an embodiment such as Embodiment 21.
[0195] Embodiment 23: For each possible source, determining an individual random hit rate such that the random hit rate increases as the number of random arrays found by one or more individual searches of the search increases, and determining the estimated source such that the individual random hit rate of the estimated source is the lowest individual random hit rate for each possible source. Any of the embodiments of Embodiments 21-22, further comprising.
[0196] Embodiment 24: Any of the embodiments of Embodiments 21-23, further comprising identifying one or more of the substituted peptide sequences having one or more prospects such that each of the one or more substituted peptide sequences having one or more prospects is associated with an estimated source.
[0197] Embodiment 25: Any of the embodiments of Embodiments 1-24, wherein the peptide sequence is a de novo peptide sequence determined by mass spectrometry.
[0198] Embodiment 26: A non-transitory computer-readable medium configured to communicate with one or more processors of a computer device, the non-transitory computer-readable medium including instructions thereon that, when executed by the processor, cause the computer device to receive, as input, a peptide sequence and, based at least in part on one or more searches of peptide sequences in one or more databases, determine an estimated source associated with the peptide sequence, wherein each individual search of the one or more searches has a random hit rate based at least in part on the number of random arrays found by the individual search, the one or more searches are performed in an order in which the random hit rate increases until the estimated source is determined, and, as output, provide the estimated source. Non-transitory computer-readable medium.
[0199] Embodiment 27: An embodiment as in Embodiment 26, wherein one or more databases include an extended human proteome database, and the extended human proteome database includes computer-readable representations of translations from messenger ribonucleic acid (RNA) and non-coding RNA.
[0200] Embodiment 28: An embodiment as in Embodiment 27, wherein the extended human proteome database includes computer-readable representations of translations from microRNA.
[0201] Embodiment 29: An embodiment as in any of Embodiments 27-28, wherein the extended human proteome database includes computer-readable representations of translations from long non-coding RNA.
[0202] Embodiment 30: An embodiment as in any of Embodiments 27-29, wherein the extended human proteome database includes computer-readable representations of translations from human endogenous retroviruses.
[0203] Embodiment 31: An embodiment as in any of Embodiments 27-29, wherein one or more searches include a linear human proteome search for peptide sequences within the extended human proteome database, and when a linear human proteome search for peptide sequences within the extended human proteome database finds a peptide sequence within the extended human proteome database, the putative source is a linear extended human proteome source.
[0204] Embodiment 32: An embodiment as in Embodiment 31, wherein when the instructions are executed by a processor, the computer device is caused to identify whether a peptide has been putatively translated from messenger RNA or non-coding RNA when the source is a linear extended human proteome source.
[0205] Embodiment 33: One or more databases include a human genome database, one or more searches include a linear human genome search of the translation of the human genome database, and when the linear human genome search finds a human genome sequence from which a peptide is presumably synthesized, the presumed source is a linear genome source, any of Embodiments 27 to 32.
[0206] Embodiment 34: A linear human genome search excludes the portions of the human genome where messenger RNA and non-coding RNA of an extended human proteome database are transcribed, and includes the remaining portions of the human genome, such as in Embodiment 33.
[0207] Embodiment 35: One or more searches include a linear mismatch search for peptides having a mismatch to a peptide sequence within an extended human proteome database, and when the linear mismatch search finds a peptide sequence having a mismatch to a peptide sequence within the extended human proteome database, the presumed source is a linear mismatch of the extended human proteome, any of Embodiments 27 to 34.
[0208] Embodiment 36: A linear mismatch search is a search for peptide sequences having only a single mismatch to a peptide sequence, such as in Embodiment 35.
[0209] Embodiment 37: One or more databases include a non-endogenous proteome database containing computer-readable representations of proteins translated from RNA derived from non-endogenous organisms and / or proteins synthesized by non-endogenous organisms, one or more searches include a linear non-endogenous search for peptide sequences within the non-endogenous proteome database, and when the linear non-endogenous search finds a peptide sequence within the non-endogenous proteome database, the presumed source is a linear non-endogenous proteome source, any of Embodiments 26 to 36.
[0210] Embodiment 38: An embodiment such as Embodiment 37, wherein the non-endogenous proteome database includes a basic local alignment search tool (BLAST) database.
[0211] Embodiment 39: One or more searches include a cis-splice search for peptide fragments that can be cis-spliced to match a peptide sequence within an extended human proteome database, and when the cis-splice search finds a peptide fragment that can be cis-spliced to match the peptide sequence within the extended human proteome database, the source is a cis-spliced human proteome source, in any of Embodiments 27-38.
[0212] Embodiment 40: One or more searches include a trans-splice search for a computer-readable representation of a peptide fragment that can be trans-spliced to match a peptide sequence within an extended human proteome database, and when the trans-splice search finds a computer-readable representation of a peptide fragment that can be trans-spliced to match the peptide sequence within the extended human proteome database, the source is a trans-spliced human proteome source, in any of Embodiments 27-39.
[0213] Embodiment 41: An embodiment such as Embodiment 40, wherein when the trans-splice search does not find a computer-readable representation of a peptide fragment that can be trans-spliced to match the peptide sequence, the putative source is determined to be unidentified.
[0214] Embodiment 42: One or more databases include the human genome, and one or more searches are sequentially ordered in a workflow as follows: a linear human proteome search for peptide sequences in an extended human proteome database, a linear human genome search for the translation of the human genome database, a linear mismatch search for peptides having a mismatch to peptide sequences in the extended human proteome database, and a cis-splice search for peptide fragments that can be cis-spliced to match a peptide sequence in the extended human proteome database, according to any of Embodiments 27 to 41.
[0215] Embodiment 43: One or more databases include a non-endogenous proteome database that includes computer-readable representations of proteins translated from RNA derived from non-endogenous organisms and / or proteins synthesized by non-endogenous organisms, and one or more searches further include a linear non-endogenous search for peptide sequences in the non-endogenous proteome database, where the linear non-endogenous search is sequentially ordered in the workflow after the linear mismatch search and before the cis-splice search, according to an embodiment such as Embodiment 42.
[0216] Embodiment 44: One or more searches further include a trans-splice search for peptide fragments that can be trans-spliced to match a peptide sequence in the extended human proteome database, where the trans-splice search is sequentially ordered in the workflow after the cis-splice search, according to any of Embodiments 42 to 43.
[0217] Embodiment 45: Instructions, when executed by a processor, cause a computer device to stop the advancement of the workflow to subsequent searches of one or more searches when an estimated source is determined for a peptide sequence, according to any of Embodiments 42 to 44.
[0218] Embodiment 46: The peptide sequence contains at least one ambiguous residue, and when the instruction is executed by a processor, causes the computer device to create a plurality of substituted peptide sequences each containing possible residues for each of the at least one ambiguous residue, determine an individual possible source for each of the plurality of substituted peptide sequences, and determine the putative source of the peptide sequence such that the putative source is the individual possible source, any of Embodiments 26 to 45.
[0219] Embodiment 47: An embodiment like Embodiment 46, wherein the possible residues for each of the at least one ambiguous residue include leucine and isoleucine.
[0220] Embodiment 48: When the instruction is executed by a processor, causes the computer device to determine an individual random hit rate for each of the individual possible sources such that the random hit rate increases as the number of random sequences found by each individual search of one or more searches increases, and determine the putative source such that the individual random hit rate of the putative source is the lowest individual random hit rate for each of the possible sources, any of Embodiments 46 to 47.
[0221] Embodiment 49: When the instruction is executed by a processor, causes the computer device to identify one or more promising substituted peptide sequences of the plurality of substituted peptide sequences such that each of the one or more promising substituted peptide sequences is associated with the putative source, any of Embodiments 46 to 48.
[0222] Embodiment 50: An embodiment according to any of Embodiments 26 to 49, wherein the peptide sequence is a de novo peptide sequence determined by mass spectrometry.
[0223] Embodiment 51: A method of ordering a peptide source assignment workflow, comprising creating a plurality of random peptide arrays, determining a plurality of peptide source search steps, searching each of the plurality of peptide source search steps for each of the plurality of random peptide arrays, for each of the plurality of peptide source search steps, determining a random hit rate for an individual search step of the plurality of peptide source search steps, at least partially based on the number of the plurality of random peptide arrays found by the individual search step, and in the peptide source assignment workflow, ordering the peptide source search steps from the lowest random hit rate to the highest random hit rate.
[0224] Embodiment 52: An embodiment such as Embodiment 51, wherein the random peptide array includes a random array that uniformly samples all amino acids.
[0225] Embodiment 53: Any of Embodiments 51-52, wherein the random peptide array includes an array having an amino acid frequency that matches the frequency of amino acids found in vertebrates.
[0226] Embodiment 54: Any of Embodiments 51-53, wherein each peptide of the random peptide array has a length of 8 to 14 amino acids.
[0227] Embodiment 55: Any of Embodiments 51-54, wherein each peptide of the random peptide array has a length of 9 to 14 amino acids, 10 to 14 amino acids, 11 to 14 amino acids, 12 to 14 amino acids, 13 to 14 amino acids, 8 to 13 amino acids, 8 to 12 amino acids, 8 to 11 amino acids, 8 to 10 amino acids, 8 to 9 amino acids, 9 to 13 amino acids, 9 to 12 amino acids, 9 to 11 amino acids, 9 to 10 amino acids, 10 to 13 amino acids, 10 to 12 amino acids, 10 to 11 amino acids, 11 to 13 amino acids, 11 to 12 amino acids, or 12 to 13 amino acids.
[0228] Embodiment 56: Any one of Embodiments 51 to 55, wherein a plurality of peptide source search steps include a linear human proteome search for peptide sequences in an extended human proteome database, and the extended human proteome database includes computer-readable representations of translations from messenger ribonucleic acid (RNA) and non-coding RNA.
[0229] Embodiment 57: An embodiment such as Embodiment 56, wherein the extended human proteome database includes computer-readable representations of translations from microRNA.
[0230] Embodiment 58: Any one of Embodiments 56 to 57, wherein the extended human proteome database includes computer-readable representations of translations from long non-coding RNA.
[0231] Embodiment 59: Any one of Embodiments 56 to 58, wherein the extended human proteome database includes computer-readable representations of translations of human endogenous retroviruses.
[0232] Embodiment 60: Any one of Embodiments 56 to 59, wherein a plurality of peptide source search steps include a linear human genome search of translations of the human genome database.
[0233] Embodiment 61: An embodiment such as Embodiment 60, wherein the linear human genome search excludes portions of the human genome where messenger RNA and non-coding RNA of the extended human proteome database are transcribed, and includes the remaining portions of the human genome.
[0234] Embodiment 62: Any one of Embodiments 51 to 61, wherein a plurality of peptide source search steps include a linear mismatch search for peptides having a mismatch with respect to peptide sequences in an extended human proteome database, and the extended human proteome database includes computer-readable representations of translations from messenger ribonucleic acid (RNA) and non-coding RNA.
[0235] Embodiment 63: Any of Embodiments 51 - 62, wherein a plurality of peptide source search steps includes a linear non - endogenous search for peptide sequences in a non - endogenous proteome database, and the non - endogenous proteome database includes a computer - readable representation of proteins translated from RNA from non - endogenous organisms and / or proteins synthesized by non - endogenous organisms.
[0236] Embodiment 64: An embodiment such as Embodiment 63, wherein the non - endogenous proteome database includes a Basic Local Alignment Search Tool (BLAST) database.
[0237] Embodiment 65: Any of Embodiments 51 - 64, wherein a plurality of peptide source search steps includes a cis - splicing search for peptide fragments that can be cis - spliced to match a peptide sequence within an extended human proteome database, and the extended human proteome database includes a computer - readable representation of translations from messenger ribonucleic acid (RNA) and non - coding RNA.
[0238] Embodiment 66: Any of Embodiments 51 - 65, wherein a plurality of peptide source search steps includes a trans - splicing search for peptide fragments that can be trans - spliced to match a peptide sequence within an extended human proteome database, and the extended human proteome database includes a computer - readable representation of translations from messenger ribonucleic acid (RNA) and non - coding RNA.
[0239] Embodiment 67: Any of Embodiments 51 - 66, wherein when a peptide is not assigned a peptide source by any of the plurality of peptide source search steps, the peptide source assignment workflow ends with the unassigned peptide.
[0240] Embodiment 68: An embodiment according to any of Embodiments 51 to 67, wherein the peptide source assignment workflow includes the following searches sequentially ordered as follows: a linear human proteome search for peptide sequences in an extended human proteome database, a linear human genome search for the translation of the human genome database, a linear mismatch search for peptides having a mismatch to the peptide sequences in the extended human proteome database, and a cis-splice search for peptide fragments that can be cis-spliced to match the peptide sequences in the extended human proteome database.
[0241] Embodiment 69: An embodiment such as Embodiment 68, wherein the peptide source assignment workflow includes a linear non-endogenous search for peptide sequences in a non-endogenous proteome database, and the linear non-endogenous search is sequentially ordered after the linear mismatch search and before the cis-splice search within the peptide assignment workflow.
[0242] Embodiment 70: An embodiment according to any of Embodiments 68 to 69, wherein the peptide source assignment workflow includes a trans-splice search for peptide fragments that can be trans-spliced to match the peptide sequences in the extended human proteome database, and the trans-splice search is sequentially ordered after the cis-splice search within the peptide assignment workflow.
[0243] Embodiment 71: A non - transitory computer - readable medium configured to communicate with one or more processors of a computer device, the non - transitory computer - readable medium including instructions thereon, which when executed by the processor cause the computer device to: receive, as input, a plurality of peptide source search steps; create a plurality of random peptide arrays; for each of the plurality of peptide source search steps, search for each of the plurality of random peptide arrays; for each of the plurality of peptide source search steps, determine the random hit rate for an individual search step of the plurality of peptide source search steps, at least in part based on the number of the plurality of random peptide arrays found by the individual search step; order the peptide source search steps in a peptide source assignment workflow from the lowest random hit rate to the highest random hit rate; and provide, as output, the peptide source assignment workflow.
[0244] Embodiment 72: An embodiment such as Embodiment 71, wherein the random peptide array includes a random array that uniformly samples all amino acids.
[0245] Embodiment 73: An embodiment according to any one of Embodiments 71 - 72, wherein the random peptide array includes an array having an amino acid frequency that matches the frequency of amino acids found in vertebrates.
[0246] Embodiment 74: An embodiment according to any one of Embodiments 71 - 73, wherein each peptide of the random peptide array has a length of 8 - 14 amino acids.
[0247] Embodiment 75: Any of Embodiments 71 to 74, wherein each peptide of the random peptide sequences has a length of 9 to 14 amino acids, 10 to 14 amino acids, 11 to 14 amino acids, 12 to 14 amino acids, 13 to 14 amino acids, 8 to 13 amino acids, 8 to 12 amino acids, 8 to 11 amino acids, 8 to 10 amino acids, 8 to 9 amino acids, 9 to 13 amino acids, 9 to 12 amino acids, 9 to 11 amino acids, 9 to 10 amino acids, 10 to 13 amino acids, 10 to 12 amino acids, 10 to 11 amino acids, 11 to 13 amino acids, 11 to 12 amino acids, or 12 to 13 amino acids.
[0248] Embodiment 76: Any of Embodiments 71 to 75, wherein the plurality of peptide source search steps includes a linear human proteome search for peptide sequences in an extended human proteome database, and the extended human proteome database includes computer-readable representations of translations from messenger ribonucleic acid (RNA) and non-coding RNA.
[0249] Embodiment 77: An embodiment such as Embodiment 76, wherein the extended human proteome database includes computer-readable representations of translations from microRNA.
[0250] Embodiment 78: Any of Embodiments 76 to 77, wherein the extended human proteome database includes computer-readable representations of translations from long non-coding RNA.
[0251] Embodiment 79: Any of Embodiments 76 to 78, wherein the extended human proteome database includes computer-readable representations of translations of human endogenous retroviruses.
[0252] Embodiment 80: Any of Embodiments 76 to 79, wherein the plurality of peptide source search steps includes a linear human genome search of the translation of the human genome database.
[0253] Embodiment 81: An embodiment like Embodiment 80, in which the linear human genome search excludes the portions of the human genome where messenger RNA and non-coding RNA of the extended human proteome database are transcribed, and includes the remaining portions of the human genome.
[0254] Embodiment 82: An embodiment according to any one of Embodiments 71 to 81, in which the plurality of peptide source search steps includes a linear mismatch search for peptides having a mismatch to peptide sequences in the extended human proteome database, and the extended human proteome database includes computer-readable representations of translations from messenger ribonucleic acid (RNA) and non-coding RNA.
[0255] Embodiment 83: An embodiment according to any one of Embodiments 71 to 82, in which the plurality of peptide source search steps includes a linear non-endogenous search for peptide sequences in a non-endogenous proteome database, and the non-endogenous proteome database includes computer-readable representations of proteins translated from RNA derived from non-endogenous organisms and / or proteins synthesized by non-endogenous organisms.
[0256] Embodiment 84: An embodiment like Embodiment 83, in which the non-endogenous proteome database includes a basic local alignment search tool (BLAST) database.
[0257] Embodiment 85: An embodiment according to any one of Embodiments 71 to 84, in which the plurality of peptide source search steps includes a cis-splice search for peptide fragments that can be cis-spliced to match a peptide sequence in the extended human proteome database, and the extended human proteome database includes computer-readable representations of translations from messenger ribonucleic acid (RNA) and non-coding RNA.
[0258] Embodiment 86: A plurality of peptide source search steps includes a trans-splice search for peptide fragments that can be trans-spliced to match a peptide sequence in an extended human proteome database, where the extended human proteome database includes a computer-readable representation of translation from messenger ribonucleic acid (RNA) and non-coding RNA, according to any of Embodiments 71-85.
[0259] Embodiment 87: When a peptide is not assigned a peptide source by any of the plurality of peptide source search steps, the peptide source assignment workflow ends with the unassigned peptide, according to any of Embodiments 71-86.
[0260] Embodiment 88: The peptide source assignment workflow includes the following searches ordered sequentially as follows: a linear human proteome search for peptide sequences in the extended human proteome database, a linear human genome search for translation of the human genome database, a linear mismatch search for peptides having a mismatch to a peptide sequence in the extended human proteome database, a linear non-endogenous search for peptide sequences in the non-endogenous proteome database, a cis-splice search for peptide fragments that can be cis-spliced to match a peptide sequence in the extended human proteome database, and a trans-splice search for peptide fragments that can be trans-spliced to match a peptide sequence in the extended human proteome database, according to any of Embodiments 71-87.
[0261] Embodiment 89: The peptide source assignment workflow includes a linear non-endogenous search for peptide sequences in the non-endogenous proteome database, where the linear non-endogenous search is sequentially ordered after the linear mismatch search and before the cis-splice search within the peptide assignment workflow, as in Embodiment 88.
[0262] Embodiment 90: The peptide source assignment workflow includes a trans-splice search for peptide fragments that can be trans-spliced to match a peptide sequence within an extended human proteome database, and the trans-splice search is sequentially ordered after a cis-splice search within the peptide assignment workflow, according to any of embodiments 88-89.
[0263] Embodiment 91: A method comprising creating a plurality of simulated random queries, determining the number of matches associated with each source based on applying the plurality of simulated random queries to each source of a plurality of sources, determining the false discovery rate associated with each source based on the number of matches associated with each source, and creating a query support data structure configured to facilitate the application of new queries to the plurality of sources based on the false discovery rate.
[0264] Embodiment 92: An embodiment such as Embodiment 91, wherein creating a plurality of simulated random queries includes at least one of creating a plurality of uniform random queries or creating a plurality of weighted random queries.
[0265] Embodiment 93: Any of Embodiments 91-92, wherein the plurality of simulated random queries includes a plurality of simulated random text strings.
[0266] Embodiment 94: Any of Embodiments 91-93, wherein the plurality of simulated random queries includes a plurality of simulated random peptide sequences.
[0267] Embodiment 95: Any of Embodiments 91-94, wherein determining the false discovery rate associated with each source based on the number of matches associated with each source includes a function of the number of matches and the number of the plurality of simulated random queries.
[0268] Embodiment 96: Any of Embodiments 91 - 95, which includes dividing the number of matches associated with each source by the number of multiple simulation random queries to determine the false discovery rate associated with each source based on the number of matches associated with each source.
[0269] Embodiment 97: A method that includes receiving a query, applying the query to one or more of a plurality of sources based on a query support data structure, determining a label associated with the source of the plurality of sources associated with the query result based on the query result, and applying the label to the query.
[0270] Embodiment 98: An embodiment such as Embodiment 97 where the query includes a text string.
[0271] Embodiment 99: Any of Embodiments 97 - 98 where the query includes a peptide sequence.
[0272] Embodiment 100: An embodiment such as Embodiment 99 where receiving the query includes receiving a peptide sequence from a mass spectrometry system.
[0273] Embodiment 101: Any of Embodiments 97 - 100 that further includes determining one or more amino acids of the peptide sequence by a mass spectrometry system.
[0274] Embodiment 102: Any of Embodiments 97 - 101 where the query support data structure indicates the order of a plurality of sources for applying the query, and the order is based on the false discovery rate associated with each source of the plurality of sources.
[0275] Embodiment 103: Any of Embodiments 97 - 102 that further includes determining one or more replacements of the query.
[0276] Embodiment 104: Based on a query support data structure, applying a query to one or more of a plurality of sources includes applying each permutation of one or more permutations of the query to one or more of the plurality of sources, stopping additional searches when an identical match for one or more permutations of the query is found in a first source of the plurality of sources, applying a linear label to one or more permutations of the query associated with the identical match, and assigning one or more permutations of the query associated with the identical match as the correct query, an embodiment such as Embodiment 103.
[0277] Embodiment 105: Applying a query to one or more of a plurality of sources includes searching for an identical match for the query in a first source of the plurality of sources and stopping additional searches when an identical match for the query is found in the first source of the plurality of sources, any of Embodiments 97 to 104.
[0278] Embodiment 106: An embodiment such as Embodiment 105, wherein the query result includes an identical match and the label associated with the source of the plurality of sources associated with the query result includes a linear label.
[0279] Embodiment 107: Applying a query to one or more of a plurality of sources includes searching for an identical match for the query in a first source of the plurality of sources and stopping additional searches when an identical match for one or more permutations of the query is found in the first source of the plurality of sources, any of Embodiments 97 to 106.
[0280] Embodiment 108: An embodiment such as Embodiment 107, wherein the query result includes an identical match and the label associated with the source of the plurality of sources associated with the query result includes a linear label.
[0281] Embodiment 109: Applying a query to one or more of a plurality of sources includes searching for an identical match to the query in any frame of a plurality of frames of a second source of the plurality of sources, and stopping additional searches if an identical match to the query is found in any frame of the plurality of frames of the second source of the plurality of sources. An embodiment like Embodiment 107.
[0282] Embodiment 110: An embodiment like Embodiment 109, wherein the query result includes an identical match and the label associated with the source of the plurality of sources associated with the query result includes a linear label.
[0283] Embodiment 111: Applying a query to one or more of a plurality of sources includes searching for a non-identical match to the query in a third source of the plurality of sources, and stopping additional searches if a non-identical match to the query is found in the third source of the plurality of sources. An embodiment like Embodiment 109.
[0284] Embodiment 112: An embodiment like Embodiment 111, wherein the query result includes a non-identical match and the label associated with the source of the plurality of sources associated with the query result includes a mismatch label.
[0285] Embodiment 113: Applying a query to one or more of a plurality of sources includes searching for a homologous match to the query in a fourth source of the plurality of sources, and stopping additional searches if a homologous match to the query is found in the fourth source of the plurality of sources. An embodiment like Embodiment 111.
[0286] Embodiment 114: An embodiment like Embodiment 113, wherein the query result includes a homologous match and the label associated with the source of the plurality of sources associated with the query result includes a homologous label.
[0287] Embodiment 115: An embodiment such as Embodiment 113, which includes applying a query to one or more of a plurality of sources, splitting the query into a set of fragments, searching for each set of fragments in a fifth source of the plurality of sources, stopping additional searches when a match for the set of fragments is found in the fifth source of the plurality of sources, and stopping additional searches when a first match for a first fragment of the set of fragments and a second match for a second fragment of the set of fragments are found in the fifth source of the plurality of sources.
[0288] Embodiment 116: An embodiment such as Embodiment 115, where the query result includes a match for a set of fragments, and the label associated with the source of the plurality of sources associated with the query result includes a cis-splice label.
[0289] Embodiment 117: An embodiment such as Embodiment 115, where the query result includes a first match for a first fragment of a set of fragments and a second match for a second fragment of the set of fragments, and the label associated with the source of the plurality of sources associated with the query result includes a trans-splice label.
[0290] Embodiment 118: Any of Embodiments 97 - 117, further including determining the source of the query based on the label.
[0291] Embodiment 119: An embodiment such as Embodiment 118, further including verifying the output of a mass spectrometry system based on the source of the query.
[0292] Embodiment 120: An apparatus including one or more processors and memory-stored processor-executable instructions that cause the apparatus to perform any of Embodiments 91 - 119 when executed by the one or more processors.
[0293] Embodiment 121: One or more non-transitory computer-readable media that store processor-executable instructions that, when executed by a processor, cause the processor to perform any of Embodiments 91-119. Embodiment 122: A computer device and a system including a plurality of sources configured to perform any of Embodiments 91-119.
[0294] Although specific configurations have been described, the configurations in this specification are intended to be configurations that are not limiting in all respects and are possible. Therefore, this scope is not intended to be limited to the specific configurations shown.
[0295] Unless otherwise explicitly stated, it is never intended that any method shown in this specification be construed as requiring that its steps be performed in a specific order. Thus, if a method claim does not actually limit the order in which the steps should be followed, or if the steps are not specifically stated to be limited to a particular order in the claim or description, in no way is the order intended to be inferred. This holds for any possible implicit basis for interpretation, including logical issues regarding the sequence of steps or the flow of operations, the plain meaning derived from grammatical construction or punctuation, the number or type of configurations described in this specification.
[0296] One of ordinary skill in the art will be able to recognize or confirm many equivalents to the specific embodiments of the methods and compositions described herein using only routine experimentation. Such equivalents are intended to be encompassed by the following claims. Table 1
Table 1-1
Table 1-2
Table 1-3
Table 1-4
Table 1-5
Table 1-6
Table 1-7
Table 1-8
Table 1-9
Table 1-10
Table 1-11
Table 1-12
Table 1-13
Table 1-14
Table 1-15
Table 1-16
Table 1-17
Table 1-18
Table 1-19
Table 1-20
Table 1-21
Table 1-22
Table 1-23
Table 1-24
Table 1-25
Table 1-26
Table 1-27
Table 1-28
Table 1-29
Table 1-30
Table 1-31
Table 1-32
Table 1-33
Table 1-34
Table 1-35
Table 1-36
Table 1-37
Table 1-38
Table 1-39
Table 1-40
Table 1-41
Table 1-42
Table 1-43
Table 1-44
Table 1-45
Table 1-46
Table 1-47
Table 1-48
Table 1-49
Table 1-50
Table 1-51
Table 1-52
Table 1-53
Table 1-54
Table 1-55
Table 1-56
Table 1-57
Table 1-58
Table 1-59
Table 1-60
Table 1-61
Table 1-62
Table 1-63
Table 1-64
Table 1-65
Table 1-66
Table 1-67
Table 1-68
Table 1-69
Table 1-70
Table 1-71
Table 1-72
Table 1-73
Table 1-74
Table 1-75
Table 1-76
Table 1-77
Table 1-78
Table 1-79
Table 1-80
Table 1-81
Table 1-82
Table 1-83
Table 1-84
Table 1-85
Table 1-86
Table 1-87
Table 1-88
Table 1-89
Table 1-90
Table 1-91
Table 1-92
Table 1-93
Table 1-94
Table 1-95
Table 1-96
Table 1-97
Table 1-98
Table 1-99
Table 1-100
Table 1-101
Table 1-102
Table 1-103
Table 1-104
Table 1-105
Table 1-106
Table 1-107
Table 1-108
Table 1-109
Table 1-110
Table 1-111
Table 1-112
Table 1-113
Table 1-114
Table 1-115
Table 1-116
Table 1-117
Table 1-118
Table 1-119
Table 1-120
Table 1-121
Table 1-122
Table 1-123
Table 1-124
Table 1-125
Table 1-126
Table 1-127
Table 1-128
Table 1-129
Table 1-130
Table 1-131
Table 1-132
Claims
1. 1. A method for sequencing a peptide source assignment workflow, comprising: generating a plurality of random peptide sequences; Determining a plurality of peptide source search steps; searching for each of the plurality of random peptide sequences by each of the plurality of peptide source searching steps; determining, for each of the plurality of peptide source searching steps, a random hit rate for an individual search step of the plurality of peptide source searching steps based at least in part on a number of the plurality of random peptide sequences found by the individual search step; and ordering said peptide source search steps from lowest random hit rate to highest random hit rate in said peptide source assignment workflow. A method comprising:
2. The method of claim 1 , wherein the random peptide sequences comprise random sequences that uniformly sample all amino acids.
3. The method of claim 1 , wherein the random peptide sequences comprise sequences having amino acid frequencies that match the frequencies of amino acids found in vertebrates.
4. The method of claim 1, wherein each peptide of the random peptide sequence comprises a length of 8 to 14 amino acids.
5. 2. The method of claim 1, wherein each peptide of the random peptide sequence comprises a length of 9-14 amino acids, 10-14 amino acids, 11-14 amino acids, 12-14 amino acids, 13-14 amino acids, 8-13 amino acids, 8-12 amino acids, 8-11 amino acids, 8-10 amino acids, 8-9 amino acids, 9-13 amino acids, 9-12 amino acids, 9-11 amino acids, 9-10 amino acids, 10-13 amino acids, 10-12 amino acids, 10-11 amino acids, 11-13 amino acids, 11-12 amino acids, or 12-13 amino acids.
6. said step of searching multiple peptide sources comprises a linear human proteome search for peptide sequences in an extended human proteome database; the expanded human proteome database includes a computer-readable representation of translations from messenger ribonucleic acid (RNA) and non-coding RNA; The method of claim 1.
7. The method of claim 6 , wherein the extended human proteome database comprises a computer-readable representation of translations from microRNAs.
8. 7. The method of claim 6, wherein the extended human proteome database comprises a computer-readable representation of translations from long non-coding RNA.
9. The method of claim 6, wherein the extended human proteome database comprises computer-readable representations of translations of human endogenous retroviruses.
10. 7. The method of claim 6, wherein the step of searching multiple peptide sources comprises a linear human genome search of a translation of a human genome database.
11. the linear human genome search excludes portions of the human genome from which the messenger RNA and the non-coding RNA are transcribed, and includes the remaining portions of the human genome in the extended human proteome database; The method of claim 10.
12. said step of searching multiple peptide sources comprises a linear mismatch search for peptides having mismatches to said peptide sequence in an extended human proteome database; the expanded human proteome database includes a computer-readable representation of translations from messenger ribonucleic acid (RNA) and non-coding RNA; The method of claim 1.
13. the step of searching multiple peptide sources comprises a linear non-endogenous search for the peptide sequence in a non-endogenous proteome database; the non-endogenous proteome database comprises computer readable representations of proteins translated from RNA derived from the non-endogenous organism and / or synthesized by the non-endogenous organism; The method of claim 1.
14. The method of claim 13 , wherein the non-endogenous proteome database comprises a basic local alignment search tool (BLAST) database.
15. said step of searching multiple peptide sources comprises a cis-splice search in an extended human proteome database for peptide fragments that can be cis-spliced to match said peptide sequence; the expanded human proteome database includes a computer-readable representation of translations from messenger ribonucleic acid (RNA) and non-coding RNA; The method of claim 1.
16. said step of searching multiple peptide sources comprises a trans-splice search in an extended human proteome database for peptide fragments that can be trans-spliced to match said peptide sequence; the expanded human proteome database includes a computer-readable representation of translations from messenger ribonucleic acid (RNA) and non-coding RNA; The method of claim 1.
17. 2. The method of claim 1, wherein if a peptide cannot be assigned a peptide source by any of the multiple peptide source searching steps, the peptide source assignment workflow ends with the peptide unassigned.
18. The peptide source assignment workflow involves the following searches, sequenced sequentially as follows: a linear human proteome search for said peptide sequence in said extended human proteome database; Linear human genome search of translations of the human genome database; a linear mismatch search for peptides having mismatches to the peptide sequence in the extended human proteome database; and a cis-splice search in the extended human proteome database for peptide fragments that can be cis-spliced to match the peptide sequence. The method of claim 1 , comprising:
19. the peptide source assignment workflow comprises a linear non-endogenous search for the peptide sequence in a non-endogenous proteome database; the linear non-endogenous search is sequentially ordered within the peptide assignment workflow after the linear mismatch search and before the cis-splice search; 20. The method of claim 18.
20. the peptide source assignment workflow comprises a trans-splice search in the extended human proteome database for peptide fragments that can be trans-spliced to match the peptide sequence; the trans-splice search is sequentially ordered after the cis-splice search within the peptide assignment workflow; 20. The method of claim 18.
21. A non-transitory computer-readable medium configured to be in communication with one or more processors of a computing device, the non-transitory computer-readable medium including instructions thereon that, when executed by the processor, cause the computing device to: It takes as input multiple peptide source search steps, Generate multiple random peptide sequences, searching each of the plurality of random peptide sequences by each of the plurality of peptide source searching steps; determining, for each of the plurality of peptide source searching steps, a random hit rate for an individual search step of the plurality of peptide source searching steps based at least in part on a number of the plurality of random peptide sequences found by the individual search step; ordering the peptide source search steps in a peptide source assignment workflow from lowest random hit rate to highest random hit rate; and A non-transitory computer readable medium providing, as an output, said peptide source assignment workflow.
22. 22. The non-transitory computer readable medium of claim 21, wherein the random peptide sequences comprise random sequences that uniformly sample all amino acids.
23. 22. The non-transitory computer readable medium of claim 21, wherein the random peptide sequences comprise sequences having amino acid frequencies that match amino acid frequencies found in vertebrates.
24. 22. The non-transitory computer readable medium of claim 21, wherein each peptide of the random peptide sequence comprises a length of 8 to 14 amino acids.
25. 22. The non-transitory computer readable medium of claim 21 , wherein each peptide of the random peptide sequence comprises a length of 9-14 amino acids, 10-14 amino acids, 11-14 amino acids, 12-14 amino acids, 13-14 amino acids, 8-13 amino acids, 8-12 amino acids, 8-11 amino acids, 8-10 amino acids, 8-9 amino acids, 9-13 amino acids, 9-12 amino acids, 9-11 amino acids, 9-10 amino acids, 10-13 amino acids, 10-12 amino acids, 10-11 amino acids, 11-13 amino acids, 11-12 amino acids, or 12-13 amino acids.
26. said step of searching multiple peptide sources comprises a linear human proteome search for peptide sequences in an extended human proteome database; the expanded human proteome database includes a computer-readable representation of translations from messenger ribonucleic acid (RNA) and non-coding RNA; 22. The non-transitory computer readable medium of claim 21.
27. 27. The non-transitory computer readable medium of claim 26, wherein the extended human proteome database comprises a computer readable representation of translations from microRNAs.
28. 27. The non-transitory computer readable medium of claim 26, wherein the extended human proteome database comprises a computer readable representation of a translation from long non-coding RNA.
29. 27. The non-transitory computer readable medium of claim 26, wherein the extended human proteome database comprises computer readable representations of translations of human endogenous retroviruses.
30. 27. The non-transitory computer readable medium of claim 26, wherein the step of searching multiple peptide sources comprises a linear human genome search of a translation of a human genome database.
31. 31. The non-transitory computer readable medium of claim 30, wherein the linear human genome search excludes portions of the human genome into which the messenger RNA and the non-coding RNA of the extended human proteome database are transcribed and includes remaining portions of the human genome.
32. said step of searching multiple peptide sources comprises a linear mismatch search for peptides having mismatches to said peptide sequence in an extended human proteome database; the expanded human proteome database includes a computer-readable representation of translations from messenger ribonucleic acid (RNA) and non-coding RNA; 22. The non-transitory computer readable medium of claim 21.
33. the step of searching multiple peptide sources comprises a linear non-endogenous search for the peptide sequence in a non-endogenous proteome database; the non-endogenous proteome database comprises computer readable representations of proteins translated from RNA derived from the non-endogenous organism and / or synthesized by the non-endogenous organism; 22. The non-transitory computer readable medium of claim 21.
34. 34. The non-transitory computer readable medium of claim 33, wherein the non-endogenous proteome database comprises a basic local alignment search tool (BLAST) database.
35. said step of searching multiple peptide sources comprises a cis-splice search in an extended human proteome database for peptide fragments that can be cis-spliced to match said peptide sequence; the expanded human proteome database includes a computer-readable representation of translations from messenger ribonucleic acid (RNA) and non-coding RNA; 22. The non-transitory computer readable medium of claim 21.
36. said step of searching multiple peptide sources comprises a trans-splice search in an extended human proteome database for peptide fragments that can be trans-spliced to match said peptide sequence; the expanded human proteome database includes a computer-readable representation of translations from messenger ribonucleic acid (RNA) and non-coding RNA; 22. The non-transitory computer readable medium of claim 21.
37. 22. The non-transitory computer readable medium of claim 21, wherein if a peptide is not assigned a peptide source by any of the multiple peptide source searching steps, the peptide source assignment workflow ends with the peptide unassigned.
38. The peptide source assignment workflow involves the following searches, sequenced sequentially as follows: a linear human proteome search for said peptide sequence in said extended human proteome database; Linear human genome search of translations of the human genome database; a linear mismatch search for peptides having mismatches to the peptide sequence in the extended human proteome database; and a cis-splice search in the extended human proteome database for peptide fragments that can be cis-spliced to match the peptide sequence.
22. The non-transitory computer readable medium of claim 21, comprising:
39. the peptide source assignment workflow comprises a linear non-endogenous search for the peptide sequence in a non-endogenous proteome database; the linear non-endogenous search is sequentially ordered within the peptide assignment workflow after the linear mismatch search and before the cis-splice search; 40. The non-transitory computer readable medium of claim 38.
40. the peptide source assignment workflow comprises a trans-splice search in the extended human proteome database for peptide fragments that can be trans-spliced to match the peptide sequence; causing the trans-splice search to be sequentially ordered after the cis-splice search within the peptide assignment workflow; 40. The non-transitory computer readable medium of claim 38.
Citation Information
Patent Citations
Identification quality evaluation method and apparatus for endogenous modified peptide
JP2019185224A