sequencing data encoded peptides according to tandem mass spectrometry
By employing a two-stage sequencing method and a sequencing method based on the highest strength tag, the problem of low sequencing efficiency of data-encoded peptides in existing technologies has been solved, enabling efficient and accurate sequencing of non-natural amino acids and random peptide sequences.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- THE HONG KONG POLYTECHNIC UNIV
- Filing Date
- 2023-04-10
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies struggle to efficiently and economically sequence data-encoded peptides, especially when using non-natural amino acids and random peptide sequences, as existing algorithms lack training data and adaptability.
A two-stage sequencing method and a sequencing method based on the highest intensity tag were used. Candidate sequences were identified and peptide sequences were estimated by preprocessing mass spectrometry data. Invalid sequences were removed by checksum and cleanup, and the peptide sequences were finally determined.
It improves the efficiency and accuracy of data-encoded peptide sequencing, reduces sequencing time and cost, and adapts to the complexity of non-natural amino acid and random peptide sequences.
Smart Images

Figure CN116895335B_ABST
Abstract
Description
[0001] Citation of relevant applications
[0002] This application claims priority and benefit to U.S. Provisional Patent Application No. 63 / 362,757, filed April 11, 2022, the disclosure of which is incorporated herein by reference in its entirety.
[0003] List of abbreviations
[0004] AAC amino acid combination
[0005] Directed Acyclic Graph (DAG)
[0006] DNA deoxyribonucleic acid
[0007] MS / MS tandem mass spectrometry
[0008] m / z mass-to-charge ratio
[0009] OFF frequency offset function
[0010] RS Reed-Solomon
[0011] TIC total ion chromatogram Technical Field
[0012] This disclosure generally relates to the sequencing of data-encoded peptides. Specifically, this disclosure relates to a method for sequencing data-encoded peptides based on experimental spectra obtained from analysis of the data-encoded peptides by mass spectrometry or MS / MS. Background Technology
[0013] Using peptides for data storage offers several advantages over using DNA. The most significant advantage is the higher storage density achievable with peptides, as non-natural amino acids can also be used to form peptides, thus expanding the alphabet size for encoding digital data into amino acids. Furthermore, due to the greater durability of peptides compared to DNA, data storage using peptides offers a longer storage time than data storage using DNA.
[0014] Peptide sequencing is a crucial element in decoding data stored in data-encoded peptides. Mass spectrometry (especially MS / MS) is particularly useful for peptide sequencing. US 11,315,023B2 discloses a technique for sequencing data-encoded peptides based on experimental spectra obtained by mass spectrometry or MS / MS. An improved technique for sequencing data-encoded peptides using mass spectrometry or MS / MS is desired. Summary of the Invention
[0015] One aspect of this disclosure provides a computer-implemented method for sequencing data-encoded peptides based on experimental spectra. The raw data from the experimental spectra are histogram data of intensity versus mass-to-charge ratio, which were initially obtained from the analysis of the data-encoded peptides by mass spectrometry.
[0016] The method includes: preprocessing raw data to remove uninterpretable peaks, thereby generating preprocessed data; identifying a first set of one or more candidate sequences of competing peptide sequences from a spectrum, wherein the spectrum is formed based on the preprocessed data rather than the raw data to generate a smaller number of candidate sequences, thereby reducing sequencing time costs; processing the first set of candidate sequences to estimate the peptide sequence; after obtaining a set of one or more estimated sequences of the peptide sequence, verifying whether each estimated peptide sequence is invalid; and cleaning the set of one or more estimated sequences of the peptide sequence to remove any invalid estimated peptide sequences found.
[0017] Preferably, preprocessing the raw data to generate preprocessed data includes: dividing the set of mass-to-charge ratios of the raw data into multiple subsets, such that each subset consists of the mass-to-charge ratios of isotopes of fragments encoding peptides, each mass-to-charge ratio having a signal peak in the experimental spectrum, and the isotopes of the fragments having the same positive integer charge value; calculating the monoisotopic mass of the fragments in each subset; calculating the intensity of the fragments based on the raw data and the mass-to-charge ratios in each subset; and generating preprocessed data. The preprocessed data includes multiple masses and multiple intensities, wherein the multiple masses are formed by the monoisotopic masses of the fragments in each subset, and wherein each of the multiple masses is associated with a corresponding intensity of the multiple intensities.
[0018] Preferably, the preprocessed data further includes a first mass set of presumed b-ions and a second mass set of presumed y-ions from the experimental spectrum. The first mass set and the second mass set are generated by distributing the corresponding masses from the plurality of masses into the first mass set and the second mass set.
[0019] Preferably, identifying the first candidate sequence set from the spectrum further includes the following steps: if the parent ion peak identified in the experimental spectrum has a charge value of 2, then each candidate sequence in the first candidate sequence set is identified by searching the second mass set; and if the charge value of the parent ion peak is 3, then each candidate sequence in the first candidate sequence set is identified by searching the first mass set and the second mass set.
[0020] In some embodiments, processing a first candidate sequence set to estimate the peptide sequence includes the following steps: if the first candidate sequence set consists of a single candidate sequence, then designating the single candidate sequence in the first candidate sequence set as an estimated sequence of the peptide sequence; otherwise, selecting one or more candidate sequences from the first candidate sequence set to form a second set of one or more candidate sequences, such that one or more candidate sequences in the second candidate sequence set are more likely to be the peptide sequence than candidate sequences not selected in the first candidate sequence set. If the second candidate sequence set consists of a single candidate sequence, then designating the single candidate sequence in the second candidate sequence set as an estimated sequence of the peptide sequence. If the second candidate sequence set consists of multiple candidate sequences that do not contain any AAC, then designating the multiple candidate sequences in the second candidate sequence set as estimated sequences of the peptide sequence. If the second candidate sequence set consists of multiple candidate sequences that contain one or more AACs, then using raw data to determine valid amino acid sequences for the one or more AACs. After determining the valid amino acid sequences, the candidate sequences in the second candidate sequence set are further subdivided to produce a third set of one or more candidate sequences. The third candidate sequence set is obtained from the second candidate sequence set by replacing one or more AACs with a determined valid amino acid sequence and then discarding any candidate sequences with an undetermined AAC. If the third candidate sequence set consists of a single candidate sequence, the single candidate sequence in the third candidate sequence set is designated as an estimated sequence of the peptide sequence; otherwise, one or more candidate sequences are selected from the third candidate sequence set to form a fourth set of one or more candidate sequences, such that one or more candidate sequences in the fourth candidate sequence set are more likely to be the peptide sequence than the unselected candidate sequences in the third candidate sequence set. If the fourth candidate sequence set consists of a single candidate sequence, the single candidate sequence in the fourth candidate sequence set is designated as an estimated sequence of the peptide sequence; otherwise, multiple candidate sequences in the fourth candidate sequence set are designated as estimated sequences of the peptide sequence.
[0021] In some embodiments, selecting one or more candidate sequences from a first candidate sequence set to form a second candidate sequence set includes the following steps: Selecting one or more candidate sequences with the longest consecutive amino acid length from the candidate sequences in the first candidate sequence set to form a fifth candidate sequence set. If the fifth candidate sequence set consists of a single candidate sequence, then the fifth candidate sequence set is designated as the second candidate sequence set; otherwise, selecting one or more candidate sequences with the largest number of amino acids from among the candidate sequences in the fifth candidate sequence set to form a sixth candidate sequence set. If the sixth candidate sequence set consists of a single candidate sequence, then the sixth candidate sequence set is designated as the second candidate sequence set; otherwise, selecting one or more candidate sequences with the smallest matching error from among the candidate sequences in the sixth candidate sequence set to form a seventh candidate sequence set. If the seventh candidate sequence set consists of a single candidate sequence, then the seventh candidate sequence set is designated as the second candidate sequence set; otherwise, selecting one or more candidate sequences with the highest average intensity value of the acquired amino acids from among the candidate sequences in the seventh candidate sequence set to form an eighth candidate sequence set. If the eighth candidate sequence set consists of a single candidate sequence, then the eighth candidate sequence set is designated as the second candidate sequence set; otherwise, one or more candidate sequences with different offsets and the most frequent occurrences of different ion types are selected from the multiple candidate sequences in the eighth candidate sequence set to form the second candidate sequence set.
[0022] In some embodiments, selecting one or more candidate sequences from the third candidate sequence set to form the fourth candidate sequence set includes: selecting one or more candidate sequences from the third candidate sequence set that have the maximum number of amino acids in the determined valid amino acid sequences to form the ninth candidate sequence set; and if the ninth candidate sequence set consists of a single candidate sequence, designating the ninth candidate sequence set as the fourth candidate sequence set, otherwise selecting one or more candidate sequences from the ninth candidate sequence set that have the minimum matching error in the determined valid amino acid sequences to form the fourth candidate sequence set.
[0023] In some embodiments, identifying a first set of candidate sequences from a spectrum includes: identifying one or more valid paths in the spectrum that have the longest length, wherein a valid path begins at a first vertex and ends at a tail vertex having the quality of the data-encoded peptide; using each of the identified one or more paths to generate a new candidate sequence; and assigning the new candidate sequence to the first set of candidate sequences.
[0024] In some embodiments, identifying a first set of candidate sequences from a spectrum includes the following steps: (a) sorting the plurality of masses in descending order of their respective intensities to form an ordered mass sequence, wherein the ordered mass sequence ranks the plurality of masses, with the highest-ranked mass associated with the highest corresponding intensity; (b) given selected masses in the ordered mass sequence and a selected number of higher-ranked masses, identifying one or more highest-intensity-based tags in the spectrum, wherein each highest-intensity-based tag is a partial sequence consisting of a first amino acid having a selected mass and a plurality of remaining amino acids having masses selected from the selected number of higher-ranked masses; (c) processing the one or more highest-intensity-based tags, wherein processing each highest-intensity-based tag... The labels include: determining the prefix and suffix of each label based on the highest strength; if the prefix and suffix are successfully determined, combining the prefix, the labels based on the highest strength, and the suffix to form a new candidate sequence; and assigning this new candidate sequence to a first candidate sequence set; (d) repeating steps (b) and (c) sequentially using consecutive values in a strictly decreasing value sequence as a selected number of higher-ranking qualities until the first candidate sequence set is not empty or the strictly decreasing sequence is exhausted; and (e) starting from the highest-ranking quality, repeating steps (b)-(d) sequentially using consecutive qualities in the ordered quality sequence as selected qualities until the first candidate sequence set is not empty or a pre-selected number of higher-ranking qualities is reached in the ordered quality sequence.
[0025] In some embodiments, determining the prefixes and suffixes of each highest-intensity-based tag includes the following steps: identifying the prefix as a valid first path in the spectrum having the longest length, wherein: the first path connects the first amino acid and the head of each highest-intensity-based tag; allowing one or more first AACs with a total length of at most a first maximum length to appear in the identified first path; and setting the first maximum length to L if the charge value of the precursor ion peak is 2. p1 And if the charge value of the mother ion peak is 3, then when searching for the first path, the first maximum length is initially set to L. p1 If the first maximum length is L p1 If no first path is found, the first maximum length is increased to L. p2 L p2 >L p1 The suffix is identified as a valid second path in the spectrum with the longest length, wherein: the second path connects the tail and tail amino acid of each tag based on the highest intensity; one or more second AACs with a total length of at most the second maximum length are allowed to appear in the identified second path; if the charge value of the parent ion peak is 2, then the second maximum length is set to L.s1 And if the charge value of the parent ion peak is 3, then when searching for the second path, the second maximum length is initially set to L. s1 If the second maximum length is L s1 If no second path is found, the second maximum length is increased to L. s2 L s2 >L s1 .
[0026] In some embodiments, invalidity of each peptide sequence estimate can be determined by: verifying the length of each peptide sequence estimate against a predetermined correct length of the peptide sequence; performing a sequence check on the amino acids “G” and “L” in each peptide sequence estimate; performing a sequence check based on a sequence check bit carried in the data-encoded peptide; or any combination thereof.
[0027] Other aspects of this disclosure are also disclosed as shown in the following embodiments. Attached Figure Description
[0028] Figure 1 A 3-bit symbol block of 4095×17 with a total of 208845 (=4095×17×3) bits is shown as the first instance of data designed to be carried on a data-encoding peptide;
[0029] Figure 2 A 4095×19 3-bit symbol block with a total of 98268 information bits, 12285 sequence check bits and 3 RS codes is shown as a second example of data designed to be carried on a data-encoding peptide.
[0030] Figure 3 A flowchart of a two-stage sequencing method according to certain embodiments of the present disclosure is shown;
[0031] Figure 4 A spectrum of path finding in a graph theory model according to certain embodiments of the present disclosure is shown;
[0032] Figure 5 A flowchart of a sequencing method based on the highest strength tag according to certain embodiments of the present disclosure is shown;
[0033] Figure 6 A conceptual diagram is shown of candidate sequences determined by a sequencing method based on the highest strength tag, according to certain embodiments of the present disclosure;
[0034] Figure 7 A flowchart is shown of a first method for finding prefixes and suffixes based on the highest strength labels to form possible candidate sequences according to certain embodiments of the present disclosure, wherein the prefixes and suffixes are identified based on the maximum length of the AAC;
[0035] Figure 8 A flowchart is shown of a second method for finding prefixes and suffixes based on the highest strength labels to form possible candidate sequences according to certain embodiments of the present disclosure, wherein the prefixes and suffixes are identified based on the maximum length of the AAC;
[0036] Figure 9 A flowchart of a method for verifying candidate sequences according to an embodiment of the present disclosure is shown;
[0037] Figure 10 The TIC of an exemplary mixture of peptides encoded from experimentally obtained data is shown;
[0038] Figure 11 A flowchart is shown of a computer-implemented method for sequencing data-encoded peptides based on an experimental spectrum, according to an exemplary embodiment of the present disclosure.
[0039] Figure 12 A flowchart illustrating exemplary steps for processing raw experimental spectrum data to produce preprocessed data is shown;
[0040] Those skilled in the art will understand that the elements in the accompanying drawings are shown for simplicity and clarity only and are not necessarily drawn to scale. For example, the dimensions of some elements in schematic diagrams, block diagrams, or flowcharts may be exaggerated relative to other elements to aid in understanding the embodiments of this disclosure. Detailed Implementation
[0041] This disclosure relates to peptide sequencing of data-encoded peptides. Two related peptide sequencing methods are disclosed: a two-stage sequencing method and a sequencing method based on the strongest tag.
[0042] Part A describes the technical terms and symbols used in this disclosure. Part B provides examples of designing peptide sequences for data-encoded peptides. The resulting peptide sequence design is occasionally mentioned while explaining these two sequencing methods. Part C establishes a model of the problem of sequencing data-encoded peptides. Parts D and E respectively illustrate two-stage sequencing methods and sequencing methods based on the highest strength tags. Part F provides validation methods for peptide sequence-based designs used to validate candidate sequences during the generation of these sequences. The embodiments of this disclosure are primarily based on the developments in Part DF.
[0043] A. Technical terms and symbols
[0044] The following terms and symbols are used in the specification and appended claims.
[0045] A peptide is a physical chain of amino acids linked together by peptide bonds.
[0046] A “data-encoded peptide” is a peptide in which the amino acids that make up the peptide are intentionally selected and positioned in a meaningful way to represent digital data.
[0047] A peptide's "peptide sequence" is a string of numbers representing an ordered sequence of amino acid residues within the peptide. Therefore, this ordered sequence specifies the assembly order of the constituent amino acids when forming the peptide.
[0048] The "first amino acid" is defined as the amino acid residue immediately adjacent to the N-terminus of the peptide. In other words, the first amino acid is the N-terminal amino acid residue of the peptide.
[0049] "Tail amino acid" is defined as the amino acid residue immediately adjacent to the C-terminus of a peptide. In other words, tail amino acid is the C-terminal amino acid residue of a peptide.
[0050] Those skilled in the art will generally understand the meaning of "b ion," "y ion," and other fragment ions resulting from the breaking of peptide bonds in a peptide. This disclosure is intended to use the generally accepted interpretations of b ions, y ions, and other fragment ions obtained by breaking peptide bonds in a peptide. A b ion of a peptide is a charged fragment resulting from the splitting of a peptide at the peptide bond between two constituent amino acids, wherein the b ion contains the N-terminus of the peptide. A y ion of a peptide is a charged fragment resulting from the splitting of a peptide at the peptide bond between two constituent amino acids, wherein the y ion contains the C-terminus of the peptide.
[0051] A "tag" is a partial sequence of an amino acid sequence, which is formed by consecutive amino acids.
[0052] "AAC" stands for amino acid sequence-independent composition. It should be noted that the assembly order of the constituent amino acids in an AAC is not required when modeling or initially constructing it. Nevertheless, in peptide sequencing, it is often necessary to resolve the assembly order of an AAC based on additional information from the AAC.
[0053] "Mass spectrometry" is a histogram of intensity versus mass-charge ratio in a chemical sample analyzed by a mass spectrometer.
[0054] "Experimental spectrum" refers to the mass spectrum obtained in an experiment.
[0055] The "raw data of the experimental spectrum" is histogram data of intensity versus mass-to-charge ratio, which was originally obtained from the analysis of the chemical sample by mass spectrometry. In practice, the initially obtained histogram is typically derived from a mass spectrometer. However, in this disclosure, signal enhancement techniques (e.g., noise reduction algorithms) can be used when preparing the initially obtained histogram based on physical measurement data obtained from the mass spectrometer.
[0056] "Matching error" is defined as the average error between the observed quality value of an amino acid obtained from the experimental spectrum and the actual quality value of the amino acid, normalized to the corresponding observed quality value.
[0057] The "mother ion peak" in mass spectrometry, also known as the molecular ion peak, is the peak produced by charged molecules when there is no fragmentation during molecule analysis in a mass spectrometer.
[0058] A = {a1, a2, ..., a} K} represents a set of amino acids that can be selected as constituent amino acids when forming a data-encoded peptide, where K is the number of different amino acids in the set, and a i Let i ∈ {1,…,K} be the i-th amino acid in A. In other words, A is an alphabet of amino acids that form the data-encoding peptide.
[0059] g = {g1, g2, ..., g} K} represents the set of mass of amino acid residues that can be selected when forming a data-encoded peptide, where g i ,i∈{1,…,K} is related to a i The mass of the corresponding amino acid residues.
[0060] P = {P1, P2, ..., P} N} represents the peptide sequence of a peptide, where N is the number of amino acids that make up the peptide, and P is the number of amino acids that make up the peptide. i ,i∈{1,…,N} is the i-th amino acid residue in the peptide sequence. Without causing confusion, peptides are also denoted by P, and it should be understood that a peptide called P is a peptide whose peptide sequence is P.
[0061] m = {m1, m2, ..., m} N} represents the set of masses of amino acid residues in peptide P, where m i ,i∈{1,…,N} is P i The quality.
[0062] m H It indicates the mass of a hydrogen atom or a proton.
[0063] m OH This indicates the mass of the hydroxyl group.
[0064] m Ngroup This represents the mass of the N-terminal functional group attached to the first amino acid of peptide P. If P has an unprotected N-terminus, the mass of the N-terminal functional group is equal to m. H .
[0065] m Cgroup This represents the mass of the C-terminal functional group attached to the tail amino acid of peptide P. If P has an unprotected C-terminus, the mass of the C-terminal functional group is equal to m. OH .
[0066] m head and m tailLet m represent the mass of the first and last amino acids, respectively. If there are no fixed amino acids, then the mass m is... head and m tail It can be zero.
[0067] M represents the mass of peptide P. Note that M is composed of... Provided.
[0068] m b ={m b,1 ,m b,2 ,…,m b,N+1 ,m b,N+2} represents the set of masses of b ions in peptide P, where for i∈{1,…,N+1}, m b,i+1 =m b,i +m i And m b,1 =m Ngroup -m H Note: m b,2 =(m Ngroup -m H )+m head ;m b,N =Mm H -m Cgroup -m tail ;m b,N+1 =Mm H -m Cgroup ;and m b,N+2 =M.
[0069] m y ={m y,1 ,m y,2 ,…,m y,N+1 ,m y,N+2} represents the set of masses of y ions in peptide P, where for i∈{1,…,N+1}, m y,i+1 =m y,i -m i And m y,1 =M-(m Ngroup -m H Note that m y =Mm b Also note: m y,2 =M-(m Ngroup -m H )-m head ;m y,N =m H +m Cgroup +m taik ;m y,N+1 =m H +m Cgroup ;and m y,N+2 =0.
[0070] Δ={Δ1,Δ2….,Δ N} represents the set of mass differences between theoretical (mass) spectra and experimental (mass) spectra, where Δ i ∈[-δ,+δ],i∈{1,…,N}, and δ is the tolerance value.
[0071] T(P) represents the theoretical spectrum produced by peptide P.
[0072] S represents the MS / MS spectrum obtained by tandem mass spectrometry.
[0073] L represents the intensity of spectrum S and the number of m / z comparisons.
[0074] z = {z1, z1, ..., z} L} represents the charge set of spectrum S. Typically, z i Selected from 1, 2, and 3.
[0075] (m / z)={(m / z)1,(m / z)2,…,(m / z) L} represents the set of mass / charge ratios of spectrum S.
[0076] ρ represents the number of subsets in the spectrum S. In each subset, all (m / z) ratios are isotopes of a specific segment, where the charge value is equal to the reciprocal of the difference between consecutive mass / charge ratios.
[0077] M out This represents the mass output of the mass spectrometer corresponding to its highest intensity. Mass output M out It could be a single isotope peak, or the next peak (mass +1.007276 Da).
[0078] G i ,i∈{1,2,…,ρ} represents the i-th subset in the spectrum S.
[0079] m′ i Representing subset G i The monoisotopic mass. It should be noted that m′ i It can be done through m′ i =(m / z) i,0 z′ i,0 -m H z′ i,0 Calculate, where (m / z) i,0 It is a subset G i The lowest value in, z′ i,0 It is (m / z) i,0 The corresponding charge.
[0080] m′ b ={m′ b,1 ,m′b,2 ,…,m′ b,L} represents the presumed mass set of b ions in spectrum S, where m′ b,i =(m / z) i z i -m H z i , i = 1, ..., L.
[0081] m′ y ={m′ y,1 ,m′ y,2 ,…,m′ y,L} represents the set of estimated y-ion equivalent b-ion masses for spectrum S, where m′ y,i =M-[(m / z)] i z i -m H z i ], i = 1, ..., L.
[0082] I = {I1, I2, ..., I} L Let} denote the set of intensities of spectrum S, where for each i, i = 1, 2, ..., L, the intensity I is... i With (m / z) i correspond.
[0083] J represents the ranking of strength.
[0084] m′ B,J The mass of the presumed b-ion with the highest intensity of the J-th spectrum S is represented.
[0085] m′ Y,J The equivalent b-ion mass of the presumed y-ion with the J-th highest intensity in spectrum S is represented.
[0086] J max This indicates the maximum number of higher-ranking qualities that can be used as a starting point when searching for tags.
[0087] n represents the number of candidate sequences.
[0088] W represents the quantity of quality with higher ranking strength used in the label-finding scheme.
[0089] V represents the maximum number of iterations in the sequencing method based on the highest strength tag.
[0090] L AAC This indicates the length of AAC.
[0091] L p1 and L p2 These represent the two maximum lengths of an AAC, which are used to limit the length of each AAC in the prefix.
[0092] L s1 and L s2 These represent the two maximum lengths of an AAC, which are used to limit the length of each AAC in the suffix.
[0093] P L This indicates the position of the last amino acid "L" in the peptide sequence P.
[0094] P G This indicates the position of the first amino acid "G" in the peptide sequence P.
[0095] L′ represents the number of symbols in a row of 4095×17 blocks or 4095×19 blocks. Note that L′ = N-2.
[0096] Seq represents the index of the sequence in the block, where Seq ranges from 1 to 4095.
[0097] S1,S2,…,S L′ The L′ symbol (excluding the first and last amino acids) is derived from the block and encoded in peptide P.
[0098] Q i,j Indicates the symbol S used for protection i and S j The order check bits are selected from (1,2), (2,3), and (L′-1,L′). For details on the order check bits, please refer to US 11,315,023B2.
[0099] A1, A2, A3, and A4 represent address symbols S6, S7, S8, and S9, respectively.
[0100] n i This indicates the code length of the i-th RS code.
[0101] k i This represents the number of 12-bit information symbols in the i-th RS code.
[0102] R represents the overall bitrate.
[0103] Single-letter symbols are occasionally used to represent different amino acids, such as when describing the structure of peptides. The amino acids and their respective single-letter symbols (enclosed in parentheses) are as follows: alanine (A); arginine (R); asparagine (N); aspartic acid (D); cysteine (C); glutamine (Q); glutamic acid (E); glycine (G); histidine (H); isoleucine (I); leucine (L); lysine (K); methionine (M); phenylalanine (F); proline (P); serine (S); threonine (T); tryptophan (W); tyrosine (Y); valine (V); selenocysteine (U); and pyrrolidone (O).
[0104] B. Design of data-encoded peptides
[0105] In the design of fixed-length sequences, each peptide sequence has L′ (L′=17) 3 symbols, which are represented as S1, S2, …, S 17 We assume that: (i) for a symbol sequence of length 17 {S1, S2, ..., S...} 17}, 4 symbol sequences {S1,S2,S3,S4}, {S 10 ,S 11 ,S 12 ,S 13} and {S 14 ,S 15 ,S 16 ,S 17} 15%, 10%, and 25% respectively could not be correctly recovered; and (ii) there were three ambiguous symbol orders: the order of S1 and S2, the order of S2 and S3, and S... 16 and S 17 The order.
[0106] Then, we propose an error correction method based on RS codes. This method uses: (i) even in any 15% of the 4-symbol sequences {S1S2S3S4}, any 10% of the 4-symbol sequences {S... 10 S 11 S 12 S 13} and any 25% of the 4-symbol sequences {S 14 S 15 S 16 S 17 The three RS codes that can recover the original data even when it cannot be correctly recovered, and (ii) the sequences of S1 and S2, S2 and S3, and S in each peptide sequence, respectively. 16 and S 17 The three sequential check bits are in the correct order.
[0107] like Figure 1As shown, a 3-bit symbol block of 4095×17 is first constructed, which has a total of 208845 (=4095×17×3) bits. Each row of this block contains 17 symbols (i.e., S1, S2, …, S…). 17 The symbol ) represents the 17-mer data encoding peptide. The four symbols (S6, S7, S8, and S9) are defined as A1, A2, A3, and A4, respectively, serving as addresses for the peptide sequence, and are assumed to be one of the following values: 0000 (octal form), 0001, 0002, ..., 0007, 0010, ..., 0777, 1000, ..., 7775, 7776. The three check bits carried in S5 are represented as Q. 1, Q 2, and Q 16, They are used to represent the order of S1 and S2, the order of S2 and S3, and S... 16 and S 17 The order. If S i The value is greater than S j The value of Q is then checked by the sequence check bit. i,j If the check bit is "1", then the sequence check bit is "0"; otherwise, the sequence check bit is "0". Then the remaining 12 symbols of each row (i.e., S1, S2, ..., S5, S...) can be used. 10 ,…,S 17 To recover the coded bits c, including the error correction code information bits and parity bits.
[0108] Next, we assume a 4-symbol sequence {S1S2S3S4}, {S... 10 S 11 S 12 S 13} and {m 14 S 15 S 16 S 17} 15%, 10%, and 25% of the sequences, respectively, could not be correctly recovered. Then, we used three different (n) methods on these three partial sequences. i ,k i RS code, where n i It is the code length of the i-th RS code, k i This is the number of 12-bit information symbols in the i-th RS code. Assume that for i = 1, 2, and 3, n... iTaking the same value of 4095, i.e., n1 = n2 = n3 = 4095, and k1 = 2867, k2 = 3275, and k3 = 2047, respectively, are used for the first, second, and third RS codes. The resulting (4095, 2867) RS code (represented as RS1), (4095, 3275) RS code (represented as RS2), and (4095, 2047) RS code (represented as RS3) can correct up to 614, 410, and 1024 12-bit symbol errors, respectively. For the entire block, there are at most (k1 + k2 + k3) × 12 = 98268 bits that can be used as data bits, and the maximum overall code rate R of the block is given by 98268 / (4095 × 17 × 3) = 0.4705.
[0109] For our dataset A, the number of information bits is 96224, of which 47965 are "0" bits and 48259 are "1" bits. Since each 4095×17 block can contain a maximum of 98268 information bits, we fill the remaining 98268-96224=2044 positions in each block with "0" and "1" bits with equal probability. Then, these 98268 bits are divided into three parts. The first part is used to fill the symbols {S1S2S3S4} in the row sequence ranging from (n1-k1+1) to n1. The second part is used to fill the symbols {S...} in the row sequence ranging from (n2-k2+1) to n2. 10 S 11 S 12 S 13 The third part is used to fill the symbol {S} in the row sequence ranging from (n3-k3+1) to n3. 14 S 15 S 16 S 17 During the encoding process, RS1 uses the bits in the first part to generate encoded symbols to fill the row sequence {S1S2S3S4} in the range from 1 to (n1-k1); RS2 uses the bits in the second part to generate encoded symbols to fill the row sequence {S1S2S3S4} in the range from 1 to (n2-k2). 10 S 11 S 12 S 13 RS3 uses the bits in the third part to generate coded symbols to fill the {S} in the row sequence ranging from 1 to (n³-k³). 14 S 15 S 16 S 17}
[0110] Finally, we count the number of 3-bit symbols in the block that are equivalent to 000, 001, ..., 111 (or represented by 0 to 7). The number of symbols 000, 001, ..., 111 are 8294, 8380, 8927, 9181, 8831, 8754, 8671, and 8577, respectively. The actual bit rate R of our dataset A is given by 96224 / (4095×17×3) = 0.4607.
[0111] Now, we consider the case where the sequence length varies, and assume it is one of the following values: 15, 16, 17, 18, and 19. Then, we construct... Figure 2 The block shown is 4095×19. We consider every 5 sequences of length 15, 16, 17, 18, and 19 as a group. The length L′ of the j-th sequence in each group is... j Through L′ j The estimation is done using 15 + modSeq - 1,5, where mod represents modulo operation, and Seq is the index of the sequence in the block, ranging from 1 to 4095. Since Seq and the address sequence {A1A2A3A4} are calculated using Seq = A1 × 8... 3 +A2×8 2 +A3×8 1 The +A4+1 is a one-to-one mapping, so the correctness of the sequence can be determined using address-based length checks.
[0112] set up It is the i-th symbol of the j-th sequence in each group. Figure 1 It can be seen that a 4095×17 block can be converted into a 4095×19 block in the following way: (i) moving sequence 1 and And insert it into sequence 5. Previously; and moved sequence 2 And insert it into sequence 4. Before.
[0113] because Figure 2 The variable-length sequence shown is achieved by shifting... Figure 1 It is constructed from symbols in the fixed-length sequence shown, therefore Figure 2 The overall bitrate R of the 4095×19 blocks in the middle is... Figure 1 The overall bitrate of the 4095×17 blocks is the same.
[0114] Next, the 3-bit symbols in the 4095×17 block (corresponding to the symbols from “0” to “7” for the 3 bits from 000 to 111) are mapped to 8 amino acids (Y, T, E, A, S, V, G, and F) to obtain 4095 sequences, where “G” corresponds to the symbol “6”.
[0115] Furthermore, if a quality conflict occurs (i.e., the quality M of one sequence is within the tolerance range (25 ppm) of other sequences), then the amino acid "G" is replaced with an unused amino acid "L". This process begins with the first "G" closest to the first amino acid. After replacing only one "G", if the updated quality M still falls within the tolerance range of other sequences, the replacement process continues by moving to the next "G" in the original sequence. After this process is complete, the sequence consists of 9 possible amino acids, where both "G" and "L" correspond to the symbol "6". Since all "L" amino acids should appear before "G", the order of "G" and "L" can be used to distinguish the validation of the estimated peptide sequence after peptide sequencing. The effect of this replacement can be seen in an exemplary dataset of 4095 data-encoded peptides, which have fixed N-terminal amino acid H and fixed C-terminal amino acid R, and the amino acids between them are encoded using the 3-symbol to 8-amino acid mapping as described above. Before the replacement, the number of data-encoded peptides with quality conflicts was 3630, and there were 732 groups of data-encoded peptides, where all data-encoded peptides in each group had the same M within the tolerance range. After the replacement, the number of data-encoded peptides with quality conflicts decreased to 2746, and there were 619 groups of data-encoded peptides, where all data-encoded peptides in each group had the same M within the tolerance range. This reduced the problem of co-elute of two or more isomeric ectopic peptides, thereby reducing the number of overlapping MS / MS spectra that interfered with sequencing.
[0116] C. Peptide sequencing issues
[0117] Typically, most existing sequencing algorithms (including Sherenga, PepNovo, NSNovo, pNovo, uniNovo, NovoHCD, and Novor) rely on training based on the characteristics of the data, given the use of the entire set of 20 natural amino acids and existing peptides containing those natural amino acids. For example, by The OFF (De Novo Peptide Sequencing via Tandem Mass Spectrometry) introduced by [Authors' Name] was initially derived empirically by counting the frequency of different ion types and feature types in the training data. After obtaining candidate paths using DAG models and dynamic programming for Sherenga, pNovo, and uniNovo, OFF was used as a reference for scoring these paths.
[0118] In peptide-based data storage, amino acids are generated by arbitrary combinations of bits "0" and "1", thus the peptide sequence is completely random. Therefore, no suitable training sequences can be used with existing sequencing algorithms. Furthermore, peptide-based data storage can use some or all of the natural amino acids, depending on the peptide design. Alternatively, non-natural amino acids can also be used.
[0119] One limitation of dynamic algorithms is the lack of training spectra for sequencing peptides carrying digital data. Therefore, de novo sequencing based on graph models is used to determine peptide sequences. In a graph model, the MS / MS spectrum is represented by a DAG, called a spectrogram. Peaks in the spectrum can serve as vertices, and an edge is added between two vertices when the mass difference between two peaks equals the mass of an amino acid. The goal of dynamic programming is to find the longest (or optimal) path from the first vertex to the last vertex in the graph. Alternative methods for sequence identification include starting from the middle of the MS / MS spectrum. For example, sequence tagging methods first infer a partial sequence called a tag, then search for the entire sequence that matches that tag. In tag-based methods, tags are first searched for in the MS / MS spectrum according to some scoring scheme. Then, sequence inference depends on peptide comparison using database search methods, or on using de novo sequencing to extend the effective path of the tag in the middle of the path.
[0120] Table 1 summarizes the peptide sequencing problem, including information on the amino acids used, the spectrum S, the overall sequence quality, and the first and last amino acids. The sequencing problem is to find the peptide P whose theoretical spectrum T(P) best matches the experimental spectrum S. In many practical cases, the length N of the peptide sequence is fixed and known. Therefore, candidate peptides with a length not equal to N are discarded. This paper discloses two sequencing methods: a two-stage sequencing method and a sequencing method based on the highest strength tag. For both methods, the sequence is estimated by first inferring a partial sequence with a small amount of reliable information, and then searching for missing parts of the sequence with less reliable data or raw data. The advantage of this method is that it increases the sequencing speed by generating a smaller number of candidate sequences. Furthermore, the sequencing method based on the highest strength tag is more effective in rejecting unlikely candidate sequences due to the introduction of tags.
[0121] Table 1. Explanation of peptide sequence issues
[0122]
[0123]
[0124] D. Peptide sequencing: A two-stage sequencing method
[0125] Figure 3 A flowchart illustrating exemplary steps of a two-stage sequencing method 100 is shown. The two-stage sequencing method 100 includes four steps: preprocessing, candidate sequence generation, sequence selection, and candidate segmentation. Figure 3 As shown, steps 110, 120, and 130 belong to the first stage (i.e., stage 1), while step 140 is processed in the second stage (i.e., stage 2). In stage 1 of the two-stage sequencing method 100, after performing step 110, preprocessed data is used to infer a partial sequence. In stage 2, the remaining portion of the sequence is determined based on the raw data.
[0126] In step 110, preprocessing is performed. Preprocessing serves two purposes. The first is to remove some unexplained peaks caused by noise and uncertainty. The second is to convert the mass-to-charge ratio set (m / z) into the corresponding mass set m′. b and m′ y Given a set of mass-to-charge ratios (m / z), at which signal peaks appear in the experimental spectrum S, these ratios are divided into ρ subsets G1, G2, ..., G... ρ In each subset G i In the subset G, i ∈ {1, 2, ..., ρ}, all (m / z) ratios are isotopes of specific segments with the same chemical composition. Furthermore, the isotopes in each subset have the same charge state, represented by positive integer values of charge (typically 1, 2, or 3). For each subset G... i , through m′i =(m / z) i,0 z′ i,0 -m H z′ i,0 Calculate the monoisotopic mass m′ for i = 1, 2, ..., ρ. i , where (m / z) i,0 It is the lowest value in this subset, z′ i,0 It is (m / z) i,0 The corresponding charges. Then, these m′ i Value in the mass set m′ b and m′ y The distribution is between [a certain value]. In some embodiments, one of the distribution criteria is based on [a certain value] and m′. i The intensity corresponding to the value. In some embodiments, one of the criteria for the distribution is G. i The isotopic pattern of the (m / z) ratio. In some embodiments, one of the criteria for distribution is based on the fact that if m′ i At m′ b,i In the middle, then in m′ y,j There exists a corresponding m′ j , where m′ j =Mm′ i In some embodiments, the distribution is determined in real time in step 120, which will be described later. In some embodiments, only those data with typical charge characteristics are retained as preprocessed data, so the preprocessed data may be more reliable than the original data. However, in cases where the data is incomplete or ambiguous (e.g., data lacking charge characteristics), some useful quality values may be discarded during preprocessing based on the above criteria. Therefore, in cases where there are missing / uncertain elements in the sequence in stage 1, the original data can be considered in stage 2.
[0127] In step 120, the preprocessed data from step 110 is used to find valid paths (sequences), and the number n of candidate sequences is counted. Please refer to [link / reference]. Figure 4 The spectrum 200 is shown as an example illustrating this disclosure. Assume m... Ngroup =m H And m Cgroup =m OH A graph theory model is used to find candidate paths (valid paths), each path in the DAG with a first mass m. head Start, and with quality Mm H -m Cgroup -m tail End. Since the mass obtained through the mass-to-charge ratio set (m / z) may be generated by b ions, y ions, or both b ions and y ions, a single mass set (m′) can be considered in the pathfinding algorithm.b,i or m′ y,i ) or two mass sets (m′) b,i and m′ y,i In a graphical model, the mass of a fragment ion can be represented by a vertex. If the mass difference between two fragment ions is equal to the mass of any amino acid, then an edge is added between those two vertices. For example... Figure 4 As shown, the tree can be expanded edge by edge. If the set of vertices for the correct path is complete, the sequencing problem can be simplified to finding the longest path in the graph. This path should include both the first and last vertices. Furthermore, only paths ending with quality M are considered candidate paths. Figure 4 In the example, only path 1 and path 2 are candidate paths.
[0128] The selection of an appropriate mass set for peptide sequencing is strongly based on peptide design. In one embodiment, F is the first amino acid (the fixed amino acid at the N-terminus), R or K is the tail amino acid (the fixed amino acid at the C-terminus), and there is no P or other essential amino acid in the middle. Previous experimental data have shown that the parent ion peak of charge 2 is the strongest, while the peak of charge 3 is very weak. When selecting the peak of charge 2 for MS / MS, γ ions are much stronger than b ions. In this case, only the MS / MS spectrum of the peak of charge 2 will be analyzed. Furthermore, the mass of γ ions is sufficient to find all sequences, so only γ ions are used. In another embodiment, H is the head, R is the tail, and there is no P or other essential amino acid in the middle. In this case, the peak of charge 3 may be much stronger than the peak of charge 2, or as strong as the peak of charge 2. In this case, the MS / MS spectrum of one or both parent peaks will be analyzed. With charge 2, the mass of γ ions is sufficient to find all sequences, so only γ ions are used. However, with charge 3, both γ and b ions are used. In another embodiment, the head and tail are not basic amino acids, the charge is typically 2, and γ and β ions are used.
[0129] Due to incomplete fragmentation in MS / MS characterization, two or three missing ions are frequently observed in a sequence. Under the publicly available model of the peptide under consideration, to ensure that the pathway starting from the head extends to the tail, it is assumed that the maximum number of missing amino acids in stage 1 is L. AAC From m b,1 Starting with a mass of 0, first try to find the mass of the first amino acid, m. b,2 =m head +Δ1. Next, the mass m is found using preprocessed data. b,i+1 This makes the mass difference between the current vertex and the next vertex approximately equal to the mass s. i or at most L AAC A person with quality s v The mass of amino acids ∈g (for v = 1, 2, ..., N) (where l≤L) AAC That is, for two consecutive vertices i and i+1, m b,i+1 =m b,i +(s i +Δ i ), or, m b,i+l =m b,i + (where Δ) i,i+l ∈[-lδ,+lδ] is a label of length l used for the distance from vertex i to vertex (i+l). Please refer to [reference]. Figure 4 Assume path 2 is the correct path, but with experimental quality, where for two consecutive vertices, m b,1 =0,m b,i+1 =m b,i +(m i +Δ i Or, for a tag of length l, If path 1 is the correct path with theoretical quality, then m b,i+1 =m b,i +m i , where m b,1 =m Ngroup -m H m b,2 =(m Ngroup -m H )+m head , ..., m b,N =Mm H -m Cgroup -m tail m b,N+1 =Mm H -m Cgroup And m b,N+2 =M. Please refer to... Figure 4 If the mass difference between two vertices is equal to the mass of one amino acid, a solid edge is added. On the other hand, if the mass difference is equal to the sum of the masses of two or more amino acids, a dashed edge is added, and the missing vertex is represented by a hollow circle.
[0130] In step 130, when obtaining scores for candidate sequences from steps 131 to 135, the influence of the following five factors is considered in conjunction: the length of the acquired consecutive amino acids (step 131), the number of acquired amino acids (step 132), the matching error (step 133), the average intensity of the acquired amino acids (step 134), and the number of occurrences of different ion types with different offsets (step 135). First, the sequence with the longest acquired consecutive amino acid length is selected (step 131). Then, the sequence with the largest number of acquired amino acids is selected from the selected sequences (step 132). For sequences with equal acquired consecutive amino acid lengths and equal acquired amino acid numbers, the matching error is evaluated. As described above, this matching error is the average error normalized to the observed quality value of the amino acid obtained from the experimental spectrum and the actual quality value of the amino acid according to the corresponding observed quality value (step 133). If more than one sequence has the same matching error, the average intensity of the acquired amino acids is further calculated, and sequences with larger average intensity values are assigned higher scores (step 134). Furthermore, multiple ion types are generally considered important factors when inferring amino acids. This means that the mass value may correspond to different types of ions in the spectrum. Generally, the more times different ion types of an amino acid appear, the more likely that amino acid is correct. Therefore, for sequences with the same score after performing the evaluations in steps 131-134 above, the occurrence frequency of different ion types is counted to determine the sequence (step 135). The mass offset sets of the N-terminal α-ion, β-ion, and β-ion type sets (i.e., {a, α-H2O, α-NH3, α-NH3-H2O}, {b, β-H2O, β-H2O-H2O, β-NH3, β-NH3-H2O}, and {c, c-H2O, c-H2O-H2O, c-NH3, c-NH3-H2O}) are {-27, -45, -44, -62}, {+1, -17, -35, -16, -34}, and {+18, 0, -18, +1, -17}, respectively. By shifting the masses of the c-ion and b-ion type sets by +27 and +18 respectively, the mass offset sets of the c-terminal x-ion and y-ion type sets can be calculated. Depending on the characteristics of the fragmentation method and the data, all or some of the above ion types can be used flexibly.
[0131] The candidate sequences obtained in step 120 are found using preprocessed data, the purpose of which is to provide more reliable information for generating partial sequences. However, when the data provided by preprocessing is insufficient, AACs may exist in the candidate sequences. In step 140, if a selected sequence with a missing mass value exists, meaning the corresponding mass difference equals the sum of at least two amino acids, then the raw data can be used in stage 2 to find as many vertices as possible for the path. For the raw data, assuming all (m / z) ratios have a chance of being generated by single-charged, double-charged, or triple-charged ions, then q mass-to-charge ratios (m / z) can be used. i The set of i = 1, 2, ..., q is transformed into the presumed mass set m′ of b ions. b and the estimated equivalent b-ion mass set m′ of the y-ion y set m′ b and m′ y Each of the values has 3q elements. Although the number of quality values increases, only the range between the head and tail qualities of the AAC is considered, and this range is relatively small compared to the range of the entire sequence. For example... Figure 4 As shown, path 1 illustrates a difference equal to the sum of four amino acids. Using further information from the raw data, the following differences can be identified: (a) the composition of amino acids and AACs, (b) the composition of two AACs, (c) the composition of the tag and AACs, and (d) a single tag. Note that a spacing equal to the sum of more amino acids effectively ensures the formation of a valid path. However, this may generate more candidate sequences, thus requiring longer peptide sequencing times.
[0132] like Figure 3 As shown, after finding the missing amino acid in AAC in step 141, the sequence with the longest consecutive amino acid length obtained from AAC is selected as a candidate sequence (step 142). If at least two candidate sequences remain after selection, a final decision is made based on the amino acid matching error of each sequence obtained from AAC (step 143).
[0133] E. Peptide sequencing: a sequencing method based on the strongest tags
[0134] First, the mass-to-charge ratio (m / z) corresponding to the first or second highest intensity is identified to further infer the tag or path. In highest-intensity tag-based sequencing methods, short tags with three amino acids, such as GutenTag, DirecTag, and NovoHCD, are used. While shorter tags avoid introducing incorrect amino acids, the number of candidate tags is large, and sometimes the information provided by the tags is insufficient for sequence inference. Note that the tag length is not fixed; if the data is complete, its maximum length can be the length of a peptide, thus helping to reduce the search space. When a tag contains incorrect amino acids, it usually cannot be expanded with valid prefixes and suffixes. In this case, the tag length is shortened by adaptively reducing the number of higher-intensity data points used in the tag-finding algorithm. Furthermore, due to data uncertainty, a vertex with the highest intensity may not necessarily exist in the correct path. When no valid path can be found, a tag with the second highest intensity can be inferred.
[0135] Figure 5 A flowchart illustrating exemplary steps used in the highest-strength tag-based sequencing method 300 is shown. Method 300 begins with step 302 for preprocessing the raw data. Step 302 is identical to step 110 of the two-stage sequencing method 100. Method 300 then proceeds from step 302 to steps 304, 306, and 308.
[0136] In steps 304, 306, and 308, the intensities of the preprocessed data are sorted from maximum to minimum, where J represents the rank of the intensity. The mass-to-charge ratio with the highest intensity is then identified and converted into the corresponding mass of the b-ion. Initially, it is set to J = 1 and i = 1, and only the rank W = w is used in the tag-finding process. i (w1>w2>…>w V The mass-to-charge ratio of ).
[0137] In steps 302 and 304, for each sequence, the preprocessing unit also outputs (i) a selection of a mass set based on charge (most commonly 2 or 3) (i.e., a set of y ions, or a set of y and b ions), and (ii) the mass M of the entire sequence based on the isotopic pattern of the parent ion (= M). out Or M out The determination of –1.007276Da) is similar to that in the preprocessing step (i.e., step 110) described above, where M out This is the mass output of the mass spectrometer corresponding to the highest intensity, which can be a single isotope peak or the next peak (mass +1.007276 Da).
[0138] The selection of an appropriate mass set for peptide sequencing is strongly based on peptide design. In one embodiment, F is the first amino acid (the fixed amino acid at the N-terminus), R or K is the tail amino acid (the fixed amino acid at the C-terminus), and there is no P or other essential amino acid in the middle. Previous experimental data have shown that the parent ion peak of charge 2 is the strongest, while the peak of charge 3 is very weak. When selecting the peak of charge 2 for MS / MS, γ ions are much stronger than b ions. In this case, only the MS / MS spectrum of the peak of charge 2 will be analyzed. Furthermore, the mass of γ ions is sufficient to find all sequences, so only γ ions are used. In another embodiment, H is the head, R is the tail, and there is no P or other essential amino acid in the middle. In this case, the peak of charge 3 may be much stronger than the peak of charge 2, or as strong as the peak of charge 2. The MS / MS spectrum of one or both parent peaks will be analyzed. With charge 2, the mass of γ ions is sufficient to find all sequences, so only γ ions are used. However, with charge 3, both γ and b ions are used. In another embodiment, the head and tail are not basic amino acids, the charge is typically 2, and γ and β ions will be used.
[0139] Method 300 then proceeds to step 310 to find the tag based on the highest intensity. This is determined from the mass m′ of the b-ion with the highest intensity. B,J Or the mass m′ of the y ion Y,J We begin by finding the label based on the highest strength by simultaneously forward-connecting vertices pointing to the tail vertex of the path and backward-connecting vertices pointing to the head vertex of the path. Here, the mass difference between vertices is the mass g of any amino acid. k (k = 1, ..., K), and the label length should be as long as possible (see...). Figure 6 The tag containing the amino acid with the highest strength is then obtained and referred to as the highest strength-based tag. Knowing the quality of the first and last amino acids of the highest strength-based tag, method 300 proceeds to step 312, where a prefix capable of forwardly linking the head of the path to the head of the tag is found using the method described in step 120 of the two-stage sequencing method 100. For tags with a valid prefix, in step 314, the suffix portion of the sequence can be further found by linking the tail of the tag to the tail of the forward path using a similar method.
[0140] In the prefix and suffix search processes in steps 312 and 314, it is crucial to appropriately set the maximum length of the AAC, which allows for the finding of more valid sequences even when some quality is unreliable. The maximum length of the AAC is set to limit the length of each AAC in the prefix or suffix. The preprocessed data is relatively reliable and provides a quality-intensity pattern for each output quality. Therefore, the main path with one or more AACs is first found using the preprocessed data. Furthermore, this path connects the first amino acid of the prefix portion, the last amino acid of the suffix portion, and the highest intensity tag. If the maximum length of the AAC is small (e.g., 3 or 4), the search time is short, but some valid paths may not be found. If no prefix or suffix portion is found, the maximum length of the AAC can be adjusted to a sufficiently large value (e.g., 10). Figure 7 and Figure 8 The details of the prefix and suffix search methods 500 for charges 3 and 2 are shown separately. Let the two maximum lengths of the prefix AAC be L. p1 and L p2 L p2 >L p1 The two maximum lengths of the suffix AAC are L. s1 and L s2 L s2 >L s1 .like Figure 4 As shown in AAC(b), after obtaining the tag with the highest strength, L is first used. p1 To find the prefix. If the prefix is not found, use L. p2 To find a prefix. After finding a valid prefix, first use L s1 Use it to find the suffix. If the search fails, use L. s2 Based on the properties of the spectrum, two pairs (L p1 ,L s1 ) and (L p2 ,L s2 Each pair in the sequence can have the same or different lengths. For a 4095 sequence dataset, use L... p1 =L s1 =3 and L s2 =10>L p2 =6. When the charge is 2, only L is considered. p1 and L s1 ,like Figure 7 As shown.
[0141] In step 316, candidate paths can be constructed by combining the following three parts: prefix, tag, and suffix. In step 318, sequences can be selected and subdivided according to steps 130 and 140 of the two-stage sequencing method 100. Note that a larger value for W may sometimes introduce one or more incorrect amino acids at the beginning and / or end of the tag, while a smaller value for W may give a more reliable tag, but the tag length may be limited. Therefore, in steps 322 and 324, if no valid candidate sequence is found, the value of W (W = w) can be reduced by incrementing i by 1 (i.e., i ← i + 1). i The label-prefix-suffix search process is repeated until a candidate sequence is found or i = V is reached (where V is the maximum number of iterations).
[0142] In steps 332 and 334, when the experimental quality with the highest intensity gives unreliable information due to noise and uncertainty, no label based on the highest intensity or a valid path with a label based on the highest intensity can be found. In this case, a quality with the second highest intensity is used by setting J←J+1 and i=1 to find a label and candidate sequence based on the second highest intensity. This process continues until a sequence is found or J=J+1 is reached. max (where J) max It is the maximum number of high-ranking qualities that are allowed as a starting point for searching for tags.
[0143] In some embodiments, as described above, γ and B ions are considered in sequencing. To reduce the search space and improve sequencing accuracy, we consider the following three steps to select the mass set: (1) selecting mass ranges for γ and B ions; (2) eliminating mass overlap between γ and B ions; and (3) avoiding the coexistence of two masses with the same peak representing the spectrum in the path.
[0144] Step 1 Sequencing is performed using the mass values of both Y and B ions. Depending on the nature of the spectrum, the following options can be selected: (i) some Y ion mass values and some B ion mass values; (ii) all Y ion mass values and some B ion mass values; (iii) some Y ion mass values and all B ion mass values; or (iv) all Y ion mass values and B ion mass values. Using a higher Y ion and / or B ion mass value introduces more noise and interference into the sequencing process. Therefore, it is best to use the fewest possible mass values to ensure that the valid sequence includes the correct sequence. In other words, option (i) above is preferred.
[0145] In case (i), the quality set m′ used to find the path is given by m′={m′ b ,m′ y}={m′ b,1 ,m′b,2 ,…,m′ b,L ,m′ y,1 ,m′ y,2 ,…,m′ y,L The following is given. For peptides in the 4095 sequence dataset as described above, it has been found that the latter half of the complete path can be obtained by using the mass of the y ion, while the first half is mainly obtained by using the mass of the b ion. If the mass set is m′={m′ b,1 ,m′ b,2 ,…,m′ b,u ,m′ y,u+1 ,m′ y,u+2 ,…,m′ y,L}, then for each i = 1, 2, ..., L, the path does not contain mass m′. b,i and m′ y,i A simple setting method is to set u such that m′ b,u It is close to the maximum value of M / 2. At this point, the quality set m′ used to find the path contains m′ below M / 2. b The mass of b ions and m′ of M / 2 or greater than or equal to M / 2 y The mass of the y ion in the matrix, i.e., for i = 1, 2, ..., u, m′ b,i <M / 2, and for i = u+1, u+2, ..., L, m′ y,i ≥M / 2.
[0146] Step 2 For some spectra, only the set {m′} is used. b,1 ,m′ b,2 ,…,m′ b,u The mass in} (where for i = 1, 2, ..., u, m′) b,i <M / 2) cannot completely cover the first half of the pathway, therefore some information about the y ions is needed to obtain one or two correct masses located near M / 2. Here we adjust the mass set used for sequencing to include the mass m′ of the y ions. y,i ≥M / 2-e (for i=v+1,v+2,…,L) (here e=200) and the mass m′ of the b ion b,i <M / 2 (for i = 1, 2, ..., u), that is, m′ = {m′ b,1 ,m′ b,2 ,…,m ′ b,u ,m′ y,v ,m′ y,v+1 ,…,m′ y,u ,m′ y,u+1 ,…,m′ y,L}, where u > v. Within the mass range of M / 2 - e to M / 2, the mass of the b ion can overlap with the mass of the y ion, meaning that different peaks in the spectrum produce the same mass, i.e., m′. b,i =m′ y,i We assume that at m′ b,i =m′ y,i and m′ b,i+1 =m′ y,i+1 When there exists a mass pair P1={m′ in the spectrum y,i ,m′ y,i+1}、P2={m′ y,i ,m′ b,i+1}、P3={m′ b,i ,m′ y,i+1} and P4={m′ b,i ,m′ b,i+1 Then we merge these four pairs P1 through P4 into one pair. Generally, there is a deviation between experimental quality and theoretical quality. We then prioritize according to P1 > P4 > P3 > P2. After finding all four pairs, we select the quality pair P1 and remove the quality pairs P2, P3, and P4. Similarly, if only two quality pairs P2 and P4 (or P3 and P4) are found, we keep only P4.
[0147] Step 3 According to the antisymmetric requirement [Sherenga], only mass m′ b,i or Mm′ b,i (m′ y,i or Mm′ y,i A peak can appear in the path to avoid the same peak appearing twice in the spectrum. Assume the mass is m′. b,j (or m′) y,j ) is one of the qualities of the label. We remove quality Mm′ from set m′. b,j (or Mm′) v,j To further search for the prefix, we assume the mass is m′. b,k (or m′) y,k ) is one of the qualities of the prefix part. We remove the quality Mm′ from the set m′. b,k (or Mm′) y,k ), and continue searching for the suffix.
[0148] In the above descriptions of steps 304 and 306, it is assumed that all masses of the peptide and the corresponding calculations are accurate. In reality, the masses given by the spectra have tolerances, therefore all calculations are performed within the tolerance range.
[0149] F. Validation methods based on sequence design
[0150] Figure 9A schematic diagram of the candidate sequence validation method 600 is shown. For each candidate sequence obtained using the highest-strength tag sequencing method 300, we perform the following three validation steps based on the peptide design.
[0151] First, the sequence length can be 15, 16, 17, 18, or 19. As described above in the peptide sequence design section, the sequence length is directly related to its address within the entire block. Therefore, the validity of the sequence can be determined by (i) evaluating the expected length of the sequence based on its address and (ii) comparing the expected length with the actual length. This verification is performed in... Figure 9 The diagram shows verification step 610.
[0152] Secondly, if an output sequence of valid length exists and contains the amino acid "L", the order of amino acids "G" and "L" is further verified. This verification is performed in... Figure 9 This is shown as verification step 620. First, locate the position P of the last amino acid "L". L And the position P of the first amino acid "G" G Then compare P L and P G If P L >P G If the sequence is invalid, it should be discarded.
[0153] Third, if a sequence can be successfully found, then an order check is performed based on the sequence's check bits. This check is performed in... Figure 9 This is shown as verification step 630. Based on the estimated sign pairs {S1,S2}, {m2,S3}, and {S... L′-1 ,S L′ The sequence of S1 and S2 is used to generate a check bit based on predefined rules. The generated check bit is compared with the corresponding bit in the estimated symbol S5 to check if they are the same. For example, if the order of S1 and S2 is correct, the check bit generated for {S1, S2} should match the first bit of the estimated symbol S5. If the generated check bit does not match any of the three bits of S5 in the estimated sequence, the estimated sequence should be removed from that block.
[0154] Finally, if the sequence passes all three verification processes, it is output as a valid sequence.
[0155] It should be noted that the three verification steps 610, 620, and 630 can be performed in any order, without strictly following the prescribed sequence. Figure 9 The above-described order is executed as shown. Furthermore, not all three verification steps 610, 620, and 630 can be used together in the verification. Any combination of one or more of the three verification steps 610, 620, and 630 can be used.
[0156] Although validation method 600 is shown for use in the highest-strength tag-based sequencing method 300, it can also be used in the two-stage sequencing method 100. However, some validation steps 610, 620, and 630 may need to be modified for the two-stage sequencing method 100. For example, even during the search for a valid path in step 120, a requirement for a specific length may be observed. In some practical cases, validation step 610 may not be necessary.
[0157] G. Sample Data
[0158] The peptide mixture encoded by the first 50 addresses in the 4095 peptide block described above was sequenced using the method described above, where F represents the head and R represents the tail (sequence details are shown in Table 3). The TIC of this mixture was... Figure 10 As shown in the image, all peptides were correctly sequenced.
[0159] Table 3. Detailed sequence information of 50 peptides
[0160]
[0161]
[0162]
[0163]
[0164] H. Detailed Description of Embodiments of this Disclosure
[0165] The embodiments of this disclosure will now be described in detail, examples, and applications based on the two-stage sequencing method 100, the sequencing method 300 based on the highest strength tag, and the verification method 600 disclosed above.
[0166] One aspect of this disclosure provides a computer-implemented method for sequencing data-encoded peptides based on experimental spectra. The method provides one of three outputs. In the first output, corresponding to the optimal case, the peptide sequence estimate is uniquely determined during sequencing of the data-encoded peptide. In the second output, multiple estimates of the peptide sequence are generated. In the third output, no estimate is found.
[0167] Figure 11 A flowchart illustrating exemplary steps of the disclosed computer-implemented method (referred to as 1100 for convenience) is shown. Method 1100 includes steps 1110, 1120, 1130, 1140, and 1150.
[0168] In step 1110, the raw data of the experimental spectrum is preprocessed to remove uninterpretable peaks, resulting in preprocessed data. As mentioned above, the raw data of the experimental spectrum is histogram data of intensity versus mass-to-charge ratio, which was initially obtained from the analysis of the data-encoded peptides by mass spectrometry. Step 1110 corresponds to step 110 of the two-stage sequencing method 100 and step 302 of the highest intensity tag-based sequencing method 300. Now, by means of... Figure 12 Further explanation of step 1110, Figure 12 A flowchart of an exemplary step performed in step 1110 is shown.
[0169] For example, step 1110 includes steps 1210, 1220, 1230, and 1250. In step 1210, the set of mass-to-charge ratios of the raw data is divided into multiple subsets, where each mass-to-charge ratio has a signal peak in the experimental spectrum. Each subset consists of the mass-to-charge ratio of the isotopes of the corresponding fragments of the data-encoded peptide. Furthermore, the isotopes of the fragments have the same positive integer charge value. Typically, the charge has a value of 1, 2, or 3. Then, in step 1220, the monoisotopic mass of the fragments is calculated for each subset. In step 1230, the intensity of the fragments is also calculated based on the intensity associated with the mass-to-charge ratios in each subset. For example, the intensity of the fragments can be calculated by summing the intensity values associated with different mass-to-charge ratios in each subset. Steps 1220 and 1230 are repeated for multiple subsets. In step 1250, preprocessed data is generated. The preprocessed data includes multiple masses and multiple intensities. The multiple masses are formed by the monoisotopic masses of the fragments in each subset. Each mass in a plurality of masses is associated with a corresponding strength in a plurality of strengths.
[0170] Preferably, the preprocessed data further includes a first mass set of presumed b ions and a second mass set of presumed y ions from the experimental spectrum. Step 1110 further includes a step 1240 of distributing the corresponding masses from the plurality of masses into the first mass set and the second mass set before performing step 1250.
[0171] Please refer to Figure 11 After obtaining the preprocessed data in step 1110, in step 1120, a first set of one or more candidate sequences of peptide sequences competing for the peptide is identified from the spectra, wherein the spectra are formed based on the preprocessed data rather than the raw data. As a result, fewer candidate sequences are generated, thereby advantageously reducing the time cost in sequencing. Step 1120 corresponds to step 120 of the two-stage sequencing method 100 and steps 304, 306, 308, 310, 312, 314, and 316 of the highest-strength tag-based sequencing method 300.
[0172] After obtaining the first candidate sequence set in step 1120, the first candidate sequence set is processed in step 1130 to estimate the peptide sequence.
[0173] After obtaining a set of one or more estimated peptide sequences, it is necessary to verify the validity of each estimated sequence so that any invalid peptide sequence estimates can be eliminated. Step 1140 is to verify whether each peptide sequence estimate is invalid. In step 1140, whether each peptide sequence estimate is invalid can be determined by: verifying the length of each peptide sequence estimate against a predetermined correct length of the peptide sequence; performing a sequence check on the amino acids "G" and "L" in each peptide sequence estimate; performing a sequence check based on the sequence check bits carried in the data-encoded peptide; or any combination thereof. In step 1150, the set of one or more estimated peptide sequences is cleaned to remove any invalid peptide sequence estimates found in step 1140.
[0174] Preferably, in step 1120, each candidate sequence in the first candidate sequence set is identified based on a search of a suitable mass set. If the precursor ion peak identified in the experimental spectrum has a charge value of 2, then preferably, each candidate sequence in the first candidate sequence set is identified or determined by searching a second mass set. If the precursor ion peak has a charge value of 3, then each candidate sequence in the first candidate sequence set is identified by searching both the first and second mass sets.
[0175] Step 1120 includes steps 130 and 140, and is detailed below.
[0176] If the first candidate sequence set consists of a single candidate sequence, then the single candidate sequence in the first candidate sequence set is designated as an estimated sequence of the peptide sequence. Otherwise, one or more candidate sequences are selected from the first candidate sequence set to form a second set of one or more candidate sequences, such that one or more candidate sequences in the second candidate sequence set are more likely to be the peptide sequence than the unselected candidate sequences in the first candidate sequence set (step 130).
[0177] If the second candidate sequence set consists of a single candidate sequence, then the single candidate sequence in the second candidate sequence set is designated as an estimated sequence of the peptide sequence.
[0178] If the second candidate sequence set consists of multiple candidate sequences that do not contain any AAC, then the multiple candidate sequences in the second candidate sequence set are designated as the estimated sequences of the peptide sequence.
[0179] If the second candidate sequence set consists of multiple candidate sequences containing one or more AACs, the original data is used to determine the valid amino acid sequences for the one or more AACs. After the valid amino acid sequences are determined, the candidate sequences in the second candidate sequence set are subdivided to produce a third set of one or more candidate sequences. The third candidate sequence set is obtained by replacing the one or more AACs from the second candidate sequence set with the determined valid amino acid sequences and then discarding any candidate sequences containing undetermined AACs (step 141).
[0180] If the third candidate sequence set consists of a single candidate sequence, then the single candidate sequence in the third candidate sequence set is designated as an estimated sequence of the peptide sequence. Otherwise, one or more candidate sequences are selected from the third candidate sequence set to form a fourth set of one or more candidate sequences, such that one or more candidate sequences in the fourth candidate sequence set are more likely to be the peptide sequence than the unselected candidate sequences in the third candidate sequence set (steps 142 and 143).
[0181] If the fourth candidate sequence set consists of a single candidate sequence, then that single candidate sequence in the fourth candidate sequence set is designated as an estimated sequence of the peptide sequence. Otherwise, multiple candidate sequences in the fourth candidate sequence set are designated as estimated sequences of the peptide sequence.
[0182] In some embodiments, step 130 is described as follows: First, one or more candidate sequences with the longest consecutive amino acid length are selected from the candidate sequences in the first candidate sequence set to form a fifth candidate sequence set (step 131). If the fifth candidate sequence set consists of a single candidate sequence, then the fifth candidate sequence set is designated as the second candidate sequence set. Otherwise, one or more candidate sequences with the largest number of amino acids are selected from the multiple candidate sequences in the fifth candidate sequence set to form a sixth candidate sequence set (step 132). If the sixth candidate sequence set consists of a single candidate sequence, then the sixth candidate sequence set is designated as the second candidate sequence set. Otherwise, one or more candidate sequences with the smallest matching error are selected from the multiple candidate sequences in the sixth candidate sequence set to form a seventh candidate sequence set (step 133). If the seventh candidate sequence set consists of a single candidate sequence, then the seventh candidate sequence set is designated as the second candidate sequence set. Otherwise, one or more candidate sequences with the highest average intensity value of the acquired amino acids are selected from the multiple candidate sequences in the seventh candidate sequence set to form an eighth candidate sequence set (step 134). If the eighth candidate sequence set consists of a single candidate sequence, then the eighth candidate sequence set is designated as the second candidate sequence set. Otherwise, select one or more candidate sequences with different offsets and the highest occurrence frequency from multiple candidate sequences in the eighth candidate sequence set to form the second candidate sequence set (step 135).
[0183] In some embodiments, step 140 is described as follows: One or more candidate sequences having the maximum number of amino acids in the determined valid amino acid sequences are selected from the third candidate sequence set to form a ninth candidate sequence set (step 142). If the ninth candidate sequence set consists of a single candidate sequence, then the ninth candidate sequence set is designated as the fourth candidate sequence set. Otherwise, one or more candidate sequences having the minimum matching error in the determined valid amino acid sequences are selected from the ninth candidate sequence set to form the fourth candidate sequence set (step 143).
[0184] Step 1120 can be implemented as step 120. In some embodiments, step 1120 is described as follows.
[0185] First, one or more valid paths with the longest length are identified in the spectrum, where a valid path begins at a head vertex and ends at a tail vertex with the quality of the peptide encoded by the data. Then, each of the identified paths is used to generate new candidate sequences. These new candidate sequences are assigned to a first set of candidate sequences, unless they are discarded for some other consideration.
[0186] Step 1120 can be implemented as steps 304, 306, 308, 310, 312, 314, and 316. In some embodiments, step 1120 includes the following steps.
[0187] Step A The plurality of masses are sorted in descending order of their corresponding intensities to form an ordered mass sequence, wherein the ordered mass sequence ranks the plurality of masses, and the highest-ranked mass is associated with the highest corresponding intensity (step 304).
[0188] Step B Given a selected quality and a selected number of higher-ranking qualities in the ordered quality sequence, identify one or more tags based on the highest intensity in the spectrum (step 310). Each tag based on the highest intensity is a partial sequence consisting of a first amino acid with a selected quality and multiple remaining amino acids, each with a quality selected from a selected number of higher-ranking qualities.
[0189] Step C Processing the one or more tags based on the highest strength. Specifically, processing each tag based on the highest strength includes the following tasks: determining the prefix and suffix of each tag based on the highest strength (steps 312 and 314); if the prefix and suffix are successfully determined, combining the prefix, the tags based on the highest strength, and the suffix to form a new candidate sequence (step 316); and assigning this new candidate sequence to the first candidate sequence set.
[0190] When determining the prefix, it is preferable that the prefix is identified as a valid first path in the spectrum with the longest length, wherein the first path connects the first amino acid and the heads of each tag based on the highest intensity, and allows for one or more first AACs with a total length of at most a first maximum length in the identified first path. If the charge value of the parent ion peak is 2, it is preferable to set the first maximum length to L when determining the prefix. p1 If the charge value of the parent ion peak is 3, then when searching for the first path, it is preferable to initially set the first maximum length to L. p1 If the first maximum length is L p1 If no first path is found, the first maximum length is increased to L. p2 L p2 >L p1 .
[0191] When determining the suffix, it is preferred that the suffix be identified as a valid second path in the spectrum with the longest length, wherein the second path connects the tail and tail amino acids of each tag based on the highest intensity, and that one or more second AACs with a total length of at most the second maximum length are allowed in the identified second path. If the charge value of the parent ion peak is 2, then the second maximum length is preferably set to L. s1 If the charge value of the parent ion peak is 3, then when searching for the second path, it is preferable to initially set the second maximum length to L. s1 If the second maximum length is L s1 If no second path is found, the second maximum length is increased to L. s2 L s2 >L s1 .
[0192] Step D Steps B and C are repeated sequentially using consecutive values from a strictly decreasing sequence of values as a selected number for higher-ranking quality, until the first candidate sequence set is not empty or the strictly decreasing sequence is exhausted (steps 322 and 324).
[0193] Step E Starting with the highest-ranked quality, step BD is repeated sequentially using consecutive qualities from the ordered quality sequence as selected qualities until the first candidate sequence set is not empty or a pre-selected number of higher-ranked qualities is reached in the ordered quality sequence (steps 332 and 334). The pre-selected number is J. max .
[0194] The disclosed computer implementation method can be implemented using a computer system with appropriate programming. This computer system can be implemented using one or more computers. These computers can be general-purpose computers, workstations, computing servers, distributed servers in a computing cloud, laptops, mobile computing devices, etc.
[0195] This disclosure may be implemented in other specific forms without departing from the spirit or essential characteristics of this disclosure. Therefore, the embodiments of this disclosure are to be considered exemplary rather than restrictive in all respects. The scope of the invention is defined by the appended claims rather than by the foregoing description; therefore, all changes made within the meaning and scope of the equivalents of the claims should be included within the scope of this invention.
Claims
1. A computer-implemented method for sequencing data-encoded peptides based on an experimental spectrum, wherein the raw data of the experimental spectrum is data from an intensity-to-mass-charge ratio histogram initially obtained from mass spectrometry analysis of the data-encoded peptides, the method comprising: The raw data is preprocessed to remove uninterpretable peaks, resulting in preprocessed data. A first set of one or more candidate sequences for identifying competing peptide sequences from a spectrum, wherein the spectrum is formed based on preprocessed data rather than raw data to generate a smaller number of candidate sequences, thereby reducing the time cost of sequencing. Process the first set of candidate sequences to estimate the peptide sequence; After obtaining a set of one or more estimated sequences of the peptide sequence, verify whether each estimated peptide sequence is invalid; as well as Clean the set of one or more estimated peptide sequences to remove any invalid peptide sequence estimates found. The preprocessing of the original data to generate preprocessed data includes: The mass-to-charge ratio set of the original data is divided into multiple subsets, such that each subset consists of the mass-to-charge ratio of the isotopes of the data-encoded peptide fragments, each mass-to-charge ratio having a signal peak in the experimental spectrum, and the isotopes of the fragments having the same positive integer charge. Calculate the single isotopic mass of the fragments in each subset; The intensity of the fragment is calculated based on the original data and the mass-to-charge ratio of each subset; and Generate preprocessed data, including: The preprocessed data includes multiple masses and multiple intensities, wherein the multiple masses are formed by the monoisotopic masses of fragments from various subsets, and wherein each mass is associated with a corresponding intensity; and The preprocessed data also includes a first mass set of estimated b ions and a second mass set of estimated y ions from the experimental spectra, wherein the first mass set and the second mass set are generated by distributing corresponding masses from the plurality of masses into the first mass set and the second mass set. The identification of the first candidate sequence set from the spectrum includes the following steps: (a) The plurality of masses are sorted in descending order of their respective intensities to form an ordered mass sequence, wherein the ordered mass sequence ranks the plurality of masses, and the highest-ranked mass is associated with the highest corresponding intensity. (b) Given a selected mass and a selected number of higher-ranked masses in the ordered mass sequence, identify one or more tags based on the highest intensity in the spectrum, wherein each tag based on the highest intensity is a partial sequence consisting of a first amino acid having a selected mass and a plurality of remaining amino acids having masses selected from the selected number of higher-ranked masses; (c) Processing the one or more tags based on the highest strength, wherein processing each tag based on the highest strength includes: Determine the prefix and suffix for each tag based on the highest strength; If the prefix and suffix are successfully determined, the prefix, the respective tags based on the highest strength, and the suffix are combined to form a new candidate sequence; and The new candidate sequence is assigned to the first candidate sequence set; (d) Repeat steps (b) and (c) sequentially using consecutive values from a strictly decreasing sequence of values as a selected number for higher-ranking quality, until the first candidate sequence set is not empty or the strictly decreasing sequence is exhausted; and (e) Starting with the highest-ranked quality, repeat steps (b)-(d) sequentially using consecutive qualities in the ordered quality sequence as selected qualities until the first candidate sequence set is not empty or a pre-selected number of higher-ranked qualities are reached in the ordered quality sequence.
2. The method according to claim 1, wherein, Identifying the first candidate sequence set from the spectral graph further includes: If the parent ion peak identified in the experimental spectrum has a charge value of 2, then each candidate sequence in the first candidate sequence set is identified by searching the second mass set; and If the charge value of the parent ion peak is 3, then each candidate sequence in the first candidate sequence set is identified by searching the first mass set and the second mass set.
3. The method of claim 1, wherein processing the first candidate sequence set to estimate the peptide sequence comprises: If the first candidate sequence set consists of a single candidate sequence, then the single candidate sequence in the first candidate sequence set is designated as an estimated sequence of the peptide sequence; otherwise, one or more candidate sequences are selected from the first candidate sequence set to form a second set of one or more candidate sequences, such that one or more candidate sequences in the second candidate sequence set are more likely to be the peptide sequence than the unselected candidate sequences in the first candidate sequence set. If the second candidate sequence set consists of a single candidate sequence, then the single candidate sequence in the second candidate sequence set is designated as an estimated sequence of the peptide sequence; If the second candidate sequence set consists of multiple candidate sequences that do not contain any amino acid combination AAC, then the multiple candidate sequences in the second candidate sequence set are designated as the estimated sequences of the peptide sequence; If the second candidate sequence set consists of multiple candidate sequences containing one or more AACs, then the original data is used to determine the valid amino acid sequences for the one or more AACs; After the effective amino acid sequence is determined, the candidate sequences in the second candidate sequence set are subdivided to generate a third set of one or more candidate sequences, wherein the third candidate sequence set is obtained from the second candidate sequence set by replacing the one or more AACs with the determined effective amino acid sequence and then discarding any candidate sequences with an undetermined AAC. If the third candidate sequence set consists of a single candidate sequence, then the single candidate sequence in the third candidate sequence set is designated as an estimated sequence of the peptide sequence; otherwise, one or more candidate sequences are selected from the third candidate sequence set to form a fourth set of one or more candidate sequences, such that one or more candidate sequences in the fourth candidate sequence set are more likely to be the peptide sequence than the unselected candidate sequences in the third candidate sequence set. as well as If the fourth candidate sequence set consists of a single candidate sequence, then the single candidate sequence in the fourth candidate sequence set is designated as an estimated sequence of the peptide sequence; otherwise, multiple candidate sequences in the fourth candidate sequence set are designated as estimated sequences of the peptide sequence.
4. The method of claim 3, wherein selecting one or more candidate sequences from the first candidate sequence set to form the second candidate sequence set comprises: One or more candidate sequences with the longest consecutive amino acid length are selected from the candidate sequences in the first candidate sequence set to form the fifth candidate sequence set; If the fifth candidate sequence set consists of a single candidate sequence, then the fifth candidate sequence set is designated as the second candidate sequence set; otherwise, one or more candidate sequences with the largest number of amino acids are selected from the multiple candidate sequences in the fifth candidate sequence set to form the sixth candidate sequence set. If the sixth candidate sequence set consists of a single candidate sequence, then the sixth candidate sequence set is designated as the second candidate sequence set; otherwise, one or more candidate sequences with the minimum matching error are selected from the multiple candidate sequences in the sixth candidate sequence set to form the seventh candidate sequence set. If the seventh candidate sequence set consists of a single candidate sequence, then the seventh candidate sequence set is designated as the second candidate sequence set; otherwise, one or more candidate sequences with the highest average intensity value of the obtained amino acids are selected from the multiple candidate sequences in the seventh candidate sequence set to form the eighth candidate sequence set. as well as If the eighth candidate sequence set consists of a single candidate sequence, then the eighth candidate sequence set is designated as the second candidate sequence set; otherwise, one or more candidate sequences with different offsets and the most frequent occurrences of different ion types are selected from the multiple candidate sequences in the eighth candidate sequence set to form the second candidate sequence set.
5. The method of claim 3, wherein selecting one or more candidate sequences from the third candidate sequence set to form the fourth candidate sequence set comprises: One or more candidate sequences with the largest number of amino acids among the determined valid amino acid sequences are selected from the third candidate sequence set to form the ninth candidate sequence set; as well as If the ninth candidate sequence set consists of a single candidate sequence, then the ninth candidate sequence set is designated as the fourth candidate sequence set; otherwise, one or more candidate sequences with the smallest matching error in the determined valid amino acid sequences are selected from the ninth candidate sequence set to form the fourth candidate sequence set.
6. The method of claim 1, wherein determining the prefix and suffix of each tag based on the highest strength comprises: The prefix is identified as the first path with the longest length in the spectrum, where: The first pathway connects the first amino acid to the heads of each tag based on the highest strength. One or more first amino acid combinations AACs with a total length of at most the first maximum length are allowed to appear in the identified first pathway; If the charge value of the parent ion peak determined in the experimental spectrum is 2, then the first maximum length is set to L. p1 ;as well as If the charge value of the parent ion peak is 3, then when searching for the first path, the first maximum length is initially set to L. p1 If the first maximum length is L p1 If no first path is found, the first maximum length is increased to L. p2 L p2 >L p1 ; as well as The suffix is identified as the second path with the longest length in the spectrum, wherein: The second pathway connects the tail and tail amino acids of each tag based on the highest strength; One or more second AACs with a total length of at most the second maximum length are allowed to appear in the identified second path; If the charge value of the parent ion peak is 2, then the second maximum length is set to L. s1 ;as well as If the charge value of the parent ion peak is 3, then when searching for the second path, the second maximum length is initially set to L. s1 If the second maximum length is L s1 If no second path is found, the second maximum length is increased to L. s2 L s2 >L s1 .
7. The method according to claim 1, wherein, Verifying whether the estimated sequence of each peptide is invalid includes checking the length of each estimated sequence against the predetermined correct length of the peptide sequence.
8. The method according to claim 1, wherein, Verifying whether the estimated sequences of each peptide are invalid includes checking the order of amino acids "G" and "L" in each estimated sequence.
9. The method according to claim 1, wherein, Verifying whether the estimated sequence of each peptide is invalid includes performing a sequence check based on the sequence check bits carried in the data-encoded peptide.
Citation Information
Patent Citations
Data storage using peptides
US11315023B2
Data storage using peptides
US20210090688A1