Method for non-targeted identification of peptide sequences
A method using a device with a peptide library and mass spectrometry identifies small peptides in complex mixtures by comparing mass/charge ratios, addressing the challenge of characterizing peptides of 2 to 6 amino acids, achieving rapid and accurate identification.
Patent Information
- Application Number
- FR2024001353
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-12
- Publication Date
- 2025-08-15
AI Technical Summary
Current methods fail to efficiently characterize small peptides of 2 to 6 amino acids in complex mixtures due to their small size and diverse physicochemical properties, making separation and detection difficult.
A method utilizing a device with a circuit and memory, containing a library of predetermined peptides, identifies peptide sequences by comparing mass/charge ratios and fragment ions, allowing non-targeted identification of peptides without prior hypotheses on their structure or content.
Enables systematic identification of all peptides in a sample, including those with modified side chains, without prior assumptions about their sequence, facilitating rapid and accurate characterization.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method for non-targeted identification of peptide sequences Technical field
[0001] The present invention relates to a method for non-targeted identification of peptide sequences of peptides comprising between 2 and 6 amino acids present in a sample, the method being implemented by a device comprising a circuit and a memory, as well as a method for analyzing peptides of a sample. The present invention also relates to a method for analyzing peptides having between 2 and 6 amino acids of a sample implementing the identification method of the invention. Prior art
[0002] Given their physicochemical properties and potential activities, small peptides derived from proteins are of great interest to the food, pharmaceutical and cosmetic industries. However, enzymatic hydrolysis of proteins leads to complex mixtures of peptides of varying sizes and properties.
[0003] Analysis of the overall peptide / protein content and analysis of free amino acids in a sample is well established. Protein analysis by proteomics is also known.
[0004] However, small peptides are not covered by conventional proteomic or peptide analysis methods, and a finer characterization of small peptides in complex mixtures remains a challenge due to their small size and extremely different physicochemical characteristics, which make their separation by liquid chromatography (LC) and their detection by mass spectrometry (MS) difficult.
[0005] Typically, the Applicant company develops yeast extracts, food matrices obtained by the hydrolysis of yeast cream. During the hydrolysis step, the proteins of the yeast cream are fragmented into complex mixtures comprising peptides having between 2 and 35 amino acids and free amino acids. The proteins of the yeast cream can be analyzed by proteomic analysis methods. There are also peptidomic analysis techniques for characterizing peptides comprising at least 5 or 6 amino acids. However, current techniques do not allow for rapid characterization of all the short peptides or small peptides in a yeast extract sample, even though these small peptides constitute a major constituent of most yeast extracts.
[0006] There is therefore a need for a method allowing systematic and non-targeted identification of all the small peptides in a sample.
[0007] Summary
[0008] A method for non-targeted identification of peptide sequences of peptides comprising between 2 and 6 amino acids present in a sample is provided, the method being implemented by a device comprising a circuit and a memory, the memory comprising: - at least one library of predetermined peptides comprising between 2 and 6 amino acids, each predetermined peptide is defined by a peptide sequence, the library comprising for each predetermined peptide: • the amino acids constituting the peptide sequence with at least one terminal amino acid labeled with an end marker, and a total mass / charge ratio (m / z) for the peptide sequence, • fragment ions calculated from the peptide sequence, said calculated fragment ions being respectively associated with mass / charge ratios, at least one or a plurality of calculated fragment ions comprise the end marker, - sets of information respectively associated with peptides present in the sample, called measured peptides, measured by a mass spectrometer, each measured peptide comprises a peptide sequence having at least one end amino acid labeled with an end marker, each set of information of a measured peptide comprises at least one MS / MS spectrum comprising a total mass / charge ratio of the measured peptide and one or a plurality of mass / charge ratios of the measured fragment ions, the measured peptides and the predetermined peptides being labeled at the same end by the same end marker, for each set of information of a measured peptide,the method comprises: / a / determining a selection of predetermined peptides from the library and said at least one MS / MS spectrum, each predetermined peptide of the selection having a total mass / charge ratio equal to the total mass / charge ratio of the peptide measured within an error range less than a predetermined value, , / b / a determination of the labeled end amino acid of the measured peptide by comparing the mass / charge ratios of the measured fragment ions with the calculated mass / charge ratios of the fragment ions of the predetermined peptides of the selection within an error range less than a predetermined value, / c / a determination of at least one reconstructed peptide sequence for the measured peptide from the calculated fragment ions of the selection, the determined labeled end amino acid and the measured fragment ions, - a rendering of an identification list of peptides measured in said sample in which each measured peptide in the list is defined by a reconstructed peptide sequence.
[0009] Advantageously, it is thus possible to obtain at the output of the method an identification list which includes all the peptides of the sample having a peptide sequence having a length (or a number of amino acids) of between 2 and 6 amino acids, typically between 2 and 4 amino acids, without there having been (having been) prior a priori hypotheses on the structure and the content of the different peptides contained in the sample. As will be detailed below, the peptides present in the sample may comprise modified amino acids, in particular modified amino acids containing modified side chains, the modification being able to be of natural or synthetic origin.
[0010] Each measured fragment ion may be representative of the presence of one or a plurality of linked amino acids in the peptide sequence of the measured peptide.
[0011] By total mass / charge ratio, it can be understood the mass / charge ratio of the entire measured peptide, i.e. comprising all the amino acids constituting the peptide sequence of the measured peptide labeled with an end marker.
[0012] In one or more embodiments, the library may be divided into two separate libraries or a plurality of separate libraries. More specifically, the data in the library may be distributed across one or a plurality of separate libraries.
[0013] According to one or more examples, the comparison in / b / may relate to the mass / charge ratios of the measured fragment ions having the end marker (detected in the MS / MS spectrum of the measured peptide) which are compared with the mass / charge ratios of the fragment ions of the selection, and / or with the mass / charge ratios of the fragment ions of the selection having the end marker. This comparison may be carried out within an error range less than a predetermined value.
[0014] In one or more embodiments, the calculated fragment ions may be calculated from the peptide sequence by applying a fragmentation method to the peptide sequence.
[0015] According to one or more embodiments, the fragmentation method may be defined by one or a plurality of fragmentation rules, which may be included in a library of fragmentation rules. The calculated fragment ions may be calculated from a predetermined peptide by applying the fragmentation rule(s) of the library.
[0016] In one or more embodiments, the fragmentation method may be identical to the fragmentation method of the mass spectrometer.
[0017] In one or more embodiments, the calculated fragment ions (or a set of calculated fragment ions) may be calculated from a predetermined peptide by applying a plurality of fragmentation methods, for example defined in the fragmentation rule library. More specifically, sets of calculated fragment ions may be calculated from a predetermined peptide to which a plurality of fragmentation methods are respectively applied, each set of calculated fragment ions being obtained from a respective fragmentation method.
[0018] Thus, advantageously, the same library can be used to identify peptide sequences in a sample which would have been measured by different mass spectrometers or the same mass spectrometer comprising a plurality of mass analyzers, preferably by mass spectrometers capable of carrying out a tandem analysis using a specific fragmentation method or capable of using different fragmentation methods.
[0019] In one or more embodiments, the library of predetermined peptides comprises all possible peptide sequences of 2 to 6 amino acids obtained from a group of N amino acids, each sequence of the predetermined library defines a predetermined peptide of the library and further comprises a terminal amino acid which is labeled with the terminal tag.
[0020] By all possible peptide sequences, it can be understood all possible combinations forming peptide sequences of a length between 2 and 6 amino acids obtained from a group of N amino acids, and for which an end amino acid is labeled with the end marker.
[0021] In one or more embodiments, N may comprise at least 22 proteinogenic amino acids and one or a plurality of modified amino acids. In one or more embodiments, a modified amino acid may be modified from a proteinogenic amino acid. In one or more embodiments, N may be greater than 2, preferably greater than 20. In one or more embodiments, N may be less than 500, 400, 300 or 200. N may typically be between 2 and 100, for example between 2 and 50.
[0022] In one or more embodiments, the library comprises at least one predetermined peptide which is broken down into a plurality of predetermined charged peptides, each predetermined charged peptide having the same peptide sequence as said at least one predetermined peptide and a respective charge.
[0023] By predetermined charged peptide, it can be understood a predetermined peptide having one or more charge(s), chosen from one or more positive charge(s) (e.g. +1, +2, or +3) or one or more negative charge(s) (e.g. -1, -2, -3).
[0024] Thus advantageously, it is possible to determine a peptide sequence of a measured peptide which comprises one or more charges, for example which is monocharged, dicharged, or tricharged.
[0025] In one or more embodiments, the end labeled with the end tag is the same for all peptide sequences in the predetermined library.
[0026] In one or more embodiments, the labeled end of the labeled end amino acid is A-terminal and / or C-terminal, preferably A-terminal.
[0027] In one or more embodiments, the total mass / charge ratio (m / z) of a predetermined peptide is determined from the mass / charge ratios of the amino acid residues constituting the sequence of the predetermined peptide.
[0028] In one or more embodiments, the mass spectrometer is a mass spectrometer with tandem analysis capability (MS / MS), preferably coupled to liquid chromatography (LC).
[0029] In one or more embodiments, in / c / , said at least one peptide sequence of a measured peptide is reconstructed from: - the determined labeled end amino acid,
[0030] - comparisons between the mass / charge ratios of the measured fragment ions and the calculated fragment ions.
[0031] Comparisons between the mass / charge ratios of the measured fragment ions and the calculated fragment ions can be made taking into account an error range less than a predetermined value.
[0032] In one or more embodiments, the memory comprises at least a first library and a second library, the first library comprising the amino acids constituting the peptide sequence with at least one end amino acid labeled with the end tag, and the total mass / charge ratio (m / z) for the peptide sequence, and the second library being determined only for the predetermined peptides of the selection, the second library comprising at least the predetermined peptides of the selection, and comprising fragment ions calculated respectively from the peptide sequences of the predetermined peptides of the selection, the calculated fragment ions being associated respectively with mass / charge ratios, at least one or a plurality of calculated fragment ions comprise the end tag.
[0033] Thus, advantageously, it is possible to reduce the calculation time for determining a reconstructed peptide sequence.
[0034] In one or more embodiments, the second library can be determined from at least one fragmentation method applied to the predetermined peptides of the selection.
[0035] In one or more embodiments, each set of information relating to a measured peptide may also include mass spectrum information (referred to as MS spectrum), wherein the MS spectrum information may include a mass-to-charge ratio of the measured peptide (i.e., parent ion), a retention time, and an intensity of the measured peptide (i.e., intensity of the parent ion).
[0036] The retention time included in the MS spectrum (or information set) of a measured peptide can be obtained from the measurement carried out following separation by liquid chromatography.
[0037] In one or more embodiments, the selection may be alternatively determined from the MS spectrum information of the measured peptide and the library.
[0038] In one or more embodiments, the one or more reconstructed peptide sequences for each measured peptide may be determined using a sequencing method derived from the methods used in so-called "de novo" sequencing.
[0039] The peptide sequence of a measured peptide is reconstructed using the principles of the methods used in so-called "de novo" sequencing using the MS / MS spectrum information of the measured peptide, the determined labeled end amino acid, and the calculated fragment ions of the predetermined peptides of the selection.
[0040] In one or more embodiments, when a plurality of reconstructed peptide sequences is determined for the same peptide measured in / c / , the method further comprises:
[0041] / dZ calculate, for each reconstructed peptide sequence, a score from the set of information associated with the measured peptide and the reconstructed peptide sequence,
[0042] / e / select the reconstructed peptide sequence for the peptide measured from from a comparison of the scores of the reconstructed peptide sequences, the peptide sequence with the highest score being selected.
[0043] Thus, advantageously, the use of score comparison in the case of co-elution of peptides makes it possible to ensure that the peptide sequence determined for a measured peptide corresponds to the measured peptide and to the set of information obtained for this measured peptide.
[0044] In one or more embodiments, each mass / charge ratio of the measured fragment ions is respectively associated with an intensity in the MS / MS spectrum, and in which the score of each reconstructed peptide sequence for the same measured peptide is determined from the following equation:
[0045] ^j of the intensities of the fragment ions measured from the reconstructed peptide sequence ^measured fragment ion intensities of the plurality of reconstructed peptide sequences
[0046] In one or more embodiments, only the measured fragment ions having said end marker are used to calculate a score of a peptide sequence.
[0047] In one or more embodiments, in / b / one or a plurality of labeled end amino acids is determined, and in / c / , for each determined labeled end amino acid, at least one reconstructed peptide sequence for the measured peptide is determined from the calculated fragment ions of the selection, the determined labeled end amino acid and the measured fragment ions.
[0048] In one or more embodiments, the predetermined value of the error range is less than 20 ppm, preferably less than 15 ppm, preferably less than 10 ppm, and even more preferably less than 5 ppm.
[0049] In one or more embodiments, the list may comprise retention times respectively associated with the measured peptides, each retention time being obtained from the set of information, in particular the MS spectrum for example, used to determine a reconstructed peptide sequence of a measured peptide.
[0050] Each measured peptide in the list can be associated with a respective retention time obtained from the information set (eg MS spectrum) used to determine the reconstructed peptide sequence of the measured peptide.
[0051] In one or more embodiments, each measured peptide in the list is associated, by an identification means, with the set of information used to determine the reconstructed peptide sequence of the measured peptide.
[0052] Advantageously, the use of an identification means makes it possible to quickly order the sets of information and to ensure that a set of information (e.g. MS / MS spectrum(s) and / or MS spectrum) is correctly associated with the reconstructed peptide sequence for which it was used. In addition, this facilitates the post-processing of the data (e.g. identification list, sets of information, etc.) obtained.
[0053] In one or more embodiments, the list comprises intensities respectively associated with the measured peptides, each intensity being obtained from the set of information (eg MS spectrum) used to determine a reconstructed peptide sequence of a measured peptide, and each measured peptide of the list comprises a proportion value of the peptide in the mixture, the proportion value “Qp”, for each measured peptide of the list, is calculated from:
[0054] „ Intensity of a measured peptide from list X / the intensities of the measured peptides from the list
[0055] Advantageously, it is thus possible to obtain values of relative proportions of the peptides of the mixture.
[0056] The intensity of each measured peptide in the list may have been obtained from the MS spectrum of the measured peptide that is included in the information set.
[0057] In one or more embodiments, the peptides are natural peptides and / or synthetic peptides.
[0058] In one or more embodiments, the peptides contain at least one labeled end amino acid and one or more amino acids selected from proteinogenic amino acids or modified amino acids.
[0059] In one or more embodiments, the labeled end amino acid is an A-terminal amino acid labeled with an A-terminal end tag selected from dansyl, dabsyl, 1-naphthyl isocyanate (IsoC), or 1-naphthyl isothiocyanate, each of these end tags optionally being labeled with one or more carbon 13 (13C), deuterium (D), nitrogen 15 (15N), and / or sulfur 34 (34S) atoms. The labeling makes said end tags distinguishable. The A-terminal end marker may for example be selected from dansyl, dansyl labeled with one or more carbon 13 atoms (13 C), dansyl labeled with deuterium (D), dabsyl, 1-naphthyl isocyanate (IsoC), or 1-naphthyl isothiocyanate.
[0060] In one or more embodiments, the labeled end amino acid is a C-terminal amino acid labeled with a C-terminal label selected from 2-(A,A'-dimethylamino)-1-ethylamine (DMED) or 3-(AA'-dimethylamino)-1-propylamine (DMAPA), each of these end labels optionally being labeled with one or more carbon 13 (13C), deuterium (D), and / or nitrogen 15 (15N) atoms. The labeling makes it possible to make said end labels differentiable. The C-terminal label may for example be selected from DMED or DMAPA.
[0061] In one or more embodiments, the peptides have been subjected to a protective treatment of amino acids containing a free thiol group such as cysteine, typically an alkylation treatment, the mass / charge ratios of the predetermined peptides of the library taking into account the protective treatment of the free thiol group of cysteine. This protective treatment is commonly carried out by carbamidomethylation typically using chloroacetamide or iodoacetamide. This protective treatment can also be carried out by other reagents known to those skilled in the art such as 4-hydroxyphenyl)ethyl iodoacetamide, iodoacetic acid and α-ethyl maleimide.
[0062] In one or more embodiments, the liquid chromatography is reversed-phase liquid chromatography, preferably carried out with a gradient of at least one hour.
[0063] In one or more embodiments, the mass spectrometer is equipped with a mass analyzer preferably selected from a quadrupole trap mass analyzer and a triple quadrupole mass analyzer, more preferably is equipped with a high-resolution mass analyzer selected from an orbital trap analyzer, a Q-TOF analyzer, a TIMS-TOF analyzer and an FT-ICR analyzer.
[0064] In one or more embodiments, the peptides present in the sample are natural peptides and / or synthetic peptides.
[0065] In one or more embodiments, the sample is a mixture of small natural peptides or a mixture of peptides derived from peptide synthesis. The sample is preferably a protein hydrolyzate, for example a yeast protein hydrolyzate, a plant protein hydrolyzate, or an animal protein hydrolyzate.
[0066] According to another aspect, there is provided a computer program comprising instructions for implementing the method according to the present disclosure when this program is executed by a processor.
[0067] According to another aspect, there is provided a non-transitory recording medium readable by a computer on which is recorded a program for implementing the method according to the present disclosure when this program is executed by a processor.
[0068] According to another aspect, there is provided a method for analyzing peptides having between 2 and 6 amino acids of a sample comprising: i) a step of treating the sample comprising at least one step ii) of labeling one end of the peptides of the sample by bringing the sample into contact with an end marker, ii) a step of analysis by liquid chromatography coupled to a mass spectrometer, preferably to a mass spectrometer with a tandem analysis capacity (MS / MS), iii) a data processing step implementing a method of non-targeted identification of peptide sequences as described in the present document.
[0069] In one or more embodiments, the step of treating the sample further comprises a step of alkylating the peptides of the sample by contacting with an alkylating agent, and wherein the alkylation step preferably precedes the step of treating the sample.
[0070] In one or more embodiments, the liquid chromatography is reversed-phase liquid chromatography, preferably carried out with a gradient of at least one hour.
[0071] In one or more embodiments, the mass spectrometer is equipped with a mass analyzer which is preferably a quadrupole trap mass analyzer or a triple quadrupole mass analyzer, preferably a high resolution mass analyzer selected from an orbital trap analyzer, a Q-TOF analyzer, a TIMS-TOF analyzer, and an FT-ICR analyzer.
[0072] In one or more embodiments, the sample contains a protein hydrolysate, for example a yeast protein hydrolysate, a plant protein hydrolysate, or an animal protein hydrolysate. Brief description of the drawings
[0073] Other characteristics, details and advantages will appear on reading the detailed description below, and on analyzing the attached drawings, in which: Fig.l
[0074] [Fig.l] illustrates a diagram of the method for non-targeted identification of peptide sequences of peptides from a sample. Fig. 2
[0075] [Fig.2] schematically illustrates an example of a library of predetermined peptides. Fig. 3
[0076] [Fig.3] schematically illustrates the phenomenon of co-elution of two peptides. Fig. 4
[0077] [Fig.4] illustrates a device for implementing the method of the present disclosure. Description of the embodiments
[0078] The terms "peak" and "line" may be interpreted in the same way and are interchangeable in the present disclosure. Similarly, the terms "intensity" and "amplitude" may be interpreted in the same way and are interchangeable in the present disclosure.
[0079] [Fig.l] illustrates a diagram of the method for non-targeted identification of peptide sequences of peptides from a sample.
[0080] In this document, the term "sample" means a representative portion of a larger population, selected for the purpose of carrying out observations, measurements or analyses to draw conclusions about the original population. In the context of the invention, the sample contains several peptides, which may have different amino acid sequence lengths, and for the same amino acid sequence length, different amino acid sequences.
[0081] The peptides may be natural peptides and / or synthetic peptides.
[0082] When the sample is a portion of a natural source or a portion of the product of one or more chemical or biochemical reaction(s) applied to a natural source, the peptides are referred to as "natural peptides". In one or more embodiments, the sample contains natural peptides.
[0083] When the sample is a portion of a population of peptides artificially manufactured in the laboratory by chemical or biotechnological synthesis methods, such as solid phase peptide synthesis (SPPS), liquid phase peptide synthesis or enzymatic synthesis, the peptides are referred to as "synthetic peptides" or "synthetic peptides". In one or more embodiments, the sample contains synthetic peptides.
[0084] In one or more embodiments, the sample containing natural peptides is a protein hydrolysate.
[0085] As used herein, protein hydrolysate, or simply hydrolysate, is a product derived from the breakdown of proteins into smaller peptides or even amino acids by a process called hydrolysis. Hydrolysis is a chemical reaction that involves the addition of water to break the peptide bonds between the amino acids that make up proteins. This process of protein fragmentation leads to the formation of smaller molecules that are more easily digested and absorbed by the body. The goal of producing protein hydrolysates is often to obtain products that can be quickly assimilated by the body, as the peptides and amino acids resulting from hydrolysis are generally more soluble and easier to absorb than intact proteins.These hydrolysates are commonly used in the food industry, dietary supplements, and medical preparations, particularly for people with specific needs regarding protein digestion or absorption. They can also be used in the cosmetics industry. Protein hydrolysis can be achieved by various methods, usually using chemical agents, enzymes, or a combination of both. The main hydrolysis methods used in industry, particularly the food industry, are acid hydrolysis and enzymatic hydrolysis. The hydrolyzate can be prepared from various sources, preferably natural. The hydrolyzate can thus be a hydrolyzate of microbial proteins, plant proteins, or animal proteins.The protein hydrolyzate may, for example, be a hydrolyzate of yeast or other microorganism proteins, a whey protein hydrolyzate, a casein protein hydrolyzate, a soy protein hydrolyzate, a fish protein hydrolyzate, a collagen hydrolyzate, a wheat protein hydrolyzate, a rice protein hydrolyzate.
[0086] In one or more embodiments, the sample of synthetic peptides is a sample of peptides derived from a chemical synthesis, preferably derived from a solid phase synthesis (SPPS) or liquid phase synthesis.
[0087] In one or more embodiments, one or more peptides present in the sample are bioactive.
[0088] As used herein, "small bioactive peptides" means small peptides exhibiting specific biological activities. These peptides may be of natural or synthetic origin and are capable of modulating various biological processes when they interact with receptors, enzymes or other target molecules in a living organism. These peptides may, for example, exhibit antimicrobial, anti-inflammatory, or antioxidant properties (such as the tripeptide glutathione).
[0089] The peptides identified by the method of the invention are advantageously short peptides (or in other words small peptides) having between 2 and 6 amino acids, typically between 2 and 5 amino acids, or between 2 and 4 amino acids. The peptides identified by the method of the invention can thus be chosen from dipeptides, tripeptides, tetrapeptides, pentapeptides, and hexapeptides, and mixtures thereof.
[0090] In one or more embodiments, the sample contains short peptides having between 2 and 6 amino acids, at least one end of which is labeled with an end tag.
[0091] Preferably, the sample contains a large number of short peptides having different sequences, preferably comprises at least 100 short peptides having different sequences, at least 200 short peptides having different sequences, at least 300 short peptides having different sequences, or at least 500 short peptides having different sequences.
[0092] In some embodiments, the sample contains at most 10,000 short peptides having different sequences, at most 5,000 short peptides having different sequences, at most 2,000 short peptides having different sequences, or at most 1,000 short peptides having different sequences.
[0093] In some embodiments, the sample contains between 100 and 2,000 short peptides having different sequences, between 100 and 1,500 short peptides having different sequences, or between 100 and 1,000 short peptides having different sequences.
[0094] In some embodiments, the sample contains between 150 and 2,000 short peptides having different sequences, between 200 and 2,000 short peptides having different sequences, between 250 and 2,000 short peptides having different sequences, between 150 and 1,000 short peptides having different sequences, between 200 to 1,000 short peptides with different sequences, or between 250 and 1,000 short peptides with different sequences.
[0095] A soy hydrolyzate typically contains around 150 short peptides with different sequences.
[0096] A casein hydrolysate typically contains around 300 peptides.
[0097] A yeast hydrolysate typically contains around 500 peptides.
[0098] As detailed below, the peptides present in the sample may comprise modified amino acids, regardless of the origin and nature of the modification. The modified amino acids may in particular have modified side chains.
[0099] In one or more embodiments, the labeled end of the labeled end amino acid, for example of one or a plurality of peptides of the sample, may be the A^-terminal and / or C-terminal. Preferably, the labeled end is the V-terminal.
[0100] Identification method
[0101] The method is preferably a non-targeted identification method.
[0102] The expression "non-targeted", relating to the identification method, means that the method is carried out without a priori hypothesis on the sequence and content of the different peptides contained in the sample. The nature and content of the different peptides contained in the sample are not known in advance. In other words, the composition of the sample in peptides and their peptide sequences is not known in advance.
[0103] The method may be implemented by a device that includes a memory and a processor.
[0104] The memory may comprise at least one library of predetermined peptides comprising between 2 and 6 amino acids, typically between 2 and 5 amino acids, or between 2 and 4 amino acids.
[0105] Each predetermined peptide can be defined by a peptide sequence, and the library can comprise for each predetermined peptide: • the amino acids constituting the peptide sequence with at least one terminal amino acid labeled with an end marker, and a total mass / charge ratio (m / z) for the peptide sequence, • fragment ions calculated from the peptide sequence, said calculated fragment ions being respectively associated with mass / charge ratios, at least one or a plurality of calculated fragment ions comprise the end marker.
[0106] The total mass / charge ratio (m / z) of a predetermined peptide can be determined from the amino acid sequence of the predetermined peptide. Indeed, the sum of all the mass / charge ratios of the amino acid residues of a peptide sequence of the library can correspond to the total mass / charge ratio of the predetermined peptide defined by this peptide sequence.
[0107] In one or more embodiments, the fragment ions (or set of fragment ions) can be calculated from the peptide sequence by applying a fragmentation method to the peptide sequence.
[0108] Note that, in one or more embodiments, the library may be divided into two separate libraries or a plurality of separate libraries. For example, a first library may comprise the amino acids constituting the peptide sequences with at least one end amino acid labeled with an end marker, and the total mass / charge ratios (m / z) for the peptide sequences, and a second library may comprise the predetermined peptides of the first library, the total mass / charge ratios of the predetermined peptides as well as fragment ions calculated for each predetermined peptide from their peptide sequence, the fragment ions being associated respectively with mass / charge ratios, and at least one or a plurality of calculated fragment ions comprise the end marker. This second library may thus be determined from the first library.In such a configuration, the method according to the present application may use both libraries or the plurality of libraries.
[0109] In this document, "peptide sequence" or "amino acid sequence" or "sequence of amino acid residues" refers to the description of the sequence of amino acids (monomers) that make up a peptide. Peptides are in fact ordered linear polymers, made up of amino acids. The sequence is generally represented in the form of a character string that is stored in a computer file in text format.
[0110] In this document, "amino acid" means a residue comprising a central carbon atom, the α-carbon, on which are articulated the carboxyl group (-COOH), the amino group (-NH2), a hydrogen atom (-H) and a side group (noted -R). It is the nature of this side group (also called side chain) which differentiates amino acids from each other. The central α-carbon is asymmetric in all cases except that of glycine (because glycine has a hydrogen in place of a side chain). It is possible to have two different conformations of the same amino acid, conformations which cannot be interconverted without breaking and reassembling covalent bonds. These conformations are called L and D.Although both can be synthesized in a single chemical reaction, only the L-form is used for protein synthesis (proteinogenic), while the non-proteinogenic D-form is incorporated into the synthesis of carbohydrates or lipids in living organisms. Amino acids can be proteinogenic amino acids or modified amino acids, among others.
[0111] By proteinogenic amino acid, or "unmodified amino acid" is meant the amino acids that are the basic building blocks of proteins. Proteinogenic amino acids are the basic building blocks of proteins. They polymerize to form linear polypeptides in which the amino acid residues are joined by peptide bonds. Protein biosynthesis takes place on ribosomes, which carry out the translation of messenger RNA into proteins. The order in which the amino acids are linked to the polypeptide chain is specified by the succession of codons carried by the sequence of the messenger RNA, which is a copy of the DNA of the cell nucleus; these codons, which are triplets of nucleotides, are translated into amino acids by transfer RNA according to the genetic code.This directly specifies 20 standard proteinogenic amino acids (A, R, N, D, C, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, V) to which two other amino acids are added through a more complex mechanism involving, for selenocysteine (U), a SECIS element which recodes the UGA stop codon and, for pyrrolysine (O), a PYLIS element which recodes the UAG stop codon. There are therefore a total of 22 proteinogenic amino acids. Pyrrolysine is only present in the proteins of certain methanogenic archaea, so that eukaryotes and bacteria only use 21 proteinogenic amino acids.
[0112] The most common amino acids are represented by a 1 or 3 letter code as shown in the table below. The symbols B, Z and J are used when the amino acid is not completely determined (ambiguous amino acid). Symbol Three-letter code Definition A Ala Alanine R Arg Arginine N Asn Asparagine D Asp Aspartic acid (aspartate) C Cys Cysteine Q Gin Glutamine E Glu Glutamic acid (glutamate) G Gly Glycine H His Histidine I Ile Isoleucine L Leu Leucine K Lys Lysine M Met Methionine F Phe Phenylalanine P Pro Proline O Pyl Pyrrolysine S Ser Serine U Sec Selenocysteine T Thr Threonine W Trp Tryptophan [Table 1]: Amino acids - symbols, three-letter code, and definitions
[0113] In this document, the term "peptide" refers indifferently to a peptide present in the sample, in other words to a measured peptide, or to a peptide from the library, in other words to a predetermined peptide.
[0114] In one or more embodiments, the peptides contain amino acids selected from proteinogenic acids and modified amino acids.
[0115] The peptides contain at least one terminal amino acid labeled with an end tag.
[0116] In one or more embodiments, the peptides contain at least one labeled end amino acid and one or more amino acids selected from proteinogenic amino acids and modified amino acids.
[0117] In one or more embodiments, the labeled terminal amino acid may be a proteinogenic amino acid or a modified amino acid. For example, the labeled terminal amino acid may be an amino acid with a modified side chain.
[0118] Modified amino acid means any non-proteinogenic amino acid. The modified amino acid may be the product of a chemical modification of endogenous or exogenous origin applied to a proteinogenic amino acid. The modified amino acid may also be a synthetic amino acid.
[0119] In this document, an endogenous reaction is a reaction that occurs naturally within an organism or a biological system, for example a post-translational modification type reaction, whereas an exogenous reaction is a reaction that is triggered or caused by factors external to the organism or biological system. Examples of exogenous reactions are the oxidation of peptides that occurs outside the organism, under laboratory or industrial conditions (eg chemical or photochemical oxidation), or a protection reaction with a protecting group, for example to protect an amino acid from oxidation.
[0120] In one or more embodiments, the modified amino acid is an amino acid whose side chain, carboxyl group, or amino group is modified.
[0121] In one or more embodiments, the modified amino acid is an amino acid containing a modified side chain. The modified side chain may, for example, be the product of an endogenous chemical modification such as a post-translational modification, the product of one or more exogenous chemical or biochemical reaction(s), such as the product of a reaction with a side chain protecting group. The modified amino acid may also be the product of a reaction of replacing an atom of the side chain with an isovalent atom, for example the replacement of a hydrogen atom with a fluorine atom, or of a sulfur atom with a selenium atom or of a nitrogen atom with a phosphorus atom.
[0122] Any chemical modification of amino acids, in particular of amino acid side chains, whatever the origin, can be detected by the method of the invention provided that it is implemented in the library.
[0123] Examples of modification of the side chains of cysteine, lysine and methionine are detailed below.
[0124] Cysteine is an amino acid that contains a thiol group (-SH), which makes it reactive. It may be necessary to protect cysteine from oxidation or other unwanted reactions. Protection of cysteine usually involves the formation of a derivative that makes it less reactive.
[0125] In some embodiments, the amino acid containing a modified side chain is a cysteine whose side chain is modified by a chemical reaction to render it non-reactive. For example, the cysteine may be carbamidomethylated. Carbadomethylation of the cysteine may be accomplished by contacting the sample with an alkylating agent, such as iodoacetamide, typically for the purpose of protecting the free thiol group of the cysteine from oxidation. Cysteine can be protected with other types of protecting groups, such as α-ethyl maleimide (Schilling, D., Barayeu, U., Steimbach, RR, Talwar, D., Miller, AK, & Dick, TP (2022). Commonly used alkylating agents limit persulfide detection by converting protein persulfides into thioethers. Angewandte Chemie, 134(30), e202203684. DOI: 10.1002 / anie.202203684. PMID: 35506673).
[0126] In some embodiments, the amino acid containing a modified side chain is a lysine whose side chain is acetylated, typically is an acetyllysine. Acetyllysine can be formed endogenously in certain proteins by post-translational modification of a lysine residue. This amino acid can thus be found in the peptides of the sample although not being a proteinogenic acid.
[0127] In some embodiments, the amino acid containing a modified side chain is a methionine whose side chain is oxidized, typically is a methionine sulfoxide (also called methionine sulfoxide).
[0128] Methionine sulfoxide can be formed in some proteins by post-translational modification of a methionine residue. This amino acid can thus be found in the peptides of the sample although it is not a proteinogenic acid.
[0129] In some embodiments, the amino acid is a selenium-containing amino acid. Selenium-containing amino acids are analogs of standard amino acids that contain selenium instead of sulfur in their structure. These selenium-containing amino acids can be incorporated into proteins during protein biosynthesis, in place of the corresponding sulfur-containing amino acids. The principal selenium-containing amino acids are selenomethionine (Sem) and selenocysteine (Sec), which are the selenium-containing equivalents of methionine and cysteine, respectively. Peptides comprising selenium-containing amino acids can be prepared by chemical peptide synthesis, recombinant expression, biocatalysis, or via postsynthetic modification techniques.
[0130] Importantly for the method of the invention, the peptides contain at least one terminal amino acid labeled with an end tag. The end tag amino acid may be an amino acid whose amine group (-NH2) or carboxyl group (-COOH) is modified to label the peptide of which it constitutes the V-terminus or C-terminus respectively.
[0131] As used herein, end markers are understood to mean agents or chemical groups specifically used to mark or label the N-terminal or C-terminal end of proteins or peptides. A distinction is therefore made between N-terminal end markers and C-terminal end markers. Such end markers are known in the fields of molecular biology research, proteomics, and in the fields of protein identification, purification and analysis.
[0132] In one or more embodiments, the amino acid is an iV-terminal amino acid whose amine group (-NH2) is modified with an iV-terminal tag.
[0133] The V-terminal end marker may in particular be selected from dansyl, dabsyl, 1-naphthyl isocyanate (IsoC), or 1-naphthyl isothiocyanate, each of these end markers optionally being labeled with one or more carbon 13 (13C), deuterium (D), nitrogen 15 (15N), and / or sulfur 34 (34S) atoms.
[0134] Dansyl optionally labeled with one or more carbon 13 (13C), deuterium (D), nitrogen 15 (15N), and / or sulfur 34 (34S) atoms is preferred.
[0135] In one or more embodiments, the terminal amino acid is a C-terminal amino acid whose carboxyl group (-COOH) is modified with a C-terminal tag.
[0136] The C-terminal end marker may in particular be selected from 2-(N,N'-dimethylamino)-1-ethylamine (DMED) and 3-(A,A'-dimethylamino)-1-propylamine (DMAPA), each of these end markers optionally being labeled with one or more carbon 13 (13C), deuterium (D), and / or nitrogen 15 (15N) atoms.
[0137] In one or more embodiments, the peptides present in the sample are peptides of which only one terminal amino acid is labeled with an end tag, in other words are peptides of which only one terminal is labeled with an end tag.
[0138] In one or more embodiments, the peptides are peptides whose two terminal amino acids are labeled with an end tag, in other words are peptides whose two ends are labeled.
[0139] The library may be calculated (or generated) such that the predetermined peptides contain amino acids selected from proteinogenic amino acids, modified amino acids, in particular amino acids containing a modified side chain, and mixtures thereof, it being understood that the predetermined peptides contain at least one end amino acid labeled with an end marker.
[0140] The library of predetermined peptides (or libraries) may be generated before each measurement of a sample by the device for implementing the method, or may be received by any means of communication. For example, the library may be generated on a third-party device and transmitted by it to the device implementing the method.
[0141] In one or more embodiments, the library of predetermined peptides may comprise all possible peptide sequences of 2 to 6 amino acids obtained from a group of N amino acids.
[0142] By all possible peptide sequences, it can be understood all possible combinations forming peptide sequences of a length between 2 and 6 amino acids, for example between 2 and 4 amino acids, obtained from a group of N amino acids.
[0143] For example, N can be in a range of 2 to 50 amino acids. There are only about 20 naturally occurring amino acids. However, as detailed below, the library can be constructed to account for the presence of acids modified amino acids, in the measured peptides whether they are modified amino acids of natural or synthetic origin. Thus, in one or more embodiments, N may preferably comprise at least the 22 proteinogenic amino acids and modified amino acids of said proteinogenic amino acids. In one or more embodiments, N may be greater than 2, preferably greater than 20. In one or more embodiments, N may be less than 500, 400, 300 or 200. N may typically be between 2 and 100, for example between 2 and 50.
[0144] Thus, each sequence of the predetermined library may define a predetermined peptide of the library and may further comprise a terminal amino acid which is labeled with the marker.
[0145] Furthermore, in one or more embodiments, the memory may comprise sets of information associated respectively with peptides present in the sample, called measured peptides, measured by a mass spectrometer, each measured peptide may comprise a peptide sequence having at least one terminal amino acid labeled with an end marker. Each set of information of a measured peptide may comprise at least one MS / MS spectrum comprising a total mass / charge ratio of the measured peptide and one or a plurality of mass / charge ratios of the measured fragment ions.
[0146] Each fragment ion may be representative of the presence of one or a plurality of amino acids (or amino acid residues) linked in the peptide sequence of the measured peptide.
[0147] Furthermore, according to one or more embodiments, each set of information relating to a measured peptide may also comprise mass spectrum information (referred to as MS spectrum), the MS spectrum information being able to comprise a mass / charge ratio of the measured peptide (i.e. of the parent ion), a retention time, and an intensity of the measured peptide (i.e. an intensity of the parent ion).
[0148] Further, the measured peptides and at least a plurality of predetermined peptides from the predetermined peptide library may be labeled at the same end with the same end tag.
[0149] According to one or more examples, the mass spectrometer may be a mass spectrometer with tandem analysis capability (MS / MS), preferably coupled to liquid chromatography (LC), for example using a liquid chromatographic column.
[0150] Such a mass spectrometer, optionally coupled to liquid chromatography, makes it possible to obtain sets of information on peptides measured in a sample such as mass spectra (MS) of the peptide (also called parent ion), and / or MS / MS fragmentation spectra giving information on the structure of the peptide (or parent ion) from the fragment ions present in the MS / MS spectra.
[0151] For example, the information sets relating to these spectra (MS and MS / MS) may include mass / charge ratios and peak intensities relating to fragment ions and / or parent ions constituting the measured peptides, as well as retention times of the measured peptides.
[0152] Furthermore, a set of information of a measured peptide may comprise an MS spectrum and one or a plurality of MS / MS spectra, generally close in their content. In these cases, it may be decided to use a single MS / MS spectrum (e.g. that of the parent ion with the highest intensity) or the plurality of MS / MS spectra.
[0153] Liquid chromatography (LC), using a chromatographic column for example, can be used to separate peptides before their measurement by the mass spectrometer, and also allows a retention time to be determined.
[0154] This retention time refers to the time a peptide from the sample remains in the chromatographic column (of the chromatograph) before being eluted, i.e. until the time the peptide reaches the mass spectrometer. This retention time (for each peptide) can therefore be measured from the time the sample is injected into the chromatograph until the time a peptide reaches the mass spectrometer.
[0155] [Fig.2] schematically illustrates an example of a library of predetermined peptides.
[0156] In this schematic example, the generated library can comprise all possible peptide sequences (SEQ_PEP) (twelve in this case) comprising 2 to 3 amino acids that can be obtained from a group of N=3 amino acids (A, B, C), and where each peptide sequence of the library can comprise an end amino acid that is marked by an end marker.
[0157] As previously described for the sample peptides, the labeled end of the terminal amino acid of the library peptide sequences may be the A-terminal and / or C-terminal end. Preferably, the labeled end is of the A-terminal (or A-terminal type) end.
[0158] For example, in the case of the library of [Fig. 2], the labeled end is the A-terminus. Thus, the end amino acid labeled with the end tag * for the predetermined peptide PP2 is amino acid B, the end amino acid labeled with the end tag * for the predetermined peptide PP8 is amino acid A, etc.
[0159] Each peptide sequence (SEQ_PEP) may respectively define a predetermined peptide PP1.. .PP12 of the library and may be used to determine or calculate the total mass / charge ratio (m / z) ml.. .ml2 of the predetermined peptide as described previously (R_seq_pep). For example, the mass / charge ratio (m / z) m7 may be calculated (by the sum for example) from the peptide sequence A*-BC of the peptide PP7 with A* the labeled end amino acid (i.e. by the sum of the mass / charge ratios of the amino acid residues constituting the peptide sequence).
[0160] It should also be noted that in this example of [Fig.2], the ratios ml and m2 are equal, since the peptide sequences of the predetermined peptides PP1 and PP2 comprise the same amino acids and in the same number. This reasoning can also be applied in the case of the predetermined peptides PP3 and PP4, the predetermined peptides PP5 and PP6, as well as the predetermined peptides PP7 to PP12.
[0161] Note that the mass / charge ratio of a labeled end amino acid takes into account the end marker in its mass / charge ratio. Indeed, an amino acid A does not have the same mass / charge ratio as an amino acid labeled A* with the end marker *.
[0162] Of course, in the example of [Fig.2], the labeled end is A-terminal (or A-terminal type) but it can also be C-terminal (or C-terminal type). For example, the predetermined library can comprise both predetermined peptides having an A-terminal type labeled end and predetermined peptides having a C-terminal type labeled end, and / or predetermined peptides having a C-terminal type labeled end and an A-terminal type labeled end.
[0163] According to one or more examples, the measured peptides and the predetermined peptides of the predetermined peptide library can be labeled at the same end with the same end tag.
[0164] For example, the measured peptides may all have an A-terminal (or A-terminal-like) labeled end, and the library of predetermined peptides comprises at least a plurality of predetermined peptides having an A-terminal labeled end among predetermined peptides having an A-terminal, and / or C-terminal, labeled end, and / or a combination of both.
[0165] In one or more embodiments, the peptide sequences of the library may all have the same labeled end, for example A-terminal or C-terminal.
[0166] In one or more embodiments, the amino acids of a group of N amino acids may be proteinogenic amino acids or amino acids modified. For example, modified amino acids can be amino acids whose side chain is modified.
[0167] Thus, the predetermined peptides may comprise a labeled A-terminal or C-terminal amino acid, proteinogenic amino acids and amino acids whose side chain is modified. Preferably, the amino acids whose side chain is modified are selected from the group consisting of cysteine, lysine and methionine.
[0168] According to one or more embodiments, the cysteine of the predetermined peptides is a cysteine whose side chain is modified, preferably by carbamidomethylation. This modification makes it possible in particular to avoid the unwanted formation of disulfide bonds during the preparation and analysis of the sample.
[0169] Further, according to one or more examples, the library may comprise at least one predetermined peptide of the library which is declined into a plurality of predetermined charged peptides, each predetermined charged peptide having the same peptide sequence as said at least one predetermined peptide and a respective charge.
[0170] By predetermined charged peptide, it can be understood a predetermined peptide having one or more positive or negative charges.
[0171] Thus, for example, it is possible to decline one or more predetermined peptides from the library into a multitude of charged versions, each charged version having its own charge (or respective charge) but having the same peptide sequence of the predetermined peptide from which it declines.
[0172] Thus advantageously, it is possible to determine a peptide sequence of a measured peptide which comprises one or more charges, for example which is monocharged, dicharged, or tricharged.
[0173] With reference to [Fig.2], there is also illustrated the fragment ions (or a set of fragment ions) which can be calculated from the amino acid sequence of the predetermined peptide. Each fragment ion (or fragment or peptide fragment ion) of the fragment ions (Ions_Frag) can be associated with a respective mass / charge ratio (R_ion_frag).
[0174] For example, the predetermined peptide PP1 may comprise a peptide sequence (Seq-pep) A* - B, where A* is the end acid labeled with the end tag *, and the fragment ions of the sequence (or set of fragment ions) A*, B as well as other fragments calculated (or obtained) from the sequence A* - B (by a fragmentation method). The fragment ions (or set of fragment ions) calculated for the predetermined peptide PP2 may comprise the fragment ion B*, the fragment ion A and other fragments of the sequence B* - A, the fragment ion B* comprises the end tag *.
[0175] According to one or more examples, the fragmentation method may be identical to the fragmentation method of the mass spectrometer or a mass spectrometer with tandem analysis capability (MS / MS) (or liquid chromatography-tandem mass spectrometry system).
[0176] Indeed, the fragment ions in the predetermined peptide library and which are obtained (or calculated) from a respective peptide sequence may be equivalent to a peptide (or parent ion) which would be fragmented by tandem mass spectrometry to give measured fragment ions.
[0177] According to one example, the fragmentation methods can be: - Collisional Activation Dissociation (CID), - Electron Transfer Induced Dissociation (ETD), - Electron Capture Induced Dissociation (ECD), - Electron Impact Induced Fragmentation (EID), - Infrared Multiphoton Absorption Induced Dissociation (IRMPD), - UV Induced Photodissociation (UVPD).
[0178] Further, according to one or more examples, fragment ions can be calculated from a predetermined peptide by applying a plurality of fragmentation methods.
[0179] For example, each fragmentation method may be specific to a type of mass spectrometer (or a tandem mass spectrometry system) using a specific fragmentation method. According to one example, in the library of predetermined peptides, a peptide sequence may comprise fragments (or fragment ions), for example grouped as sets of fragment ions, a first set of fragment ions being obtained according to a collisional activation dissociation (CID) fragmentation method, and a second set of fragment ions may be obtained according to an electron transfer dissociation method (see previous example list).
[0180] Thus advantageously, the same library can be used to identify peptides in a sample which would have been measured by different mass spectrometers (or tandem mass spectrometry systems) respectively using a specific fragmentation method or which can use different fragmentation methods (see previous list).
[0181] With reference to [Fig.l], the method may comprise a determination 110 of a selection of predetermined peptides from the library and said at least one MS / MS spectrum.
[0182] Each predetermined peptide of the selection may have a total mass / charge ratio equal to the total mass / charge ratio of the peptide measured within an error range less than a predetermined value.
[0183] Thus, for example, determining the selection for a measured peptide may include comparing the total mass-to-charge ratio of the measured peptide, i.e., the mass-to-charge ratio of the parent ion in the MS / MS spectrum of the measured peptide, with each predetermined total peptide mass-to-charge ratio present in the library, and selecting a predetermined peptide from the library whenever the total mass-to-charge ratio present in the library is equal to (or close to) the total mass-to-charge ratio (mass-to-charge ratio of the parent ion) of the measured peptide within an error range less than a predetermined value.
[0184] According to one example, and with reference to [Fig.2], if the mass / charge ratio of the measured peptide is equal to or close to the ml ratio, then the selection (shaded part in [Fig.2]) will include the predetermined peptide PP1, but also the predetermined peptide PP2, since the m2 ratio is equal to the ml ratio. Indeed, the predetermined peptides PP1 and PP2 include the same amino acids as well as the same number of amino acids and the same end marker.
[0185] According to one or more examples, the predetermined value of the error range may be less than 20 ppm (parts per million), preferably less than 15 ppm, preferably less than 10 ppm, and even more preferably less than 5 ppm.
[0186] Once a selection is determined for a measured peptide, the method may include determining 120 the labeled end amino acid of the measured peptide by comparing the measured fragment ion mass / charge ratios with the calculated fragment ion mass / charge ratios of the predetermined peptides of the selection (i.e., fragment ions of the shaded portion in the example of [Fig. 2]) within an error range less than a predetermined value.
[0187] According to one or more examples, the comparison may concern the mass / charge ratios of the measured fragment ions having the end tag (detected in the MS / MS spectrum of the measured peptide) which are compared with the mass / charge ratios of the fragment ions of the selection, and / or with the mass / charge ratios of the fragment ions of the selection having the end tag.
[0188] Thus, determining the labeled end amino acid, prior to any reconstruction of the sequence, can make it possible to determine the end amino acid of the sequence which can then be used as a starting point for reconstructing the peptide sequence of the measured peptide.
[0189] The method may then comprise 130 determining at least one reconstructed peptide sequence for the measured peptide from the calculated fragment ions of the selection (i.e. fragment ions of the shaded portion in the example of [Fig.2]), the determined labeled end amino acid and the measured fragment ions. The method may then comprise rendering 140 an identification list of peptides measured in said sample in which each measured peptide in the list is defined by the reconstructed peptide sequence.
[0190] According to one or more examples, each measured peptide in the list may be associated, by an identification means, with the information set, for example, the MS / MS spectrum, used to determine the reconstructed peptide sequence of the measured peptide. For example, the identification means may correspond to a numbering of each information set, and each measured peptide in the list may be associated with a respective number that corresponds to the number of the information set used to determine the reconstructed peptide sequence of the measured peptide.
[0191] Note that when a plurality of libraries are used, for example two libraries as described above, a first library can be used to determine the selection, and the second library can be used to reconstruct the sequence, the predetermined peptides of the first library (and therefore of the selection) can also be present in the second library.
[0192] According to one or more examples, preferably, it is also possible to use a first library as described previously, and calculate a second library, as described previously, which contains only the predetermined peptides of the selection (as well as the calculated fragment ions and the mass / charge ratios of the calculated fragment ions).
[0193] According to one or more embodiments, the peptide sequence (i.e. at least one peptide sequence) of a measured peptide can thus be reconstructed from: - the determined labeled end amino acid, - comparisons between the mass / charge ratios of measured fragment ions and calculated fragment ions.
[0194] When reconstructing a peptide sequence for a measured peptide, i.e., after determining the labeled end amino acid (starting point), comparisons are made between the measured fragment ions and the calculated fragment ions (from the selection), in order to determine the reconstructed peptide sequence of a measured peptide. For example, the comparisons may involve the different mass-to-charge ratios of the available fragment ions, i.e., the measured and / or calculated fragment ions.
[0195] In one or more embodiments, the peptide sequence of each measured peptide may be reconstructed using a method derived from the methods used in so-called "de novo" sequencing. Briefly, "de novo" sequence reconstruction relies on assembling sequences of overlapping sequence fragments to deduce complete sequences. The reconstruction method may be based on the following articles: - Biemann, K. (1992). Mass spectrometry of peptides and proteins. Annual review of biochemistry, 61(1), 977-1010, DOI: 10.1146 / annurev.bi.6L070192.004553. PMID: 1497328. - Wysocki, VH, Resing, KA, Zhang, Q., & Cheng, G. (2005). Mass spectrometry of peptides and proteins. Methods, 35(3), 211-222. DOI: 10.1016 / j.ymeth.2004.08.013. PMID: 15722218. - Seidler J, Zinn N, Boehm ME, Lehmann WD. De novo sequencing of peptides by MS / MS. Proteomics. 2010 Feb ;10(4):634-49. doi:10.1002 / pmic.200900459. PMID: 19953542.
[0196] The peptide sequence of a measured peptide can be reconstructed using "de novo" sequencing principles from the MS / MS spectrum(s) information of the measured peptide, the determined labeled end amino acid, and the fragment ions of the predetermined peptides of the selection.
[0197] In the process of measuring a peptide by mass spectrometry, it may happen that the collected data include irrelevant elements. These spurious data may be due to artifacts generated by the mass spectrometer or by the simultaneous presence of another peptide (different from the peptide to be measured). This latter situation may arise, in particular, when two peptides with similar retention times co-elute in the spectrometer.
[0198] [Fig.3] schematically illustrates this phenomenon of co-elution of two peptides.
[0199] Note that the co-elution may comprise more than two peptides, for example three or four peptides, or even more, which may have different peptide sequences, and / or may have identical peptide sequences but are differently charged.
[0200] As shown in [Fig.3], each retention time RT1 and RT2 can correspond to the retention time of a peptide of the sample having migrated in the chromatographic column to the mass spectrometer (or tandem mass spectrometry system). As can be seen in [Fig.3], in the case of co-elution, the chromatographic peaks of the two peptides can overlap for example, partially or totally, because of their retention times which are very close. In such a situation, the tandem mass spectrometer can then record a set of information, that is to say a combined spectrum, also called chimeric, of the co-eluted peptides.
[0201] Such a situation can make it difficult to determine the peptide sequence truly corresponding to the measured peptide, since a plurality of peptide sequences can be included in the information set (eg MS / MS spectrum(s)) of the measured peptide, and each peptide sequence can potentially correspond to the peptide sequence of the measured peptide.
[0202] Thus, in one or more embodiments, when a plurality of reconstructed peptide sequences is determined for the same measured peptide, for example in the presence of parasitic data due to co-elution of one or more peptides, the method may further comprise:
[0203] - calculate, for each reconstructed peptide sequence, a score from the set of information associated with the measured peptide and the reconstructed peptide sequence,
[0204] - select the reconstructed peptide sequence for the measured peptide from from a comparison of the scores of the reconstructed peptide sequences, the peptide sequence with the highest score being selected
[0205] The reconstructed peptide sequence with the highest score can then truly correspond to the measured peptide, and can be rendered (provided) in the identification list with the score obtained. Only the peptide sequence with the highest score is retained in the identification list.
[0206] However, the reconstructed sequence(s) for a measured peptide having presented a lower score, and which correspond to one or more eluted peptides, may also correspond to peptides in the mixture requiring measurement and identification, for example because they also have peptide sequences with a number of amino acids between 2 and 6 amino acids. These eluted peptides (for example of the same peptide sequence but charged differently and / or of different peptide sequences) with a measured peptide may therefore also correspond to measured peptides, and may generally be determined before or after depending on the time resolution of the mass spectrometer (or tandem mass spectrometry system).
[0207] For example, with reference to [Fig.3], the peptide sequence of the measured peptide PP1 which elutes at retention time RT1 (close to retention time RT2), and the peptide sequence which elutes at retention time RT2 can be determined in the same two retention times. In this case, i.e. in the presence of co-elution, the reconstructed sequence with the highest score will correspond to the peptide PP1 at retention time RT1, while it will correspond to PP2 at retention time RT2. Both peptide sequences will be included in the identification list with the same retention time.
[0208] During a co-elution, which can result in the presence of a plurality of measured peptides from the identification list having close or equal retention times, an alert can be notified, e.g. to a user or registered in the identification list, and a post-processing (e.g. a verification of the results) of the identification list can be carried out. For example, the data for a retention time of a measured peptide obtained from the different data sources (LC chromatogram, MS spectrum etc.) can be compared to determine that the retention time corresponds to that measured peptide (e.g. using an extracted ion chromatogram, XIC, with MS detection and comparing peak intensities and / or retention times in the MS and MS / MS spectra). As another example, when there are several peptides of the same sequence but charged differently, post-processing may consist of keeping only the peptide from the list with the highest intensity, for example.
[0209] In one or more embodiments, each mass / charge ratio of the measured fragment ions can be respectively associated with an intensity in the MS / MS spectrum, and the score of each reconstructed peptide sequence for the same measured peptide can be determined from the following equation: ^jthe measured fragment ion intensities of the reconstructed peptide sequence kl^the measured fragment ion intensities of the plurality of reconstructed peptide sequences
[0210] By intensity of a fragment ion, it can also be understood the intensity of the peak of the fragment ion.
[0211] According to one or more examples, only fragment ions having the end tag can be used to calculate the score of a peptide sequence.
[0212] Thus, for example, with reference to [Fig.2], if the fragment ions of the MS / MS spectrum of a measured peptide correspond to the predetermined peptides PP1 and PP2 of the library (i.e. of the selection, grayed-out part in the example of [Fig.2]), only the fragments A* - A*B and B* - B*A of PP1 and PP2 can be used to calculate the scores of each.
[0213] In one or more embodiments, one or a plurality of labeled end amino acids may be determined, and for each determined labeled end amino acid, at least one reconstructed peptide sequence for the measured peptide may be determined from the calculated fragment ions of the selection, the determined labeled end amino acid, and the measured fragment ions.
[0214] Indeed, as described above, in the presence of a co-elution of one or more peptides (for example of the same peptide sequence but charged differently and / or of different peptide sequences), a plurality of reconstructed peptide sequences can be determined. These reconstructed peptide sequences can have in common the same determined labeled end amino acid or have different determined labeled end amino acids. For example, in the peptide sequences A* - B - C and A* - C - B having the same mass / charge ratio (or very close), A* is the determined labeled end amino acid which is common to both sequences. According to another example, in the sequences A* - B -C and B* - A - C having the same mass / charge ratio (or very close), A* and B* are the determined labeled end amino acids which are different between the two sequences. In the latter case, when reconstructing peptide sequences, each determined labeled end amino acid is considered as a starting point for at least one peptide sequence to be reconstructed.
[0215] Furthermore, in one or more embodiments, the list may comprise retention times respectively associated with the measured peptides, each retention time being obtained from the set of information, in particular from the MS spectrum (and / or also the MS / MS spectrum) for example, used to determine a reconstructed peptide sequence of a measured peptide.
[0216] Each measured peptide in the list can be associated with a respective retention time obtained from the information set (eg MS spectrum) used to determine the reconstructed peptide sequence of the measured peptide.
[0217] According to one or more examples, the retention time for each measured peptide can be obtained following liquid chromatography (LC) coupled with mass spectrometry detection such as a tandem mass spectrometry (MS / MS) system.
[0218] As described above, the retention time can be obtained from the measurement carried out by liquid chromatography, and transcribed (or indicated) in the MS spectrum of the parent ion (and / or the MS / MS spectrum(s) from the parent ion).
[0219] According to one or more examples, the list may further comprise intensities respectively associated with the measured peptides, each intensity being obtained from the set of information (for example in the MS spectrum of the measured peptide) used to determine a reconstructed peptide sequence of a measured peptide. Each measured peptide of the list may comprise a proportion value of the peptide in the mixture, the proportion value "Qp", for each measured peptide of the list, may be calculated from: Intensity of a measured peptide from list JLdes intensities of measured peptides from the list
[0220] Advantageously, it is thus possible to obtain values of relative proportions of the peptides of the mixture.
[0221] [Fig.4] illustrates a device for implementing the method of the present disclosure.
[0222] In this embodiment, the device 400 may include a circuit 403 and a memory 402 for storing program instructions loadable into the circuit, and adapted to cause the circuit 403 to execute the method of the present disclosure when the program instructions are managed by the circuit 403.
[0223] The memory 402 may also store data and information useful for carrying out the method of the present disclosure as described above.
[0224] The circuit 403 can be for example: - a processor or processing unit capable of interpreting instructions in a computer language, the processor or processing unit may comprise, be associated with or be attached to a memory comprising the instructions, or - the combination of a processor / processing unit and a memory, the processor or processing unit being adapted to interpret instructions in a computer language, the memory comprising said instructions, or - an electronic card in which the process sequence is described in silicon, or - a programmable electronic chip such as an FPGA chip (for “Field-Programmable Gate Array”),
[0225] - a graphics processor, or GPU (from the English Graphics Processing Unit).
[0226] This device may comprise an input interface 405 for receiving input data and an output interface 407 for providing a set of useful data and / or controlling a characterization system for example, such as a tandem mass spectrometry system coupled to liquid chromatography.
[0227] For example, the input interface 405 may receive input data such as a library or a plurality of libraries (as described above for example) of predetermined peptides or parameters for generating a library of predetermined peptides, or even sets of information associated respectively with peptides present in the sample measured by a mass spectrometer. In addition, optionally, the input interface 405 may be connected to a mass spectrometer such as a tandem mass spectrometry system (mass spectrometer with tandem analysis capability (MS / MS), coupled to liquid chromatography so as to allow the measurement of the sample and receive the sets of information associated respectively with peptides present in the sample directly from the mass spectrometer.
[0228] By way of example, the output interface 407 may provide, for example, a library or a plurality of libraries (for example as described previously) generated from predetermined peptides, or even an identification list of peptides measured in a sample. Furthermore, optionally, the output interface may be connected to a mass spectrometer such as a tandem mass spectrometry system coupled to liquid chromatography so as to allow the control of the mass spectrometer for example, or to provide the sets of information associated respectively with peptides present in the sample directly from the mass spectrometer, for example sending them to a third-party device such as a computer.
[0229] According to one example, the device 400 may be a computer 401 comprising the circuit 403 and the memory 402. According to another example, the device may be the computer of the mass spectrometer. The computer of the mass spectrometer (MS or LC-MS / MS) may be used, for example, to control the mass spectrometer to perform measurements on a sample, to generate or receive one or more libraries, or to implement the method of the present disclosure.
[0230] To facilitate interaction with the device 400 or the computer 401, a screen 411, a keyboard 412, and a mouse 413 may be provided and connected to the computer circuit 403.
[0231] Analysis method
[0232] The disclosure also relates to a method for analyzing peptides having between 2 and 6 amino acids from a sample comprising: i) a step of treating the sample comprising at least one step ii) of labeling one end of the peptides of the sample by bringing the sample into contact with an end marker, ii) a liquid chromatography analysis step coupled with a mass spectrometer, preferably a mass spectrometer with tandem analysis capability (MS / MS), iii) a data processing step implementing a method for non-targeted identification of peptide sequences as described in this document.
[0233] The peptides, amino acids, end tags, and samples analyzed are as described in the aspect above.
[0234] In one or more embodiments, the peptides are natural and / or synthetic peptides.
[0235] In one or more embodiments, the analyzed peptides have between 2 and 4 amino acids.
[0236] In one or more embodiments, the peptides subjected to step ii) contain at least one labeled end amino acid and one or more amino acids selected from proteinogenic amino acids and modified amino acids.
[0237] In one or more embodiments, the sample is a mixture of small natural peptides or a mixture of peptides derived from peptide synthesis. The sample is preferably a protein hydrolyzate, for example a yeast protein hydrolyzate, a plant protein hydrolyzate, or an animal protein hydrolyzate.
[0238] In one or more embodiments, step i) of treating the sample further comprises a step i2) of alkylating the peptides of the sample by contacting with an alkylating agent, and step i2) preferably preceding step ii).
[0239] This optional alkylation step is carried out to protect amino acids carrying a free thiol group (-SH), such as cysteine, from undesirable chemical reactions which may occur during subsequent steps of the method such as oxidations.
[0240] The person skilled in the art knows how to select compatible liquid chromatography techniques for coupling with mass spectrometry.
[0241] In some embodiments, the liquid chromatography is reversed-phase liquid chromatography, preferably performed with a gradient of at least one hour.
[0242] In some embodiments, the liquid chromatography (LC) implemented is nanoscale liquid chromatography (Nano LC) or microscale liquid chromatography (Micro LC).
[0243] Nano LC is a high-performance liquid chromatography (HPLC) technique used to separate and analyze samples with flow rates in the hundreds of nanoLiters / min (nL / min). It shares similar principles with conventional liquid chromatography, but operates with much lower flow rates and sample volumes, typically in the nanoLiters range.
[0244] Micro LC, or microscale liquid chromatography or capillary LC, is a liquid chromatography (LC) technique that uses smaller columns and reduced flow rates of the order of hundreds of microLiters / min (pL / min) compared to conventional liquid chromatography (HPLC). Micro LC lies between conventional HPLC and nano LC in terms of flow rates and column size.
[0245] In some embodiments, the mass spectrometer is equipped with a mass analyzer preferably selected from a quadrupole trap mass analyzer and a triple quadrupole mass analyzer, more preferably is equipped with a high-resolution mass analyzer selected from an orbital trap mass analyzer, a time-of-flight (Q-TOF) mass analyzer, a Trapped Ion Mobility Spectrometry (TIMS-TOF) mass analyzer and a Fourier Transform Cyclotron Resonance (FT-ICR) mass analyzer.
[0246] The present disclosure is illustrated in a non-limiting manner by the following examples. Examples
[0247] Example 1: Method for detecting short peptides in protein hydrolysates.
[0248] Objectives
[0249] In this example, we seek to characterize the short peptides present in a yeast extract obtained by hydrolysis of a yeast cream. During hydrolysis, the proteins of the yeast cream are fragmented into peptides which can be decomposed as follows: - oligopeptides comprising between 5 and 35 amino acids, - short peptides having between 2 and 5 amino acids, - free amino acids.
[0250] Short peptides and free amino acids form the majority components. Proteins and oligopeptides can be analyzed by conventional proteomics and peptidomics analyses. In a first step, the peptide mixture is decomplexified using chromatographic techniques (liquid chromatography in particular) or by capillary electrophoresis, detected and then subjected to fragmentation in the mass spectrometer in order to obtain the amino acid sequence information. The data are typically obtained by analysis by liquid chromatography coupled with tandem mass spectrometry (LC-MS / MS). The MS and MS / MS data are then exploited using software for the identification of peptides by searching by homology of masses (m / z) and sequences in protein data catalogs.However, these techniques are not reliable for small peptides, which is why most software does not allow the analysis of peptides comprising less than 5 or 6 amino acids.
[0251] A method is proposed here for analyzing the protein content of hydrolyzed yeast extract. We have identified that short peptides, which are among the major components of yeast extracts, present analytical difficulties by LC-MS / MS related to their physicochemical properties whether in terms of separation and recovery or in detection by mass spectrometry (short peptides do not fit into the spectrometer detection windows). In addition, their small size reduces the reliability of their identification by sequence homology. We have developed a systematic and non-targeted method allowing the identification of a large number of short peptides in complex mixtures such as food matrices, protein hydrolysates.
[0252] The method developed is based on two components: the implementation of an analytical method comprising a treatment of the sample aimed at labeling the short peptides by dansylation, and an instrumental analysis by LC-MS / MS, followed by a data processing aimed at identifying the sequence of the short dansylated peptides in the sample from the data set generating the short peptides dansylated on their AMcrminalc part. The dansylation step is preferably preceded by an alkylation step because without this step, we found that it was difficult to detect amino acids comprising free thiol groups sensitive to oxidation, in particular cysteine. This alkylation step protects the free thiol group and makes it easier to detect after labeling. The data processing step is based on the use of a suitable and calculated catalog of short peptides, which does not correspond to a catalog comprising a set of peptides resulting from the digestion of endogenous or unlabeled proteins (directly accessible online or generated by in silico digestion from proteins contained in a database) as is conventionally the case, but to a catalog comprising all possible combinations of natural and modified amino acids leading to dipeptides, tripeptides, and tetrapeptides, dansylated at their A-terminus, i.e. a catalog comprising approximately 6,250,000 predetermined peptides for N=50.
[0253] Materials and methods
[0254] Sample processing
[0255] Peptides in a given sample, typically a yeast protein hydrolysate, are alkylated and then dansylated. Alkylation
[0256] The peptides are alkylated to protect the peptides comprising free thiol groups.
[0257] The protocol is as follows: The sample is incubated for 1 h at room temperature and in the dark with 5 mM iodoacetamide in 100 mM ammonium bicarbonate (ABC) at pH 8.8.
[0258] Labeling of the N-terminus with dansyl chloride
[0259] The protocol is as follows: . Suspend 1 mg of the sample in 1 mL of 100 mM ammonium bicarbonate (ABC) buffer, pH 8.8. . Suspend 1 mg of the sample in 1 mL of 100 mM ammonium bicarbonate (ABC) buffer, pH 8.8. . Take a 25 pL aliquot of the suspended sample and mix with 12.5 pL of acetonitrile (ACN) and 12.5 pL of 100 mM ABC buffer pH 8.8. . Vortex the mixture. . Add 25 μL of a solution of 18 mg / mL of dansyl chloride solubilized in ACN and homogenize by vortexing. . Incubate for 1 hour at 40°C. Stop the reaction by adding 5 µL of 250 mM NaOH (to remove excess dansyl chloride) (quench step) and homogenize by vortexing. . Heat the mixture for 10 minutes at 40°C. . Add 25 pL of 425 mM formic acid (FA) (diluted in a 50 / 50 ACN / water solution) and homogenize by vortexing. . Evaporate in SpeedVac. . Reconstitute the dry product in 100 pL of water. . Filter the sample onto a 96-well plate with a PVDF membrane (pore size 0.45 pm) (after wetting the wells with 100 μL of 70% EtOH and washing them twice with 200 μL of milliQ water), in order to remove aggregates. Do not dry. . Start the instrumental analysis by LC-MS / MS by diluting the filtrate 1 / 10 in water + 0.1% AF and injecting 1 pL of the dilution into the instrument.
[0260] Instrumental analysis by LC-MS / MS
[0261] The dansylated peptides are analyzed by liquid chromatography coupled with tandem mass spectrometry. The steps of this analysis are: a. Separation of the dansylated peptides by liquid chromatography b. Detection of dansylated (parent) peptide ions by MS (parent ion spectrum) c. Fragmentation in MS / MS (fragment ion spectrum from a parent ion). Liquid chromatography
[0262] The liquid chromatography performed is nanoscale liquid chromatography (nano-LC) with a C18 column (Acclaim™ PepMap™ 100 C18) conventionally used in proteomics.
[0263] The gradient is more than one hour, for example two hours, to improve the resolution of the chromatography. Detection
[0264] Detection is performed by tandem mass spectrometry. Precursor ions are detected by MS and fragmented by MS / MS. A high-resolution orbital trap mass analyzer in positive ion mode is used.
[0265] The detector makes it possible to obtain sets of information for each measured peptide, such as the sets of information described previously.
[0266] Data processing
[0267] The method for non-targeted identification of peptide sequences previously presented can be implemented, for example by a device such as presented in [Fig.4], using the sets of information obtained above.
[0268] Results
[0269] The protein content of short peptides within the sample was characterized. The sequences of all short peptides present in the sample (between 350 and 500 short peptides of different sequences) were identified. The implementation of the method allows the characterization of short peptides in complex mixtures, and their effects in terms of technological and organoleptic properties of the hydrolysates or food matrices in which they are present.
[0270] Expressions such as "comprise", "include", "incorporate", "contain", "be" and "have" are to be interpreted in a non-exclusive manner when construing the description and its associated claims.
[0271] The method is not limited to the examples of embodiments described above, only by way of example, but it encompasses all the variants that a person skilled in the art may envisage within the framework of the claims below.
[0272] Although described through a number of detailed exemplary embodiments, the proposed method and apparatus for implementing an embodiment of the method include various variations, modifications, and improvements that will be apparent to those skilled in the art, it being understood that these various variations, modifications, and improvements are within the scope of the present disclosure, as defined by the following claims. In addition, different aspects and features described above may be implemented together, or separately, or substituted for each other, and all different combinations and subcombinations of the aspects and features are within the scope of the present disclosure. Furthermore, some systems and equipment described above may not incorporate all of the modules and functions described for the preferred embodiments.
Claims
1. Claims Method for non-targeted identification of peptide sequences of peptides comprising between 2 and 6 amino acids present in a sample, the method being implemented by a computer comprising a circuit and a memory, the memory comprising: - at least one library of predetermined peptides comprising between 2 and 6 amino acids, each predetermined peptide is defined by a peptide sequence, the library comprising for each predetermined peptide: • the amino acids constituting the peptide sequence with at least one end amino acid labeled with an end marker, and a total mass / charge ratio (m / z) for the peptide sequence, • fragment ions calculated from the peptide sequence, said calculated fragment ions being associated respectively with mass / charge ratios, at least one or a plurality of calculated fragment ions comprise the end marker, - sets of information associated respectively with peptides present in the sample, called measured peptides, measured by a mass spectrometer, each measured peptide comprises a peptide sequence having at least one end amino acid labeled with an end marker, each set of information of a measured peptide comprises at least one MS / MS spectrum comprising a total mass / charge ratio of the measured peptide and one or a plurality of mass / charge ratios of the measured fragment ions, the measured peptides and the predetermined peptides being labeled at the same end by the same end marker, for each set of information of a measured peptide, the method comprises: / a / a determination (110) of a selection of predetermined peptides from the library and said at least one MS / MS spectrum, each predetermined peptide of the selection having a total mass / charge ratio equal to the total mass / charge ratio of the peptide measured according to an error range less than a predetermined value, / b / a determination (120) of the labeled end amino acid of the measured peptide by comparing the mass / charge ratios of the measured fragment ions with the mass / charge ratios of the calculated fragment ions of the predetermined peptides of the selection according to an error range less than a predetermined value, / c / a determination (130) of at least one reconstructed peptide sequence for the measured peptide from the calculated fragment ions of the selection, the determined labeled end amino acid and the measured fragment ions, - a rendering (140) of an identification list of peptides measured in said sample in which each measured peptide of the list is defined by a reconstructed peptide sequence.
2. Method according to the preceding claim, the calculated fragment ions are calculated from the peptide sequence by applying a fragmentation method to the peptide sequence, preferably the fragmentation method is identical to the fragmentation method of the mass spectrometer.
3. A method according to any preceding claim, wherein the library of predetermined peptides comprises all possible peptide sequences of 2 to 6 amino acids obtained from a group of N amino acids, each sequence of the predetermined library defines a predetermined peptide of the library and further comprises a terminal amino acid which is labeled with the terminal label, preferably the library comprises at least one predetermined peptide which is declined into a plurality of charged predetermined peptides, each charged predetermined peptide having the same peptide sequence as said at least one predetermined peptide and a respective charge.
4. A method according to any preceding claim, wherein the labeled end of the labeled end amino acid is AMcrminalc and / or C-terminal, preferably AMcrminalc.
5. A method according to any one of claims, wherein the total mass / charge ratio (m / z) of a predetermined peptide is determined from the mass / charge ratios of the amino acid residues constituting the sequence of the predetermined peptide.
6. A method according to any preceding claim, wherein the mass spectrometer is a mass spectrometer with tandem analysis capability (MS / MS), preferably coupled to liquid chromatography (LC).
7. Method according to any one of the preceding claims, wherein in / c / , said at least one peptide sequence of a measured peptide is reconstructed from: - the determined labeled end amino acid, - comparisons between the mass / charge ratios of the measured fragment ions and the calculated fragment ions.
8. A method according to any preceding claim, wherein the memory comprises at least a first library and a second library, the first library comprising the amino acids constituting the peptide sequence with at least one end amino acid labeled with the end tag, and the total mass / charge ratio (m / z) for the peptide sequence, and the second library being determined only for the predetermined peptides of the selection, the second library comprising at least the predetermined peptides of the selection, and comprising fragment ions calculated respectively from the peptide sequences of the predetermined peptides of the selection, the calculated fragment ions being associated respectively with mass / charge ratios, at least one or a plurality of calculated fragment ions comprise the end tag.
9. Method according to the preceding claim in combination with one of claims 2 and 3, in which the second library is determined from at least one fragmentation method applied to the predetermined peptides of the selection.
10. Method according to any one of the preceding claims, wherein when a plurality of reconstructed peptide sequences is determined for the same peptide measured in / c / , the method further comprises: / d / calculating, for each reconstructed peptide sequence, a score from the set of information associated with the measured peptide and the reconstructed peptide sequence, / d / selecting the reconstructed peptide sequence for the measured peptide from a comparison of the scores of the sequences reconstructed peptide sequences, with the peptide sequence with the highest score being selected.
11. Method according to the preceding claim, each mass / charge ratio of the measured fragment ions is respectively associated with an intensity in the MS / MS spectrum, and in which the score of each reconstructed peptide sequence for the same measured peptide is determined from the following equation: £of the measured fragment ion intensities of the reconstructed peptide sequence ^of the measured fragment ion intensities of the plurality of reconstructed peptide sequences preferably, only the measured fragment ions having said end marker are used to calculate a score of a peptide sequence.
12. A method according to any preceding claim, wherein in / b / one or a plurality of labeled end amino acids is determined, and wherein in / c / , for each determined labeled end amino acid, at least one reconstructed peptide sequence for the measured peptide is determined from the calculated fragment ions of the selection, the determined labeled end amino acid and the measured fragment ions.
13. A method according to any preceding claim, wherein the list comprises retention times respectively associated with the measured peptides, each retention time being obtained from the information set used to determine a reconstructed peptide sequence of a measured peptide.
14. A method according to any preceding claim, wherein each measured peptide in the list is associated, by an identification means, with the set of information used to determine the reconstructed peptide sequence of the measured peptide.
15. A method according to any preceding claim, wherein the list comprises intensities respectively associated with the measured peptides, each intensity being obtained from the information set used to determine a reconstructed peptide sequence of a measured peptide, and each measured peptide in the list comprises a proportion value of the peptide in the mixture, the proportion value "Qp", for each measured peptide in the list, is calculated from: Intensity of a measured peptide in list £Xes intensities of the measured peptides in the list
16. A method according to any preceding claim, wherein the peptides are natural peptides and / or synthetic peptides.
17. A method according to any preceding claim, wherein the peptides contain at least one labeled end amino acid and one or more amino acids selected from proteinogenic amino acids or modified amino acids.
18. A method according to any preceding claim, wherein the sample is a protein hydrolyzate, preferably a yeast protein hydrolyzate, a vegetable protein hydrolyzate, or an animal protein hydrolyzate.
19. A method according to any preceding claim, wherein the labeled end amino acid is an N-terminal amino acid labeled with an N-terminal end label selected from dansyl, dabsyl, 1-naphthyl isocyanate (IsoC), or 1-naphthyl isothiocyanate, each of which end labels may optionally be labeled with one or more carbon 13 (13C), deuterium (D), nitrogen 15 (15N), and / or sulfur 34 (34S) atoms.
20. A method according to any preceding claim, wherein the labeled end amino acid is a C-terminal amino acid labeled with a C-terminal label selected from 2-(N,N'-dimethylamino)-1-ethylamine (DMED) and 3-(MA'-dimethylamino)-1-propylamine (DMAPA), each of these end labels optionally being labeled with one or more carbon 13 (13C), deuterium (D), and / or nitrogen 15 (15N) atoms.
21. A method according to any preceding claim, wherein the peptides have been subjected to a protective treatment of the free thiol group, typically an alkylation treatment, the mass / charge ratios of the predetermined peptides of the library taking into account the protective treatment of the free thiol group of cysteine.
22. A method of analyzing peptides having between 2 and 6 amino acids from a sample comprising: i) a step of processing the sample comprising at least one step ii) of labeling one end of the peptides of the sample by bringing the sample into contact with an end marker, ii) a step of analysis by liquid chromatography coupled to a mass spectrometer, preferably to a mass spectrometer with a tandem analysis capacity (MS / MS), iii) a step of data processing implementing a method of non-targeted identification of peptide sequences as defined in any one of the preceding claims.
23. Method according to the preceding claim, in which step i) of treating the sample further comprises a step i2) of alkylation of the peptides of the sample by contacting with an alkylating agent, and in which step i2) preferably precedes step ii).
24. A method according to any one of claims 22 or 23, wherein the liquid chromatography is reverse phase liquid chromatography, preferably carried out with a gradient of at least one hour.
25. A method according to any one of claims 23 to 24, wherein the mass spectrometer is equipped with a mass analyzer preferably selected from a quadrupole trap mass analyzer or a triple quadrupole mass analyzer, more preferably is equipped with a high-resolution mass analyzer selected from an orbital trap mass analyzer, a Q-TOF mass analyzer, a TIMS-TOF mass analyzer, and an FT-ICR mass analyzer.
26. Computer program comprising instructions for implementing the method according to any one of claims 1 to 21 when this program is executed by a processor.
27. A non-transitory recording medium readable by a computer on which is recorded a program for implementing the method according to any one of claims 1 to 21 when this program is executed by a processor.
Citation Information
Patent Citations
Compounds and methods for double labelling of polypeptides to allow multiplexing in mass spectrometric analysis
CN101542291A
Method and apparatus for mass spectrometry of biomolecular samples with data independent acquisition
CN114965728A
Method and apparatus for analysing samples of biomolecules using mass spectrometry with data-independent acquisition
EP4047371A1
Method for identifying and characterizing a microbial population by mass spectrometry
FR3106414A1
Protein expression profile database
US20050048564A1