Method for non-targeted identification of peptide sequences

A method using a device with a peptide library and mass spectrometer for non-targeted identification of small peptides in complex samples addresses the challenge of characterizing peptides of 2 to 6 amino acids, achieving rapid and comprehensive sequence determination.

WO2025172665A1PCT designated stage Publication Date: 2025-08-21LESAFFRE & CIE
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
PCT/FR2025/050116
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-12
Filing Date
2025-02-11
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Current methods fail to efficiently characterize small peptides of 2 to 6 amino acids in complex mixtures due to their small size and diverse physicochemical properties, making separation by liquid chromatography and detection by mass spectrometry difficult.

Method used

A method utilizing a device with a circuit and memory that includes a library of predetermined peptides, labeled with end markers, and a mass spectrometer for non-targeted identification of peptides, reconstructing sequences through fragment ion comparisons within an error range, enabling the determination of peptide sequences without prior hypotheses on the sample content.

Benefits of technology

Enables systematic and rapid identification of all peptides in a sample, including modified amino acids, without prior knowledge of their structure, providing a comprehensive list of peptide sequences between 2 and 6 amino acids.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FR2025050116_21082025_PF_FP_ABST
    Figure FR2025050116_21082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure describes a method for the non-targeted identification of peptide sequences from peptides comprising between 2 and 6 amino acids present in a sample, the method being implemented by a device comprising a circuit and a memory, wherein the memory comprises at least one library of predetermined peptides comprising between 2 and 6 amino acids, wherein each predetermined peptide is defined by a peptide sequence, and wherein sets of information are respectively associated with peptides present in the sample, which peptides are referred to as measured peptides, the measured peptides and the predetermined peptides being labelled at the same end with the same end marker, and a method for analysing peptides comprising between 2 and 6 amino acids in a sample, the method comprising a data processing step that implements the method for the non-targeted identification of peptide sequences.
Need to check novelty before this filing date? Find Prior Art

Description

Description Title: Method for non-targeted identification of peptide sequences Technical field

[0001] The present invention relates to a method for the non-targeted identification of peptide sequences of peptides comprising between 2 and 6 amino acids present in a sample, the method being implemented by a device comprising a circuit and a memory, as well as a method for analyzing the peptides of a sample. The present invention also relates to a method for analyzing peptides having between 2 and 6 amino acids of a sample implementing the identification method of the invention. Prior art

[0002] Given their physicochemical properties and potential activities, small peptides derived from proteins are of great interest to the food, pharmaceutical, and cosmetic industries. However, enzymatic hydrolysis of proteins leads to complex mixtures of peptides of varying sizes and properties.

[0003] Analysis of the overall peptide / protein content and free amino acid analysis of a sample is well established. Protein analysis by proteomics is also known.

[0004] However, small peptides are not covered by classical proteomic or peptide analysis methods, and a finer characterization of small peptides in complex mixtures remains a challenge due to their small size and extremely different physicochemical characteristics, which make their separation by liquid chromatography (LC) and their detection by mass spectrometry (MS) difficult.

[0005] Typically, the Applicant company develops yeast extracts, food matrices obtained by hydrolysis of yeast cream. During the hydrolysis step, the yeast cream proteins are fragmented into complex mixtures comprising peptides having between 2 and 35 amino acids and free amino acids. Yeast cream proteins can be analyzed by proteomic analysis methods. There are also peptidomic analysis techniques for characterizing peptides comprising at least 5 or 6 amino acids. However, current techniques do not allow for rapid characterization of all the short peptides or small peptides in a yeast extract sample, even though these small peptides constitute a major constituent of most yeast extracts.

[0006] There is therefore a need for a method allowing systematic and non-targeted identification of all small peptides in a sample. Summary

[0007] A method for non-targeted identification of peptide sequences of peptides comprising between 2 and 6 amino acids present in a sample is provided, the method being implemented by a device comprising a circuit and a memory, the memory comprising: - at least one library of predetermined peptides comprising between 2 and 6 amino acids, each predetermined peptide is defined by a peptide sequence, the library comprising for each predetermined peptide: • the amino acids constituting the peptide sequence with at least one end amino acid labeled with an end marker, and a total mass / charge ratio (m / z) for the peptide sequence, • fragment ions calculated from the peptide sequence, said calculated fragment ions being associated respectively with mass / charge ratios, at least one or a plurality of calculated fragment ions comprise the end marker, - sets of information associated respectively with peptides present in the sample, called measured peptides, measured by a mass spectrometer, each measured peptide comprises a peptide sequence having at least one end amino acid labeled with an end marker, each set of information of a measured peptide comprises at least one MS / MS spectrum comprising a total mass / charge ratio of the measured peptide and one or a plurality of mass / charge ratios of the measured fragment ions, the measured peptides and the predetermined peptides being labeled at the same end by the same end marker, for each set of information of a measured peptide, the method comprises: / a / a determination of a selection of predetermined peptides from the library and said at least one MS / MS spectrum, each predetermined peptide of the selection having a total mass / charge ratio equal to the total mass / charge ratio of the peptide measured within an error range less than a predetermined value, / b / a determination of the labeled end amino acid of the measured peptide by comparing the mass / charge ratios of the measured fragment ions with the calculated mass / charge ratios of the fragment ions of the predetermined peptides of the selection within an error range less than a predetermined value, Herein a determination of at least one reconstructed peptide sequence for the measured peptide from the calculated fragment ions of the selection, the determined labeled end amino acid and the measured fragment ions, - a rendering of a list of identification of peptides measured in said sample in in which each measured peptide in the list is defined by a reconstructed peptide sequence.

[0008] Advantageously, it is thus possible to obtain at the output of the method an identification list which includes all the peptides in the sample having a peptide sequence having a length (or a number of amino acids) of between 2 and 6 amino acids, typically between 2 and 4 amino acids, without there having been (having been) prior a priori hypotheses on the structure and content of the different peptides contained in the sample. As will be detailed below, the peptides present in the sample may include modified amino acids, in particular modified amino acids containing modified side chains, the modification being able to be of natural or synthetic origin.

[0009] Each measured fragment ion may be representative of the presence of one or a plurality of linked amino acids in the peptide sequence of the measured peptide.

[0010] By total mass / charge ratio, it can be understood the mass / charge ratio of the entire measured peptide, i.e. including all the amino acids constituting the peptide sequence of the measured peptide labeled with an end marker.

[0011] In one or more embodiments, the library may be divided into two or more distinct libraries. More specifically, the library data may be distributed across one or more distinct libraries.

[0012] According to one or more examples, the comparison in / b / may relate to the mass / charge ratios of the measured fragment ions having the end tag (detected in the MS / MS spectrum of the measured peptide) which are compared with the mass / charge ratios of the fragment ions of the selection, and / or with the mass / charge ratios of the fragment ions of the selection having the end tag. This comparison may be carried out within an error range less than a predetermined value.

[0013] In one or more embodiments, the calculated fragment ions may be calculated from the peptide sequence by applying a fragmentation method to the peptide sequence.

[0014] According to one or more embodiments, the fragmentation method may be defined by one or a plurality of fragmentation rules, which may be included in a fragmentation rule library. The calculated fragment ions may be calculated from a predetermined peptide by applying the fragmentation rule(s) of the library.

[0015] In one or more embodiments, the fragmentation method may be the same as the fragmentation method of the mass spectrometer.

[0016] In one or more embodiments, the calculated fragment ions (or a set of calculated fragment ions) may be calculated from a predetermined peptide by applying a plurality of fragmentation methods, for example defined in the fragmentation rule library. More specifically, sets of calculated fragment ions may be calculated from a predetermined peptide to which a plurality of fragmentation methods are respectively applied, each set of calculated fragment ions being obtained from a respective fragmentation method.

[0017] Thus, advantageously, the same library can be used to identify peptide sequences in a sample which would have been measured by different mass spectrometers or the same mass spectrometer comprising a plurality of mass analyzers, preferably by mass spectrometers capable of carrying out a tandem analysis using a specific fragmentation method or capable of using different fragmentation methods.

[0018] In one or more embodiments, the predetermined peptide library comprises all possible peptide sequences of 2 to 6 amino acids obtained from a group of N amino acids, each sequence of the predetermined library defines a predetermined peptide of the library and further comprises a terminal amino acid which is labeled with the terminal tag.

[0019] By all possible peptide sequences can be understood all possible combinations forming peptide sequences of a length between 2 and 6 amino acids obtained from a group of N amino acids, and for which an end amino acid is labeled with the end marker.

[0020] In one or more embodiments, N may comprise at least 22 proteinogenic amino acids and one or a plurality of modified amino acids. In one or more embodiments, a modified amino acid may be modified from a proteinogenic amino acid. In one or more embodiments, N may be greater than 2, preferably greater than 20. In one or more embodiments, N may be less than 500, 400, 300, or 200. N may typically be between 2 and 100, for example between 2 and 50.

[0021] In one or more embodiments, the library comprises at least one predetermined peptide which is broken down into a plurality of predetermined charged peptides, each predetermined charged peptide having the same peptide sequence as said at least one predetermined peptide and a respective charge.

[0022] By predetermined charged peptide, it may be understood a predetermined peptide having one or more charges, chosen from one or more positive charges (e.g. +1, +2, or +3) or one or more negative charges (e.g. -1, -2, -3).

[0023] Thus advantageously, it is possible to determine a peptide sequence of a measured peptide which comprises one or more charges, for example which is monocharged, dicharged, or tricharged.

[0024] In one or more embodiments, the end labeled with the end tag is the same for all peptide sequences in the predetermined library.

[0025] In one or more embodiments, the labeled end of the labeled end amino acid is A / -terminal and / or C-terminal, preferably A / -terminal.

[0026] In one or more embodiments, the total mass / charge ratio (m / z) of a predetermined peptide is determined from the mass / charge ratios of the amino acid residues constituting the sequence of the predetermined peptide.

[0027] In one or more embodiments, the mass spectrometer is a mass spectrometer with tandem analysis capability (MS / MS), preferably coupled to liquid chromatography (LC).

[0028] In one or more embodiments, herein, said at least one peptide sequence of a measured peptide is reconstructed from: - the determined labeled end amino acid, - comparisons between the mass / charge ratios of measured fragment ions and calculated fragment ions.

[0029] Comparisons between measured fragment ion mass-to-charge ratios and calculated fragment ions can be made taking into account an error range less than a predetermined value.

[0030] In one or more embodiments, the memory comprises at least a first library and a second library, the first library comprising the amino acids constituting the peptide sequence with at least one terminal amino acid labeled with the terminal label, and the total mass / charge ratio (m / z) for the peptide sequence, and the second library being determined only for the predetermined peptides of the selection, the second library comprising at least the peptides predetermined peptides of the selection, and comprising fragment ions calculated respectively from the peptide sequences of the predetermined peptides of the selection, the calculated fragment ions being associated respectively with mass / charge ratios, at least one or a plurality of calculated fragment ions comprise the end marker.

[0031] Thus, advantageously, it is possible to reduce the computation time for determining a reconstructed peptide sequence.

[0032] In one or more embodiments, the second library can be determined from at least one fragmentation method applied to the predetermined peptides of the selection.

[0033] In one or more embodiments, each set of information relating to a measured peptide may also include mass spectrum information (referred to as MS spectrum), wherein the MS spectrum information may include a mass-to-charge ratio of the measured peptide (i.e., parent ion), a retention time, and an intensity of the measured peptide (i.e., intensity of the parent ion).

[0034] The retention time included in the MS spectrum (or information set) of a measured peptide can be obtained from the measurement made following separation by liquid chromatography.

[0035] In one or more embodiments, the selection may be alternatively determined from the MS spectrum information of the measured peptide and the library.

[0036] In one or more embodiments, the one or more reconstructed peptide sequences for each measured peptide may be determined using a sequencing method derived from methods used in so-called "de novo" sequencing.

[0037] The peptide sequence of a measured peptide is reconstructed using the principles of methods used in so-called "de novo" sequencing using the MS / MS spectrum information of the measured peptide, the determined labeled end amino acid, and the calculated fragment ions of the predetermined peptides in the selection.

[0038] In one or more embodiments, when a plurality of reconstructed peptide sequences is determined for the same measured peptide in Here, the method further comprises: ldi calculating, for each reconstructed peptide sequence, a score from the set of information associated with the measured peptide and the reconstructed peptide sequence, lel select the reconstructed peptide sequence for the measured peptide from a comparison of the scores of the reconstructed peptide sequences, with the peptide sequence with the highest score being selected.

[0039] Thus, advantageously, the use of score comparison in the case of co-elution of peptides makes it possible to ensure that the peptide sequence determined for a measured peptide corresponds to the measured peptide and to the set of information obtained for this measured peptide.

[0040] In one or more embodiments, each mass / charge ratio of the measured fragment ions is respectively associated with an intensity in the MS / MS spectrum, and in which the score of each reconstructed peptide sequence for the same measured peptide is determined from the following equation:

[0041] In one or more embodiments, only measured fragment ions having said end tag are used to calculate a score of a peptide sequence.

[0042] In one or more embodiments, in / b / one or a plurality of labeled end amino acids is determined, and in Here, for each determined labeled end amino acid, at least one reconstructed peptide sequence for the measured peptide is determined from the calculated fragment ions of the selection, the determined labeled end amino acid and the measured fragment ions.

[0043] In one or more embodiments, the predetermined value of the error range is less than 20 ppm, preferably less than 15 ppm, preferably less than 10 ppm, and even more preferably less than 5 ppm.

[0044] In one or more embodiments, the list may comprise retention times respectively associated with the measured peptides, each retention time being obtained from the set of information, in particular the MS spectrum for example, used to determine a reconstructed peptide sequence of a measured peptide.

[0045] Each measured peptide in the list can be associated with a respective retention time obtained from the information set (eg MS spectrum) used to determine the reconstructed peptide sequence of the measured peptide.

[0046] In one or more embodiments, each measured peptide in the list is associated, by an identification means, with the set of information used to determine the reconstructed peptide sequence of the measured peptide.

[0047] Advantageously, the use of an identification means makes it possible to quickly order the information sets and to ensure that a set of information (e.g. MS / MS spectrum(s) and / or MS spectrum) is correctly associated with the reconstructed peptide sequence for which it was used. In addition, this facilitates the post-processing of the data (e.g. identification list, information sets, etc.) obtained.

[0048] In one or more embodiments, the list comprises intensities respectively associated with the measured peptides, each intensity being obtained from the set of information (eg MS spectrum) used to determine a reconstructed peptide sequence of a measured peptide, and each measured peptide of the list comprises a proportion value of the peptide in the mixture, the proportion value "Qp", for each measured peptide of the list, is calculated from: Intensity of a measured peptide from list S of the intensities of the measured peptides from the list

[0049] Advantageously, it is thus possible to obtain values ​​of relative proportions of the peptides in the mixture.

[0050] The intensity of each measured peptide in the list may have been obtained from the MS spectrum of the measured peptide that is included in the information set.

[0051] In one or more embodiments, the peptides are natural peptides and / or synthetic peptides.

[0052] In one or more embodiments, the peptides contain at least one labeled end amino acid and one or more amino acids selected from proteinogenic amino acids or modified amino acids.

[0053] In one or more embodiments, the labeled end amino acid is an A / -terminal amino acid labeled with an A / -terminal label selected from dansyl, dabsyl, 1-naphthyl isocyanate (IsoC), or 1-naphthyl isothiocyanate, each of which end labels may optionally be labeled with one or more carbon atoms 13 ( 13 C), deuterium (D), nitrogen 15 ( 15 N), and / or sulfur 34 ( 34 S). The labeling makes it possible to make said end markers differentiable. The A / -terminal end marker can for example be chosen from dansyl, dansyl labeled with one or more carbon atoms 13 ( 13 C), deuterium-labeled dansyl (D), dabsyl, 1-naphthyl isocyanate (IsoC), or 1-naphthyl isothiocyanate.

[0054] In one or more embodiments, the labeled end amino acid is a C-terminal amino acid labeled with a C-terminal tag. selected from 2-(A / ,A / -dimethylamino)-1-ethylamine (DMED) or 3-(N,N'- dimethylamino)-1 -propylamine (DMAPA), each of these end markers optionally being labeled with one or more carbon 13 atoms ( 13 C), deuterium (D), and / or nitrogen 15 ( 15 N). The labeling makes it possible to make said end markers differentiable. The C-terminal end marker can, for example, be chosen from DMED or DMAPA.

[0055] In one or more embodiments, the peptides have been subjected to a protective treatment of amino acids containing a free thiol group such as cysteine, typically an alkylation treatment, the mass / charge ratios of the predetermined peptides of the library taking into account the protective treatment of the free thiol group of cysteine. This protective treatment is commonly carried out by carbamidomethylation typically using chloroacetamide or iodoacetamide. This protective treatment can also be carried out by other reagents known to those skilled in the art such as 4-hydroxyphenyl)ethyl iodoacetamide, iodoacetic acid and α-ethyl maleimide.

[0056] In one or more embodiments, the liquid chromatography is reversed-phase liquid chromatography, preferably operating with a gradient of at least one hour.

[0057] In one or more embodiments, the mass spectrometer is equipped with a mass analyzer preferably selected from a quadrupole trap mass analyzer and a triple quadrupole mass analyzer, more preferably is equipped with a high-resolution mass analyzer selected from an orbital trap analyzer, a Q-TOF analyzer, a TIMS-TOF analyzer and an FT-ICR analyzer.

[0058] In one or more embodiments, the peptides present in the sample are natural peptides and / or synthetic peptides.

[0059] In one or more embodiments, the sample is a mixture of small natural peptides or a mixture of peptides derived from peptide synthesis. The sample is preferably a protein hydrolyzate, for example a yeast protein hydrolyzate, a plant protein hydrolyzate, or an animal protein hydrolysate.

[0060] According to another aspect, there is provided a computer program comprising instructions for implementing the method according to the present disclosure when this program is executed by a processor.

[0061] According to another aspect, there is provided a non-transitory recording medium readable by a computer on which is recorded a program for implementing the method according to the present disclosure when this program is executed by a processor.

[0062] According to another aspect, there is provided a method for analyzing peptides having between 2 and 6 amino acids of a sample comprising: i) a step of processing the sample comprising at least one step ii) of labeling one end of the peptides of the sample by bringing the sample into contact with an end marker, ii) a step of analysis by liquid chromatography coupled to a mass spectrometer, preferably to a mass spectrometer with a tandem analysis capacity (MS / MS), iii) a data processing step implementing a method of non-targeted identification of peptide sequences as described in the present document.

[0063] In one or more embodiments, the step of treating the sample further comprises a step of alkylating the peptides of the sample by contacting with an alkylating agent, and wherein the alkylating step preferably precedes the step of treating the sample.

[0064] In one or more embodiments, the liquid chromatography is reversed-phase liquid chromatography, preferably operating with a gradient of at least one hour.

[0065] In one or more embodiments, the mass spectrometer is equipped with a mass analyzer which is preferably a quadrupole trap mass analyzer or a triple quadrupole mass analyzer, preferably a high resolution mass analyzer selected from an orbital trap analyzer, a Q-TOF analyzer, a TIMS-TOF analyzer, and an FT-ICR analyzer.

[0066] In one or more embodiments, the sample contains a protein hydrolysate, for example a yeast protein hydrolysate, a plant protein hydrolysate, or an animal protein hydrolysate. Brief description of the drawings

[0067] Other features, details and advantages will become apparent upon reading the detailed description below, and upon analyzing the attached drawings, in which: Fig. 1

[0068] [Fig. 1] illustrates a flowchart of the process for non-targeted identification of peptide sequences of peptides from a sample. Fig. 2

[0069] [Fig. 2] schematically illustrates an example of a predetermined peptide library. Fig. 3

[0070] [Fig. 3] schematically illustrates the phenomenon of co-elution of two peptides. Fig. 4

[0071] [Fig. 4] illustrates a device for implementing the method of the present disclosure.

[0072] [Fig. 5] illustrates the number of short peptide sequences identified in various protein hydrolysates derived from soybean (SH), casein (CH) and yeast (YE), with the method of the invention. The values ​​represent the mean of three replicates ± standard deviation.

[0073] [Fig. 6] illustrates a comparative qualitative analysis of short peptide sequences identified in different soy protein hydrolysates (SH, Figure 6A), casein (CH, Figure 6B) and yeast (YE, Figure 6C). Description of the embodiments

[0074] The terms "peak" and "line" may be interpreted in the same way and are interchangeable in the present disclosure. Similarly, the terms "intensity" and "amplitude" may be interpreted in the same way and are interchangeable in the present disclosure.

[0075] Figure 1 illustrates a flowchart of the process for non-targeted identification of peptide sequences of peptides from a sample.

[0076] In this document, the term "sample" means a representative portion of a larger population, selected for the purpose of carrying out observations, measurements or analyses to draw conclusions about the original population. In the context of the invention, the sample contains several peptides, which may have different amino acid sequence lengths, and for the same amino acid sequence length, different amino acid sequences.

[0077] Peptides can be natural peptides and / or synthetic peptides.

[0078] When the sample is a portion of a natural source or a portion of the product of one or more chemical or biochemical reaction(s) applied to a natural source, the peptides are referred to as "natural peptides." In one or more embodiments, the sample contains natural peptides.

[0079] When the sample is a portion of a population of peptides artificially manufactured in the laboratory by chemical or biotechnological synthesis methods, such as solid-phase peptide synthesis (SPPS), solid-phase peptide synthesis liquid or enzymatic synthesis, the peptides are referred to as "synthetic peptides" or "synthetic peptides". In one or more embodiments, the sample contains synthetic peptides.

[0080] In one or more embodiments, the sample containing natural peptides is a protein hydrolysate.

[0081] In this document, protein hydrolysate, or simply hydrolysate, is a product derived from the breakdown of proteins into smaller peptides or even amino acids, through a process called hydrolysis. Hydrolysis is a chemical reaction that involves the addition of water to break the peptide bonds between the constituent amino acids of proteins. This process of protein fragmentation leads to the formation of smaller molecules that are more easily digested and absorbed by the body. The goal of producing protein hydrolysates is often to obtain products that can be quickly assimilated by the body, as the peptides and amino acids resulting from hydrolysis are generally more soluble and easier to absorb than intact proteins.These hydrolysates are commonly used in the food industry, dietary supplements, and medical preparations, particularly for people with specific needs regarding protein digestion or absorption. They can also be used in the cosmetics industry. Protein hydrolysis can be achieved by various methods, usually using chemical agents, enzymes, or a combination of both. The main hydrolysis methods used in industry, particularly the food industry, are acid hydrolysis and enzymatic hydrolysis. The hydrolyzate can be prepared from various sources, preferably natural. The hydrolyzate can thus be a hydrolyzate of microbial proteins, plant proteins, or animal proteins.The protein hydrolyzate may, for example, be a hydrolyzate of yeast or other microorganism proteins, a whey protein hydrolyzate, a casein protein hydrolyzate, a soy protein hydrolyzate, a fish protein hydrolyzate, a collagen hydrolyzate, a wheat protein hydrolyzate, a rice protein hydrolyzate.

[0082] In one or more embodiments, the sample of synthetic peptides is a sample of peptides derived from chemical synthesis, preferably derived from solid phase synthesis (SPPS) or liquid phase synthesis.

[0083] In one or more embodiments, one or more peptides present in the sample are bioactive.

[0084] In this document, “small bioactive peptides” are defined as small peptides exhibiting specific biological activities. These peptides may be of natural or synthetic origin and are capable of modulating various biological processes when interacting with receptors, enzymes, or other target molecules in a living organism. These peptides may, for example, exhibit antimicrobial, anti-inflammatory, or antioxidant properties (such as the tripeptide glutathione).

[0085] The peptides identified by the method of the invention are advantageously short peptides (or in other words small peptides) having between 2 and 6 amino acids, typically between 2 and 5 amino acids, or between 2 and 4 amino acids. The peptides identified by the method of the invention can thus be chosen from dipeptides, tripeptides, tetrapeptides, pentapeptides, and hexapeptides, and mixtures thereof.

[0086] In one or more embodiments, the sample contains short peptides having between 2 and 6 amino acids, at least one end of which is labeled with an end tag.

[0087] Preferably, the sample contains a large number of short peptides having different sequences, preferably comprises at least 100 short peptides having different sequences, at least 200 short peptides having different sequences, at least 300 short peptides having different sequences, or at least 500 short peptides having different sequences.

[0088] In some embodiments, the sample contains at most 10,000 short peptides having different sequences, at most 5,000 short peptides having different sequences, at most 2,000 short peptides having different sequences, or at most 1,000 short peptides having different sequences.

[0089] In some embodiments, the sample contains between 100 and 2,000 short peptides having different sequences, between 100 and 1,500 short peptides having different sequences, or between 100 and 1,000 short peptides having different sequences.

[0090] In some embodiments, the sample contains between 150 and 2,000 short peptides having different sequences, between 200 and 2,000 short peptides having different sequences, between 250 and 2,000 short peptides having different sequences, between 150 and 1,000 short peptides having different sequences, between 200 and 1,000 short peptides having different sequences, or between 250 and 1,000 short peptides having different sequences.

[0091] A soy hydrolysate typically contains around 150 short peptides with different sequences.

[0092] A casein hydrolysate typically contains around 300 peptides.

[0093] A yeast hydrolysate typically contains around 500 peptides.

[0094] As detailed below, peptides present in the sample may include modified amino acids, regardless of the origin and nature of the modification. Modified amino acids may include modified side chains.

[0095] In one or more embodiments, the labeled end of the labeled end amino acid, e.g., of one or a plurality of peptides of the sample, may be the A / -terminus and / or the C-terminus. Preferably, the labeled end is the A / -terminus. Identification process

[0096] The method is preferably a non-targeted identification method.

[0097] The term "non-targeted" in relation to the identification process means that the process is carried out without any a priori hypothesis on the sequence and content of the different peptides contained in the sample. The nature and content of the different peptides contained in the sample are not known in advance. In other words, the composition of the sample in peptides and their peptide sequences is not known in advance.

[0098] The method may be implemented by a device that includes a memory and a processor.

[0099] The memory may comprise at least one library of predetermined peptides comprising between 2 and 6 amino acids, typically between 2 and 5 amino acids, or between 2 and 4 amino acids.

[0100] Each predetermined peptide can be defined by a peptide sequence, and the library can include for each predetermined peptide: • the amino acids constituting the peptide sequence with at least one end amino acid labeled with an end marker, and a total mass / charge ratio (m / z) for the peptide sequence, • fragment ions calculated from the peptide sequence, said calculated fragment ions being respectively associated with mass / charge ratios, at least one or a plurality of calculated fragment ions comprise the end marker.

[0101] The total mass-to-charge ratio (m / z) of a predetermined peptide can be determined from the amino acid sequence of the predetermined peptide. Indeed, the sum of all the mass-to-charge ratios of the amino acid residues of a peptide sequence in the library can correspond to the total mass-to-charge ratio of the predetermined peptide defined by this peptide sequence.

[0102] In one or more embodiments, the fragment ions (or set of fragment ions) can be calculated from the peptide sequence by applying a fragmentation method to the peptide sequence.

[0103] Note that, in one or more embodiments, the library may be divided into two separate libraries or a plurality of separate libraries. For example, a first library may comprise the amino acids constituting the peptide sequences with at least one end amino acid labeled with an end marker, and the total mass / charge ratios (m / z) for the peptide sequences, and a second library may comprise the predetermined peptides of the first library, the total mass / charge ratios of the predetermined peptides as well as fragment ions calculated for each predetermined peptide from their peptide sequence, the fragment ions being respectively associated with mass / charge ratios, and at least one or a plurality of calculated fragment ions comprise the end marker. This second library may thus be determined from the first library.In such a configuration, the method according to the present application may use both libraries or the plurality of libraries.

[0104] In this document, "peptide sequence" or "amino acid sequence" or "sequence of amino acid residues" refers to the description of the chain of amino acids (monomers) that make up a peptide. Peptides are in fact ordered linear polymers, made up of amino acids. The sequence is generally represented as a string of characters that is stored in a computer file in text format.

[0105] In this document, "amino acid" refers to a residue comprising a central carbon atom, the α-carbon, on which are articulated the carboxyl group (-COOH), the amino group (-NH2), a hydrogen atom (-H) and a side group (noted -R). It is the nature of this side group (also called a side chain) that differentiates amino acids from each other. The central α-carbon is asymmetric in all cases except that of glycine (because glycine has a hydrogen in place of a side chain). It is possible to have two different conformations of the same amino acid, conformations that cannot be interconverted without breaking and reassembling covalent bonds. designates these L and D conformations. Although both can be synthesized in the same chemical reaction, only the L form is used for protein synthesis (proteinogenic), while the non-proteinogenic D form is integrated into the synthesis of carbohydrates or lipids in living organisms. Amino acids can notably be proteinogenic amino acids or modified amino acids.

[0106] Proteinogenic amino acids, or "unmodified amino acids," are amino acids that are the basic building blocks of proteins. Proteinogenic amino acids are the basic building blocks of proteins. They polymerize to form linear polypeptides in which amino acid residues are joined by peptide bonds. Protein biosynthesis takes place on ribosomes, which translate messenger RNA into proteins. The order in which amino acids are linked to the polypeptide chain is specified by the sequence of codons carried by the messenger RNA sequence, which is a copy of the DNA in the cell nucleus; these codons, which are triplets of nucleotides, are translated into amino acids by transfer RNA according to the genetic code.This directly specifies 20 standard proteinogenic amino acids (A, R, N, D, C, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, V) to which two other amino acids are added through a more complex mechanism involving, for selenocysteine ​​(U), a SECIS element that recodes the UGA stop codon and, for pyrrolysine (O), a PYLIS element that recodes the UAG stop codon. There are therefore a total of 22 proteinogenic amino acids. Pyrrolysine is only present in the proteins of certain methanogenic archaea, so that eukaryotes and bacteria only use 21 proteinogenic amino acids.

[0107] The most common amino acids are represented by a 1 or 3 letter code as shown in the table below. The symbols B, Z and J are used when the amino acid is not completely determined (ambiguous amino acid). Table 1]: Amino acids - symbols, three-letter code, and definitions

[0108] In this document, the term "peptide" refers indifferently to a peptide present in the sample, in other words to a measured peptide, or to a peptide in the library, in other words to a predetermined peptide.

[0109] In one or more embodiments, the peptides contain amino acids selected from proteinogenic acids and modified amino acids.

[0110] Peptides contain at least one terminal amino acid labeled with an end tag.

[0111] In one or more embodiments, the peptides contain at least one labeled end amino acid and one or more amino acids selected from proteinogenic amino acids and modified amino acids.

[0112] In one or more embodiments, the labeled end amino acid may be a proteinogenic amino acid or a modified amino acid. For example, the labeled end amino acid may be an amino acid with a modified side chain.

[0113] A modified amino acid is any non-proteinogenic amino acid. A modified amino acid may be the product of an endogenous or exogenous chemical modification applied to a proteinogenic amino acid. A modified amino acid may also be a synthetic amino acid.

[0114] In this document, an endogenous reaction is a reaction that occurs naturally within an organism or biological system, for example, a post-translational modification type reaction, whereas an exogenous reaction is a reaction that is triggered or caused by factors external to the organism or biological system. Examples of exogenous reactions are the oxidation of peptides that occurs outside the body, under laboratory or industrial conditions (e.g. chemical or photochemical oxidation), or a protective reaction with a protecting group, for example to protect an amino acid from oxidation.

[0115] In one or more embodiments, the modified amino acid is an amino acid whose side chain, carboxyl group, or amino group is modified.

[0116] In one or more embodiments, the modified amino acid is an amino acid containing a modified side chain. The modified side chain may, for example, be the product of an endogenous chemical modification such as a post-translational modification, the product of one or more exogenous chemical or biochemical reaction(s), such as the product of a reaction with a side chain protecting group. The modified amino acid may also be the product of a reaction of replacing an atom of the side chain with an isovalent atom, for example, the replacement of a hydrogen atom with a fluorine atom, or of a sulfur atom with a selenium atom or of a nitrogen atom with a phosphorus atom.

[0117] Any chemical modification of amino acids, in particular of amino acid side chains, whatever their origin, can be detected by the method of the invention provided that it is implemented in the library.

[0118] Examples of modification of the side chains of cysteine, lysine, and methionine are detailed below.

[0119] Cysteine ​​is an amino acid that contains a thiol group (-SH), making it reactive. It may be necessary to protect cysteine ​​from oxidation or other unwanted reactions. Protecting cysteine ​​usually involves forming a derivative that makes it less reactive.

[0120] In some embodiments, the amino acid containing a modified side chain is a cysteine ​​whose side chain is modified by a chemical reaction to render it non-reactive. For example, the cysteine ​​may be carbamidomethylated. Carbadomethylation of the cysteine ​​may be accomplished by contacting the sample with an alkylating agent, such as iodoacetamide, typically for the purpose of protecting the free thiol group of the cysteine ​​from oxidation. Cysteine ​​can be protected with other types of protecting groups, such as α-ethyl maleimide (Schilling, D., Barayeu, U., Steimbach, RR, Talwar, D., Miller, AK, & Dick, TP (2022). Commonly used alkylating agents limit persulfide detection by converting protein persulfides into thioethers. Angewandte Chemie, 134(30), e202203684. DOI: 10.1002 / anie.202203684. PMID: 35506673).

[0121] In some embodiments, the amino acid containing a modified side chain is a lysine whose side chain is acetylated, typically an acetyllysine. Acetyllysine can be formed endogenously in certain proteins by post-translational modification of a lysine residue. This amino acid can thus be found in the peptides of the sample although it is not a proteinogenic acid.

[0122] In some embodiments, the amino acid containing a modified side chain is a methionine whose side chain is oxidized, typically a methionine sulfoxide (also called methionine sulfoxide).

[0123] Methionine sulfoxide can be formed in some proteins by post-translational modification of a methionine residue. This amino acid can thus be found in the peptides in the sample even though it is not a proteinogenic acid.

[0124] In some embodiments, the amino acid is a selenium-containing amino acid. Selenium-containing amino acids are analogs of standard amino acids that contain selenium instead of sulfur in their structure. These selenium-containing amino acids can be incorporated into proteins during protein biosynthesis, in place of the corresponding sulfur-containing amino acids. The principal selenium-containing amino acids are selenomethionine (Sem) and selenocysteine ​​(Sec), which are the selenium-containing equivalents of methionine and cysteine, respectively. Peptides comprising selenium-containing amino acids can be prepared by chemical peptide synthesis, recombinant expression, biocatalysis, or via post-synthesis modification techniques.

[0125] Importantly for the method of the invention, the peptides contain at least one terminal amino acid labeled with an end tag. The end tag amino acid may be an amino acid whose amine group (-NH2) or carboxyl group (-COOH) is modified to label the peptide of which it constitutes the A / -terminus or C-terminus respectively.

[0126] In this document, end markers are understood to mean agents or chemical groups specifically used to mark or label the N-terminus or C-terminus of proteins or peptides. A distinction is therefore made between A-terminus markers and C-terminus markers. Such end markers are known in the fields of molecular biology research, proteomics, and in the fields of protein identification, purification, and analysis.

[0127] In one or more embodiments, the amino acid is an A / -terminal amino acid whose amino group (-NH2) is modified with an A / -terminal tag.

[0128] The A / -terminal end marker may in particular be selected from dansyl, dabsyl, 1-naphthyl isocyanate (IsoC), or 1-naphthyl isothiocyanate, each of these end markers optionally being labeled with one or more carbon atoms 13 ( 13 C), deuterium (D), nitrogen 15 ( 15 N), and / or sulfur 34 ( 34 S).

[0129] Dansyl optionally labeled with one or more carbon atoms 13 ( 13 C), deuterium (D), nitrogen 15 ( 15 N), and / or sulfur 34 ( 34 S) is preferred.

[0130] In one or more embodiments, the terminal amino acid is a C-terminal amino acid whose carboxyl group (-COOH) is modified with a C-terminal tag.

[0131] The C-terminal end marker may in particular be selected from 2-(A / ,A / -dimethylamino)-1-ethylamine (DMED) and 3-(A / ,A / -dimethylamino)-1-propylamine (DMAPA), each of these end markers optionally being labeled with one or more carbon atoms 13 ( 13 C), deuterium (D), and / or nitrogen 15 ( 15 N).

[0132] In one or more embodiments, the peptides present in the sample are peptides of which only one terminal amino acid is labeled with an end tag, in other words are peptides of which only one terminal is labeled with an end tag.

[0133] In one or more embodiments, the peptides are peptides whose two terminal amino acids are labeled with an end tag, in other words are peptides whose two ends are labeled.

[0134] The library may be calculated (or generated) such that the predetermined peptides contain amino acids selected from proteinogenic amino acids, modified amino acids, in particular amino acids containing a modified side chain, and mixtures thereof, it being understood that the predetermined peptides contain at least one terminal amino acid labeled with an end tag.

[0135] The predetermined peptide library (or libraries) may be generated before each measurement of a sample by the device for implementing the method, or may be received by any means of communication. For example, the library may be generated on a third-party device and transmitted by it to the device implementing the method.

[0136] In one or more embodiments, the library of predetermined peptides may comprise all possible peptide sequences of 2 to 6 amino acids obtained from a group of N amino acids.

[0137] By all possible peptide sequences is meant all possible combinations forming peptide sequences of a length between 2 and 6 amino acids, for example between 2 and 4 amino acids, obtained from a group of N amino acids.

[0138] For example, N may be in a range of 2 to 50 amino acids. There are only about twenty naturally occurring amino acids. However, as detailed below, the library may be constructed to account for the presence of modified amino acids in the measured peptides, whether they are naturally occurring or synthetically derived modified amino acids. Thus, in one or more embodiments, N may preferably comprise at least the 22 proteinogenic amino acids and modified amino acids of said proteinogenic amino acids. In one or more embodiments, N may be greater than 2, preferably greater than 20. In one or more embodiments, N may be less than 500, 400, 300 or 200. N may typically be between 2 and 100, for example between 2 and 50.

[0139] Thus, each sequence of the predetermined library may define a predetermined peptide of the library and may further comprise a terminal amino acid which is labeled with the marker.

[0140] Furthermore, in one or more embodiments, the memory may comprise sets of information associated respectively with peptides present in the sample, called measured peptides, measured by a mass spectrometer, each measured peptide may comprise a peptide sequence having at least one terminal amino acid labeled with an end marker. Each set of information of a measured peptide may comprise at least one MS / MS spectrum comprising a total mass / charge ratio of the measured peptide and one or a plurality of mass / charge ratios of the measured fragment ions.

[0141] Each fragment ion may be representative of the presence of one or a plurality of amino acids (or amino acid residues) linked in the peptide sequence of the measured peptide.

[0142] Furthermore, according to one or more embodiments, each set of information relating to a measured peptide may also comprise mass spectrum information (referred to as MS spectrum), the MS spectrum information being able to comprise a mass / charge ratio of the measured peptide (i.e., of the parent ion), a retention time, and an intensity of the measured peptide (i.e., an intensity of the parent ion).

[0143] Further, the measured peptides and at least a plurality of predetermined peptides from the predetermined peptide library may be labeled at the same end with the same end tag.

[0144] According to one or more examples, the mass spectrometer may be a mass spectrometer with tandem analysis capability (MS / MS), preferably coupled to liquid chromatography (LC), for example using a liquid chromatographic column.

[0145] Such a mass spectrometer, optionally coupled to liquid chromatography, makes it possible to obtain sets of information on peptides measured in a sample such as mass spectra (MS) of the peptide (also called parent ion), and / or MS / MS fragmentation spectra giving information on the structure of the peptide (or parent ion) from the fragment ions present in the MS / MS spectra.

[0146] For example, the information sets relating to these spectra (MS and MS / MS) may include mass / charge ratios and peak intensities relating to fragment ions and / or parent ions constituting the measured peptides, as well as retention times of the measured peptides.

[0147] Furthermore, a set of information of a measured peptide may comprise one MS spectrum and one or a plurality of MS / MS spectra, generally close in content. In these cases, it may be decided to use a single MS / MS spectrum (e.g., that of the parent ion with the highest intensity) or the plurality of MS / MS spectra.

[0148] Liquid chromatography (LC), using a chromatographic column for example, can be used to separate peptides before their measurement by the mass spectrometer, and also allows a retention time to be determined.

[0149] This retention time refers to the length of time a peptide in the sample remains in the chromatographic column (of the chromatograph) before being eluted, i.e. until the peptide reaches the mass spectrometer. This retention time (for each peptide) can therefore be measured from the time the sample is injected into the chromatograph until the time a peptide reaches the mass spectrometer.

[0150] Figure 2 schematically illustrates an example of a predetermined peptide library.

[0151] In this schematic example, the generated library may include all possible peptide sequences (SEQ_PEP) (twelve in this case) comprising 2 to 3 amino acids that can be obtained from a group of N=3 amino acids (A, B, C), and where each peptide sequence in the library may include a terminal amino acid that is labeled with an end tag.

[0152] As previously described for the sample peptides, the labeled end of the terminal amino acid of the library peptide sequences may be the A / -terminal and / or C-terminal. Preferably, the labeled end is the N-terminal (or A / -terminal type).

[0153] For example, in the case of the library in Figure 2, the labeled end is the A / -terminus. Thus, the end amino acid labeled with the end tag * for the predetermined peptide PP2 is amino acid B, the end amino acid labeled with the end tag * for the predetermined peptide PP8 is amino acid A, etc.

[0154] Each peptide sequence (SEQ_PEP) can respectively define a predetermined peptide PP1...PP12 of the library and can be used to determine or calculate the total mass / charge ratio (m / z) m1...m12 of the predetermined peptide as described previously (R_seq_pep). For example, the mass / charge ratio (m / z) m7 can be calculated (by the sum for example) from the peptide sequence A*-BC of the peptide PP7 with A* the terminal amino acid labeled (i.e. by the sum of the mass / charge ratios of the amino acid residues constituting the peptide sequence).

[0155] Note also that in this example of Figure 2, the ratios m1 and m2 are equal, since the peptide sequences of the predetermined peptides PP1 and PP2 comprise the same amino acids and in the same number. This reasoning can also be applied in the case of the predetermined peptides PP3 and PP4, the predetermined peptides PP5 and PP6, as well as the predetermined peptides PP7 to PP12.

[0156] Note that the mass / charge ratio of a labeled end amino acid takes into account the end marker in its mass / charge ratio. Indeed, an A amino acid does not have the same mass / charge ratio as an A* labeled amino acid with the * end marker.

[0157] Of course, in the example of Figure 2, the labeled end is A / -terminal (or A / -terminal type) but it can also be C-terminal (or C-terminal type). For example, the predetermined library can include both predetermined peptides having an A / -terminal type labeled end and predetermined peptides having a C-terminal labeled end, and / or predetermined peptides having a C-terminal labeled end and an A-terminal labeled end.

[0158] According to one or more examples, the measured peptides and the predetermined peptides of the predetermined peptide library may be labeled at the same end with the same end tag.

[0159] For example, the measured peptides may all have an A / -terminal (or A / -terminal-like) labeled end, and the predetermined peptide library includes at least a plurality of predetermined peptides having an A / -terminal labeled end among predetermined peptides having an A / -terminal, and / or C-terminal, and / or a combination of both.

[0160] In one or more embodiments, the peptide sequences of the library may all have the same labeled end, for example A / -terminal or C-terminal.

[0161] In one or more embodiments, the amino acids in a group of N amino acids may be proteinogenic amino acids or modified amino acids. For example, the modified amino acids may be amino acids with a modified side chain.

[0162] Thus, the predetermined peptides may comprise a labeled A-terminal or C-terminal amino acid, proteinogenic amino acids, and side chain modified amino acids. Preferably, the side chain modified amino acids are selected from the group consisting of cysteine, lysine, and methionine.

[0163] According to one or more embodiments, the cysteine ​​of the predetermined peptides is a cysteine ​​whose side chain is modified, preferably by carbamidomethylation. This modification makes it possible in particular to avoid the unwanted formation of disulfide bonds during the preparation and analysis of the sample.

[0164] Further, according to one or more examples, the library may comprise at least one predetermined peptide of the library which is broken down into a plurality of predetermined charged peptides, each predetermined charged peptide having the same peptide sequence as said at least one predetermined peptide and a respective charge.

[0165] By predetermined charged peptide, it can be understood a predetermined peptide having one or more positive or negative charges.

[0166] Thus, for example, it is possible to decline one or more predetermined peptides from the library into a multitude of loaded versions, each version charged having its own charge (or respective charge) but having the same peptide sequence of the predetermined peptide from which it declines.

[0167] Thus advantageously, it is possible to determine a peptide sequence of a measured peptide which comprises one or more charges, for example which is monocharged, dicharged, or tricharged.

[0168] Referring to Figure 2, it is also illustrated the fragment ions (or a set of fragment ions) that can be calculated from the amino acid sequence of the predetermined peptide. Each fragment ion (or fragment or peptide fragment ion) of the fragment ions (lons_Frag) can be associated with a respective mass / charge ratio (RJon_frag).

[0169] For example, the predetermined peptide PP1 may comprise a peptide sequence (Seq-pep) A* - B, where A* is the end acid labeled with the end tag *, and the fragment ions of the sequence (or set of fragment ions) A*, B as well as other fragments calculated (or obtained) from the sequence A* - B (by a fragmentation method). The fragment ions (or set of fragment ions) calculated for the predetermined peptide PP2 may comprise the fragment ion B*, the fragment ion A and other fragments of the sequence B* - A, the fragment ion B* comprises the end tag *.

[0170] According to one or more examples, the fragmentation method may be identical to the fragmentation method of the mass spectrometer or a mass spectrometer with tandem analysis (MS / MS) capability (or liquid chromatography-tandem mass spectrometry system).

[0171] Indeed, the fragment ions in the predetermined peptide library that are obtained (or calculated) from a respective peptide sequence may be equivalent to a peptide (or parent ion) that would be fragmented by tandem mass spectrometry to yield measured fragment ions.

[0172] For example, fragmentation methods can be: - Dissociation by Collisional Activation (CID), - Electron transfer induced dissociation (ETD), - Electron capture induced dissociation (ECD), - Electron impact-induced fragmentation (El D), - Dissociation induced by infrared multiphoton absorption (IRMPD), - LIV-induced photodissociation (IIVPD).

[0173] Further, according to one or more examples, fragment ions can be calculated from a predetermined peptide by applying a plurality of fragmentation methods.

[0174] For example, each fragmentation method may be specific to a type of mass spectrometer (or a tandem mass spectrometry system) using a specific fragmentation method. In one example, in the library of predetermined peptides, a peptide sequence may comprise fragments (or fragment ions), for example grouped as sets of fragment ions, a first set of fragment ions being obtained according to a collisional activation dissociation (CID) fragmentation method, and a second set of fragment ions may be obtained according to an electron transfer dissociation method (see previous example list).

[0175] Thus advantageously, the same library can be used to identify peptides in a sample that would have been measured by different mass spectrometers (or tandem mass spectrometry systems) respectively using a specific fragmentation method or being able to use different fragmentation methods (see previous list).

[0176] With reference to Figure 1, the method may comprise a determination 110 of a selection of predetermined peptides from the library and said at least one MS / MS spectrum.

[0177] Each predetermined peptide in the selection may have a total mass-to-charge ratio equal to the total mass-to-charge ratio of the measured peptide within an error range less than a predetermined value.

[0178] Thus, for example, determining the selection for a measured peptide may involve comparing the total mass-to-charge ratio of the measured peptide, i.e., the mass-to-charge ratio of the parent ion in the MS / MS spectrum of the measured peptide, with each predetermined total peptide mass-to-charge ratio present in the library, and selecting a predetermined peptide from the library whenever the total mass-to-charge ratio present in the library is equal to (or close to) the total mass-to-charge ratio (mass-to-charge ratio of the parent ion) of the measured peptide within an error range less than a predetermined value.

[0179] According to an example, and with reference to Figure 2, if the mass / charge ratio of the measured peptide is equal to or close to the ratio m1, then the selection (grayed part in Figure 2) will include the predetermined peptide PP1, but also the predetermined peptide PP2, since the ratio m2 is equal to the ratio m1. Indeed, the predetermined peptides PP1 and PP2 include the same amino acids as well as the same number of amino acids and the same end tag.

[0180] According to one or more examples, the predetermined value of the error range may be less than 20 ppm (parts per million), preferably less than 15 ppm, preferably less than 10 ppm, and even more preferably less than 5 ppm.

[0181] Once a selection is determined for a measured peptide, the method may include determining 120 the labeled end amino acid of the measured peptide by comparing the measured fragment ion mass-to-charge ratios with the calculated fragment ion mass-to-charge ratios of the predetermined peptides in the selection (i.e., fragment ions in the shaded portion in the example of Figure 2) within an error range less than a predetermined value.

[0182] According to one or more examples, the comparison may relate to the mass / charge ratios of the measured fragment ions exhibiting the end tag (detected in the MS / MS spectrum of the measured peptide) which are compared with the mass / charge ratios of the fragment ions of the selection, and / or with the mass / charge ratios of the fragment ions of the selection exhibiting the end tag.

[0183] Thus, determining the labeled end amino acid, prior to any sequence reconstruction, can allow the determination of the end amino acid of the sequence which can then be used as a starting point for reconstructing the peptide sequence of the measured peptide.

[0184] The method may then comprise 130 determining at least one reconstructed peptide sequence for the measured peptide from the calculated fragment ions of the selection (i.e., fragment ions of the shaded portion in the example of Figure 2), the determined labeled end amino acid, and the measured fragment ions. The method may then comprise rendering 140 an identification list of peptides measured in said sample wherein each measured peptide in the list is defined by the reconstructed peptide sequence.

[0185] According to one or more examples, each measured peptide in the list may be associated, by an identification means, with the information set, e.g., the MS / MS spectrum, used to determine the reconstructed peptide sequence of the measured peptide. For example, the identification means may correspond to a numbering of each information set, and each measured peptide in the list may be associated with a respective number that corresponds to the number of the information set used to determine the reconstructed peptide sequence of the measured peptide.

[0186] Note that when a plurality of libraries are used, for example two libraries as described above, a first library can be used to determine the selection, and the second library can be used to reconstruct the sequence, the predetermined peptides of the first library (and therefore of the selection) can also be present in the second library.

[0187] According to one or more examples, preferably, it is also possible to use a first library as described above, and calculate a second library, as described above, which contains only the predetermined peptides of the selection (as well as the calculated fragment ions and the mass / charge ratios of the calculated fragment ions).

[0188] According to one or more embodiments, the peptide sequence (i.e. at least one peptide sequence) of a measured peptide can thus be reconstructed from: - the determined labeled end amino acid, - comparisons between the mass / charge ratios of measured fragment ions and calculated fragment ions.

[0189] When reconstructing a peptide sequence for a measured peptide, i.e., after determining the labeled end amino acid (starting point), comparisons are made between the measured fragment ions and the calculated fragment ions (from the selection), in order to determine the reconstructed peptide sequence of a measured peptide. For example, the comparisons may involve the different mass-to-charge ratios of the available fragment ions, i.e., the measured and / or calculated fragment ions.

[0190] In one or more embodiments, the peptide sequence of each measured peptide may be reconstructed using a method derived from methods used in so-called "de novo" sequencing. Briefly, "de novo" sequence reconstruction relies on assembling sequences of overlapping sequence fragments to deduce complete sequences. The reconstruction method may be based on the following articles: - Biemann, K. (1992). Mass spectrometry of peptides and proteins. Annual review of biochemistry, 61 (1), 977-1010, DOI: 10.1146 / annurev.bi.61.070192.004553. PMID: 1497328. - Wysocki, VH, Resing, KA, Zhang, Q., & Cheng, G. (2005). Mass spectrometry of peptides and proteins. Methods, 35(3), 211-222. DOI: 10.1016 / j.ymeth.2004.08.013. PMID: 15722218. - Seidler J, Zinn N, Boehm ME, Lehmann WD. De novo sequencing of peptides by MS / MS. Proteomics. 2010 Feb ;10(4):634-49. doi:10.1002 / pmic.200900459. PMID: 19953542.

[0191] The peptide sequence of a measured peptide can be reconstructed using "de novo" sequencing principles from the MS / MS spectrum(s) information of the measured peptide, the determined labeled end amino acid, and the fragment ions of the predetermined peptides in the selection.

[0192] In the process of measuring a peptide by mass spectrometry, it may happen that the collected data include irrelevant elements. These spurious data may be due to artifacts generated by the mass spectrometer or by the simultaneous presence of another peptide (different from the peptide to be measured). The latter situation can occur, in particular, when two peptides with similar retention times co-elute in the spectrometer.

[0193] Figure 3 schematically illustrates this phenomenon of co-elution of two peptides.

[0194] Note that the co-elution may include more than two peptides, for example three or four or more peptides, which may have different peptide sequences, and / or may have identical peptide sequences but are differently charged.

[0195] As shown in Figure 3, each retention time RT1 and RT2 can correspond to the retention time of a peptide from the sample that has migrated through the chromatographic column to the mass spectrometer (or tandem mass spectrometry system). As can be seen in Figure 3, in case of co-elution, the chromatographic peaks of the two peptides can overlap, for example, partially or completely, because of their retention times that are very close. In such a situation, the tandem mass spectrometer can then record a set of information, i.e. a combined spectrum, also called chimeric, of the co-eluted peptides.

[0196] Such a situation can make it difficult to determine the peptide sequence truly corresponding to the measured peptide, since a plurality of peptide sequences may be included in the information set (eg MS / MS spectrum(s)) of the measured peptide, and each peptide sequence may potentially correspond to the peptide sequence of the measured peptide.

[0197] Thus, in one or more embodiments, when a plurality of reconstructed peptide sequences is determined for the same measured peptide, for example in the presence of spurious data due to co-elution of one or more peptides, the method may further comprise: - calculate, for each reconstructed peptide sequence, a score from the set of information associated with the measured peptide and the reconstructed peptide sequence, - select the reconstructed peptide sequence for the measured peptide from a comparison of the scores of the reconstructed peptide sequences, the peptide sequence with the highest score being selected

[0198] The reconstructed peptide sequence with the highest score can then truly correspond to the measured peptide, and can be returned (provided) in the identification list with the obtained score. Only the peptide sequence with the highest score is retained in the identification list.

[0199] However, the reconstructed sequence(s) for a measured peptide that had a lower score, and that correspond to one or more eluted peptides, may also correspond to peptides in the mixture that need to be measured and identified, for example because they also have peptide sequences with a number of amino acids between 2 and 6 amino acids. These eluted peptides (for example of the same peptide sequence but charged differently and / or of different peptide sequences) with a measured peptide may therefore also correspond to measured peptides, and can generally be determined before or after depending on the time resolution of the mass spectrometer (or tandem mass spectrometry system).

[0200] For example, referring to Figure 3, the peptide sequence of the measured peptide PP1 that elutes at retention time RT 1 (close to retention time RT2), and the peptide sequence that elutes at retention time RT2 can be determined in the same two retention times. In this case, i.e. in the presence of co-elution, the reconstructed sequence with the highest score will correspond to peptide PP1 at retention time RT1, while it will correspond to PP2 at retention time RT2. Both peptide sequences will be included in the identification list with the same retention time.

[0201] Upon co-elution, which may result in the presence of a plurality of measured peptides from the identification list having similar or equal retention times, an alert may be notified, e.g. to a user or entered in the identification list, and post-processing (e.g. verification of results) of the identification list may be performed. For example, data for a retention time of a measured peptide obtained from the different data sources (LC chromatogram, MS spectrum, etc.) may be compared to determine that the retention time corresponds to this measured peptide (e.g. using an extracted ion chromatogram, XIC, with the MS detection and comparing peak intensities and / or retention times in MS and MS / MS spectra). As another example, when there are several peptides of the same sequence but charged differently, post-processing may consist of keeping only the peptide from the list with the highest intensity, for example.

[0202] In one or more embodiments, each mass / charge ratio of the measured fragment ions can be respectively associated with an intensity in the MS / MS spectrum, and the score of each reconstructed peptide sequence for the same measured peptide can be determined from the following equation:

[0203] By intensity of a fragment ion, it can also be understood the intensity of the fragment ion peak.

[0204] According to one or more examples, only fragment ions exhibiting the end tag can be used to calculate the score of a peptide sequence.

[0205] Thus, for example, referring to Figure 2, if the fragment ions of the MS / MS spectrum of a measured peptide correspond to the predetermined peptides PP1 and PP2 of the library (i.e. of the selection, grayed part in the example of Figure 2), only the fragments A* - A*B and B* - B*A of PP1 and PP2 can be used to calculate the scores of each.

[0206] In one or more embodiments, one or a plurality of labeled end amino acids may be determined, and for each determined labeled end amino acid, at least one reconstructed peptide sequence for the measured peptide may be determined from the calculated fragment ions of the selection, the determined labeled end amino acid, and the measured fragment ions.

[0207] Indeed, as described above, in the presence of a co-elution of one or more peptides (for example of the same peptide sequence but charged differently and / or of different peptide sequences), a plurality of reconstructed peptide sequences can be determined. These reconstructed peptide sequences can have in common the same determined labeled end amino acid or have different determined labeled end amino acids. For example, in the peptide sequences A* - B - C and A* - C - B having the same mass / charge ratio (or very close), A* is the determined labeled end amino acid that is common to both sequences. According to another example, in the sequences A* - B - C and B* - A - C having the same mass / charge ratio (or very close), A* and B* are the determined labeled end amino acids that are different between the two sequences. In the latter case, when reconstruction of peptide sequences, each determined labeled end amino acid is considered as a starting point for at least one peptide sequence to be reconstructed.

[0208] Furthermore, in one or more embodiments, the list may comprise retention times respectively associated with the measured peptides, each retention time being obtained from the information set, in particular from the MS spectrum (and / or also the MS / MS spectrum) for example, used to determine a reconstructed peptide sequence of a measured peptide.

[0209] Each measured peptide in the list can be associated with a respective retention time obtained from the information set (eg MS spectrum) used to determine the reconstructed peptide sequence of the measured peptide.

[0210] According to one or more examples, the retention time for each measured peptide can be obtained following liquid chromatography (LC) coupled with mass spectrometry detection such as a tandem mass spectrometry (MS / MS) system.

[0211] As described previously, the retention time can be obtained from the measurement performed by liquid chromatography, and transcribed (or indicated) in the MS spectrum of the parent ion (and / or the MS / MS spectrum(s) from the parent ion).

[0212] According to one or more examples, the list may further comprise intensities respectively associated with the measured peptides, each intensity being obtained from the set of information (e.g. in the MS spectrum of the measured peptide) used to determine a reconstructed peptide sequence of a measured peptide. Each measured peptide in the list may comprise a proportion value of the peptide in the mixture, the proportion value "Qp", for each measured peptide in the list, may be calculated from: Intensity of a measured peptide from list P of intensities of measured peptides from list

[0213] Advantageously, it is thus possible to obtain values ​​of relative proportions of the peptides in the mixture.

[0214] Figure 4 illustrates a device for implementing the method of the present disclosure.

[0215] In this embodiment, the device 400 may include a circuit 403 and a memory 402 for storing program instructions that can be loaded into the circuit, and capable of causing the circuit 403 to execute the method of the present disclosure when the program instructions are managed by the circuit 403.

[0216] The memory 402 may also store data and information useful for carrying out the method of the present disclosure as described above.

[0217] Circuit 403 can be for example: - a processor or processing unit capable of interpreting instructions in a computer language, the processor or processing unit may comprise, be associated with or be attached to a memory comprising the instructions, or - the combination of a processor / processing unit and a memory, the processor or processing unit being adapted to interpret instructions in a computer language, the memory comprising said instructions, or - an electronic card in which the process sequence is described in silicon, or - a programmable electronic chip such as an FPGA chip (for “Field-Programmable Gate Array”), - a graphics processor, or GPU (from the English Graphics Processing Unit).

[0218] This device may comprise an input interface 405 for receiving input data and an output interface 407 for providing a set of useful data and / or driving a characterization system for example, such as a tandem mass spectrometry system coupled with liquid chromatography.

[0219] For example, the input interface 405 may receive input data such as a library or a plurality of libraries (as described above for example) of predetermined peptides or parameters for generating a library of predetermined peptides, or even sets of information associated respectively with peptides present in the sample measured by a mass spectrometer. Furthermore, optionally, the input interface 405 may be connected to a mass spectrometer such as a tandem mass spectrometry system (mass spectrometer with tandem analysis capability (MS / MS), coupled to liquid chromatography so as to allow the measurement of the sample and receive the sets of information associated respectively with peptides present in the sample directly from the mass spectrometer.

[0220] For example, the output interface 407 may provide, for example, a library or a plurality of libraries (e.g., as described above) generated from predetermined peptides, or an identification list of peptides measured in a sample. Further, optionally, the output interface may be connected to a mass spectrometer such as a tandem mass spectrometry system coupled with liquid chromatography so as to allow the control of the mass spectrometer for example, or to provide the sets of information associated respectively with peptides present in the sample directly from the mass spectrometer, for example sending them to a third-party device such as a computer.

[0221] In one example, the device 400 may be a computer 401 comprising the circuit 403 and the memory 402. In another example, the device may be the computer of the mass spectrometer. The computer of the mass spectrometer (MS or LC-MS / MS) may be used, for example, to control the mass spectrometer to perform measurements on a sample, to generate or receive one or more libraries, or to implement the method of the present disclosure.

[0222] To facilitate interaction with the device 400 or the computer 401, a screen 411, a keyboard 412, and a mouse 413 may be provided and connected to the computer circuit 403. Analysis method

[0223] The disclosure also relates to a method for analyzing peptides having between 2 and 6 amino acids of a sample comprising: i) a step of processing the sample comprising at least one step h) of labeling one end of the peptides of the sample by bringing the sample into contact with an end marker, ii) a step of analysis by liquid chromatography coupled to a mass spectrometer, preferably to a mass spectrometer with a tandem analysis capacity (MS / MS), iii) a data processing step implementing a method of non-targeted identification of peptide sequences as described in the present document.

[0224] The peptides, amino acids, end tags, and samples analyzed are as described in the aspect above.

[0225] In one or more embodiments, the peptides are natural and / or synthetic peptides.

[0226] In one or more embodiments, the analyzed peptides have between 2 and 4 amino acids.

[0227] In one or more embodiments, the peptides subjected to step ii) contain at least one labeled end amino acid and one or more amino acids selected from proteinogenic amino acids and modified amino acids.

[0228] In one or more embodiments, the sample is a mixture of small natural peptides or a mixture of peptides derived from peptide synthesis. The sample is preferably a protein hydrolyzate, for example a yeast protein hydrolyzate, a plant protein hydrolyzate, or an animal protein hydrolysate.

[0229] In one or more embodiments, step i) of treating the sample further comprises a step i2) of alkylating the peptides of the sample by contacting with an alkylating agent, and step i2) preferably preceding step h).

[0230] This optional alkylation step is performed to protect amino acids carrying a free thiol group (-SH), such as cysteine, from undesirable chemical reactions that may occur during subsequent steps of the method such as oxidations.

[0231] The person skilled in the art knows how to select compatible liquid chromatography techniques for coupling with mass spectrometry.

[0232] In some embodiments, the liquid chromatography is reversed-phase liquid chromatography, preferably performed with a gradient of at least one hour.

[0233] In some embodiments, the liquid chromatography (LC) implemented is nanoscale liquid chromatography (Nano LC) or microscale liquid chromatography (Micro LC).

[0234] Nano LC is a high-performance liquid chromatography (HPLC) technique used to separate and analyze samples at flow rates in the hundreds of nanoliters / min (nL / min). It shares similar principles with conventional liquid chromatography, but operates at much lower flow rates and sample volumes, typically in the nanoliter range.

[0235] Micro LC, or microscale liquid chromatography or capillary LC, is a liquid chromatography (LC) technique that uses smaller columns and reduced flow rates in the hundreds of microliters / min (pL / min) compared to conventional liquid chromatography (HPLC). Micro LC falls between conventional HPLC and nano LC in terms of flow rates and column size.

[0236] In some embodiments, the mass spectrometer is equipped with a mass analyzer preferably selected from a quadrupole trap mass analyzer and a triple quadrupole mass analyzer, more preferably is equipped with a high-resolution mass analyzer selected from an orbital trap mass analyzer, a time-of-flight (Q-TOF) mass analyzer, a TIMS- mass analyzer TOF (“Trapped Ion Mobility Spectrometry” - “Time-of-Flight”) and a Fourier Transform Cyclotron Resonance (FT-ICR) mass analyzer.

[0237] This disclosure is illustrated, without limitation, by the following examples. Examples Example 1: Method for detecting short peptides in protein hydrolysates. Goals

[0238] In this example, we are trying to characterize the short peptides present in a yeast extract obtained by hydrolysis of yeast cream. During hydrolysis, the proteins in the yeast cream are fragmented into peptides which can be broken down as follows: - oligopeptides comprising between 5 and 35 amino acids, - short peptides with between 2 and 5 amino acids, - free amino acids.

[0239] Short peptides and free amino acids form the majority components. Proteins and oligopeptides can be analyzed by classical proteomics and peptidomics analyses. In a first step, the peptide mixture is decomplexed using chromatographic techniques (especially liquid chromatography) or capillary electrophoresis, detected and then subjected to fragmentation in the mass spectrometer to obtain amino acid sequence information. The data are typically obtained by liquid chromatography analysis coupled with tandem mass spectrometry (LC-MS / MS). The MS and MS / MS data are then analyzed using software for peptide identification by mass (m / z) and sequence homology searches in protein data catalogs.However, these techniques are not reliable for small peptides, which is why most software does not allow the analysis of peptides comprising less than 5 or 6 amino acids.

[0240] A method is proposed here to analyze the protein content of hydrolyzed yeast extract. We have identified that short peptides, which are among the major components of yeast extracts, present analytical difficulties by LC-MS / MS related to their physicochemical properties whether in terms of separation and recovery or in detection by mass spectrometry (short peptides do not fit into the spectrometer detection windows). In addition, their small size reduces the reliability of their identification by sequence homology. We have developed a systematic and non-targeted method allowing the identification of a large number of short peptides in complex mixtures such as food matrices, protein hydrolysates.

[0241] The method developed is based on two components: the implementation of an analytical method including sample treatment to label short peptides by dansylation, and instrumental analysis by LC-MS / MS, followed by data processing to identify the sequence of short dansylated peptides in the sample from the dataset generating short peptides dansylated on their A / -terminal part. The dansylation step is preferably preceded by an alkylation step because without this step, we found it difficult to detect amino acids comprising free thiol groups sensitive to oxidation, in particular cysteine. This alkylation step protects the free thiol group and makes it better detected after labeling.The data processing step is based on the use of a suitable and calculated catalogue of short peptides, which does not correspond to a catalogue comprising a set of peptides resulting from the digestion of endogenous or unlabelled proteins (directly accessible online or generated by in silico digestion from proteins contained in a database) as is classically the case, but to a catalogue comprising all possible combinations of natural and modified amino acids leading to dipeptides, tripeptides, and tetrapeptides, dansylated at their A / -terminal end, i.e. a catalogue comprising approximately 6,250,000 predetermined peptides for N=50. Materials and methods Sample processing

[0242] Peptides in a given sample, typically a yeast protein hydrolysate, are alkylated and then dansylated. Alkylation

[0243] Peptides are alkylated to protect peptides containing free thiol groups.

[0244] The protocol is as follows: The sample is incubated for 1 h at room temperature and in the dark with 5 mM iodoacetamide in 100 mM ammonium bicarbonate (ABC) at pH 8.8. N-terminal labeling with dansyl chloride

[0245] The protocol is as follows: . Suspend 1 mg of the sample in 1 mL of 100 mM ammonium bicarbonate (ABC) buffer, pH 8.8. . Suspend 1 mg of the sample in 1 mL of 100 mM ammonium bicarbonate (ABC) buffer, pH 8.8. . Take a 25 pL aliquot of the suspended sample and mix with 12.5 pL of acetonitrile (ACN) and 12.5 pL of 100 mM ABC buffer pH 8.8. . Vortex the mixture. . Add 25 μL of a solution of 18 mg / mL of dansyl chloride solubilized in ACN and homogenize by vortexing. . Incubate for 1 hour at 40°C. Stop the reaction by adding 5 µL of 250 mM NaOH (to remove excess dansyl chloride) (quench step) and homogenize by vortexing. . Heat the mixture for 10 minutes at 40°C. . Add 25 pL of 425 mM formic acid (FA) (diluted in a 50 / 50 ACN / water solution) and homogenize by vortexing. . Evaporate in SpeedVac. . Reconstitute the dry product in 100 pL of water. . Filter the sample onto a 96-well plate with a PVDF membrane (pore size 0.45 pm) (after wetting the wells with 100 μL of 70% EtOH and washing them twice with 200 μL of milliQ water), in order to remove aggregates. Do not dry. . Start the instrumental analysis by LC-MS / MS by diluting the filtrate 1 / 10 in water + 0.1% AF and injecting 1 pL of the dilution into the instrument. Instrumental analysis by LC-MS / MS

[0246] Dansylated peptides are analyzed by liquid chromatography coupled with tandem mass spectrometry. The steps of this analysis are: a. Separation of dansylated peptides by liquid chromatography b. Detection of dansylated peptide ions (parent) by MS (parent ion spectrum) c. Fragmentation in MS / MS (fragment ion spectrum from a parent ion). Liquid chromatography

[0247] The liquid chromatography performed is nanoscale liquid chromatography (nano-LC) with a C18 column (Acclaim™ PepMap™ 100 C18) classically used in proteomics.

[0248] The gradient is more than one hour for example two hours to improve the resolution of the chromatography. Detection

[0249] Detection is performed by tandem mass spectrometry. Precursor ions are detected by MS and fragmented by MS / MS. A high-resolution orbital trap mass analyzer in positive ion mode is used.

[0250] The detector allows to obtain sets of information for each measured peptide, such as the sets of information described previously. Data processing

[0251] The method of non-targeted identification of peptide sequences previously presented can be implemented, for example by a device such as presented in figure 4, using the sets of information obtained above. Results

[0252] The protein content of short peptides within the sample was characterized. The sequences of all short peptides present in the sample (between 350 and 500 short peptides of different sequences) were identified. The implementation of the method allows the characterization of short peptides in complex mixtures, and their effects in terms of technological and organoleptic properties of the hydrolysates or food matrices in which they are present. Example 2: Study of short peptides in soybean hydrolysates, casein and yeast extracts. 1. Objective of the study

[0253] In the present Example, the method of the invention is used to characterize short peptides present in soy, casein and yeast extract hydrolysates from various manufacturers, intended for different applications, obtained from different sources following various hydrolysis processes. 2. Materials and methods 2.1. Chemicals, reagents and samples

[0254] Ammonium bicarbonate (A6141), acetonitrile ACN for LC-MS LiChrosolv® (1000292500), glycine (G7403), formic acid FA (1.00264), hydrochloric acid HCl 37% (320331), iodoacetamide (11149), sodium 1,2-naphthoquinone-4-sulfonate NQS (70382), sodium hydroxide NaOH (221465), trifluoroacetic acid TFA (T6508) were purchased from Sigma Aldrich (St. Louis, MO, USA). Dansyl chloride DNS-CI (D0656) was supplied by TCI chemicals Europe (Zwijndrecht, Belgium). The Milli-Q ZRQSVP3WW purification system from Merck Millipore (Burlington, MA, USA) was used for the preparation of ultrapure water. The AccQ-Tag derivatization kit (186003836), amino acid standard (WAT088122), eluent A (186003839), eluent B (186003838) for Determination of total and free amino acids were obtained from Waters Corporation TM (Milford, MA, USA).

[0255] The different hydrolyzed samples available in the market were selected according to the starting raw material (plant, animal, microorganism) as well as the hydrolysis process (fermentation, enzymatic hydrolysis or autolysis) and are presented in Table 2: Table 2]: Hydrolyzed products included in this study a Kikkoman: Noda-Shi, Japan; b Suzi Wan: a Mars company, McLean, VA, USA; c Sigma Aldrich: St. Louis, MO, USA; d ITW Reagents: Monza, Italy; e Gibco: Life Technologies Miami, FL, USA. 2.2. Analysis by size exclusion chromatography

[0256] To characterize the molecular size profile of proteins and peptides in the hydrolysates, SEC size exclusion chromatography analysis was performed using a Prominence-i LC-2030 Compact HPLC system (228-65800-58, Shimadzu™' Kyoto, Japan) equipped with a Shodex protein KW-802.5 silica column (563781, 8 mm x 300 mm, Phenomenex™, Torrance, CA, USA) at 25 °C and coupled to a UV detector (210 nm and 280 nm). 25 pL of the suspended samples as 1% (wt / wt) solutions and filtered with 0.45 µm cellulose acetate filters in a syringe are injected into the column. Elution was performed with a mobile phase containing 20% of ACN in water with 0.1% TFA at a flow rate of 0.5 mg / mL at 25 °C with a maximum pressure of 5 MPa. Various standards of different molecular weights commonly used for size exclusion chromatography analysis in 2 mg / mL solution were analyzed to determine the molecular weight groups and their retention volumes. Data processing was performed using LabSolutions v5.106 software, and calibration curves were established using GPC postrun software (Shimadzu™, Kyoto, Japan). 2.3. Sample preparation procedure

[0257] Liquid samples of GibcoTM Bacto™ yeast extract and naturally fermented soy sauces were freeze-dried using the Alpha 2-4 LGS plus instrument (CHRIST, Osterode am Harz, Germany). 1 mg of each sample was weighed and dissolved in 100 μL of a 50 mM iodoacetamide solution in ammonium bicarbonate (100 mM, pH 8.8) for alkylation of the free thiol groups of cysteine. This reaction was left for one hour at room temperature and was stopped by dilution with 900 μL of ammonium bicarbonate solution (100 mM, pH 8.8). 25 μL of the mixture was transferred into a new 1.5 mL Eppendorf tube and diluted with 12.5 μL of ACN and 12.5 μL of ammonium bicarbonate solution (100 mM, pH 8.8). Then, 25 μL of DNS-Cl solution (18 mg / mL in ACN) was added to the mixture and incubated for 1 h at 40 °C. The derivatization reaction was stopped with 5 μL of NaOH solution (250 mM) and incubated for 10 minutes at 40 °C.NaOH neutralization was performed by adding 25 μL of formic acid solution (425 mM, in ACN / H2O v / v). The mixture was dried using the Concentrator plus (Concentrator Savant ISS110) from Eppendorf™ (Hamburg, Germany). The labeled samples were then resuspended in 100 μL of 0.1% formic acid solution and then filtered through a hydrophobic PVDF membrane filter plate (MultiScreen HTS® IP, 0.45 μm, MSIPS4510, Merck Milliprore) using the PlatePrep 96-well vacuum filtration station (Supelco™, Bellefonte, PA, USA), according to the manufacturer's instructions. The filtrate was diluted (1 / 10) in 0.1% formic acid solution for LC-MS / MS analysis. 2.4. LC-MS / MS analysis method

[0258] The labeled peptides were separated by reversed-phase liquid chromatography using a nanoflow HPLC system (U3000 RSLC Thermo Fisher ScientificTM , Waltham, MA, USA) equipped with a C18 column (Acclaim PepMaplOO C18, 3 pm, 75 mm id x 500 mm, Thermo Fisher Scientific™), after loading them by partial loop injection for 5 min at a flow rate of 10 pL.min' 1 with Buffer A (5% ACN, 0.1% FA) on a preconcentration trap (Thermo Scientific™, Acclaim PepMapI 00 C18, 5 μm, 300 μm id x 5 mm). A 120 min linear gradient of 5-50% Buffer B (75% ACN and 0.1% FA) at 45°C with a flow rate of 250 nL.min' 1was implemented for separation. The end of the gradient included a 10-minute wash step with 99% Buffer B, before re-equilibration with Buffer A. This system is coupled by a nano-electrospray ion source to a Q Exactive plus mass spectrometer (Thermo Scientific™) for tandem mass spectrometry detection. A DDA acquisition method was applied, with a mass range of 350-1500 m / z at a resolution of 70,000, an AGC target of 1e6, and a maximum ion injection time of 90 ms. 10 MS / MS spectra were acquired for one MS scan (TopN), with an AGC target of 5e5, a maximum ion injection time for MS / MS of 140 ms and a resolution of 17,500. Other parameters, including the m / z isolation range, dynamic exclusion as well as the normalized collision energy, were set to 2 m / z, 30 s and 30 respectively, with a default charge state z set to 1. 2.5. Data processing and statistical analyses

[0259] The raw files corresponding to the analyzed samples were analyzed and processed using an in-house developed parallel software, Short_pept, written in Python language. Briefly, this software extracts data related to mass / charge ratio, retention time, signal intensity and spectrum number from the raw files using ThermoFisher Scientific's RawFile Reader (https: / / github.com / thermofisherlsms / RawFileReader). It calculates a library of DNS-labeled samples, which are then analyzed by the Short_pept software. It calculates a library of DNS-labeled short peptides with chemical formulas and mass / charge ratios for precursor ion identification, in addition to a fragment ion library, based on a predetermined library formed by a combination of natural and modified amino acids (25 amino acids) as well as chemical elements.

[0260] Table 3 represents the dictionary of natural and modified amino acids used for calculating short peptide precursor masses and fragment ion libraries in the Short_pept software: Table 3]

[0261] Table 4 represents the dictionary of element masses used for the composition-based library calculation in the Short_pept software: Table 4

[0262] Table 5 represents the fragmentation functions implemented for the calculation of fragment ions from MS / MS spectra in the Short_pept software: Table 5]

[0263] De novo sequencing of fragment precursor ions corresponding to the theoretical library is performed by searching their MS / MS spectra (Seidler et al., 2010. De novo sequencing of peptides by MS / MS. Proteomics, 10(4), 634-649). Corresponding information on retention time, signal intensity, and spectrum number of annotated peptides is extracted from the MS detection combined with the chromatographic profile after peak fitting using the Savitsky-Golay filter (Savitzky, et al., 1964. Smoothing and Differentiation of Data by Simplified Least Squares Procedures. Analytical Chemistry, 36(8), 1627-1639). For confident annotation of sequences, a score based on fragment ion intensities in MS / MS spectra is calculated. The maximum length of short peptides is set to 4 amino acids. The mass tolerance for peptide precursor (MS) and fragment ions (MS / MS) was set at 10 ppm.The research focused on variable modifications of lysine acetylation, methionine oxidation, and cysteine ​​carbamidomethylation. The distinction between isoleucine and leucine was made manually based on fragment ions. specific in the low m / z region (Jiang et al., 2020. A computational and experimental study of the fragmentation of l-leucine, l-isoleucine and l-allo-isoleucine under collision-induced dissociation tandem mass spectrometry. Analyst (Cambridge, II. K.), 145(20), 6632-6638). The experiments were carried out in technical triplicate. 3. Results and discussion

[0264] Generally, protein hydrolysates are characterized by the size of their macromolecules, the peptide content calculated based on total nitrogen and amino nitrogen, and the hydrolysis rate. In industry, calculating the ratio of amino nitrogen to total nitrogen is a conventional practice to obtain a rough idea of ​​protein hydrolysis (Yi et al., 2021. Estimation of Degree of Hydrolysis of Protein Hydrolysates by Size Exclusion Chromatography. Food Analytical Methods, 14(4), 805-813). As shown in Table 2 below, naturally fermented soy sauces (SH-1 and SH-2) are the most hydrolyzed substrates compared to other types of hydrolysates, with a hydrolysis rate reaching approximately 60%, regardless of the manufacturer. In this case, the result of the high hydrolysis rate converges with those of the SEC profile, where the small molecular weight fraction (< 0.5 kDa) is the most abundant representing 60-65% of the molecular profile.However, peptone from enzymatic digestion of soybean (SH-3) has a low hydrolysis rate that does not exceed 29% correlated with the percentage of the small molecular weight fraction (< 0.5 kDa) (30.5%). For SH-3, the most represented fraction in the SEC profile is that of 1-5 kDa, corresponding to oligopeptides.

[0265] The results are presented in Table 6. Table 6]: Results of global analyses applied to protein hydrolysates a The percentage shown for SEC fractions represents the relative area of ​​each fraction using UV detection at 210 nm. b The degree of hydrolysis is estimated by the ratio of amino nitrogen to total nitrogen.

[0266] For casein hydrolysates, the results clearly demonstrate a marked difference between the two products in terms of hydrolysis (Table 6). SEC profiles indicate a significantly higher fraction of small molecular weights for the pancreatic digest of casein CH-2 (64%) compared to the enzymatic hydrolysate of casein CH-1 (35%). This result is correlated with the hydrolysis rate, the values ​​of which are higher for CH-2 (40.6%) than for CH-1 (21.7%). This result can be explained by a difference in processing linked to the use of distinct enzymatic preparations, leading to two well-defined products but with contrasting compositions. For yeast extracts (YE-1 to YE-4), subtle differences were observed between the products in terms of hydrolysis rate (see Table 6), which is approximately 40% regardless of the manufacturer.However, the molecular weight (MW) distribution shown by the SEC profiles is quite dissimilar: even if the small molecular weight fraction is the most represented in all samples, this fraction is more intense in samples YE-1 (64.3%) and YE-3 (57.5%) than in YE-2 (49.5%) and YE-4 (49.8%), which suggests that these products are derived from processes that are not identical.

[0267] Although they provide insight into compositional differences with a significant representation of small peptides, these aggregate analyses cannot be correlated with the functional properties of the products and their applications. Further molecular characterization could map the disparity at a finer level, directly linked to activity.

[0268] The protein hydrolysates studied contain high amounts of small peptides (see Table 6) which could be analyzed with the method of the invention.

[0269] As detailed above, the method is based on an untargeted identification of peptide sequences of short peptides and allows in particular their relative quantification. In this study, an amine labeling of the molecules with DNS-CI, followed by an nLC-MS / MS analysis using a high-resolution mass spectrometer were implemented. The results for each of the samples SH-1, SH-2, SH-3, CH-1, CH-2, YE-1, YE-2, YE-3 and YE-4, obtained in CSV format, include the identified short peptide sequences, their corresponding retention times as well as their signal intensity (not shown here).

[0270] The results are presented in Figure 5 and Figure 6 and in Table 7 below: Table 7]: Distribution of short peptide content in various protein hydrolysates, based on signal intensities. Values ​​not included in parentheses represent the mean relative signal intensity (%) for three technical replicates per sample analyzed, while those in parentheses correspond to the standard deviation of the relative signal intensity (%) for three technical replicates per sample analyzed. Soybean hydrolysates SH-1, SH-2, and SH-3

[0271] Application of the method allowed the identification of approximately 196, 152, and 216 unique short peptides (i.e., dipeptides, tripeptides, and tetrapeptides) in SH-1, SH-2, and SH-3 soybean hydrolysates, respectively (Figure 5). Comparison of the identified sequences revealed considerable overlap with a large number of common sequences between SH-1 and SH-2 soybean hydrolysates naturally brewed for culinary applications (Figure 6A). Both products share fewer common sequences with SH-3 soybean hydrolysate obtained by enzymatic hydrolysis and commercialized for biotechnological applications. Most of the identified sequences correspond to dipeptides (90–92% of the signal intensity), followed by tripeptides (9–10% of the signal intensity) (Table 7). Tetrapeptides were detected, but they represented the minor component of soy hydrolysates (0.1-0.2% of signal intensity) (Table 7).The enzymatic hydrolysate contained more tripeptides and fewer dipeptides than the fermented soybean hydrolysates. These results converge with the results of the degree of hydrolysis which are respectively 59.7%, 60.3% and 28.8% for SH-1, SH-2 and SH-3.

[0272] Casein hydrolysates CH-1 and CH-2

[0273] For casein hydrolysates, the sequences of 309 and 310 unique short peptides (i.e., dipeptides, tripeptides, and tetrapeptides) were identified for C-1 and CH-2, respectively (Figure 5), among which 154 sequences were common to both products (Figure 6B). However, the distribution of peptides according to their length was totally different between the two casein hydrolysates (Table 7). While the enzymatically derived casein hydrolysate produced a greater proportion of dipeptides (62% of signal intensity), followed by tripeptides (30% of signal intensity) and tetrapeptides (8% of signal intensity), pancreatic casein hydrolysate resulted in a greater representation of tripeptides (52% of signal intensity), followed by dipeptides (38% of signal intensity) and tetrapeptides (10% of signal intensity). This indicates that peptones from enzymatic or pancreatic digestion of casein have a distinct short peptide composition and potentially suggests different functional properties.

[0274] YE-1, YE-2, YE-3 and YE-4 yeast extracts

[0275] For yeast extracts, the sequences of 350 to 450 unique short peptides (i.e., dipeptides, tripeptides, and tetrapeptides) were identified, highlighting compositional differences that could be related to supplier-specific processes (Figure 5). YE-1 was the extract that showed the greatest diversity in terms of short peptide sequences. It also shared the fewest common peptides with other yeast extracts, as shown in Figure 6C. All yeast extracts were primarily enriched in dipeptides (83 to 88% of signal intensity), followed by tripeptides (12 to 16% of signal intensity) and tetrapeptides (0.3 to 0.6% of signal intensity) (Table 7).

[0276] To our knowledge, these results represent the most comprehensive characterization of short peptides in hydrolyzed matrices forming complex mixtures.

[0277] From these results, further exploration and processing of the data can be carried out using other tools, for example to study the theoretical physicochemical characteristics and the theoretical biological or biochemical activities of the products studied (not described here).

[0278] Expressions such as "comprise," "include," "incorporate," "contain," "be," and "have" are to be interpreted in a non-exclusive manner when construing the description and its associated claims.

[0279] The method is not limited to the examples of embodiments described above, only by way of example, but it encompasses all the variants that may be envisaged by those skilled in the art within the framework of the claims below.

[0280] Although described through a number of detailed exemplary embodiments, the proposed method and the apparatus for implementing an embodiment of the method include various variations, modifications and improvements which will be obvious to those skilled in the art, it being understood that these various variations, modifications and improvements are part of the scope of the present disclosure, as defined by the claims which follow. In addition, various aspects and characteristics described above may be implemented together, or separately, or substituted for each other, and all various combinations and subcombinations of the aspects and features are within the scope of this disclosure. Furthermore, some systems and equipment described above may not incorporate all of the modules and functions described for the preferred embodiments.

Claims

Claims

1. A method for non-targeted identification of peptide sequences of peptides comprising between 2 and 6 amino acids present in a sample, the method being implemented by a device comprising a circuit and a memory, the memory comprising: - at least one library of predetermined peptides comprising between 2 and 6 amino acids, each predetermined peptide is defined by a peptide sequence, the library comprising for each predetermined peptide: • the amino acids constituting the peptide sequence with at least one terminal amino acid labeled with an end marker, and a total mass / charge ratio (m / z) for the peptide sequence, • fragment ions calculated from the peptide sequence, said calculated fragment ions being associated respectively with mass / charge ratios, at least one or a plurality of calculated fragment ions comprise the end marker, - sets of information associated respectively with peptides present in the sample, called measured peptides, measured by a mass spectrometer, each measured peptide comprises a peptide sequence having at least one end amino acid labeled with an end marker, each set of information of a measured peptide comprises at least one MS / MS spectrum comprising a total mass / charge ratio of the measured peptide and one or a plurality of mass / charge ratios of the measured fragment ions, the measured peptides and the predetermined peptides being labeled at the same end by the same end marker, for each set of information of a measured peptide, the method comprises: / a / a determination (110) of a selection of predetermined peptides from the library and said at least one MS / MS spectrum, each predetermined peptide of the selection having a total mass / charge ratio equal to the total mass / charge ratio of the peptide measured according to an error range less than a predetermined value, / b / a determination (120) of the labeled end amino acid of the measured peptide by comparing the mass / charge ratios of the measured fragment ions with the mass / charge ratios of the calculated fragment ions of the predetermined peptides of the selection according to an error range less than a predetermined value, Here a determination (130) of at least one reconstructed peptide sequence for the measured peptide from the calculated fragment ions of the selection, the determined labeled end amino acid and the measured fragment ions, - a rendering (140) of an identification list of peptides measured in said sample in in which each measured peptide in the list is defined by a reconstructed peptide sequence.

2. Method according to the preceding claim, the calculated fragment ions are calculated from the peptide sequence by applying a fragmentation method to the peptide sequence, preferably the fragmentation method is identical to the fragmentation method of the mass spectrometer.

3. A method according to any preceding claim, wherein the library of predetermined peptides comprises all possible peptide sequences of 2 to 6 amino acids obtained from a group of N amino acids, each sequence of the predetermined library defines a predetermined peptide of the library and further comprises a terminal amino acid which is labeled with the terminal label, preferably the library comprises at least one predetermined peptide which is declined into a plurality of charged predetermined peptides, each charged predetermined peptide having the same peptide sequence as said at least one predetermined peptide and a respective charge.

4. A method according to any one of the preceding claims, wherein the labeled end of the labeled end amino acid is A / -terminal and / or C-terminal, preferably A / -terminal.

5. A method according to any one of claims, wherein the total mass / charge ratio (m / z) of a predetermined peptide is determined from the mass / charge ratios of the amino acid residues constituting the sequence of the predetermined peptide.

6. A method according to any preceding claim, wherein the mass spectrometer is a mass spectrometer with tandem analysis capability (MS / MS), preferably coupled to liquid chromatography (LC).

7. A method according to any preceding claim, wherein in Here, said at least one peptide sequence of a measured peptide is reconstructed from: - the determined labeled end amino acid, - comparisons between the mass / charge ratios of measured fragment ions and calculated fragment ions.

8. A method according to any preceding claim, wherein the memory comprises at least a first library and a second library, the first library comprising the amino acids constituting the peptide sequence with at least one end amino acid labeled with the end tag, and the total mass / charge ratio (m / z) for the peptide sequence, and the second library being determined only for the predetermined peptides of the selection, the second library comprising at least the predetermined peptides of the selection, and comprising fragment ions calculated respectively from the peptide sequences of the predetermined peptides of the selection, the calculated fragment ions being associated respectively with mass / charge ratios, at least one or a plurality of calculated fragment ions comprise the end tag.

9. Method according to the preceding claim in combination with one of claims 2 and 3, in which the second library is determined from at least one fragmentation method applied to the predetermined peptides of the selection.

10. A method according to any preceding claim, wherein when a plurality of reconstructed peptide sequences are determined for a same peptide measured in

10. the method further comprises: Zd / calculate, for each reconstructed peptide sequence, a score from the set of information associated with the measured peptide and the reconstructed peptide sequence, ZeZ select the reconstructed peptide sequence for the measured peptide from a comparison of the scores of the reconstructed peptide sequences, with the peptide sequence with the highest score being selected.

11. Method according to the preceding claim, each mass / charge ratio of the measured fragment ions is respectively associated with an intensity in the MS / MS spectrum, and in which the score of each reconstructed peptide sequence for the same measured peptide is determined from the following equation: preferably, only measured fragment ions having said end marker are used to calculate a score of a peptide sequence.

12. A method according to any preceding claim, wherein in Zb / one or a plurality of labeled end amino acids is determined, and wherein in Here, for each determined labeled end amino acid, at least one reconstructed peptide sequence for the measured peptide is determined from the calculated fragment ions of the selection, the determined labeled end amino acid and the measured fragment ions.

13. A method according to any preceding claim, wherein the list comprises retention times respectively associated with the measured peptides, each retention time being obtained from the information set used to determine a reconstructed peptide sequence of a measured peptide.

14. A method according to any preceding claim, wherein each measured peptide in the list is associated, by an identification means, with the set of information used to determine the reconstructed peptide sequence of the measured peptide.

15. A method according to any preceding claim, wherein the list comprises intensities respectively associated with the measured peptides, each intensity being obtained from the information set used to determine a reconstructed peptide sequence of a measured peptide, and each measured peptide in the list comprises a proportion value of the peptide in the mixture, the proportion value "Qp", for each measured peptide in the list, is calculated from: Intensity of a measured peptide from list Qp= — - X of the intensities of the measured peptides from the list

16. A method according to any preceding claim, wherein the peptides are natural peptides and / or synthetic peptides.

17. A method according to any preceding claim, wherein the peptides contain at least one labeled end amino acid and one or more amino acids selected from proteinogenic amino acids or modified amino acids.

18. A method according to any preceding claim, wherein the sample is a protein hydrolyzate, preferably a yeast protein hydrolyzate, a vegetable protein hydrolyzate, or an animal protein hydrolyzate.

19. A method according to any preceding claim, wherein the labeled end amino acid is an A / -terminal amino acid labeled with an A / -terminal label selected from dansyl, dabsyl, 1-naphthyl isocyanate (IsoC), or 1-naphthyl isothiocyanate, each of which end labels may optionally be labeled with one or more carbon atoms 13 ( 13 C), deuterium (D), nitrogen 15 ( 15 N), and / or sulfur 34 ( 34 S).

20. A method according to any preceding claim, wherein the labeled end amino acid is a C-terminal amino acid labeled with a C-terminal tag selected from 2-(N,N'- dimethylamino)-1-ethylamine (DMED) and 3-(A / ,A / -dimethylamino)-1-propylamine (DMAPA), each of these end-labels optionally being labeled with one or more carbon 13 atoms ( 13 C), deuterium (D), and / or nitrogen 15 ( 15 N).

21. A method according to any preceding claim, wherein the peptides have been subjected to a protective treatment of the free thiol group, typically an alkylation treatment, the mass / charge ratios of the predetermined peptides of the library taking into account the protective treatment of the free thiol group of cysteine.

22. Method for analyzing peptides having between 2 and 6 amino acids of a sample comprising: i) a step of processing the sample comprising at least one step h) of labeling one end of the peptides of the sample by bringing the sample into contact with an end marker, ii) a step of analysis by liquid chromatography coupled to a mass spectrometer, preferably to a mass spectrometer with a tandem analysis capacity (MS / MS), iii) a step of data processing implementing a method of non-targeted identification of peptide sequences as defined in any one of the preceding claims.

23. Method according to the preceding claim, in which step i) of treating the sample further comprises a step i2) of alkylation of the peptides of the sample by contacting with an alkylating agent, and in which step i2) preferably precedes step h).

24. A method according to any one of claims 22 or 23, wherein the liquid chromatography is reverse phase liquid chromatography, preferably carried out with a gradient of at least one hour.

25. A method according to any one of claims 23 to 24, wherein the mass spectrometer is equipped with a mass analyzer preferably selected from a quadrupole trap mass analyzer or a triple quadrupole mass analyzer, more preferably is equipped with a high-resolution mass analyzer selected from an orbital trap mass analyzer, a Q-TOF mass analyzer, a TIMS-TOF mass analyzer, and an FT-ICR mass analyzer.

26. Computer program comprising instructions for implementing the method according to any one of claims 1 to 21 when this program is executed by a processor.

27. ​​Non-transitory recording medium readable by a computer on which is recorded a program for implementing the method according to any one of claims 1 to 21 when this program is executed by a processor.

Citation Information

Patent Citations

  • Compounds and methods for double labelling of polypeptides to allow multiplexing in mass spectrometric analysis

    CN101542291A

  • Method and apparatus for mass spectrometry of biomolecular samples with data independent acquisition

    CN114965728A

  • Method and apparatus for analysing samples of biomolecules using mass spectrometry with data-independent acquisition

    EP4047371A1

  • Method for identifying and characterizing a microbial population by mass spectrometry

    FR3106414A1

  • Protein expression profile database

    US20050048564A1