Methods for extracting information about protein sequence modifications
Patent Information
- Application Number
- JP2024537871
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-23
- Filing Date
- 2022-12-23
- Publication Date
- 2026-02-27
AI Technical Summary
Current methods for identifying protein sequence modifications, such as those caused by amino acid substitutions, are labor-intensive, time-consuming, and prone to human errors, with high false positive rates due to the complexity of mass spectrometry data analysis.
A computer-implemented method using at least two different enzyme digestions of protein samples, followed by mass spectrometry, to identify candidate sequence modifications with high accuracy by filtering out false positives based on peptide coverage, enzyme specificity, and other criteria.
Reduces human effort and error by automating the identification of protein sequence modifications, significantly lowering false positive rates and improving the reliability of protein analysis.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to extracting information about protein sequence modifications in proteins. [Background technology]
[0002] Complex biotechnological production methods can introduce various modifications into therapeutic proteins, potentially resulting in highly heterogeneous products. Depending on the location and type, these modifications can have a significant effect on the structure, stability, immunogenicity and biological activity of the protein. Thus, extensive characterization of therapeutic proteins is the basis for providing safe and effective medicines to patients.
[0003] A frequently used technique to identify modifications in therapeutic proteins is proteolytic digestion coupled with chromatographic peptide separation and mass spectrometry (LC-MS / MS). The proteolytic enzyme trypsin is the gold standard for this approach. Other proteases such as chymotrypsin, LysC, LysN, AspN, GluC, and ArgC are also used in proteomics, but to a lesser extent. Multienzyme strategies have been proposed to maximize sequence coverage through the use of combinations of parallel or sequential proteolytic digestions.
[0004] These techniques lead to large amounts of mass spectrometry (MS) data. This is especially true when looking for the presence of sequence variants (SVs). SVs represent amino acid substitutions in the primary structure of a protein, which can arise due to mutations and misincorporation. To identify SVs in the MS raw data, special software such as Mascot Error Tolerance Search (Matrix Science Inc.) or Byonic (Protein Metrics Inc.) can be used. These software solutions can identify unexpected mass shifts and annotate these mass shifts as modifications or sequence variants.
[0005] Because there are multiple possibilities for the SV of each amino acid in a peptide sequence, the probability of a random match identified by the software MS / MS algorithm (a "false positive hit") is high compared to regular database searches for chemical modifications or post-translational modifications (PTMs) such as glycans.
[0006] Distinguishing true positives from large numbers of false positives is a challenge: this validation method is currently performed manually, can take days or even weeks to complete, and is prone to human error.
[0007] In a typical sequence variant analysis experiment, sample preparation, sample analysis by LC-MS / MS instrumentation, and software searching takes approximately 2-3 days, but the subsequent manual "hit validation" by examining various criteria such as retention time, mass accuracy, and MS / MS spectra can take days or even weeks.
[0008] It is an object of the present invention to provide improved methods for extracting information about protein sequence modifications, such as SVs. Summary of the Invention
[0009] According to one aspect of the present invention, there is provided a computer-implemented method for extracting information about protein sequence modifications in a protein, the computer-implemented method comprising: receiving protein data derived at least in part from mass spectrometry measurements performed on peptides obtained by at least two different enzymatic digests performed on each sub-sample of a representative sample of the protein; identifying candidate sequence modifications in the protein using the received protein data; determining a subset of candidate sequence modifications that represent true sequence modifications with a higher average probability than the remaining candidate sequence modifications; and outputting data representing the determined subset of candidate sequence modifications, wherein determining the subset of candidate sequence modifications comprises selecting candidate sequence modifications depending on the candidate sequence modifications being at amino acid sequence positions that are covered by at least two different peptide species, each of which contains the modification.
[0010] Thus, a method is provided for identifying a subset of candidate sequence modifications that is more likely to be accurate via a computer-automated procedure, thereby saving human effort / time and / or reducing error. Receiving protein data from peptides obtained by at least two different enzymatic digests increases the reliable coverage of the protein sequence and provides a basis for additional filtering criteria that may further reduce false positives, as described below.
[0011] In one embodiment, determining the subset of candidate sequence modifications comprises selecting candidate sequence modifications dependent on the candidate sequence modifications occurring at amino acid sequence positions where the ratio of the number of peptide species covering the amino acid sequence position and containing the candidate sequence modification to the number of peptide species covering the amino acid sequence position and not containing the candidate sequence modification is equal to or greater than a predefined threshold ratio. This criterion excludes candidate modifications in peptide species where the corresponding modified peptide species occurs relatively rarely compared to the corresponding unmodified (wild type) peptide species. The inventors have found this filtering approach to be effective in reducing false positives.
[0012] In one embodiment, the method includes pre-processing the protein data to identify data associated with a selected subset of peptide species, and excluding the selected subset of peptide species from use in determining the subset of candidate sequence modifications. Pre-processing further improves performance.
[0013] Pre-processing may include excluding peptide species for which a) a candidate modification is present and b) the highest intensity in the mass spectrometric measurement of the corresponding peptide species ("wild type") with the same amino acid sequence but without the candidate modification is below a predefined threshold intensity. This filtering approach is based on the recognition that some peptides are less intense than others due to physicochemical properties, which are defined by their corresponding sequences. Modifications to such sequences typically do not completely change the ionization properties. Thus, low intensity "wild type" peptides tend to be associated with low intensity (and therefore relatively unreliably identified) modified peptides. Thus, excluding such peptide species from subsequent analysis effectively contributes to reducing false positives.
[0014] Pre-processing may include filtering out peptide species for which a) a candidate modification is present, and b) the highest quality score in mass spectrometry measurements of corresponding peptide species with the same amino acid sequence but without the candidate modification is below a predefined threshold score. The quality score represents the degree of matching between theoretical fragments and observed fragments in mass spectrometry measurements. This filtering approach is based on the recognition that low quality scores of wild-type peptide species indicate that the information obtained from each peptide species with modifications is relatively unreliable. Thus, filtering out such peptide species from subsequent analysis effectively contributes to reducing false positives.
[0015] Pre-processing may include excluding each peptide species that has a candidate modification at the cleavage site of the enzymatic digestion that generated the peptide species. This filtering approach is based on the recognition that modifications can affect the digestion behavior of an enzyme, meaning that peptide species with a start or end point that corresponds to the position of a modification are not optimal for detecting that modification. Thus, excluding these peptide species from subsequent analysis contributes to reducing false positives.
[0016] Pre-processing may include filtering out peptide species having a length greater than a predefined threshold length. This filtering approach is based on the recognition that modifications identified in long peptide species are generally more difficult to verify and therefore less reliable. Thus, filtering out peptide species longer than a predefined threshold length is effective in reducing false positives.
[0017] In one embodiment, determining the subset of candidate sequence modifications comprises selecting candidate modifications reliant on the candidate sequence modifications being at amino acid sequence positions that are covered by at least two different peptide species derived from peptides that each contain the modification and that were obtained using different enzymatic digests. The inventors have found this approach to be highly effective at eliminating false positives.
[0018] According to a further aspect of the present invention there is provided a computer-implemented method for extracting information about protein sequence modifications in a protein, comprising: receiving protein data derived at least in part from mass spectrometry measurements performed on peptides obtained by at least two different enzymatic digestions performed on each sub-sample of a representative sample of the protein; identifying candidate sequence modifications in the protein using the received protein data; determining a subset of candidate sequence modifications which represent true sequence modifications with a higher average probability than the remaining candidate sequence modifications; and outputting data representative of the determined subset of candidate sequence modifications, wherein determining the subset of candidate sequence modifications comprises selecting a candidate sequence modification depending on the candidate sequence modification being at an amino acid sequence position where a ratio of the number of peptide species covering the amino acid sequence position and containing the candidate sequence modification to the number of peptide species covering the amino acid sequence position and not containing the candidate sequence modification is equal to or greater than a predefined threshold ratio.
[0019] Thus, a method is provided for identifying a subset of candidate sequence modifications through a computer-automated procedure, which is more likely to be accurate by saving human effort / time and / or reducing error. Filtering out candidate modifications according to the determined ratio filters out candidate modifications that occur relatively rarely in the corresponding modified peptide species compared to the corresponding unmodified (wild type) peptide species. The inventors have found that this filtering approach is effective in reducing false positives.
[0020] In some embodiments, the received protein data is derived from mass spectrometry measurements performed on peptides obtained by 5 or 6 different enzymatic digests performed on each subsample of a representative sample of proteins, preferably each of the 5 or 6 enzymatic digests using a different one of the following: trypsin, thermolysin, AspN, pronase, pepsin, ProAlanase. The inventors have found that these particular number of digestions provides an advantageous balance of high sensitivity against few false positives.
[0021] In some embodiments, the method further comprises: identifying one or more groups of peptide species in the received protein data, each group of peptide species containing only peptide species that all have the same candidate sequence modification, the candidate sequence modification being in the determined subset of candidate sequence modifications and different in each group; and outputting data representing the peptide species in each of the identified groups. Identifying groups of peptide species that all have the same candidate sequence modification allows information to be presented to the user in a more organized manner, facilitating efficient assessment of candidate sequence modifications. This approach helps users avoid the duplication of assessing the same candidate sequence modification multiple times in different peptide species.
[0022] According to a further aspect of the present invention there is provided a computer-implemented method for extracting information about protein sequence modifications in a protein, the method comprising: receiving protein data derived at least in part from mass spectrometry measurements performed on peptides obtained by at least two different enzymatic digests performed on each sub-sample of a representative sample of the protein; identifying candidate sequence modifications in the protein using the received protein data; determining a subset of the candidate sequence modifications which represent true sequence modifications with a higher average probability than the remaining candidate sequence modifications; and outputting data representative of the determined subset of candidate sequence modifications, wherein determining the subset of candidate sequence modifications comprises selecting the candidate sequence modifications in dependence on the candidate sequence modifications satisfying a quantification condition, the quantification condition indicating that the amount detected by mass spectrometry measurements of at least a selected subset of peptide species having the candidate sequence modification exceeds a predefined quantification threshold compared to the total amount detected by mass spectrometry measurements of the same peptide species with and without the candidate sequence modification.
[0023] Thus, methods are provided for identifying a subset of candidate sequence modifications through computer-automated procedures, which are more likely to be accurate by saving human effort / time and / or reducing error. Eliminating candidate modifications based on whether a quantification condition is met has been found to be particularly effective in reducing false positives.
[0024] In one embodiment, the at least two different enzymatic digestions include one or more sequence-specific enzymatic digestions and one or more non-specific enzymatic digestions, and the selected subset of peptide species is selected for at least the candidate sequence modifications covered by at least one peptide species derived using sequence-specific enzymatic digestion, so as to exclude the peptide species derived using one or more non-specific enzymatic digestions. It has been found that preferentially or exclusively considering peptide species derived from sequence-specific enzymatic digestion provides a particularly high performance and allows further reduction of false positives.
[0025] Embodiments of the present disclosure will now be further described, by way of example only, with reference to the accompanying drawings, in which: [Brief description of the drawings]
[0026] [Figure 1] 1 is a flowchart depicting a method for extracting information about a protein sequence. [Diagram 2] 1 shows a schematic depiction of the coverage of a portion of a protein sequence by peptide species derived from different digests. [Diagram 3] Data demonstrating the performance of the disclosed method are shown in Figure 3. Figure 3 is based on samples from an actual product development project. [Figure 4] Data demonstrating the performance of the disclosed method are shown in Figure 4. Figure 4 is based on artificially generated samples with known amounts of sequence variance arbitrarily "spiked" into the samples. [Diagram 5] Data demonstrating the performance of the disclosed method are shown in Figure 5. Figure 5 is based on artificially generated samples with known amounts of sequence variance arbitrarily "spiked" into the samples. [Figure 6] Data demonstrating the performance of the disclosed method are shown in Figure 6. Figure 6 is based on artificially generated samples with known amounts of sequence variance arbitrarily "spiked" into the samples. [Figure 7] Data demonstrating the performance of the disclosed method are shown in Figure 7. Figure 7 is based on artificially generated samples with known amounts of sequence variance arbitrarily "spiked" into the samples. [Figure 8] Data demonstrating the performance of the disclosed method are shown in Figure 8. Figure 8 is based on artificially generated samples with known amounts of sequence variance arbitrarily "spiked" into the samples. [Figure 9] 13A-13C are schematic depictions of examples of detection amounts of peptide species in samples with and without candidate modifications of different digestion and charge states. [Figure 10] Shown are data demonstrating the performance of the disclosed method using a quantification threshold, data derived using artificially generated samples with known amounts of sequence variance arbitrarily "spiked" into the samples. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0027] Various embodiments of the present disclosure relate to computer-implemented methods. Each step of the method of the present disclosure can be performed by a computer in the most general sense of the term, which means any device capable of performing the data processing steps of the method, including dedicated digital circuitry. The computer can include various combinations of known computer elements, including, for example, a CPU, RAM, SSD, motherboard, network connections, firmware, software, and / or other elements known in the art that allow the computer to perform the necessary computational operations. The necessary computational operations can be defined by one or more computer programs. The one or more computer programs can be provided in the form of a medium or data carrier, optionally a non-transitory medium that stores computer-readable instructions. When the computer-readable instructions are read by the computer, the computer performs the necessary method steps. The computer can be a self-contained unit, such as a general-purpose desktop computer, a laptop, a tablet, a mobile phone, or other smart devices. Alternatively, the computer can be a distributed computing system with several different computers connected to each other via a network, such as the Internet or an intranet.
[0028] An embodiment of the present disclosure relates to extracting information about protein sequence modifications in proteins. An exemplary method framework is depicted diagrammatically in Figure 1 and described below.
[0029] The method starts with the preparation of a representative sample of the protein to be analyzed (step S1). The protein can be, for example, a therapeutic protein. The sample can be provided in any of a variety of forms known in the art. For example, a typical protein sample can be obtained directly from the supernatant of a cell culture, recovered by cell discard, by solubilization and renaturation of inclusion bodies of bacterial cells, or by extraction from tissues. The sample can also be subjected to further mechanical or chemical purification steps, such as filtration, diafiltration, dialysis, centrifugation, precipitation, or chromatography. In one embodiment, the protein sample is essentially free of other proteins, i.e. contains less than 20%, optionally less than 10%, optionally less than 5%, optionally less than 2%, optionally less than 1%, optionally less than 0.5%, optionally less than 0.2% of other proteins. Typically, a sample of approximately 300-500 μg can be prepared.
[0030] The sample may be initially processed as a single sample. For example, the sample may be subjected to general digestion (step S2), e.g., enzymatic deglycosylation using PNGase. It may be desirable to remove glycans, as this would result in additional fragments that would make translation of the peptide pattern more difficult.
[0031] In step S3, the representative sample is divided into sub-samples, and each sub-sample is subjected to a different enzymatic digestion. Each digestion uses a different enzyme or a different combination of enzymes. Other conditions may also vary between the different digestions, such as the time the digestion process is allowed to proceed before being stopped. Typically, each digestion uses a single enzyme, but the use of a combination of enzymes in a single sub-sample may sometimes be appropriate (e.g., to obtain peptides of a suitable length for a subsequent mass spectrometry step). In the arrangements of the present disclosure, at least two different enzymatic digestions are used (for the two corresponding sub-samples). In some arrangements, the number of enzymatic digestions is high, for example, at least three, optionally at least four, optionally at least five, optionally at least six, optionally at least seven, optionally at least eight, optionally at least nine. In one arrangement, the number of enzymatic digestions is between five and nine, preferably five or six. The processing of the sub-samples through the different digestions is depicted diagrammatically in FIG. 1 by boxes labeled "digestion 1", "digestion 2", etc. In principle, any number N of such digestions can be performed.
[0032] Each enzymatic digestion uses a different one or combination of the following: trypsin; thermolysin; AspN; elastase; chymotrypsin; LysC; LysN; GluC; ArgC; pronase; pepsin; proalanase. In one particular embodiment, all nine of the following digestions are used: trypsin only; thermolysin only; AspN only; elastase only; chymotrypsin only; a combination of LysC+GluC (e.g., in a 1:20 ratio); pronase only; pepsin only; proalanase only. The digestions may be allowed to proceed for 0.5 to 4 hours, depending on, for example, the enzymes used.
[0033] The enzymes used in each digestion belong to the class of proteases and drive proteolysis, which is the breakdown of proteins by cleaving peptide bonds into smaller polypeptides. The location of cleavage depends on the protein being digested and the enzyme or combination of enzymes present. Thus, each digestion produces a different population of peptide species.
[0034] In step S4, the peptides obtained by digestion are processed by mass spectrometry measurements. The outputs of the individual digestions can be processed separately from each other or in combination. In the example shown in FIG. 1, the output of digestion 1 (peptides obtained by applying digestion 1 to the subsample) is processed by a first mass spectrometry method MS, the output of digestion 2 (peptides obtained by applying digestion 2 to the subsample) is processed by a second mass spectrometry method MS, and so on. In some arrangements, each mass spectrometry method comprises liquid chromatography tandem mass spectrometry (LC-MS / MS). LC-MS / MS is a well-known analytical chemistry technique for analyzing peptide species. The liquid chromatography column is connected to the ion source of the mass spectrometry system, which allows the components of the sample separated by liquid chromatography to be directly fed to the mass spectrometry system. To obtain sequence information corresponding to the peptide m / z signals, the mass spectrometry system operates in tandem mode (MS / MS) to obtain comprehensive information about the sample composition. The components received from the liquid chromatography column are ionized and subsequently separated according to their mass-to-charge ratio in the first mass spectrometer. The separated ions are then split into smaller fragment ions, sometimes referred to as peptide fragments, which are separated and detected by a second mass spectrometer operation (e.g., by a second mass analysis step, either by the same or a different mass spectrometer).
[0035] The subsequent step S5 may be computer implemented.
[0036] In a computer-implemented process, each peptide species may be defined by at least: i) an amino acid sequence; ii) an enzymatic digestion that generated the peptide species; and iii) a modification state indicating whether a candidate modification is present, and if so, the nature and amino acid sequence location of the modification. In some arrangements, each peptide species is further defined by iv) a charge state in a mass spectrometry measurement.
[0037] In step S5, protein data derived at least in part from the mass spectrometry measurements of step S4 are received. Thus, the protein data are derived from mass spectrometry measurements performed on peptides obtained by the different enzymatic digestions applied to each subsample. The protein data can take any of the various forms known in the art that represent information about the peptides analyzed in this way. For example, the protein data may include information obtained by comparing the measured masses of peptides (MS) or peptide fragments (MS / MS) with the predicted theoretical values of the corresponding peptide species with and without modifications. Special software such as Mascot Error Tolerance Search (Matrix Science Inc.) or Byonic (Protein Metrics Inc.) may be used. These software solutions are able to identify unexpected mass shifts and annotate these mass shifts as sequence modifications.
[0038] In step S6, the protein data received from step S5 is used to identify candidate sequence modifications in the protein. This can be achieved by analyzing the protein data to identify unexpected mass shifts as candidate sequence modifications, or the protein data may already be annotated with this information, as described above. The resulting candidate sequence modifications usually need to be manually reviewed and checked for validity (i.e., to reduce the number of false positives). The steps of the method described below replace manual review with automatic review, or greatly reduce the number of candidate modifications that require manual review.
[0039] In step S7, a subset of candidate sequence modifications is determined that represents true sequence modifications with a higher average probability than the remaining candidate sequence modifications. The determined subset thus represents a list of modifications that are likely to be true modifications. The list is obtained via a computer-implemented process, thus reducing or avoiding the need for manual review. The list can be output (step S8) to the user according to user selection (e.g., as an output data stream or file, or as instructions on a computer display). The determination of the subset of candidate sequence modifications includes filtering based on coverage of amino acid positions by peptide species (of peptides derived by digestion), optionally based on coverage by peptide species derived from different enzymatic digestions.
[0040] The coverage of the amino acid positions of an exemplary segment of a protein sequence is depicted diagrammatically in FIG. 2, where the horizontal lines below the sequence (labeled 10) represent the different peptide species. The peptide species are grouped into groups 11-17 according to which enzymatic digestion was used to obtain them. Thus, group 11 represents the portion of peptide species obtained using the first enzymatic digestion, group 12 represents the portion of peptide species obtained using the second enzymatic digestion, group 13 represents the portion of peptide species obtained using the third enzymatic digestion, and so on. As can be seen, the coverage by the different peptide species varies according to the position along the amino acid sequence 10. At position A, for example, coverage is provided by peptide species of groups 11, 12, 13, 16 and 17 (with no coverage for groups 14 and 15), while at position B, coverage is provided by peptide species of groups 11-15 and 17 (with no coverage for group 16).
[0041] In the arrangement, the determination of the subset in step S7 involves selecting candidate sequence modifications depending on the presence of the candidate sequence modifications at amino acid sequence positions covered by at least two different peptide species (derived from peptides obtained by at least two enzymatic digestions), each of which contains the modification. Thus, candidate sequence modifications at positions that are not covered by a sufficiently large number of different peptide species (from the same or different digestions) carrying the modification are excluded. Thus, the determination step S7 acts as a filter that excludes candidate modifications based at least on how the corresponding sequence positions are covered by the peptide species. A filter setting corresponding to this type of filtering is, for example, referred to below as "noPep". The observation of the same modification in many different peptide species indicates a higher probability that the modification is a genuine modification. Thus, the removal of modifications present in only a single peptide species (or a small number of different peptide species) effectively reduces false positives (i.e., candidate sequence modifications that are not genuine). The effectiveness of this filtering approach is demonstrated by the data shown in Figures 4 and 5, described below. As described below, step S7 may additionally include filters based on other exclusion criteria.
[0042] Filtering based on coverage by different peptide species can be made more stringent, thereby increasing the exclusion of false positives. An optimal balance may be struck between the reliability of eliminating false positives and avoiding or minimizing the exclusion of true positives. In some arrangements, the filtering is strengthened to exclude candidate sequence modifications that are not covered by at least three, optionally at least four, optionally at least five, and optionally at least six different peptide species carrying the modification. The effect of such strengthened filtering is shown in FIG. 5B.
[0043] In some embodiments, the determination of the subset in step S7 involves selecting candidate sequence modifications depending on the candidate sequence modifications being at amino acid sequence positions that are covered by at least two different peptide species derived from peptides that each contain the modification and that were obtained using different enzymatic digestions. Thus, candidate sequence modifications at positions that are not covered by peptide species derived from at least two different enzymatic digestions are eliminated. Thus, the determination step S7 acts as a filter that eliminates candidate modifications based at least on how the corresponding sequence positions are covered by peptide species derived from different enzymatic digestions. A filter setting corresponding to this type of filtering is, for example, referred to as "number of digestions" below. This approach is highly effective at removing false positives, since such false positives are mainly random and unlikely to be present in peptide species generated by multiple different digestions. On the other hand, true positives are found in peptides generated by different enzymes. The effectiveness of this filtering approach is demonstrated by the data shown in FIG. 3, described below.
[0044] Filtering based on coverage by different enzyme digests can be made more stringent, thereby increasing the exclusion of false positives. An optimal balance may be struck between the reliability of eliminating false positives and avoiding or minimizing the exclusion of true positives. In some arrangements, filtering is enhanced to exclude candidate sequence modifications that are not covered by peptide species with at least three, optionally at least four, and optionally at least five different enzyme digests. The effect of such enhanced filtering is shown in FIG. 5A.
[0045] Pretreatment before step S7 In some arrangements, the protein data is pre-processed prior to the analysis of step S7. Pre-processing of the protein data includes identifying data associated with a selected subset of peptide species and excluding the selected subset of peptide species from use in determining the subset of candidate sequence modifications of step S7.
[0046] In some arrangements, a peptide species is excluded if a) a candidate modification is present, and b) the highest intensity (peak height of the apex) in the mass spectrometry measurement of a corresponding peptide species having the same amino acid sequence but without the candidate modification is below a predefined threshold intensity. A filter setting corresponding to this type of filtering is, for example, referred to below as "WT intensity". The low intensity in the mass spectrometry measurement of such a "wild type" peptide species (i.e., a peptide species without a modification) indicates that the information obtained from the corresponding peptide species with the modification is relatively unreliable. In essence, the mass spectrometry measurement is relatively inefficient in measuring the peptide species with this particular sequence (with or without the modification). In other words, some peptides are less intense than others due to the physicochemical properties defined by their respective sequences. Modifications to such sequences typically do not completely change the ionization properties. Thus, low intensity "wild type" tends to be associated with low intensity (and therefore relatively unreliable) modified peptides. Thus, excluding such peptide species from subsequent analysis contributes to effectively reduce false positives. This is demonstrated by the data presented in Figure 7, discussed below.
[0047] In some arrangements, a peptide species is excluded if a) a candidate modification is present, and b) the highest precision score in the mass spectrometry measurement of the corresponding peptide species having the same amino acid sequence but without the candidate modification is below a predefined threshold score. A filter setting corresponding to this type of filtering is, for example, referred to below as "WT score". "Precision score" in this context refers to the result of applying an algorithm that compares a theoretical fragmentation pattern with a fragmentation pattern generated by a mass spectrometry measurement. The precision score can be a metric that quantifies the degree of matching / correlation of theoretical and observed fragments. A higher score indicates a higher correlation (better matching). Similar to the case of low intensity discussed above, a low precision score for a wild-type peptide species indicates that the information obtained from the corresponding peptide species with the modification is relatively unreliable. Therefore, excluding such peptide species from subsequent analysis contributes to effectively reduce false positives.
[0048] In some arrangements, pre-processing includes excluding each peptide species that has a candidate modification at the cleavage site of the enzyme that generated the peptide species. A filter setting corresponding to this type of filtering is, for example, referred to below as "exclude cleavage sites". Modifications can affect the digestion behavior of an enzyme, meaning that peptide species that have a start or end point that corresponds to the position of a modification are not optimal for detecting that modification. Thus, excluding these peptide species from subsequent analysis contributes to reducing false positives.
[0049] In some arrangements, preprocessing involves filtering out peptide species with lengths above a predefined threshold length. Filter settings corresponding to this type of filtering are, for example, referred to below as "peptide length". Modifications specified for long peptide species are generally more difficult to validate and therefore less reliable. Long peptide species (e.g., >3500 Da) are typically highly charged (4-6). Highly charged peptides typically cause poor MS / MS coverage. The reason for this is that highly charged peptides are accelerated much faster towards the collision cell of the mass analyzer, which results in fewer fragments compared to less charged peptide species. It is possible to obtain "good" MS / MS data (i.e., high fragment ion coverage, high intensity) from highly charged peptides using different fragmentation modes. However, such fragmentation modes (e.g., electron transfer dissociation (ETD)) tend to result in lower MS / MS scores than alternatives (e.g., collision-induced dissociation (CID)) for small peptide species that typically have less charge. Therefore, excluding peptide species longer than a predefined threshold length is effective in reducing false positives.
[0050] Other selection criteria in step S7 In some arrangements, determining the subset of candidate sequence modifications in step S7 comprises selecting the candidate sequence modifications in dependence on the candidate sequence modifications being present at amino acid sequence positions where the ratio of the number of peptide species covering the amino acid sequence position and containing the candidate sequence modification to the number of peptide species covering the amino acid sequence position and not containing the candidate sequence modification is equal to or greater than a predefined threshold ratio. A filter setting corresponding to this type of filtering is, for example, referred to below as a "ratio filter". This criterion filters out candidate modifications in peptide species where the corresponding modified peptide species occurs relatively rarely compared to the corresponding unmodified (wild type) peptide species. The effectiveness of this filtering approach is demonstrated by the data shown in FIG. 6 described below. In some embodiments, the minimum ratio (known threshold ratio) is set in the range of 2-10%, preferably in the range of 2-5%, preferably in the range of 2-4%, preferably in the range of 2.5-3.5%, preferably about 3%.
[0051] Experiments to demonstrate performance Experiments demonstrating performance are described below with reference to "filter settings," which define configuration options for the selection or discounting of candidate sequence modifications to obtain a subset of candidate sequence modifications output by the method. Filter settings include the following, some of which are discussed above:
[0052] "SV Score" - a precision score that represents the degree of correlation between theoretical and observed fragments (peptide species) in mass spectrometric measurements of fragments containing a candidate modification.
[0053] "WT Score" - a precision score that represents the degree of correlation between theoretical and observed fragments (peptide species) in mass spectrometry measurements of fragments that do not contain the candidate modification.
[0054] "ppm" - the mass accuracy of the mass spectrometry measurement.
[0055] "MS1 Corr" - a metric that represents the best MS1 correlation (comparing theoretical and measured isotope patterns) for all identified peptide species with the same sequence, digest type, charge state, modification and modification position. A score of 1 is a perfect match.
[0056] "Peptide length" - the number of amino acids in a peptide species.
[0057] "WT Intensity" - indicates the highest intensity (apex peak height) in the mass spectrometry measurement for all identifications of unmodified peptide species corresponding to modified peptide species containing the candidate modification, having the same digest type and considering all charge states present.
[0058] "Exclude Cleavage Sites" - a "yes" or "no" setting that indicates whether candidate sequence modifications that correspond to enzyme cleavage sites are excluded.
[0059] "Ratio Filter" - the ratio of the number of peptide species that cover an amino acid sequence position and contain the candidate sequence modification to the number of peptide species that cover an amino acid sequence position and do not contain the candidate sequence modification.
[0060] "noPep" - the number of peptide species that cover the amino acid sequence position of the candidate sequence modification and contain the modification.
[0061] "Number of digests" - the number of peptide species covering the amino acid sequence position of the candidate sequence modification, containing the modification, and derived from peptides obtained using different enzymatic digests.
[0062] "XIC Ratio" - indicates the minimum XIC ratio for the candidate sequence modification using the ratio of the XIC area of the candidate sequence modification to the XIC area of the corresponding peptide species that does not contain the modification.
[0063] FIG. 3 is a table depicting the application of the method to seven different product development projects.
[0064] To generate the data shown in Figure 3, purified protein samples (usually 350 μg) from seven different projects were denatured in 8 M guanidine hydrochloride at pH 7.0, reduced by addition of DTT (dithiothreitol) and incubated for 1 h at 37°C. S-carboxymethylation of the reduced samples was performed by addition of iodoacetyl acid. Before digestion with several enzymes, the buffer was exchanged into digestion buffer (50 mM Tris, 2 mM CaCl2, pH 7.5) using a NAP5 column. The samples were divided into nine equal fractions and the different enzymes were added. The digestion conditions were enzyme dependent. The following conditions were used:
[0065] [Table 1]
[0066] The resulting digests were separated by RP-LC coupled to a Thermo Fisher Scientific Orbitrap Fusion mass spectrometer (data-dependent setup). Separation was performed using a 120 min gradient (mobile phase A: water with 0.1% v / v formic acid (FA); mobile phase B: acetonitrile with 0.1% v / v FA) on an ACQUITY UPLC CSH C18 column (water, 130 Å, 1.7 μm, 2.1 mm×150 mm). Data-dependent setup was used for the Fusion-Orbitrap.
[0067] The first column (Projekt) in FIG. 3 contains a number representing the project. The second column represents the number of candidate sequence modifications identified in step S6 of the method. The third column shows the number of sequence modifications in the subset determined in step S7. The numbers in parentheses indicate the reduction rate of false positives, e.g., 9 / 2447=0.4% in the first row means that 99.6% of the hits were filtered out. Filter settings were as follows: WT intensity>1e6; SV score>140; WT score>220; peptide length: 5-32; MS1 Corr>0.95; ppm<4; number of digests≧2 (data applied by Orbitrap Fusion LC-MS / MS data-dependent top time; CID in IonTrap; 120 min gradient, 9 digests). These results demonstrate that the method can confidently identify true positives from a large number of candidate sequence modifications, as well as confidently filter out all, or substantially all, false positives from such a large number of candidate sequence modifications. Note that the data in Figure 3 is not an artificially created data set by spiked-in sequence modifications. The data in Figure 3 is derived from the analysis of samples from an actual project using the method. Since it is unknown what SVs are actually present in the sample, no results on sensitivity (i.e., true positives) can be provided in this case. In contrast to the spike-in data underlying Figures 4-8 (discussed below), here it is possible to assess how many known and expected SVs in the sample are identified using the method.
[0068] 4-8 and 10 are bar graphs showing the results of experiments demonstrating the performance of embodiments of the present disclosure.
[0069] To quantitatively determine whether all sequence modifications present in the samples could be detected, nine different test samples (1-9), each containing a first antibody, were spiked by adding a second antibody to the solution at a 0.5% level, as shown in Table 2. An additional set of nine test samples (10-18) was generated by spiking the second antibody with the first antibody at a 0.5% level.
[0070] [Table 2]
[0071] A total of 18 such spiked samples were generated, whose digestion and subsequent analysis using RP-LC coupled to an Orbitrap Fusion mass spectrometer were performed as described above for FIG.
[0072] In these 18 samples, a theoretical total of 140 positions were expected to differ by only one amino acid between the first and second antibodies in otherwise identical sequences from the position to ≥7 amino acids N- and C-termini, mimicking naturally occurring sequence variants. The resulting LC-MS / MS files were processed into protein data by using the Software-Tool Byonic (Protein Metrics, San Carlos, CA / USA). This software compared the experimental data (peptide mass (MS) and peptide fragments (MS / MS)) with the theoretical data of the amino acid sequence of the corresponding "primary antibody" (i.e., the antibody present at 99.5 percent in the sample) generated in silico. The results were ranked and visualized by the software Byologic (Protein Metrics) using the calculated dataset generated by Byonic. The protein data of the 18 samples exported from Byologic were pooled and then analyzed. By comparing the number of correctly determined sequence variants in spiked samples with the theoretically expected sequence variants, it was possible to quantitatively assess the impact of individual filter settings on the overall sensitivity (true positives) of the method.
[0073] In Figures 4-8 and 10, in each case the bars are presented in pairs, each pair including a solid white bar on the left and a shaded bar on the right. The height of each solid white bar represents the sensitivity % and corresponds to the scale presented on the left vertical axis. Sensitivity is defined as the ratio between the obtained true candidate sequence modifications and the calculated sum of true candidate sequence modifications based on the sequence of the antibody used for spiking. The height of each shaded bar represents the total number of false positives and corresponds to the scale presented on the right vertical axis. False negatives are candidate sequence modifications that do not correspond to the true sequence modifications introduced by spiking and are presented in absolute numbers. In Figures 4B and 5-8, the number of false positives represents the number of false positive candidate sequence modifications. In Figure 4A, the number of false positives corresponds to the number of peptide species with false positive sequence modifications, because the method used to generate these results (unlike the method of the present disclosure) does not include a step of pooling different peptide species with the same candidate modification into one hit.
[0074] 4A and 4B are graphs depicting the results of comparative experiments demonstrating improved performance compared to alternative approaches. The graphs depict a pair of bars for each of six different configurations (referred to as Settings A through F, respectively).
[0075] Settings A-C (shown in FIG. 4A) refer to configurations where the disclosed method is not used. In each case, trypsin is used as the primary enzyme. In the configuration of setting A, which may be referred to as "trypsin_old", only trypsin is used to obtain peptides for mass spectrometry. In the configuration of setting B, which may be referred to as "3 enzymes_old", trypsin was used first, with Asp-N and thermolysin used as alternative enzymes to fill gaps in the sequence coverage where the trypsin "wild type" peptide was not detected. Only the most potent wild type peptide of the alternative enzymes was used to fill each gap. In the configuration of setting C, which may be referred to as "9 enzymes_old", the approach of setting B was expanded to use eight alternative enzymes / enzyme mixtures in addition to trypsin to fill gaps in the sequence coverage where the trypsin "wild type" peptide was not detected. Only the most potent wild type peptide of the alternative enzymes was used to fill each gap. The eight alternative enzymes / enzyme mixtures were Asp-N, thermolysin, chymotrypsin, Glu-C / Lys-C, pronase, pepsin, proalanase, and elastase. In settings A-C, sequence variants were filtered at >0.1% of the XIC ratio. Other filter settings were ppm=±4ppm, size=750-3100 daltons, precision score>140, and intensity>1e7.
[0076] Settings D-F (Setting D is shown in Figure 4A, Settings E and F are shown in Figure 4B) refer to configurations in which nine enzymes / enzyme mixtures (trypsin, Asp-N, thermolysin, chymotrypsin, Glu-C / Lys-C, pronase, pepsin, proalanase, and elastase) are used equally. All sequence variants corresponding to wild-type peptides that were retained after filtering were used.
[0077] In the configuration of Set D, which may be referred to as "9 enzymes_new", all nine enzymes were used equally (not only the most potent wild-type peptides of the alternative enzymes were used to fill their respective gaps), but filtering according to embodiments of the present disclosure was not applied. The "wild-type peptides" were prefiltered using the following prefiltering settings: ppm = ± 4 ppm, size = 750-3100 daltons, precision score > 140, intensity > 1e6.
[0078] In the configuration of setting E, which may be referred to as "9 enzymes_new+filter 1", all nine enzymes were used with filtering according to embodiments of the present disclosure. Filtering was performed with the following settings to ensure that no sequence variants were missed: SV score > 140, WT score > 140 ppm ± 4 ppm, MS1 Corr > 0.95, peptide length = 5-32 amino acids, WT intensity > 1e6, exclude cleavage sites = "yes", ratio filter > 2%, number of digests > 1, noPep > 2, XIC ratio > 0.1%.
[0079] In configuration setting F, which may be referred to as "9 enzymes_new+filter 2", all nine enzymes were used with filtering according to embodiments of the present disclosure. Filtering was performed with the following settings to achieve sensitivity similar to the old method (i.e., similar to setting B), but with a reduction in false positives of >90%: SV score >180, WT score >260 ppm ± 3 ppm, MS1 Corr >0.96, peptide length = 7-32 amino acids, WT intensity >1e6, excluded cleavage sites = "yes", ratio filter >10%, noPep >2, number of digests >1, XIC ratio >0.1%.
[0080] 4A and 4B that setting D achieves high sensitivity, but also high false positives, while settings E and F demonstrate that embodiments of the present disclosure provide an improved balance between sensitivity and false positives. Setting E achieves the same high sensitivity as setting D, but has significantly fewer false positives. Setting F achieves similar high sensitivity as either of the older approaches represented by settings A-C, but has significantly fewer false positives (>92% reduction).
[0081] FIG. 5A is a graph depicting the results of a further experiment demonstrating how sensitivity and false positives vary in the method of the disclosed embodiment as a function of the required coverage of an amino acid sequence by different peptide species, each of which contains a modification and is each derived from peptides obtained using a different enzymatic digestion (corresponding to the filter setting "number of digests"). Before the filtering step, the total number of predicted modifications (including true and false positive hits) was 25694. The first pair of bars, labeled "1", corresponds to the case where a minimum coverage by a single peptide species is required for candidate selection. The second pair of bars, labeled "2", corresponds to the case where a minimum coverage by two peptide species from two different digestions is required for candidate selection. The third pair of bars, labeled "3", corresponds to the case where a minimum coverage by three peptide species from three different digestions is required for candidate selection. The fourth pair of columns, labeled "4", corresponds to the case where the minimum coverage is by four peptide species from four different digestions. Other filter settings were held constant in each case and were as follows: SV score >140, WT score >140 ppm ± 4 ppm, MS1 Corr >0.95, peptide length = 5–32 amino acids, WT intensity >1e6, exclude cleavage sites = “yes”, ratio filter >0%, noPep ≥1, XIC ratio >0%.
[0082] It can be seen from Figure 5A that increasing the minimum coverage with different peptide species derived from different enzymatic digests rapidly reduces false positives while having a limited adverse effect on sensitivity. Requiring a minimum coverage of two peptides results in a significant reduction in false positives and no measurable change in sensitivity. Requiring a minimum coverage of two peptides derived from two different digests provides a good balance between high sensitivity and few false positives.
[0083] FIG. 5B is a graph depicting the results of a further experiment demonstrating how sensitivity and false positives vary in the method of the disclosed embodiment as a function of the required coverage of an amino acid sequence by different peptide species (corresponding to the filter setting "noPep"). Before the filtering step, the total number of predicted modifications (including true and false positive hits) was 25694. The first pair of bars labeled "1" corresponds to the case where a minimum coverage by a single peptide species is required for candidate selection. The second pair of bars labeled "2" corresponds to the case where a minimum coverage by two peptide species is required for candidate selection. The third pair of bars labeled "3" corresponds to the case where a minimum coverage by three peptide species is required for candidate selection. The fourth pair of columns labeled "4" corresponds to the case where a minimum coverage by four peptide species is required for candidate selection. Other filter settings were kept the same in each case and were as follows: SV score >140, WT score >140 ppm ± 4 ppm, MS1 Corr >0.95, peptide length = 5–32 amino acids, WT intensity >1e6, exclude cleavage sites = “yes”, ratio filter >0%, number of digests ≥1, noPep ≥1, XIC ratio >0%.
[0084] It can be seen from FIG. 5B that increasing the minimum coverage with different peptide species rapidly reduces false positives while having a limited adverse effect on sensitivity. Requiring a minimum coverage of two peptides results in a significant reduction in false positives and no measurable change in sensitivity. Requiring a minimum coverage of three peptides provides a good balance between high sensitivity and few false positives.
[0085] FIG. 6 is a graph depicting the results of a further experiment demonstrating how sensitivity and false positives vary in the method of the disclosed embodiment as a function of the minimum value of the ratio (corresponding to the filter setting "ratio filter") between the number of peptide species covering the amino acid sequence position and containing the candidate sequence modification and the number of peptide species covering the amino acid sequence position and not containing the candidate sequence modification. The first pair of bars corresponds to a ratio value that must be at least 1%. The second to fifth pairs of bars represent 2%, 3%, 5% and 10% of the minimum ratio value, respectively. Other filter settings were kept the same in each case and were as follows: SV score > 140, WT score > 140 ppm ± 4 ppm, MS1 Corr > 0.95, peptide length = 5 to 32 amino acids, WT intensity > 1e6, excluded cleavage sites = "yes", number of digests > 1, noPep > 2, XIC ratio > 0%.
[0086] It can be seen from Figure 6 that increasing the value of the minimum ratio results in a quicker decrease in false positives and a slower significant decrease in sensitivity. As in Figures 5A and 5B, the total number of predicted modifications (including true and false positive hits) before the filtering step here was 25694. Increasing the value of the minimum ratio from 1% to 2% significantly reduces false positives while not significantly affecting sensitivity. Increasing the value of the minimum ratio to 3% results in a good balance between high sensitivity and very few false positives. Higher values of the minimum ratio result in very few false positives while keeping sensitivity at a useful level.
[0087] FIG. 7 is a graph depicting the results of a further experiment demonstrating how sensitivity and false positives vary in the method of the disclosed embodiment as a function of the required minimum intensity of the mass spectrometry measurement (corresponding to the filter setting "WT intensity"). Pairs of bars are shown, each corresponding to a minimum intensity increasing from 1e+05 to 1e+08 from left to right. Other filter settings were kept the same in each case and were as follows (no ratio filter): SV score > 140, WT score > 140 ppm ± 4 ppm, MS1 Corr > 0.95, peptide length = 5-32 amino acids, excluded cleavage sites = "yes", ratio filter > 0%, number of digests ≥ 1, noPep ≥ 2, XIC ratio > 0.1%.
[0088] It can be seen from Figure 7 that increasing the value of the minimum intensity results in a quicker decrease in false positives and a slower significant decrease in sensitivity. As in Figures 5 and 6, the total number of predicted modifications (including true and false positive hits) before the filtering step here was 25694. Increasing the value of the minimum intensity from 1e+05 to 1e+06% significantly reduces false positives while not significantly affecting sensitivity. Increasing the value of the minimum intensity to 1e+07 results in a good balance between high sensitivity and very few false positives. A high value of the minimum intensity results in very few false positives while keeping the sensitivity at a useful level.
[0089] 8 is a graph depicting the results of a further experiment demonstrating how sensitivity and false positives vary in the method of the disclosed embodiments as a function of increasing numbers of enzymes. Pairs of bars corresponding to enzyme groups containing from 2 to 9 enzymes are shown as follows: Group i = trypsin, pepsin; group ii = trypsin, pepsin, proalanase; Group iii = trypsin, thermolysin, pepsin, proalanase; Group iv = trypsin, thermolysin, pronase, pepsin, proalanase; Group v = trypsin, thermolysin, AspN, pronase, pepsin, proalanase; Group vi = trypsin, thermolysin, AspN, elastase, pronase, pepsin, proalanase; Group vii = trypsin, thermolysin, AspN, elastase, GluC, pronase, pepsin, proalanase; Group viii = trypsin, thermolysin, AspN, elastase, chymotrypsin, GluC, pronase, pepsin, proalanase.
[0090] As a control, the sensitivity and false positives of only one enzyme (pepsin) are shown in group ix. Of the nine used for enzymatic digestion, pepsin showed the highest sensitivity when using the filter settings described below. Therefore, pepsin was used here as a comparison instead of the gold standard trypsin used for enzymatic digestion.
[0091] The same filter settings were applied in each case and were as follows: SV score >140, WT score >140 ppm ± 4 ppm, MS1 Corr >0.95, peptide length = 5–32 amino acids, WT intensity >1e6, excluded cleavage sites = “yes”, ratio filter >2%, number of digests ≥1, noPep ≥2, XIC ratio >0.1%.
[0092] It can be seen from Figure 8 that increasing the number of enzymes increases the sensitivity, but also increases the false positives. As in Figures 5-7, the total number of predicted modifications (including true and false positive hits) before the filtering step here was 25694 when all nine enzymes were used. Compared to the control using only data from a single digestion, a fairly good sensitivity can already be achieved by using two enzymes. A particularly good balance between high sensitivity and a low number of false positives is achieved with 5-6 enzymes, while a higher number of enzymes achieves even better sensitivity with moderate false positives.
[0093] In some embodiments, the method includes identifying one or more groups of peptide species in the received protein, each group of peptide species containing only peptide species that all have the same candidate sequence modification. The candidate sequence modifications are in a determined subset of the candidate sequence modifications and are different in each group. Thus, the grouping method can be performed after any combination of the filtering steps described above. The method can include outputting data representing which peptide species are in each of the identified groups. The data can include a list of the peptide species in each identified group. The data can be adapted for display as a graph or other non-text based representation. The method can be performed to provide selectable options to define one or more groups. The selectable options can include a specification of the filter settings to be applied (e.g., corresponding to WT intensity > 1e6, SV score > 140, peptide length 5-32, MS1 score > 0.95, etc.). The selectable options can include a specification of the modification (e.g., an alanine to serine SV exchange). The selectable options can include a specification for the position of the modification. A position can include specification of either or both the chain (eg, light chain LC) and the amino acid (eg, amino acid 25).
[0094] In some embodiments, the determination of the subset in step S7 includes selecting a candidate sequence modification depending on whether the candidate sequence modification satisfies a quantification condition. The quantification condition generally requires that the detected relative amount of a peptide species having the candidate sequence modification (i.e., relative to the total detected amount of the corresponding peptide species with and without the modification) should be relatively high. The inventors reasoned that the satisfaction of such a quantification condition may strongly indicate that the candidate sequence modification is a genuine sequence modification, and thus provide a useful basis for filtering. The filter setting corresponding to this type of filtering may be referred to as "Quant" in this specification.
[0095] The challenge in effectively implementing such a filter is to formulate an appropriate metric to accurately represent the detected relative abundance of peptide species bearing the candidate sequence modification. The complexity of the situation is illustrated diagrammatically in Figure 9, which depicts an exemplary result of a mass spectrometry measurement.
[0096] FIG. 9 depicts peptide species as boxes, with and without a candidate sequence modification. The right column of boxes represents peptide species with the candidate modification (schematically indicated by a black vertical bar in each box in the right column). The left column of boxes represents the corresponding peptide species without the candidate modification (i.e., wild-type peptide species). Nine rows (a)-(i) are presented, with each row containing a pair of corresponding peptide species (with and without the candidate sequence modification). Within each box, the charge state is indicated by Z2 (doubly charged), Z3 (triple charged), or Z4 (quadruple charged). The portion of the time-intensity curve in the mass spectrometry measurement corresponding to each peptide species is represented by counts. *The area is shown after the "area" with s. The boxes marked "nd" correspond to undetected peptide species (with modifications). The rows correspond to different peptide species or different charge states of peptide species. In the example shown, the peptide species are derived from three different digests (designated Dig1, Dig2 and Dig3). In each of Dig1 and Dig2, a single peptide species in two different charge states (z2 and z3) is shown (labeled Pep1 in Dig1 and Pep2 in Dig2). Rows (a) and (b) correspond to Pep1 from Dig1, and rows (c) and (d) correspond to Pep2 from Dig2. In Dig3, two peptide species are shown (labeled Pep3 and Pe4). Rows (e) and (f) correspond to Pep3 in charge states z2 and z3. Rows (g), (h) and (i) correspond to Pep4 in charge states z2, z3 and z4. In this particular example, Dig1 was a chymotrypsin digest and Pep1 was a peptide species corresponding to the amino acid sequence range 613-629. Dig2 was a trypsin digest and Pep2 was a peptide species corresponding to the amino acid sequence range 608-623. Dig3 was a LysC+GluC digest and Pep3 was a peptide species corresponding to the amino acid sequence range 604-620 and Pep4 was a peptide species corresponding to the amino acid sequence range 604-623 derived from the same LysC+GluC digest.
[0097] In principle, various metrics can be formulated to express the detected relative amounts of peptide species bearing sequence modifications.
[0098] Examples of metrics contemplated by the inventors are described below and are referred to as Metrics 1 to 7. The determined value of each metric may be compared to a pre-defined quantification threshold to perform Quant filtering (i.e. to determine whether the subset from step S7 contains a candidate sequence modification).
[0099] In an exemplary arrangement, the mass spectrometry measurement, which may be liquid chromatography tandem mass spectrometry (LC-MS / MS), outputs a time-intensity curve. For a given peptide species, the amount of detection of the peptide species correlates strongly to both the maximum (peak intensity) of the curve portion corresponding to the peptide species and the area under the curve portion corresponding to the peptide species. In the metric discussion below, when discussing the amount of detection of an individual peptide species, reference is made to "area". The method can also be performed using the maximum (peak intensity) instead of the area, or indeed any other suitable parameter extracted from the mass spectrometry measurement that correlates to the amount of detection of the peptide species.
[0100] The percentages listed in column 21 of FIG. 9 represent the ratio, expressed as a percentage, of the area of the peptide species in the right column (i.e., with the modification) to the sum of the peptide species in the right and left columns (i.e., with and without the modification) in each row. The percentages in column 21 thus provide information relevant to determining the relative detection abundance of peptide species with a candidate sequence modification. However, it can be seen that the percentages vary significantly with size. The percentages listed in column 22 represent the ratio (expressed as a percentage) of the sum of the areas of all charge states with the modification (i.e., the area of all boxes in the right column of the peptide species considered) to the sum of the areas of all charge states of the peptide species with and without the modification (i.e., the sum of the areas of all boxes in the right and left columns of the peptide species considered). Thus, for example, for Pep1, a value of 0.15% is 100%×4.3×10 6 / (4.3×10 6 +4.4×10 8 +2.4×10 9 ) Again, significant variation is seen in the percentages in column 22.
[0101] For metrics 1 and 3, only the most intense (largest area) of all wild type corresponding peptide species carrying the modification was used to calculate the relative amount of the modification. Thus, the metric is calculated using only one of the rows of FIG. 9 (i.e. one of the values in column 21). In the case of metric 1, the selection is limited to peptide species originating from one selected digestion. Typically, the selected digestion is a trypsin digestion, but this is not essential. In the case of metric 3, peptide species originating from all digestions are considered.
[0102] In one example, the digestion selected for metric 1 can be Dig2 (trypsin). One peptide species (Pep2) with two charge states was derived using Dig2 in the example of FIG. 9. According to metric 1, the maximum area wild type corresponding to the peptide species carrying the Pep2 modification is in row (c) because no charge state corresponding to row (d) is detected in the peptide species with the modification (i.e., the box in the right column is "nd"). Metric 1 is thus 0.33% in this example.
[0103] If the digest selected for metric 1 is Dig3 instead of Dig2, the maximum area wild type corresponding to the peptide species carrying the modification is that in row (f), which corresponds to the z3 charge state of Pep3. Metric 1 is therefore 0.13% in this example.
[0104] In metric 3, peptide species from all digests are considered, such that the maximum area wild type corresponding to the peptide species carrying the modification used in calculating metric 3 is also that of row (f) since this is row (f) that has the maximum area wild type corresponding to the peptide species carrying the modification in all digests. Row (b) has the maximum area wild type, but the version with the modification (right column) was not detected ("nd"), so row (b) is not used.
[0105] Metric 2 is a variation of Metrics 1 and 3, where all charge states of peptide species with modifications are considered, regardless of whether an individual charge state has a modification. Metric 2 is then calculated according to the method described in the documentation for column 22. Metric 2 is the percentage of column 22 that corresponds to the peptide species that contains the most area wild type (considering all charge states). Metric 2 may consider only peptides derived from the selected digest, or peptide species derived from all digests. If the selected digest is Dig2, the output of Metric 2 is 0.07%, since only one peptide species is derived from Dig2. If the selected digest is Dig3, the output of Metric 2 is 0.30%, since Pep3 is the peptide species derived using Dig3 that contains the most area wild type. If all digests are considered, the output of Metric 2 is 0.15%, since overall Pep1 has the most area wild type (row (b)).
[0106] For metrics 4-7, the metrics are calculated based on a combination of information about the areas of peptide species derived from multiple different digests.
[0107] Metric 4 is derived from multiple (e.g., all) digests, but uses pre-filtering to determine whether the area of the wild-type peptide species (left column) is greater than or equal to a threshold (e.g., 10 7 count * s) (i.e., all rows in FIG. 9 where the right hand column is not "nd"). Thus, only rows where the left hand column area exceeds the threshold and the right hand column is not "nd" are considered. The metric is then calculated as the average of the corresponding percentages of columns 21. Thus, if the threshold area is 10 7 count * s, then all rows that do not have "nd" in the right column are considered. Metric 4 is the average of the values in column 21 of rows (a), (c), (e), (f) and (h), yielding 0.66%. The threshold area is 3×10 8count * When raised to s, row (e) is filtered out and metric 4 becomes instead the average of the values in column 21 of rows (a), (c), (f) and (h), yielding 0.52%.
[0108] Metrics 5-7 use a methodology referred to herein as weighted quantification. In this type of approach, information about the area is combined so that the quantification metric can be expressed as a percentage according to the formula: TIFF2025502710000004.tif15128 where Σ Area of Modification is the sum of the areas of the modified peptide species and Σ Area of Wild Type is the sum of the areas of the wild type peptide species that correspond to the modified peptide species (correspondence also requires corresponding charge states). In each case, only the charge states in which the detected peptide species with the modification is present are considered. Thus, in the example of FIG. 9, rows (b), (d), (g) and (i) have no contribution.
[0109] Metric 5 considers peptide species from all digests. Thus, in the example of Figure 9, the output of Metric 5 considers rows (a), (c), (e), (f) and (h). The output of Metric 5 is the sum of all areas of the right columns of rows (a), (c), (e), (f) and (h) divided by the sum of all areas of both columns of rows (a), (c), (e), (f) and (h), which is equal to 0.47%.
[0110] In Metric 6, only peptide species derived from sequence-specific enzymes are considered. In the example of FIG. 9, this results in the exclusion of peptide species derived from Dig1, which was performed using chymotrypsin. The output of Metric 6 is the sum of all areas of the right columns of rows (c), (e), (f) and (h) divided by the sum of all areas of both columns of rows (c), (e), (f) and (h), which is equal to 0.39%.
[0111] Metric 7 uses a combination of the approaches of metrics 5 and 6 to provide more complete coverage. Thus, only sequence-specific enzyme digestion is used where there is coverage by sequence-specific enzyme digestion and gaps are filled by non-specific enzyme digestion. In other words, for candidate modifications for which sequence-specific enzyme digestion provides coverage, sequence-specific enzyme digestion is used as described above for metric 6. If the modification is not covered by any sequence-specific enzyme digestion, non-specific enzyme digestion must be used, which means that all peptide species available for this candidate modification and protein are used as described above for metric 5. This case is special because only non-specific enzyme digestion is used. Metric 5 can use sequence-specific and non-specific enzyme digestion.
[0112] The effectiveness of the various metrics was tested using a sample spiked with 140 sequence differences at a rate of 1%. The results are shown in Table 1 below. The digestion chosen for metrics 1 and 2 was trypsin. The threshold for metric 4 was 10 7 count * It was set to s.
[0113] [Table 1]
[0114] The best performing methods are found to be those based on metrics 5, 6 and 7, which all achieve very low average deviation from 1% of the target. Metric 6 achieves better performance than metric 5 for an increase in "5x too low", but worse for "SV(%)". Metric 7 achieves the best overall performance.
[0115] Based on the above insight, the quantification condition is configured in such a way that the amount detected by mass spectrometry of at least a selected subset of peptide species having the candidate sequence modification exceeds a predefined quantification threshold relative to the total amount detected by mass spectrometry of the same peptide species (which may be multiple peptide species, if the subset includes multiple peptide species) with and without the candidate sequence modification. In the configuration, the selected subset includes multiple or all peptide species (e.g., represented by metrics 5, 6 or 7 discussed above) derived from multiple or all of the at least two different enzymatic digestions used. Thus, the quantification condition is: TIFF2025502710000006.tif14128 Using the formula above, or a similar or equivalent formula, eg, a corresponding value, the metric can be calculated and compared to a predefined quantification threshold.
[0116] In some arrangements, the at least two different enzymatic digestions include one or more sequence-specific enzymatic digestions and one or more non-specific enzymatic digestions. This is the case in the example illustrated in FIG. 9. In such an arrangement, the selected subset of peptide species may be selected to exclude peptide species derived using one or more non-specific enzymatic digestions, as in metric 6. Thus, the selected subset may consist of peptide species derived from multiple sequence-specific enzymatic digestions. For candidate sequence modifications that are not covered by at least one peptide species derived using sequence-specific enzymatic digestion, the selected subset of peptide species may be selected to include peptide species derived using one or more non-specific enzymatic digestions. Thus, gaps that do not have specific enzymatic digestion coverage can be filled using non-specific enzymatic digestion (e.g., metric 7). Alternatively, all relevant peptide species may be included, and no subset of peptide species is selected based on the nature of the enzymatic modification used (e.g., metric 5).
[0117] Sequence-specific enzymatic digestion within the meaning of the present disclosure may be digestion performed by at least one proteolytic enzyme (protease) that cleaves the protein N-terminus or C-terminus of a specific amino acid or sequence of adjacent amino acids in a protein sequence in a predictable manner, for example, trypsin cleaves the protein C-terminus of amino acid K or R in a protein amino acid sequence. Other enzymatic digestions, i.e., enzymatic digestions that are not sequence-specific, may be referred to as non-specific enzymatic digestion in the present disclosure. Cleavage sites are created when the non-specific enzymatic digestion used may be less predictable or unpredictable, but the digestion of a particular protein using a particular protease is reproducible.
[0118] In the configuration, the sequence-specific enzymatic digestion comprises or consists of one or more of the following enzymes: trypsin, endoproteinase ASpN, endoproteinase LysC, endoproteinase GluC.
[0119] In the configuration, the non-specific enzymatic digestion comprises or consists of one or more of the following enzymes: thermolysin, elastase, pronase, proalanase, pepsin, chymotrypsin.
[0120] FIG. 10 is a graph depicting the results of a further experiment demonstrating how sensitivity and false positives vary in the method of the present disclosure embodiment as a function of increasing number of quantification thresholds (corresponding to filter setting "Quant" and using metric 7 described above). Pairs of bars are shown corresponding to quantification thresholds (Quant) of 0, 0.1, 0.2, 0.3 and 0.5, respectively, increasing from left to right. Increasing the quantification threshold improves the suppression of false positives, but may also reduce sensitivity. The quantification threshold may be selected according to requirements, and the above values are merely for illustration. In some embodiments, the quantification threshold is selected to be equal to or greater than 0.05, 0.1, 0.2 or 0.3. In some embodiments, the quantification threshold is additionally (as required) or alternatively selected to be equal to or less than 0.6, 0.5, 0.4, 0.3, 0.2 or 0.1. Other filter settings were kept the same in each case and were as follows: SV score >140, WT score >140ppm ± 4ppm, MS1 Corr >0.95, peptide length = 5-32 amino acids, WT intensity >1e6, exclude cleavage sites = "yes", ratio filter >2%, noPep ≥ 2. Hit selection for each candidate modification (defined by protein and type of modification) was performed using only modified peptide species resulting from sequence-specific enzymatic digestion (i.e., Asp-N, trypsin and GluC+LysC in this case). In cases where no specific enzymatic digestion was available, all hits from non-specific enzymatic digestion were considered for quantification.
[0121] It can be seen from Figure 10 that increasing the Quant filter setting results in a fast reduction in false positives and a tolerable slower reduction in sensitivity. For example, with a Quant filter setting of 0.1, there was a useful 33% (from 3586 to 2405) reduction in the number of false positives, and only a 1.4% (from 100.0% to 98.6%) reduction in the true positives obtained. The results obtained with Quant filter setting = 0 (no Quant filter) already correspond to the results of applying the filters described above with reference to Figures 5B, 6 and 7: - at least two peptides per modification (Figure 5B), -Ratio filter >2% (Figure 6), - Only high intensity WT (>1e6) (Figure 7).
[0122] Due to this filtering, the number of false positives is already highly reduced even before applying the Quant filter. In this context, the observed improvement of 33% is particularly significant.
Claims
1. 1. A computer-implemented method for extracting information about protein sequence modifications in a protein, comprising: receiving protein data derived at least in part from mass spectrometry measurements performed on peptides obtained by at least two different enzymatic digestions performed on each sub-sample of the representative sample of proteins; using said received protein data to identify candidate sequence modifications in said protein; determining a subset of said candidate sequence modifications that represent true sequence modifications with a higher average probability than the remaining said candidate sequence modifications; and outputting data representing said determined subset of candidate sequence modifications. Including, said determining said subset of candidate sequence modifications selecting a candidate sequence modification depending on the location of said candidate sequence modification in an amino acid sequence position that is covered by at least two different peptide species, each of which contains said modification; Including, The computer-implemented method.
2. Each peptide species contains at least: i) amino acid sequence; ii) the enzymatic digestion that produced the peptide species; and iii) a modification status indicating whether the candidate modification is present, and if so, the nature and amino acid sequence location of said modification; The method of claim 1 , wherein:
3. Each peptide species is as follows: iv) the charge state in the mass spectrometry measurement 3. The method of claim 2, further defined by:
4. pre-processing the protein data to identify data associated with a selected subset of peptide species, and excluding the selected subset of peptide species from use in determining the subset of candidate sequence modifications; The method according to any one of claims 1 to 3, comprising:
5. The pretreatment a) a peptide species in which a candidate modification is present, and b) the highest intensity in said mass spectrometry measurement of a corresponding peptide species having the same amino acid sequence but without said candidate modification is below a predetermined threshold intensity. The method of claim 4, comprising excluding:
6. The pretreatment a) a peptide species in which a candidate modification is present, and b) the highest accuracy score in said mass spectrometry measurement of a corresponding peptide species having the same amino acid sequence but without said candidate modification is less than a predetermined threshold score. including excluding The accuracy score represents the degree of match between the theoretical fragments and the observed fragments in the mass spectrometry measurement. The method of claim 4.
7. The pretreatment eliminating each peptide species having a candidate modification at the cleavage site of the enzymatic digestion that generated the peptide species; and / or filtering out peptide species having lengths exceeding a predetermined threshold length; The method of claim 4, comprising:
8. said determining said subset of candidate sequence modifications selecting a candidate sequence modification depending on whether said candidate sequence modification is located at an amino acid sequence position that is covered by at least two different peptide species, each of which contains said modification and is derived from peptides obtained using different enzymatic digestions; The method according to any one of claims 1 to 3, comprising:
9. said determining said subset of candidate sequence modifications selecting a candidate sequence modification depending on whether the candidate sequence modification is located at an amino acid sequence position where the ratio of the number of peptide species covering the amino acid sequence position and containing the candidate sequence modification to the number of peptide species covering the amino acid sequence position and not containing the candidate sequence modification is equal to or exceeds a predetermined threshold ratio; The method of claim 1 , comprising:
10. 1. A computer-implemented method for extracting information about protein sequence modifications in a protein, comprising: receiving protein data derived at least in part from mass spectrometry measurements performed on peptides obtained by at least two different enzymatic digestions performed on each sub-sample of the representative sample of proteins; using said received protein data to identify candidate sequence modifications in said protein; determining a subset of said candidate sequence modifications that represent true sequence modifications with a higher average probability than the remaining said candidate sequence modifications; and outputting data representing said determined subset of candidate sequence modifications. Including, said determining said subset of candidate sequence modifications selecting a candidate sequence modification depending on whether the candidate sequence modification is located at an amino acid sequence position where the ratio of the number of peptide species covering the amino acid sequence position and containing the candidate sequence modification to the number of peptide species covering the amino acid sequence position and not containing the candidate sequence modification is equal to or exceeds a predetermined threshold ratio; Including, The computer-implemented method.
11. The method according to claim 9 or 10, wherein the predetermined threshold ratio is in the range of 2 to 10%.
12. said determining said subset of candidate sequence modifications selecting a candidate sequence modification dependent on said candidate sequence modification satisfying a quantification condition; Including, the quantification condition indicates that the amount detected by the mass spectrometry measurement of at least a selected subset of peptide species having the candidate sequence modification exceeds a predetermined quantification threshold relative to the total amount detected by the mass spectrometry measurement of the same peptide species with and without the candidate sequence modification. The method of claim 1.
13. 1. A computer-implemented method for extracting information about protein sequence modifications in a protein, comprising: receiving protein data derived at least in part from mass spectrometry measurements performed on peptides obtained by at least two different enzymatic digestions performed on each sub-sample of the representative sample of proteins; using said received protein data to identify candidate sequence modifications in said protein; determining a subset of said candidate sequence modifications that represent true sequence modifications with a higher average probability than the remaining said candidate sequence modifications; and outputting data representing said determined subset of candidate sequence modifications. Including, said determining said subset of candidate sequence modifications selecting a candidate sequence modification dependent on said candidate sequence modification satisfying a quantification condition; Including, the quantification condition indicates that the amount detected by the mass spectrometry measurement of at least a selected subset of peptide species having the candidate sequence modification exceeds a predetermined quantification threshold relative to the total amount detected by the mass spectrometry measurement of the same peptide species with and without the candidate sequence modification. The computer-implemented method.
14. 14. The method of claim 12 or 13, wherein the selected subset comprises some or all of the peptide species from some or all of the at least two different enzymatic digestions.
15. 14. The method of claim 12 or 13, wherein the at least two different enzymatic digestions comprise one or more sequence-specific enzymatic digestions and one or more non-specific enzymatic digestions.
16. 16. The method of claim 15, wherein the selected subset of peptide species is selected for candidate sequence modifications covered by at least one peptide species derived using sequence-specific enzymatic digestion, to exclude peptide species derived using the one or more non-specific enzymatic digestions.
17. 17. The method of claim 16, wherein the selected subset of peptide species is selected to include peptide species derived using one or more non-specific enzymatic digestions for candidate sequence modifications not covered by at least one peptide species derived using sequence-specific enzymatic digestion.
18. 17. The method of claim 16, wherein the sequence-specific enzymatic digestion comprises one or more of the following: trypsin, endoproteinase AspN, endoproteinase LysC, endoproteinase GluC.
19. 17. The method of claim 16, wherein the non-specific enzymatic digestion comprises one or more of the following: thermolysin, elastase, pronase, proalanase, pepsin, chymotrypsin.
20. 14. The method of claim 12 or 13, wherein the mass spectrometry measurement comprises liquid chromatography tandem mass spectrometry.
21. For each peptide species, the amount detected by the mass spectrometry measurement is the area under the time-intensity curve portion in the mass spectrometry measurement corresponding to the peptide species; or a maximum value in the time-intensity curve portion in the mass spectrometry measurement corresponding to the peptide species; 21. The method of claim 20, wherein:
22. 14. The method of any one of claims 1 to 3, 9, 10, 12, and 13, wherein each enzymatic digestion uses a different one or combination of the following: trypsin; thermolysin; AspN; elastase; chymotrypsin; LysC; LysN; GluC; ArgC; pronase; pepsin; proalanase.
23. 14. The method of any one of claims 1 to 3, 9, 10, 12 and 13, wherein the received protein data is derived from mass spectrometry measurements performed on peptides obtained by 5 or 6 different enzymatic digestions performed on sub-samples of each of the representative samples of the proteins, preferably each of the 5 or 6 enzymatic digestions using a different one of the following: trypsin, thermolysin, AspN, pronase, pepsin, proalanase.
24. identifying one or more groups of peptide species in the received protein data, each group of peptide species containing only peptide species that all have the same candidate sequence modification, the candidate sequence modification being within the determined subset of candidate sequence modifications and different for each group; and outputting data representing peptide species in each of said identified groups; The method of any one of claims 1 to 3, 9, 10, 12, and 13, further comprising:
25. A computer program, or a computer readable medium or data carrier signal carrying said computer program, comprising: The computer program comprises instructions that, when the program is executed by a computer, cause the computer to perform the method according to any one of claims 1 to 3, 9, 10, 12 and 13. said computer program or a computer readable medium or data carrier signal carrying said computer program.