Compound assembly
The method improves mass spectral data processing by grouping features with consistent retention times and applying isotope removal and fragmentation analysis to enhance compound identification accuracy and reliability.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- THERMO FISHER SCI BREMEN
- Filing Date
- 2024-07-19
- Publication Date
- 2026-06-02
AI Technical Summary
Existing methods for processing mass spectral data are inadequate in accurately assembling mass-to-charge ratio signals, which include isotopes, solvent adducts, homodimers, different charge states, and fragments, complicating qualitative and quantitative analysis.
A method for processing mass spectral data involves detecting groups of features with specific retention times, applying isotope removal algorithms, and identifying compounds through MS2 or higher-order fragmentation spectra, using chromatographic separation and mass spectrometry to form clusters of features with consistent retention times and masses, resolving conflicts through candidate ion types and relationships.
This approach enhances the reliability and consistency of compound identification by accurately grouping features and resolving conflicts, improving the interpretation of complex mass spectral data.
Smart Images

Figure 0007869266000001 
Figure 0007869266000002 
Figure 0007869266000003
Abstract
Description
[Technical Field]
[0001] This invention relates to the field of mass measurement, and more particularly to a method for processing mass spectral data. [Background technology]
[0002] A considerable amount of time has passed since the discovery of soft ionization techniques, which enabled the in-situ mass measurement of small molecules, peptides, and proteins. These techniques severely limit fragmentation within the source, greatly simplifying the obtained spectra. However, a single compound still manifests as a series of mass-to-charge ratio (m / z) signals. These signals include not only isotopes, but also solvent adducts, homodimers or heterodimers, different charge states, and any remaining fragments within the source. Accurately "assembling" these multiple m / z signals is a crucial step in both qualitative and quantitative analysis.
[0003] There is still room for improvement in the methods for processing mass spectral data. [Overview of the project]
[0004] A method for processing mass spectral data for a sample is provided. The mass spectral data is obtained from multiple MS 1 Mass spectra and multiple MSs N Includes mass spectra, multiple MS 1 Mass spectra and multiple MSs N Each mass spectrum has its own associated retention time. The method involves multiple MS 1 The method includes detecting groups of features in a mass spectrum, wherein each feature in the group has its own mass, and the features in the group have a corresponding retention time. The method also includes identifying one or more compounds in a sample based on the group of features.
[0005] As will be described in more detail below, various embodiments provide an improved method of processing mass spectral data.
[0006] In multiple embodiments, the mass spectral data is generated by an analytical instrument that includes a chromatographic separation device (such as a liquid chromatography (LC) separation device or a gas chromatography (GC) separation device) and a mass spectrometer. The chromatographic separation device can separate the sample in a chromatographic separation scan, and the mass spectrometer can acquire the mass spectral data during the chromatographic separation scan. Thus, multiple MS 1 mass spectra and multiple MS N Each mass spectrum of the mass spectra will have its respective associated retention time, that is, each mass spectrum is acquired at its respective chromatographic retention time during the chromatographic separation scan.
[0007] As will be further described below, in some embodiments, each MS N mass spectrum is an MS2 mass spectrum. However, each MS N can instead be a higher-order fragmentation spectrum such as an MS3 spectrum. Generally, N is an integer ≥ 2.
[0008] The mass spectral data may be composed of (i.e., formed from) at least one sample file, and each sample file is data output from the mass spectrometer from a respective chromatographic separation scan. The mass spectral data of each sample file includes multiple MS 1 mass spectra out of multiple MS 1 mass spectra, and multiple MS N mass spectra out of multiple MS NThis may include mass spectra. Mass spectral data may include a single such sample file or multiple sample files. If multiple sample files exist, each sample file may be obtained from chromatographic separation scans for different fractions of the same sample.
[0009] The method involves multiple MS 1 The method involves detecting groups of features within a mass spectrum. Each feature in the group may have its own mass and associated retention time. Features in a group may have corresponding (e.g., equal within a first tolerance) retention times but may have different masses. The method involves multiple masses. 1 This may involve detecting one or more further groups of features within the mass spectrum, each distinct group having its own distinct retention time. Processing only a single group of features is described in detail below, but it will be understood that each further group of features can be processed in a similar manner.
[0010] In some embodiments, each group of features initially comprises multiple MS 1 Features are detected in the mass spectrum, and then the detected features with corresponding (i.e., equal within a first tolerance) retention times are grouped together. The features are then first analyzed in the MS of each sample file. 1 By detecting file-specific features within the mass spectrum, and then grouping the file-specific features (feature-per-file) across sample files with corresponding mass and retention times to form a feature, multiple MS 1 It can be detected in the mass spectrum.
[0011] Therefore, multiple MS 1 The step of detecting groups of features in the mass spectrum is: For each sample file, the characteristics of each individual file within that sample file are detected, and these characteristics include their respective mass and retention time. The process involves forming multiple features from the features of each file, where each feature has its own mass and retention time. This may include forming a group of features by grouping features that have corresponding retention times.
[0012] In these embodiments, the characteristics of each file in the sample file are first the multiple MS of the sample file 1 This can be detected by constructing a chromatogram for each intrinsic mass-to-charge ratio (m / z) within the mass spectrum and determining the characteristic retention time for each chromatogram. The characteristic retention time for a chromatogram may be the retention time at the center or peak of the chromatogram and can be determined, for example, using a peak detection algorithm or similar. The chromatograms can then be grouped into sets according to their characteristic retention times, and an isotope removal algorithm can be applied to each set of chromatograms to form a set of features per file.
[0013] Therefore, the characteristics of each file are isotope-detoxified MS, each having a unique combination of mass and retention time. 1 These are features that appear in the mass spectral data. Then, each file-specific feature in each group of file-specific features will have its own mass and its own retention time, and each file-specific feature in each group will have a corresponding (i.e., the same within the second tolerance) retention time.
[0014] Similarly, the step of detecting the features of each of the multiple files within the sample file is: Multiple MS files in the sample files 1The process involves constructing multiple chromatograms from a mass spectrum, each chromatogram having its own mass-to-charge ratio (m / z). To determine the characteristic retention times for each chromatogram, Grouping chromatograms with corresponding characteristic retention times into one or more sets of chromatograms, This may include applying an isotope removal algorithm to each set of chromatograms to form a group of features for each file.
[0015] As described above, features within each group of features may have equal retention times within a first tolerance range, while features within each group of file-specific features may have equal retention times within a second tolerance range. The second tolerance range may be smaller than the first tolerance range. As will be further explained below, this has the effect of adequately explaining the retention time differences between different sample files.
[0016] Therefore, the step of grouping features having corresponding retention times may include grouping features having equal retention times within a first tolerance range. The step of grouping chromatograms having corresponding characteristic retention times may include grouping chromatograms having equal retention times within a second acceptable range. The second tolerance range may be smaller than the first tolerance range.
[0017] In some embodiments, the method, for each of one or more features in a group of features, (i) obtains an identification result for that feature, corresponding MS NThe process includes (i) running the mass spectrum through a mass spectrum search engine, and (ii) determining a candidate ion type (such as a candidate adsorption ion type) for a feature based on any mass difference between the mass associated with the feature and the mass predicted from the identification result. In this case, the step of identifying one or more compounds in the sample may be based on both the group of features and the candidate (adsorption) ion type.
[0018] In these embodiments, steps (i) and (ii) are performed by the corresponding MS N If data is available in mass spectral data, those characteristics of the group can be performed. If multiple sample files exist, corresponding MS N The mass spectrum can be obtained from any of the sample files. Therefore, the MS corresponding to the features N The mass spectrum is for any one of the file-specific features corresponding to that feature (multiple MS N MS (from mass spectrum) N It could be a spectrum.
[0019] As mentioned above, mass spectral data can include only a single sample file, in which case each of the multiple features is formed from one feature per file among the features per file.
[0020] In these embodiments, the step of identifying one or more compounds based on a group of features is to (i) determine one or more clusters of features, each cluster of features containing one or more features of the group and possibly corresponding to each (single) compound; (ii) determine one or more arrangements of feature clusters for a group of features, each arrangement containing one or more non-conflicting clusters of features; and (iii) select a preferred arrangement from one or more arrangements of feature clusters for a group of features; and then, This may include identifying one or more compounds based on a preferred arrangement of characteristic clusters.
[0021] However, in certain embodiments, the mass spectral data includes multiple sample files, and each of the multiple features is formed by grouping the file-specific features across sample files having corresponding masses and corresponding retention times. In this case, for each group of features, there exists a corresponding group of file-specific features within each of the multiple sample files.
[0022] Next, the step of identifying one or more compounds based on a group of characteristics is: For each group of file characteristics within each sample file, (i) Determining one or more clusters of file-specific features, where each cluster of file-specific features includes one or more file-specific features of the group and, in some cases, corresponds to each compound. (ii) For each group of features per file, determine the placement of one or more clusters of features per file, where each placement includes one or more non-conflicting clusters of features per file. (iii) For each file feature group, select a preferred arrangement from one or more arrangements of the feature clusters for each file, Next, based on the preferred arrangement of multiple sample files, (iv) Determining one or more arrangements of feature clusters for a group of features, wherein each feature cluster contains one or more features of the feature group, and in some cases corresponds to each compound, and each arrangement contains one or more non-competing feature clusters. (v) For a group of features, select a preferred arrangement from one or more arrangements of feature clusters, and then, This may include identifying one or more compounds based on a preferred arrangement of characteristic clusters.
[0023] As will be explained in more detail below, this two-step process, which first determines the preferred arrangement for each sample file and then resolves any conflicts between preferred arrangements from different sample files, has the effect of increasing the reliability and consistency of compound identification when multiple sample files are present.
[0024] In some embodiments, the step of determining one or more clusters of file-specific features is performed for each group of file-specific features within each sample file. Assign one or more candidate ion types to the characteristics of each file in the group, Determine one or more candidate relationships between the features of each file in the group, This may include resolving any conflicts between candidate ion types and candidate relationships for a group.
[0025] In these embodiments, the step of assigning one or more candidate ion types to the features of each file in the group may include assigning one or more candidate ion types from several different categories to the features of each file in the group. Candidate ion type categories may include, for example, (i) identified ion types, (ii) user-defined base ion types, (iii) default ion types, and (iv) in-source fragment ion types.
[0026] Similarly, the step of determining one or more candidate relationships (or "transitions") between the features of each file in a group may involve determining one or more candidate relationships from several different categories of relationships. Categories of relationships may include, for example, (i) in-source fragment relationships and (ii) adduct relationships.
[0027] Therefore, for example, each file feature in a group may have its own charge, and the step of assigning one or more candidate ion types to each file feature in a group is, Assign the identified ion type to any file-by-file feature in the group corresponding to the feature from which the identification result was obtained, and / or This may include assigning a user-defined base ion type or a default ion type to each file feature in a group, based on the charge of each feature in each file.
[0028] The additional or alternative step of assigning one or more candidate ion types to each file-per-feature of a group may include assigning a fragment ion type in the source to any file-per-feature of the group having a mass corresponding to the mass of the expected source fragment of another file-per-feature in the group.
[0029] In this case, the step of determining one or more candidate relationships between file-per-file features of a group may include determining a source fragment relationship between a file-per-file feature of a group and another file-per-file feature of the group if the file-per-file feature has a mass corresponding to the expected mass of source fragments of other file-per-file features (or, in a similar sense, if other file-per-file features have a mass corresponding to the expected mass of source fragments of file-per-file features).
[0030] In these embodiments, the mass of one or more expected source fragments of a file-by-file feature corresponds to the MS of that feature per file. N From the mass spectrum (multiple MS N It can be determined from the mass spectrum.
[0031] Alternatively, the method may include providing the mass of one or more expected source fragments of a feature as part of the identification result for that feature. In this case, the mass of one or more expected source fragments of a feature per file can be derived from the provided mass.
[0032] Furthermore, the mass provided as part of the identification results is configured to simulate in-source fragmentation with one or more MSs. NThis can be determined from the mass spectrum. As will be explained in more detail below, this is how specialized MS is obtained. N By using mass spectrometry, the identification of fragmentation within a source can be significantly improved.
[0033] In some embodiments, the step of determining one or more candidate relationships between file-per-file features of a group may include determining one or more candidate adduct relationships between file-per-file features of a group based on an allowable mass shift between the masses of the file-per-file features of the group.
[0034] In the embodiment, any conflicts are resolved once all possible candidate ion types and candidate relationships are determined for each group of features in each file.
[0035] In some embodiments, the step of resolving any conflict between candidate ion types and candidate relationships is: To remove any candidate relationships that conflict with the file-specific features assigned as the identified ion type (e.g., to remove any candidate adduct relationships that conflict with the file-specific features assigned as the identified ion type, and / or to remove any candidate in-source fragment relationships), and / or This may include removing any violating candidate source fragment relationships.
[0036] As described above, in some embodiments, the method includes determining one or more clusters of file-by-file features for each group of file-by-file features within each sample file. Each such determined cluster may potentially correspond to each compound.
[0037] Therefore, in the embodiment, the step of determining one or more clusters of features for each file is, Determine most or all possible clusters of features for each file, then, To remove any invalid clusters, and / or This may include removing any clusters that do not contain the file-specific features assigned as user-defined base ions.
[0038] As described above, the method may include determining one or more arrangements of clusters (of file-specific features or features) for each group, and then selecting one of these arrangements as the preferred arrangement. Each such arrangement contains one or more clusters that do not conflict with each other. Any cluster that does not have conflicting clusters can be used immediately in the preferred arrangement without further processing.
[0039] Therefore, the step of determining the placement of one or more clusters (of features per file, or of features) is, To determine whether a cluster is in conflict with one or more other clusters, This may include using a cluster in a preferred cluster arrangement for a group when it is determined that the cluster does not conflict with any other clusters.
[0040] Conversely, if competing clusters exist, they must be resolved from each other. In embodiments, this is done by determining several different arrangements of clusters, assigning a score to each arrangement, and selecting the arrangement with the highest score as the preferred arrangement.
[0041] Therefore, the step of selecting a preferred arrangement of clusters (of features per file, or of features) from one or more arrangements of clusters is: Determining the score for each cluster configuration, This may include selecting the arrangement that yields the highest score.
[0042] The step of determining the score for each cluster configuration is: Determining the cluster score for each cluster in the arrangement, by (i) assigning weight coefficients to each candidate ion type assignment in the cluster, (ii) assigning relation scores to each candidate relation in the cluster, and (iii) calculating the cluster score for the cluster by dividing the sum of the weight coefficients and relation scores by the number of features in the cluster or the number of features per file. This may include determining the score for each arrangement by dividing the sum of the cluster scores for the arrangement by the number of clusters in the arrangement.
[0043] Further embodiments provide a method for measuring mass, and this mass spectrometry method is Multiple MS 1 Mass spectra and multiple MSs N The process involves analyzing a sample to obtain mass spectral data, including a mass spectrum, where each mass spectrum has its own associated retention time. This includes processing mass spectral data using the method described above.
[0044] A further embodiment provides a non-temporary computer-readable storage medium for storing computer software code that performs the above-described method when executed on a processor.
[0045] A further embodiment provides a control system for analytical instruments such as mass measuring instruments, the control system configured to cause the analytical instrument to carry out the method described above.
[0046] A further embodiment provides analytical instruments such as mass measuring instruments equipped with the control system described above.
[0047] Next, various embodiments will be described in more detail with reference to the attached drawings. [Brief explanation of the drawing]
[0048] [Figure 1]A schematic diagram of a mass measuring instrument that can operate according to the embodiment is shown. [Figure 2] The method according to the embodiment is outlined below. [Figure 3] This shows a representation of the main relationship graph, which includes multiple features. [Figure 4] This shows a relationship graph representing each file, including the characteristics of each individual file. [Figure 5A] This shows the characteristics of each file, representing the identified ion type. [Figure 5B] This shows the characteristics of each file, labeled as one of two different user-defined base ion types. [Figure 5C] This shows the characteristics of each file, which are labeled as the default ion type. [Figure 5D] This represents the characteristics of each file, labeled as either one of two different user-defined base ion types or one of the fragment ions in the source. [Figure 6] This shows a relationship graph for each file, representing the two characteristics of each file and including various possible ion types and transitions. [Figure 7A] This example illustrates the process of removing invalid transitions from the initial relationship graph for each file. [Figure 7B] This example illustrates the process of removing invalid transitions from the initial relationship graph for each file. [Figure 8] This shows a relationship graph for each file, representing the eight characteristics of each file and including various possible ion types and transitions. [Figure 9] Figure 8 shows various possible clusters formed from the characteristics of each file. [Figure 10] This illustrates the process of removing invalid clusters from an initial set of possible clusters. [Figure 11]This illustrates the process of removing invalid, isolated clusters from an initial set of possible clusters. [Figure 12] Figure 8 shows various possible clusters formed from the characteristics of each file after invalid clusters have been removed. [Figure 13] Figure 12 illustrates the process of determining various possible descriptions of the feature groups for each file from the various possible clusters shown. [Figure 14] A final description of the group of features for each file, as determined according to the embodiment, is illustrated below. [Figure 15] This illustrates the process of breaking down fragments. [Figure 16A] The results of experiments using conventional assembly methods are illustrated as an example. [Figure 16B] Experimental results using various embodiment assembly methods are illustrated. [Modes for carrying out the invention]
[0049] Figure 1 schematically illustrates analytical instruments, such as a mass analyzer, that may be used in conjunction with the methods described herein. As shown in Figure 1, the instrument includes an ion source 10, a mass filter 20, a fragmentation device 30, and a mass spectrometer 40.
[0050] The ion source 10 is configured to generate ions from a sample. The ion source 10 may be coupled to a chromatographic separation device (not shown), such as a liquid chromatography (LC) separation device, a gas chromatography (GC) separation device, or a capillary electrophoresis separation device, so that the sample to be ionized in the ion source 10 is supplied from the separation device. The ion source 10 may be any suitable ion source, such as an electrospray ionization (ESI) ion source, an atmospheric pressure ionization (API) ion source, a chemical ionization ion source, an electron impact (EI) ion source, or the like.
[0051] The mass filter 20 is positioned downstream of the ion source 10 and is configured to receive ions from the ion source 10. The mass filter 20 is configured to filter the received ions according to their mass-to-charge ratio (m / z). The mass filter 20 may be configured such that received ions with m / z within the mass filter's m / z transport window are transported forward by the mass filter, while received ions with m / z outside the m / z transport window are attenuated by the mass filter, i.e., not transported forward by the mass filter. The width and / or center m / z of the transport window can be controlled (variable), for example, by suitable control of the RF and / or DC voltages applied to the electrodes of the mass filter 20. Therefore, for example, the mass filter 20 may be able to operate in a transport operation mode in which most or all ions within a relatively wide m / z window are transported forward by the mass filter 20, or it may be able to operate in a filtering operation mode in which only ions within a relatively narrow m / z window (centered at a desired m / z) are transported forward by the mass filter 20. The mass filter 20 may be any preferred type of mass filter, such as a quadrupole mass filter.
[0052] The fragmentation device 30 is positioned downstream of the mass filter 20 and is configured to receive almost all or all of the ions transported by the mass filter 20. The fragmentation device 30 may be configured to selectively fragment some or all of the received ions, i.e., to generate fragment ions. The fragmentation device 30 may be able to operate in a fragmentation operating mode in which almost all or all of the received ions are fragmented to generate fragment ions (which can then be transported forward from the fragmentation device 30), and in a non-fragmentation operating mode in which almost all or all of the received ions are transported forward without being (intentionally) fragmented. The non-fragmentation operating mode may also be implemented by causing the ions to avoid the fragmentation device 30. The fragmentation device 30 may also be able to operate in one or more intermediate operating modes in which the degree of fragmentation is controllable (variable). The fragmentation device 30 may also be able to operate in higher-order (MS) fragmentation in which the fragment ions are further fragmented one or more times by the fragmentation device 30. N It may also be possible to operate in fragmented mode.
[0053] The fragmentation device 30 may be any suitable type of fragmentation device, such as a collision-induced dissociation (CID) fragmentation device, an electron-induced dissociation (EID) fragmentation device, or a photodissociation fragmentation device. Numerous other types of fragmentation are possible.
[0054] In some embodiments, the fragmentation device 30 is a collision-induced dissociation (CID) fragmentation device. Therefore, the fragmentation device may include a collision cell, which may be filled with a collision gas maintained at a relatively high pressure, for example. Ions can be selectively fragmented within the collision cell by controlling (variing) the kinetic energy at which they enter the collision cell. In the fragmentation operating mode, ions may be accelerated to enter the collision cell with relatively high kinetic energy, which may fragment most or all of the accelerated ions. In the non-fragmentation operating mode, ions may be entered with relatively low kinetic energy, which may be insufficient to fragment most or all of the ions. In the intermediate mode, ions may be entered with intermediate kinetic energy.
[0055] The mass spectrometer 40 is positioned downstream of the fragmentation device 30 and is configured to receive ions from the fragmentation device 30. Therefore, depending on the operating mode of the fragmentation device 30, the mass spectrometer 40 can receive unfragmented precursor ions or fragmented ions. The mass spectrometer 40 is configured to analyze the received ions to determine their mass-to-charge ratio (m / z) and / or mass, i.e., to generate a mass spectrum of the ions. The mass spectrometer 40 may be any suitable type of mass spectrometer, such as an ion trap mass spectrometer, an electrostatic orbital trap mass spectrometer (such as the Orbitrap® FT mass spectrometer from Thermo Fisher Scientific), or a time-of-flight (ToF) mass spectrometer such as a multi-reflecting time-of-flight (MR-ToF) mass spectrometer.
[0056] Figure 1 is merely schematic, and it should be noted that the instrument may include, and certainly does, any number of one or more additional components. For example, the instrument may include one or more ion transfer stages positioned between any of the exemplified components, including, for example, an atmospheric pressure interface and / or one or more ion guides, lenses, and / or other ion optical devices configured so that some or all of the ions can be properly transferred through the instrument. These ion transfer stages may include any preferred number and configuration of ion optical devices, for example, optionally one or more ion guides, lenses, and / or other ion optical devices.
[0057] In some embodiments, the instrument may include two or more mass spectrometers. For example, the instrument may be a dual mass spectrometer hybrid mass analyzer of the type described in European Patent No. 3,410,463, the details of which are incorporated herein by reference.
[0058] As shown in Figure 1, the device is under the control of a control unit 50, such as a appropriately programmed computer, which controls the operation of various components of the device, for example, setting the voltages to be applied to various components of the device. The control unit 50 can also receive and process data from various components, including analytical devices in the methods of various embodiments.
[0059] The instrument may be capable of operating in various operating modes. For example, the instrument may be a tandem mass measuring instrument capable of operating in MS1 and MS2 operating modes.
[0060] In MS1 (or "total mass scan") operating mode, the mass filter 20 operates in its transfer mode, and the fragmentation device 30 operates in its non-fragmentation mode, resulting in unfragmented ("precursor" or "parent") ions over a wide m / z range (e.g., the total mass range) being analyzed by the analyzer 40 to generate an MS1 spectrum.
[0061] In MS2 operation mode, the mass filter 20 operates in its filtering mode, and the fragmentation device 30 operates in its fragmentation mode. As a result, for example, precursor ions in a selected narrow m / z range are fragmented, and the resulting fragments ("products" or "daughters") ions are analyzed by the analyzer 40 to generate an MS2 spectrum.
[0062] The instrument may also be capable of operating in one or more higher-order fragmentation operating modes, such as the MS3 operating mode, where the precursor ions are fragmented, and at least some of the resulting fragment ions are themselves fragmented, and the second-generation fragment ions ("granddaughter ions") are analyzed by the analyzer 40 to generate an MS3 spectrum. Generally, the instrument operates in any order of fragmentation operating modes, i.e., MS3 where N≧2. N It may be possible to operate in the operating mode.
[0063] The method for operating the analytical instrument includes providing the sample to a chromatographic (e.g., LC or GC) separation device so that the sample is chromatographically separated, ionizing the eluate from the chromatographic separation device in an ion source 10, and analyzing the resulting ions. Different compounds in the sample experience different retention times (RT) within the chromatographic separation device and therefore elute (and ionize) from the chromatographic separation device at different times. The chromatographic separation device typically takes tens of seconds or several minutes to complete each chromatographic separation scan.
[0064] During each chromatographic separation scan, multiple MS2 spectra (or, more generally, multiple MS2 spectra) are extracted. N The spectrum can be obtained, for example, by sequentially changing the center of the (narrow) m / z window of the mass filter among several different m / z values so that each of several different precursor ions having a different m / z is sequentially selected (and fragmented).
[0065] In data-dependent acquisition (DDA) operating mode, multiple different m / z values may correspond to multiple different precursor ions identified from the corresponding MS1 data (i.e., total mass scan). Therefore, a typical data-dependent acquisition (DDA) method involves, during a chromatographic separation scan, (i) acquiring an MS1 spectrum over the m / z range of interest, (ii) identifying one or more precursor ions of interest within the MS1 spectrum, and (iii) performing an MS2 (or MS) scan for each identified precursor ion of interest. N Step (iii) includes repeating the step of obtaining a spectrum. For each identified precursor ion, step (iii) includes isolating the precursor ion using a mass filter 20, fragmenting the isolated precursor ion in a fragmentation device 30, and mass spectrometry the fragment ion using a mass spectrometer 40.
[0066] Data Independent Acquisition (DIA) MS2 (or MS N In this operating mode, multiple different m / z values can be obtained from a predetermined (fixed) list, i.e., without referring to MS1 data. For example, a narrow m / z separation window may be progressively advanced across the entire m / z range of interest, as described, for example, in European Patent No. 3,410,463.
[0067] In any case, each chromatographic separation scan performed by the analytical instrument will generate a sample file. Each sample file will contain data for multiple MS1 spectra, each with its associated retention time, and multiple MS2 (or MS NRegarding the spectrum, each includes data with associated retention time and associated mass filter separation window or precursor ion m / z. The prepared sample may be fractionated, and a chromatographic separation scan may be performed for each fraction to generate multiple such sample files for the same sample.
[0068] The repeat rate of the DDA / DIA method can be fast enough to sample each compound of interest that has been chromatographically separated during its chromatographic elution. Therefore, for each sample file, a chromatogram can be constructed for each specific m / z of interest from multiple MS1 spectra. Such chromatograms typically appear as peaks corresponding to the chromatographic elution peaks of each chromatographically separated compound.
[0069] Next, the center (e.g., vertex) retention time (RT) of each chromatogram can be determined, for example, using a suitable peak detection algorithm. Chromatograms having the same center retention time (RT) (e.g., within a specific tolerance range) can be grouped together to obtain one or more sets of m / z values, where all m / z values in the set are determined to have the same retention time (RT).
[0070] Each such set of multiple m / z values can result from multiple different compounds co-eluting from the chromatographic separation device, making it highly complex and difficult to interpret. Therefore, each set of multiple m / z values needs to be "assembled" into one or more clusters, where each cluster is a subset of m / z values from a set (or, as appropriate, a complete set) of multiple m / z values determined to have the same RT, and all m / z values within a cluster belong to the same single compound. In other words, each single chromatographic separation compound can result in multiple m / z clusters in the MS1 data (with the same RT), and the presence of multiple co-eluting compounds introduces complexity into the MS1 data.
[0071] A single compound can produce multiple different m / z values in MS1 (at the same RT) due to, for example, different charge states (z) and different isotopes. Sets of isotopes appear in the MS1 data as a series of distinctive peaks separated from each other, such as 1 m / z for monovalent species, 1 / 2 m / z for divalent species, and 1 / 3 m / z for trivalent species. Understanding this means that the MS1 data can be "isotope-removed" by grouping the m / z values corresponding to a set of isotopes (in sets of m / z determined to have the same RT) into so-called "file-specific features," each file-specific feature having a single distinctive m / z value (e.g., corresponding to the lightest isotope). Understanding this also means that the precise charge state can be determined for each such file-specific feature (from the m / z separation between isotopic peaks). (Any "singlet" peaks appearing in MS1 data (i.e., peaks where the corresponding isotope is not present at all, if the corresponding isotope is below the detection limit—usually low isotope abundance peaks) can be assumed to be monovalent (or have some other default charge depending on the sample type, etc.).) Then, knowledge of the charge state (z) and mass-to-charge ratio (m / z) means that a single characteristic mass (m) can be determined for each file's characteristic. Algorithms for such isotope removal and charge state determination are known in the art.
[0072] However, despite such isotope removal algorithms, multiple file-per-file features may still exist for each compound. In other words, when each set of multiple m / z isotopes is removed, a group of multiple file-per-file features may remain (here, each such group of multiple file-per-file features may arise from a single compound or from multiple different co-eluting compounds), and therefore each group of multiple file-per-file features may be further assembled into one or more clusters, all of which must belong to the same single compound.
[0073] The presence of multiple file-specific features for a single compound can have several different causes. In particular, it can result from the following: (i) Solvent adducts; i.e., adducts during the ionization process may result in a peak in MS1 with a characteristic mass difference relative to the neutral mass M of the compound. For example, protonated ions [M+H] + It has a mass difference of approximately 1, but the ammonia adduct [M+NH4] + For example, they have a mass difference of approximately 18. (ii) Homodimer or heterodimer; that is, two proteins bound together. These may again manifest as characteristic mass differences in MS1. (iii) Fragments within the source; that is, unintended fragmentation of ions can result in fragmented ions appearing in the MS1 spectrum (not just the desired MS2 spectrum), which makes interpretation of the MS1 spectrum difficult.
[0074] Similarly, if multiple sample files exist for the same sample, multiple "features" may exist for each compound (where a feature is a set of multiple corresponding features from each of the multiple sample files). Ideally, all corresponding features from each file should be identical with respect to m / z and RT, but in reality, RT may differ slightly between different chromatographic separation runs. Therefore, each group of features must be assembled into one or more clusters, so that all features within a cluster belong to the same single compound.
[0075] Various embodiments focus on methods for assembling such features or file-specific features into clusters, that is, methods for determining which features or file-specific features relate to the same compound.
[0076] Known methods for performing such composite assemblies rely solely on the expected mass shifts between features on a file-by-file basis. However, we now recognize that these methods have several problems. (i)(a) Compounds with intrinsic charges cannot be precisely identified. Some compounds exist with intrinsic charges, and these usually do not form additional adducting ions. Therefore, there is no observable mass shift in MS1 (between at least two different ions), which provides a clue for their precise assignment (i.e., the ion is, for example, a protonated ion [M+H] + or ammonia adduct ion [M+NH4] + It is not possible to determine whether or not this is the case from MS1 data alone, and this requires knowing that the neutral mass M of the compound can be reliably determined. In known methods, these "lone ions" are, for example, defined by the default ion definition (e.g., protonated ions [M+H] for electrospray ionization). + ) may be inaccurately assigned. As will be further described below, various embodiments may use its fragmentation (e.g., MS2 or MS) against a standard database (e.g., the mzCloud(trademark) database). N This problem is addressed by performing preliminary identification of compounds through data retrieval. (i)(b) In relation to this, compounds that do not form (de)protonated ions are often incorrectly identified. Some compounds may differ from their default ion (e.g., protonated ion [M+H] for electrospray ionization). + It is preferable to create one type of ion (different from the others). For example, some lipids produce sodium adduct ions [M+Na] + It tends to form only fragments. This leads to the same problem as described above. As will be further explained below, various embodiments address this problem again by performing preliminary identification of the compound by searching its fragment data against a standard database. (ii) Compound-specific in-source fragments are not identified. Many compounds undergo small amounts of fragmentation during ionization (or somewhere in the instrument during MS1 scanning), and therefore these fragments appear in the MS1 spectrum. These fragment ions can then be inaccurately treated by the assembly algorithm as additional MS1 features, producing false positive hits. One example of this is PEG (polyethylene glycol) that has lost one or more units of EG (ethylene glycol). Upon analysis, it may appear as if the sample contains both PEGn6 and PEGn5, but the latter is actually a fragment created in the instrument. Various embodiments, as described below, (e.g., MS2 or MS in the sample file) N (from the data) Criterion MS2 (or MS) for each feature N By acquiring data, or by using an identification step (e.g., mzCloud®) (which may be configured to simulate in-source fragmentation), low-energy MS2 or MS2 can be used. N From spectral acquisition to standard MS2 or MS N Return the spectrum, then the reference MS2 or MS N This problem is addressed by searching the MS1 data for fragments that match the fragments within the data. (iii) Inconsistent adduct type assignments across multiple sample files. Several reasons exist that different fractions of the same sample may produce different results with respect to adduct type assignments. This could be due to variations in intensity / isotope abundance between fractions, for example, one or more adducts falling below the detection limit, resulting in entire adduct clusters being assigned differently. As described below, embodiments address this problem by placing multiple potentially competing individual fraction descriptions (or "arrangements") into a final graph, which is resolved at the end of the process.
[0077] Figure 2 illustrates the method according to the embodiment. As shown in Figure 2, in the first step (step 100), file-specific features are detected in each sample file (in the manner described above), and these are then grouped into features across multiple sample files by matching their RT and m / z values. These integrated features are inserted into the main relational graph 510, for example, as nodes.
[0078] Figure 3 shows an explanatory diagram of an example of such a main relation graph 510. In Figure 3, each shaded circle (node) represents an identified feature.
[0079] Returning to Figure 2, in step 200, MS2 (or MS N ) The data is MS2 or MS in any of the sample files for the same sample. N If available for any of the features detected from the data, this data is subjected to a database search algorithm for identification. For example, any suitable database search algorithm can be used, such as the mzCloud™ spectral library. From the returned results (which will include the precursor ion neutral mass M), the possible adduct type assignments are determined. In particular, the ion type of a feature can be calculated using the mass shift between the measured mass and the extracted identification mass. For example, if the measured mass in MS1 is (M+18), the feature is an ammonia adduct [M+NH4]. + It is identified as such, and this information is inserted into the main relational graph 510 as one possible identification of the ion type.
[0080] This process not only enables highly reliable assignment of ionic types to features but also provides a method for accurately identifying compounds with intrinsic charges. Possible ionic types may be obtained directly from identification metadata, from one or more predefined lists, or, if necessary, automatically generated for a given charge state, for example, by (de)protonation.
[0081] Therefore, in the embodiment, the library search 200 is performed in step 500 before the final compound assembly. The results of the search 200 are not treated as a definitive identification of a feature (as conventionally), but rather as one strong possibility that must be resolved later against any other competing explanations and identifications (e.g., from other sample files).
[0082] Steps 300 and 400 in Figure 2 are performed to group the file-by-file features for each group of file-by-file features (in the same RT) into possible clusters of related file-by-file features. To address problem (iii) above, this process works on a file-by-file basis, and then all possible descriptions are added to the main graph 510, and any conflicts between descriptions from different sample files are resolved in step 500 (which will be further described below).
[0083] In step 300, the file-specific features within each sample file are grouped in step 310 by matching them (for example, in the method described above) with the RT value and, optionally, one or more other characteristics such as peak shape. Because this process operates on a file-by-file basis at this stage, a very narrow tolerance can be used to group file-specific features by RT. Mapping to features is performed in step 320, as shown.
[0084] Figure 4 shows an explanatory diagram of an example of the resulting file-by-file relationship graph. In this file-by-file context, each shaded circle in Figure 4 represents a file-by-file feature within a group of file-by-file features determined to have the same RT.
[0085] In step 400 of Figure 2, the characteristics of each grouped file can be analyzed to determine all possible relationships between the various characteristics of each file within each group.
[0086] This initially involves tentatively assigning all possible ion types to the characteristics of each file in the group. Various categories of possible ion types, including identified ions (Figure 5A), base ions (Figure 5B), default ions (Figure 5C), and fragment ions (Figure 5D), are illustrated in Figures 5A–5D. At this stage, each file characteristic may have multiple possible tentative ion type assignments, which can be resolved later in the process.
[0087] Ion type identification information (from step 200) is added to the corresponding features for each file identified in step 200. For example, Figure 5A shows the protonated ion [M+H] in step 200. + This represents the file-specific characteristics of the features identified as such.
[0088] Each file-specific feature is also provisionally assigned as one or more user-defined base ions (Figure 5B) depending on the known charge of the file-specific feature (where the charge of each file-specific feature is known from the isotope removal algorithm as described above (or from some other preprocessing of the data)), or provisionally assigned as a default ion (Figure 5C) for which the user has not defined a base ion for the ion with the known charge of that file-specific feature. User-defined base ions can be set as desired by the user, for example, depending on the sample chemistry and / or the configuration of the ion source. As will be further described below, in some embodiments, at least one user-defined base ion may need to be present in a cluster in which its cluster is considered valid. The default ion may be the most common ion formed during the ionization of the ion with the charge of the file-specific feature (e.g., by (de)protonation). For example, in the case of electrospray ionization, the default ion for an ion with a monovalent positive charge is the protonated ion [M+H] + It may also be set as follows, and for ions with a divalent positive charge, [M+2H] 2+ It may be set as such.
[0089] Therefore, for example, Figure 5B shows one of the two user-defined base ions, i.e., the protonated ion [M+H] + or sodium adduct ion [M+Na] + This represents the characteristics of each file with a monovalent positive charge, which are provisionally assigned as either of the following. Figure 5C shows the default divalent positive charge ion, i.e., [M+2H] 2+ This represents the characteristics of each file that has a divalent positive charge, which has been provisionally assigned as such.
[0090] Returning to Figure 2, in step 410, the features of each file within the group are analyzed to determine whether any of the features of each file could be a source fragment. MS2 (or MS N )Data is available, and for each file in the group, describe the characteristics of representative MS2 (or MS N )The data is retrieved.
[0091] Typical MS2 (or MS N )The data is the corresponding MS2 (or MS) from the sample file. N ) data may be acceptable, or more usefully, standard MS2 (or MS) from a library of low-energy collision spectra. N ) may be a spectrum. Such a spectrum may be returned along with identification (e.g., from mzCloud®) and may be configured to accurately simulate in-source fragmentation. Thus, in embodiments, a new feature is added to the search engine (e.g., mzCloud®) that returns the mass from a low-energy collision spectrum (e.g., HCD10) as part of each identification hit.
[0092] If identification 200 is unsuccessful, and / or if low-energy collision data is unavailable, MS2 (or MS) obtained from the raw file will be used. N )The spectrum is instead a representative MS2 (or MS N It may also be used as a reference. However, the assignment confidence is higher when fragmented data collected based on reliable standards is used.
[0093] Next, a typical MS2 (or MS N The data is compared with the characteristics of other files within the group, for example, a representative MS2 (or MS NBy searching for file-by-file features within the group that have a mass consistent with the mass of the source fragment in the data, it is determined whether any of the other file-by-file features could potentially be the source fragment of the file-by-file feature under consideration.
[0094] Any of the file-by-file features within the group that have such a consistent mass are labeled as potentially in-source fragments. Thus, for example, Figure 5D shows two user-defined base ions ([M+H] + Alternatively, [M+Na] + This represents a file-specific feature with a monovalent positive charge, which is provisionally assigned as either one of the following, or as a source fragment ion. The ion type of a source fragment is generally unknown, so a general ion type assignment (e.g., [Me]) is not used. + ) may be used.
[0095] For possible source fragments, the relationship between the possible source fragment and its parent ion is also recorded in the graph. Thus, for example in Figure 6, this potential relationship is shown by the dotted line connection between the nodes (the file-specific feature labeled "P" is potentially the precursor ion of source fragment "F").
[0096] Returning to Figure 2, in the next step (step 420), the possible adduct relationships between file-by-file features within a group are determined by finding the allowable mass shifts between the masses of the file-by-file features. Using a predefined list of ion types (e.g., adducts, common neutral loss, simple polymers, and charge states), the expected mass shifts are generated, and these expected mass shifts are applied to each group to find all possible relationships between file-by-file features in the group. Thus, for example, if two file-by-file features have a mass difference of 17 (=18-1), then those two features are protonated ions [M+H] + and its ammonia addition [M+NH4] +If the difference is 22 (=23-1), the characteristic of each file is [M+H]. + and [M+Na] + It is possible, for example.
[0097] Any such possible relationship is recorded in the graph. For example, in Figure 6, potential adduct relationships are shown by solid line connections between nodes, where the line represents a possible adduct relationship (i.e., in this example, [M+H]). + ⇔[M+NH4] + They are labeled according to the following criteria.
[0098] Figure 6 shows a simple example of two feature groups per file with various possible relationships that are labeled. In particular, the first feature (left) is the type of ion identified [M+H] + or source fragment [Me] + It is labeled as either of the following. The second feature (on the right) is two different base ions, namely [M+H] + Or [M+NH4] + It is labeled as either one of the following. In addition, in Figure 6, [M+H] + ⇔[M+NH4] + Solid lines labeled "P" represent possible adduct transitions, while dotted lines represent possible source-in-shard relationships (features labeled "P" are precursor ions of source-in-shard "F").
[0099] As can be seen from the example in Figure 6, the process described above can result in multiple possible potentially conflicting relationships between the features of each file within the group. These conflicts are then resolved in step 430 of Figure 2.
[0100] To do this, first any conflicts between ion type assignments (loops) and transitions (connections between nodes) are resolved. This process is illustrated by Figure 7 and includes (i) removing any transitions that conflict with the identified ion (Figure 7A), and (ii) removing fragment transitions that are in violation (Figure 7B).
[0101] Figure 7A shows again the example of Figure 6, but with an additional solid line labeled + ⇔[M+H] + which represents a second possible adduct transition (in addition to the transition + ⇔[M+NH4] + ). As seen in Figure 7A, the in-source fragment label [M-e] for the first feature per file + has been removed because it conflicts with the identified ion type [M+H] for the first feature per file + . Next, the second possible adduct transition ([M+H-NH3] + ⇔[M+H] + ) has also been removed because the [M+H-NH3] + required for that transition also conflicts with the identified ion type [M+H] for the first feature per file + .
[0102] Figure 7B shows an example where violating fragment transitions are removed. Specifically, the fragment transition [M+H] + ⇔[M+Na] + is removed because the [M+H] + fragment cannot be created from the [M+Na] + precursor (i.e., a fragment transition must not "create" new elements). In contrast, the possible fragment transition [M+H] + ⇔[M+NH4] + is retained because the [M+H] + fragment can be created from the [M+NH4] + precursor.
[0103] Figure 8 shows a simplified, exemplary per-file graph that may result from the process described above, where all remaining possible ion type assignments and transitions are indicated by various loops and connections between nodes. Several different explanations for assembling the per-file features are still possible. From the graph in Figure 8, all possible clusters (each cluster corresponding to a single compound M) can be determined.
[0104] Figure 9 shows the result of this decision. In Figure 9, each possible cluster is shown as a shaded box. The leftmost box 431 represents a cluster that has no other competing explanations (i.e., no file-specific features appear in any other possible clusters). Thus, this represents a successful identification of a cluster that does not require further processing (at the file-specific level). This cluster 431 is added to the main graph 510 (which will be resolved later in step 500 for any other competing explanations from other sample files).
[0105] Each remaining cluster is a possible cluster that competes with at least one other possible cluster. These various competitions are illustrated by the vertical dashed lines in Figure 9. These competitions are then resolved.
[0106] As illustrated by Figure 10, in an embodiment, if the user provides one or more base ion types (as described above), any of the possible clusters that do not contain one or more per-file features that are potentially assigned as one or more of the base ion types are rejected. (If the user does not provide any required base ion types, this step may be skipped.) Thus, as illustrated by Figure 10A, clusters containing base ions may be retained. As illustrated by Figure 10B, clusters containing only identified ions may be rejected. As illustrated by Figure 10C, clusters containing only default ions may be rejected. As illustrated by Figure 10D, clusters containing no ions may be rejected.
[0107] Referring again to Figure 9, in the illustrated example, only two transitions are found without any possible assignment (loop) of a base-added ion or identified adduct to any of the nodes in cluster 432; therefore, the possible cluster labeled 432 is eliminated. Thus, cluster 432 is considered invalid and excluded from further consideration.
[0108] Next, as illustrated in Figure 11, any invalid ion clusters are rejected, and any remaining isolated clusters (i.e., single per-file features that never transition to any other per-file feature) are considered. As illustrated in Figure 11A, assignments to any isolated clusters with valid ion type assignments (i.e., base ion assignments, identified ion assignments, or default ion assignments) are retained. As illustrated in Figure 11B, for isolated clusters labeled as potentially in-source fragments, ion type assignments are retained only for identified ions and generic fragment ions (whereas ion type assignments are removed for base ions and default ions).
[0109] Accordingly, in the embodiment, step 430 includes (i) removing ion type assignments having mismatched charges, (ii) removing ion type assignments that compete with any pre-identified ion types, (iii) removing any clusters that are missing one or more user-specified base ion types, and (iv) removing any invalid fragment-precursor relationships (e.g., where additional atoms need to be added to the fragment).
[0110] Figure 12 shows the resulting graph after processing the graph in Figure 9. As seen in Figure 12, clusters 431 and 432 from Figure 9 have been removed (for the reasons given above). Furthermore, after removing cluster 432, the orphaned node 433 from Figure 9 no longer conflicts with any other nodes and is therefore considered resolved and has also been removed from the graph in Figure 12.
[0111] As can be seen in Figure 12, at this stage, there may still be multiple competing clusters. However, the graph has been decomposed into two separate subgraphs, and since these subgraphs do not compete with each other, they can each be processed separately.
[0112] Next, the remaining conflicts are resolved using a scoring system. This process is illustrated in Figure 13, which shows the process for the leftmost subgraph in Figure 12. A similar process will be carried out separately for the other subgraphs in Figure 12.
[0113] In the scoring process, the largest of the possible clusters is selected first and assumed to be correct ("valid"). Thus, as shown in Figure 13A, the cluster labeled 434 is initially enabled. Any other clusters that compete with this largest cluster are then disabled. Now, the next largest possible cluster (which is neither valid nor disabled) is enabled, and so on, until the first description (or "assignment" or "placement") is reached. Thus, in the example in Figure 13A, the image at the bottom represents this initial description, where the group of features per file is described by two clusters, 434 and 435.
[0114] Next, this description is given an assignment score. To calculate the assignment score, each ion type assignment (loop) is given a relative weighting coefficient. Any suitable weighting coefficient can be used. For example, the highest relative weight may be used for the most common adducts (e.g., [M+H] or [MH]), a slightly lower weight may be used for common adducts (e.g., [M+Na] or [M+K]), and an even lower weight may be used for the rest (e.g., uncommon adducts). The relative magnitudes of the various weighting coefficients may be configured as desired, or set by the user, to reflect, for example, the probability of a particular adduct appearing under the specific experimental conditions / sample chemistry used to acquire the data.
[0115] In addition, each transition is scored as twice the sum of its two ion weight coefficients. The coefficient 2 is used here because a transition explains two nodes, while a loop assignment explains only one node. The default assignment (loop) is usually the highest weighted adduct, so the algorithm otherwise prefers the orphan assigned as the default adduct, which is undesirable.
[0116] Next, each cluster is assigned a cluster score, which is calculated by dividing the sum of all loop scores and transition scores for the cluster by the number of nodes in the cluster.
[0117] Finally, an assignment score is calculated for the description, which is the sum of all cluster scores for all enabled clusters in the description divided by the number of clusters. The assignment score is designed to generate high scores for situations with many similar clusters (and to avoid descriptions containing, for example, only one large cluster and many isolated clusters).
[0118] Returning to Figure 13, one or more other possible explanations are then tested. To do this, as shown in Figure 13B, one of the clusters that was disabled in the first step (i.e., Figure 13A) is enabled, and the above process is repeated to generate a conflicting explanation. Thus, for example, in Figure 13B, the image on the right represents this conflicting explanation, where the group of features per file is now explained by three clusters 436, 437, and 438. This explanation is then scored again using the same scoring system.
[0119] Again, another possible explanation is tested by enabling one of the different clusters that were disabled in the first step, and the above process is repeated to generate another competing explanation. Thus, for example, in Figure 13C, the image on the right represents this further competing explanation, where the group of features per file is now explained by two clusters 437 and 439. This explanation is scored again using the same scoring system.
[0120] This process can be repeated, for example, until all possible explanations are scored. The explanation with the highest assigned score can then be selected as the final explanation for the output.
[0121] Therefore, possible clusters are recursively evaluated to arrive at the best possible assignment score, where the score design reflects the properties of the analyzed sample, for example, that most compounds behave similarly under specific chromatographic conditions and thus create similar ion clusters.
[0122] Figure 14 shows an example of the final description, where the group of file-specific features from Figure 9 is explained by five clusters 431, 433, 437, 439, and 440. This final description is added to the main graph 510.
[0123] If multiple sample files exist, a similar process is performed for each sample file, and one explanation is added to the main graph 510 for each sample file (with any duplicates removed). This may result in the main graph 510 having several competing explanations (i.e., similar to Figure 12).
[0124] Therefore, returning to Figure 2, in step 500, any conflicting explanations in the main graph 510 are resolved. This process is carried out using the same algorithm as described above (see Figures 12-14). This step ensures that information from all sample files is taken into consideration and that a consistent assignment is achieved across the files.
[0125] Each resulting subcluster represents a unique compound. Therefore, ultimately, compounds can be identified from each cluster in the final main graph (step 520 in Figure 2).
[0126] While various specific embodiments have been described above, various alternative and additional embodiments are possible.
[0127] For example, one possible further step in the processing pipeline might be quantitative analysis of the results. This requires knowing which features within the cluster can be used and which should be ignored (e.g., to sum their areas for quantification). Typically, any fragment is ignored, and only ions assigned as meaningful adducts are used for this purpose. However, fragments may be used for quantification if, for example, they relate to actual adduct ions. For instance, a common fragment of [M+NH4] is [M+H]. Therefore, an additional step can be applied to mark the actual adduct ions that should be used for further processing, while any common fragments should be ignored.
[0128] This process is illustrated by Figure 15, where several fragments can be shared by multiple precursors, as illustrated in Figure 15A. If this is the case, and the relationship to one of the precursors is also a valid adduct transition, as illustrated in Figure 15B, then any relationship to any other precursor is rejected. This means that the fragment can be retained for further use, while avoiding its reuse. If, for a fragment, no fragment-precursor relationship is detected, and several other adduct transitions exist from the feature, as illustrated in Figure 15C, then their assignments are rejected, and the fragment is assigned as a general fragment ion.
[0129] The methods described herein have general applicability to small molecule mass measurement analyses, such as metabolic pathway analysis, degradation product analysis, and forensic analysis.
[0130] Figure 16 shows exemplary results from artificial samples of eight compounds mixed in known ratios, along with their respective
[13] C-labeled analogues. Due to the unique labeling, it was possible to identify the compounds along with all possible analogues and to evaluate the method.
[0131] As shown in Figure 16A, the conventional technique clearly exhibits the expected isotope exchange profile, but reported three additional compounds with retention times close to one of the expected compounds. Considering the assigned molecular formulas, and especially the differences in formulas between the compounds at the same retention time, an inaccurate interpretation was suggested.
[0132] As shown in Figure 16B, the same data was analyzed by the compound assembly method according to the embodiment. In this case, all three additional compounds were precisely interpreted as in-source fragments and bound to their respective precursors.
[0133] Compound assembly methods of various embodiments enable the complete and consistent assembly of multiple diverse forms of ions produced from a single compound in multiple samples. Unlike conventional techniques that rely solely on expected mass shifts in MS1 spectra, embodiments additionally utilize MS2 fragmentation spectra and online identification tools for initial adduct assignment and for detection of untargeted in-source fragments. The results from these steps are integrated into unique ion clusters using a set of chemical and heuristic rules. This strategy significantly reduces the number of false positive identifications.
[0134] According to the embodiment, this is done in particular by the following: (i) Apply a two-step grouping mechanism to integrate potentially relevant features across multiple different sample files. This allows for consistent ion assignment across multiple different sample files. (ii) Pre-identify compounds using acquired MS2 data and provide highly reliable adduct assignment candidates based on spectral library consistency. This avoids further ion misassignment of identifiable compounds. (iii) Utilize the curated low-energy collision spectra of already identified compounds in the data, and then search for specific in-source fragment candidates. This helps reduce the number of intrinsic compounds that produce false positives. (iv) Use the acquired MS2 spectra to search for candidate fragments within the source. This helps reduce the number of false positives for intrinsic compounds when low-energy collision spectra are not available. (v) Apply a set of rules to the network of all possible ion relationships to eliminate assignments with low confidence. This reduces the chance of incorrect adduct assignments. (vi) Apply a custom scoring and evaluation mechanism to an integrated network of ion relationships to resolve possible conflicting assignments. This creates individual ion clusters corresponding to specific compounds.
[0135] While the present invention has been described with reference to various embodiments, it will be understood that various modifications can be made without departing from the scope of the invention as described in the appended claims.
Claims
1. A method for processing mass spectral data, wherein the mass spectral data is a plurality of MS 1 Mass spectra and multiple MS N The method includes (N≧2) mass spectra, each mass spectrum having its respective associated retention time, and the method is The aforementioned multiple MS 1 A step of detecting a group of features in a mass spectrum, wherein each feature in the group has its own mass, and the features in the group have equal retention times within a first tolerance range. For each of the one or more features of the group, (i) in order to obtain the identification result for that feature, the corresponding MS N (ii) a step of running the mass spectrum through a mass spectrum search engine, and (ii) a step of determining a candidate ion type for the feature based on the mass difference between the mass associated with the feature and the mass predicted from the identification result, and then, A method comprising the step of identifying one or more compounds based on the group of features and the candidate ion type.
2. (ii) The method of claim 1, wherein the step of determining a candidate ion type for the feature includes determining a candidate addition ion type for the feature based on the mass difference between the mass of the feature and the expected mass from the identification result.
3. The mass spectral data includes at least one sample file, each sample file corresponding to a chromatographic separation scan and containing multiple MS 1 Mass spectra and multiple MS N The mass spectrum includes the multiple MS 1 The step of detecting groups of features in the mass spectrum is: For each sample file, the step of detecting multiple file features within that sample file, wherein each file feature has its own mass and its own retention time, A step of forming a plurality of features from the aforementioned file features, wherein each feature has its own mass and its own retention time, The method according to claim 1 or 2, comprising the step of forming a group of features by grouping features having equal retention times within a first tolerance range.
4. The step of detecting multiple file features within a sample file is: The multiple MS of the sample file 1 A step of constructing multiple chromatograms from a mass spectrum, wherein each chromatogram has its own mass-to-charge ratio (m / z), The steps include determining the characteristic retention time for each chromatogram, The steps include: grouping chromatograms having equal retention times within a second acceptable range into one or more sets of chromatograms; The method according to claim 3, comprising the step of applying an isotope removal algorithm to each set of chromatograms so as to form a group of file features.
5. The method according to claim 4, wherein the second tolerance range is smaller than the first tolerance range.
6. The method according to claim 4, wherein the mass spectral data comprises a plurality of sample files, and each of the plurality of features is formed by grouping file features having a corresponding mass and a corresponding retention time.
7. Each group of features is formed from a corresponding group of file features within each sample file, and the step of identifying one or more compounds based on the group of features is as follows: For each group of file features in each sample file, (i) A step of determining one or more clusters of file features, wherein each cluster of file features includes one or more file features of the group and, if applicable, corresponds to each compound. (ii) A step of determining, with respect to the group of file features, one or more arrangements of the clusters of file features, wherein each arrangement includes one or more non-conflicting clusters of file features, (iii) With respect to the group of file features, the step of selecting a preferred arrangement from the one or more arrangements of the cluster of file features, Next, based on the preferred arrangement of the plurality of sample files, (iv) A step of determining the arrangement of one or more clusters of features for the group of features, wherein each cluster of features includes one or more features of the group of features, and optionally corresponds to each compound, and each arrangement includes one or more non-competing clusters of features. (v) For the group of features, the step of selecting a preferred arrangement from the one or more arrangements of feature clusters, and then, The method according to claim 6, comprising the step of identifying one or more compounds based on the preferred arrangement of feature clusters.
8. A method for processing mass spectrometry data, wherein the mass spectrometry data includes a plurality of sample files, and each sample file includes a plurality of MS 1 mass spectra and a plurality of MS N (N≥2) mass spectra, and each mass spectrum has its respective associated retention time, and the method comprises For each sample file, the MS of that sample file 1 A step of detecting multiple file features in a mass spectrum, wherein each file feature has its own mass and its own retention time, The steps include forming a plurality of features from the file features by grouping file features having corresponding mass and corresponding retention time, The steps include forming groups of features by grouping features having equal retention times within a first acceptable range and forming corresponding groups of file features within each sample file, and then, For each group of file features in each sample file, (i) A step of determining one or more clusters of file features, wherein each cluster of file features includes one or more file features of the group and, if applicable, corresponds to each compound. (ii) A step of determining, with respect to the group of file features, one or more arrangements of the clusters of file features, wherein each arrangement includes one or more non-conflicting clusters of file features, (iii) With respect to the group of file features, the step of selecting a preferred arrangement from the one or more arrangements of the cluster of file features, Next, based on the preferred arrangement of the plurality of sample files, (iv) A step of determining the arrangement of one or more clusters of features for the group of features, wherein each cluster of features includes one or more features of the group of features, and optionally corresponds to each compound, and each arrangement includes one or more non-competing clusters of features. (v) For the group of features, the step of selecting a preferred arrangement from the one or more arrangements of feature clusters, and then, A method comprising the step of identifying one or more compounds based on the preferred arrangement of feature clusters.
9. The step of determining one or more clusters of file features is to determine each group of file features in each sample file, A step of assigning one or more candidate ion types to each file feature of the group, The steps include determining one or more candidate relationships between the file features of the group, The method according to claim 8, comprising the step of resolving any conflict between the candidate ion type and the candidate relationship.
10. Each file feature has its own charge, and the step of assigning one or more candidate ion types to each file feature of the group is, The steps of assigning the identified ion type to any file feature in the group corresponding to the feature from which the identification result was obtained, and / or The method according to claim 9, comprising the step of assigning a user-defined base ion type or a default ion type to each file feature of the group based on the respective charges of the file features.
11. The step of assigning one or more candidate ion types to each file feature of the group includes the step of assigning a source fragment ion type to any file feature of the group having the mass corresponding to the mass of a predicted source fragment of another file feature in the group, and / or The method according to claim 9 or 10, wherein the step of determining one or more candidate relationships between file features of the group includes the step of determining an in-source fragment relationship between a file feature of the group and another file feature of the group, when the file feature has the mass corresponding to the mass of an expected in-source fragment of another file feature.
12. A step of obtaining the mass of a source fragment expected to have file features, The mass of the expected source fragment of the file feature is the MS corresponding to the file feature. N The method according to claim 11, further comprising the step of obtaining by determining from a mass spectrum.
13. A step of obtaining the mass of a source fragment expected to have file features, The method according to claim 11, further comprising the step of obtaining, by providing the mass of one or more expected in-source fragments of the feature as part of the identification result of the feature.
14. A method for processing mass spectral data, wherein the mass spectral data is a plurality of MS 1 Mass spectra and multiple MS N The method includes (N≧2) mass spectra, each mass spectrum having its respective associated retention time, and the method is The aforementioned multiple MS 1 A step of detecting a group of features in a mass spectrum, wherein each feature in the group has its own mass, and the features in the group have equal retention times within a first tolerance range. For each of the one or more features of the group, (i) in order to obtain the identification result for that feature, the corresponding MS N (ii) the step of putting the mass spectrum into a mass spectrum search engine, and (ii) providing the mass of one or more expected in-source fragments of the feature as part of the identification result, The steps include assigning a source fragment ion type to any feature of the group having the mass corresponding to the mass of a predicted source fragment of another feature in the group, A method comprising the step of identifying one or more compounds based on the group of features and the fragment ion type in the source.
15. The mass provided as part of the identification result is one or more MS configured to simulate in-source fragmentation. N The method according to claim 14, determined from the mass spectrum.
16. The step of determining one or more candidate relationships between file features of the group is: The method according to claim 9 or 10, further comprising the step of determining one or more candidate adduct relationships between file features of a group based on an allowable mass shift between file features of the group.
17. The step of selecting a preferred arrangement of the cluster from the one or more arrangements of the cluster is: The steps include determining the score for each cluster configuration, The method according to claim 8, comprising the step of selecting the arrangement having the highest score.
18. The step of determining a score for each cluster arrangement is: A step of determining the cluster score for each cluster in the arrangement, comprising: (i) assigning weight coefficients to each candidate ion type assignment of the cluster; (ii) assigning relation scores to each candidate relation of the cluster; and (iii) calculating the cluster score for the cluster by dividing the sum of the weight coefficients and relation scores by the number of features or file features within the cluster. The method according to claim 17, comprising the step of determining a score for each arrangement by dividing the sum of the cluster scores for the arrangement by the number of clusters in the arrangement.
19. A method for measuring mass, Analyze the sample and use multiple MS 1 Mass spectra and multiple MS N A step of acquiring mass spectral data including a mass spectrum, wherein each mass spectrum has its respective associated retention time, A method comprising the step of processing the mass spectral data using the method according to any one of claims 1, 2, 8, 9, 10, 14, 15, 17, and 18.
20. A non-temporary computer-readable storage medium for storing computer software code that, when executed on a processor, performs the method according to any one of claims 1, 2, 8, 9, 10, 14, 15, 17, and 18.
21. A control system for an analytical instrument, wherein the control system is configured to cause the analytical instrument to perform the method according to any one of claims 1, 2, 8, 9, 10, 14, 15, 17, and 18.
22. An analytical instrument comprising the control system described in claim 21.