Ms2-driven feature detection in DIA data

WO2026167600A1PCT designated stage Publication Date: 2026-08-13DH TECH DEVMENT PTE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-05
Publication Date
2026-08-13

Smart Images

  • Figure IB2026051122_13082026_PF_FP_ABST
    Figure IB2026051122_13082026_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and computer program products for processing data from a mass spectrometer. A mass spectrum dataset with data in a first dimension, a second dimension, and a third dimension is obtained and divided into one or more slices across the first dimension. An intensity distribution is generated for each of the one or more slices across at least a second dimension and, for each intensity distribution generated, one or more features are located to produce a slice feature list. One or more slice feature lists are consolidated to produce a combined feature list across all slices.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 18768.0123WOU1 / 2024-25201-P-USMS2-DRIVEN FEATURE DETECTION IN DIA DATACROSS REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 755,818, filed February 7, 2025, the disclosure of which is hereby incorporated by reference in its entirety.BACKGROUND

[0002] Feature detection in mass spectrometry (MS) data is the process of identifying and characterizing ion signals that correspond to analytes within a sample. Features are defined by attributes such as retention time, mass-to-charge ratio (m / z), and intensity. Detection algorithms analyze raw MS data to distinguish meaningful signals from noise, resolve overlapping peaks, and group related data points, such as isotopic patterns or adducts, into cohesive features. Accurate feature detection is essential for downstream analysis in applications like proteomics and metabolomics, enabling researchers to identify, quantify, and compare analytes across samples with precision and reliability.SUMMARY

[0003] Examples presented herein relate to a method for processing data from a mass spectrometer. The method includes obtaining a mass spectrum dataset with data in a first dimension, a second dimension, and a third dimension and dividing the mass spectrum dataset into one or more slices across the first dimension. An intensity distribution is generated for each of the one or more slices across at least a second dimension and, for each intensity distribution generated, one or more features are located to produce a slice feature list. One or more slice feature lists are consolidated to produce a combined feature list across all slices.

[0004] In other examples presented herein, the data comprises an intensity measurement defined by the first and second dimensions. In further examples presented herein, an intensity at any slice is determined as a combination of the intensity measurements across the first dimension slice at a coordinate in the second dimension.In further examples presented herein, the combination is sum of at least a portion of the measurements.

[0005] In other examples presented herein, the intensity distribution is generated across both the second dimension and a third dimension. In further examples presented herein, an intensity at any slice is determined as a combination of the intensity measurements across the first dimension slice at a coordinate in each of the second dimension and the third dimension. In further examples presented herein, the combination comprises a selection based on measurement intensity distribution across the first dimension at each intersection between coordinates of the second and third dimensions. In other further examples presented herein, each slice feature list comprises one or more features, wherein each feature of the one or more features comprises a local maximum of the corresponding intensity distribution. In still further examples presented herein, the combined feature list comprises one or more fragment groups, each fragment group comprising member fragment slices, each member fragment slice comprising a feature having a common coordinate in at least one of the second and third dimensions with each other member slice. In still further examples presented herein, at least one of the second and third dimensions differs among the member fragment slices by no more than a predefined amount. In other further examples presented herein, each fragment group comprises fragment slices includes features with a common property in the second dimension. In still further examples presented herein, the common features are defined by a common coordinate with a predefined delta in at least one of the second and third dimensions. In still further examples presented herein, the method further includes performing a library search using a subset of the one or more features associated with a fragment group. In other further examples presented herein, the method further includes performing a sample alignment using a subset of one of more slices and grouping information to align subset of features across multiple samples, wherein the grouping information defines one or more slices as belonging to one group in particular sample. In still further examples presented herein, the method further includes performing a statistical analysis using the subset of one or more slices and grouping information, and further using a corresponding feature intensity to determine differences between fragment groups within a sample or feature intensity trends across an experimental dimension.

[0006] In other examples presented herein, the first dimension is a fragment m / z dimension. In further examples presented herein, the one or more slices are equal sized slices. In still further examples presented herein, the one or more slices are 1 Da or narrower slices. In other examples presented herein, the second dimension is one of a QI dimension and a mobility dimension. In further examples presented herein, a third dimension is a time dimension. In still further examples presented herein, the time dimension comprises a retention time. In other examples presented herein, the second dimension is a QI dimension and a third dimension is a mobility dimension.

[0007] In other examples presented herein, the combined feature list comprises a list of fragment groups, wherein each group is associated with at least one precursor ion. In further examples presented herein, a particular fragment is associated with two or more of the fragment groups. In still further examples presented herein, the method further includes determining a confidence measure for each fragment assigned to a particular fragment group.

[0008] In other examples presented herein, the mass spectrum dataset further comprises a fourth dimension, wherein the first, second, third, and fourth dimension are each selected from a group comprising fragment m / z, precursor m / z, retention time, and mobility. In other examples presented herein, at least one of the one or more slices has a predefined center. In further examples presented herein, one or more of the one or more slices has a predefined center. In other examples presented herein, the method further includes deconvoluting the mass spectrum dataset using the combined feature list. In further examples presented herein, any fragment group is selected as a region of interest. In other further examples presented herein, the deconvolution uses a subset of the one or more slices associated with the fragment group.

[0009] Other examples presented herein relate to a system for processing data from a mass spectrometer. The system includes at least one processor and a memory in communication with the processor and having instructions. The instructions, when executed by the processor, cause the processor to: obtain a mass spectrum dataset with data in a first dimension, a second dimension, and a third dimension; divide the mass spectrum dataset into one or more slices across the first dimension; generate, for each of the one or more slices in the mass spectrum dataset, a intensity distribution across at least one of the second and third dimensions; locate, for each intensity distribution generated, one or more features to produce a slice feature list; and consolidate one or more slicefeature lists to produce a combined feature list across all slices. In other examples presented herein, the intensity distribution is generated across each of the second and third dimensions. In other examples presented herein, the second dimension is one of a QI dimension and a mobility dimension. In further examples presented herein, the third dimension is a time dimension. In other examples presented herein, the combined feature list comprises a list of fragment groups, wherein each group is associated with at least one detected precursor.

[0010] Other examples presented herein relate to a non-transitory computer-readable medium having stored thereon sequences of instructions, the sequences of instructions including instructions that when executed by a computer system causes the computer system to perform: obtaining a mass spectrum dataset with data in a first dimension, a second dimension, and a third dimension; dividing the mass spectrum dataset into one or more slices across the first dimension; generating, for each of the one or more slices in the mass spectrum dataset, a intensity distribution across at least one of the second and third dimensions; locating, for each intensity distribution generated, one or more features to produce a slice feature list; and consolidating one or more slice feature lists to produce a combined feature list across all slices.

[0011] Other examples presented herein relate to a method for processing data from a mass spectrometer. The method includes obtaining a mass spectrum dataset with data in a first dimension, a second dimension, and a third dimension; dividing the mass spectrum dataset into one or more slices across the first dimension using a fragment m / z of interest, wherein the first dimension is a fragment m / z dimension; generating, for each of the one or more slices in the mass spectrum dataset, an intensity distribution across at least one of the second and third dimensions; locating, for each intensity distribution generated, one or more features to produce a target slice feature list; and combining the target slice feature lists to create a combined feature list. In other examples presented herein, the method further includes combining one or more slices into a target slice.

[0012] Other examples presented herein relate to processing mass spectrometry data acquired by a data independent acquisition mode of the mass spectrometer. The processing including:i) evaluating the MS2 spectra to locate fragment ions detected for the DIA experiment. If required, the evaluating further includingdistinguishing fragment ions from residual precursor ions that may be captured in the MS2 spectra.ii) grouping located fragment ions to associate with an expected precursor ion transmitted across a set of precursor ion selection windows (the grouping may be further based on retention time and / or mobility selection time).iii) identifying and / or associating a transmitted precursor ion with the grouped fragment ions.

[0013] In general, fragment ions may be expected to be generated from precursor ions that may pass through a precursor selection window at a given time in the MS experiment. In this case, precursors may be identified by evaluating the intensity profiles of fragment ions that have been commonly grouped. The time and corresponding intensity profiles of the grouped fragment ions may be utilized to evaluate potential precursor ions available to transmit through the precursor selection windows operative at that time. Conveniently, the fragment ion spectra may be evaluated to identify groups of fragment ions, at a given experiment time, that may be associated with one another as potentially deriving from a common precursor ion transmitted by a precursor selection window at that experiment time. Once a group of common fragment ions has been identified from the fragment ion spectra, i.e., MS2 spectra, a potential precursor ion that is eligible to pass within the corresponding set of overlapping precursor ion selection windows may be assigned to that group of fragment ions. In this way, potential precursor ions may be identified and assigned based on the fragment ion profiles associated with a set of overlapping precursor ion selection windows.

[0014] Contrary to prior methods, which considered defined sets of selection windows to define specific precursor ions, the current method is data driven and evaluates detected fragment ions to identify a set, or sets, of precursor selection windows that may be relevant to identify the originating precursor ion.

[0015] A variety of additional inventive aspects will be set forth in the description that follows. The inventive aspects can relate to individual features and to combinations of features. It is to be understood that both the forgoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the broad inventive concepts upon which the embodiments disclosed herein are based.BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which are incorporated in and constitute a part of the description, illustrate several aspects of the present disclosure. A brief description of the drawings is as follows:

[0017] FIG. 1 is an example encoding of MSMS data suitable for application of the methods disclosed herein.

[0018] FIG. 2A is an example intensity distribution for a slice taken for the range of ~54-55 Da in the MS2 dimension

[0019] FIG. 2B is an example intensity distribution for a slice taken for the range of -142-143 Da in the MS2 dimension.

[0020] FIG. 3 is an example table demonstrating a combined feature list, according to embodiments of the present disclosure.

[0021] FIG. 4 is an example plot of the combined feature list of FIG. 3, according to embodiments of the present disclosure.

[0022] FIG. 5 is a block diagram of an example mass spectrometry system, which embodiments of the present disclosure may be implemented upon or in association with.

[0023] FIG. 6 is a block diagram of an example data processing system for processing and MS2 feature detection of mass spectrum data.

[0024] FIG. 7 is a flowchart of an example method for processing data from a mass spectrometer for MS2 feature detection.

[0025] FIG. 8 is a flowchart of another example method for processing data from a mass spectrometer for MS2 feature detection.

[0026] FIG. 9 illustrates an example block diagram of a virtual or physical computing system upon which embodiments of the present disclosure can be implemented.DETAILED DESCRIPTION

[0027] Reference will now be made in detail to exemplary aspects of the present disclosure that are illustrated in the accompanying drawings. Wherever possible, the same reference numbers will be used throughout the drawings to refer to the same or like parts.

[0028] As discussed herein, features generally refer to distinct signal patterns that represent ions or analytes within the data. A feature may be characterized by one or moredimensions of the data, including retention time, mass-to-charge ratio (m / z), mobility, and intensity, and often corresponds to an ionized molecular species or its adducts. Feature detection algorithms can identify these signals by filtering out noise, resolving overlapping peaks, and grouping related data points to define the analyte’s presence and abundance. Features may serve as the fundamental units for downstream analysis in applications like metabolomics, proteomics, and lipidomics. Discussed in further detail below are particular feature types relevant to different aspects of the disclosed methods and systems. For example, TP features refer to true positive features and FP features refer to false positive features. This categorization may apply in various implementations. In some examples, a TP feature relates to whether a feature describes or identifies a liquid chromatography (LC) peak or not. In other examples, a TP feature could be true association of a compound to a particular feature (or fragments found to be related to the feature) and an FP feature would be an incorrect association. In examples presented herein, features often refer to a local maxima of an intensity distribution.

[0029] As discussed herein, MS refers to mass spectrometry generally and MS / MS, MSMS, or tandem MS refers to tandem mass spectrometry. Tandem mass spectrometry is an analytical technique that involves multiple stages of mass analysis to study molecular ions. In MSMS, ions generated in the first stage are selected based on their mass-to-charge ratio (m / z) and may be designated as precursor ions. These precursor ions are fragmented in a controlled manner, and the resulting fragment ions are analyzed in a second mass analyzer of the mass spectrometer. The sequential stages of filtering and fragmentation may be repeated and, in some cases, filtering may be based on a mass filter or another chemical and / or physical property filter. This process provides detailed structural information about the molecule, enabling identification, characterization, and quantification of complex analytes in fields like proteomics, metabolomics, and drug discovery.

[0030] As discussed herein, MSI generally refers to the first stage of an MSMS mass analysis, where all ions generated from the sample are separated and detected based on their m / z. In some examples, an MSMS analysis may be applied with another chemical and / or physical property filter, which may occur prior to the first stage. As discussed herein, filter refers to both separation and filtration as these terms are conventionally used to discuss the division of the constituents of a sample into discrete data elements. This first stage provides an overview of the sample’s molecularcomposition. MSI features include, for example, the m / z and intensities associated with precursor ions from the first stage MSMS analysis. In another example, the MSI features are defined to further account for a retention time based on separation performed before or in conjunction with the MSI mass analysis. In another example, where mobility and time separation are used together, the MSI features are defined in dimensions of mobility, retention time, and precursor m / z and intensity.

[0031] As discussed herein, MS2 refers to the second stage of analysis in tandem mass spectrometry. A specific ion (a precursor ion) from an MSI scan is isolated and fragmented, and its fragments are analyzed to gain structural or compositional information. MS2 features are, for example, the fragment m / z and intensities associated with the fragment ions from the second stage MSMS analysis.

[0032] As discussed herein LCMS or LC / MS refers to a liquid chromatographymass spectrometry assay and resulting data. LCMS combines the separation capabilities of liquid chromatography (LC) with the detection and identification power of mass spectrometry (MS). In LC, a liquid mobile phase carries the sample through a chromatographic column, where analytes are separated based on their chemical properties, such as polarity or size. The separated compounds are then ionized and introduced into the mass spectrometer, where they are analyzed based on their mass-to-charge ratio (m / z). LCMS is widely used in fields like proteomics, metabolomics, pharmacology, and environmental analysis for its ability to analyze complex mixtures with high sensitivity and specificity. An LC dimension, as used herein, refers to the separation or time dimension associated with the LC separation and may often be expressed as a retention time (RT).

[0033] As discussed herein, data dependent acquisition (DDA) and data independent acquisition (DIA) refer to two methods of acquiring MSMS data. In DDA, the instrument selects precursor ions from an MSI scan based on predefined criteria, such as intensity, for fragmentation and analysis in MS2. This approach often focuses on the most abundant ions, providing detailed information about specific analytes but may miss low-abundance signals. DIA, on the other hand, fragments all ions within a selected m / z ranges in a systematic and unbiased manner, capturing comprehensive MS2 data for all detectable ions. The key difference is that DDA targets specific ions, while DIA collects data on all ions within defined precursor m / z windows, making DIA more suitable for reproducible, high-throughput, and comprehensive analyses.

[0034] Some DIA methods relevant to the present disclosure include sequential window DIA, where the precursor or QI m / z range is divided into discrete, fixed-width acquisition windows, and all ions within each window are fragmented and analyzed systematically, providing comprehensive MS / MS data. Another relevant method is scanning DIA, which uses a dynamic, continuously scanning MSI or QI m / z acquisition window that overlaps with adjacent acquisition windows. Information from a set of the overlapping acquisition windows may be evaluated to produce acquisition data that allows for higher resolution and smoother transitions across the m / z range. In order to evaluate the set of overlapping acquisition windows it is necessary to consider an additional dimension of data - the precursor or QI m / z position of each of the overlapping acquisition windows. The key difference is that sequential window DIA uses fixed windows, offering simplicity and data consistency, while scanning DIA employs a moving acquisition window that overlaps with adjacent acquisition windows, for improved resolution and detail across the spectrum.

[0035] As discussed herein, convolution refers to the overlapping of signals from different ions in the same mass-to-charge ratio (m / z) range due to limited resolution or the presence of ions with similar m / z values. This phenomenon can arise from isobaric compounds, isotopic variants, and co-eluting species during chromatographic separation. Convolution in MSI data can make it difficult to accurately resolve and quantify individual ions, particularly in complex mixtures. While convolution in MS2 data is associated with similar difficulties, MS2 data is generally less complex than MSI and fragment ion readings are less frequently convolved.

[0036] For example, isobaric precursors refer to ions that have nearly identical masses (same or very close mass-to-charge ratios, m / z) but originate from different molecular species. These ions present unique challenges in mass spectrometric analysis, including in methods that rely on precursor ion selection, such as tandem mass spectrometry (MS / MS) because they produce signals that tend to overlap and are difficult to resolve from one another. When isobaric ions are fragmented simultaneously in tandem mass spectrometry (e.g., during DDA), their resulting fragment ions may have their signals superimposed in one or more subsets of the MS2 spectrum. This convolution complicates the interpretation of the data, as it becomes challenging to attribute fragment ions to a specific precursor.

[0037] Most existing data processing of data dependent acquisition (DDA) and data independent acquisition (DIA) data is based on MSI features. For example, an LC / MS peak finder is a computational tool designed to identify and quantify MSI peaks within LCMS data. By analyzing the signal intensity across retention time and mass-to-charge (m / z) dimensions, it detects features that correspond to individual analytes. Peak finders filter noise, resolve overlapping signals, and refine peak shapes to ensure accurate identification. MSI data is often highly complex and therefore requires a sophisticated LCMS peak finder to detect features and, even with such refined tools, is limited by resolution of the LC dimension for resolving convolution, e.g., due to isobaric precursors.

[0038] Disclosed herein is novel data processing for MS data and, in some embodiments, for DIA data in particular. As disclosed herein and discussed in further detail throughout, using the MS2 data of MSMS scans to detect MSI features provides a number of benefits over the traditional methods which look directly at complex MSI data. For example, MS2 data generally has less convolution enabling more true positive (TP) and less false positive (FP) feature detections, even with simple peak finding algorithms.

[0039] In addition, some features might be below a detection limit in MSI due to acquisition constraints that protect the detector in a case of overabundant ions; for example, the entire ion current may be suppressed in some cases to maintain the ion rate within a safe region for the detector. While these constraints may be necessary for safe operation of the instrument, they have the consequence of losing ions with low abundance that go below the detection limit, such that some MSI features which are low abundance are not detectable given the ion statistics implied by acquisition setting.

[0040] As is disclosed herein, LCMS and / or MSI feature regions can be detected from MS2 data, instead of from complex LCMS or MSI data. In embodiments, a precursor feature region is defined by dimensions including time, precursor m / z (or a QI dimension in some DIA implementations), and associated MS2 fragments. The accuracy of the detected feature precursor m / z follows from the settings used during the data acquisition process and the encoding of the data for storage. The accuracy of the precursor m / z can be further improved in downstream data processing based on a fragment’s QI and LC dimension intensity distribution shape and / or by discovering, such as by use of an MSI or LCMS peak finder, an accurate precursor m / z when available.

[0041] FIG. 1 is an example encoding 10 of MSMS data suitable for application of the methods disclosed herein. Plot 12 shows a precursor ion 14 at m / z 15 within a precursor ion mass range. Plot 12 is one example of MSI data as discussed herein, demonstrating a precursor ion 14 signal as an intensity detected at a precursor ion m / z 15. A precursor ion transmission window 16 with an m / z width W is stepped with a step size S m / z across the mass range, producing a series of overlapping transmission windows. The example of FIG. 1 relates particularly to scanning DIA data, but those of skill in the art will recognize the broader applicability of the processes and principles disclosed herein to MS data acquired using other methods. In FIG. 1, a first appearance of a precursor ion 18 occurs in first appearance overlapping window 19, for example. The same precursor ion 18 may transmit through any of the overlapping windows 19 that include that precursor m / z.

[0042] Group 20, for example, of G overlapping windows of the series immediately preceding first appearance overlapping window 19 is selected so that group 20 spans at least the precursor ion uncertainty interval, which is the width W of transmission window 16. The number of overlapping windows G in group 20 is calculated according to G>W / S and is 8, in this example. A sum of counts or intensities of unique product ion 18 detected from each window of the G overlapping windows of group 20 is calculated. The sum 21 calculated for group 20 is shown plotted in plot 22. The sum 21 is associated with the position of a first overlapping window of group 20 and is plotted at the position of the first overlapping window of group 20 in plot 22.

[0043] A sum of counts or intensities of unique fragment ion 18 detected from each window of the G overlapping windows of group 20 is calculated. The sum 21 calculated for group 20 is shown plotted in plot 22. The sum 21 is associated with the position of a first overlapping window 19 of group 20 and is plotted at the position of the first overlapping window 19 of group 20 in plot 22. Plot 22 shows that there is a precursor ion 14 at m / z 23 (which coincides with m / z 15) within a precursor ion mass range. Plot 22 is an example ofMS2 data as discussed herein, demonstrating a fragment ion 18 signal as an intensity detected at one or more windows 16. Plot 22 presents the intensities as a sum of measured intensity across a group G, but fragment ion intensity may be recorded in other ways, such as direct measurement. Function 24 describes how summed counts or intensities of unique product ion 14 vary with the position of group 20.

[0044] FIG. 2A is an example intensity distribution 50 for a slice taken for the range of -54-55 Da in the MS2 dimension. As disclosed herein and discussed in further detail below, the MS2 dimension of the MS data is divided into one or more “slices,” with a slice covering a discrete range of the data in the MS2 dimension. For example, a range of fragment ion m / z, such as the example of FIG. 2A, where the slice covers a range of data for fragment ions with an m / z between 54 and 55 Da. The data for each slice is processed to produce an intensity distribution, like intensity distribution 50, across additional dimensions of the MS data. In the example intensity distribution 50, the additional dimensions include a precursor m / z (the y-axis) and a retention time (the x-axis). Each point of the intensity distribution, for example each pixel of the graphical example in FIG. 2A, represents a combination of the measured values across the range covered by the slice.

[0045] In embodiments, the combination may be a sum. For example, each point of intensity distribution represents a sum of the intensities measured across the range covered by the slice, e.g., the intensity measured at 54.0 Da plus the intensity measured at 54.5 Da, etc. In embodiments, the combination may be a measured maximum. For example, each point of intensity distribution represents the maximum intensity measured across the range covered by the slice, e.g., the intensity measured at 54.0 Da is greater than the intensity measured at 54.5 Da at a particular point and only the data for 54.0 Da is represented at that particular point.

[0046] Each “slice” represents a particular range of fragment m / z values. A particular slice may be associated with multiple fragment ions, each of which falls within the represented range of fragment m / z values. In other words, because each slice covers a range of fragment m / z values, rather than being limited to a particular fragment’s precise m / z, two or more fragments of different m / z may produce a feature within the same slice if their different fragment m / z’s fall within the range covered by the slice. In some cases, two (or more) fragment ions within a single slice are attributable to two (or more) different precursor ions. Fragment slice intensity is the sum of all (both in this example) fragments in the slice. In some cases, the QI position of a feature within the slice will be skewed toward a more intense fragment ion QI position.

[0047] Each ellipse 52, 58, 59 denotes a region of the intensity distribution assigned as a feature region. Ellipses 52, 58, 59 are example feature regions highlighted in this example for discussion, but other feature regions may be located in intensity distribution50. In embodiments, a feature is defined by a feature position point. The feature position point may be located at coordinates corresponding to a most intense portion in the feature region. For example, feature position point 57 is located at the most intense point of the feature region located at the ellipse 52. In embodiments, the feature position point 57 may be used to assign properties to an associated feature. In this example, feature position point 57 has a precursor m / z property of ~338 Da, based on the feature position point 57 relative to the y-axis which is associated with the precursor m / z dimension. Likewise, feature position point 57 has a retention time property of -13.72 min based on the x-axis which is associated with the time dimension. A feature associated with feature position point 57 is attributed these properties (precursor m / z of -338 Da and retention time of -13.72 min). Other properties which may be associated with the feature include, but are not limited to, a fragment m / z of 54-55 Da, based on the slice in which the feature is located, and an intensity, based on the distribution used to locate the feature position point 57, e.g., the intensity value assigned to feature position point 57 which may be a sum or a maximum, as discussed above.

[0048] Individual feature regions within an intensity distribution may be determined to have overlapping boundaries, for example as shown by the feature regions located at ellipses 58 and 59. Multiple overlapping feature regions within an intensity distribution may be distinguished from one another by the determination of distinct centers. For example, the boundaries of ellipses 58 and 59 overlap, but feature position point 58a of ellipse 58 is clearly distinguished from feature position point 59a of ellipse 59, demonstrating that two distinct features are present.

[0049] Together, all features located within a particular slice make up a slice feature list. In the example of FIG. 2A, a slice feature list associated with the slice taken for the range of 54-55 would include at least three features including features located at ellipses 52, 58, and 59. Those of skill in the art will recognize that more than three features may be located in the intensity distribution 50 depending, for example, on the method of feature detection and thresholds applied.

[0050] FIG. 2B is an example intensity distribution 60 for a slice taken for the range of -142-143 Da in the MS2 dimension. Example intensity distribution 60 includes at least two feature regions with which a feature may be associated, depicted as ellipses 61 and 62 in FIG. 2B. In this example, a feature region associated with ellipse 62 has a feature position point 63. Feature position point 63 has properties including a precursorm / z of -338 and a retention time of -13.72 min, which are substantially similar to the properties associated with feature position point 57 of FIG. 2A. These two features, one associated with feature position point 57 from the 54-55 Da slice of FIG. 2A and another associated with feature position point 63 of the 142-143 Da slice of FIG. 2B, may then be placed together in a fragment group based on these substantially similar properties. In other words, these two features together are indicative of a precursor at 338 Da and 13.72 min that produced at least two fragments — one within 54-55 Da m / z and one within 142-143 Da m / z.

[0051] In this example, the two features are grouped together, into a fragment group, on the basis of having collocated feature position points. In other examples, grouping of features across slices could be done on a basis of corresponding feature intensity distributions or other properties or characteristics.

[0052] A set of grouped features across multiple slices form a fragment group, where each slice within the group contains a feature associated a common QI and / or RT value across each slice in the group. Member features of this fragment group may then be defined by the common QI value and / or the common RT value among the fragment slices in the group. As discussed herein, a value or property of a feature is referred to as “common” with another feature (which may be located on a different slice from the first feature) if the value or property is substantially similar, e.g., differing no more than a predetermined delta or tolerance.

[0053] FIG. 3 is an example table 65 demonstrating the combined feature lists. In this example, fragment intensity distribution slices of IDa each from 40 to 1000 Da were processed together. Fragments were aligned by precursor m / z and retention time and in the resulting table 65 there is a row for each potential feature group and each column is a (nominal) fragment mass (many columns not shown). An example of a data subset which may be characterized as a fragment group is illustrated in the highlighted row 67 in FIG. 3. The table 65 provides a common feature QI value in column “m / z” 69, common feature RT value in column “Ret. Time” 71, and list of slice m / z’s (shown in column headers 73a-d) and feature intensities (shown in the cells) in the corresponding slice columns 73a-d. In the figure, for highlighted row 67, only cells with non-zero intensity (74b, 74d) are part of the corresponding feature group. The QI value in column m / z 69 can be interpreted as a low resolution precursor m / z and the RT value in column Ret. Time 71 can be interpreted as a precursor retention time. Each slice m / z in slicecolumns 73a-d with an intensity greater than zero can be interpreted as a low resolution fragment m / z and fragment intensity reading.

[0054] Each feature group corresponding to a QI and / or retention time-region may be processed by data deconvolution, e.g., untargeted data deconvolution. In this way, “blind” untargeted deconvolution can be performed much faster than is possible with current approaches. Because the majority of analytes will contain at least one fragment without any interference, this fragment can provide a starting point (e.g., an associated retention time, precursor m / z, etc.) for each precursor in the sample. In contrast, MSI data may often have interference of varying degrees and so no clear starting point is present.

[0055] FIG. 4 is an example plot 70 of an MS2-driven feature group according to embodiments of the present disclosure. Plot 70 displays an example graphical representation of the data associated with an MS2-driven feature group, such as the fragment group data associated with highlighted row 67 of table 65 (as discussed above). This example MS2-driven feature group has a precursor m / z at about 243.9, known with approximately nominal mass accuracy with this analysis, in the top left header 75 and corresponding feature intensities (y-axis) at corresponding feature slice m / z (x-axis). Each intensity of plot 70 represents the value from a particular slice, e.g., 1 Da range, intensity value associated with the corresponding feature such as seen in FIGS. 2A and 2B.

[0056] Output of the disclosed MS2-feature finder can be used, for example, for improved or accelerated deconvolution of MS data. In embodiments, MS2-feature finder results are used to detect a region of interest (ROI) in the data. This ROI provides a starting point and enables a faster and more efficient blind deconvolution compared to using all possible Ql-RT coordinates in the data.

[0057] In embodiments, deconvolution of the detected ROI itself is done using fragment slice information in the deconvolution process. For example, the fragment slice information may be used to reduce the size of fragment arrays used as input for deconvolution. In another example, the fragment slice information is used to find representative Ql-RT distributions that can be used in turn to assist deconvolution. For example, representative Ql-RT distributions can be used either as a postprocessing step of blind deconvolution at a Ql-RT site or to create targets for deconvolution.

[0058] Results of the MS2-feature finding can also be used directly as a low resolution MSMS spectra for conducting a library search. For example, for each feature detected, a list of fragments that contributed to the corresponding feature region may be created and used for subsequent analysis. The list of fragments could function as an equivalent to a low resolution DDA spectra. The library search may then be conducted on the list of features using appropriate precursor and fragment m / z tolerances. This is advantageously faster than conventional search methods and, in many cases, may be sufficient for compound identification.

[0059] While some examples presented herein focus on use of fragment groups themselves, in embodiments, further grouping is possible. Two or more fragment groups can be aligned using, for example, a substantially common retention time into “RT groups.” For a given RT group, the collected of precursor m / z’s, associated with each of the fragment groups within the RT group, may be characterized as a low-resolution precursor (MSI) spectrum. In embodiments, this spectrum is used to find chemical relationships amongst the collected precursors. Non-limiting examples of chemical relationship include adducts and in-source fragmentation (e.g., a mass difference of ~22 Da between two peaks may indicate that the lower m / z is from an M+H+ ion and the other M+Na+). Compared to the conventional method of using MSI directly, an advantage of this approach is a cleaner precursor spectrum which only contains m / z(s) which are correlated with one another in the RT dimension. Another advantage is the spectrum produced as disclosed herein has fragments formed at low collision energies eliminated. This analysis can also take advantage of the fragment m / z for each fragment group. For example, referring back to the M+Na+ example above, if the putative M+H+ and M+Na+ have at least some fragment m / z in common the hypothesis is more likely to be correct. This analysis can be done directly with the low-resolution spectrum as discussed above or in combination with a separate high resolution MSI dataset to enable a more accurate precursor m / z to be assigned to the precursor spectrum. It would also be possible to interrogate any high-resolution MS2 data associated with the separate high resolution MSI dataset to assign accurate fragment m / z’s.

[0060] Results of the MS2-feature finder can be used to align samples for a particular purpose, such as statistical analysis. For example, a list of all fragment groups across all samples or a subset of samples may be created, where each fragment group is represented by a QI value, an RT value, and a set of fragment slices and their intensities(or at least one fragment slice from that fragment group). Alignment may use a tolerance (e.g., greater than or equal to 0) for each coordinate and a representative coordinate, where the representative coordinate relates to a candidate precursor ion species. The representative coordinate could be selected as a coordinate from a representative sample or a virtual coordinate obtained from statistical analysis across multiple samples.

[0061] In embodiments, starting from a most intense feature among a set of features (e.g., all features or a subset) and defining a most intense fragment group coordinate may simplify the selection of a representative coordinate. All fragment groups, or a subset, within a coordinate tolerance to the most intense fragment group coordinate may then be extracted. In this case, the tolerance could be larger but, because it is done on fragment level and in a three dimensional space (including fragment m / z), accuracy is improved, even with the large tolerance, versus a traditional two dimensional space (e.g., using only precursor m / z and retention time). Another benefit is that even if a fragment group is not found in any particular sample, there will generally be enough samples to select a representative coordinate from at least one sample or calculate the representative coordinate based on coordinate distribution and their attributes, like the intensity at those coordinates.

[0062] Another benefit of this MS2-fragment group based alignment is that every fragment group is represented with one or more fragments or fragment slices. Instead of considering all MS2 features, alignment is done on representative features only, while grouping information is used to manage representative MS2 fragments in the entire set of samples. This is beneficial to manage the size of the list of representative coordinates, as well as utilizing only the best (e.g., most intense, most pure, most unique, etc.) for a particular analysis or application.

[0063] In embodiments, statistical analysis using feature intensities across samples (e.g., after alignment) may also be performed. Some non-limiting examples of statistical analysis include a sample normalization step in combination with uni- or multi-variate analysis or sample clustering.

[0064] FIG. 5 is a block diagram of an example mass spectrometry system 100, which embodiments of the present disclosure may be implemented upon or in association with. In embodiments, other systems for determining compound structure and identity may be integrated with a mass spectrometry system or, in other embodiments, may be independent of the mass spectrometry system. Example system 100 includes an ionsource 110, a first mass separator or mass filter 120, a fragmentation device 130, a second mass separator or a mass analyzer 140, and a computing system 150.

[0065] In embodiments, system 100 further includes a sample introduction device 160. Sample introduction device 160 introduces one or more compounds of interest from a sample to ion source 110 over time. Sample introduction device 160 performs techniques that include, but are not limited to, direct injection, liquid chromatography, gas chromatography, capillary electrophoresis, or ion mobility.

[0066] Mass filter 120 and fragmentation device 130 are shown as different stages of a quadrupole and mass analyzer 140 is shown as a time-of-flight (TOF) device. Those of ordinary skill in the art will appreciate that either of mass filter 120 and mass analyzer 140 may include other types of mass separator and analysis devices including, but not limited to, ion traps, orbitraps, ion mobility devices, time-of-flight (TOF) devices, or Fourier transform ion cyclotron resonance (FT-ICR) devices. In embodiments, mass filter 120 and mass analyzer 140 are respective examples of a first and a second mass separator, arranged in a series. A system may be configured according to the present disclosure with a quadrupole for the first mass separator, or mass filter 120, and a TOF device for the second mass separator, or mass analyzer 140. Each mass separator is configured to receive a set of ions, perform a detection of the set of ions, and generate a set of detection signals corresponding to detection of the set of ions.

[0067] Ion source device 110 transforms a sample or compounds of interest from a sample into an ion beam. Ion source device 110 can perform ionization using techniques that include, but are not limited to, matrix assisted laser desorption / ionization (MALDI) or electrospray ionization (ESI).

[0068] Mass filter 120 receives the ion beam. In embodiments, mass filter 120 is configured by a user for a particular precursor ion transmission window based on the experimental goals for the sample being run. The precursor ion transmission window, as discussed herein, refers to the range of precursor ions that are allowed to pass through a specific selection step and into the subsequent stages of mass analysis or fragmentation. In many tandem mass spectrometry (MS / MS) experiments, the precursor ions are first filtered based on their m / z (mass-to-charge ratio) in order to isolate a specific ion or group of ions of interest for further analysis or fragmentation. The precursor ion selection process employs a mass filter or a specific set of voltages that allow only ions within acertain m / z range (the precursor ion transmission window) to pass through to the next stage.

[0069] Fragmentation device 130 of tandem mass spectrometer 102 fragments or transmits the precursor ions transmitted at each overlapping step by mass filter 120. In examples related to scanning DIA, one or more resulting fragment ions are produced for each overlapping window of the series. Fragmentation device 130 fragments the precursor ions when a collision energy high enough to fragment ions is used. Fragmentation device 130 transmits the precursor ions when a collision energy low enough not to fragment ions is used. As a result, the resulting fragment ions can include precursor ions.

[0070] In mass spectrometry, dissociation mechanisms are techniques used to fragment molecules into smaller pieces to facilitate their analysis and identification. For example, collision energy is used to influence the fragmentation of ions during collision-induced dissociation (CID) or collision-induced fragmentation (CIF). This process is commonly used in MS / MS to provide structural information about a molecule by inducing the dissociation of precursor ions into fragment ions. The setting of the collision energy impacts the resulting fragmentation pattern. Some other examples of dissociation mechanisms include electron transfer dissociation, electrospray ionization, matrix-assisted laser desorption / ionization, infrared multiphoton dissociation, and sustained off-resonance irradiation. Each dissociation method offers varying advantages depending on the type of analysis being performed and the nature of the sample. The choice of dissociation technique may impact the type of information obtained from mass spectrometric analysis.

[0071] Mass analyzer 140 of tandem mass spectrometer 102 detects intensities or ion counts for each of the one or more resulting fragment ions for each overlapping window of the series that form mass spectrum data for each overlapping window of the series.

[0072] Computing system 150 can be, but is not limited to, a computer, a microprocessor, the computing system of FIG. 9, or any device capable of sending and receiving control signals and data from a tandem mass spectrometer and processing data. Computing system 150 is in communication with ion source device 110, mass filter 120, fragmentation device 130, and mass analyzer 140. Computing system 150 is shown as a separate device but can be a processor or controller of tandem mass spectrometer 102 oranother device. Computing system 150 may store in a memory device (not shown) mass spectrum data for each precursor ion window analysis is performed for, including for each overlapping window of the series in examples performing scanning DIA.

[0073] In embodiments, computing system 150 instead performs an encoding and storing step, and encodes and stores each unique fragment ion detected by mass analyzer 140 in real-time during data acquisition. Prior to storing mass spectrum data, computing system 150 performs one or more processing steps on the raw mass spectrum data received to prepare the data for viewing, analysis, and storage. Raw mass spectrum data includes the counts or intensities of fragment ions at different m / z ratios over time. Computing system 150 may be further configured to retrieve stored mass spectrum data for subsequent, additional processing and analysis, and / or transmission to additional, remote, or otherwise separate computing systems.

[0074] FIG. 6 is a block diagram of an example data processing system 200 for processing and MS2 feature detection of mass spectrum data. Data processing system 200 is implemented, in embodiments, by computing system 150. Data processing system 200 may be among multiple subsystems or software executed by computing system 150 in operating mass spectrometer 102 and handling of data output by the mass spectrometer. The present disclosure is directed to data processing for MS2 feature detection and analysis, but those of skill in the art will readily understand that other subsystems and / or modules may exist within and be executed by computing system 150. Computing system 150 is presented in examples herein as a single device, but in embodiments may be one or more processing devices networked or otherwise in communication. Functions may be divided among individual devices or shared across the collective processing capability of the one or more processing devices. In embodiments, data processing system 200 forms part of or serves as a controller. The controller may be configured to transmit operation commands to the mass spectrometer and / or to receive unprocessed mass data from second mass separator including the set of detection signals and perform one more data processing actions on the unprocessed mass data. In embodiments, data processing system 200 may be fully remote from and independent of the mass spectrometry system 100.

[0075] In the example of FIG. 6, data processing system 200 includes storage 204 and feature finder 210, which further includes data divider 212, slice feature finder 214, feature consolidator 216, and display 250. In embodiments, data processing system 200may further include components, not shown, for additional functions. Some non-limiting examples of additional functionality and components which data processing system 200 may incorporate include a confidence measurer and a deconvoluter.

[0076] Data, including mass spectrometry data 202 and derivative processing outputs of mass spectrometry data 202, may be passed freely among the components of data processing system 200. Various outputs of the analysis and processing components may be stored in storage 204, or pass to display 250.

[0077] Storage 204 receives mass spectrometry data 202 from mass spectrometer 102 for storage. In embodiments, mass spectrometry data 202 may instead be received from a computing system, such as another computing system 150, and undergo processing and / or preprocessing before being received by storage 204. Mass spectrometry data 202 may be processed and encoded, including being compressed, prior to being received and stored at storage 204.

[0078] Feature finder 210 may receive mass spectrometry data 202 from storage 204, or from mass spectrometer 102 or another computing system, and process the mass spectrometry data 202 for MS2 feature detection. Feature finder 210 may further include data divider 212, slice feature finder 214, and feature consolidator 216 to perform operations as part of the MS2 feature detection process. Feature finder 210 may further process MS2 feature detection analysis and results for display on display 250 to provide a user visualization of the quantification analysis and results.

[0079] Data divider 212 receives, retrieves, or otherwise obtains mass spectrometry data 202, such as by retrieval from storage 204. Data divider 212 divides the mass spectrometry data 202 into one or more slices within a first dimension. In embodiments, the first dimension is the MS2 or fragment m / z dimension, such that each of the one or more slices is defined by a range of fragment m / z values, e.g., Da. For example, slices may be of equal sizes and each slice may be 5 Da, 2 Da, 1 Da, 0.5 Da, etc. In some examples, slices may not be of equal sizes and data 202 may be divided into one or more slices of differing sizes. In embodiments, slices may be defined by an upper and lower boundary or by a central point and a plus / minus delta. For example, all MS2 data of the data 202 may be divided into slices or a number of slices may be taken to coincide with a number of fragments of interest. In the latter case, for example, a pattern of five fragments may be of interest and five slices, each with a central point that coincides with a fragment m / z of one of the five fragments of the pattern of interest may be taken by thedata divider 212. Each of the five slices may have a same plus / minus delta such that each of the five slices is equal in size, or deltas may vary among the five slices based on, for example, expected or observed data in a particular slice. The size of each slice may be determined automatically or may be set by user input. In embodiments, slice sizes may be automated by default and responsive to adjustment by manual user input.

[0080] Slice feature finder 214 generates intensity distributions, such as intensity distributions 50 and 60 of FIGS. 2A and 2B, for each slice produced by data divider 212. Intensity distributions map combined fragment ion intensity across a second and third dimension, different from the first dimension along which the slices are divided. For example, where slices are divided along the MS2 dimension, intensity distributions may be generated with precursor ion m / z as a first axis and retention time as a second axis of the distribution. Combined fragment ion intensity may be a sum of the measured intensities at each point in the intensity distribution, or each point may represent a max value measured across the fragment ion m / z range of the slice.

[0081] Slice feature finder 214 may further perform a feature detection analysis on the intensity distribution. For example, slice feature finder 214 may scan each slice or one or more target slices for local maximums and identify each such maximum as a feature region. For example, a feature region may be defined as an area of high intensity in a given slice, as multiple adjacent points of high intensity, as a pattern of increasing intensity, etc. Slice feature finder 214 further generates a feature list for each slice and associates the list with the slice. The feature list includes a summary of the features, the feature regions, and / or the feature position points identified by slice feature finder 214 for the particular slice.

[0082] Feature consolidator 216 combines features from one more feature lists prepared by slice feature finder 214. For example feature consolidator 216 reviews a range of feature lists for features or feature position points with common coordinates, e.g., a same precursor m / z and / or a same retention time (within a specified tolerance). Features with such common coordinates are grouped into fragment groups, along with properties of the data of the slice from which the feature is drawn. In embodiments, a fragment group is associated with a particular precursor ion.

[0083] FIG. 7 is a flowchart of an example method 300 for processing data from a mass spectrometer for MS2 feature detection. Method 300 may be executed by acomputing device 150 or may be executed by a separate or integrated processing system independent of computing device 150.

[0084] At operation 302, a mass spectrum dataset is obtained. In embodiments, the dataset includes data in a first dimension, a second dimension, and a third dimension. The multidimensional mass spectrum dataset of FIG. 1 is an example of a dataset appropriate for processing by method 300.

[0085] At operation 304, the mass spectrum dataset is divided into one or more slices across the first dimension. In embodiments, the first dimension is the MS2 or fragment m / z dimension. In embodiments, the one or more slices are equal sized slices. For example, the one or more slices may be 1 Da slices.

[0086] In embodiments, at least one of the one or more slices has a predefined center or one or more of the one or more slices has a predefined center. Rather than defining an upper and lower boundary of the slice by a range, a desired central point of the slice may be defined and each of the upper and lower boundary set an equal and predetermined delta from the desired central point. For example, the example of FIG. 2A is discussed above as an intensity distribution 50 slice defined by a range of ~55-56 Da. Alternatively, this intensity distribution 50 slice may be defined with a predefined center of ~55.5 Da and a delta of 0.5 Da.

[0087] At operation 306, an intensity distribution is generated, for each of the one or more slices in the mass spectrum dataset, across at least one of the second and third dimensions. In embodiments, the intensity distribution is generated across each of the second and third dimensions for each slice. In embodiments, the second dimension is one of a QI dimension and a mobility dimension and the third dimension is a time dimension. For example, the time dimension in some cases is a retention time. See, for example, the example intensity distribution 50 of FIG. 2A which is generated across a QI dimension on the y-axis and a time dimension on the x-axis. In embodiments, the intensity distribution is generated across three dimensions, each of a QI dimension, a mobility dimension, and a time dimension.

[0088] At operation 308, one or more feature regions are located, for each intensity distribution generated, to produce a slice feature list. In embodiments, each slice feature list includes one or more features, wherein each feature of the one or more features is defined using a coordinate calculation method. In an example, each feature region and / or a feature position point is located as a local maximum of the corresponding intensitydistribution. In other examples, each feature is a calculated as a centroid, using an intensity weighted average, or using another statistical method or heuristic rule. In embodiments, as seen in the example intensity distribution 50 of FIG. 2A, the intensity distribution can be presented as a contour or heat map with intensities represented by a depth or hue of color so that local maximums are visually apparent due to their color. In some embodiments, a contour or heat map is generated as part of underlying data processing of the method, but the map is not displayed to a user.

[0089] At operation 310, one or more slice feature lists are consolidated to produce a combined feature list across all slices. In embodiments, the combined feature list includes one or more fragment groups. In embodiments, each group is associated with at least one precursor ion. As discussed herein, a fragment group relates fragments associated with a common precursor ion or precursor ion range. While, in some cases, “precursor ion” alone may indicate a positive and accurate identification of a precursor ion m / z, in some cases fragment groups may instead relate to a coarsely defined precursor ion m / z, or precursor ion range, which may not provide a final accurate identification of precursor ion m / z.

[0090] Each fragment group is made up of one or more features with common feature properties. Together, the common features of a fragment group define a set of data characteristics indicative of a precursor ion. Each feature in the fragment group has at least one coordinate, in at least one of the second and third dimensions (e.g., QI or RT dimensions), which is the same as, within a range or tolerance, the other features in the fragment group. In other words, a first feature is located on a first slice and associated with a precursor m / z and a retention time based on its orientation in the slice. A second feature located in a second slice is similarly oriented in relation to the QI dimension, and therefore is associated with a same (within a tolerance) precursor m / z as the first feature. In embodiments, the second feature may also be associated with a same RT as the first feature. The first and second slices are each defined by a different fragment m / z range, and each of the first and second features are associated with the respective different fragment m / z range based on the respective slice. Based on this substantially similar orientation in the QI and RT dimensions, the first and second feature are defined as a common feature. As a result, the associated first and second features are placed together in a fragment group, which may in turn be associated with a particular precursor ion. Inembodiments, each fragment group includes features (from the slices) having common coordinates in one or both the second and third dimensions.

[0091] In embodiments, the common features are defined by a common coordinate, within a predefined delta, in at least one of the second and third dimensions. In other words, a target may be defined with coordinates in terms of one or both of the second and third dimensions, e.g., a precursor m / z and / or a retention time, and a range established around that target with a plus and / or minus of a predefined delta. Features with a coordinate falling within that range are determined to be common features with are presumed to originate from a same precursor ion, and in response the associated slices are placed together in a fragment group. In embodiments, at least one of the second and third dimensions differs from the other by a predefined amount.

[0092] In embodiments, a particular fragment may be associated with two or more of the fragment groups. For example, based on a feature’s coordinates in either of the second or third dimension, e.g., QI or time dimensions, a particular feature may fall within a target range or predefined delta associated with more than one precursor ion.

[0093] At operation 312, a confidence measure is optionally determined for each fragment assigned to a particular fragment group. The confidence measure may provide a quantification of the likelihood that a fragment associated with the fragment group, and by extension the associated precursor ion, is correct. The confidence measure may be determined, for example, by applying weights according to measured coordinates of a feature and their nearness to expected coordinates associated with the precursor ion.

[0094] At operation 314, the mass spectrum dataset is optionally deconvoluted using the combined feature list. As discussed above, locations of a fragment or MS2 feature which appear without convolution can provide a starting point for removing convolution from more complex portions of the data. In embodiments, deconvolution may be performed by selecting any fragment group as a region of interest. The selected fragment group will have a certain number of the fragment slices associated with the group, referred to herein as a subset of the slices. The deconvolution may then proceed using the subset of the one or more slices associated with the fragment group. The subset may also, or instead, be used to perform a library search to determine candidate identities for the precursor and / or fragments associated with the subset.

[0095] FIG. 8 is a flowchart of another example method 400 for processing data from a mass spectrometer for targeted MS2 feature detection. Targeted here means thegoal is not to find all MS2 features or precursors in the data, but instead to focus on a subset of features or data with relevance to particular biological question / experiment. For example, one might be interested in finding all precursors associated with characteristic fragment or combination of fragments. Method 400 may be executed to detect presence in a sample or distribution across sample, of a specific combination of fragments. As discussed herein, a “specific combination” could be used to determine how a slice is defined or the way features are combined to create feature groups in combined feature list. Method 400 may be executed by a computing device 150 or may be executed by a separate or integrated processing system independent of computing device 150.

[0096] At operation 402, a mass spectrum dataset is obtained. In embodiments, the dataset includes data in a first dimension, a second dimension, and a third dimension. The multidimensional mass spectrum dataset of FIG. 1 is an example of a dataset appropriate for processing by method 400.

[0097] At operation 404, the mass spectrum dataset is divided into one or more slices using fragment m / z’s of interest to produce slices of interest. In embodiments, fragment m / z’s of interest may be selected based on relevance to particular biological question / experiment and / or responsive to analyzed characteristics of the dataset. In embodiments, fragment m / z’s of interest may be selected based on manual input from a user or may be automatically selected by the system based on a predetermined criteria. The predetermined criteria may be manually defined by a user or may be determined by the system based on one or more characteristics of the dataset.

[0098] In embodiments, the mass spectrum dataset is divided into a number of slices to coincide with the number of fragments of interest. In the latter case, for example, a pattern of five fragments may be of interest and five slices, each with a central point that coincides with a fragment m / z of one of the five fragments of the pattern of interest may be taken. These five slices may be sized to span the full mass spectrum dataset or may only cover a subset of the mass spectrum dataset. Each of the five slices may have a same plus / minus delta such that each of the five slices is equal in size, or deltas may vary among the five slices based on, for example, expected or observed data in a particular slice. The size of each slice may be determined automatically or may be set by user input. In embodiments, slice sizes may be automated by default and responsive to adjustment by manual user input.

[0099] In embodiments, the one or more slices are equal sized slices. For example, the one or more slices may be 1 Da slices. In some examples, slices may not be of equal sizes and data may be divided into one or more slices of differing sizes. In embodiments, slices may be defined by an upper and lower boundary or by a central point and a plus / minus delta.

[0100] At operation 405, one or more slices of interest are optionally combined to form target slices. Target slices may be single slices of interest or a combination of data from two or more slices of interest. At operation 406, an intensity distribution is generated, for the target slices, across at least one of the second and third dimensions. In embodiments, the intensity distribution is generated across each of the second and third dimensions for each target slice. In some examples, the intensity distribution will contain data from one or more of the slices of interest forming the combined target slice. An intensity distribution spanning multiple slices of interest may provide useful perspective on a pattern or relationship among the fragments of interest.

[0101] In embodiments, the second dimension is one of a QI dimension and a mobility dimension and the third dimension is a time dimension. In embodiments, the intensity distribution is generated across three dimensions, each of a QI dimension, a mobility dimension, and a time dimension. See, for example, the example intensity distribution 50 of FIG. 2A which is generated across a QI dimension on the y-axis and a time dimension on the x-axis.

[0102] At operation 408, one or more feature regions are located, for each intensity distribution generated, to produce a target slice feature list. In embodiments, each target slice feature list includes one or more features, wherein each feature of the one or more features is defined using a coordinate calculation method. In an example, each feature region and / or a feature position point is located as a local maximum of the corresponding intensity distribution. In other examples, each feature is a calculated as a centroid, using an intensity weighted average, or using another statistical method or heuristic rule. In embodiments, as seen in the example intensity distribution 50 of FIG. 2A, the intensity distribution can be presented as a contour or heat map with intensities represented by a depth or hue of color so that local maximums are visually apparent due to their color. In some embodiments, a contour or heat map is generated as part of underlying data processing of the method, but the map is not displayed to a user.

[0103] At operation 409, one or more target slice feature lists are consolidated to produce a combined feature list across all target slices. In embodiments, the combined feature list includes one or more fragment groups. In embodiments, each group is associated with at least one precursor ion or precursor compound. As discussed herein, a fragment group relates fragments associated with a common precursor ion or precursor ion range. Multiple fragment groups, associated with different precursor ions, may be considered together to characterize more complex compounds within the sample.

[0104] FIG. 9 illustrates an example block diagram of a virtual or physical computing system upon which embodiments of the present disclosure can be implemented. The example computing system 150 includes one or more processors 152, a system memory 158, and a system bus 172 that couples the system memory 158 to the one or more processors 152. The system memory 158 includes RAM (Random Access Memory) 161 and ROM (Read-Only Memory) 162. A basic input / output system that contains the basic routines that help to transfer information between elements within the computing system 150, such as during startup, is stored in the ROM 162. The computing system 150 further includes a mass storage device 164. The mass storage device 164 is able to store software instructions and data. Mass storage device 164 may correspond to storage 204 of FIG. 6. The one or more processors 152 can be one or more central processing units or other processors.

[0105] The mass storage device 164 is connected to the one or more processors 152 through a mass storage controller (not shown) connected to the system bus 172. The mass storage device 164 and its associated computer-readable data storage media provide nonvolatile, non-transitory storage for the computing system 150. Although the description of computer-readable data storage media contained herein refers to a mass storage device, such as a hard disk or solid state disk, it should be appreciated by those skilled in the art that computer-readable data storage media can be any available non-transitory, physical device or article of manufacture from which the central display station can read data and / or instructions.

[0106] Computer-readable data storage media include volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storage of information such as computer-readable software instructions, data structures, program modules or other data. Example types of computer-readable data storage media include, but are not limited to, RAM, ROM, EPROM, EEPROM, flash memory or othersolid state memory technology, CD-ROMs, DVD (Digital Versatile Discs), other optical storage media, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by the computing system 150.

[0107] According to various embodiments of the invention, the computing system 150 may operate in a networked environment using logical connections to remote network devices through the network 148. The network 148 is a computer network, such as an enterprise intranet and / or the Internet. The network 148 can include a LAN, a Wide Area Network (WAN), the Internet, wireless transmission mediums, wired transmission mediums, other networks, and combinations thereof. The computing system 150 may connect to the network 148 through a network interface unit 154 connected to the system bus 172. It should be appreciated that the network interface unit 154 may also be utilized to connect to other types of networks and remote computing systems. The computing system 150 also includes an input / output controller 156 for receiving and processing input from a number of other devices, including a touch user interface display screen, or another type of input device. Similarly, the input / output controller 156 may provide output to a touch user interface display screen or other type of output device.

[0108] As mentioned briefly above, the mass storage device 164 and the RAM 161 of the computing system 150 can store software instructions and data. The software instructions include an operating system 168 suitable for controlling the operation of the computing system 150. The mass storage device 164 and / or the RAM 161 also store software instructions, that when executed by the one or more processors 152, cause one or more of the systems, devices, or components described herein to provide functionality described herein. For example, the mass storage device 164 and / or the RAM 161 can store software instructions that, when executed by the one or more processors 152, cause the computing system 150 to receive and execute managing network access control and build system processes.

[0109] The software instructions further include one or more software applications 166. Software applications may include dedicated systems and algorithms for performing specific tasks or actions or providing specific interfaces. One or more of data processing system 200 and / or one or more component of data processing system 200 may be encompassed by software applications 166.

[0110] Illustrative examples of the systems and methods described herein are provided below. An embodiment of the system or method described herein may include any one or more, and any combination of, the aspects described below.

[0111] Aspect 1. A method for processing data from a mass spectrometer, the method including: obtaining a mass spectrum dataset with data in a first dimension, a second dimension, and a third dimension; dividing the mass spectrum dataset into one or more slices across the first dimension; generating, for each of the one or more slices in the mass spectrum dataset, an intensity distribution across at least a second dimension; locating, for each intensity distribution generated, one or more features to produce a slice feature list; and consolidating one or more slice feature lists to produce a combined feature list across all slices.

[0112] Aspect 2. The method of aspect 1, wherein the data includes an intensity measurement defined by the first and second dimensions.

[0113] Aspect 3. The method of aspect 2, wherein an intensity at any slice is determined as a combination of the intensity measurements across the first dimension slice at a coordinate in the second dimension.

[0114] Aspect 4. The method of aspect 3, wherein the combination is sum of at least a portion of the measurements.

[0115] Aspect s. The method of any one of aspects 1-4, wherein the intensity distribution is generated across both the second dimension and a third dimension.

[0116] Aspect 6. The method of aspect 5, wherein an intensity at any slice is determined as a combination of the intensity measurements across the first dimension slice at a coordinate in each of the second dimension and the third dimension.

[0117] Aspect ?. The method of aspect 6, wherein the combination includes a selection based on measurement intensity distribution across the first dimension at each intersection between coordinates of the second and third dimensions.

[0118] Aspect 8. The method of any one of aspects 5-7, wherein each slice feature list includes one or more features, wherein each feature of the one or more features includes a local maximum of the corresponding intensity distribution.

[0119] Aspect 9. The method of aspect 8, wherein the combined feature list includes one or more fragment groups, each fragment group includes member fragment slices, each member fragment slice including a feature having a common coordinate in at least one of the second and third dimensions with each other member slice.

[0120] Aspect 10. The method of aspect 9, wherein at least one of the second and third dimensions differs among the member fragment slices by no more than a predefined amount.

[0121] Aspect 11. The method of any one of aspects 8-10, wherein each fragment group includes fragment slices includes features with a common property in the second dimension.

[0122] Aspect 12. The method of aspect 11, wherein the common features are defined by a common coordinate with a predefined delta in at least one of the second and third dimensions.

[0123] Aspect 13. The method of aspect 11 or 12, further comprising performing a library search using a subset of the one or more features associated with a fragment group.

[0124] Aspect 14. The method of any one of aspects 11-13, further including performing a sample alignment using a subset of one of more slices and grouping information to align subset of features across multiple samples, wherein the grouping information defines one or more slices as belonging to one group in particular sample.

[0125] Aspect 15. The method of aspect 14, further comprising performing a statistical analysis using the subset of one or more slices and grouping information, and further using a corresponding feature intensity to determine differences between fragment groups within a sample or feature intensity trends across an experimental dimension.

[0126] Aspect 16. The method of any one of aspects 1-15, wherein the first dimension is a fragment m / z dimension.

[0127] Aspect 17. The method of aspect 16, wherein the one or more slices are equal sized slices.

[0128] Aspect 18. The method of aspect 17, wherein the one or more slices are 1 Da or narrower slices.

[0129] Aspect 19. The method of any one of aspects 1-18, wherein the second dimension is one of a QI dimension and a mobility dimension.

[0130] Aspect 20. The method of claim 19, wherein a third dimension is a time dimension.

[0131] Aspect 21. The method of claim 20, wherein the time dimension includes a retention time.

[0132] Aspect 22. The method of any one of aspects 1-21, wherein the second dimension is a QI dimension and a third dimension is a mobility dimension.

[0133] Aspect 23. The method of any one of aspects 1-22, wherein the combined feature list includes a list of fragment groups, wherein each group is associated with at least one precursor ion.

[0134] Aspect 24. The method of aspect 23, wherein a particular fragment is associated with two or more of the fragment groups.

[0135] Aspect 25. The method of aspect 24, further including determining a confidence measure for each fragment assigned to a particular fragment group.

[0136] Aspect 26. The method of any one of aspect 1 -25, wherein the mass spectrum dataset further includes a fourth dimension, wherein the first, second, third, and fourth dimension are each selected from a group comprising fragment m / z, precursor m / z, retention time, and mobility.

[0137] Aspect 27. The method of any one of aspects 1-26, wherein at least one of the one or more slices has a predefined center.

[0138] Aspect 28. The method of aspect 27, wherein one or more of the one or more slices has a predefined center.

[0139] Aspect 29. The method of any one of aspects 1-28, further including deconvoluting the mass spectrum dataset using the combined feature list.

[0140] Aspect 30. The method of aspect 29, wherein any fragment group is selected as a region of interest.

[0141] Aspect 31. The method of aspect 29 or 30, wherein the deconvolution uses a subset of the one or more slices associated with the fragment group.

[0142] Aspect 32. A system for processing data from a mass spectrometer, the system including: at least one processor; a memory in communication with the processor and having instructions which, when executed by the processor, cause the processor to: obtain a mass spectrum dataset with data in a first dimension, a second dimension, and a third dimension; divide the mass spectrum dataset into one or more slices across the first dimension; generate, for each of the one or more slices in the mass spectrum dataset, a intensity distribution across at least one of the second and third dimensions; locate, for each intensity distribution generated, one or more features to produce a slice feature list; and consolidate one or more slice feature lists to produce a combined feature list across all slices.

[0143] Aspect 33. The system of aspect 32, wherein the intensity distribution is generated across each of the second and third dimensions.

[0144] Aspect 34. The system of aspect 32 or 33, wherein the second dimension is one of a QI dimension and a mobility dimension.

[0145] Aspect 35. The system of aspect 34, wherein the third dimension is a time dimension.

[0146] Aspect 36. The system of any one of aspects 32-35, wherein the combined feature list includes a list of fragment groups, wherein each group is associated with at least one detected precursor.

[0147] Aspect 37. A non-transitory computer-readable medium having stored thereon sequences of instructions, the sequences of instructions including instructions that when executed by a computer system causes the computer system to perform: obtaining a mass spectrum dataset with data in a first dimension, a second dimension, and a third dimension; dividing the mass spectrum dataset into one or more slices across the first dimension; generating, for each of the one or more slices in the mass spectrum dataset, a intensity distribution across at least one of the second and third dimensions; locating, for each intensity distribution generated, one or more features to produce a slice feature list; and consolidating one or more slice feature lists to produce a combined feature list across all slices.

[0148] Aspect 38. A method for processing data from a mass spectrometer, the method including: obtaining a mass spectrum dataset with data in a first dimension, a second dimension, and a third dimension; dividing the mass spectrum dataset into one or more slices across the first dimension using a fragment m / z of interest, wherein the first dimension is a fragment m / z dimension; generating, for each of the one or more slices in the mass spectrum dataset, an intensity distribution across at least one of the second and third dimensions; locating, for each intensity distribution generated, one or more features to produce a target slice feature list; and combining the target slice feature lists to create a combined feature list.

[0149] Aspect 39. The method of aspect 38, further including combining one or more slices into a target slice.

[0150] This disclosure described some examples of the present technology with reference to the accompanying drawings, in which only some of the possible examples were shown. Other aspects can, however, be embodied in many different forms andshould not be construed as limited to the examples set forth herein. Rather, these examples were provided so that this disclosure was thorough and complete and fully conveyed the scope of the possible examples to those skilled in the art.

[0151] Although specific examples were described herein, the scope of the technology is not limited to those specific examples. One skilled in the art will recognize other examples or improvements that are within the scope of the present technology. Therefore, the specific structure, acts, or media are disclosed only as illustrative examples. Examples according to the technology may also combine elements or components of those that are disclosed in general but not expressly exemplified in combination, unless otherwise stated herein. The scope of the technology is defined by the following claims and any equivalents therein.

Claims

What is claimed is:

1. A method for processing data from a mass spectrometer, the method comprising:obtaining a mass spectrum dataset with data in a first dimension, a second dimension, and a third dimension;dividing the mass spectrum dataset into one or more slices across the first dimension;generating, for each of the one or more slices in the mass spectrum dataset, an intensity distribution across at least a second dimension;locating, for each intensity distribution generated, one or more features to produce a slice feature list; andconsolidating one or more slice feature lists to produce a combined feature list across all slices.

2. The method of claim 1, wherein the data comprises an intensity measurement defined by the first and second dimensions.

3. The method of claim 2, wherein an intensity at any slice is determined as a combination of the intensity measurements across the first dimension slice at a coordinate in the second dimension.

4. The method of claim 3, wherein the combination is sum of at least a portion of the measurements.

5. The method of any one of claims 1-4, wherein the intensity distribution is generated across both the second dimension and a third dimension.

6. The method of claim 5, wherein an intensity at any slice is determined as a combination of the intensity measurements across the first dimension slice at a coordinate in each of the second dimension and the third dimension.

7. The method of claim 5 or 6, wherein each slice feature list comprises one or more features, wherein each feature of the one or more features comprises a local maximum of the corresponding intensity distribution.

8. The method of claim 7, wherein the combined feature list comprises one or more fragment groups, each fragment group comprising member fragment slices, each member fragment slice comprising a feature having a common coordinate in at least one of the second and third dimensions with each other member slice.

9. The method of claim 7 or 8, wherein each fragment group comprises fragment slices includes features with a common property in the second dimension and the method further comprises performing a sample alignment using a subset of one of more slices and grouping information to align the subset of features across multiple samples, wherein the grouping information defines one or more slices as belonging to one group in a particular sample.

10. The method of claim 9, further comprising performing a statistical analysis using the subset of one or more slices and grouping information, and further using a corresponding feature intensity to determine differences between fragment groups within a sample or feature intensity trends across an experimental dimension.

11. The method of any one of claims 1-10, wherein the first dimension is a fragment m / z dimension and the one or more slices are equal sized slices.

12. The method of any one of claims 1-11, wherein the second dimension is one of a QI dimension and a mobility dimension.

13. The method of claim 12, wherein a third dimension is a time dimension.

14. The method of any one of claims 1-13, wherein the combined feature list comprises a list of fragment groups, wherein each group is associated with at least one precursor ion.

15. The method of claim 14, wherein a particular fragment is associated with two or more of the fragment groups.

16. The method of any one of claims 1-15, wherein at least one of the one or more slices has a predefined center.

17. The method of any one of claims 1-16, further comprising deconvoluting the mass spectrum dataset using the combined feature list.

18. A system for processing data from a mass spectrometer, the system comprising:at least one processor;a memory in communication with the processor and having instructions which, when executed by the processor, cause the processor to:obtain a mass spectrum dataset with data in a first dimension, a second dimension, and a third dimension;divide the mass spectrum dataset into one or more slices across the first dimension;generate, for each of the one or more slices in the mass spectrum dataset, an intensity distribution across at least one of the second and third dimensions;locate, for each intensity distribution generated, one or more features to produce a slice feature list; andconsolidate one or more slice feature lists to produce a combined feature list across all slices.

19. A method for processing data from a mass spectrometer, the method comprising:obtaining a mass spectrum dataset with data in a first dimension, a second dimension, and a third dimension;dividing the mass spectrum dataset into one or more slices across the first dimension using a fragment m / z of interest, wherein the first dimension is a fragment m / z dimension;generating, for each of the one or more slices in the mass spectrum dataset, an intensity distribution across at least one of the second and third dimensions;locating, for each intensity distribution generated, one or more features to produce a target slice feature list;andcombining the target slice feature lists to create a combined feature list.

20. The method of claim 19, further comprising combining one or more slices into a target slice.