Data structures for rapid extraction of spectra and chromatograms
By constructing data structures that include transposed matrices as an index, the method addresses the limitations of existing data structures in handling high-dimensional mass spectrometry data, achieving efficient storage and rapid data extraction for enhanced analysis capabilities.
Patent Information
- Application Number
- PCT/IB2024/061367
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-16
- Filing Date
- 2024-11-14
- Publication Date
- 2025-05-22
AI Technical Summary
Existing mass spectrometry data structures are inadequate for efficiently storing and extracting high-dimensional data, such as 4D mobility and scanning data, which limits the ability to rapidly access and analyze complex multi-dimensional data sets.
The method involves constructing data structures by recording mass spectrometry data in a first matrix with dimensions corresponding to time and mass-to-charge, transposing it into a second matrix with reversed dimensions, and storing both matrices as an index, allowing for rapid extraction of data across multiple dimensions.
This approach enables efficient storage and rapid extraction of high-dimensional mass spectrometry data, improving performance and enabling access to numerous dimensions of raw data, thereby facilitating advanced data analysis and interpretation.
Smart Images

Figure IB2024061367_22052025_PF_FP_ABST
Abstract
Description
DATA STRUCTURES FOR RAPID EXTRACTION OF SPECTRA AND CHROMATOGRAMSCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application is being filed as a PCT International application and claims the benefit of and priority to U.S. Provisional Application No. 63 / 599,614, filed November 16, 2023, the disclosure of which is hereby incorporated by reference in its entirety.BACKGROUND
[0002] Mass spectrometry is a powerful analytical technique used to identify and quantify molecules based on their mass and charge. It works by ionizing a sample, separating the resulting ions based on their mass-to-charge ratio, and then detecting and measuring the abundance of these ions, typically represented by peaks in the data. This information can be used to determine the composition and structure of molecules, making mass spectrometry a valuable tool in various scientific fields, including chemistry, biochemistry, and environmental science.
[0003] Mass spectrometry generates a wealth of data, and preserving it accurately is essential for reproducibility, quality control, and further investigations. The specific format and tools used for storing mass spectrometry data may vary depending on the research field, laboratory setup, or data volume.SUMMARY
[0004] Examples presented herein relate to a method of constructing mass spectrometry data structures for increased high-dimensional extraction. The method includes obtaining mass spectrometry data including intensities for one or more product ions and an offset for each of the intensities and recording the intensities for the one or more product ions in a first m x n matrix with an n-dimension corresponding to a time dimension and an m-dimension corresponding to a mass-to-charge dimension. The method further includes transposing the first m x n matrix into a second m x n matrix with an n-dimension corresponding to a mass-to-charge dimension and an m-dimensioncorresponding to a time dimension and storing each of the first and second m x n matrices as an index of the mass spectrometry data.
[0005] In other examples presented herein, the method further includes receiving a selection of one of the first and second m x n matrices. In further examples presented herein, the method further includes, in response to receiving the selection of one of the first and second m x n matrices, using the selected matrix to extract a portion of the mass spectrometry data. In other further examples presented herein, the method further includes receiving a request to extract a portion of the mass spectrometry data relating to a particular product ion of the one or more product ions, in response to receiving the request to retrieve the portion of the mass spectrometry data, searching one of the first and second m x n matrices, based on the received request, for the portion of the mass spectrometry data, and retrieving the portion of the mass spectrometry data based on the searching.
[0006] In other examples presented herein, the request to retrieve a portion of the mass spectrometry data includes a search criteria. In further examples presented herein, one of the first and second m x n matrices is selected to be searched based on the search criteria defining the particular product ion.
[0007] In yet other examples presented herein, the intensities are recorded sequentially within the n-dimension of the first m x n matrix. In further examples presented herein, the time dimension comprises the offset. In other further examples presented herein, the intensities are recorded in a mxl matrix that is stored in the n- dimension of the first m x n matrix.
[0008] In still other examples presented herein, the method further includes transposing the first m x n matrix into a third m x n matrix with an n-dimension corresponding to a mass-to-charge dimension and an m-dimension corresponding to one of collision energy, charge series offset, and mobility. In other examples presented herein, the method further includes transposing the first m x n matrix into a third m x n matrix with an n-dimension corresponding to a mass-to-charge dimension and an m- dimension corresponding to a quadrupole dimension, wherein the quadrupole dimension characterizes a series of overlapping precursor ion transmission windows.
[0009] In other examples presented herein, the mass spectrometry data is obtained from a mass spectrometer. In still other examples presented herein, the mass spectrometry data is obtained from a storage device. In yet other examples presented herein, the mass spectrometry data is obtained from an independent processing system.
[0010] Other examples presented herein relate to a method for extracting high dimensional features from mass spectrometry data. The method includes obtaining mass spectrometry data including intensities for one or more product ions and an offset for each of the intensities and recording the intensities for the one or more product ions in a first m x n matrix with an n-dimension corresponding to a time dimension and an m- dimension corresponding to a mass-to-charge dimension. The method further includes transposing the first m x n matrix into a second m x n matrix with an n-dimension corresponding to a mass-to-charge dimension and an m-dimension corresponding to a time dimension and storing each of the first and second m x n matrices as an index of the mass spectrometry data. The method also includes receiving a request to retrieve a portion of the mass spectrometry data relating to a particular product ion of the one or more product ions, the request including a search criteria, in response to receiving the request to retrieve the portion of the mass spectrometry data, extracting the portion of data from one of the first and second m x n matrices, selected based on the search criteria, and retrieving the portion of the mass spectrometry data based on the searching.
[0011] In other examples presented herein, the search parameter is a time parameter and, based on the search parameter being a time parameter, the second m x n matrix is used for extracting the portion of the data. In yet other examples presented herein, the search parameter is a mass-to-charge parameter and, based on the search parameter being a mass-to-charge parameter, the first m x n matrix is used for extracting the portion of the data.
[0012] Other examples presented herein relate to a system for constructing mass spectrometry data structures for increased high-dimensional extraction. The system includes a processor and a non-transitory memory in communication with the processor and storing instructions. The instructions, when executed, cause the processor to: obtain mass spectrum data including intensities for each of one or more product ions for each window of a series of overlapping precursor ion transmission windows; discretize the mass spectrum data such that the series of overlapping precursor ion transmission windows is divided into a series of discrete precursor ion transmission windows; record the intensities for the one or more product ions in a first m x n matrix with an n-dimension corresponding to a time dimension and an m-dimension corresponding to a mass-to- charge dimension; transpose the first m x n matrix into a second m x n matrix with an n- dimension corresponding to a mass-to-charge dimension and an m-dimensioncorresponding to a time dimension; and store each of the first and second m x n matrices as an index of the mass spectrometry data.
[0013] In other examples presented herein, at least one of the first and second m x n matrices is a sparse matrix. In yet other examples presented herein, each of the first and second m x n matrices is a sparse matrix.
[0014] A variety of additional inventive aspects will be set forth in the description that follows. The inventive aspects can relate to individual features and to combinations of features. It is to be understood that both the forgoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the broad inventive concepts upon which the embodiments disclosed herein are based.BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings, which are incorporated in and constitute a part of the description, illustrate several aspects of the present disclosure. A brief description of the drawings is as follows:
[0016] FIG. 1 is a block diagram of an example system for construction and use of data structures for a mass spectrometer.
[0017] FIG. 2 is block diagram of an example data processing system for constructing data structures for storing mass spectrometry data for high-dimensional extraction.
[0018] FIG. 3 is a block diagram of an example layered architecture of an example system for constructing data structures for storing mass spectrometry data for highdimensional extraction.
[0019] FIG. 4 is a flowchart of an example method of constructing mass spectrometry data structures for increased high-dimensional extraction.
[0020] FIG. 5 is a flowchart of an example method of high-dimensional extraction of mass spectrometry data.
[0021] FIG. 6 illustrates an example computing system with which aspects of the present disclosure may be implemented.DETAILED DESCRIPTION
[0022] Disclosed herein are methods and systems for structuring mass spectrometry data for rapid and multi-dimensional extraction of spectra and chromatograms. According to the present disclosure, a matrix used to record data from one or more massspectrometer experiments is transposed or otherwise transformed. The transformed matrix and the original matrix are linked through an index, and each in turn provide indexing of unique dimensions of the mass spectrometry data. The combined indexing allows rapid and effective searching of numerous dimensions of the mass spectrometry data. Multiple matrices may be generated according to desired dimensionality of the data including, for example, dimensions for collision energy and mobility.
[0023] Data output from a mass spectrometer is generally stored in a linear manner which requires the serial reading of data files. Raw mass spectrometry data is generally linearly streamed to disk where the scan data is compressed and written into a direct file. Each specific “scan,” or record of information about the masses and intensities of ions present in a sample, is nominally written to an index in the format of an offset from the start of the file. The offset refers to a numerical value that specifies a position within a file. It indicates the distance from the beginning of the file to a specific location where the data for a particular scan begins, identifying where within the file the data for a specific scan can be found. In this way, the index information (e.g., the offset) is used to locate and access the data for a specific scan within the data file. This is a common storage format for all forms of mass spectrometry data, and is used across many vendorspecific formats, as well as for vendor-neutral formats such as mzML or mzXML.
[0024] The index structure of mass spectrometry data is generally organized in a two-dimensional matrix, where one axis represents mass-to-charge ratios (m / z values), and the other axis represents the offset, which identifies the specific scan and relates to the retention time of the data points. This matrix structure is commonly used in various mass spectrometry techniques, such as liquid chromatography-mass spectrometry (LC- MS) and gas chromatography-mass spectrometry (GC-MS).
[0025] Each column in the matrix corresponds to a specific m / z value or mass range. These columns represent the range of masses that the mass spectrometer is scanning during the analysis. Each row in the matrix represents a specific time point or scan number during the analysis as an offset. This axis represents the progression of time or scan events as data is acquired continuously. In liquid chromatography mass spectrometry, this typically corresponds to retention time, while in gas chromatography mass spectrometry, it might represent scan number or time. Each cell in the matrix contains the intensity value or count of the ions detected at a specific m / z value and retention time (or scan number). These intensity values or counts represent the abundance or signal strength of ions detected at that specific mass and time point.
[0026] The matrix structure is particularly useful for handling complex data sets generated by mass spectrometry experiments, especially in metabolomics, proteomics, and other fields directed to the detailed analysis of complex molecules. It allows researchers to efficiently manage and analyze large amounts of data, identify peaks corresponding to specific compounds, and perform statistical analyses to draw meaningful conclusions about the composition of samples or the behavior of analytes over time.
[0027] However, with advances in mass spectrometry methods has come the inception of increasingly multi-dimensional data, referring to dimensions in the data beyond the classically applied dimension of time, m / z, and intensity. The recent burst in complexity of mass spectrometry data generation methods for data independent acquisition, and other methods, requires that the data be stored in a manner which will allow for the rapid expansion of dimensions of data.
[0028] This complex multi-dimensional data, such as 4-dimensional (4D) mobility and scanning data, require that the mechanism by which data is stored is reconsidered. Mobility data in mass spectrometry pertains to ion mobility measurements, which help resolve ions based on their shape and size, while scanning data involves the systematic measurement of ion intensities at different m / z values, facilitating the identification and quantification of chemical compounds in a sample. These two types of data are complementary and are often used in combination to gain comprehensive insights into complex mixtures. 4D mobility data is an enhancement of standard mobility data which incorporates an additional dimension, typically related to m / z or another property, to provide more detailed information about the ions being analyzed.
[0029] The rigidity of the present data format restricts the use of higher order dimensionality of data extraction. To address this and other limitations of the existing data structures for mass spectrometry data, disclosed herein are systems and method for modulation of data storage through an index of multiple linked matrices with diverse dimensionality, making it possible to ensure that the access to data is as fast as possible and that the raw data is stored in the most efficient manner. Further, the present disclosure enables the extraction of data in a higher performance manner through the creation of higher dimensional data matrices.
[0030] Numerous advantages are achieved by the data construction and structures of the present disclosure. Overall performance of data extraction and manipulation is improved, and extractions other than time based linear extraction of data is enabled inorder to provide rapid access to numerous dimensions of the raw data. This will provide a meta data record of associations in a manner which will allow for rapid result extraction.
[0031] FIG. 1 is a block diagram of an example system 100 for construction and use of data structures for a mass spectrometry. Example system 100 includes an ion source 110, a first mass separator 120, a fragmentation device 130, a second mass separator or a mass analyzer 140, and a computing system 150.
[0032] In embodiments, system 100 further includes a sample introduction device 170. Sample introduction device 170 introduces one or more compounds of interest from a sample to ion source 110 over time. Sample introduction device 170 performs techniques that include, but are not limited to, direct injection, liquid chromatography, gas chromatography, capillary electrophoresis, or ion mobility.
[0033] Mass filter 120 and fragmentation device 130 are shown as different stages of a multiple quadrupole device and mass analyzer 140 is shown as a time-of-flight (TOF) device. Those of ordinary skill in the art will appreciate that either of mass filter 120 and mass analyzer 140 may include other types of mass separator and analysis devices including, but not limited to, ion traps, orbitraps, ion mobility devices, time-of- flight (TOF) devices, or Fourier transform ion cyclotron resonance (FT-ICR) devices. In embodiments, mass filter 120 and mass analyzer 140 are respective examples of a first and a second mass separator, arranged in a series. The second mass separator may configured to be faster than the first mass separator. For example, a system may be configured according to the present disclosure with a quadrupole for the first mass separator, or mass filter 120, and a TOF device for the second mass separator, or mass analyzer 140. Each mass separator is configured to receive a set of ions, perform a detection of the set of ions, and generate a set of detection signals corresponding to detection of the set of ions.
[0034] Ion source device 110 transforms a sample or compounds of interest from a sample into an ion beam. Ion source device 110 can perform ionization techniques that include, but are not limited to, matrix assisted laser desorption / ionization (MALDI) or electrospray ionization (ESI).
[0035] Mass filter 120 receives the ion beam. In embodiments, mass filter 120 is configured by a user for a particular precursor ion transmission window based on the experimental goals for the sample being run. The precursor ion transmission window, as discussed herein, refers to the range of precursor or parent ions that are allowed to pass through a specific selection step and into the subsequent stages of mass analysis orfragmentation. In many tandem mass spectrometry (MS / MS) experiments, the precursor ions are first selected based on their m / z (mass-to-charge ratio) in order to isolate a specific ion of interest for further analysis or fragmentation. The precursor ion selection process employs a mass filter or a specific set of voltages that allow only ions within a certain m / z range (the precursor ion transmission window) to pass through to the next stage.
[0036] The precursor ion transmission window is typically defined by setting specific parameters within the mass spectrometer’s control software. The width of the precursor ion transmission window affects the specificity of the analysis with a narrower window providing higher specificity, while a wider window may allow more ions to pass through but with potentially less selectivity. Balancing these factors is crucial for achieving the desired level of analytical sensitivity and specificity in a given experiment. Further, ensuring accurate execution of the selected window by the mass spectrometer is essential to accurate results, as discussed in further detail below.
[0037] In embodiments, such as those implementing a scanning data independent acquisition such as for example, where a mass filter 120 filters the ions by moving a precursor ion transmission window with a precursor ion mass-to-charge ratio (m / z) width in overlapping steps across a precursor ion mass range of R m / z with a step size S m / z. A series of overlapping transmission windows are produced across the mass range. Mass filter 120 transmits precursor ions within the transmission window at each overlapping step. In some embodiments, this refers to a scanning sequential window acquisition of all theoretical fragment ion spectra (SWATH) method.
[0038] Fragmentation device 130 of tandem mass spectrometer 102 fragments or transmits the precursor ions transmitted at each overlapping step by mass filter 120. In examples related to scanning data independent acquisition method such as scanning SWATH, one or more resulting product ions are produced for each overlapping window of the series. Fragmentation device 130 fragments the precursor ions when a collision energy high enough to fragment ions is used. Fragmentation device 130 transmits the precursor ions when a collision energy low enough not to fragment ions is used. As a result, the resulting product ions can include precursor ions.
[0039] Mass analyzer 140 of tandem mass spectrometer 102 detects intensities or counts for each of the one or more resulting product ions for each overlapping window of the series that form mass spectrum data for each overlapping window of the series. Mass analyzer 140 detects counts if it is a TOF device as shown in the example of FIG.1. If mass analyzer 140 is instead another type of mass analyzer, such as a quadrupole, for example, it detects intensities.
[0040] Computing system 150 can be, but is not limited to, a computer, a microprocessor, the computing system of FIG. 6, or any device capable of sending and receiving control signals and data from a tandem mass spectrometer and processing data. Computing system 150 is in communication with ion source device 110, mass filter 120, fragmentation device 130, and mass analyzer 140. Computing system 150 is shown as a separate device but can be a processor or controller of tandem mass spectrometer 102 or another device. Computing system 150 may store in a memory device (not shown) mass spectrum data for each precursor ion window analysis is performed for, including for each overlapping window of the series in examples performing scanning data independent acquisition including scanning SWATH. In embodiments, computing system 150 instead performs an encoding and storing step, and encodes and stores each unique product ion detected by mass analyzer 140 in real-time during data acquisition. Prior to storing mass spectrum data, computing system 150 performs one or more processing steps on the raw mass spectrum data received to prepare the data for viewing, analysis, and storage. Raw mass spectrum data includes the counts or intensities of product ions at different m / z ratios over time.
[0041] FIG. 2 is block diagram of an example data processing system 200 for constructing data structures for storing mass spectrometry data for high-dimensional extraction. In embodiments, data processing system 200 is a component of or otherwise operates within computing system 150. Data processing system 200 may be a separate and / or independent system which communicates with computing system 150 via a network. Data processing system 200 includes one or more components to perform processing, storage, and extraction operations on mass spectrometry data 202. Example components, which may be implemented as physical hardware, software, or some combination, includes preprocessor 204, matrix generator 206, and extractor 208. Data processing system 200 communicates with a storage device 212, where index 214 and one or more matrices 216 are stored.
[0042] Raw mass spectrometer data is typically large and complex, and it undergoes extensive data processing and analysis to extract meaningful information. Data processing may be carried out in a series or group of actions. Actions may be performed collectively by the data processing system 200 or individual steps or portions of the processing may be executed by individual components or modules of the data processingsystem 200. In some embodiments, some components or features shown as integrated with data processing system 200 may instead or in addition be executed elsewhere on an external component.
[0043] Preprocessor 204 acts as a raw data processor and performs one or more processing operations on mass spectrometry data 202 to prepare the data for indexing, storage, and subsequent analysis. Before indexing, the mass spectrometry data often undergoes preprocessing, which includes, by example, data conversion, noise reduction, peak picking, and deconvolution. This step helps simplify the data and enhances the quality of the information to be indexed. Peaks, representing ions and their corresponding intensities at specific m / z values, are identified and quantified. The detected peaks are typically used as the basis for indexing.
[0044] Matrix generator 206 generates one or more matrices 216 of the mass spectrometry data. The matrices may be interlinked by index 214 and stored in storage device 212. In embodiments, matrix generator represents a matrix processor.
[0045] In the example of FIG. 2, matrix generator 206 receives preprocessed mass spectrometry data from preprocessor 204. However, those of skill in the art will understand that, in embodiments, matrix generator 206 and preprocessor 204 may be a same component or algorithm. In embodiments, some or all of the functions of preprocessor 204 are performed in a physically distinct or otherwise separate device from matrix generator 206 and matrix generator 206 obtain the processed mass spectrometry data via a wired or wireless network connection. In embodiments, matrix generator 206 stores the raw spectrometry data, prior to preprocessing.
[0046] In an example, a mass spectrometry data matrix has one or more rows, with each row representing a single linearly recorded mass spectrum or MS scan. The matrix has one or more columns, with each column corresponding to a specific m / z value. The matrix elements contain the intensities or counts of ions at their respective m / z values for a given mass spectrum or MS scan.
[0047] In this example, each row represents a different scan by the mass spectrometer, which is some cases may each represent a different sample, analyte, or compound, and each column represents the intensity of ions at specific m / z values. The actual data values may be numerical, representing the intensity of ions at those m / z values for each analyte. In embodiments, data values may include pointers, such as byte-based pointers, indicating an original scan from which the data was extracted.
[0048] An offset may be included as metadata to indicate the interval between data points. This offset is used to capture the time element of each scan. The data is stored in a time-based manner with the offset, or the location in the file, recording a time associated with the particular scan and / or measurement taken. In this way, the rows also relate to a time-dimension within the matrix. While the data is extracted on a time basis, each scan is generally treated as an independent time stream, or an independent one row matrix of the m / z values. Functionality for looking across multiple scans or experiments, rather than the range of m / z values detected, is limited.
[0049] In some cases of more complex data, such as SWATH and scanning SWATH, this format provides inefficient search capabilities. In SWATH, a series of scans across multiple precursor ion transmission windows are conducted in sequence. The term "windows" refers to specific mass-to-charge ratio (m / z) ranges or segments that are sequentially isolated and analyzed by the mass spectrometer. The mass spectrometer systematically cycles through each of the selected windows one at a time, isolating the precursor ions within the defined m / z range. The isolated precursor ions are then fragmented and the resulting fragment ions detected. Counts and / or intensities are recorded in a data-independent manner where all fragment ions generated within a specific window are collected without bias. The process is repeated for each window in a serial manner, covering the entire mass range of interest. This sequential acquisition of data in multiple windows generates a comprehensive dataset that includes information about all detectable ions. The collected data can then be used for various purposes, such as protein identification, quantification, and profiling in proteomics or metabolite identification and quantification in metabolomics.
[0050] The sequential isolation of the windows is referred to herein as a “QI dimension,” referring to the common practice of using a first quadrupole in a tandem mass spectrometer to execute the scan across the range of windows. However, as will be understood by those of skill in the art, other types of mass filters are able to perform thesequential window scan and therefore the QI dimension refers to a first mass filter in a tandem mass spectrometer but is not limited to being a quadrupole.
[0051] The QI dimension is a particularly useful dimension of SWATH data, however existing data structures do not support rapid or efficient extraction of desirable data profiles, such as spectra or chromatograms from across a SWATH scan, e.g., the series of ion transmission windows. As is disclosed herein, the matrix as discussed above is transposed, with the transposed matrix having one or more rows with each row corresponding to the m / z dimension. The transposed matrix also has one or more columns with each column corresponding to a window of the QI dimension.
[0052] In this example implementation, where in one dimension of a matrix is tandem mass spectrometry m / z and another dimension of the matrix is a quadrupole dimension. This transposed index enables the software to extract a different dimension of data.
[0053] This example looks in particular at SWATH and generating a matrix with a QI dimension, but those of skill in the art will understand the same principles are applicable to transforming the matrix or generating additional matrices to capture other dimensions of the mass spectrometry data. The QI dimension is one transformation which may be utilized but there are a number of different transpositions which could be applied to expand the scan matrix which provide a linear view of data in a new dimension which is useful for result extractions. Another implementation is to create a potential charge series offset which provide a link direct to the charge states of different compounds in the data files, or to create another matrix index which provides direct evidence of clustering of spectra. Other or additional matrices with dimensions are possible, such as dimensions of mobility or collision energy.
[0054] In embodiments, matrix generator 206 further generates an index 214 of the one or more matrices 216. Index 214 interrelates the individual matrices 216. Index 214 provides for access to each, all, of some combination of matrices 216 as appropriate for a particular data extraction.
[0055] Extractor 208 provides access to each of the matrices 216 and index 214. In embodiments, extractor 208 further provides search capabilities for targeted extraction of data from matrices 216. Data processing software is used to extract information from the matrix. Peak detection algorithms are applied to identify and quantify the peaks in the data, which represent the presence of specific compounds or analytes in the sample. The software may integrate the peak areas, perform background subtraction, and apply various corrections to improve data quality. Visualization tools are used to display the mass spectrometry data in a format that researchers can interpret easily. Heatmaps, chromatograms, and 2D plots are common ways to visualize the matrix data, making it easier to identify patterns and analyze the data.
[0056] FIG. 3 is a block diagram of an example layered architecture 300 of an example system for constructing data structures for storing mass spectrometry data for high-dimensional extraction. A layered architecture assists in the organization of complex computing systems by organizing components and distinct functionalities together. Architecture 300 includes an analysis layer 302, a data layer 304, and a utility layer 306. Other layers not shown include various support functionalities such as a hardware layer, an operating system layer, and one or more application layers. Each layer includes one or more components or coding, such as Application Programming Interface (APIs), to execute the assigned functionality.
[0057] Analysis layer 302 provides tools and algorithms for the analysis of the mass spectrometry data. This layer includes modules for various tasks, such as preprocessing 322, peak identification and detection 324, and charge state determination 326. It also offers functionalities for data processing and transformation, such as data filtering and baseline correction 328. These modules are presented as examples of the functionality of analysis layer 302, but it will be understood that other modules, algorithms, and functions are compatible with envisioned to be included within analysis layer 302. For example, analysis layer 302 may contain further modules to provide support for the identification and quantification of peptides and proteins in mass spectrometry experiments. Features for searching and scoring peptide sequences against mass spectrometry data may be included.
[0058] Data layer 304 provides tools and algorithms for managing and manipulating the mass spectrometry data files. It may support a wide range of mass spectrometry file formats, including mzML, mzXML, mzData, and more. This layer enables storage of mass spectrometry data and further enables users to read and write data in these formats,making it possible to interface with various mass spectrometry instruments and software. In embodiments, it also includes features for data conversion, data format conversion, and data access. Data layer 304 of the example of FIG. 3 includes at least modules or coding for matrix formation342, indexing 344, and extraction 346.
[0059] Matrix formation 342 generates one or more matrices of the mass spectrometry data based on the dimensions of the data. In some cases, a matrix may be generated in response to a requests for data extraction, in which the matrix is generated to provide a linear record of the data in a dimensions associated with the request. In embodiments, some of all of the matrices are generated from linearly recorded data. Indexing 344 records and interrelates the one or more matrices. Extraction 346 retrieves and returns portions of the mass spectrometry data in response to user requests. One or more of the matrices may be searched and the data extracted, based on the request.
[0060] Utility layer 306 is a set of auxiliary tools and libraries that support various functionalities within architecture 300. These utilities assist with user-initiated tasks such as file input / output 362, data visualization 364, and integration with external software 366. Utility layer 306 includes libraries 368 for interfacing with programming languages, allowing development of custom applications and scripts for mass spectrometry data analysis. It may also offer utilities for logging, error handling, and other common software development tasks. Some optional tools in utility layer 306 may include file format conversion and use of command-line access to mass spectrometry data.
[0061] FIG. 4 is a flowchart of an example method 400 of constructing mass spectrometry data structures for increased high-dimensional extraction. Method 400 may be executed by a system such as data processing system 200 of FIG. 2.
[0062] At operation 402, mass spectrometry data is obtained. The mass spectrometry data may be obtained directly from a mass spectrometer. For example, data processing system 200 of FIG. 2 demonstrates mass spectrometry data received from a mass spectrometer and incorporates data processing for the raw data. In embodiments, the data is obtained from a storage device or from a separate independent system. For example, preprocessing may occur externally by a separate, independent system and may then be transmitted to the system executing method 400. In embodiments, the data is retained in storage to be retrieved by the system executing method 400. In embodiments, the mass spectrometry data includes counts and / or intensities for one or more product ions and an offset for each of the intensities.
[0063] At operation 404, the mass spectrometry data is recorded in a first matrix. In embodiments, recording the mass spectrometry data includes recording the intensities for the one or more product ions. The first matrix, in embodiments, is a m x n matrix, with n rows and m columns. In some embodiments, an n-dimension corresponds to a time dimension and an m-dimension corresponds to a mass-to-charge dimension. In embodiments, the intensities are recorded sequentially within the m-dimension of the first m x n matrix. The intensities may be recorded in a 1 x m matrix that is stored in the n-dimension of the first m x n matrix. In some implementations, the time dimension comprises an offset within the file. The offset may be recorded as metadata associated with each row or each column, or each data cell in the matrix.
[0064] At operation 406, the first matrix is transposed into a second matrix. In embodiments, the first m x n matrix is transposed into a second m x n matrix with different dimensions. For example, in the second m x n matrix, an n-dimension corresponds to a mass-to-charge dimension and an m-dimension corresponds to a time dimension. Advantageously, transposing the first matrix is accomplished without requiring new code to be written to enable generation of the second matrix.
[0065] In embodiments, at least one of the first and second m x n matrices is a sparse matrix. In some cases, each of the first and second m x n matrices is a sparse matrix. An advantage of the disclosed system and method for providing multiple matrices to linearly record numerous dimensions of the mass spectrometry data, is a reduced need to store a dense matrix of the mass spectrometry data. Because individual matrices provide a record of the various dimensions of the data, individual matrices can be recorded as a sparse matrix without data loss. Searching and extraction therefore occurs primarily in sparse matrices, reducing time and processing power needed for each extraction.
[0066] The example method 400 presented a first matrix and second transformed matrix, but in embodiments either of the first and second matrix may be transformed further into a third m x n matrix. The third matrix may have, for example, an n-dimension corresponding to a mass-to-charge dimension and an m-dimension corresponding to one of collision energy, charge series offset, mobility, etc.
[0067] In another example, the third matrix has an n-dimension corresponding to a mass-to-charge dimension and an m-dimension corresponding to a quadrupole dimension. The quadrupole dimension refers to the QI dimension of a scanning SWATH experiment as discussed above, where the quadrupole dimension characterizes a series of overlapping precursor ion transmission windows.
[0068] At operation 408, each of the first and second m x n matrices are stored to an index of the mass spectrometry data. This index and matrices may be stored together in a common storage device or may be distributed across more than once storage device.
[0069] FIG. 5 is a flowchart of an example method 500 of high-dimensional extraction of mass spectrometry data from data structures. Method 500 may be executed by a system such as data processing system 200 of FIG. 2.
[0070] At operation 502, a request is received to extract a portion of the mass spectrometry data. In an example, the portion of the mass spectrometry data relates to a particular product ion of one or more product ions in the mass spectrometry data. In embodiments, one or one or more mass spectrometry data storage matrices are searched according to method 500, in response to receiving the request to retrieve the portion of the mass spectrometry data.
[0071] The searching may be based upon the received request, such that a search criteria is selected or used based on the request or a particular storage matrix is targeted for the search based on the request or the search parameter. The portion of the mass spectrometry data is then retrieved based on the searching, as discussed further below. In embodiments, the request to retrieve a portion of the mass spectrometry data includes a search criteria.
[0072] At operation 504, a selection of a mass spectrometry data storage matrix is received. In embodiments, the mass spectrometry data storage matrix is one of the first and second m x n matrices, as discussed above in reference to FIG. 4. Selection of the matrix to be searched may be based on a matrix selection received from a user. In embodiments, selection of the matrix to be searched may an automatic determination of the system based on a search criteria in the request. For example, one of the first and second m x n matrices may be selected to be searched based on the search criteria such that if the criteria is a time-based criteria the second matrix is selected.
[0073] At operation 506, the requested portion of the mass spectrometry data is extracted using the selected matrix. In embodiments, the use of the selected matrix to extract the mass spectrometry data is in response to receiving the selection of the mass spectrometry data storage matrix.
[0074] In embodiments, method 400 of constructing the matrices and index, and method 500 of extracting data from the matrices may each be performed separately by an independent system, or both may be performed by a common system. For example, a single system or, in some cases, a single component or layer may execute all necessaryoperations for generating, storing, and extracting one or more matrices of mass spectrometry data. A method for extracting high dimensional features from mass spectrometry data may include obtaining mass spectrometry data including intensities for one or more product ions and an offset for each of the intensities and recording the intensities for the one or more product ions in a first m * n matrix. Then, transposing the first m x n matrix into a second m x n matrix with an n-dimension corresponding to a mass-to-charge dimension and an m-dimension corresponding to a time dimension and storing each of the first and second m x n matrices as an index of the mass spectrometry data. A request may then be received by the same system to retrieve a portion of the mass spectrometry data, such as data relating to a particular product ion of the one or more product ions. The request may include a search parameter and, response to the request, one of the first and second m x n matrices is searched, the matrix to be searched selected based on the search parameter, for the portion of the mass spectrometry data.
[0075] For example, if a user submits a search parameter which is a time parameter then, based on the search parameter being a time parameter, the second m x n matrix is used for extracting the portion of the data. For example, if data related to a particular precursor ion transmission window is requested, such as in a SWATH experiment, the transposed matrix with the time dimension in the m-dimension may be preferred. As another example, if the user submits a search parameter which is a mass-to-charge parameter and, based on the search parameter being a mass-to-charge parameter, the first m x n matrix is used for extracting the portion of the data. For example, if data related to a particular product ion or fragment is requested, the original matrix with the m / z dimension in the m-dimension may be preferred.
[0076] The portion of the mass spectrometry data is then retrieved or extracted and returned to the user based on the searching. Returning the extracted data may include providing a visualization of the data or may include moving the data to a cache to be further processed for analysis.
[0077] FIG. 6 illustrates an example block diagram of a virtual or physical computing system 150. One or more aspects of the computing system 150 can be used to implement the systems and methods for constructing data structures for high- dimensionality extraction. In particular, the computing system 150 may be used to implement the data processing system 200 and underlying or integrated components, such as matrix generator 206 and extractor 208.
[0078] In the embodiment shown, the computing system 150 includes one or more processors 152, a system memory 158, and a system bus 172 that couples the system memory 158 to the one or more processors 152. The system memory 158 includes RAM (Random Access Memory) 160 and ROM (Read-Only Memory) 162. A basic input / output system that contains the basic routines that help to transfer information between elements within the computing system 150, such as during startup, is stored in the ROM 162. The computing system 150 further includes a mass storage device 164. The mass storage device 164 is able to store software instructions and data. Mass storage device 164 may correspond to storage device 212 of FIG. 2, and store index 214 and matrices 216 as described in greater detail above. The one or more processors 152 can be one or more central processing units or other processors.
[0079] The mass storage device 164 is connected to the one or more processors 152 through a mass storage controller (not shown) connected to the system bus 172. The mass storage device 164 and its associated computer-readable data storage media provide nonvolatile, non-transitory storage for the computing system 150. Although the description of computer-readable data storage media contained herein refers to a mass storage device, such as a hard disk or solid state disk, it should be appreciated by those skilled in the art that computer-readable data storage media can be any available non- transitory, physical device or article of manufacture from which the central display station can read data and / or instructions.
[0080] Computer-readable data storage media include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable software instructions, data structures, program modules or other data. Example types of computer-readable data storage media include, but are not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other solid state memory technology, CD-ROMs, DVD (Digital Versatile Discs), other optical storage media, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by the computing system 150.
[0081] According to various embodiments of the invention, the computing system 150 may operate in a networked environment using logical connections to remote network devices through the network 148. The network 148 is a computer network, such as an enterprise intranet and / or the Internet. The network 148 can include a LAN, a Wide Area Network (WAN), the Internet, wireless transmission mediums, wired transmissionmediums, other networks, and combinations thereof. The computing system 150 may connect to the network 148 through a network interface unit 154 connected to the system bus 172. It should be appreciated that the network interface unit 154 may also be utilized to connect to other types of networks and remote computing systems. The computing system 150 also includes an input / output controller 156 for receiving and processing input from a number of other devices, including a touch user interface display screen, or another type of input device. Similarly, the input / output controller 156 may provide output to a touch user interface display screen or other type of output device.
[0082] As mentioned briefly above, the mass storage device 164 and the RAM 160 of the computing system 150 can store software instructions and data. The software instructions include an operating system 168 suitable for controlling the operation of the computing system 150. The mass storage device 164 and / or the RAM 160 also store software instructions, that when executed by the one or more processors 152, cause one or more of the systems, devices, or components described herein to provide functionality described herein. For example, the mass storage device 164 and / or the RAM 160 can store software instructions that, when executed by the one or more processors 152, cause the computing system 150 to receive and execute managing network access control and build system processes.
[0083] The software instructions further include one or more software applications 166. Software applications may include dedicated systems and algorithms for performing specific tasks or actions or providing specific interfaces. One or more of data processing system 200 and / or one or more component of data processing system 200 may be encompassed by software applications 166.
[0084] Referring to the figures and examples presented herein generally, the disclosed environment provides a physical environment with which aspects of the outcome prediction and trade-off analysis systems can be implemented.
[0085] Having described the preferred aspects and implementations of the present disclosure, modifications and equivalents of the disclosed concepts may readily occur to one skilled in the art. However, it is intended that such modifications and equivalents be included within the scope of the claims which are appended hereto.
Claims
What is claimed is:
1. A method of constructing mass spectrometry data structures for increased highdimensional extraction, the method comprising: obtaining mass spectrometry data including intensities for one or more product ions and an offset for each of the intensities; recording the intensities for the one or more product ions in a first m x n matrix with an n-dimension corresponding to a time dimension and an m-dimension corresponding to a mass-to-charge dimension; transposing the first m x n matrix into a second m x n matrix with an n-dimension corresponding to a mass-to-charge dimension and an m-dimension corresponding to a time dimension; and storing each of the first and second m x n matrices as an index of the mass spectrometry data.
2. The method of claim 1 , further comprising receiving a selection of one of the first and second m x n matrices.
3. The method of claim 2, further comprising, in response to receiving the selection of one of the first and second m x n matrices, using the selected matrix to extract a portion of the mass spectrometry data.
4. The method of claim 2, further comprising: receiving a request to extract a portion of the mass spectrometry data relating to a particular product ion of the one or more product ions; in response to receiving the request to retrieve the portion of the mass spectrometry data, searching one of the first and second m x n matrices, based on the received request, for the portion of the mass spectrometry data; and retrieving the portion of the mass spectrometry data based on the searching.
5. The method of claim 4, wherein the request to retrieve a portion of the mass spectrometry data includes a search criteria.
6. The method of claim 5, wherein one of the first and second m x n matrices is selected to be searched based on the search criteria defining the particular product ion.
7. The method of any preceding claim, wherein the intensities are recorded sequentially within the n-dimension of the first m x n matrix.
8. The method of claim 7, wherein the time dimension comprises the offset.
9. The method of claim 7, wherein the intensities are recorded in a m x 1 matrix that is stored in the n-dimension of the first m x n matrix.
10. The method of any preceding claim, further comprising transposing the first m x n matrix into a third m x n matrix with an n-dimension corresponding to a mass-to-charge dimension and an m-dimension corresponding to one of collision energy, charge series offset, and mobility.
11. The method of any preceding claim, further comprising transposing the first m x n matrix into a third m x n matrix with an n-dimension corresponding to a mass-to-charge dimension and an m-dimension corresponding to a quadrupole dimension, wherein the quadrupole dimension characterizes a series of overlapping precursor ion transmission windows.The method of any preceding claim, wherein the mass spectrometry data is obtained from a mass spectrometer.The method of any preceding claim, wherein the mass spectrometry data is obtained from a storage device.The method of any preceding claim, wherein the mass spectrometry data is obtained from an independent processing system.
15. A method for extracting high dimensional features from mass spectrometry data, the method comprising:obtaining mass spectrometry data including intensities for one or more product ions and an offset for each of the intensities; recording the intensities for the one or more product ions in a first m x n matrix with an n-dimension corresponding to a time dimension and an m-dimension corresponding to a mass-to-charge dimension; transposing the first m x n matrix into a second m x n matrix with an n-dimension corresponding to a mass-to-charge dimension and an m-dimension corresponding to a time dimension; storing each of the first and second m x n matrices as an index of the mass spectrometry data; receiving a request to retrieve a portion of the mass spectrometry data relating to a particular product ion of the one or more product ions, the request including a search criteria; in response to receiving the request to retrieve the portion of the mass spectrometry data, extracting the portion of data from one of the first and second m x n matrices, selected based on the search criteria; and retrieving the portion of the mass spectrometry data based on the searching.
16. The method of claim 15, wherein the search parameter is a time parameter and, based on the search parameter being a time parameter, the second m x n matrix is used for extracting the portion of the data.
17. The method of claim 15, wherein the search parameter is a mass-to-charge parameter and, based on the search parameter being a mass-to-charge parameter, the first m x n matrix is used for extracting the portion of the data.
18. A system for constructing mass spectrometry data structures for increased highdimensional extraction, the system comprising: a processor; a non-transitory memory in communication with the processor and storing instructions that, when executed, cause the processor to: obtain mass spectrum data including intensities for each of one or more product ions for each window of a series of overlapping precursor ion transmission windows;discretize the mass spectrum data such that the series of overlapping precursor ion transmission windows is divided into a series of discrete precursor ion transmission windows; record the intensities for the one or more product ions in a first m x n matrix with an n-dimension corresponding to a time dimension and an m- dimension corresponding to a mass-to-charge dimension; transpose the first m x n matrix into a second m x n matrix with an n- dimension corresponding to a mass-to-charge dimension and an m-dimension corresponding to a time dimension; and store each of the first and second m x n matrices as an index of the mass spectrometry data.
19. The system of claim 18, wherein at least one of the first and second m x n matrices is a sparse matrix.
20. The system of claim 19, wherein each of the first and second m x n matrices is a sparse matrix.