Method for solving charge state fuzzy problem in mass spectrum in high-quality and ultrahigh-quality ranges
By identifying and removing false positive compounds in the deconvolution mass spectrometry data, the fuzzy charge state problem of macromolecular compounds is solved, and the accurate correction and identification of compound abundance is achieved.
Patent Information
- Application Number
- CN202380081240.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-06
- Filing Date
- 2023-12-08
- Publication Date
- 2025-07-04
AI Technical Summary
In mass spectrometry analysis, for macromolecular compounds with molecular weight close to or exceeding 450 kDa, there is ambiguity in determining the charge state, resulting in false positives and abundance judgment errors in the identification of the compound.
By determining the charge state distribution group in the deconvolution mass spectrometry data, identifying false positive compounds and removing their signals, correcting compound abundance, and calculating molecular weight using weighting factors to ensure the accuracy of charge state distribution.
Effectively eliminate false positive recognition, correct compound abundance, and improve the recognition accuracy of macromolecular compounds and the accuracy of abundance judgment.
Smart Images

Figure CN120266109A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims priority to co - pending U.S. Non - Provisional Application No. 18 / 151,321, filed on January 6, 2023. Technical Field
[0002] The present invention relates to mass spectrometry and mass spectrometers. More specifically, the present invention relates to the mass spectrometric analysis of macromolecular compounds having a molecular weight equal to or greater than 450 kilodaltons (450 kDa). Background Art
[0003] Over the past few decades, mass spectrometry has advanced significantly and has now become one of the most widely applicable analytical tools for detecting and characterizing numerous classes of molecules. Mass spectrometry is applicable to almost all molecular species that can be ionized to form ions in the gas phase. Thus, mass spectrometry is perhaps the most generally applicable quantitative analysis method. In addition, mass spectrometry is a highly selective technique, especially suitable for analyzing complex mixtures composed of compounds at different concentrations. Mass spectrometry has extremely high detection sensitivity, with detection limits for some molecular species approaching parts per trillion. Due to these beneficial properties, a great deal of attention has been focused over the past few decades on developing mass spectrometry methods for analyzing complex mixtures of biological molecules such as peptides, proteins, carbohydrates, oligonucleotides, and complexes of these molecules.
[0004] A common application of mass spectrometry in the analysis of natural samples involves the characterization and / or quantification of the components of complex mixtures of biological molecules. Many such biological molecules are biopolymers such as polynucleotides (RNA and DNA), polypeptides, and polysaccharides. Generally, the chemical composition (related to the specific set of monomers that make up the polymer) and the monomer sequence are the distinguishing analytical features of a particular class of biopolymer molecules. However, since biopolymer molecules of a given class typically have high molecular weights and can produce ions with multiple charge states, the process of differentiating the various molecules in such a mixture by mass spectrometry can be quite challenging.
[0005] Biomacromolecules are typically dissolved in a mixture of water or aqueous buffer and an organic solvent and then introduced into the ionization source of a mass spectrometer. For example, a solid-phase extraction device can be used to separate soluble sample-derived compounds from insoluble compounds. The soluble fraction generally consists of a variety of compounds, many of which are macromolecules such as peptides and proteins. These fractions can then be further fractionated using reverse-phase chromatography. The various soluble biomacromolecules are either chromatographically separated or simply injected and can then be conveniently ionized by electrospray ionization. The resulting ion species are herein referred to as "primary" ion species. The mass analyzer of the mass spectrometer then detects and quantifies the primary ion species or specifically generated fragment ion species (or other product ion species) according to the respective mass-to-charge ratio (m / z) of the species.
[0006] An important feature of electrospray ionization is that it tends to preserve the molecular structure without excessive fragmentation because the ionization mechanism proceeds by adding solvent-derived charged units to each molecular framework. Thus, each organic molecular species in the un-ionized sample may generate a large number of ion species after ionization where each such i th ion species contains its respective charge state z i . Thus, a mass spectrum is typically a highly complex record of multiple ion species generated by each of the various compounds.
[0007] More specifically, the isotope-unresolved mass spectrum of electrospray ions of an organic molecular species usually shows a set of peaks at different mass-charge (m / z) values, where a set of peaks are all related to the same molecular weight value but are distributed over different m / z value ranges because different peaks correspond to ion species with a variety of different charge states z i . Here, the symbol refers to the mass-to-charge ratio of a member of the ion species group with a total charge value z, which is equal to i, and this charge value is also denoted as z i . In addition, the symbol refers to the maximum charge state observed in the mass spectrum of the organic molecular species Assuming that the adduct species carrying the charge is mainly single-charged (such as a proton), the members of this set of peaks can be expressed by the formula: Equation (1), m A is the mass of the adduct species. Each such set of peaks is referred to herein as “charge state distribution”. When using mass spectrometry to identify intact molecular ions of large molecular weight compounds (such as compounds with a molecular weight greater than 450 kDa or up to several hundred megaDaltons), the ability to identify mass spectral peaks corresponding to multiply charged ion species is very useful. For such large molecules, the m / z ratio of low charge state ions is usually greater than the maximum m / z value measurable by many mass spectrometer systems. Increasing the charge state z results in the observation of mass spectral peaks at m / z values within the instrument's mass analysis range.
[0008] In fact, mass analysis of a single sample may yield many overlapping charge state distributions of the above type. Accordingly, many computer software programs and algorithms capable of separating (“deconvoluting”) and identifying the various overlapping distributions are well known and / or commercially available. Generally, the output of such deconvolution programs includes: (i) a list of bar graphs of the identified peaks; (ii) grouping of the bar graph values into possible charge state distributions and assignment of a possible charge state to each peak in each group; and (iii) for each determined charge state distribution, calculation of the molecular species molecular weight
[0009] For example, one such computer program is described in U.S. Patent No. 10,217,619. FIG. 1 shows the deconvolution results of a five-component protein mixture consisting of cytochrome c, lysozyme, myoglobin, trypsin inhibitor, and carbonic anhydrase as described in U.S. Patent No. 10,217,619. The top display panel 103 of the display shows data obtained from a mass spectrometry represented as bar graphs. The central main display panel 101 shows the respective peaks with corresponding symbols. The horizontally placed mass-to-charge (m / z) scale 107 for the top panel 103 and the central panel 101 is shown below the central panel. Each horizontal line in the main panel 101 connects the bar graph symbols of the peaks that are assigned to their respective single charge state distributions. The values on the diagonal dashed line in the panel 101 are the assigned charge states. The panel 105 on the left side of the display shows the calculated molecular weight (in Daltons) of the protein molecules. The molecular weight (MW) scale of the side panel 105 is vertically oriented on the display, perpendicular to the horizontal direction m / z scale 107 of the detected ions.
[0010] The above traditional quality analysis methods are applicable to proteins with medium and low molecular weights. However, the present inventors have recognized that with the continuous improvement of the performance of mass spectrometers and the expansion of the mass analysis range to larger m / z values, a hitherto unrecognized problem may arise. Specifically, when the molecular weight approaches and exceeds approximately 450 kDa, the determination of the charge state may become ambiguous, leading to false positives in compound identification and incorrect judgments of the abundances of compounds actually present in the sample. This ambiguity is due to the uncertainty σ z (e.g., the standard deviation of the assigned charge state), and its relationship with the uncertainty of P can be shown by the following equation: In Equation (2), σ P is the uncertainty of the peak mass-to-charge ratio P. As z and σ P increase, the charge state uncertainty σ z of the high molecular weight macromolecular peak may approach or exceed 1, resulting in incorrect charge state assignments. This uncertainty is a natural consequence of the decrease in the m / z spacing between adjacent peaks in the charge state distribution as the z value increases. In this case, some signals in the charge state distribution related to the molecular species may be incorrectly assigned to the charge state distributions of different, unapparent or non-existent molecular species . Therefore, the abundance of the true (actually present) species will be underreported because some of its mass spectrometry signals will be incorrectly assigned to the incorrectly determined species In addition, false positive identifications may lead to errors in subsequent operations or decisions (such as medical diagnoses) that rely on mass spectrometry data. Summary of the Invention
[0011] This article introduces a method for solving the problem of ambiguous charge state determination that occurs during the deconvolution of ultra-high quality (equal to or exceeding 450 kDa) mass spectrometry data. This method can identify the component compounds determined by the deconvolution program, which share one or more assigned m / z peaks, but in different component compounds, the shared m / z peaks are assigned different charge states. This method can further determine which components are real and which are false positives. Then, this method eliminates the false positives, combines their signals with the signals of the actually present components to correct the abundances of the various components, and selectively generates the final spectrum of the component molecular weights.
[0012] According to one aspect of the present invention, there is provided a method for eliminating false positive identifications and correcting the abundances of sample component compounds with molecular weights greater than or equal to approximately 450 kDa determined by the deconvolution of mass spectrometry data, the method comprising: Determine a set of charge state distributions determined by deconvolution in deconvoluted mass spectrometry data, wherein determining all charge state distributions in the set includes at least one commonly assigned mass spectral peak, and in each charge state distribution of the set, each commonly assigned mass spectral peak is assigned a respective different charge state; Identify, in the determined set of charge state distributions, the charge state distributions corresponding to false positive compound identifications; distributions; Add the peak intensities of the commonly assigned mass spectral peaks of the determined set of charge state distributions determined to correspond to false positive compound identifications to the peak intensities of the target charge state distributions of the set not determined to correspond to false positive compound identifications; Exclude all charge state distributions corresponding to false positive compound identifications from the determined set of charge state distributions; and Calculate the component compound abundances corresponding to the target charge state distributions using the sum of the peak intensities.
[0013] In some cases, the step of identifying, in the determined set of charge state distributions, the charge state distributions corresponding to false positive compound identifications may include: Assigning respective weighting factors to each charge state distribution in the determined set of charge state distributions; Calculating, in the determined set of charge state distributions, the fractional weighted average molecular weight of each true or hypothetical component molecular species corresponding to the respective charge state distributions in the determined set of charge state distributions, the calculation using the assigned weighting factors; and Finding, in the determined set of charge state distributions, the target charge state distribution whose charge state distribution is closest to the calculated average molecular weight value. In some cases, the assignment of the weighting factors may be at least partially based on the assigned or calculated molecular weights or calculated mean squared errors of the actual or hypothetical component molecular species corresponding to the charge state distributions. In some cases, the assignment of the weighting factors may be at least partially based on the mass spectral peak intensities of the respective charge state distributions.
[0014] According to another aspect of the present invention, there is provided a mass spectrometer system, comprising: An electrospray ion source configured to receive a sample composed of one or more component compounds, the molecular weights of these compounds being greater than or equal to 450 kilodaltons (kDa); A mass analyzer configured to receive ions generated by ionization of the sample component compounds; A detector configured to detect the ions output from the mass analyzer and generate mass spectrometry data; A data storage device configured to receive mass spectrometry data from a detector; and A programmable processor device configured to receive mass spectrometry data from the detector or the data storage device and comprising operative computer-readable instructions to: Perform conventional deconvolution on the mass spectrometry data; and Automatically detect and eliminate false positive compound identifications resulting from the conventional deconvolution, wherein the false positive compound identifications are caused by a standard deviation of charge state assignments equal to or greater than 1. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above description of the present invention and various other aspects will become apparent from the following description given by way of example only and with reference to the drawings which are not drawn to scale, in which:
[0016] FIG. 1 is a graphical illustration of the results of mass spectrometry deconvolution of a five-component protein mixture consisting of cytochrome c, lysozyme, myoglobin, trypsin inhibitor, and carbonic anhydrase;
[0017] Figure 2A is a chromatogram showing the elution curves of various multimers of the globular protein apoferritin;
[0018] Figure 2B is the corresponding Figure 2A low-resolution mass spectrum of the eluate corresponding to the elution curve;
[0019] Figure 3A is a calculated molecular weight spectrum of the elution fraction, the molecular weight of which is obtained by applying a deconvolution algorithm to the Figure 2B data;
[0020] Figure 3B is a graphical description of the charge state uncertainty that may cause Figure 2B a single peak in the mass spectrometry data to be assigned to multiple charge state distributions, the graphical description including a main display panel located centrally with corresponding symbols representing the assignment of each peak; a top display panel showing the original mass spectrometry data; a left panel showing the calculated molecular weight along the vertical molecular weight axis; and a mass-to-charge (m / z) scale horizontally placed between the top panel and the central panel;
[0021] Figure 3C is a graphical description similar to Figure 3B showing the result of correcting the mass spectrometry data to a single molecular weight value using the method of the present invention;
[0022] Figure 3D is the corrected Figure 3A calculated molecular weight spectrum of the elution fraction according to the present invention;
[0023] Figure 4 is a flowchart of a method for resolving mass spectrometry charge state ambiguity according to the present invention; and
[0024] Figure 5 is a schematic diagram of a chromatography / mass spectrometry spectral generation and automated analysis system that can be used in conjunction with the method of the present invention. DETAILED DESCRIPTION
[0025] The following description is presented to enable a person skilled in the art to make and use the invention, and the following description is provided in the context of a particular application and its requirements. Various modifications to the described embodiments will be readily apparent to those skilled in the art, and the general principles herein can be applied to other embodiments. Thus, the invention is not limited to the embodiments and examples shown, but is to be accorded the widest possible scope consistent with the features and principles shown and described. For a more complete understanding of the features of the present invention in detail, please refer to FIGS. 1, 2A, 2B, 3A, 3B, 3C, 3D, 4 and 5 and the following description.
[0026] In the description of the present invention herein, it should be understood that unless otherwise implicitly or explicitly understood or stated, words presented in the singular form cover their plural counterparts, and words presented in the plural form cover their singular counterparts. In addition, it should be understood that unless otherwise implicitly or explicitly understood or stated, for any given component or embodiment described herein, any possible candidates or alternatives listed for the component can generally be used individually or in combination with each other. Further, it should be understood that the figures shown herein are not necessarily drawn to scale, where only some elements may be drawn for clarity of the present invention. Additionally, reference numerals may be repeated in the various figures to show corresponding or similar elements.
[0027] Figures 2A - 2B and FIGS. 3A - 3D illustrate the occurrence of the hitherto unrecognized problems discussed above. Figure 2A is a chromatogram showing the elution curve 61 of various globular protein apoferritin polymers, including a broad elution peak that reaches a maximum at retention time 62. Figure 2B is related to Figure 2A is a low - resolution mass spectrum of the eluate corresponding to the elution curve of, with an m / z range of 4000 - 15000 Th. Figure 3A is a molecular weight spectrum of the chemical components eluted from the eluate, which is obtained by Figure 2BThe data was determined using traditional deconvolution algorithms. The molecular weights of the chemical components calculated by the deconvolution routine were 471679.1 Da (molecular component peak 71), 493192.1 Da (molecular component peak 72), 507043.7 Da (molecular component peak 73), and 517193.1 Da (molecular component peak 74), respectively. The calculated molecular weight of molecular component peak 73 is close to the molecular weight of the known apoferritin 24-polymer. However, as can be clearly seen from the charge state assignments listed in Table 1 below, some of the mass spectrometry signal intensities of the apoferritin 24-polymer peaks (i.e., Figure 2B the data shown in Table 1. Charge State Assignments
[0028] Figure 3B is a graphical description similar to Figure 1, showing how charge state uncertainty leads to Figure 2B a single peak in the mass spectrometry data being assigned to more than one charge state distribution. Figure 3B The top display panel 103 in Figure 2B is an expanded version of the mass spectrometry data in i with reference to the horizontal mass-to-charge (m / z) scale 107 at the bottom of the figure. The left panel 105 shows a rotational view of the molecular weights calculated by the standard deconvolution software package, with the mass scale (i.e., the molecular weight scale) in the vertical direction. Finally, the central panel 101 relates the mass spectrometry peaks in Table 1 to the assigned charge states, as well as the horizontally placed m / z scale and the vertically placed molecular weight scale. The bar graphs of the peaks in Table 1 are shown as single points in the central panel 101. Similar to the graphical description of Figure 1, the diagonal dashed lines in the central panel 101 are constant charge state lines, and each diagonal is labeled according to the charge state z i it represents. The slope of each diagonal dashed line is 1 / z
[0029] Located Figure 3B The asterisks above the mass spectrometry in the top panel 103 indicate the positions of the seven different peaks used for assignment, as shown in Table 1. (For reference, Figure 3BThe peaks marked with diamonds in the top panel 103 represent the charge state distribution of another apoferritin multimer with a molecular weight of 493 kDa, in Figures 3A - 3D which is represented by the molecular component peak 72 in Figure 3B Although only seven peaks are listed in Table 1, these seven peaks are plotted at 11 different points in the molecular weight versus charge ratio plot of the central panel 101 because, as described above, four of these peaks (peaks 2 - 5 in Table 1) are assigned to two hypothetical peaks 73, 74 corresponding to different molecular weights through deconvolution. The horizontal line 173 connects peak 73 to the corresponding charge state assignment of the mass spectrometry line in the table, and the horizontal line 174 connects peak 74 to its corresponding charge state assignment, and both sets of assignments are obtained through deconvolution calculations.
[0030] As described above (e.g., Equation 2), when studying compound ions with molecular weights close to or exceeding 500 kDa using mass spectrometry analysis, the uncertainty in charge assignment may approach or exceed one charge unit. The charge assignment error line can be defined in various ways, for example, defined by σ p in Equation 2 or a multiple thereof. The exact molecular weight threshold at which the charge uncertainty exceeds one charge unit depends on the way the error line is defined and various instrument factors such as the reproducibility and / or precision of the mass spectrometry m / z measurement, the m / z detection range, the detected charge amount (e.g., ≥50, ≥100, ≥150), etc. Figure 3B The hypothetical charge state error line 111 described in
[0031] Figure 3C shows the statistically credible range of variation of the isocharge state lines such as the diagonal dashed line when the charge state uncertainty (however defined) is ±1 charge unit. Thus, without using other information, one or both sets in the charge state assignment (e.g., either aligned with the 174 line or aligned with the 173 line) must be considered statistically credible. Figure 3B is a graphical description similar to
[0032] Figure 4FIG. 200 is a flow chart of a method 200 for resolving such mass spectrometry charge state ambiguity according to the present invention. Method 200 is applicable to analyzing macromolecules with a molecular weight greater than or equal to 450 kDa, which are ionized by an ionization technique, and during the ionization process, the ions of the macromolecules are generated by the adduct of multiple charged particles with the molecule to be analyzed. Generally speaking, this method is applicable to analyzing organic macromolecules ionized by electrospray or thermospray ionization, but this is not necessarily the case.
[0033] When preparing to execute method 200, before executing method 100, the mass spectrum of a sample containing a macromolecular component compound is measured with a mass spectrometer and / or retrieved from a data memory. In addition, after measuring or retrieving the mass spectrum, a traditional "deconvolution" procedure is adopted to logically organize the various mass-to-charge ratios P of the observed mass spectrum peaks into ...... such peak groups, which are initially considered to contain charge state distributions corresponding to the initially determined component molecular species and so on. As described above, as part of the deconvolution procedure, a charge state is assigned to each member in each charge state distribution, and a molecular weight is assigned to each component molecular species. In the practice of high-resolution mass spectrometry and ultra-high-resolution mass spectrometry, the molecular weights of various component molecular species can be between 0.45 MDa (megadaltons) and 3 MDa, while the m / z values of the detected ion species can be between 4000 - 20000 Th. Therefore, the charge states of the detected ions can be between 50 and 150. Therefore, due to the increased uncertainty in the above charge assignment, the charge state assignment by the deconvolution procedure must be regarded as tentative. Therefore, some of the identified component molecular species may be false positive identifications, and the abundances of some of the actually existing species that have been determined may be underestimated.
[0034] In the initial step 203 of method 200, a search is performed on the result list generated by the deconvolution procedure to identify all cases where the deconvolution routine assigns one or more observed mass spectrum peaks to multiple charge state distributions (i.e., "groups" of charge state distributions), and assigns different charge states to each charge state distribution in this group. Based on the result of this search, one or more groups of charge state distributions are determined in the deconvolved mass spectrometry data, and the criterion for determining the group is that all charge state distributions in the group include at least one commonly assigned mass spectrum peak (i.e., a peak shared by all charge state distributions in the group), and each commonly assigned mass spectrum peak is assigned a different charge state in each charge state distribution.
[0035] In step 205, a weighting factor is assigned to each charge state distribution in each set of determined charge state distributions (step 203), each weighting factor being related to the mass spectrometry data quality of the mass spectrometry peaks corresponding to the species determined according to certain selected metrics. The weighting factor can be based on mass spectrometry peak intensity, mean square error of mass, or other mass metrics suitable for the analytical instrument.
[0036] In step 207, the fractional weighted average molecular weight of each set of charge state distributions determined in step 203 is calculated using the weighting factors. However, before performing step 207, step 206 includes first calculating the molecular weight of the component compound corresponding to the charge state distribution using all the P values and z values of the mass spectrometry peaks within each charge state distribution in each determined set, whether the component compound is an actual compound or a hypothetical compound, and whether the component compound actually exists in the sample. Then, in step 207, the average of these individual molecular weights is calculated using the weighting factors assigned in step 205.
[0037] Although each fractional weighted average molecular weight calculated in step 207 generally does not correspond to any real component species, it will generally be close to the real molecular weight of a particular species that does exist in the sample. Thus, in step 209, the component molecular species whose initially calculated molecular weight (step 206) is closest to the intensity weighted average molecular weight (step 207) is found from each set of charge state distributions. In each set, this so - located single molecular component species is the most likely true positive species in that set, referred to herein as the "target" component species.
[0038] In subsequent step 211, all component species other than the determined target components are removed from each determined set. The removed molecular components are considered false positives. Thus, for each co - assigned peak of each set of charge state distributions, the signal intensity initially assigned to the removed species is added to the signal intensity of the target component species found in the previous step. Finally, in step 213, the abundances of the various determined target components are recalculated based on the respective total peak signal intensities.
[0039] Figure 3B Shows the results when this method is applied to Figure 3A the deconvolution results shown. Specifically, the false positive at 517 kDa ( Figure 3A peak 74 in Figure 3B ) is identified and removed, and its signal is combined with the signal of the true component at 507 kDa to obtain the corrected mass and intensity, shown as molecular component peak 75 in
[0040] Figure 5FIG. 0 is a schematic diagram of a system 10 for generating and automatically analyzing chromatography / mass spectrometry spectra according to the present invention. According to well-known chromatography principles, a chromatograph 33, such as a liquid chromatograph, a high-performance liquid chromatograph, or an ultra-high-performance liquid chromatograph, receives a sample 32 of an analyte mixture and separates the analyte mixture at least partially into individual chemical components. The at least partially separated chemical components are transferred to a mass spectrometer 34 at different respective times for mass analysis. When the mass spectrometer receives each chemical component, the chemical component is ionized by an ionization source 34a of the mass spectrometer. The ionization source can generate a plurality of ions (i.e., a plurality of precursor ions), including different charges or masses for each chemical component. Thus, a plurality of ion species with different mass-to-charge ratios (e.g., charge state distributions) can be generated for each chemical component, and each such component elutes from the chromatograph at its own characteristic time. These different ion species are analyzed by a mass analyzer 34b of the mass spectrometer system 34 and detected by an ion detector 35. As is well known, through the combined action of mass analysis and ion detection, the mass spectrometer can generate mass spectrometry data that records the detected ion intensity as a function of the mass-to-charge ratio of the ions generated in the sample. Using the mass spectrometry data, various ion species can be appropriately determined according to different mass-to-charge ratios. The mass spectrometer can include a mass filtering device (not shown) for isolating ion species within certain selected m / z ranges and a fragmentation unit (not shown) for generating product ions by fragmentation of the selected ion species.
[0041] Still referring to Figure 5 , the programmable processor 37 of the system 10 is electrically coupled to the detector of the mass spectrometer and receives the data generated by the detector during the chromatographic / mass spectrometric analysis of the sample. The programmable processor can include a separate stand-alone computer or can include only a circuit board or any other programmable logic device operated by firmware or software. Alternatively, the programmable processor can also be electrically coupled to the chromatograph and / or the mass spectrometer to transmit electronic control signals to one or the other of these instruments to control their operation. The nature of such control signals may be determined according to the data transmitted from the detector to the programmable processor or by the analysis of such data. The programmable processor can also be electrically coupled to a display or other output 38 to directly output the data or the results of the data analysis to the user or to an electronic data storage 36.
[0042] Figure 5 The programmable processor 37 of the illustrated system 10 generally includes computer-readable instructions that can be used to: control the individual operations and sequences of operations of the chromatograph 33; control the operations and sequences of operations of the mass spectrometer 34; receive a mass spectrum from the detector 35; perform a conventional deconvolution program on the mass spectrometry data; perform the logical steps of method 100 ( Figure 4 ) to eliminate false positive identifications and correct the compound abundances provided by the conventional deconvolution program.
[0043] The discussions included in this application are intended to serve as a basic description. The scope of the present invention is not limited to the specific embodiments described herein, which are intended to be illustrative of individual aspects of the present invention. Functionally equivalent methods and components are within the scope of the present invention. Various other modifications of the present invention will become apparent to those skilled in the art in addition to what is shown and described herein, based on the foregoing description and drawings.
Claims
1. A method for eliminating false positive identifications and correcting the abundances of component compounds in a sample having a molecular weight greater than or equal to 450 kDa as determined from deconvoluted mass spectrometry data, the method comprising: (a) determining in the deconvoluted mass spectrometry data a set of charge state distributions identified by deconvolution, wherein all charge state distributions in the determined set include at least one commonly assigned mass spectral peak, and in each charge state distribution of the set, each commonly assigned mass spectral peak is assigned a respective different charge state; (b) identifying in the determined set of charge state distributions the charge state distributions corresponding to false positive compound identifications; (c) adding the peak intensities of the commonly assigned mass spectral peaks of the determined set of charge state distributions determined to correspond to false positive compound identifications to the peak intensities of the target charge state distributions of the set not determined to correspond to false positive compound identifications; (d) removing from the determined set of charge state distributions all charge state distributions corresponding to false positive compound identifications; (e) calculating the abundances of the component compounds corresponding to the target charge state distributions using the sum of the peak intensities.
2. The method according to claim 1, wherein, The step (b) of identifying in the determined set of charge state distributions the charge state distributions corresponding to false positive compound identifications includes: (b1) assigning to each charge state distribution in the determined set of charge state distributions a respective weighting factor; (b2) calculating, in the determined set of charge state distributions, the fractional weighted average molecular weight of each true or hypothesized component molecular species corresponding to the respective charge state distributions in the determined set of charge state distributions, the calculation using the assigned weighting factors; (b3) finding in the determined set of charge state distributions the target charge state distribution whose charge state distribution is closest to the calculated average molecular weight value.
3. The method according to claim 1, wherein the assignment of each weighting factor is at least partially based on the intensity of the mass spectral peaks of the respective charge state distributions.
4. The method according to claim 1, wherein the assignment of each weighting factor is at least partially based on the assigned or calculated mean squared error of the molecular weights of the true or hypothesized component molecular species corresponding to the respective charge state distributions.
5. A mass spectrometer system, comprising: an electrospray ionization source configured to receive a sample composed of one or more component compounds having a molecular weight greater than or equal to 450 kilodaltons (kDa); a mass analyzer configured to receive ions generated by ionization of the sample component compounds; a detector configured to detect ions output from the mass analyzer and generate mass spectrometry data; a data storage device configured to receive the mass spectrometry data from the detector; and a programmable processor device configured to receive the mass spectrometry data from the detector or the data storage device and comprising operable computer-readable instructions to: perform conventional deconvolution on the mass spectrometry data; and automatically detect and eliminate false positive compound identifications resulting from the conventional deconvolution, wherein the false positive compound identifications are caused by a standard deviation of charge state assignments equal to or greater than 1.
6. The mass spectrometer system according to claim 5, wherein the computer-readable instructions are operative to automatically detect and eliminate false positive compound determinations resulting from conventional deconvolution: determine a set of charge state distributions resulting from deconvolution in the deconvolved mass spectrometry data, wherein all charge state distributions in the determined set include at least one commonly assigned mass spectrometry peak, and in each charge state distribution of the set, each commonly assigned mass spectrometry peak is assigned a respective different charge state; within the determined set of charge state distributions, assign respective weighting factors to each charge state distribution of the set; within the determined set of charge state distributions, calculate a fraction weighted average molecular weight for each true or hypothesized component molecular species corresponding to the respective determined charge state distributions in the set using the weighting factors; in the determined set of charge state distributions, identify a single target charge state distribution corresponding to the molecular weight closest to the calculated average molecular weight value; and exclude all charge state distributions from the determined set of charge state distributions other than the single target charge state distribution.
7. The mass spectrometer system according to claim 5, wherein the charge state of one or more detected component compound ions is greater than or equal to 50.
8. The mass spectrometer system according to claim 5, wherein the charge state of one or more detected component compound ions is greater than or equal to 100.
Citation Information
Patent Citations
Methods for data-dependent mass spectrometry of mixed intact protein analytes
US10217619B2
Cited By
Multi-sample joint deconvolution method applied to GC-MS (Gas Chromatography-Mass Spectrometer) technology
CN122045583A