An apparatus and method for identifying a sample therein
Patent Information
- Application Number
- PCT/US2025/029060
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-16
- Filing Date
- 2025-05-13
- Publication Date
- 2026-08-27
Smart Images

Figure US2025029060_27082026_PF_FP_ABST
Abstract
Description
Reference No. P230046-711 -WO017240135PCTTITLE AN APPARATUS AND METHOD FOR IDENTIFYING A SAMPLE THEREINPRIORITY
[0001] The present application claims priority under 35 U.S.C. Section 119(e) from Provisional Application 63 / 648,453, entitled "AN APPARATUS AND METHOD FOR IDENTIFYING A SAMPLE THEREIN,” filed on May 16, 2024, the entire contents of which are incorporated herein by reference.GOVERNMENT SUPPORT
[0002] This invention was made with government support under Government Contract No. 20CWMDARI 00039-01-00 and CWMD 1916-004 awarded by the U.S. Department of Homeland Security. The government has certain rights in the invention.FIELD
[0003] This disclosure relates generally to the field of machine learning and more particularly relates to classifying and / or identifying unknown substances using machine learning and mass spectrometry.BACKGROUND
[0004] Two-dimensional Tandem Mass Spectrometry (2D MS / MS, or 2DMS) enables the collection of the fragmentation patterns and intact masses of all components in a sample in a matter of seconds using a single scan of a linear ion trap mass spectrometer and a single ion injection.SUMMARY
[0005] In part, in one aspect, the disclosure relates to an apparatus comprising a database. The database comprises a plurality of sets of data, wherein each set of data comprises an identity of known substances and tandem mass spectrometry spectrum data for each of the known substances. The apparatus comprises a control circuit comprising a processor and a memory. The memory stores instructions executable by the processor to: obtain a first portion of data from the database; input the first portion of the data into a machine learning model; and train the machine learning model with the first portion of the data to form a trained machine learning model. The machine learning model is trained to identify the identity of substances based on the tandem mass spectrometry spectrum data of 1323430828.1Reference No. P230046-711 -WO017240135PCTthe substance. The memory further stores instructions executable by the processor to input a second portion of data from the database into the machine learning model and evaluate the machine learning model using the second portion of the data. The second portion and the first portion comprise different sets of data from the database.
[0006] In yet another aspect, the disclosure relates to an apparatus comprising: a processor; and a memory storing a plurality of instructions and a pre-trained machine learning model executable by the processor. The pre-trained machine learning model is trained using a database comprising a plurality of known substances and a plurality of tandem mass spectrometry spectrum data for each of the known substances. The pretrained machine learning model is trained to classify the data with one of the substances in the database. The instructions are configured to cause the processor to: receive tandem mass spectrometry data for an unknown sample; and using the pre-trained machine learning model, determine an identity of the sample. The identity is one of a substance stored in a database to train the pre-trained machine learning model or a substance not used to train the machine learning model.
[0007] In yet another aspect, the disclosure relates to a method comprising obtaining data from a database. The database comprises a plurality of sets of data and each set of data comprises an identity of a known substance and tandem mass spectrometry spectrum data for the known substance. The method comprises inputting a first portion of the data into a machine learning model and training the machine learning model with the first portion of the data. The machine learning model is trained to identify the identity of substances based on the tandem mass spectrometry spectrum data of the substance. The method further comprises inputting a second portion of data from the database into the machine learning model. The second portion and the first portion comprise different sets of data from the database. The method further comprises: evaluating the machine learning model using the second portion of the data; determining the machine learning model is accurate based on the evaluation; based on the machine learning model being accurate, receiving tandem mass spectrometry data from an unknown sample; determining, using the machine learning model stored in a memory of a computing device, the identity of an unknown sample; and providing an output indicating the identity of the sample.
[0008] Although the present disclosure relates to different aspects and embodiments, it is understood that the different aspects and embodiments disclosed herein can be integrated, combined, or used together as a combination system, or in part, as separate components, devices, and systems, as appropriate. Thus, each embodiment disclosed2323430828.1Reference No. P230046-711 -WO017240135PCTherein can be incorporated in each of the aspects to varying degrees as appropriate for a given implementation.
[0009] These and other features of the applicant’s teachings are set forth herein.BRIEF DESCRIPTION OF THE FIGURES
[0010] Unless specified otherwise, the accompanying drawings illustrate aspects of the innovations described herein. Referring to the drawings, wherein like numerals refer to like parts throughout the several views and this specification, several embodiments of presently disclosed principles are illustrated by way of example, and not by way of limitation. The drawings are not intended to be to scale. A more complete understanding of the disclosure may be realized by reference to the accompanying drawings in which:
[0011] FIG. 1 is a flow diagram illustrating a non-limiting embodiment of a method for analyzing a sample in a mass spectrometer according to the present disclosure;
[0012] FIG. 2 is a schematic diagram of a non-limiting embodiment of a mass spectrometer according to the present disclosure;
[0013] FIG. 3 is a perspective view of a non-limiting embodiment of an ion trap according to the present disclosure;
[0014] FIG. 4 is a cross-sectional view of the ion trap of FIG. 3, taken along line 3-3;
[0015] FIGs. 5A-5C are graphs illustrating aspects of a non-limiting embodiment of a scan function according to the present disclosure, wherein the time scale in each graph is the same;
[0016] FIGs. 6A-6C are graphs illustrating aspects of a non-limiting embodiments of a scan function according to the present disclosure, wherein the time scale in each graph is the same;
[0017] FIG. 7 is a perspective view of non-limiting embodiments of an ion trap and two detectors according to the present disclosure;
[0018] FIG. 8 is a cross-sectional view of FIG. 7, taken along line 7-7;
[0019] FIGs. 9A-9D are graphs illustrating a non-limiting embodiments of a scan function according to the present disclosure, wherein the time scale in each graph is the same;3323430828.1Reference No. P230046-711 -WO017240135PCT
[0020] FIG. 10 is a schematic diagram of a non-limiting embodiment of a mass spectrometer according to the present disclosure;
[0021] FIG. 11 is a method of training a machine learning model to identify a sample, according to the present disclosure;
[0022] FIG. 12 illustrates examples of 2DMS spectra collected after calibration and processing according to the present disclosure;
[0023] FIG. 13 illustrates the validation of a machine learning model, according to the present disclosure;
[0024] FIG. 14 is a method of determining the identity of a substance, according to the present disclosure;
[0025] FIG. 15 illustrates a convolutional neural network to identify fentanyl and fentanyl-variant containing samples, according to the present disclosure; and
[0026] FIG. 16 illustrates an example of identifying an unknown sample, according to the present disclosure.DETAILED DESCRIPTION
[0027] The emergence of novel psychoactive substances, for example, molecular variants of fentanyl and nitazenes, are difficult to identify by conventional analytical means (colorimetric tests, Raman, FTIR, mass spectrometry). Amongst the methods most suitable to address the opioid epidemic is mass spectrometry. Conventional mass spectrometry methodologies rely on comprehensive mass spectral libraries such as the NIST El and MS / MS libraries and the SWGDRUG El library for compound identification. This strategy accurately identifies known threats that already exist in the library but fails to identify or even classify molecular variants of known threats, which have different masses than the base molecular threat (e.g., fentanyl) but often similar fragmentation patterns.
[0028] This disclosure relates in part to a method of library-less identification of existing threats and their near-neighbor variants by AI / ML analysis of 2DMS data, to identify previously unseen threats (including narcotics, chemical warfare agents, explosives, toxic industrial chemicals, etc.).
[0029] This disclosure also relates to the application of mass spectrometry to identify unknown substances. Molecular components of virtually any kind of sample (solid / liquid / gas / aerosol) such as pathogens (bacteria, virus, or microorganism) or drug4323430828.1Reference No. P230046-711 -WO017240135PCTanalogs can be identified through a combination of soft ambient ionization, producing pseudomolecular ions that, through a single stage of mass analysis, inform on the molecular weight of the sample’s constituents. Several (i.e., two or more) stages of mass analysis and gas-phase fragmentation, termed tandem mass spectrometry (or “MSn”, where n is an integer), increase both molecular specificity and analytical signal-to-noise ratios that can be beneficial for differentiating similar threats or preventing false positive identifications.
[0030] Identification of samples may include, for example, classifying a class of molecules or pathogens such as a sample X as containing a fentanyl variant or identifying the exact molecule such as, for example, classifying the sample as fentanyl-like or identifying the sample as fluorofentanyl.
[0031] Identification of pathogens may use MS methods such as, for example, matrix-assisted laser desorption / ionization time-of-flight MS, liquid chromatography MS, ion mobility, and pyrolysis MS.
[0032] Two-dimensional tandem mass spectrometry (“2D MS / MS”, or “2DMS”) enables the collection of the fragmentation patterns and intact masses of all components in a sample in a matter of seconds using a single scan of a linear ion trap mass spectrometer and a single ion injection. 2DMS enables fast chemical and biological threat detection.
[0033] This disclosure encompasses methods and systems enabling greater accuracy in chem / bio threat detection as well as lower limits of detection and fewer false positives. The combination of 2D MS / MS and artificial intelligence (machine learning, in particular) -collectively referred to as AI / ML - may also provide a method for detecting variants of known threats, providing a potential path toward library-free identification of unknown malicious materials and other unknown materials.
[0034] With reference now to the accompanying figures, FIG. 1 schematically illustrates a non-limiting method 100 for analyzing a sample using a mass spectrometer. Method 100 can be performed using, for example, a mass spectrometer 200 as described with respect to FIG. 2 below. Referring again to FIG. 1, the method 100 can ionize the sample and measure a mass-to-charge ratio (m / z) of the ions generated. The sample can comprise a homogenous or heterogeneous mixture of chemical compounds. The sample can be introduced to the mass spectrometer for example, by injection, actively sample the surrounding environment, and / or otherwise receive a sample provided to it or that it encounters.5323430828.1Reference No. P230046-711 -WO017240135PCT
[0035] As illustrated in FIG. 1, method 100 can optionally (boxed in broken lines) comprise ionizing at least a portion of the sample at step 102. In various non-limiting embodiments, ionization can comprise electrospray ionization and / or other ionization technique. In certain non-limiting embodiments of method 100, precursor ions can be produced by ionizing the sample and can be unfragmented. Ionization electrically charges a molecule and thereby generates an ion from the molecule through gain or loss of one or more electrons and / or charged particle (e.g. proton, sodium ion, chloride ion) from the molecule. For example, the precursor ions can be the individual molecules in the sample modified by addition of an electrical charge.
[0036] The method 100 can optionally comprise fragmenting at least a portion of the sample and / or ions, at step 103. For example, fragmenting precursor ions forms first-generation product ions, fragmenting first-generation product ions forms second-generation product ions, and fragmenting second-generation product ions forms third-generation product ions. The product ions can be fragmented a number of times based on the desired application and may be fragmented to third-generation product ions or further. Fragmenting is a chemical disassociation caused by, for example, the removal of at least one electron from an ion, collision with a gas molecule and / or solid surface, electron capture or transfer, and / or ultraviolet and / or infrared photon absorption. The m / z values of the ions generated may be the same or different. In various non-limiting embodiments, fragmenting at least the portion of the sample can comprise at least one of in-source collision induced dissociation, beam-type collision induced dissociation, collision induced dissociation by resonance excitation, surface-induced dissociation, infrared multiphoton dissociation, ultraviolet photodissociation, electron capture dissociation, electron transfer dissociation, and electron impact dissociation.
[0037] In various non-limiting embodiments, fragmenting at least a portion of the sample and / or ions at step 103 can occur prior to introducing the ions and / or sample to the ion trap. In various other non-limiting embodiments, fragmenting at least a portion of the sample and / or ions at step 103 can occur within the ion trap. In various non-limiting embodiments, fragmenting at least a portion of the sample and / or ions at step 103 can occur before the ion trap in a collision device. For example, fragmenting at least the portion of the sample and / or ions at step 103 can comprise beam type collision induced dissociation within the collision cell prior to the ion trap. In various non-limiting embodiments, fragmenting at least the portion of the sample and / or ions at step 103 can comprise at least one of beam type collision induced dissociation within the ion trap and collision induced dissociation by resonance excitation within the ion trap.6323430828.1Reference No. P230046-711 -WO017240135PCT
[0038] In certain non-limiting embodiments, fragmenting at least the portion of the sample and / or ions at step 103 is performed over a range of different collision energies (e.g., by varying the ion’s acceleration prior to the collision device) within the collision cell. In various non-limiting embodiments, fragmenting at least a portion of the sample and / or ions at step 103 can occur both in a collision device and within the ion trap. The ions can be fragmented based on their m / z value and applied electric fields such that an m / z value of the ion that was fragmented can be determined based the applied electric field and thus based on a time of detection.
[0039] As illustrated, the method 100 comprises storing the ions in an ion trap of the mass spectrometer, at step 104. For example, at least one of precursor ions produced from the sample, first-generation product ions produced from precursor ions, second-generation product ions produced from the first-generation product ions, third-generation product ions produced from the second-generation product ions, and further generation product ions can be stored in the ion trap.
[0040] As illustrated in FIG. 1, the method 100 comprises applying a scan function to the ion trap, at step 106. Applying the scan function can comprise exciting at least a portion of the ions in the ion trap selectively over time to fragment ions into product ions and ejecting ions from the ion trap. In various non-limiting embodiments, at least one of a precursor ion can be fragmented into first-generation product ions, a first-generation product ion can be fragmented into second-generation product ions, a second-generation product ion can be fragmented into third-generation product ions, or further generation product ion can be fragmented by the scan function. In various non-limiting embodiments, fragmenting is performed on ions with relatively low m / z values first, and fragmenting is next performed on ions with successively greater m / z values. In certain non-limiting embodiments, at least one of a precursor ion, a first-generation product ion, a second-generation product ion, a third-generation product ion, or a further generation product ion can be ejected from the ion trap by the scan function.
[0041] The scan function can produce one further generation of ions from ions and / or sample stored in the ion trap. For example, a single scan function may fragment a first-generation product ion into a first set of second-generation product ions, but the single scan function may not further fragment the second-generation product ions. In various nonlimiting embodiments, the second-generation product ions can be stored in the ion trap and fragmented into third-generation product ions. Certain scan functions can fragment ions7323430828.1Reference No. P230046-711 -WO017240135PCTwithin the ion trap and may not eject ions from the ion trap. Thus, the scan functions can be applied until a desired generation of product ions is achieved.
[0042] The scan function comprises exciting at least a portion of the ions in the ion trap selectively over time to fragment precursor ions into product ions and ejecting at least one of the product ions and the precursor ions from the ion trap. For example, the scan function may comprise a radio frequency (RF) voltage (e.g., a trapping voltage), an excitation frequency, and an ejection frequency. The RF voltage, the excitation frequency, and the ejection frequency can be applied to electrodes in the ion trap to form a dynamic electric field within the ion trap configured to control (e.g., store, eject) the ions as desired. For example, the scan function can store ions based on an m / z value of the ions and / or eject ions based on an m / z of the ions such that the ions can be sorted based on the m / z value of the respective ion.
[0043] The RF voltage can store ions in the ion trap. For example, the RF voltage can store ions in the ion trap based on the secular frequency of ions in the ion trap such that the ions oscillate within a potential well in the ion trap. The RF voltage can be selected based on the desired m / z values to be stored in and / or ejected from the ion trap. In various nonlimiting embodiments, the RF voltage can be in a range of 50 to 5,000 volts.
[0044] The excitation frequency can excite ions in the ion trap to fragment the ions. For example, the excitation frequency can bring the secular frequency of ions into resonance with the excitation frequency based on the m / z value of the ion. Changing the excitation frequency and / or the RF voltage can change the ions that are in resonance with the excitation frequency. Resonance can occur when the secular frequency of the respective ion matches (e.g., or otherwise comes close to matching) the excitation frequency, which can cause the ions to become kinetically excited and collide with other molecules in the ion trap. The other molecules in the ion trap can be molecules of a gas, such as, for example, helium, nitrogen, air, and / or other gas. The collision can cause the fragmentation of the ions. For example, a precursor ion can be fragmented into first-generation product ions, a first-generation product ion can be fragmented into second-generation product ions, a second-generation product ion can be fragmented into third-generation product ions, and further generation product ions can be fragmented. In various non-limiting embodiments, the excitation frequency can be in a range of 20 kHz to 2 MHz, such as, a range of 1MHz to 2MHz or 20kHz to 1MHz. In certain non-limiting embodiments, the ejection frequency can be in a range of 20 kHz to 1 MHz.8323430828.1Reference No. P230046-711 -WO017240135PCT
[0045] In various non-limiting embodiments, a scan function comprises a constant RF voltage (e.g., FIG. 5A), a sweeping excitation frequency (e.g., FIG. 5B), and successive ejection frequency scans of variable duration and variable frequency range based on the sweeping excitation frequency (e.g., 50). As illustrated in FIG. 50, the constant RF voltage can be a high voltage signal (e.g., relative to the excitation frequency and ejection frequency, a range of 100 volts to greater than 1000 volts) configured to store the ions in the ion trap. As used herein, a “constant” means the frequency or voltage does not deviate from a predetermined value by more than 0.2%, such as, for example, no more than 0.1%. As illustrated in FIG. 5A, the sweeping excitation frequency can be a low AC voltage (e.g., relative to the RF voltage, 100 mV to less than 100 voltages) auxiliary waveform configured to excite the ions to fragment the ions in the ion trap. The sweeping excitation frequency may excite lower m / z value ions by starting at a first frequency and may proceed to excite higher m / z value ions by lowering the applied frequency over time. In various non-limiting embodiments, the mass scan rate can be constant.
[0046] In various non-limiting embodiments, the sweeping excitation frequency sweeps from a first frequency to a second frequency that is less than the first frequency. The sweep from the first frequency to the second frequency can be linear, nonlinear, parabolic, inverse Mathieu (e.g., such that time is proportional to 1 / q, where q is the excited ion’s Mathieu q value), or other suitable sweep. In various non-limiting embodiments, the sweeping excitation frequency spans a Mathieu q range for the ions in the ion trap (e.g., q = 0.908 to 0.1). In various non-limiting embodiments wherein the sweep from the first frequency to the second frequency is an inverse Mathieu q range sweep, the relationship between the excited precursor ion m / z value and time of detection can be linear, which can simplify mass calibration and increase processing speed.
[0047] As illustrated in FIG. 5B, in certain embodiments the successive ejection frequency scans of variable duration and variable frequency range can be low AC voltage (e.g., relative to the RF voltage, 100 mV to less than 100 voltages) auxiliary waveforms. The successive ejection frequency scans can excite ions based on their respective m / z values such that the ions are ejected from the ion trap and detected. One scan of the ejection frequency scans is a sinusoidal frequency sweep, which can be from a high frequency to a low frequency (e.g., from a low m / z to a high m / z), and a new scan is started when the ejection frequency reaches a terminal value (e.g., equal to the excitation frequency) and is then returned (e.g., increased) to a starting position (e.g., 1 the RF frequency). For example, as the sweeping excitation frequency decreases (e.g., as ions with larger m / z values are excited), the frequency range of each scan of the successive frequency scans9323430828.1Reference No. P230046-711 -WO017240135PCTmay increase to span the m / z value range of product ions (e.g., larger m / z value range) that would be expected to form by excitation and dissociation of an ion based on the excitation frequency. The duration of each scan may be adjusted based on the frequency range of each scan and can be nonlinear. For example, as the frequency range of each scan increases, the duration of each scan may increase. In various non-limiting embodiments, each scan of the successive ejection frequency scans comprises a frequency range greater than a frequency range of an immediately preceding scan of the successive ejection frequency scans. In certain non-limiting embodiments, the frequency range of each scan can span from half of the RF voltage down to a precursor ion excitation frequency. In various non-limiting embodiments, half of the RF voltage corresponds to the ions at the Mathieu stability boundary of q = 0.908 in a sinusoidally driven ion trap. In various nonlimiting embodiments, the excitation frequency can occur for a time period in a range of 100 milliseconds (ms) to 10 seconds.
[0048] The scan function embodiment illustrated in FIGs. 5A-C can be configured to analyze only the expected m / z range of each precursor ion and fragments thereof and may not include an entire m / z value range. The precursor ions can be analyzed more efficiently and the product ions can be analyzed more often. Thus, increases in density of data for the precursor ions, in sensitivity of detection, and in the signal-to-noise ratio can be achieved.
[0049] For example, in certain embodiments the sweeping excitation frequency can excite a first precursor ion at a first excitation frequency to form a first set of product ions from dissociation of the first precursor ion. A first scan of the successive ejection frequency scans can sequentially eject the first set of product ions from the ion trap based on their respective m / z value. As the sweeping excitation frequency changes, a second precursor ion can be excited at a second excitation frequency to form second set of product ions from dissociation of the second precursor ion. A second scan of the successive ejection frequency scan can sequentially eject the second set of product ions from the ion trap based on their respective m / z value. In various non-limiting embodiments, the first precursor ion can comprise a first m / z value that is less than a second m / z value of the second precursor ion, and the first scan can be performed for a first duration that is less than a second duration of the second scan. In certain non-limiting embodiments, the first scan can be performed over a first frequency range that is less than a second frequency range of the second scan.
[0050] In certain non-limiting embodiments, the scan function comprises a sweeping RF voltage (e.g., FIG. 6C), an excitation frequency that can vary with time or be constant (e.g.,10323430828.1Reference No. P230046-711 -WO017240135PCTFIG. 6A), and successive ejection frequency scans (e.g., FIG. 6B). As illustrated in FIG. 6A, in certain embodiments the excitation frequency can be constant. The excitation frequency can be set at a Mathieu q value in a range of, for example, 0 to 0.908, such as, for example, 0.1 to 0.4.
[0051] As illustrated in FIG. 60, in certain embodiments the sweeping RF voltage can be ramped from a first voltage to a second voltage greater than the first voltage. The ramp of the sweeping RF voltage can be linear, nonlinear, parabolic, or other suitable sweep. In various non-limiting embodiments, the ramp from the first voltage to the second voltage is linear. Changing the RF voltage can change the ions that oscillate within the ion trap, thereby causing ions to resonate with the excitation frequency. For example, if the RF voltage is ramped linearly, the relationship between the m / z value of the excited ion and time can be linear. In various embodiments, the RF voltage can be ramped over a time period in a range of 100 ms to 2 seconds, such as, for example, 200 ms to 1 second or 300 ms to 900 ms.
[0052] In certain embodiments, each scan of the successive ejection frequency scans can comprise the identical predetermined duration and / or identical predetermined frequency range. As the ions are excited, a scan of the ejection frequency scans can eject the product ions as they are formed from the excitation and resulting dissociation of the ions. For example, each scan can span a frequency range from half the RF voltage to the excitation frequency (e.g., FIG. 6B). Each scan can be linear, nonlinear, parabolic, inverse Mathieu, or other suitable sweep. In various embodiments, each scan can be configured such that the relationship between the m / z of each product ion and time is linear to reduce the complexity of the calibration needed. The duration of each scan be in a range of 0.1 ms to 10 ms, such as, for example, 0.2 ms to 5 ms or 0.5 ms to 2 ms.
[0053] In various non-limiting embodiments, the scan function can eject ions in at least two directions. For example, the scan function can comprise an RF voltage (e.g., FIG. 9D), an excitation frequency (e.g., 9A), first ejection frequency scans for the first direction (e.g., 9B), and second ejection frequency scans for the second direction (e.g., 9C). As illustrated in FIG. 9A and 9D, the excitation frequency can be applied substantially similarly to the excitation frequency discussed with respect to FIG. 6A above, and the RF voltage can be applied substantially similarly to the RF voltage discussed with respect to FIG. 6C above. In various non-limiting embodiments, the excitation frequency can be applied in a manner substantially similar to the excitation frequency discussed with respect to FIG. 5A above, and11323430828.1Reference No. P230046-711 -WO017240135PCTthe RF voltage can be applied in a manner substantially similar to the RF voltage discussed with respect to FIG. 50 above.
[0054] The first ejection frequency and the second ejection frequency can be applied to an ion trap at the same time or at different times. In various embodiments, the first ejection frequency and the second ejection frequency can be configured to span different frequency ranges and may not overlap, thereby ejecting ions with different m / z values, which may increase speed of the analysis. For example, the first ejection frequency can eject ions having m / z values in a first m / z value range in a first direction towards a first detector, and the second ejection frequency can eject ions having m / z values in a second m / z value range in a second direction towards a second detector. In various non-limiting embodiments, the first m / z range can be lower than the second m / z range. In various non-limiting embodiments, the scan function comprises applying a first ejection frequency to a first set of two opposing electrodes of the ion trap, applying a second ejection frequency to a second set of two opposing electrodes of the ion trap, and applying an excitation frequency to at least one of the first set and the second set, wherein the first set comprises the first electrode and the second set comprises the second electrode and no electrode in the first set is in the second set.
[0055] In certain embodiments, the first ejection frequency and the second ejection frequency can be configured to at least partially overlap, thereby creating orthogonal detection of ions with the same m / z values, which may increase data quality. In various nonlimiting embodiments, the first ejection frequency and the second ejection frequency can be configured to wholly overlap.
[0056] Noise can originate from ejection of unfragmented precursor ions or production of fragment ions with m / z values less than a lower m / z cutoff (e.g., ions that are immediately ejected from the ion trap due to trajectory instability). The noise can be reduced, for example, by the scan function described with respect to FIGs. 5A-5C, the scan function described with respect to FIGs. 6A-6C, or the scan function described with respect to FIGs.9A-9D. In various non-limiting embodiments, the noise can be reduced by selecting which of pair of opposing electrodes the excitation frequency is applied to. For example, the excitation frequency can be combined on the same rods as one or both excitation frequency.
[0057] Referring back to FIG. 1 , the method 100 can optionally comprise detecting ions ejected from the ion trap and generating spectrum data based on the detected ions, at step 108. The ions can be detected with a single detector or two or more detectors. For example, at least one of a precursor ion, a first-generation product ion, a second-generation12323430828.1Reference No. P230046-711 -WO017240135PCTproduct ion, a third-generation product ion, or a further generation product ion can be detected at step 108. The spectrum data can be, for example, a data set of ion abundance versus time. In various non-limiting embodiments, the spectrum data can be, for example, a two-dimensional data set with respect to mass (e.g., two dimensions of m / z).
[0058] The data can be analyzed at step 110. The analysis can comprise correlating and / or combining spectrum data for precursor ions with the spectrum data for respective product ions. For example, based on the scan function, the time of detection can be correlated to an m / z value and / or a generation of product (e.g., precursor, first-generation, etc.). For example, spectrum data can be generated for a precursor ion of a sample, the sample data can comprise an m / z value and abundance for the precursor ion, and the m / z value and abundance for any product ion generated from the precursor ion, whether first-generation, second-generation, or other generation.
[0059] Spectrum data can be combined from a first spectrum data, a second spectrum data, and optionally other spectrum data to form combined spectrum data. For example, first spectrum data can be generated from detecting a first composition comprising at least a portion of a first set of precursor ions produced from a sample, a first set of first-generation product ions produced from the first set of precursor ions, and a first set second-generation product ions produced from the first set of first-generation product ions. Second spectrum data can be generated from detecting a second set of precursor ions produced from the sample and a second set of first-generation product ions produced from the second set of precursor ions. The first and second spectrum data can be combined to correlate the precursor ions to the first-generation product ions and the second-generation product ions. For example, the second spectrum data may provide the identity of the precursor ions, the first spectrum data may provide the identity of the second-generation product ions, and the precursor ions can be matched with the second-generation product ions by the overlap of the first-generation product ions. The second spectrum data can be generated prior to, at least partially concurrently with, or after generating the first spectrum data.
[0060] The method 100 can enable efficient and rapid characterization of complex mixtures on a molecular level by providing direct measurements of m / z values (e.g., correlating to molecular weight) of the intact ionized molecules (e.g., precursor ions). The detected product ions can be used to deduce structural information (e.g., molecular substructure) based on the fragmentation patterns. The method 100 can be performed on a mass spectrometer in a time period ranging from 10 ms to 10 seconds, such as, for example, 100 ms to 5 seconds, 200 ms to 2 seconds, 200 ms to 1 second, or 300 ms to 90013323430828.1Reference No. P230046-711 -WO017240135PCTms. In various embodiments, the method 100 can be performed with a single ion injection, and the entire ion population can be characterized (by measuring precursor m / z and product m / z simultaneously) in a single analysis scan. In various non-limiting embodiments, the method 100 can comprise two ion injections and multiple spectrum data can be combined. The method 100 can increase the m / z value ranges detected and increase the consistency of product ion m / z values, while using simple scan functions.
[0061] FIG. 2 shows a non-limiting embodiment of a mass spectrometer 200 that can be used to analyze. For example, the mass spectrometer 200 can be configured to measure the m / z values of various ions generated from a sample. For example, the mass spectrometer 200 can be configured to perform the method according to FIG. 1. Referring back to FIG. 2, the mass spectrometer 200 can comprise an ionizer 220, an ion trap 222 in ion communication with the ionizer 220, a detector 224 in ion communication with the ion trap 222, and a controller 226 in signal communication with the ion trap 222 and the detector 224. In various non-limiting embodiments, the mass spectrometer comprises a single ion trap 222.
[0062] As used herein, “ion communication” means that the recited elements are configured with features such that ions, gas molecules, and / or the like can be transmitted between the two elements. In various non-limiting embodiments, the ion communication can be generation by control of ions between the recited elements by an electric field or by a physical structure (e.g., a tube, walls defining a bore).
[0063] The ionizer 220 can be in ion communication with the ion trap 222 via an ion conduit 228 suitable to transfer the sample and / or ions from the ionizer 220 to the ion trap 222 such that gases, vapors, particles entrained in a gas, and / or ions can be transferred from the ionizer 220 to the ion trap 222 through the ion conduit 228. The ion trap 222 can be in ion communication with the detector 224 by an ion conduit 230 suitable to transfer ions ejected from the ion trap 222 to the detector 224 and so that gases, vapors, particles entrained in a gas, and / or ions can flow from the ion trap 222 to the detector 224 through the ion conduit 230. The ion conduits 228 and 230 can be a suitable ion pathway, which may be generated by an electric field and / or comprise a physical structure.
[0064] The ionizer 220, as supported by the controller 226, is configured to ionize a sample, thereby generating a precursor ion or precursor ions from the sample. The ionizer 220 can be configured to perform at least one of electron ionization, photo ionization, chemical ionization, collisionally activated disassociation, electrospray ionization, and other suitable methods of ionization on the sample.14323430828.1Reference No. P230046-711 -WO017240135PCT
[0065] The controller 226 can be configured to control the functionality of the mass spectrometer 200. The controller 226 can be in signal communication with the ionizer 220, the ion trap 222, and the detector 224. For example, the controller 226 can communicate with the ionizer 220, the ion trap 222, and the detector 224 through physical wires and / or wireless signals. The controller 226 can comprise, for example, a processor operatively coupled to memory, a DC voltage source, an AC voltage source, and a rectifier, and may comprise other hardware components. For example, the DC voltage source, AC voltage source, and rectifier may be used to generate the scan function. The memory can comprise machine executable instructions for controlling the ionizer 220, ion trap 222, and detector 224. In various non-limiting embodiments, the memory can include instructions for analyzing the data received from the detector 224. The processor can be a microprocessor, microcontroller, or other basic computing device that incorporates the functions of a computer’s central processing unit (CPU) on an integrated circuit.
[0066] The ion trap 222, as supported by the controller 226, can be configured to receive the sample and / or ions, ionize the sample, store the ions in the ion trap 222, excite at least a portion of the ions in the ion trap 222 selectively over time to fragment the ions, and / or eject ions from the ion trap 222. For example, the controller 226 can be configured to apply a scan function to the ion trap 222 to control ions within the ion trap 222. The controller 226 can generate and apply to the ion trap 222 an RF voltage, an excitation frequency, and / or an ejection frequency to create an electric field, such as, for example, an oscillating potential well within the ion trap 222 that can selectively store, excite, and / or eject ions. As the RF voltage, the excitation frequency, and / or the ejection frequency changes, the electric field within the ion trap 222 can change, which can affect the ions stored, excited, and / or ejected by the ion trap 222. In various non-limiting embodiments, the ion trap 222, in conjunction with the controller 226, can be configured to ionize the sample, thereby generating precursor ions from the sample within the ion trap 222.
[0067] In various embodiments, the ion trap 222 can be a quadrupole ion trap, such as, for example, a 3D quadrupole ion trap, a linear quadrupole ion trap, a toroidal ion trap, a cylindrical ion trap, or a rectilinear ion trap. For example, the ion trap 222 can be a linear quadrupole ion trap 322 as illustrated in FIGs. 3 and 4. Referring to FIG. 4, the ion trap 322 can comprise four electrodes 340, 342, 344, 346 defining a cavity 360. Ions can be introduced to, stored in, excited in, and / or ejected from the cavity 360. Each electrode 340, 342, 344, and 346 can be in electrical communication with the controller 226. In various non-limiting embodiments, the ions are collisionally cooled in the ion trap 222 before performing any scan functions on the ions. Collisional cooling can reduce the kinetic energy15323430828.1Reference No. P230046-711 -WO017240135PCTof the ions after injection into the trap or after a first fragmentation event but prior to the scan function being applied, which can increase data quality. Collision cooling can comprise storing ions in the ion trap 222 for a period time prior to the scan function such that the ions collide with background gas molecules (e.g., nitrogen, air, helium) and lose kinetic energy.
[0068] Referring to FIG. 3, in certain embodiments the ion trap 322 can comprise caps 348 (e.g., endcaps), 350, which can be configured to confine ions along a longitudinal axis, Al, extending from cap 348 to cap 350. The caps 348, 350 can comprise bores 352, 354, respectively, that can be suitable to form ion pathways into the cavity 360 of the ion trap 322 such that gases, vapors, particles entrained in a gas, and / or ions can flow into the cavity 360 of the ion trap 322.
[0069] Referring again to FIG. 4, in various embodiments the electrodes 340, 342, 344, and 346 can be configured in sets. For example, a set of electrodes can comprise pairs of opposing electrodes, such as, for example, a first set of electrodes including electrodes 340 and 344 and a second set of electrodes including electrodes 342 and 346. In certain nonlimiting embodiments, the RF voltage can be applied in a quadrupolar fashion such that a positive phase of the RF voltage can be applied to first set of opposing electrodes and a negative phase of the RF voltage can be applied to second set of opposing electrodes, where no electrode in the first set is in the second set. In various non-limiting embodiments, a single phase of the RF voltage can be applied to the first set of opposing electrodes and the second set of opposing electrodes can be virtually grounded. In certain non-limiting embodiments regarding AC signals (e.g., ejection frequency, excitation frequency), the AC signal can be applied in a dipolar fashion such that to the positive phase of the AC signal can be applied to one electrode of a first set of opposing electrodes and a negative phase of the AC signal can be applied to a second electrode in the first set of opposing electrodes.
[0070] To control the ions, the controller 226 can apply a scan function to electrodes 340, 342, 344, and 346 of the ion trap 322, such as, for example, the scan function shown in FIGs. 5A-5C, the scan function shown in FIGs. 6a-6C, and / or the scan function shown in FIGs. 9A-9D. For example, the RF voltage can be applied in quadrupolar fashion to all of electrodes 340, 342, 344, and 346 of the ion trap 322 or to one set of electrodes. In various non-limiting embodiments, the excitation frequency can be applied to all of electrodes 340, 342, 344, and 346 of the ion trap 322. The excitation frequency and ejection frequency may be applied in a dipolar fashion (e.g., 180 degrees out of phase) either on the same set of electrodes or on orthogonal electrodes pairs. In certain non-limiting embodiments, the controller 226 can be configured to apply the excitation frequency and the ejection frequency16323430828.1Reference No. P230046-711 -WO017240135PCTscans to a first set of electrodes. In various non-limiting embodiments, the controller 226 can be configured to apply the excitation frequency to a first set of electrodes, and to apply the ejection frequency scans to a second set of electrodes, wherein no electrode in the first set is in the second set.
[0071] In certain non-limiting embodiments, the ions are excited in the first dimension of the ion trap 322, and the excited ions are ejected in the second dimension, which differs from different the first dimension, towards a detector. As illustrated, the ion trap 322 has an opening 332 on the electrode 340, which can be aligned with the detector 224 such that the opening 332 can form the ion conduit 230.
[0072] Referring back to FIG. 2, in various non-limiting embodiments, the ion trap 222 can be configured to fragment ions within the ion trap 222. For example, the controller 226 can apply a scan function to the ion trap 222 and introduce a gas that causes the ions to collide with gas molecules within the ion trap 222, thereby creating product ions. For example, the ion trap 222, as supported by the controller 226, can fragment a precursor ion into first-generation product ions, a first-generation product ion into second-generation product ions, a second-generation product ion into a third-generation product ions, and / or another product ion.
[0073] Referring to FIGs. 7 and 8, in various embodiments the ion trap 222 can be a linear quadrupole ion trap 422. The controller 226 can be configured to apply a scan function to the ion trap 422 to eject ions from the ion trap 422 in a first direction 470 and a second direction 472 different than the first direction 470. The controller 226 can be configured to eject ions from the ion trap 422 in the first direction 470 during a first time period, and to eject ions in the second direction 472 during a second time period. The first time period can at least partially overlap with the second time period, or the first time period may not overlap with the second time period.
[0074] The ion trap 422 comprises at least two openings. The openings may include, for example, a first opening 432a in the electrode 440 of the ion trap 422 configured to receive ejected ions from the ion trap 422 along a first path in the first direction 470, and a second opening 432b in the electrode 442 of the ion trap 422 configured to receive ejected ions from the ion trap 422 along a second path in the second direction 472. In various non-limiting embodiments, the first opening 432 can be oriented substantially orthogonal to the second opening 432b. As illustrated, in certain embodiments first opening 432a and second opening 432b can form the ion conduit 230.17323430828.1Reference No. P230046-711 -WO017240135PCT
[0075] In certain embodiments the ion trap 422 can comprise caps 448 (e.g., endcaps), 450, which can be configured to confine ions along a longitudinal axis, Ai, extending from cap 448 to cap 450. The caps 448, 450 can comprise bores 452, 454, respectively, that can be suitable to form ion pathways into the cavity 460 of the ion trap 422 such that gases, vapors, particles entrained in a gas, and / or ions can flow into the cavity 460 of the ion trap 422.
[0076] In various non-limiting embodiments, the detector 224 can comprise at least two detectors, such as, for example, first detector 424a and second detector 424b, as illustrated in FIG. 8. For example, the first detector 424a can be in a first location and aligned with the first opening 432a such that ions ejected from the ion trap 422 in the first direction 470 are received and detected by the first detector 424a. The second detector 424b can be in a second location and aligned with second opening 432b such that ions ejected from the ion trap 422 in the second direction 472 are received and detected by the second detector 424b.
[0077] To control the ions, the controller 226 can apply a scan function to electrodes 440, 442, 444, and 446 of the ion trap 422, such as, for example, the scan function as shown in FIG. 9. For example, a first ejection frequency can be applied to a first set of two opposing electrodes of the electrodes 440, 442, 444, and 446 of the ion trap 422, a second ejection frequency can be applied to a second set of two opposing electrodes of the electrodes 440, 442, 444, and 446 of the ion trap 422, and an excitation frequency can be applied to at least one of the first set and the second set, wherein no electrode in the first set is in the second set. In various non-limiting embodiments, in order to eject ions in the first direction 470 and the second direction 472, the first ejection frequency scan is applied to electrodes 440 and 444, and the second ejection frequency scan is applied to electrodes 442 and 446.
[0078] The RF voltage can be applied in a quadrupolar fashion to all of electrodes 440, 442, 444, and 446 of the ion trap 422 or a single phase of RF voltage can be applied to one set of electrodes. In various non-limiting embodiments, the excitation frequency can be applied to all of electrodes 440, 442, 444, and 446 of the ion trap 322. The excitation frequency and ejection frequencies may be applied in a dipolar manner.
[0079] Referring yet again to FIG. 2, the detector 224, as supported by the controller 226, can be configured to detect ions ejected from the ion trap 222. For example, the detector 224, as supported by the controller 226, can generate data (e.g., a mass spectrum data) comprising the m / z value of the detected ions and an abundance (e.g., intensity) of the detected ions at the m / z value. In various non-limiting embodiments, the detector 224 can18323430828.1Reference No. P230046-711 -WO017240135PCTcomprise at least one of an electron multiplier, a Faraday cup collector, a photographic and stimulation-type detector, and other detector type.
[0080] Referring back to FIG. 8, in various non-limiting embodiments comprising at least two detectors, the first detector 424a can be configured to detect ions ejected from the ion trap 422 in the first direction 470 during a first time period, and the second detector 424b can be configured to detect ions ejected from the ion trap 422 in the second direction 472 during a second time period, and the first time period at least partially overlaps with the second time period. In certain non-limiting embodiments comprising at least two detectors, the first time period does not overlap with the second time period.
[0081] Referring back to FIG. 2, the detector 224 is configured to produce data based on the ions received and detected from the ion trap 222. The detector 224 can send the data to the controller 226. In various non-limiting embodiments, the data can be processed separately and may be combined to form combined spectrum data. For example, referring to FIG. 8, a spectrum can be generated by combining first data received from the first detector 424a and second data received from the second detector 424b. The first data can be based on ions ejected from the ion trap 422 in the first direction 470, and the second data can be based on ions ejected from the ion trap 422 in the second direction 472. In various non-limiting embodiments, the controller 226 is further configured to correlate and / or combine data received from the first detector 424a and data received from the second detector 424b to produce combined spectrum data.
[0082] For a given sample, three dimensions of mass information can be acquired. For example, when a precursor ion fragments, first-generation product ions are formed. The first-generation product ions can be fragmented into second-generation product ions. The correlation of precursor ions, first-generation product ions, and second-generation product ions can be collected and analyzed to provide the three dimensions of mass information.
[0083] As shown in FIG. 10, the mass spectrometer 200 can optionally comprise a collision cell 521 in fluid communication with the ion trap 222 (via an ion conduit 328b) and the ionizer 220 (via an ion conduit 328a). The collision cell 521 can be configured for fragmenting at least a portion of a sample and / or ions, thereby forming a first mixture of ions prior to the ion trap 222. The first mixture can comprise one or more precursor ions and first-generation product ions formed from fragmentation of the precursor ions. In various nonlimiting embodiments, the collision cell 521 can be configured for beam-type collision induced dissociation.19323430828.1Reference No. P230046-711 -WO017240135PCT
[0084] The first mixture can be provided to and received by the ion trap 222. The ion trap 222 can be configured to apply a scan function to the first mixture. The scan function can comprise, for example, exciting at least a portion of the first mixture selectively over time to fragment the first-generation product ions into second-generation product ions, and ejecting at least a portion of the precursor ions, the first-generation product ions, and the second-generation product ions formed from first-generation product ions from the ion trap 222. The detector 224 can be configured to detect at least a portion of the precursor ions, the first-generation product ions, and the second-generation product ions ejected from the ion trap 222 and generate spectrum data.
[0085] In various non-limiting embodiments, the ion trap 222 can be configured to fragment at least a portion of a sample by applying a first scan function, thereby forming the first mixture within the ion trap 222, and then applying a second scan function to the first mixture. The ion trap 222 can be configured for at least one of beam-type collision induced dissociation within the ion trap 222 and collision-induced dissociation by resonance excitation within the ion trap 222. In various non-limiting embodiments, the ion trap 222 and / or collision cell 521 can be configured to fragment at least a portion of the sample until at least third-generation product ions are formed from the sample.
[0086] FIG. 11 is a method 600 of developing a machine learning model for identifying substances, according to a non-limiting embodiment of this disclosure. The method 600 may use a database comprising a plurality of (i.e. , more than one) sets of data. Each set of data may include an identity of a known substance and tandem mass spectrometry spectrum data corresponding to the known substance. The database comprises a plurality of sets of data for each of the known substances. The database may be formed from the data received at the controller 226 in FIG. 2. The data may be received from, for example and without limitation, any of the mass spectrometers described in FIG. 2 or FIG. 10.
[0087] The method 600 may include collecting 2DMS data from a plurality of known samples to form the database. For example, the 2DMS data of the sample may be collected from the mass spectrometer of FIG. 2 or FIG. 10 or any mass spectrometer capable of generating 2DMS data. The 2DMS data may include a mass to charge ratio of precursor ions, a mass to charge ratio of product ions, and intensity. The intensity may be between 0 and 1. For example, the spectrum data can be generated for a product ion and a precursor ion of a sample. The spectrum data may comprise the m / z value and the abundance (intensity).20323430828.1Reference No. P230046-711 -WO017240135PCT
[0088] In certain embodiments, the 2DMS data may be calibrated along the precursor and product mass to charge axes. The 2DMS data may be normalized to calculate the intensity between 0 and 1. The data may also be subject to digital processing and / or smoothing, for example using a 1D or 2D Gaussian filter or a Savitzky-Golay filter. The data may also be subject to processing to remove noise from the data.
[0089] For example, the precursor ions can be the individual molecules in the sample modified by addition of an electrical charge. The product ions may be formed through fragmentation of the product ions. The m / z values of the ions generated may be the same or different.
[0090] To form the 2DMS data, the spectrum data can be, for example, a data set of ion abundance versus time. In various non-limiting embodiments, the spectrum data can be, for example, a two-dimensional data set with respect to mass (e.g., two dimensions of m / z). For example, this can comprise correlating and / or combining spectrum data for precursor ions with the spectrum data for respective product ions.
[0091] For example, spectrum data can be generated for a precursor ion of a sample, the sample data can comprise an m / z value and abundance for the precursor ion, and the m / z value and abundance for any product ion generated from the precursor ion, whether first-generation, second-generation, or other generation. An example of this process is shown in FIG. 3.
[0092] Optionally, the method comprises forming the database with the spectrum data collected for known substances using mass spectrometers. The method may optionally comprise augmenting the sets of data in the database. For example, augmenting occurs by making copies of each 2DMS spectrum with random intensity and mass shifts. Mass shifts are assumed to affect both the precursor and product axes, though to different extents since the scan rate for the latter typically is much higher than for the former. The augmented sets of data are added to the database. The database may comprise both augmented 2DMS data and 2DMS data from a mass spectrometer. The database may comprise, alone or in combination with the augmented 2DMS data and 2DMS from a mass spectrometer, artificial 2DMS data such as 2DMS data created using spectral libraries of known compounds. The database may include 2DMS data, CID-2DMS data, or a combination of 2DMS and CID-2DMS data. The database may include 2DMS data (i.e. , a 2D dataset consisting of MS1and MS2ions) or 2DMS data generated from collections of MS1, MS2, and more generally any set of MSndata.21323430828.1Reference No. P230046-711 -WO017240135PCT
[0093] FIG. 15 illustrates a convolutional neural network to identify fentanyl and fentanyl-variant containing samples, in one aspect of the present disclosure. The input data is either 2DMS spectra data 1000 or CID-2DMS spectra data 1002. The example spectra were generated from a mixture of o-isopropyl furanyl fentanyl, p-fluoro tetrahydrofuran fentanyl, p-methoxy furanyl fentanyl, m-methyl furanyl fentanyl, furanylethyl fentanyl, furanyl fentanyl 3-furancarboxamide, tetrahydrofuran fentanyl 3-tetrahydrofurancarboxamide, p-methoxy tetrahydrofuran fentanyl, N-benzyl furanyl norfentanyl, and p-methyl furanyl fentanyl.
[0094] The 2DMS spectra 1000 and the CID-2DMS spectra 1002 can be used to identify variants. The accuracy chart 1004 shows that the machine learning model with 2DMS data correctly identified the samples, as the true label and predicted matched. The accuracy chart 1006 shows that the machine learning model with CID-2DMS data correctly identified the samples 94-99% of the time, as the true label and predicted matched.
[0095] The accuracy shown in accuracy chart 1006 in the CID-2DMS analysis is likely due to the overall lower signal-to-noise ratio due to loss of signal during the CID process. However, in some situations machine learning analysis of CID-2DMS data may provide increased accuracy due to the presence of MS1, MS2, and MS3 ions in the 2DMS pattern. In this case the peaks at precursor m / z 188, 134, and 105 are common amongst many fentanyl variants and create a common pattern by which new fentanyl variants may be accurately identified through artificial intelligence.
[0096] The method may include calibrating the 2DMS data in the database. Calibration may involve calibrating the mass spectrometry data along the precursor and product m / z axes. Optionally, the data may be rectangularized by interpolation and resampling.Calibration may be performed before the model is trained.
[0097] The method may include normalizing the data in the database. The intensity of the precursor and product ions are between, for example, 0 and 1 after normalization.Normalization may be performed before the model is trained.
[0098] For example, FIG. 12 illustrates examples of 2DMS spectra 700 collected after calibration and processing. As shown, three 2DMS spectra were collected for each of four bacterial substances: E. Faecalis, B. thuringiensis, C. koseri, and S. aureus. It will be understood, however, that the number of substances may be less or greater than the four shown and the number of 2DMS spectra data collected for each substance may be less or greater than three.22323430828.1Reference No. P230046-711 -WO017240135PCT
[0099] The sets of data include the identity, E. faecalis, B. thuringiensis, C. koseri, and S. aureus, and the spectra data. Thus, as noted, each identity has three different spectra collected.
[0100] The method 600 comprises obtaining 602 a first portion of data from the database. The first portion of data may comprise a portion of the sets of data in the database. The first portion may comprise, for example, more than half of the sets of data. For example, the first portion may comprise 70% of the data sets. The database, for example, may be split into a first portion and a second portion, and the first portion and the second portion may or may not be the same size.
[0101] The method 600 includes inputting 604 the first portion of data into a machine learning model. The machine learning model is trained to identify pathogens based on the 2DMS data. The machine learning model may comprise one of a convolutional neural network, a random forest classifier, a recurrent neural network, K-nearest neighbors, multilayer perceptron, autoencoders, neuromorphic networks, recurrent neural networks, transformer architecture, long short-term memory neural network, weighted bi-directional feature pyramid network, or any type of artificial intelligence or machine learning model.
[0102] The method 600 includes training 606 the machine learning model with the first portion of the data to identify the identity of substances based on the tandem mass spectrometry data of the substance.
[0103] For example, the method may include repeating the training step. The model may be re-trained with additional data from the mass spectrometer (such as mass spectrometer 200). The model may be able to “learn” (i.e., be re-trained) as new data comes in from a mass spectrometer. The model can be re-trained onboard the mass spectrometer, or the data can be sent back to a computing device at a central location where the model can be re-trained and then the re-trained model is pushed to the instrument (e.g., mass spectrometer) in the field.
[0104] The method may include receiving data from numerous different mass spectrometers. The data from mass spectrometers may continuously update and re-train the model. The re-trained model may be uploaded to the controllers in the numerous mass spectrometers.
[0105] The method 600 further comprises inputting 608 a second portion of data from the database into the machine learning model. The second portion of the data may23323430828.1Reference No. P230046-711 -WO017240135PCTcomprise the remaining portion of data in the database. The first and second portion of the data comprise different data sets of data from the database.
[0106] The method 600 further includes evaluating 610 the machine learning model using the second portion of the data. Evaluating the machine learning model may include determining, using the trained machine learning model, the identity of each sample in the second portion of the data based on the tandem mass spectrometry spectrum data. The method then determines whether the determined identity of each sample matches the identity of each sample stored in the database for the corresponding mass spectrum data. Based on the determined identity of the sample matching the identity of the sample in the database, the accuracy of the machine learning model is determined. If the identity determined for the sample does not match the identity of the sample in the database, the machine learning model may need to be re-trained with a different data set or with additional data sets and / or using a different machine learning model type.
[0107] For example, FIG. 13 illustrates validation of a machine learning model according to a non-limiting embodiment of this disclosure. The validation table 800 comprises three dimensions: the predicted label, the true label, and the number of substances at each point. The true label is the identity of the substance stored in the database. The predicted label is the identity of the substance predicted from the mass spectrum data.
[0108] When the predicted label and the true label match, the machine learning model has identified the correct substance for the mass spectrum data. When the predicted label and the true label do not match, the machine learning model has incorrectly identified the substance based on the mass spectrum data. Thus, the table shown in FIG. 13 illustrates the number of accurate determinations made applying the machine learning model to the mass spectrum data.
[0109] In various example, the method 600 may be implemented by the controller 226 of FIG. 2 or on a control circuit separate from the controller 226 of FIG. 10.
[0110] For example, developing a neural network trained on 2DMS data for identifying samples may involve utilizing a plurality of 2DMS spectra per substance, including biological replicates, replicates at varying concentrations, and under varying growth conditions. The spectra data may also include replicates from different mass spectrometers such as the different types of mass spectrometers shown in FIG. 2 and FIG. 10 or from varying instruments of the same general type. Depending on the machine learning model, the data must then be calibrated and processed. In the case of the convolutional neural network24323430828.1Reference No. P230046-711 -WO017240135PCTused in the present example, the data is reformed into an image, smoothed, and normalized (precursor m / z, product m / z, and intensities) between 0 and 1 prior to being sent to the model. The model is then trained on the dataset, the accuracy and loss are calculated, and, if the accuracy is high (>95%), the model is saved and can be used to identify unknowns from 2D MS / MS spectra not yet seen by the model.
[0111] For example, as shown in FIG. 13, the machine learning model can be trained using 2DMS data files consisting of 18 analytical replicates for each of 3 biological replicates for 8 bacteria species (BTU = B. thuringiensis, Bsub = B. subtilis, CFR = C. freundii, CK = C. koseri, EFM = E. faecalis, Psaer = P. aeruginosa, Saur = S. aureus, Scap = S. capitis). A set of programs for classification of samples using a convolutional neural network may be used. The programs consist of three major components. First, the data is mass calibrated along both the precursor and product m / z axes and subsequently rectangularized by interpolation and resampling. Next, the entire dataset is imported, and each spectrum is given an appropriate label based on the known contents of the sample (e.g., BTU, Bsub, etc.). A 2D Gaussian filter smooths the data to improve signal-to-noise ratios, giving, for example, the spectra in FIG. 2. The dataset may be pre-processed by x / y resampling. The dataset is then augmented by making copies of each 2D MS / MS spectrum with random intensity and mass shifts. Mass shifts are assumed to affect both the precursor and product axes, though to different extents since the scan rate for the latter is much higher than for the former. In total, the data sets comprise original files and augmented files for training and validation.
[0112] Next a convolutional neural network is constructed using several layers of a 2D convolutional layer (Conv2d) and rectified linear activation unit (ReLU) activation functions as well as 2D max pooling over an input signal composed of several input planes (MaxPool2D) for pooling the convolution results. The final activation function was a sigmoid and a binary cross entropy loss function was chosen to enable multilabel classification. The model was trained for 6 epochs using a batch size of 20 files and a 70 / 30 data split (i.e., 70% of the data is used to train the model and 30% of the data is used to validate the model). The model may be trained several times on different training sets to achieve different accuracies. For example, the system may vary the first portion of data and second portion of data to achieve better accuracies.
[0113] One considering the present disclosure will understand that the example convolution neural network discussed herein is only a single example of applying machine25323430828.1Reference No. P230046-711 -WO017240135PCTlearning to 2DMS data and that other models and methods of processing the data may be appropriate for different types or identities of samples / threats.
[0114] Once the model is trained, the method 600 may further include receiving the tandem mass spectrometry data from the mass spectrometer for the sample. The control circuit that implements the method 600 may be coupled to a mass spectrometer to receive the data. The machine learning model may identify the sample using the machine learning model and the tandem mass spectrometry data.
[0115] The sample provided to the mass spectrometer for identification by the machine learning model may or may not be a substance within the database. The machine learning model may be able to identify substances not within the database.
[0116] FIG. 14 is a method of determining the identity of a substance, according to an embodiment of this disclosure. The method 900 may be implemented on a control circuit or by a processor and memory. The memory stores a plurality of instructions and a pre-trained machine learning model executable by the processor to cause the processor to execute the method.
[0117] The method 900 includes receiving 902 tandem mass spectrometry data for an unknown sample. The method includes storing a pre-trained machine learning model. The pre-trained machine learning model is trained using a database comprising a plurality of known substances and a plurality of tandem mass spectrometry spectrum data for each of the known substances. The pre-trained machine learning model is trained to classify the data with one of the substances in the database. The pre-trained model may be trained by the method 600 of FIG. 6.
[0118] For an unknown sample, the machine learning model, when trained with comprehensive 2DMS datasets to identify analytes and their near-neighbor variants, can also then identify unseen variants (variants not present in a spectral library and also not previously seen by the model). For example, as shown in FIG. 16, all data files (and copies) containing m-fluorofentanyl and p-fluoro furanyl fentanyl were left out of the training and validation sets (i.e., they were unseen by the model at both training and validation stages) during the initial train / test stage 1010. The model, a convolutional neural network (CNN), was then instructed to classify 1012 the 486 “unseen” files as fentanyl-containing (“Fent”) or not fentanyl-containing (“Other”). In this example, the CNN was able to correctly classify these near-neighbor variants with 95% accuracy. The accuracy could be further improved by increasing the size and variety of the dataset and by signal-to-noise filtering (removing26323430828.1Reference No. P230046-711 -WO017240135PCTpoor quality spectra). In any case, this example demonstrates that a machine learning model can accurately classify unseen near-neighbor variants of a particular molecule based on 2DMS data.
[0119] The method 900 further includes using the pre-trained machine learning model and determining 904 an identity of the sample. The identity is one of a substance stored in a database to train the machine learning model or a substance not used to train the machine learning model.
[0120] Identification of samples may include classifying the sample in a class of molecules or pathogens such as a sample X as containing a fentanyl variant or being an E. coli variant, or identifying the exact molecule such as classifying the sample as fentanyl-like or identifying the sample as fluorofentanyl or E. coli. Identification may be classifying the sample in a class.
[0121] In some embodiments, the models are also able to identify near-neighbor molecular variants of molecules (e.g., molecule Y, which is similar in structure to fentanyl and p-fluorofentanyl but not identical, is a fentanyl). These new near-neighbor variants can then be added to the model’s training dataset so that new instances of the new variant can be identified.
[0122] The method may further include providing an output indicating the identity of the sample.
[0123] The method 900 may be implemented by the controller 226 of FIG. 2 or may be implemented on a control circuit separate from the controller 226 of FIG. 10.
[0124] The method 600 and the method 900 may be implemented on the same controller or control circuit. For example, the method 600 of training the model may be implemented before the method 900 of identifying substances. The methods described herein may be implemented on a control circuit or on controller 226 of FIG. 2 or FIG. 10.
[0125] The following numbered clauses are directed to various non-limiting embodiments according to the present disclosure:
[0126] Clause 1: An apparatus comprising:a database comprising a plurality of sets of data, wherein each set of data comprises an identity of a known substances and tandem mass spectrometry spectrum data for the known substance;27323430828.1Reference No. P230046-711 -WO017240135PCTa control circuit comprising a processor and a memory, wherein the memory stores instructions executable by the processor to:obtain a first portion of data from the database;input the first portion of the data into a machine learning model; train the machine learning model with the first portion of the data to form a trained machine learning model, wherein the machine learning model is trained to identify the identity of substances based on the tandem mass spectrometry spectrum data of the substance;input a second portion of data from the database into the machine learning model, wherein the second portion and the first portion comprises different sets of data from the database; andevaluate the machine learning model using the second portion of the data.
[0127] Clause 2. The apparatus of clause 1 , wherein the processor is further configured to: receive tandem mass spectrometry data for a sample; and determine, using the trained machine learning model, the identity of the sample.
[0128] Clause 3. The apparatus of any of clauses 1-2, wherein the tandem mass spectrometry spectrum data for the known substance comprises: a precursor mass to charge ratio; a product mass to charge ratio; and an intensity.
[0129] Clause 4. The apparatus of any of clauses 1-3, wherein the machine learning model comprises one of a convolutional neural network, a random forest classifier, a recurrent neural network, K-nearest neighbors, multilayer perceptron, transformer architecture, long short-term memory neural network, or weighted bi-directional feature pyramid network.
[0130] Clause 5. The apparatus of any of clauses 1-4, wherein the machine learning model is configured to identify a known sample based on the tandem mass spectrometry spectrum data provided.
[0131] Clause 6. The apparatus of any of clauses 1-5, further comprising: a mass spectrometer to collect tandem mass spectrometry spectrum data from a sample; and wherein the control circuit is to: receive the tandem mass spectrometry data from the mass spectrometer for the sample; and identify the sample using the machine learning model and the tandem mass spectrometry data.28323430828.1Reference No. P230046-711 -WO017240135PCT
[0132] Clause 7. The apparatus of clause 4, wherein the control circuit is configured to: create the database by: receiving tandem mass spectrometry data for known substances.
[0133] Clause 8. The apparatus of any of clauses 1-7, wherein the instructions executable by the processor to evaluate the machine learning model comprise: determining the identity of each sample in the second portion of the data based on the tandem mass spectrometry spectrum data; determining whether the determined identity of each sample matches the identity of each sample stored in the database for the corresponding mass spectrum data; and based on the determined identity of the sample matching the identity of the sample in the database, determining the machine learning model is accurate.
[0134] Clause 9. The apparatus of any of clauses 1-8, wherein the control circuit is further configured to: augment the database by copying the sets of data and adding random intensity and mass shifts to the tandem mass spectrometry spectrum data to create augmented data; and add the augmented data to the database.
[0135] Clause 10. An apparatus comprising:a processor; anda memory storing a plurality of instructions and a pre-trained machine learning model executable by the processor, wherein the pre-trained machine learning model is trained using a database comprising a plurality of known substances and a plurality of tandem mass spectrometry spectrum data for each of the known substances, wherein the pre-trained machine learning model is trained to classify the data with one of the substances in the database, wherein the instructions are configured to cause the processor to:receive tandem mass spectrometry data for an unknown sample; and using the pre-trained machine learning model, determine an identity of the sample, wherein the identity is one of a substance stored in a database to train the pretrained machine learning model or a substance not used to train the machine learning model.
[0136] Clause 11. The apparatus of clause 10, wherein the tandem mass spectrometry spectrum data is received from a mass spectrometer.
[0137] Clause 12. The apparatus of any of clauses 10-11, wherein tandem mass spectrometry spectrum data comprises: a precursor mass to charge ratio; a product mass to charge ratio; and an intensity.29323430828.1Reference No. P230046-711 -WO017240135PCT
[0138] Clause 13. The apparatus of clause 12, wherein processor is configured to: calibrate the tandem mass spectrometry spectrum data along the precursor and product mass to charge axes.
[0139] Clause 14. The apparatus of any of clauses 10-13, wherein processor is configured to: normalize the tandem mass spectrometry spectrum data with a 2D Gaussian filter.
[0140] Clause 15. The apparatus of any of clauses 10-14, wherein the machine learning model comprises one of a convolutional neural network, a random forest classifier, a recurrent neural network, K-nearest neighbors, multilayer perceptron, transformer architecture, long short-term memory neural network, or weighted bi-directional feature pyramid network.
[0141] Clause 16. A method comprising:obtaining data from a database, wherein the database comprises a plurality of sets of data, wherein each set of data comprises an identity of a known substances and tandem mass spectrometry spectrum data for each of the known substances;inputting a first portion of the data into a machine learning model; training the machine learning model with the first portion of the data, wherein the machine learning model is trained to identify the identity of substances based on the tandem mass spectrometry spectrum data of the substance;inputting a second portion of data from the database into the machine learning model, wherein the second portion and the first portion comprise different sets of data from the database;evaluating the machine learning model using the second portion of the data;determining the machine learning model is accurate based on the evaluation;based on the machine learning model being accurate, receiving tandem mass spectrometry data from an unknown sample;determining, using the machine learning model stored in a memory of a computing device, the identity of an unknown sample; andproviding an output indicating the identity of the sample.30323430828.1Reference No. P230046-711 -WO017240135PCT
[0142] Clause 17. The method of clause 16 wherein the tandem mass spectrometry spectrum data for each of the known substances comprises: a precursor mass to charge ratio; a product mass to charge ratio; and an intensity.
[0143] Clause 18. The method of any of clauses 16-17, wherein the machine learning model comprises one of a convolutional neural network, a random forest classifier, a recurrent neural network, K-nearest neighbors, multilayer perceptron, transformer architecture, long short-term memory neural network, or weighted bi-directional feature pyramid network.
[0144] Clause 19. The method of any of clauses 16-18, wherein the tandem mass spectrometry data is received from a mass spectrometer.
[0145] Clause 20. The method of any of clauses 16-19, wherein the identity of the sample is stored within the database.
[0146] Having thus described several aspects and embodiments of the technology of this application, it is to be appreciated that various alterations, modifications, and improvements will readily occur to those of ordinary skill in the art. Such alterations, modifications, and improvements are intended to be within the scope of the technology described in the application. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may be practiced otherwise than as specifically described. In addition, any combinations of two or more features, systems, articles, materials, and / or methods described herein, if such features, systems, articles, materials, and / or methods are not mutually inconsistent, are included within the scope of the present disclosure.
[0147] In most embodiments, a processor may be a physical or virtual processor. In other embodiments, a virtual processor may be spread across one or more portions of one or more physical processors. In certain embodiments, one or more of the embodiments described herein may be embodied in hardware such as a Digital Signal Processor (DSP). In certain embodiments, one or more of the embodiments herein may be executed on a DSP. One or more of the embodiments herein may be programmed into a DSP. In some embodiments, a DSP may have one or more processors and one or more memories. In certain embodiments, a DSP may have one or more computer readable storages. In many embodiments, a DSP may be a custom designed ASIC chip. In other embodiments, one or31323430828.1Reference No. P230046-711 -WO017240135PCTmore of the embodiments stored on a computer readable medium may be loaded into a processor and executed.
[0148] Also, as described, some aspects may be embodied as one or more methods. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.
[0149] The phrase “and / or”, as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases.
[0150] As used herein in the specification and in the claims, the phrase “at least one”, in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified.
[0151] The terms “approximately” and “about” may be used to mean within ±20% of a target value in some embodiments, within ±10% of a target value in some embodiments, within ±5% of a target value in some embodiments, and yet within ±2% of a target value in some embodiments. The terms “approximately” and “about” may include the target value.
[0152] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. The transitional phrases “consisting of” and “consisting essentially of” shall be closed or semi-closed transitional phrases, respectively.
[0153] Where a range or list of values is provided, each intervening value between the upper and lower limits of that range or list of values is individually contemplated and is encompassed within the disclosure as if each value were specifically enumerated herein. In addition, smaller ranges between and including the upper and lower limits of a given range are contemplated and encompassed within the disclosure. The listing of exemplary values32323430828.1Reference No. P230046-711 -WO017240135PCTor ranges is not a disclaimer of other values or ranges between and including the upper and lower limits of a given range.
[0154] The use of headings and sections in the application is not meant to limit the disclosure; each section can apply to any aspect, embodiment, or feature of the disclosure. Only those claims which use the words "means for" are intended to be interpreted under 35 USC 112(f). Absent a recital of “means for” in the claims, such claims should not be construed under 35 USC 112. Limitations from the specification are not intended to be read into any claims, unless such limitations are expressly included in the claims.
[0155] Embodiments disclosed herein may be embodied as a system, method, or computer program product. Accordingly, embodiments may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc.), or an embodiment combining software and hardware aspects that may all generally be referred to herein as a "circuit", "module", or "system".Furthermore, embodiments may take the form of a computer program product embodied in one or more computer readable medium(s) having computer readable program code embodied thereon.33323430828.1
Claims
Reference No. P230046-711 -WO017240135PCTCLAIMSWhat is claimed is:
1. An apparatus comprising:a database comprising a plurality of sets of data, wherein each set of data comprises an identity of a known substance and tandem mass spectrometry spectrum data for known substances;a control circuit comprising a processor and a memory, wherein the memory stores instructions executable by the processor to:obtain a first portion of data from the database;input the first portion of the data into a machine learning model; train the machine learning model with the first portion of the data to form a trained machine learning model, wherein the machine learning model is trained to identify the identity of substances based on the tandem mass spectrometry spectrum data of the substance;input a second portion of data from the database into the machine learning model, wherein the second portion and the first portion comprises different sets of data from the database; andevaluate the machine learning model using the second portion of the data.
2. The apparatus of claim 1 , wherein the processor is further configured to:receive tandem mass spectrometry data for a sample; anddetermine, using the trained machine learning model, the identity of the sample.
3. The apparatus of claim 1, wherein the tandem mass spectrometry spectrum data for the known substance comprises:a precursor mass to charge ratio;a product mass to charge ratio; andan intensity.
4. The apparatus of claim 1, wherein the machine learning model comprises one of a convolutional neural network, a random forest classifier, a recurrent neural network, K-nearest neighbors, multilayer perceptron, transformer architecture, long short-term memory neural network, or weighted bi-directional feature pyramid network.
5. The apparatus of claim 1, wherein the machine learning model is configured to identify a known sample based on the tandem mass spectrometry spectrum data provided.34323430828.1Reference No. P230046-711 -WO017240135PCT6. The apparatus of claim 1 , further comprising:a mass spectrometer to collect tandem mass spectrometry spectrum data from a sample; andwherein the control circuit is to:receive the tandem mass spectrometry spectrum data from the mass spectrometer for the sample; andidentify the sample using the machine learning model and the tandem mass spectrometry spectrum data.
7. The apparatus of claim 4, wherein the control circuit is configured to:create the database by:receiving tandem mass spectrometry data for known substances.
8. The apparatus of claim 1 , wherein the instructions executable by the processor to evaluate the machine learning model comprise:determining the identity of each sample in the second portion of the data based on the tandem mass spectrometry spectrum data;determining whether the determined identity of each sample matches the identity of each sample stored in the database for the corresponding mass spectrum data; andbased on the determined identity of the sample matching the identity of the sample in the database, determining the machine learning model is accurate.
9. The apparatus of claim 1 , wherein the control circuit is further configured to:augment the database by copying the sets of data and adding random intensity and mass shifts to the tandem mass spectrometry spectrum data to create augmented data; andadd the augmented data to the database.
10. An apparatus comprising:a processor; anda memory storing a plurality of instructions and a pre-trained machine learning model executable by the processor, wherein the pre-trained machine learning model is trained using a database comprising a plurality of known substances and a plurality of tandem mass spectrometry spectrum data for each of the known substances, wherein the pre-trained35323430828.1Reference No. P230046-711 -WO017240135PCTmachine learning model is trained to classify the data with one of the substances in the database, wherein the instructions are configured to cause the processor to:receive tandem mass spectrometry data for an unknown sample; and using the pre-trained machine learning model, determine an identity of the sample, wherein the identity is one of a substance stored in a database to train the pre-trained machine learning model or a substance not used to train the machine learning model.
11. The apparatus of claim 10, wherein the tandem mass spectrometry spectrum data is received from a mass spectrometer.
12. The apparatus of claim 10, wherein tandem mass spectrometry spectrum data comprises:a precursor mass to charge ratio;a product mass to charge ratio; andan intensity.
13. The apparatus of claim 12, wherein processor is configured to:calibrate the tandem mass spectrometry spectrum data along the precursor and product mass to charge axes.
14. The apparatus of claim 10, wherein processor is configured to:normalize the tandem mass spectrometry spectrum data with a 2D Gaussian filter.
15. The apparatus of claim 10, wherein the machine learning model comprises one of a convolutional neural network, a random forest classifier, a recurrent neural network, K-nearest neighbors, multilayer perceptron, transformer architecture, long short-term memory neural network, or weighted bi-directional feature pyramid network.
16. A method comprising:obtaining data from a database, wherein the database comprises a plurality of sets of data, wherein each set of data comprises an identity of a known substances and tandem mass spectrometry spectrum data for each of the known substances;inputting a first portion of the data into a machine learning model;36323430828.1Reference No. P230046-711 -WO017240135PCTtraining the machine learning model with the first portion of the data, wherein the machine learning model is trained to identify the identity of substances based on the tandem mass spectrometry spectrum data of the substance;inputting a second portion of data from the database into the machine learning model, wherein the second portion and the first portion comprise different sets of data from the database;evaluating the machine learning model using the second portion of the data; determining the machine learning model is accurate based on the evaluation; based on the machine learning model being accurate, receiving tandem mass spectrometry data from an unknown sample;determining, using the machine learning model stored in a memory of a computing device, the identity of an unknown sample; andproviding an output indicating the identity of the sample.
17. The method of claim 16 wherein the tandem mass spectrometry spectrum data for each of the known substances comprises:a precursor mass to charge ratio;a product mass to charge ratio; andan intensity.
18. The method of claim 16, wherein the machine learning model comprises one of a convolutional neural network, a random forest classifier, a recurrent neural network, K-nearest neighbors, multilayer perceptron, transformer architecture, long short-term memory neural network, or weighted bi-directional feature pyramid network.
19. The method of claim 16, wherein the tandem mass spectrometry data is received from a mass spectrometer.
20. The method of claim 16, wherein the identity of the sample is stored within the database.37323430828.1