Anomaly detection in the environment using single-particle aerosol mass spectrometry

Single-particle aerosol mass spectrometry with autoencoders allows for real-time detection of unknown threats by generating simulated composite spectra, addressing the limitations of conventional methods and enhancing threat identification speed.

JP2026512562APending Publication Date: 2026-04-17ZETEO TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ZETEO TECH INC
Filing Date
2023-10-20
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing methods for detecting aerosolized biological and chemical hazards in real-time are limited by the need for time-consuming sample preparation and culture steps, which hinder rapid identification of unknown threats.

Method used

A method using single-particle aerosol mass spectrometry and autoencoders to analyze mass spectra, generating simulated composite spectra for unknown threats, and employing an autoencoder to detect anomalies in real-time by training on reference and background spectra.

Benefits of technology

Enables rapid, real-time detection of unknown biological and chemical threats by reducing analysis time to minutes, improving response times in biodefense and healthcare applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026512562000001_ABST
    Figure 2026512562000001_ABST
Patent Text Reader

Abstract

This method and system uses single-particle mass spectra of environmental aerosol samples to detect anomalies caused by unknown aerosol hazardous particles in the environment. An autoencoder is trained using a dataset of single-particle mass spectra to diagnose anomalies in the environment. To test the autoencoder's ability to predict known hazardous substances as anomalies, a reference mass spectrum dataset is generated using simulated composite mass spectra of hazardous analyte particles at various analyte concentrations in the background environment.
Need to check novelty before this filing date? Find Prior Art

Description

Related Applications

[0001] This patent application claims the benefit of, and is related to, U.S. Provisional Patent Application No. 63 / 417,953, filed Oct. 20, 2022, and U.S. Provisional Patent Application No. 63 / 544,859, filed Oct. 19, 2023, both entitled “Anomaly Detection in an Environment Using Single Particle Aerosol Mass Spectra,” the disclosures of which are hereby incorporated by reference in their entireties. Research and Development by Federal Government Funding

[0002] None.

Technical Field

[0003] The present disclosure relates to methods and systems for detecting anomalies caused by unknown biological-containing particles in an environment using the single particle mass spectra of environmental aerosol samples. More specifically, without limitation, the present disclosure relates to methods and systems for detecting anomalies caused by known and unknown airborne biological and chemical-containing particles by training an autoencoder to analyze the single particle mass spectra of aerosol samples in the atmosphere. To test the autoencoder for anomalies caused by known biological substances, a simulated composite mass spectrum of hazard or threat agent particles in the background environment at different concentrations of different hazardous substances (analytes) is used to generate a reference mass spectral dataset in silico, i.e., computationally, or using computer simulation.

Background Art

[0004] Threats from aerosolized biological and chemical hazards, or other hazardous substances, remain a significant concern for the U.S. government due to their potential to have serious impacts on life and property. Two primary threat or hazard scenarios of particular concern are: (1) releases of hazardous substances or hazardous substances within enclosed structures (e.g., office buildings, airports, public transportation) where hazardous substances can be effectively dispersed throughout the building by HVAC systems; and (2) widespread releases of hazardous substances or hazardous substances across entire residential areas, such as towns and cities. Exposure to released aerosolized hazardous substances can lead to numerous casualties. In the event of widespread releases, protecting residents from initial exposure is extremely difficult without timely information on the type, quantity, and location of the contaminants. Methods and equipment are needed to identify the composition of threat agents, hazardous substances, and other relevant biological and chemical substances in real time in order to take rapid corrective action. Samples of analyte aerosols in the air can be captured using appropriate means such as filters, sampling bags, or other similar containers designed to capture inhalable particles. The particles originate from liquid samples collected from wet-wall cyclones or similar devices, which may then be re-aerosolized. An example of a wet-wall cyclone is SpinCon II (Innovaprep, Drexel, Missouri). The particles in these aerosols may include, but are not limited to, anthrax bacteria, Ebola virus, ricin, and botulinum toxin. All of these collection methods require additional processing to extract biological particles for analysis, resulting in delays of several hours or even days before hazardous aerosols can be detected and identified.

[0005] While solutions already exist for detecting and analyzing aerosol analytes such as biological agents and hazardous substances, rapid and real-time analysis is not possible. One solution involves using microfluidic technology to purify the sample and concentrate the biological analyte. For example, specific antibodies can be used to concentrate and purify the biological analyte. This target-specific solution yields reasonable results if sufficient time is allocated for the purification and concentration of the analyte. Another solution is target specificity, which works only for bacterial analytes, while sacrificing the analysis of viruses, toxins, particulate chemicals, or other components of aerosols that cannot be cultured. This method requires, for example, spreading a sample from a patient onto a bacterial culture plate and culturing it for 8 to 24 hours. After the bacterial colonies have grown, the amplified and purified individual colonies are collected and measured by whole-cell matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI TOF-MS). Numerous studies have investigated the accuracy of this technique and have shown that it can identify clinical bacterial analytes with an accuracy of over 99%. Two commercially available systems for rapid clinical bacterial identification have been developed: the Bruker Biotyper (marketed by Becton Dickinson) and the Shimadzu Vitek MS (marketed by bioMe ieux).

[0006] In conventional MALDI mass spectrometry, the analyte sample is first mixed with a MALDI matrix and then placed on the probe tip in the vacuum chamber of the mass spectrometer. The MALDI matrix resonantly absorbs ultraviolet (337 nm) laser pulses, simultaneously desorbing the analyte as molecular ions into the gas phase. In aerosol MALDI mass spectrometry, the aerosolized analyte particles are preferably coated with the MALDI matrix on the fly. For example, U.S. Patent Application No. 15 / 755063, filed by the present applicant, discloses a method for coating aerosol particles, comprising the steps of: aerosolizing a coating material (MALDI matrix) to form a first aerosol containing liquid particles; preparing a sampled aerosol containing analyte particles to form a second aerosol; preparing an acoustic coater to receive the first aerosol (MALDI matrix) and the second aerosol; and applying an acoustic field to the acoustic coater to cause the first aerosol (MALDI matrix) to collide with the particles of the second aerosol on the fly, thereby forming coated aerosol particles on the second aerosol particles, including a coating of the first aerosol. Other aerosol particle coating methods are also available.

[0007] Examples of MALDI matrix chemicals include 2,5-dihydroxybenzoic acid, α-cyano-4-hydroxycinnamic acid, 3,5-dimethoxy-4-hydroxycinnamic acid, 2-mercapto-4,5-dialkylheteroarene, 1,8-dihydroxyanthracene-9(10H)-one, 3-methoxy-4-hydroxycinnamic acid, 2,4,6-trihydroxyacetophenone, 2-(4-hydroxyphenylazo)-benzoic acid, trans-3-indoleacrylic acid, 4-hydroxy-3-methoxybenzoic acid, 6-aza-2-thiothymine, 2-amino-4-methyl-5-nitropyridine, 4-nitroaniline, 1,5-diaminonaphthalene, 5-fluorosalicylic acid, 5-chlorosalicylic acid, 5-bromosalicylic acid, 5-iodosalicylic acid, 5-methylsalicylic acid, 5-aminosalicylic acid, and 1,8-diaminonaphthalene. Furthermore, examples of MALDI matrix solutions include at least one of acetonitrile, water, ethanol, methanol, propanol, acetone, chloroform, isopropyl alcohol, tetrahydrofuran, toluene, hydrochloric acid, trifluoroacetic acid, formic acid, and acetic acid.

[0008] When ions are generated in the MALDI process, these ions may be analyzed using a time-of-flight mass spectrometer (TOF-MS). In TOF-MS, data is recorded as a time series of voltages as the ions collide with the detector. The time of flight of the ions is converted to mass, as will be described later. Figure 1A shows a schematic of a linear TOF-MS, which is typically used in conventional MALDI mass spectrometer systems. Ions are generally generated in a short source region of length ("s") defined by the backing plate and extraction grid. The voltage ("V") applied to the backing plate applies an electric field ("E") across the entire source region, where E = V / s. This electric field accelerates these ions to the same kinetic energy ("U"). This kinetic energy is given by U = mv 2The formula is given by / 2 = ze(V) = ze(Es), where m is the mass of the ion, v is the velocity of the ion, e is the charge of the electron, and z is the charge number of the ion. Although the charge z is 1 for most ions, it can be higher for particularly large ions. As an ion passes through the extraction grid, the ion velocity is inversely proportional to the square root of the mass-to-charge ratio and is expressed by the following equation. v = (2zEs / (m / e)) 1 / 2

[0009] The ions pass through a much longer drift region of length ("D"), where they are separated in time, and the mass spectra at various times of flight t are produced as follows: m / z = 2eEs(t / D) 2 =kt 2

[0010] Typically, the relationship between the mass-to-charge ratio (m / z) and the time of flight t is calibrated using a compound of known mass m in a known charge state Z, the total mass k is determined, and the mass-time relationship in TOF-MS is obtained.

[0011] To obtain higher resolution, the reflectron TOF mass spectrometer mode shown in Figure 1B is used. In this configuration, a new ion optical element is introduced to form a potential "hill" that redirects ions. The reflectron is installed within the mass spectrometer and includes a series of plates 101 wired in series to form a potential gradient. The end of the reflectron closest to the source 102 is held at ground potential, and each plate is voltage-increased stepwise up to a final voltage (V+δ) slightly higher than the source voltage (V). Although ions are formed as described in the linear TOF mode, they are not detected linearly by the linear detector 103, but instead travel to the reflectron, where they climb the potential energy gradient, stop, redirect, and accelerate to the reflectron detector 104. The reflectron mode helps to significantly reduce the spread of the time of flight of ions of the same mass, which is caused by the spread of the kinetic energy of these ions at the exit from the ion source.

[0012] The reflectron configuration can be operated as a linear TOF instrument by switching it on or off. When the voltage to the reflectron element is off, ions pass through the grounded reflectron element and collide with the detector 103 located at the end of the spectrometer. In this case, the mass spectrometer can quickly switch between the high mass range and high sensitivity provided by the linear TOF mode and the high resolution provided by the reflectron TOF.

[0013] The Johns Hopkins University Applied Physics Laboratory (JHU-APL) reported an analysis of biological threats using conventional MALDI-TOF-MS with a top-down proteomics approach. Biological threats are exposed to laser pulses, generating reproducible molecular fragments. The precise masses of these fragments are measured by TOF-MS, generating intensity plots as a function of mass, and obtaining mass spectra. As an example of a whole-cell approach, i.e., a top-down approach, Figure 2 shows the mass spectrum of Bacillus globii (Bg) in spore form. The current nomenclature for this microorganism is Bacillus atropaeus, and many strains and related substances exist from laboratories worldwide. Historically, Bg has been a CDC Tier 1 listed substance and has been used as a non-infectious mimic of Bacillus anthracis, which is generally considered the most likely substance favored by bioterrorists or a biologically harmful substance.

[0014] Samples of Bacillus atrophaeus were mixed with MALDI matrix in a solvent, coated onto a surface, dried, and then inserted into a TOF-MS for analysis. The characteristic spectra of the obtained Bacillus atrophaeus spores are shown in Figure 2. The collective features observed in the spectrum represent the characteristics of Bacillus atrophaeus and the genetic structure of this threatening organism; they are reproducible and unique to Bacillus atrophaeus. These features are unaffected by growth conditions, culture media, preparation methods, or environmental pollutants. The numbers annotated in Figure 2 represent the m / z (Dalton) values ​​of the peaks, determined by the time the ion packet reached the detector, and were subsequently converted to mass by calibrating the instrument with a substance of known mass, as previously described. This spectrum also annotates the molecular identity of many major peaks, based on detailed mass spectrometry information using known gene sequences of the microorganism, and previously reported information regarding the presence and role of small acid-soluble proteins (SASPs) in the spore membrane of the Bacillus genus. Differences in the genetic structure of members of different Bacillus species result in differences in the mass of SASP peaks, providing a recognition mechanism by MALDI mass spectrometry. It has been reported that treating spores with strong acid allows for the rapid extraction of small acid-soluble proteins that serve as characteristic biomarkers for identifying Bacillus spores.

[0015] These systems offer superior diagnostic results compared to the 16s RNA "gold standard." However, obtaining these reliable clinical results requires culture and / or extraction steps to purify the sample. Therefore, the time from sample collection to bioanalyte identification typically ranges from 12 hours to over a day. While such delays are often acceptable in clinical laboratories, they are often unacceptable in other applications such as biodefense, where real-time identification of bioanalytes is required. Beyond biodefense, point-of-care healthcare applications demand the ability to simultaneously identify not only bacteria but also fungi, viruses, and large bioorganic molecules (proteins, peptides, lipids, etc.), including biotoxins, in real time. Furthermore, reducing analysis time for clinical applications could lead to more timely treatment and identification of optimal therapies (e.g., differentiating between viral and bacterial infections) and evaluation of treatment effectiveness, potentially improving the quality and outcomes of care.

[0016] Regarding aerosol analysis, an example of aerosol TOF-MS is disclosed in the applicant's application, International Application No. PCT / US20 / 40023, which is incorporated herein by reference. In an exemplary system 300 (Figure 3), aerosol particles, such as particles containing biological material in the air, are guided to a suitable inlet element 301 where debris and material are removed from the particles. This debris and material is removed from the particles at a rate of approximately 1000 units per second and flows into an aerosol beam generator 302, which collimates the particles into a narrow beam of single particles. Before entering the beam generator 302, the aerosol is sent to a MALDI matrix processing subsystem 313, where the particles are coated with a MALDI matrix and processed for analysis. The particles then pass through the aerosol beam generator, where differential pumping is used to reduce the pressure from atmospheric pressure to the mass spectrometer base pressure (approximately 10⁻⁵ to 10⁻⁶ Torr). The beam generator uses differential pumping to reduce the pressure to a level suitable for the high vacuum in chamber 304. The particles are indexed using a continuous laser from laser generator 303 (including, but not limited to, commercially available laser scattering devices such as IBAC and Polaron systems). Furthermore, the continuous laser is used to measure particle size, fluorescence (autofluorescence), and polarization (particle shape) to identify particles of particular interest. The particles then travel through a series of focusing lenses into vacuum chamber 304, which houses an advanced time-of-flight mass spectrometer (TOF-MS) 306 and, optionally, an optical collection component 307.

[0017] The IBAC system measures aerosol particle data using UV laser-induced fluorescence. Laser-induced fluorescence (LIF) is achieved by exciting particles or samples with a laser (a continuous laser as described above can be used), and the emitted fluorescence is detected and analyzed using a suitable photodetector, including a photomultiplier tube (PMT) detector, allowing for the distinction between biological and non-biological particles. It is well known in the art to use fluorescence to determine the concentration and size distribution of biological aerosol particles measured by ultraviolet aerodynamic particle sizing devices. Polaron classifies aerosol particles based on their shape and size using polarized elastic light scattering generated from laser-excited particles (a continuous laser as described above can be used).

[0018] As each indexed particle enters the center of the chamber 304, it is irradiated with a high-power laser pulse from the laser generator 308. In aerosol mass spectrometry, the ionization laser 308 must be irradiated when the aerosol particle enters the laser irradiation area (typically less than 150 microns in diameter). Because the pulsed ionization laser 308 emits pulses of less than 5 nanoseconds (nanoseconds), advanced knowledge is required to predict when the particle will enter the ionization area and activate the laser 308. Multiple lasers may be used to measure and track the particle to predict when it will enter the laser's field of view. In the exemplary system, at least one of the lasers from generator 303 and lasers from laser generator 312 may be used to index and detect the particle exiting beam generator 302. Since both laser beams 308 and 312 are closely aligned, a single trigger laser 312 is sufficient to predict the path of a single aerosol particle and trigger the pulsed ionization laser 308, significantly reducing the complexity of the particle timing hardware.

[0019] The pulsed ionization laser 308 may be triggered using a laser from the generator 303. Laser 308 may be triggered only when at least one of the particle size, shape, and fluorescence meets or exceeds a predetermined threshold for its properties. When periodically monitoring the composition of aerosol particles in the atmosphere, selectively triggering laser 308 in this way, and then examining the ionized fragments of each particle and analyzing the collected data, can be used to avoid collecting unnecessary data and improve data management. The timing (or trigger) laser 312 may also be used to measure the optical properties of the particles (e.g., size, shape, fluorescence). These measurements can be used to select particles to be ionized, and the data can be combined with mass spectral measurements and other optical information obtained during ionization and analyzed in the data analysis system 310 using a data fusion method. The intensity of the laser pulse from the generator 308 may be adjusted so that the particles are decomposed and ions are generated from the biochemical components that make them up. In other words, the laser vaporizes and ionizes at least a portion of the molecules being analyzed, generating ions with a specific mass-to-charge ratio (m / z). These large, beneficial ions are accelerated to TOF-MS306, where they are analyzed.

[0020] Furthermore, when the particles under analysis absorb sufficient light energy from the laser beam, they emit characteristic photons as they transition from a high-energy state to a low-energy state. This emission may also be related to transitions between vibrational states. Interaction between the high-power laser pulses generated by the generator 308 and the particles may also induce transient optical properties such as higher-order fluorescence, laser-induced breakdown spectroscopy (LIBS), Raman spectra, and infrared spectra. The chamber 304 may also include a focusing optical component 307. Particle-specific spectral data generated using the TOF-MS and optical sensor 309, as well as particle-specific data from the laser devices 303 and 312 (e.g., particle size, shape, fluorescence), can be subjected to data processing, including data fusion, in the data analysis system 310 to generate compiled spectral data related to each particle. The compiled spectral data may be compared to a training dataset containing a knowledge base of known biomolecular spectra to predict composition. The system 310 may communicate with a machine learning engine 311 to enable updating of the knowledge base-based training dataset, thereby improving composition prediction over time. The pressure inside chamber 304 is reduced to at least 10⁻⁵ torr using vacuum pump 305. In the exemplary system 300, the travel time (or residence time) of particles from beam generator 302 to laser 308 is less than 1 second.

[0021] The generation of aerosol single-particle MALDI mass spectrometry signatures is fundamentally different from signatures obtained by conventional MALDI mass spectrometry or other mass spectrometry methods that examine solid bulk samples. Conventional MALDI mass spectrometry extracts ions from bulk samples that typically contain thousands of particles. These particles include pathogens such as bacteria and viruses, pathogen components such as proteins, peptides, and lipids, reagents (e.g., MALDI matrix), and contaminants (e.g., environmental substances, sputum from breath or cough samples, and other human-related substances and by-products). These target particles are dispersed throughout the sample and mixed with other particles such as environmental contaminants. The spatial distribution of particles in the bulk sample causes a distribution in the distance and time of flight (to the detector) of ions generated from the sample when the sample is irradiated with an ionization laser. Ion diffusion can be reduced to some extent by designing the ion source region, for example, by using techniques such as delayed extraction or two-stage extraction.

[0022] Furthermore, to reduce noise and improve the signal-to-noise ratio, multiple laser shots are typically performed, and the spectra from each shot are averaged. While averaging reduces the noise associated with the spectrometer, thus improving the signal-to-noise ratio of the peaks, it does not reduce variability due to sample heterogeneity. Attempts to denoise and align feature peaks using individual measurements and average spectra have shown that using average spectra yields better results. This is because each bulk sample contains particles of varying compositions, and a single measurement from these particles generates ions associated with these different particles. Therefore, in conventional mass spectrometry of bulk samples, deconvoluting the mass signal components associated with individual particles is not advantageous. Similar challenges are observed when deconvolving the signal components associated with individual particles from aerosolized samples.

[0023] U.S. Patent Application No. 17 / 507,755 by the same applicant discloses a system for identifying the composition of aerosol particles, which is incorporated herein by reference in its entirety. The disclosed system includes an aerosol beam generator that produces a beam of single particles; a continuous timing laser generator that produces a timing laser for indexing each particle in the beam; a pulsed ionization laser generator that is triggered by the timing laser and configured to collide with each index particle when it reaches the ionization region of the ionization laser, producing an ionized fragment of each index particle and at least one photon associated with each index particle; a guide tube having an exit end and positioned between the aerosol beam generator and the ionization region, which guides the particles toward the vicinity of the longitudinal axis of the guide tube; and at least one detector that analyzes the ionized fragment and at least one photon associated with each particle and produces unique spectral data associated with each index particle. The at least one detector may include at least one of a TOF-MS detector, a fluorescence detector, a LIBS detector, and a Raman spectrometer. The ionized fragments of each index particle and the photons associated with each index particle may be analyzed using a TOF-MS detector to determine the composition of each particle. In some implementations, the pulsed ionization laser may be triggered by a continuous timing laser when each particle is incident on the continuous laser beam, generating an IR laser pulse or a UV laser pulse. The continuous timing laser generator and the pulsed ionization laser generator may be configured to generate the continuous laser beam and the pulsed ionization laser beam as overlapping beams, respectively.

[0024] The selection of which indexed particle to analyze may be performed by triggering an ionization laser step when at least one characteristic of the indexed particle satisfies a predetermined threshold for that characteristic. The composition of the particle to be analyzed can be determined by generating multiple single-particle spectra using a TOF-MS detector, aligning each single-particle spectrum, denoising each aligned single-particle spectrum, averaging the aligned and denoised single-particle spectra, and comparing the average spectrum to a reference spectrum. Alignment of single-particle spectra may include selecting one or more mass ranges based on prior information about the location of the target mass ranges, selecting one spectrum as a reference spectrum for each mass range (where the reference spectrum includes at least one of a pre-selected spectrum, a spectrum present in a reference data library, and a spectrum generated using a measured single-particle spectrum dataset), and shifting the peak window of the spectrum dataset to align with the corresponding window of the reference spectrum in the time domain.

[0025] Furthermore, the step of selecting a single spectrum as a reference spectrum generated using the measurement dataset may include the steps of selecting multiple measured single-particle spectra, calculating the Pearson correlation coefficient ("PCC") for each spectrum data file by cross-correlation with each other spectrum in the dataset and recording the average PCC score of the files, and selecting the spectrum with the highest PCC score as the reference spectrum. The aligned single-particle spectra may be denoised using single-value decomposition ("SVD"). The step of determining the composition may further include at least one of the following steps: predicting the composition by comparing the average spectrum data with a training spectrum dataset knowledge base, updating the training dataset knowledge base, and improving the composition prediction over time using a machine learning method. The machine learning method may include supervised machine learning methods.

[0026] In the example of the aerosol TOF-MS system disclosed above, "indexing each particle" means "attaching a timestamp to each particle". To generate a timing laser for indexing each particle in the beam, the aerosol TOF-MS system can be configured to assign an index or timestamp to each particle using a timing circuit and a counter, as described, for example, in U.S. Patent No. 5,681,752 and U.S. Patent Publication No. 2011 / 0071764. These disclosures are hereby incorporated by reference in their entireties. A photomultiplier tube can be used to detect scattered light by particles excited by a laser beam. Laser 308 can include one or more UV laser pulses or IR laser pulses that are triggered when a selected index particle reaches the ionization region. A single trigger laser (303 or 312) can be used to trigger the ionization laser 308. Laser 308 is triggered only when at least one of the particle size, shape, and fluorescence meets or exceeds a predetermined threshold for that property. Indexing not only enables the fusion of mass spectrometry data but also enables the data fusion of the optical properties of each particle based on the spectral data associated with each particle.

[0027] Furthermore, in the exemplary aerosol TOF-MS system disclosed above, the continuous timing laser generator may be configured to perform additional pre-ionization tasks or functions (in addition to triggering), including indexing each particle in the beam, particle size of each indexed particle, particle shape by polarization, and optical property evaluation of fluorescence, and selection of indexed particles to be ionized. The exemplary TOF-MS system may include a data system that generates a data set combining optical data and unique mass spectral data, combined unique spectral data associated with each indexed particle, a data system that processes particle size, particle shape by polarization, and fluorescence using a data fusion method and generates compiled spectral data associated with the selected indexed particles, and a data system that predicts the composition of bioaerosol particles by comparing with a training data set including a knowledge base of known biological substance spectra.

[0028] The supervised learning algorithm described above requires a library of threat signatures and labeled data to learn to classify spectra into drugs or hazardous substances. This approach is effective for known threats, but there is a need for methods and systems to detect anomalies and alert to the possibility of unknown drugs or next-generation threat agents or hazardous substances. In particular, when it is impossible to release biological substances into the environment, there is also a need for methods and systems to generate a data set of mass spectral data by mixing (silico mix) the mass spectra of known threats and background spectra in a computer. By applying the silico method, measurements of biological substances under strictly controlled conditions (e.g., in an aerosol chamber) can be combined with environmental measurements to simulate a given concentration. SUMMARY OF THE INVENTION

[0029] In some implementations, exemplary methods for predicting anomalies in the background environment may include the steps of: compiling a dataset of reference analyte (hazardous substance) mass spectra of a single aerosolized analyte particle representative of hazardous particles present in the environment; compiling a dataset of reference background mass spectra of a single aerosolized background environment particle that does not contain the analyte hazardous substance particle; using the reference analyte hazardous substance mass spectra and the reference background mass spectra, computer-generating a dataset of simulated composite mass spectra of analyte particles in the background environment at different analyte concentrations in the background environment; identifying an optimal batch size for the composite spectra at each analyte hazardous substance concentration (where the batch size represents the number of spectra to be averaged); and determining a threshold error for predicting anomalies in single-particle mass spectra generated from new aerosol samples taken from the environment.

[0030] In some implementations, composite mass spectra at different concentrations of analyte hazardous substances may be generated in silico by varying the dilution ratio, which is defined as the number of reference analyte hazardous substance spectra divided by the sum of the number of reference analyte hazardous substance spectra and reference background spectra. This exemplary method may involve diluting individual analyte hazardous substance particles with individual background particles. In some implementations, the background environment may be ambient air. Analyte hazardous substance particles may contain at least one of chemical and biological substances. Reference analyte hazardous substance spectra and background spectra may be generated using an aerosol MALDI TOF-MS system.

[0031] In some implementations, the step of determining the optimal batch size in each analyte hazardous substance concentration step may include training an autoencoder to learn the features of a reference background spectrum, generating a receiver operating characteristic curve ("ROC") based on the probability distribution of autoencoder reconstruction losses associated with individual hazardous substance spectra and background spectra, and determining the optimal batch size as the batch size that maximizes the area under the ROC curve ("AUC").

[0032] In some implementations, exemplary methods may further include a step of discarding mass spectral features below a predetermined m / z cutoff value before the training step. The predetermined m / z cutoff value may be less than approximately 3000. The step of determining the threshold error may include selecting the optimal reconstruction loss threshold from the ROC curve by maximizing Youden's J statistic. Although Youden's J statistic gives equal weight to sensitivity and specificity, this is possible when the spectrum of the hazardous material is available. Also, sensitivity and specificity may have different importance depending on the use case. For defense applications, high specificity (low false alarms) may be preferred even if sensitivity is compromised. In some implementations, the optimal batch size may be 1. In some implementations, particles may be tagged as either anomalous or non-anomalous. Anomalous particles may be subject to further analysis. Non-anomalous particles are discarded.

[0033] In some implementations, an example of a method for diagnosing anomalies in the environment may include the steps of: generating single-particle mass spectra from aerosol samples taken from the environment continuously or at predetermined intervals; inputting the single-particle mass spectra into an autoencoder trained to compress and reconstruct mass spectra based on a baseline reference mass spectrum of aerosol samples taken from the environment; examining the single-particle mass spectra in a predetermined optimal batch size; determining a reconstruction loss threshold associated with the baseline reference mass spectrum of aerosol samples taken from the environment; and diagnosing anomalies in the environment, where anomalies are detected at the single-particle level without performing batch averaging of mass spectra if the reconstruction loss of the autoencoder associated with the aerosol sample exceeds the reconstruction loss threshold. The optimal batch size may be 1. Single-particle mass spectra from aerosol samples taken from the environment may be generated using an aerosol MALDI TOF-MS system.

[0034] In some implementations, exemplary methods for predicting background environmental anomalies caused by one or more unknown aerosol analyte particles may include the steps of: compiling a dataset of reference background mass spectra of a single background particle; training an autoencoder to learn features of the reference background mass spectra in various batch sizes (where the batch size represents the number of mass spectra to be averaged); determining a background mass spectrum reconstruction loss threshold in which at least 90% of background mass spectra are discarded as non-anomalous; and predicting one or more anomalies in the single particle mass spectrum generated from an aerosol sample extracted from a background environment sample (sample spectrum) if the reconstruction loss associated with the sample spectrum exceeds the reconstruction loss threshold.

[0035] In some implementations, exemplary methods may further include a step of discarding mass spectral features below a predetermined m / z cutoff value before the training step. The predetermined m / z cutoff value may be less than approximately 3000. The mass spectra may be generated using an aerosol MALDI TOF-MS system. The autoencoder may be trained over a predetermined number of epochs (the number of times the learning algorithm is run over the entire training dataset) on at least approximately 20,000 single-particle spectra collected from environmental samples. The predetermined number of epochs may be 20 epochs, or the greater of the number of epochs in which the percentage reduction of the mean training loss is minimized to numerical precision, as used, for example, in PyTorch or in accordance with the IEEE 754 standard. Exemplary methods may further include a step of examining the averaged and processed spectra of the sample spectra to determine the presence of one or more features characteristic of an unknown analyte. Exemplary methods may further include a step of examining a heatmap of the processed sample spectra to determine the presence of one or more statistically significant features characteristic of an unknown analyte.

[0036] Other features and advantages of this disclosure are described in part in the following description and accompanying drawings, and the favorable aspects of this disclosure will become apparent to those skilled in the art through examination of the following detailed description, which is described and shown here and partially grasped in conjunction with the accompanying drawings, and will also be learned through the practice of this disclosure. The advantages of this disclosure may be realized and achieved by means and combinations specifically indicated in the accompanying claims. [Brief explanation of the drawing]

[0037] The aforementioned aspects of this disclosure and its many associated benefits will be more readily apparent, as they will be better understood by referring to the following detailed description in conjunction with the attached drawings. [Figure 1A-B]Figures 1A and 1B show (A) a linear TOF-MS and (B) a TOF-MS configuration combining linear and reflective TOF-MS. [Figure 2] Figure 2 shows characteristic whole-cell (top-down) MALDI TOF-MS spectra of Bacillus atropaeus (Bg) spores obtained by high-resolution TOF-MS following several implementation examples. [Figure 3] Figure 3 shows a schematic diagram of an exemplary system for single-particle aerosol analysis following several implementation examples. [Figure 4A-C] Figures 4A-C show (A) the raw Bg spectrum, (B) the background spectrum, and (C) a heatmap of a simulated mixture, with 10% randomly selected from the Bg dataset and 90% from the background dataset, following several implementation examples. [Figure 5A-B] Figures 5A-B show a comparison of simulated distributions and binomial distributions, illustrating the probability distributions of the number of simulated spectra at batch sizes 1 (A) and 5 (B), following several implementation examples. [Figure 6A] Figure 6A shows a schematic diagram of a conventional TOF-MS signal processing pipeline. [Figure 6B] Figure 6B shows a schematic diagram of a single-particle aerosol MALDI TOF-MS signal processing pipeline following several implementation examples. [Figure 7A-C] Figures 7A–C illustrate the effect of a signal preprocessing step in an exemplary signal processing pipeline following several implementation examples. [Figure 8A] Figure 8A shows the confusion matrix, which demonstrates the KNN classifier's ability to detect and identify live pathogens against noise. [Figure 8B] Figure 8B shows the standard deviation, which is high for missed or misidentified samples, following several implementation examples. [Figure 9] Figure 9 shows schematic diagrams of autoencoder network architectures following several implementation examples. [Figure 10]Figure 10 shows a schematic diagram of an exemplary method for selecting the optimal reconstruction loss threshold for a given threat substance or analyte concentration, following several implementation examples. [Figure 11A-C] Figures 11A–11C show the probability distributions of the reconstruction loss of an exemplary autoencoder in background versus simulated Bg spectra for (A) batch size 1, (B) batch size 5, and (C) batch size 10, following several implementation examples. [Figure 12A-C] Figures 12A–12C show receiver operating characteristic ("ROC") curves based on the probability distribution of the autoencoder's reconstruction loss for background versus simulated Bg spectrum, following several implementation examples, and the area under the ROC curve ("AUC") values ​​for each ROC curve at batch size 1, batch size 5, and batch size 10.

[0038] All reference numbers, identifiers, and callouts in the figure are incorporated herein by this reference as if they were fully described here. The absence of numbering in the figure elements is not intended as a waiver of any rights. Unnumbered references may also be identified by letters in the figure or appendices.

[0039] The following detailed description includes references to the accompanying drawings, which form part of the detailed description. The drawings illustrate specific embodiments in which the disclosed systems and methods may be carried out. These embodiments, which should be understood as “examples” or “options,” are described in sufficient detail to enable those skilled in the art to carry out the invention. Embodiments can be combined, other embodiments can be utilized, or structural or logical modifications can be made without departing from the scope of the invention. Therefore, the following detailed description should not be constrained, and the scope of the invention is defined by the accompanying claims and their legal equivalents.

[0040] In this disclosure, "aerosol" generally means a suspension of particles dispersed in air or gas. "Real-time" analysis of aerosols generally means analytical methods and apparatus that identify the aerosol analyte within minutes of the aerosol sample being introduced into the analytical instrument or system. The terms "a" or "an" are used to include one or more, and the term "or" is used to refer to a non-exclusive "or" unless otherwise specified. Furthermore, terms used herein, unless otherwise specified, are for illustrative purposes only and not limiting. Unless otherwise specified in this disclosure, to interpret the scope of the term "approximately," the error range related to disclosed values ​​(such as dimensions or operating conditions) is ±10% of the value shown in this disclosure. The error range related to values ​​disclosed as percentages is ±1% of the percentage shown. The word "substantially" used before certain words may mean "to a substantial extent with respect to the specified" or "mostly but not all of the specified." Detailed explanation

[0041] Specific aspects of this invention are described in detail below to illustrate the configuration, principles, and operation of the disclosed methods and systems. However, various modifications are possible, and the scope of this invention is not limited to the exemplary aspects described.

[0042] For example, to gain a deeper understanding of the detection capabilities of threat agents or hazardous substances using atmospheric or aerosol TOF-MS, it is important to evaluate both true and false positive rates over a wide concentration range. Since systematically preparing aerosol concentrations containing substances is difficult, one example of anomaly detection methods involves diluting individual dummy substance (threat agent or analyte hazardous substance) particles in a predetermined ratio with individual background particles in silico. This simulation takes a background spectrum dataset, a dummy substance spectrum dataset, a desired dilution ratio (percent), and an output dataset size, and randomly samples from the two datasets accordingly. As illustrated in Example 1 below (Figure 4A-C), the heatmap shows that even a 10-fold dilution can significantly reduce the human eye's ability to detect threat agents or analyte hazardous substances. While averaging spectra in batch processing (compared to single-particle spectra or single averaging) improves signal-to-noise separation, the optimal batch size depends on the hazardous substance concentration of the drug or analyte. When the dilution level is sufficiently low, averaging in batch processing may not be desirable.

[0043] For example, if only one of five spectra is a threat spectrum and four are background spectra, averaging all five will only decrease the signal-to-noise ratio. However, if four of the five spectra contain the mass spectral characteristics of the threat substance or analyte, averaging all five will improve the signal-to-noise ratio. In atmospheric analysis, it is important to test different dilution rates because the threat concentrations in the air vary. The dilution rate is defined as m / (n+m), where n is the normal case (normal background) and m is the abnormal case. The example of a data curation method for environmental monitoring of biological weapons threats in single-particle mass spectrometry disclosed herein computationally dilutes individual simulated particles with individual background particles at a specified ratio. This example of the method avoids the need to systematically adjust drug-containing aerosol concentrations at different dilution rates in the laboratory. The simulation takes a dataset of background spectra, a dataset of simulated (threat substance or analyte) spectra, a desired dilution rate (percent), and the size of the output dataset, and randomly samples from the two datasets accordingly. An example of this method is that it allows for the investigation of a defined, limited number of threat particle instances in any given background over a period of time.

[0044] The probability P(x) of "extracting" either a spectrum containing a threat or a background spectrum (a specific outcome) from a random sample spectrum of the environment in n trials can be represented by a binomial distribution.

number

[0045] Figures 5A-B show a comparison of the simulation distribution and the binomial distribution, illustrating the probability distribution of the number of simulated spectra at batch sizes 1(A) and 5(B) according to several implementation examples. It can be seen that the simulation distribution and the binomial distribution agree very well at batch sizes 1 and 5, with p=10%. Therefore, the anomaly detection method disclosed herein provides a data curation process that is, in principle, applicable to any detection system, and is particularly useful in aerosol applications such as aerosol TOF-MS where the concentration of substances or analyte hazardous substances in the air is not clearly defined.

[0046] In conventional MALDI-TOF-MS analysis, a bulk sample is deposited on a plate with a matrix, placed under vacuum in a mass spectrometer, and desorbed and ionized by a laser. Figure 6A shows a schematic diagram of a conventional TOF-MS signal processing pipeline. The laser is repeatedly irradiated (typically several hundred times) while randomly scanning the beam position along the spatial distribution of the sample. In step 601, a mass spectrum is generated by each laser shot, and the recorded spectra are averaged all at once in step 602, after which further signal preprocessing (steps 603-606) and identification are performed. The process in step 603, which converts the time of flight of charged particles to mass, has already been described. Decimation in step 604 is a process that reduces the sampling rate. In practice, this usually means low-pass filtering of the signal, which helps reduce computational requirements such as memory and processing speed by reducing the number of data points in the spectrum.

[0047] Figures 7A–C illustrate the effects of signal preprocessing steps in an exemplary signal processing pipeline following several implementation examples. Figures 7A–C show the impact of baseline subtraction (and filtering and smoothing) and normalization (steps 605–606) preprocessing steps on signal quality. Baseline drift (Figure 7A), which generates very high signal levels at low mass and decreases exponentially with increasing m / z, is caused by ions from the matrix binding to other fragments of the sample. A median filter can be used to estimate and then remove this baseline. A common noise reduction technique, the median filter is computationally efficient, effective against non-Gaussian distributed noise, and has minimal impact on information including peaks associated with the hazardous substance signature of a drug or analyte (Figure 7B). The normalization step 606 allows for the estimation of the signal-to-noise ratio (SNR) at each m / z value. This normalization process significantly improves the SNR of the signature, especially at high mass (Figure 7C).

[0048] As m / z increases, noise fluctuations decrease significantly, and small fluctuations in the signal become statistically significant. Although this exemplary method shown in Figure 6A works well in a laboratory equipped with high-performance scientific instruments and offline sample preparation techniques (e.g., centrifugation, lyophilization, culture, etc.), indiscriminately averaging the spectra generated by all laser shots can introduce unwanted chemical noise that masks weak signals. Detecting weak signals requires more detailed spectral information and information-based signal processing. The identification step 607 may include feature extraction and identification. This process is optimized to reduce broader transient noise characteristics while preserving discrete spectral peaks.

[0049] Figure 6B shows a schematic diagram of a single-particle aerosol MALDI TOF-MS signal processing pipeline following several implementation examples. In single-particle aerosol MALDI TOF-MS, the mass spectra of individual suspended particles are acquired in step 608 and processed using steps 603-606 as described above. In a noisy environment such as ambient air, threatening substances and analyte hazards are highly diluted, and the majority of these single particles become features of the background spectrum. Therefore, preliminary pre-processing of each particle spectrum generated in step 608 is performed in steps 603-606, and in step 609, it is determined whether the particle spectrum is anomalous (the spectrum has features that do not represent the background spectrum) or not. Spectra determined to be anomalous are flagged for further analysis, and the rest are discarded in step 610.

[0050] In steps 603' to 606', raw anomalous spectra may be batch-processed and identified. The process shown in Figure 6B differs from the conventional MALDI processing scheme shown in Figure 6A in that the anomalous detector in step 609 is applied as a filter to individual particle spectra. In addition to improving sensitivity to known threats, this approach can also warn of the presence of unknown threat agents or analyte hazardous substances, as the anomalous detection step 609 looks only for spectra with features that do not represent the background. Even after applying the TOF-MS spectrum preprocessing methods described above, MALDI-TOF mass spectra are inherently complex data representations of biological and chemical markers present in the sample, requiring advanced expertise for proper interpretation by humans.

[0051] Typically, spectra are classified using supervised machine learning algorithms trained on a well-documented library of drug or analyte signatures; however, this approach ignores the identification of unknown signatures. Preparing for unknown next-generation threats is crucial, necessitating the use of unsupervised anomaly detection models. Therefore, the 609 anomaly detector could be used for single-particle data analysis, for example, if particles were collected from the surrounding air of a federal office building, airport, or other critical threat area, to filter out single-particle mass spectra that are unlikely to contain biological drug signatures from a mass spectrum of aerosol particles. Human analysis is impractically resource-intensive and prone to bias, as it could potentially process thousands of spectra per second.

[0052] Machine learning may be used to rapidly and accurately classify large and diverse datasets. In machine learning, mathematical models may be trained using datasets and then used to make predictions based on new data. The complexity of these models ranges from simple statistical regression to artificial neural networks. Data processing steps, such as feature selection and identification from mass spectra, may be implemented using machine learning algorithms. In machine learning techniques, input mass spectral features must be input into the model. Feature extraction from MALDI mass spectral data associates a set of peaks with a specific biological agent or analyte. Since the peak locations are known in advance, this knowledge can be used to significantly reduce the dimensionality of the mass spectral data (reducing the number of important mass spectral features). Machine learning methods may provide information about features used for classification, unlike features discovered by unsupervised clustering methods such as principal component analysis (PCA). Background samples (e.g., air without biological or chemical threats or analyte) are assumed to have no features on the mass spectrum. After extracting prominent features from a wide range of biological and chemical threats or analyte hazardous substances, each sample may be classified as either a background sample or a threat (analyte hazardous substance) using the k-nearest neighbors (KNN) algorithm.

[0053] The KNN algorithm accepts an unlabeled test vector and classifies it by assigning the most frequent label among the n closest training samples to that test vector. Here, the degree of approximation refers to the distance between vectors in a multidimensional feature space. Figures 8A-B show (A) the confusion matrix, which indicates the KNN classifier's ability to detect and identify viable bacteria against noise, and (B) the standard deviation, which is high for missed or misidentified samples, in several implementation examples. The KNN classifier described in Example 2 and Figures 8A-B below is attractive for mass spectral data analysis because it makes no assumptions about the distribution of the data, which is particularly advantageous when the dataset is small (e.g., less than about 1000, but the dataset size can vary). When the number of features is also small (e.g., less than about 50, but can vary), it is easy to add new training data without requiring large processing resources. Because the number of features is limited, meaningful information contained in less prominent peaks is likely to be excluded. However, because the dimensionality is significantly reduced, the likelihood of overfitting, a common problem when a model is only useful as a reference to the training dataset and does not generalize to other data, is reduced.

[0054] Exemplary methods for generalized anomaly detection may include (a) estimating distribution parameters such as the mean and variance of specific features within a peak window or across the entire spectrum, and (b) extracting background features using deep learning such as an autoencoder. The distribution parameters may be computed for each spectrum and relate to quantities characterized by some difference between the threat spectrum and the background spectrum. For example, the variance of the spectrum may be higher in the presence of biological / chemical particles. Figure 9 is a schematic diagram of an autoencoder network architecture following several implementation examples. The autoencoder 900 is an unsupervised deep learning method consisting of two symmetric neural networks, namely an encoder 901 and a decoder 902. These neural networks may consist of linear or convolutional layers. The encoder compresses the input data vector into a low-dimensional space ("code" 903); that is, it compresses the input data to only the most important features (including significant intensities in the mass spectrum). The decoder 902 may reconstruct the input vector from the compressed data representation using a similar structure, in which the autoencoder learns the most efficient representation of the input data. This is because the features in the input data are not predefined as important or unimportant features.

[0055] Furthermore, in the input data, spectra are not classified as background spectra or anomalous threat spectra. The input vector 904 may include intensity vectors of mass spectra with a lower limit m / z cutoff of 3,000 to remove chemical noise associated with the MALDI matrix. The exemplary autoencoder 900 may be trained over approximately 20 epochs (the number of times the learning algorithm processes the entire training dataset), such as collecting approximately 20,000 single-particle spectra in the atmosphere and smoothing out noise spikes by, for example, averaging them five times. The number of epochs is determined by considering whether there is any gain from additional training. If the improvement in the model is slight, the training process may be stopped.

[0056] In the autoencoder 900, the model architecture is reduced from 1909 features to 9 features, consisting of a linear encoder and decoder, with each linear layer followed by a ReLU (Rectified Linear Unit) activation function. The final step of the decoder is a sigmoid activation function that outputs values ​​between 0 and 1. The network can be built using the open-source machine learning library PyTorch. The encoder starts with a linear layer that reduces from 1909 features to 128 features, followed by a ReLU activation function, another linear layer that reduces from 128 to 64 features, yet another activation function, and yet another linear layer that reduces from 64 to 36 features, and so on, until the final dimension is 9. The decoder is then scaled up to 1909 features.

[0057] During the learning step, the model learns to compress examples of the input vector from 1909 features to 9 features, and to reconstruct the output vector to be as similar as possible to the input vector while minimizing the reconstruction error. As the test vector is compressed and expanded, the error between the input and output indicates how similar the test vector is to the training data. The autoencoder typically reproduces normal trace inputs with minimal error, while anomalous trace inputs show a large difference between the input and output traces. The autoencoder attempts to minimize a loss function, which can be defined as the mean squared error ("squared L2 norm") between the target value and the estimated value. Thus, the reconstruction loss acts as a thresholdable anomaly score to filter out normal signals and flag anomalous signals, potentially attributable to threatening or analyte substances, for further analysis.

[0058] Figure 10 shows a schematic diagram of an exemplary method 1000 for selecting an optimal reconstruction loss threshold for a given threat agent or analyte concentration, following several implementation examples. The exemplary method may include the following steps: (a) In step 1001, a receiver operating characteristic curve (ROC) is generated based on the probability distribution of the autoencoder's reconstruction loss in the background spectrum and the simulated (threat) spectrum. For example, Figures 11A-C show the probability distribution of the autoencoder's reconstruction loss in the background spectrum and the simulated Bg spectrum for batch size 1, (B) batch size 5, and (C) batch size 10, based on several implementation examples. (b) In step 1002, the batch sizes at the same analyte concentration in the threat agent or aerosol are compared, and the optimal batch size that maximizes the area under the ROC curve (AUC) is selected. In Figure 10, the optimal batch size is 1. For example, Figures 12A-C show receiver operating characteristic (ROC) curves based on the probability distribution of autoencoder reconstruction loss in the background spectrum and simulated Bg spectrum, and the area under the ROC curve (AUC) values ​​for each ROC curve at (A) batch size 1, (B) batch size 5, and (C) batch size 10, according to several embodiments. (c) In step 1003, the optimal reconstruction loss threshold is selected from the ROC curve by maximizing Youden's J statistic, defined as (sensitivity + specificity - 1). This parameter gives equal weight to sensitivity and specificity. Specificity may be defined as (TN / TN + FP) and sensitivity as (TP / TP + FN), where TP, FP, TN, and FN represent true positive, false positive, true negative, and false negative results, respectively.

[0059] Finally, the model may be tested by compressing and expanding the test data vector and examining the error between the input and output. This error indicates how similar the test data vector is to the training data. Thus, the autoencoder may be trained on environmental background data to narrow down the spectrum to a few key features and determine whether the test spectrum deviates from the background spectrum, rather than reducing the size of the spectral dataset. Since spectra that the autoencoder determines to be similar to the background spectrum are discarded, the number of spectra to be further analyzed is reduced (as described above in steps 609-610, see Figure 6B). Based on the anomaly detection capabilities of the autoencoder, the exemplary method shows that for very sparse signals, such as signals from threat agents in the air or hazardous substances under analysis, it is optimal to examine a single spectrum rather than the bulk average of the spectrum, as is usually done.

[0060] As is evident from Figures 12A-C, anomalies were detected at the single-particle level even without batch averaging of mass spectra. Furthermore, the autoencoder's performance was highest at the single-particle level (batch size 1, Figure 12A), with a high anomaly classifier AUC of 0.9871, a true positive rate of 0.9667, and a very low false positive rate of 0.0056. This was an unexpected result, as single spectra are generally considered to be too noisy to be properly classified (background and threat) by certain algorithms. Consequently, averaging spectra in batches (compared to single-particle spectra or single averages) is not necessary to enhance signal-to-noise separation. Single spectra may be too noisy to classify at the strain / organism level, but they do not appear to be too noisy to classify at the background and anomaly levels. In the anomaly detection process, there are no prior assumptions about the composition or type of threat agent or hazardous substance being analyzed. The autoencoder is trained using a large training dataset to learn the mass spectral features of single particles (individual aerosol particles) in the environment (background), and if a measurement deviates from the background, it is flagged for further analysis.

[0061] The exemplary methods disclosed herein were implemented using a deep learning workstation (Lambda Labs) having the following exemplary specifications. (a) Operating System: Ubuntu 20.04 (including Lambda Stack for managing TensorFlow, PyTorch, CUDA, cuDNN, etc.) (b) Processor: AMD Threadripper 3960X: 24 cores, 3.80GHz, 128MB cache, PCIe 4.0, (c)CPU cooler: air cooling, (d) GPU: 1x RTX A4000, 16GB, (e) Memory: 32GB (2x16GB, 3200MHz, (f) Operating system drive: 1TB SSD (NVMe).

[0062] The code was written in a web-based Jupyter interface that runs Python and uses many open-source libraries, including PyTorch for machine learning, Pandas for data manipulation and analysis, and Plotly for interactive graphing.

[0063] [example] [Example 1. Simulation of dilution concentration of aerial threats] Using a single-particle MALDI-TOF mass spectrometer, two datasets were acquired: the raw spectrum of Bacillus globulus (Bg) (MRI Global) and the raw background spectrum of an air sample collected at Newark Airport. Exposure with 3,000 particles of a 10% diluted threat agent or analyte was simulated. In other words, 10% of these 3,000 spectra were simulated substances (threat agent or analyte), and the remaining 90% were background spectra. Figures 4A-C show heatmaps of (A) the raw Bg spectrum, (B) the background spectrum, and (C) a simulated mixture randomly selected from 10% of the Bg dataset and 90% of the background dataset, following several implementation examples. Figures 4A-C show heatmaps of a simulated mixture randomly selected from 10% of the Bg dataset and 90% of the background dataset. Figures 4A-C demonstrate that even a 10-fold dilution can significantly reduce the human eye's ability to detect the threat. Furthermore, as shown in Figures 5A-B, using the values ​​from the aerosol simulation described above (p=300 / 3000=0.1, q=1-p=0.9), the theoretical binomial distribution and the experimental simulation distribution closely match for batch sizes (n) of 1 (single particle) and 5, suggesting that the probability distribution of the number of simulated spectra in a given batch size is a binomial distribution. This experiment demonstrates the distribution of a limited number of threat particles in an arbitrary background over a fixed operating period. The effects of processing spectra with batches of different sizes and concentrations should also be investigated.

[0064] [Example 2. Classification of aerosol TOF-MS mass spectral features of ambient air samples using the KNN classifier] Figures 8A-B show the confusion matrix (A) and standard deviation (B), which are higher for missed or misidentified samples, illustrating the KNN classifier's ability to detect and identify living agents against noise, following several implementation examples. The KNN classifier was used to classify aerosol TOF-MS mass spectra of over 2000 atmospheric samples, including background samples (without biological agents) and samples containing biological agents or analyte hazardous substances. The agents included spore strains, vegetative bacteria, viruses, and toxins across a wide concentration range. The confusion matrix generated using living agent or analyte hazardous substance data and a subset of background samples is shown in Figure 8A. No false detections were observed. Due to the wide range of challenge concentrations, some missed weak signals (agents or analyte hazardous substances identified as background) and false identifications (agents or analyte hazardous substances misidentified) were observed. These misclassifications were due to weak responses of selected features and an overall increase in spectral noise levels. While this was not unexpected, it is noteworthy that the system was able to detect the presence of drugs or analyte hazards with a low response to specific features. Figure 8B shows that simple measurements of the signal standard deviation in the 3,000–10,000 m / z range correlate well with the presence of drugs or analyte hazards. This is because the standard deviations for missed detections and false detections were significantly higher than the standard deviation measured in the background alone. The intensity distribution of these background spectra was found to be stable. As a result, background features can be extracted using machine learning techniques, and for any given sample, it can be determined whether they match the background distribution.

[0065] The exemplary systems and methods disclosed herein are applicable to a wide range of uses, from protecting critical infrastructure to clinical diagnostics, but are not limited to these. In the biodefense field, deployment in transportation, sports and entertainment venues, or government facilities could potentially detect the intentional release of biological weapons or biologically analyte hazardous substances. In the medical field, these systems and methods could be used to rapidly analyze exhaled pathogens for individualized screening at entry points or for early point-of-care diagnosis of respiratory infections.

[0066] The abstract is provided in accordance with 37 C. FR § 1.72(b) to enable readers to quickly determine the nature and essence of the technical disclosure from a general understanding. It should not be used to interpret or limit the claims or their meaning.

[0067] Although this disclosure has been described in relation to preferred forms of implementation, those skilled in the art will understand that many modifications can be made to it without departing from the spirit of this disclosure. Therefore, the scope of this disclosure is not intended to be limited by the foregoing description.

[0068] It should be understood that various modifications can be made without deviating from the essence of this disclosure. Such modifications are implicitly included in the explanation. They remain within the scope of this disclosure. It should be understood that this disclosure is intended to result in patents that cover many aspects of the disclosure, both independently, as a system, and in both method and apparatus modes.

[0069] Furthermore, each of the various elements of this disclosure and claims may also be achieved in various ways. This disclosure should be understood to encompass each such variation, whether it be a variation of any device implementation, an implementation of a method or process, or merely a variation of any of these elements.

[0070] Specifically, it should be understood that each element's word may be expressed by equivalent apparatus or method terminology, even if only the function or result is the same. Such equivalent, broader, or more general terms should be considered included in the description of each element or action. Such terms may be replaced where necessary to indicate the implicitly broad scope to which this disclosure is entitled. It should be understood that all actions may be expressed either as means to perform that action or as elements that cause that action. Similarly, each physical element disclosed should be understood to encompass the disclosure of the action that the physical element facilitates.

[0071] Furthermore, with respect to each term used, a common dictionary definition, such as that found in at least one of the standard technical dictionaries recognized by the art and the latest edition of Random House-Webster's Unabridged Dictionary, should be understood to be incorporated herein for each term and all definitions, alternative terms, and synonyms, provided that its use in this application does not contradict such interpretation.

[0072] Furthermore, the use of the transitional phrase “comprising” or “comprise” is used to maintain the “open-ended” claims of this specification, in accordance with conventional claim interpretation. Therefore, unless otherwise required by context, “comprising” is intended to mean including the described element or step or group of elements or steps, but not to exclude other elements or steps or groups of elements or steps. Such terminology should be interpreted in the broadest manner to provide the applicant with the broadest legally permissible scope.

Claims

1. In a method for predicting anomalies in the background environment, The steps include creating a dataset of reference analyte mass spectra of a single aerosolized analyte particle that represents hazardous particles present in the background environment, and The steps include creating a dataset of reference background mass spectra of single aerosolized background environmental particles that do not contain the hazardous substance particles to be analyzed, Using the above-mentioned reference mass spectrum of the hazardous substance to be analyzed and the above-mentioned reference background mass spectrum, a computer generates a dataset of simulated composite mass spectra that represent the hazardous substance particles in the background environment at different concentrations of the hazardous substance to be analyzed in the background environment. A selection step to identify the optimal batch size of the simulated composite spectrum at each analyte hazardous substance concentration, wherein the batch size represents the number of spectra to be averaged, and A method characterized by comprising the step of determining a threshold error for predicting anomalies in the single-particle mass spectrum generated from a new aerosol sample taken from the above environment.

2. The method according to claim 1, wherein composite mass spectra at different concentrations of analyte hazardous substances are generated on a computer by changing a dilution ratio, which is defined as dividing the number of reference analyte hazardous substance spectra by the sum of the number of reference analyte hazardous substance spectra and the number of reference background spectra.

3. The method according to claim 1, wherein the background environment is ambient air.

4. The method according to claim 1, wherein the hazardous substance particles to be analyzed above include at least one of a chemical substance and a biological substance.

5. The method according to claim 1, wherein the above-mentioned standard hazardous substance spectrum and the above-mentioned standard background spectrum are generated using an aerosol MALDI TOF-MS system.

6. The above step of identifying the optimal batch size for each hazardous substance concentration to be analyzed is, The steps include training the autoencoder to learn the features of the above-mentioned reference background spectrum, The steps include generating a receiver operating characteristic curve (ROC) based on the probability distribution of the autoencoder's reconstruction loss associated with the individual analyte hazardous substance spectra and background spectra, The method according to claim 1, further comprising the step of identifying an optimal batch size as the batch size that maximizes the area under the ROC curve (AUC).

7. The method according to claim 6, further comprising the step of discarding mass spectral features below a predetermined m / z cutoff value before the above training step.

8. The method according to claim 7, wherein the above-mentioned predetermined m / z cutoff value is less than approximately 3000.

9. The method according to claim 6, wherein the step of determining the threshold error comprises selecting the optimal reconstruction loss threshold from the ROC curve by maximizing Youden's J statistic.

10. The method according to claim 1, wherein the optimal batch size is 1.

11. In methods for diagnosing abnormalities in the environment, A step of generating a single-particle mass spectrum from aerosol samples collected continuously or at predetermined intervals from the environment, The steps include inputting the single-particle mass spectrum into an autoencoder trained to compress and reconstruct the mass spectrum based on a baseline reference mass spectrum of an aerosol sample taken from the environment, The steps include: examining the above single-particle mass spectrum in a predetermined optimal batch size; The steps include determining the reconstruction loss threshold associated with the baseline reference mass spectrum of the aerosol sample taken from the above environment, A method characterized by comprising a diagnostic step for diagnosing an abnormality in the machine environment when the reconstruction loss of the autoencoder associated with the aerosol sample exceeds the reconstruction loss threshold, wherein the diagnostic step detects the abnormality at the single-particle level without performing batch averaging of mass spectra.

12. The method according to claim 11, wherein the optimal batch size is 1.

13. The method according to claim 11, wherein the single-particle mass spectrum from an aerosol sample collected from the above environment is generated using an aerosol MALDI TOF-MS system.

14. In a method for predicting background environmental anomalies caused by one or more unknown aerosol analyte particles, The steps include creating a dataset of reference background mass spectra for a single background environment particle, A training step involves training an autoencoder to learn the features of the above reference background mass spectrum with various batch sizes, wherein the batch size represents the number of spectra to be averaged. The steps include determining a background spectrum reconstruction loss threshold at which at least 90% of the above background spectrum is discarded as not being abnormal, and A method characterized by having the step of predicting one or more anomalies in a single particle mass spectrum generated from an aerosol sample taken from the above background environment sample, when the reconstruction loss related to the sample spectrum exceeds the reconstruction loss threshold.

15. The method according to claim 14, further comprising the step of discarding mass spectral features below a predetermined m / z cutoff value before the above training step.

16. The method according to claim 15, wherein the above-mentioned predetermined m / z cutoff value is less than approximately 3000.

17. The method according to claim 14, wherein the above mass spectrum is generated using an aerosol MALDI TOF-MS system.

18. The method according to claim 14, wherein the autoencoder is trained over a predetermined number of epochs on at least about 20,000 single-particle spectra collected from an environmental sample.

19. The method according to claim 18, wherein the predetermined number of epochs is the greater of 20 epochs or the number of epochs in which the percentage reduction in average training loss is minimized to numerical precision.

20. The method according to claim 14, further comprising the step of examining the averaged and processed sample spectrum to determine the presence of one or more features specific to an unknown analyte hazardous substance.

21. The method according to claim 14, further comprising the step of examining a heat map of the processed sample spectrum to determine the presence of one or more statistically significant features characteristic of an unknown analyte.