Chemical labeling and identification via chemical featurization, machine learning, and relevance analysis

A machine learning and graph-based network approach efficiently identifies molecular components in chemical mixtures using rotational spectroscopy, achieving high accuracy and reducing analysis time from days to hours.

WO2025178911A1PCT designated stage Publication Date: 2025-08-28MASSACHUSETTS INST OF TECH

Patent Information

Application Number
PCT/US2025/016408
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-24
Filing Date
2025-02-19
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing spectroscopic methods struggle to accurately and efficiently identify molecular components in complex chemical mixtures due to overlapping spectral features and the sheer number of catalogued transitions, often requiring manual analysis that can take days to months.

Method used

A machine learning-based approach using rotational spectroscopy, molecular embedding, and graph-based networks to analyze chemical mixtures, assigning molecular carriers by constructing a graph with weighted nodes based on chemical relevance and spectral data, reducing analysis time to under an hour with high accuracy.

Benefits of technology

Achieves ~95% accuracy in identifying molecular components of chemical mixtures in under an hour, correcting errors in existing methods, and is applicable to various spectroscopic techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025016408_28082025_PF_FP_ABST
    Figure US2025016408_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The determination of chemical mixture components is vital to a multitude of scientific fields. Oftentimes various spectroscopic methods are employed to decipher the molecular composition of these complex mixtures. The sheer density of spectral features of different molecules present in such observations may make unambiguous assignment to individual species using these methods challenging. Yet, components of a mixture are commonly chemically related due to environmental processes or shared precursor molecules. Therefore, along with investigating the spectroscopic signals, analysis of the structural and chemical relevance of a molecule is an important consideration when determining which species are present in a mixture. Machine-learning molecular embedding methods are used with a relevance module to determine the likelihood of a molecule being present in a mixture based on the other known species, chemical priors, and spectroscopic information. By incorporating this metric, the mixture components can be identified with extremely high accuracy (∼ 97%).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CHEMICAL LABELING AND IDENTIFICATION VIA CHEMICAL FEATURIZATION, MACHINE LEARNING, AND RELEVANCE ANALYSIS

[0002] This application claims priority of U . S . Provisional Patent Application Serial No . 63 / 557 , 524 , filed February 24 , 2024 , the disclosure of which is incorporated by reference in its entirety .

[0003] Field

[0004] This disclosure describes systems and methods of labeling and identi fying chemicals in an unknown chemical mixture .

[0005] Background

[0006] Chemical samples almost ubiquitously come in the form of complex mixtures as opposed to pure compounds . The determination of the components of such mixtures is of central importance in applications ranging from pharmaceuticals and food sciences to environmental chemistry and astrochemistry . A multitude of analysis techniques are commonly used for the qualitative or quantitative characteri zation of these mixtures , including microwave , infrared, Raman, UV-Vis , X-ray, and NMR spectroscopies along with the coupling of liquid chromatography or gas chromatography with mass spectrometry (LC-MS and GC-MS ) . The utility of each of these techniques heavily depends on the nature of the sample and the analysis . Yet , in every case, the signals generated and interpreted are uniquely molecule-speci fic . That is to say, every signal carries unique molecular information that can be used to interpret and uniquely assign these features to a molecular carrier . Doing so in an automated fashion is highly desirable both for fundamental research and for applications from industry to field measurements and monitoring .

[0007] One technique that has been fairly underutili zed with regards to mixture analysis is rotational spectroscopy . This method measures transitions in the rotational states of freely rotating molecules in the gas phase following interaction with radiation . The ef ficacy of this method with regards to mixture characterization stems from its high structural sensitivity and spectral resolution . The spacing of the rotational transitions directly depends on the molecules ' moment of inertia along three principal axes . Therefore , even structurally similar conformers , isomers , and isotopologues have completely distinctive rotational spectra since the mass distributions are unique . The highly resolved and oftentimes narrow spectral features that result from this method limits peak overlapping, theoretically making identi fication of individual carriers straightforward .

[0008] The advent of broadband and semi-automated rotational spectroscopy techniques has resulted in an explosion in the number of molecules characteri zed and their spectra catalogued in the literature and public databases . While rotational transition frequencies are entirely unique to each molecule , and uncertainties are often very small , the sheer number of catalogued transitions means that often many potentially viable molecular carriers have rotational transitions nearby any measured frequency . Consequently, when analyzing a mixture using rotational spectroscopy, considering only the closeness of frequency match is oftentimes not suf ficient for accurate characteri zation . The same is true of data collected using other types of spectroscopy . As a result , assigning experimental mixture spectra may require further analysis .

[0009] Therefore , it would be beneficial i f there were a system and method to identify the various molecular components within a chemical mixture . It would be advantageous i f this technique was highly accurate and consumed less time than is currently required .

[0010] Summary

[0011] The determination of chemical mixture components is vital to a multitude of scienti fic fields . Oftentimes various spectroscopic methods are employed to decipher the molecular composition of these complex mixtures . The sheer density of spectral features of di f ferent molecules that are often present in such observations may make unambiguous assignment to individual species using these methods challenging . Yet, components of a mixture are commonly chemically related due to environmental processes or shared precursor molecules . Therefore , along with investigating the spectroscopic signals , analysis of the structural and chemical relevance of a molecule is an important consideration when determining which species are present in a mixture . Machinelearning molecular embedding methods are used with a relevance module to determine the likelihood of a molecule being present in a mixture based on the other known species , chemical priors , and spectroscopic information . By incorporating this metric in a rotational spectroscopy mixture analysis algorithm, the mixture components can be identi fied with extremely high accuracy ( ~ 97 % ) in a much more ef ficient manner than manual analysis . According to one embodiment, a method of identifying molecules in a chemical mixture is disclosed . The method comprises : a . obtaining a data set , the data set comprising a plurality of signals , each signal having an attribute with a value and being caused by a molecule in the chemical mixture ; b . selecting one of the plurality of signals ; c . identifying candidate molecules based on the value of the attribute of the selected signal ; d . determining a relevance score of each candidate molecule based on a list of known molecules in the chemical mixture ; e . adj usting the relevance score of each candidate molecule based on other criteria to obtain likelihood scores for each candidate molecule ; f . assigning one of the candidate molecules to the selected signal if the likelihood scores are above a predetermined threshold; g . adding the assigned candidate molecule to the list of known molecules ; and h . repeating steps b . - g . for each of the plurality of signals .

[0012] In some embodiments , identifying candidate molecules comprises comparing the value of the attribute of the signal to a database of known molecules ; and selecting the known molecules having a signal with the attribute having a value within a certain threshold of the value of the attribute of the signal as the candidate molecules . In some embodiments , determining the relevance score comprises using molecular feature vectors generated using a molecular embedder . In some embodiments , determining the relevance score comprises : a . using a molecular embedder to create molecular feature vectors for a plurality of molecules ; b . creating a graph-based network by plotting each molecular feature vector in a chemical vector space ; c . identifying a distance between pairs of molecular feature vectors ; d . connecting a pair of molecular feature vectors with bidirectional edges i f the distance between the pair of molecular feature vectors is less than a predetermined threshold; e . assigning a weight to the molecular feature vectors associated with the list of known molecules ; and f . calculating weights for a remainder of the molecular feature vectors in the chemical vector space based on the weights assigned to the molecular feature vectors associated with the list of known molecules and the bidirectional edges , wherein the weight of a candidate molecule is used to determine its relevance score .

[0013] In certain embodiments , the relevance score of a candidate molecule is a percentile score of the weight of the candidate molecule as compared to all of the weights of all of the molecules in the graph-based network .

[0014] In some embodiments , determining a relevance score comprises : a . using a molecular embedder to create molecular feature vectors for the list of known molecules ; b . creating a chemical vector surface by plotting each molecular feature vector as a function having an amplitude and a width; c . using the molecular embedder to create a molecular feature vector for each candidate molecule ; and d . calculating an amplitude for each candidate molecule based on a superposition of the plotted functions , wherein the amplitude is used to determine the relevance score .

[0015] In certain embodiments , additional amplitudes are determined for a plurality of other molecules ; and the relevance score of a candidate molecule is a percentile score of the value of the candidate molecule as compared to all of the additional amplitudes . In certain embodiments , the function comprises a Gaussian function .

[0016] In some embodiments , the relevance score is adj usted based on a difference between the value of the attribute of the signal and a value of the attribute in a closest signal in the candidate molecule . In some embodiments , the relevance score is adj usted based on an alignment of other signals in the candidate molecule to other signals in the data set . In some embodiments , the relevance score is adj usted based on a presence of unexpected elements . In some embodiments , the relevance score is adj usted i f the candidate molecule is isotopically substituted and a main isotopologue is not present . In some embodiments , the likelihood score of each candidate molecule is provided to a Softmax function and outputs from the Softmax function are also used to determine whether one of the candidate molecules should be assigned to the selected signal . In some embodiments , the attribute of the signal comprises a frequency, a mass / charge ratio, a retention time or a chemical shift. In some embodiments, the data set is generated from rotational spectroscopy, Nuclear Magnetic Resonance (NMR) data, high-performance liquid chromatography (HPLC) , mass spectroscopy (MS) , gas-chromatography MS (GC-MS) , electron paramagnetic resonance (ERR) , or infrared, visible, ultraviolet, or x-ray spectroscopies.

[0017] According to another embodiment, a computer system for performing the method of claim 1 is disclosed. The computer system comprises a processing system; computer storage accessible to the processing system, and computer program instructions encoded on the computer storage, wherein when the computer program instructions are processed by the processing system, the computer system is configured to receive the data set of the chemical mixture; and execute the method of claim 1 to identify the molecules in the chemical mixture.

[0018] Brief Description of the Drawings

[0019] For a better understanding of the present disclosure, reference is made to the accompanying drawings, in which like elements are referenced with like numerals, and in which:

[0020] FIG. 1 is a general schematic of the automated assignment process ;

[0021] FIG. 2 shows known rotational transitions within 0.3 MHz of a transition observed in the benzene / O2discharge experiment;

[0022] FIG. 3 is a schematic of the line assignment process for a mixture studied with rotational spectroscopy; FIG . 4A-4C are a depiction showing the generation of the graph-based network over time ;

[0023] FIG . 5 shows experimental rotational spectrum collected in the benzene discharge experiment overlaid with the simulated spectrum of cyclohexadiene at 4 K;

[0024] FIG . 6 shows an example of a computer system that may execute the disclosed method;

[0025] FIG . 7 shows the ranking algorithm according to one embodiment ; and

[0026] FIG . 8 shows the continuous chemical vector surface according to one embodiment .

[0027] Detailed Description

[0028] The disclosure herein includes the analysis and quantification of the molecular components of chemical mixtures . These chemical mixtures range from atmospheric samples , to human breath, wastewater, and even interstellar nebulae . This disclosure describes an approach to encode a molecule ' s chemical composition and structure into a computer-readable form, and, using machine learning and relevance analysis , rapidly, accurately, and in an automated fashion, both determine the chemical composition of mixtures based on experimental measurements and predict likely undetected trace components . The general process by which this is accomplished (which is application agnostic ) is described herein . The chemical composition and structure of every molecule may be uniquely described by a string of alpha-numeric characters . Using machine learning techniques , including but not limited to natural language processing algorithms , these text representations of molecules are converted into mathematical representations . Through this process , the algorithms learn the " language" of molecules , where elements , structures , chemical bonds , and other properties form the "words" of the " sentences" that are molecules . These mathematical representations signify a molecule ' s location in "chemical space , " where the distances between two molecular vector representations correlates with similarities and differences in molecules . Once encoded, these representations are then used in connection with experimental measurements to intelligently determine which molecules are most likely to be responsible for an observed experimental signal .

[0029] The proof-of-concept study was performed using rotational spectra of molecules . In rotational spectroscopy, the patterns of light each molecule absorbs or emit as it tumbles end over end in space are observed . Each molecule possesses a unique spectrum, or fingerprint , that once measured, may be used to identify that molecule in the future . In this work, analysis was performed on the data collected when a chemical reaction was initiated by electri fying the molecule benzene by itsel f or in combination with nitrogen or oxygen gas . The spectrum includes hundreds of signals from the overlapping spectra ( fingerprints ) of every molecule in the mixture . Each signal ( also referred to as a spike or transition) has an associated amplitude and frequency . A computer was tasked with identi fying which molecules were present by analyzing each individual signal in the spectrum and comparing with previously measured fingerprint spectra . The challenge , however, is experimental accuracy : both in the cataloged spectra and those newly measured . While there is only one possible molecule in the universe that produces each of the signals seen, many molecules produce signals very close - often too close for an experiment to distinguish . Given these uncertainties , using a computer to automatically assign the carrier from all possibilities that fall within the experimental errors is usually no more accurate than random chance . In a typical experiment , therefore , a spectroscopist will use their "chemical intuition" and knowledge of what molecules were involved in the reaction to make an educated j udgment call about which assignment is likely correct . The end result of all of this is that fully disentangling a chemical mixture , even with the assistance of previously state- of-the-art computer algorithms , could take a single trained spectroscopist days to as much as months . This disclosed approach has reduced this number to less than an hour, with almost no human input .

[0030] To achieve this , it is assumed that molecules that are nearby each other in chemical space are more likely to be in a chemical mixture with each other than those that are not . Thus , in one embodiment, the method begins with the construction of a graph, with each molecule represented by a node . Each node that is within a certain distance (which is a tunable hyper-parameter ) in chemical space to another is then connected by a spoke, also referred to as a bidirectional edge . Weights are given to the nodes that represent molecules known to be present in the system, because they were known precursors (which may be benzene , nitrogen, and oxygen in this speci fic example ) . This information then cascades through the graph, with nodes (molecules ) connected to these known molecules gaining weight , tertiary nodes connected to those then gaining a reduced weight , and so forth .

[0031] Subsequently, a speci fic spectral line is selected and compared to the library of known frequencies . Based on the comparison, multiple molecules are suggested as the carrier . A variety of weighted heuristics are employed, based on the importance of each candidate molecule in the graph in cooperation with quantum mechanics and statistical mechanics . Based on this the spectral line is assigned to a particular molecule .

[0032] I f the spectral line is assigned to this particular molecule that was not already a known part of the mixture , that node in the graph is given an inj ection of weight and the graph is re-generated with the new information . This process is repeated until all the signals in the sample are analyzed . In this proof-of-concept study, this process took less than an hour, whereas the original analysis of this dataset took a team of expert spectroscopists about 6 months . In addition, this approach proved more accurate , correcting a number of errors in the original analysis . An accuracy of ~95% was achieved in this proof-of-concept work, with little additional ef fort at refinement .

[0033] This technique is general and extensible to any system which presents features that are speci fic to an individual molecule . For example , with minimal ef fort, the approach may be used to assign features and mixture components from Nuclear Magnetic Resonance (NMR) data, from high-performance liquid chromatography (HPLC ) , mass spectroscopic (MS ) , gas-chromatography MS ( GC-MS ) , electron paramagnetic resonance (ERR) , and infrared, visible , ultraviolet , and x-ray spectroscopies , as j ust a few examples . This analysis platform may be used to augment and enhance existing, widely used, and commercially available instruments already ubiquitous in industrial and laboratory settings . In lay terms , this is a real advance toward a "CSI-like" machine where a sample is inserted into a black box and the computer, with high accuracy, simply reports what molecules are in the sample a short while later .

[0034] Disclosed herein is an Automated Mixture Analysis via Structural Evaluation (AMASE ) system and method that may be applied to the analysis of chemical mixtures using rotational spectroscopy . However, it is understood that AMASE is techniqueagnostic . First , the overall workflow is described, highlighting where application-speci fic customi zations are made .

[0035] The general schematic for AMASE is shown in FIG . 1 . AMASE works with any dataset that includes distinct signals that arise from molecule-speci fic properties . This may be an NMR spectrum consisting of chemical shi fts , a mass spectrum consisting of m / z ratios , a rotational / vibrational / electronic absorption or emission spectrum, and so forth . AMASE interfaces with a database , which may be online and / or local , that includes the known signal s of molecules of interest . The system then attempts to confidently assign the carrier of every observed experimental signal as either belonging to a speci fic molecule in the database or one that is currently unknown .

[0036] The experimental signals from the mixture 10 are sorted (by either absolute strength or signal-to-noise ratio ) and iterated through one at a time . The database or catalog 11 is queried for all signals from known molecules which fall reasonably close to the spectra of the observed signal . Note that the selection criteria and uncertainty tolerances are tunable hyperparameters . This then results in a list of candidate molecules 12 and their associated candidate signal assignments .

[0037] Each of these candidate molecules and signals will eventually be assigned a ranked score by AMASE, indicating the confidence that this is the carrier of the experimental signal . A variety of factors , each with their own weight, are considered, many of which are tuned to the speci fic experimental technique . For example , in spectroscopic investigations , contributions to the match can include how close the signal falls in frequency to the experimental line, the uncertainty of the catalog line, whether there are other transitions of the species which should be present in the data and whether those are missing or present . The weights given to each of these factors are tunable hyperparameters .

[0038] Additionally, in one embodiment , AMASE employs a graph-based network 13 to leverage chemical information and incorporate it into the matching criteria . Before the graph can be constructed, a large training set 14 ( typically a few million) of species is gathered from various databases and used along with a molecular embedding algorithm to construct molecular feature vectors 15 . These molecular feature vectors 15 are numerical representations of molecules that encode learned structural and chemical information for each species in the context of the full training set .

[0039] A graph-based network 13 is then constructed from a subset of the species ( typically a few hundred thousand) that are likely to be chemically relevant to the mixture 10 under study . Each species is represented as a node in the graph, and the connectivity and structure of the graph is then informed by the relationships between the molecular feature vectors of each species . I f any species are known to be present in the mixture a pri ori or become known ( as explained below) , these known priors 16 are initiali zed with a large weight 17 to that node in the graph to represent their importance . Weights are then iteratively distributed throughout the graph-based network 13 , based on connectivity, such that species nearby those known to be present have increased weight as being more likely carriers of signals .

[0040] The AMASE system and method then compiles a final match score for each potential assignment based on the experiment-speci fic criteria outlined above combined with the likelihood of each species from the graph-based network 13 . It then generates a rank- ordered list 18 of candidate signals and carriers and assesses whether the top-ranked candidate passes a threshold to be confidently assigned 19 . The threshold for a confident assignment is a tunable hyperparameter .

[0041] I f a new species is assigned to be present , it is added to the list of known priors 16 and the weights 17 in the graph are updated to reflect the new importance of this node . The entire process is then repeated for the next line in the experimental signal in the spectrum of the mixture 10 . After each assignment , all prior assignments are iteratively re-checked, to ensure that the increasingly large pool of information gleaned from the process does not suggest a prior misassignment . In the end, the AMASE system produces for each signal either a confident assignment , a suggested assignment with potential alternatives , or a label of unidenti fied, indicating that the species is either not in the database containing known signal assignments or insuf ficient evidence exists to make an assignment .

[0042] Various steps shown in FIG . 1 are now described in more detail .

[0043] When ranking molecular candidates for each signal or transition in an observed rotational spectrum, the structural / chemical relevance of the molecule is cons idered along with the frequency and intensity match of the simulated spectrum for a given species . FIG . 2 illustrates the importance of considering more than j ust the frequency match when deciding the correct molecular carrier . This plot shows a spectral peak 20 that appears in the benzene / O; discharge dataset that is described below . Overlaid on top of it are eight known rotational transitions within 0 . 3 MHz of the center frequency corresponding to six unique molecules , labelled A-F . In this case , the catalog line with the closest frequency to the observed transition (molecule C ) is actually not the correct molecular carrier . Rather, 4- ethenylidene-cyclopent-2-en- l-one (molecule D) is the true carrier of this line in the mixture . In fact , all of the transitions have error bars ( shown as hori zontal lines ) that overlap with the center frequency of the observed peak . These error bars are 10 times the statistical catalog uncertainties . This scaling factor was employed because it is well known that the statistical uncertainties in rotational spectroscopy database commonly underestimate true values . Therefore , it would not be statistically defensible to select any of these lines based solely on the analysis of the line frequencies . The Python package molsim was used to visuali ze and work with the experimental signals in the spectrum of the mixture 10 . The first step in this algorithm i s to determine the frequencies of the peaks in the mixture spectrum using a peak-finder algorithm . Once the peak frequencies and intensities are determined, they are sorted by intensity ( from strongest to weakest ) .

[0044] Next , the potential molecular carriers are determined for each peak by querying molecular spectroscopy databases and / or local catalogs 11 for nearby transitions to the observed peak frequencies . For this experiment, which used rotational spectroscopy, the database that was predominately queried was the Splatalogue database . This database was searched for all transitions within 0 . 5 MHz of the peak frequency . On average , there were approximately seven molecules that were candidates for each observed line in the mixture spectra . Of course , other catalogs and databases may be used . Additionally, the range on either side of the peak frequency (which was ±0 . 5 MHz in thi s embodiment ) may also be varied .

[0045] Once each candidate molecule 12 is identi fied, a molecular embedder is used for the graph-based network construction process . In some embodiments , the mol2vec model is used . This model may be selected because the application in this work is focused on a mixture of small hydrocarbon molecules that are relevant to astrochemistry . Previously, several studies have shown that this mol2vec algorithm is ef fective for molecular embedding in astrochemical machine learning applications . That being said, this embedder has also been successfully applied in proj ects pertaining to various other fields , such as pharmaceutical sciences and drug candidate analysis . This technique is embedder-agnostic, and di f ferent approaches may yield more optimal representations for applications such as NMR .

[0046] To create the molecular feature vectors 15 , the molecular embedder, mol2vec, may be used . mol2vec is an unsupervised machine learning method that re-purposes the word2vec algorithm to create molecular feature vectors . Using a training set of molecules ( in the form of SMILES strings ) , the Morgan algorithm is first employed to create a dictionary of substructures within a certain radius of each atom . This dictionary is then fed into a multi-layer perceptron, which is trained to map each substructure to the surrounding substructures in every molecule . Consequently, similar vector representations are generated for substructures that appear in comparable chemical contexts . These substructure representations are subsequently summed to form molecular feature vectors . This model generates multi-dimensional vectors , such as 70-dimensional vectors and was trained using a dataset of 3 , 634 , 046 molecules collected from various online databases like Pubchem, Z INC and the NASA PAH database . In this embodiment , the molecules used to train this model were tailored for astrochemical relevance , as they mostly contained less than 12 non-hydrogen atoms and were only composed of atoms in the first four rows of the periodic table . In other embodiments , more expansive training sets , or those with molecules tailored for other applications , may be used with this approach .

[0047] Furthermore, in other embodiments , a di f ferent molecular embedder, such as VICGAE , may be used . VICGAE produces dense 32- dimensional molecular feature vectors . With these molecular embeddings , the graph-based network 13 required for the relevance generation module may be constructed . The time required for the overall analysis heavily depends on the number of nodes in the graph-based network 13 . Therefore , the number of graph nodes is selected to balance the time required while ensuring that the graph is large enough to consider an ample number of molecular candidates . In one embodiment, the graph included a subset of 288 , 487 molecules from the 3 . 6 million molecules used to train the mo!2vec model . The molecular nodes were then connected with bidirectional edges if the Euclidean distance between their mol2vec-generated vector representations was less than an optimi zed threshold value , which may be tuned . In one embodiment , each node has on average ~121 connections .

[0048] Note that some of the parameters described herein may be tuned . In one embodiment, to determine which hyperparameters , such as the distance threshold, are critical to the graph generation and get an initial set of values for them, a 5- fold cross- validation grid search is used on a training dataset from a typical experiment . In this case , the molecules detected in a dataset from a benzene / 02 discharge mixture was used and were then separated into five splits . In each step, four of the five splits were inputted into the algorithm as "detected" molecules ( also referred to as known priors ) . The hyperparameters were then determined as those that maximi zed the ranking of the molecules in the final split . Time ef ficiency was also considered when determining these parameters . This hyperparameter tuning step need only be performed once for a particular type of experiment , with minor tuning steps then made as part of the assignment process , so long as the molecular contents are not wildly chemically dissimilar in future situations . For example, the hyperparameter set determined from the benzene / Cf discharge is likely sufficient for virtually any astrochemically relevant mixture , and likely for terrestrial mixtures , atmospheric, organic, and / or combustion mixtures . A mixture dominated by very large or inorganic species , on the other hand, would possibly necessitate a new hyperparameter search .

[0049] Once the graph-based network 13 is constructed, the relevance generation module is executed . At the beginning of the process , some molecular priors are provided to the algorithm as the "detected" molecules , also referred to as known priors 16 . For example , the mixture spectra analyzed in one embodiment were generated via the electrical discharge of several species . Therefore , these precursor molecules are initially entered as the known priors 16 . The relevance generation module initiali zes all weights 17 in the graph-based network 13 on the nodes corresponding to the known priors 16 . Following this , the weight of each node is iteratively updated . The exact formalism and the individual weighting parameters may be tuned according to the speci fic experimental technique . In one embodiment, the following equations were used :

[0050] In these equations , R (pl ft) is the score of node p2at time-step t . M (pi ) is the collection of all of the nodes that are connected via a bidirectional edge to node pi . L (pj ) is the number of nodes that are connected to node Pj . FIGs . 4A-4C shows how weights are added to the nodes during each time-step . FIG . 4A shows only the known priors, which each have a weight of 10. FIG. 4B shows the nodes that are connected to the known priors . The two nodes that are connected to one known prior have a weight of 3.33. However, the node in the middle of the graph has two connections to known priors, and its weight is therefore 6.67. FIG. 4G shows the nodes that are two hops from the known priors. These nodes are each connected to at least one one-hop node. Due to a greater number of connections, one of these nodes has a weight of 2.59, while the other one has a weight of 1.11. In summary, during each iteration, the weight of each node is updated based on the weight of its connected nodes. Therefore, in the first iteration, the nodes that are directly connected to the known priors 16 gain weight. The nodes that are two connections from the known priors 16 then gain a further reduced weight in the second time-step. This iterative process continues until the weight of each node converges. In one embodiment, convergence is determined when the weight of each node does not change by more than IO-10between two iterations. Of course, other definitions of convergence may also be used. For example, convergence may be based on a certain number of iterations. As currently constructed, the algorithm typically takes approximately 15 - 30 iterations to converge. During each iteration, the weights of the known priors 16 are also refreshed to their initial value, thus ensuring that the known priors are consistently the highest ranked in the graph. The division by I<(p7) is included so as to limit any significant imbalances in the graph construction. For example, if there are certain nodes that have far more connections than most others, this division ensures that they will not have a disproportionate contribution to the relevance generation module. This overall process results in the nodes that are connected to several of the known priors 16 (and thus being close in chemical vector space to these species ) becoming highly relevant .

[0051] The process of creating the final score using a ranking module according to one embodiment is shown in FIG . 7 . After the graphbased network 13 converges , the resulting relevance scores of the specific candidate molecules with a transition within 0 . 5 MHz from the strongest line in the spectrum of the mixture 10 are collected, as shown in Box 600 . Then, as shown in Box 610 , the raw relevance scores are converted into ranking percentiles based on the relevance scores of all 288 , 487 nodes in the graph ( i . e . i f the score of the molecule is higher than 99% of all other nodes in the graph, it is given a percentile score of 99 ) . This percentile may be referred to as its initial score .

[0052] Following the calculation and collection of the structural / chemical relevance scores , the algorithm proceeds to investigate how likely each candidate molecule is based on the quantum mechanical properties of the molecular transitions in question . The following steps will differ based on the speci fic spectroscopy that is being used for the analysis . The next section will describe the steps taken for analyzing mixtures using rotational spectroscopy . The main considerations are the closeness of the frequency match along with the likelihood of the observed line intensity .

[0053] The score is then adj usted based on how close the catalog frequency is to the peak frequency, as shown in Box 620 . This may be done in various ways . In one embodiment , the frequency match is considered in a linear manner by multiplying the previously mentioned percentile or initial scores by the following scaling f ctor :

[0054] The division by five in this equation was a hyperparameter determined in conj unction with other threshold hyperparameter values . This value was set so as to balance the requirement for a strong frequency match with the potentially large catalog frequency uncertainties . In other embodiments , the score may be adj usted using a di fferent algorithm . For example , rather than multiplying the score by a scalar, a correction value may be calculated based on the frequency di f ference and that correction value may be subtracted from the current score .

[0055] Next , in order to investigate the line intensity likelihood, the spectra for each molecular candidate are simulated using the molsim Python package at the temperature of the mixture . In one embodiment, this temperature may be 4 K, which is often the approximate rotational temperature reached in microwave spectroscopy experiments involving supersonic expansions . ( I f available , the simulations can be forward modeled to include instrument response functions and other instrument-speci fic ef fects . In this initial analysis , the instrument response is assumed to be flat . ) These checks are based on the quantum mechanics properties of the various candidate molecules , as shown in Box 630 . The following checks are then performed : a . I f this line is the strongest occurrence of this molecule in the mixture spectrum, the transition should correspond to one of the strongest simulated transitions at the experimental temperatures . b . The simulated catalog intensities are then scaled such that the catalog intensity of the line in question matches the observed line intensity in the mixture spectrum . There should not be any unreasonably strong predicted molecular transitions ( i . e . much stronger than the most intense observed line in the mixture spectrum) . c . I f other lines of this molecule have been observed and assigned in the mixture , the relative strengths of the lines should be reasonable at the expected rotational temperatures . d . At least hal f of the predicted lines of the molecule that are well above the noise level should be present in the mixture spectrum.

[0056] For each of these conditions that is violated, the score of the molecule is multiplied by a scaling factor . This scaling factor, which is less than 1 , may be 0 . 5 , for example . This scaling factor was chosen to suf ficiently diminish the score of the molecule so that it is no longer in consideration for assignment . Again, in other embodiments , the scaling factor may be di f ferent , or a predetermined constant may be subtracted from the score . In other words , the score is adj usted based on an alignment of other frequencies associated with the candidate molecule to other transitions or peak frequencies in the spectrum .

[0057] The importance of investigating the line intensity is demonstrated in FIG . 5 . This figure shows the experimental spectrum 400 . Note that there are various peaks , including one at 13179 MHz . The graph also shows the simulated lines of cyclohexadiene 410 . Note that the peak of the experimental spectrum 400 at 13179 MHz almost perfectly matches one of the frequencies of a cyclohexadiene catalog transition . However, when that transition is simulated to match the intensity of the line in the experimental spectrum 400 , an even stronger cyclohexadiene transition is expected to be seen around 13167 MHz . Since this line clearly does not appear in the experimental spectrum 400 , there is reason to rule out cyclohexadiene as the molecular carrier for the 13179 MHz line .

[0058] The last considerations of the algorithm pertain to additional structural / chemical factors , as shown in Box 640 . Firstly, i f the candidate molecule is isotopically substituted and the main isotopologue has not been observed in the spectrum or has been observed with line intensities that suggest unrealistic isotopic ratios , the molecular score is multiplied by a scaling factor . This scaling factor, which is less than 1 , may be 0 . 5, for example . Finally, a candidate molecule can be further ruled out i f it contains atoms that are unlikely to be present in the mixture . This is an adj ustable parameter that can be inputted by the user . For example , for the datasets analyzed in this embodiment, since the precursor molecules only contain carbon, hydrogen, nitrogen, and oxygen, the molecular scores of species containing other atoms (besides common contaminants such as sul fur ) are multiplied by the scaling factor . Other or additional heuristics may also be applied . For example, in another embodiment , i f the sample is known to consist of only one phase ( such as the gas phase ) , and the candidate molecule is a solid at this temperature , the molecular score may be multiplied by the scaling factor . This can be tailored by the user because it may also depend on various environmental factors or contaminants from previous trials within an experimental setup . Following the numerous aforementioned considerations , each molecular candidate has a final likelihood score . In order to turn these raw scores into percentages for each molecule , a Softmax function is applied to the final scores , as shown in Box 650 . I f both the raw final score and the Softmax percentage score of the top ranked molecule are above certain thresholds , the algorithm then confidently assigns the transition to the top-ranked molecule , as shown in Box 660 . In this embodiment , these threshold values are 93% and 70% , respectively . These threshold values are hyperparameters that may be optimi zed for the datasets encountered in the particular application . For example , they may be further adj usted for di f ferent applications . I f the raw score threshold is not surpassed by the top ranked molecule , this indicates that the molecule is likely not the correct molecular carrier due to either a frequency, intensity, or structural relevance mismatch . In this case , the line is listed as "unidentified . " This label can either stem from the true molecular carrier not being present in spectroscopic databases or the structural / spectroscopic match not being convincing enough for confident assignment . I f the Softmax score is not above the threshold, this indicates that there may be two or more viable candidates .

[0059] Thus , the ranking module received a list of candidate molecules , each with a relevance score , and creates a rank-ordered list 18 of candidate molecules , each with a likelihood score .

[0060] I f a molecule is confidently assigned, it is then added to the list of known priors 16 . Each time this list is updated, either by the addition of a molecule or, as discussed momentarily, i f one is removed following subsequent analysis , the graph-based network 13 is updated, and new node weights are calculated now accounting for the additional ( or removed) known priors 16 . An example of this process is shown in FIG . 3 .

[0061] The same process as described above is then run on the next strongest observed line in the mixture spectrum . For this line , the known priors 16 are more informed since they also include the molecule assigned to the strongest transition . Starting with the assignment of the second line and any updates to the graph-based network 13 , after each assignment the algorithm then proceeds to re-evaluate every previously assigned line to ensure that the assigned molecular carrier stays consistent given the refined and more informed known priors .

[0062] Finally, after each line in the mixture spectrum is investigated and assigned either to a molecular carrier or given an unidentified label , the structural relevance algorithm is run on the entire graph-based network 13 using the final list of known priors 16 . From this , identi fication of which molecules in the graph are most structurally / chemically relevant to the mixture can be made . These can then be starting points for further study ( through calculation and / or experiment ) either in real-time in parallel or in a follow-up investigation to assist in assigning the remaining unidentified features .

[0063] The process shown in FIG . 1 may be modi fied . For example, the choice of database or catalog 11 may vary based on the area of interest .

[0064] In some embodiments , the choice of molecular embedder may be varied . While the present disclosure described the use of mo!2vec and VICGAE, any suitable molecular embedder may be used to create the molecular feature vectors 15.

[0065] One additional modification is the use of a di f ferent mechanism to generate the relevance for the various candidate molecules . In the embodiment described above , a graph-based network 13 is created . This may be considered to be a discrete relevance generation module , in that nodes are either neighbors of other nodes , or they are not . Thus , however, in another embodiment, a continuous relevance generation module may be used .

[0066] In this embodiment, rather than creating a graph-based network, a continuous chemical vector surface is used . Specifically, all of the known priors 16 are plotted in the chemical vector space . However, these known priors 16 are plotted as Gaussian functions , having an amplitude and a Gaussian width . Of course , other types of functions which also have an amplitude and a width, such as Poisson, Lorentzian or others , may be used as well . The amplitude of any other location in the chemical vector space can then be readily computed based on the superposition of the Gaussian functions of the known priors 16 . FIG . 8 shows a simpli fied illustration that shows this continuous relevance generation module . In this figure , it is assumed that the output of the molecular embedder is two dimensional , and the known priors 16 are plotted in two dimensions , each with a predetermined amplitude and Gaussian width . In this figure, there are 5 known priors 16 , where several of them are grouped very close together, creating a taller, wider peak 800 . This chemical vector surface is then used to generate the relevance score of each candidate molecule . Speci fically, the vector position of the candidate molecule is used as a coordinate position . The relevance score of this particular candidate molecule is then calculated by determining the amplitude of the combined Gaussian chemical space surface at its coordinates . As more known priors 16 are generated, these are added to the chemical vector surface as new Gaussian functions , having the same amplitude and Gaussian width as the other known priors 16. Note that in this continuous relevance generation module , it is not necessary to calculate the relevance of all molecules . The relevance score obtained using this continuous relevance generation module may then be provided to the ranking module to generate the rank-ordered list 18 , as described above . In certain embodiments , the relevance score is converted to a percentile before being provided to the ranking module . In one embodiment , a plurality of randomly selected chemical molecules , such as greater than 1000 , are selected and their relevance scores are computed using the chemical vector surface as described above . These additional relevance scores are used in combination with the relevance scores of the candidate molecules to convert the relevance scores of the candidate molecules into percentile scores .

[0067] It is noted that this disclosure describes the experimental signals as a spectrum of frequencies , however the experimental signals may take other forms . For example , a gas chromatograph mass spectroscopy measurement generates data whose defining attribute is a mass-to-charge ratio (m / z ) , rather than a frequency . However, in all cases , a data set is presented, wherein the data set comprises a plurality of signals , each signal representative of a molecule in the chemical compound . In many embodiments , the data set is a spectrum and the signals are the individual signal transitions (or frequency peaks ) . In other embodiments , the attribute of the signals may be mass / charge ratios , or chemical shi fts , or retention times , and so forth . Thus , each signal in the data set may have an attribute with a value , which may be one of these types of attributes or others depending on the speci fic application, but in each case, will be speci fic to an individual molecule . Note that each signal may have more than one attribute , wherein one attribute is freguency, mass / charge ratio , chemical shi ft, or retention time and a second attribute may be intensity, signal / noise ratio , amplitude , or another parameter .

[0068] Thus , in summary, the system includes a sorting module , which receives a data set comprising a plurality of experimental signals for a mixture and generates a sorted list of signals that are arranged based on the value of some attribute of the signal , such as intensity or signal to noise ratio . The system includes a selection module that receives the next signal from the sorted list and generates a list of candidate molecules based on proximity to the attribute of that signal . The system also includes a relevance generation module , that calculated the relevance of each candidate molecule based on the known molecules in the mixture and molecular feature vectors generated using a molecular embedder . The system also includes a ranking module that uses the relevance scores from the relevance generation module and other heuristics to determine the likelihood that each candidate module is the correct molecule . The relevance generation module and the ranking module are executed repeatedly until the entire sorted list is processed . Note that each of these modules may be implemented as software programs that execute on a computer system. As noted above, one or more computers can be used to implement such a computational process , using one or more general-purpose computers , such as client devices including mobile devices and client computers , one or more server computers , or one or more database computers , or combinations of any two or more of these , which can be programmed to implement the functionality such as described in the example implementations .

[0069] FIG . 6 is a block diagram of a general-purpose computer which processes computer programs using a processing system. Computer programs on a general-purpose computer generally include an operating system and applications . The operating system is a computer program running on the computer that manages access to resources of the computer by the applications and the operating system . The resources generally include memory, storage , communication interfaces , input devices and output devices . Examples of such general-purpose computers include, but are not limited to , larger computer systems such as server computers , database computers , desktop computers , laptop and notebook computers , as well as mobile or handheld computing devices , such as a tablet computer, handheld computer, smart phone , media player, personal data assistant , audio and / or video recorder, or wearable computing device .

[0070] With reference to FIG . 6, an example computer 500 comprises a processing system including at least one processing unit 502 and a memory 504 . The computer can have multiple processing units 502 and multiple devices implementing the memory 504 . A processing unit 502 can include one or more processing cores (not shown) that operate independently of each other . Additional co-processing units , such as graphics processing unit 520 , also can be present in the computer. The memory 504 may include volatile devices (such as dynamic random-access memory (DRAM) or other random-access memory device) , and non-volatile devices (such as a read-only memory, flash memory, and the like) or some combination of the two, and optionally including any memory available in a processing device. Other memory such as dedicated memory or registers also can reside in a processing unit. Such a memory configuration is delineated by the dashed line 504 in FIG. 6. The computer 500 may include additional storage (removable and / or non-removable) including, but not limited to, solid state devices, or magnetically recorded or optically recorded disks or tape. Such additional storage is illustrated in FIG. 6 by removable storage 508 and nonremovable storage 510. The various components in FIG. 6 are generally interconnected by an interconnection mechanism, such as one or more buses 530.

[0071] A computer storage medium is any medium in which data can be stored in and retrieved from addressable physical storage locations by the computer. Computer storage media includes volatile and nonvolatile memory devices, and removable and nonremovable storage devices. Memory 504, removable storage 508 and non-removable storage 510 are all examples of computer storage media. Some examples of computer storage media are RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optically or magneto-optically recorded storage device, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices. Computer storage media and communication media are mutually exclusive categories of media. The computer 500 may also include communications connection ( s ) 512 that allow the computer to communicate with other devices over a communication medium . Communication media typically transmit computer program code , data structures , program modules or other data over a wired or wireless substance by propagating a modulated data signal such as a carrier wave or other transport mechanism over the substance . The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal , thereby changing the configuration or state of the receiving device of the signal . By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media include any non-wired communication media that allows propagation of signals , such as acoustic, electromagnetic, electrical , optical , infrared, radio frequency and other signals . Communication connections 512 are devices , such as a network interface or radio transmitter, that interface with the communication media to transmit data over and receive data from signals propagated through communication media .

[0072] The communication connections may include one or more radio transmitters for telephonic communications over cellular telephone networks , and / or a wireless communication interface for wireless connection to a computer network . For example , a cellular connection, a Wi-Fi connection, a Bluetooth connection, and other connections may be present in the computer . Such connections support communication with other devices , such as to support voice or data communications .

[0073] The computer 500 may have various input device ( s ) 514 such as various pointer (whether single pointer or multi-pointer ) devices , such as a mouse , tablet and pen, touchpad and other touch-based input devices , stylus , image input devices , such as still and motion cameras , audio input devices , such as a microphone . The computer 500 may have various output device ( s ) 516 , such as a display, speakers , printers , and so on, that also may be included . These devices are well known in the art and need not be discussed at length here .

[0074] The various storage 510 , communication connections 512 , output devices 516 and input devices 514 can be integrated within a housing of the computer, or can be connected through various input / output interface devices on the computer, in which case the reference numbers 510 , 512 , 514 and 516 can indicate either the interface for connection to a device or the device itsel f as the case may be .

[0075] An operating system of the computer typically includes computer programs , commonly called drivers , which manage access to the various storage 510 , communication connections 512 , output devices 516 and input devices 514 . Such access generally includes managing inputs from and outputs to these devices . In the case of communication connections , the operating system also may include one or more computer programs for implementing communication protocols used to communicate information between computers and devices through the communication connections 512 .

[0076] Any of the foregoing aspects may be embodied as a computer system, as any individual component of such a computer system, as a process performed by such a computer system or any individual component of such a computer system, or as an article of manufacture including computer storage in which computer program code is stored and which, when processed by the processing system ( s ) of one or more computers , configures the processing system ( s ) of the one or more computers to provide such a computer system or individual component of such a computer system.

[0077] Each component (which also may be called a "module" or "engine" or "computational model" or the like ) , of a computer system such as described herein, and which operates on one or more computers , can be implemented as computer program code processed by the processing system ( s ) of one or more computers . Computer program code includes computer-executable instructions and / or computer-interpreted instructions , such as program modules , which instructions are processed by a processing system of a computer . Generally, such instructions define routines , programs , obj ects , components , data structures , and so on, that, when processed by a processing system, instruct the processing system to perform operations on data or configure the processor or computer to implement various components or data structures in computer storage . A data structure is defined in a computer program and specifies how data is organi zed in computer storage , such as in a memory device or a storage device, so that the data can accessed, manipulated, and stored by a processing system of a computer .

[0078] This system and method has many uses . In some embodiments , this system and method is used to determine the molecular composition of the emissions from a factory or engine . For example , the known priors may be the fuel used by the engine or factory, which may be used by the relevance module . This information may then be used to modi fy one or more parameters of the factory or engine to reduce the emission of harmful chemicals , or to introduce a treatment stage to process the emissions before releasing them into the atmosphere. This modification may occur quickly, such as within one hour, or may be implemented over a longer time period.

[0079] In another embodiment, this system and method may be used to analyze the molecular composition of human breath. This may be performed to detect illicit substances, such as alcohol or drugs, or the presence of pathogens indicative of disease. In response, a corrective action may be taken. This corrective action may include initiating treatment (in the case of disease) , arrest (in the case of illicit substances) , or another action.

[0080] In another embodiment, this system and method may be used to enforce security protocols. This system and method may be used to detect forbidden materials from entering a public space, such as an airport, train station or sports stadium. Detection of a forbidden material may initiate a warning or alert so as to prevent the material from entering the facility.

[0081] In another embodiment, this system and method is used to detect the molecular composition of interstellar entities. The understanding of the molecular composition of interstellar sources drives an understanding of the physical properties of these sources in space, such as the temperatures, velocities, and evolutionary history. They can also trace specific interstellar events (such as shocks) .

[0082] In each of these scenarios, the identification of the various molecular constituents may result in an action, which may be an alert, a change in the operation of a factory or engine , or another action .

[0083] This system and method has many advantages . As previously stated, this method has high accuracy and significantly reduces the time required to determine the constituent molecules in a mixture . For example , testing of chemical mixtures that included more than 400 peaks was performed in less than 1 hours using the graph-based network and in less than 20 minutes using the chemical vector surface technique . In some cases , once the catalogs were queried, the remainder of the analysis took less than 30 minutes and less than 2 minutes , respectively . To demonstrate these attributes , several tests were performed . The first set of tests described below were performed using the graph-based network as the relevance generation module .

[0084] Previously collected and analyzed mixture spectra from McCarthy et al . ( 2020 ) were used to test the disclosure system and method . In McCarthy et al . ( 2020 ) , the authors used a combination of chirped-pulse and cavity enhanced microwave spectroscopy to study the rotational spectrum of three mixtures produced by subj ecting pure benzene as well as combinations of benzene with molecular oxygen and molecular nitrogen to an electrical discharge . This discharge fragments some precursor molecules , generates radicals and ions , and drives some populations into excited states , the combination of which results in a reaction mixture that produces a wide variety of product species . After measuring the broadband rotational spectrum of the mixture , the authors ultimately identi fied over 160 molecular products . Due to the large number of product molecules , the completeness of the line assignment in these spectra, and the wealth of spectroscopic catalogs made available by this work, these mixture spectra were ideal candidates to test the algorithm. Spectra of each mixture were collected in two frequency regions : 8 -18 GHz and 18 -26 GHz . The line intensities in these two regions di f fer notably, almost certainly due to differing response functions in the two microwave circuits used to perform the measurements . Since relative line intensities is a factor in the assignment algorithm, and because the instrument responses for these are not known, these frequency regions need to be analyzed separately . The analysis presented in this Disclosure focused on the 8 -18 GHz portions of each spectrum .

[0085] For each mixture, the known priors 16 inputted into the algorithm ( see FIG . 1 ) were the precursor species ( i . e . benzene and molecular oxygen for the benzene / O2discharge ) . Because there were three available mixtures , the benzene and benzene / O2mixtures were used to determine the appropriate hyperparameters for the various components . The benzene / N2mixture was used purely to validate the disclosed method .

[0086] The algorithm classi fies spectral lines into the following categories : confidently assigned, more than one potential carrier, or unidenti fied . Lines that are confidently assigned have exactly one molecule for which the required calculation thresholds are met . Lines with multiple possible carriers have several molecules for which the thresholds are met . Unidenti fied lines have no molecules that meet these criteria .

[0087] Benzene Discharge

[0088] In total , the dataset contained 419 spectral lines assigned by McCarthy et al . ( 2020 ) from 6-20 GHz . Spectroscopic catalogs were available for the molecular carriers of all but 13 lines. Therefore, this analysis mainly focused on the remaining 406 lines. Of these, the algorithm confidently assigned 384 (94.6%) to a single molecular carrier. 373 of these agree with those determined by McCarthy et al. (2020) . For the lines for which there was a disagreeing assignment, attempts to manually verify which molecule was in fact correct were made. For these additional checks, we sought to ensure that most other strong predicted transitions were present in the mixture spectrum and that the observed relative intensities were reasonable at the expected experimental temperatures. For 8 of the 11 disagreeing lines, it was verified that the algorithm suggested the correct molecular assignment. For example, the algorithm assigned the three transitions at 8772 MHz, 12706 MHz, and 16311 MHz to cyclohexa-2 , 4-dienone, while they were originally assigned to three distinct molecules. Nine additional transition are also assigned to cyclohexa-2 , 4-dienone, and the aforementioned three all match the relative intensities reasonably well. In fact, these three transitions are some of the strongest predicted transitions of the molecule. Therefore, it is very likely that this molecule is accountable for these three lines.

[0089] The algorithm labelled five lines (1.2%) as having multiple possible molecular carriers. However, for four of these five lines, the algorithm suggested that the most likely carrier is the species assigned by McCarthy et al. (2020) . Finally, 17 lines were labelled as unidentified. As was done for the disagreeing confident assignments, the rotational spectra of the assigned molecules were manually simulated to verify why the algorithm labelled the line as unassigned. For eight of these lines, the assigned molecules have several simulated strong transitions that are not present in the experimental spectrum or was assigned based on an unrealistically weak predicted transition . Therefore , the

[0090] "unidenti fied" assignment of the algorithm for these lines was valid .

[0091] Overall , through this manual analysis , it was found that there was only a total of 12 clear misassignments from the algorithm . Thus , i f it is assumed that the assignments for which McCarthy et al . ( 2020 ) and the present algorithm agree are correct, the overall accuracy rate of the algorithm was ~97 . 0% . Further, following the database querying and catalog scraping, the time required to assign the molecules was only 14 minutes .

[0092] Benzene / O2Discharge

[0093] For the discharge of benzene and molecular oxygen, the entire dataset contained 899 lines . Spectroscopic catalogs were available for the molecular carriers of 859 of these lines . Of these 859 lines , the disclosed system and method provided confident assignments for 769 lines . The algorithm and McCarthy et al . ( 2020 ) provided the same molecular carrier for all but 34 of these 769 lines . These 34 disagreements were manually investigated . Ultimately, 24 of the disagreements were believed to be correctly identi fied by the present algorithm. 29 lines were then assigned as having multiple possible molecular carriers . However, in 23 of these 29 instances , the highest ranked of these multiple possible carriers agreed with the assignment of McCarthy et al . ( 2020 ) .

[0094] Furthermore, the algorithm suggested that 61 of these 859 lines were "unassigned, " indicating that none of the molecular candidates were a convincing match . Each of these lines was manually investigated . It was ultimately found that there were 11 clear misassignments by the algorithm. However, in the remaining cases, the algorithm's output was correct. In most of these instances, the algorithm does not assign the molecular carrier because of issues with the relative intensity. For example, if a significant number of stronger or similar intensity simulated lines of a molecule do not appear in the observed spectrum, the algorithm does not have enough evidence to assign the line to the molecule in question. Of note, the majority of these disagreeing assignments occur in the weakest 200 lines of the spectrum. This is fairly unsurprising since it is more difficult to compare relative line intensities when even the strongest lines are hardly above the noise level.

[0095] Overall, through the aforementioned manual analysis of the disagreeing lines, there were only 21 total obvious misassignments by the algorithm. Therefore, the overall accuracy rate was 838 / 859 (97.6%) . Further, following the database querying and catalog scraping, the time required to assign the molecules was only 29 minutes .

[0096] Benzene / N2Discharge

[0097] The prior two mixture analyses were used to tune the model, for instance to arrive at reasonable hyper-parameters. As described earlier, however, once done, the algorithm is applicable to similar mixture analyses with little to no tuning. To test this, the third mixture presented in McCarthy et al. (2020) was used and the algorithm was executed using the hyperparameters previously predetermined. For this mixture, the dataset contained a total of 717 lines that were identified by McCarthy et al. (2020) . Of these lines, spectroscopic catalogs were available for the carriers of all but 24, leaving 693 remaining transitions. Of these, the algorithm confidently assigned 638 (~ 92%) to a single molecular carrier. In 585 of these instances, the algorithm assigned the same molecular carrier as McCarthy et al. (2020) . Following the same procedure as the previous two datasets, the lines for which there was a disagreeing assignment were manually investigated through spectral simulation. Ultimately, of these 53 lines, there were only five clear mis-assignments by the algorithm. In the other instances, the molecule suggested by the algorithm is a feasible match .

[0098] For 24 of the transitions in this mixture, the algorithm suggested that there was more than one possible molecular carrier. However, in all but five instances, the highest ranked of these possible carriers agreed with the assignment of McCarthy et al. (2020) . Finally, the algorithm listed 31 of the 693 lines as "unidentified." Through further manual analysis, there were only six clear misassignments from the algorithm.

[0099] Overall, in the dataset of 693 lines, only 11 lines were clearly misassigned by the algorithm, thus giving an accuracy rate of 98.4%. Further, following the database querying and catalog scraping, the time required to assign the molecules was only 18 minutes. As mentioned previously, this dataset was used solely for validation, while the others were utilized to tune the hyperparameters. It is therefore notable that the assignment accuracy on this dataset is very similar to the benzene and benzene / 02 experiments. This provides confidence that the current algorithm architecture can be successfully generalized to additional mixtures. Cold molecular cloud TMC- 1

[0100] Additional tests were executed using the continuous relevance generation module . This variation of the algorithm was testing using the GOTHAM observations of the cold molecular cloud TMC- 1 . Many of the observational details of this line survey are presented in Sita et al . ( 2022 ) . For this proof-of-concept investigation, data ranging from 18 to 36 . 4 GHz was considered . The algorithm analyzed 438 lines of significance .

[0101] Of the 438 lines , the algorithm uniquely assigned 422 to a single molecular carrier, 3 were listed as having several possible carriers , and 13 were unassigned . 47 unique molecular species were identi fied . All of these molecules have been previously detected toward TMC- 1 , thus suggesting that there were no " false positive" assignments of molecules that are not truly present in the data . For 11 of these species , one or more isotopologues were also assigned . Therefore, in all , 79 molecules / isotopologues / isotopomers were identified in the data . Further, following the database querying and catalog scraping, the time required to assign the molecules with the continuous relevance generation module was less than two minutes .

[0102] In nine instances , a molecule that is not present in the data was only ruled out due to a low structural relevance score . For these molecules , the spectroscopic analysis did not provide suf ficient evidence to dismiss them. This highlights the importance of incorporating structural relevance in the assignment algorithm, as the addition of these nine misassignments would significantly impair the algorithm' s performance and lead to considerably less useful results .

[0103] The present disclosure is not to be limited in scope by the specific embodiments described herein . Indeed, other various embodiments of and modi fications to the present disclosure , in addition to those described herein, will be apparent to those of ordinary skill in the art from the foregoing description and accompanying drawings . Thus , such other embodiments and modifications are intended to fall within the scope of the present disclosure . Further, although the present disclosure has been described herein in the context of a particular implementation in a particular environment for a particular purpose , those of ordinary skill in the art will recognize that its usefulness is not limited thereto and that the present disclosure may be beneficially implemented in any number of environments for any number of purposes . Accordingly, the claims set forth below should be construed in view of the full breadth and spirit of the present disclosure as described herein .

Claims

What is claimed is:

1. A method of identifying molecules in a chemical mixture, comprising : a. obtaining a data set, the data set comprising a plurality of signals, each signal having an attribute with a value and being caused by a molecule in the chemical mixture; b. selecting one of the plurality of signals; c. identifying candidate molecules based on the value of the attribute of the selected signal; d. determining a relevance score of each candidate molecule based on a list of known molecules in the chemical mixture; e. adjusting the relevance score of each candidate molecule based on other criteria to obtain likelihood scores for each candidate molecule; f. assigning one of the candidate molecules to the selected signal if the likelihood scores are above a predetermined threshold; g. adding the assigned candidate molecule to the list of known molecules; and h. repeating steps b.- g. for each of the plurality of signals .

2. The method of claim 1, wherein identifying candidate molecules comprises: a. comparing the value of the attribute of the signal to a database of known molecules; andb . selecting the known molecules having a signal with the attribute having a value within a certain threshold of the value of the attribute of the signal as the candidate molecules .3 . The method of claim 1 , wherein determining the relevance score comprises using molecular feature vectors generated using a molecular embedder .4 . The method of claim 1 , wherein determining the relevance score comprises : a . using a molecular embedder to create molecular feature vectors for a plurality of molecules ; b . creating a graph-based network by plotting each molecular feature vector in a chemical vector space ; c . identifying a distance between pairs of molecular feature vectors ; d . connecting a pair of molecular feature vectors with bidirectional edges i f the distance between the pair of molecular feature vectors is less than a predetermined threshold; e . assigning a weight to the molecular feature vectors associated with the list of known molecules ; and f . calculating weights for a remainder of the molecular feature vectors in the chemical vector space based on the weights assigned to the molecular feature vectors associated with the list of known molecules and the bidirectional edges , wherein the weight of a candidate molecule is used to determine its relevance score .5 . The method of claim 4 , wherein the relevance score of a candidate molecule is a percentile score of the weight ofthe candidate molecule as compared to all of the weights of all of the molecules in the graph-based network .6 . The method of claim 1 , wherein determining a relevance score comprises : a . using a molecular embedder to create molecular feature vectors for the list of known molecules ; b . creating a chemical vector surface by plotting each molecular feature vector as a function having an amplitude and a width; c . using the molecular embedder to create a molecular feature vector for each candidate molecule ; and d . calculating an amplitude for each candidate molecule based on a superposition of the plotted functions , wherein the amplitude is used to determine the relevance score .7 . The method of claim 6, wherein additional amplitudes are determined for a plurality of other molecules ; and the relevance score of a candidate molecule is a percentile score of the value of the candidate molecule as compared to all of the additional amplitudes .8 . The method of claim 6 , wherein the function comprises a Gaussian function .9 . The method of claim 1 , wherein the relevance score is adj usted based on a difference between the value of the attribute of the signal and a value of the attribute in a closest signal in the candidate molecule .10 . The method of claim 1 , wherein the relevance score is adj usted based on an alignment of other signals in the candidate molecule to other signals in the data set .

11. The method of claim 1, wherein the relevance score is adjusted based on a presence of unexpected elements.

12. The method of claim 1, wherein the relevance score is adjusted if the candidate molecule is isotopically substituted and a main isotopologue is not present.

13. The method of claim 1, wherein the likelihood score of each candidate molecule is provided to a Softmax function and outputs from the Softmax function are also used to determine whether one of the candidate molecules should be assigned to the selected signal.

14. The method of claim 1, wherein the attribute of the signal comprises a frequency.

15. The method of claim 1, wherein the attribute of the signal comprises a mass / charge ratio.

16. The method of claim 1, wherein the attribute comprises a retention time.

17. The method of claim 1, wherein the attribute comprises a chemical shift.

18. The method of claim 1, wherein the data set is generated from rotational spectroscopy, Nuclear Magnetic Resonance (NMR) data, high-performance liquid chromatography (HPLC) , mass spectroscopy (MS) , gaschromatography MS (GC-MS) , electron paramagnetic resonance (ERR) , or infrared, visible, ultraviolet, or x- ray spectroscopies.

19. A computer system for performing the method of claim 1, comprising: a processing system; computer storage accessible to the processing system, and computer program instructions encoded onthe computer storage, wherein when the computer program instructions are processed by the processing system, the computer system is configured to: a. receive the data set of the chemical mixture; and b. execute the method of claim 1 to identify the molecules in the chemical mixture.

Citation Information

Patent Citations

  • Sensor arrangement and method for the qualitative and quantitative detection of chemical substances and / or mixtures of substances in an environment

    US20070094179A1

  • Systems and methods for provisioning training data to enable neural networks to analyze signals in NMR measurements

    WO2022263026A1

Cited By

  • Food processing quality inspection method and system based on fingerprint spectrum

    CN121352631A