Machine learned chemical intuition for the ai generation of molecules within specific chemical spaces

WO2026183312A1PCT designated stage Publication Date: 2026-09-03MASSACHUSETTS INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/016805
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-26
Filing Date
2026-02-26
Publication Date
2026-09-03

Smart Images

  • Figure US2026016805_03092026_PF_FP_ABST
    Figure US2026016805_03092026_PF_FP_ABST
Patent Text Reader

Abstract

A method may include reading spectroscopic data of a source corresponding to a known and unknown chemical constituent, generating a chemical space surface having a distribution of weights corresponding to known chemical constituent; sampling weighted points of the chemical space surface; providing a first set of vectors corresponding to the sampled weighted points to a trained ANN, trained to produce a vector encoding the presence of molecular subunits corresponding to the sampled weighted points, for the first set, receiving from the ANN a second set of vectors corresponding one of the first set, each of the second set indicating the presence of molecular subunits, for the second set, generating strings of molecular subunits by concatenating random selections from its molecular subunits, filtering the strings of molecular subunits for chemical validity, mapping the strings to the weighted chemical vector space, and outputting molecular representations corresponding to the unknown chemical constituent.
Need to check novelty before this filing date? Find Prior Art

Description

MTV-25725MACHINE LEARNED CHEMICAL INTUITION FOR THE Al GENERATION OF MOLECULES WITHIN SPECIFIC CHEMICAL SPACESCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This Application claims the benefit of priority to US Application Serial No.63 / 763,880, filed on February 26, 2025, the entire contents of which are hereby incorporated by reference.BACKGROUND

[0002] Modem radio telescopes generate vast amounts of observational data, offering valuable insights into the molecular composition of interstellar sources. Identifying the molecules within these datasets typically involves time-consuming and labor-intensive manual analysis. This problem is also pervasive in other areas other than interstellar spectroscopic research, and thus the various systems and methods described herein are readily applicable to other fields in which large amounts of molecular or other data need be analyzed automatedly.

[0003] In one exemplary field, such as astrochemical research, understanding interstellar molecular inventories is crucial. These detected molecules can provide unique insight into the conditions and properties of interstellar sources. They also probe specific physical processes, such as outflows and shocks. The vast majority (>90%) of interstellar molecular detections have been achieved using radio astronomical observations. With the continual development of state-of-the-art radio astronomy facilities, such as the Green Bank Telescope (GBT) and Atacama Large Millimeter / submillimeter Array (ALMA), extremely high sensitivity and high-resolution interstellar observations can be collected over broad frequency regions. These resulting line surveys can contain thousands of different spectral features, corresponding to molecular emission or absorption. While this wealth of observational data has resulted in a boom in new detected interstellar species (which is increasing annually at a non-linear rate), the assignment of these extremely dense line surveys can be a challenging task. The observed spectral features primarily correspond to rotational emission from gas-phase molecules. The assignment of these surveys is generally tackled by manually comparing the observed astronomical spectral peaks to rotational spectra cataloged in online databases such as the Cologne Database for Molecular Spectroscopy, Jet Propulsion Laboratory (JPL) Database, and Splatalogue.1FH13297226.4MTV-25725

[0004] However, due to the wealth of cataloged molecules, there are often a large number of molecular candidates in the various databases that have rotational transitions nearby (within experimental uncertainty) any observed peak in the observational data. As a result, tedious manual analysis of such data can take weeks or even months for a team of trained scientists. The potential to automate this process is therefore appealing, but is complicated by the need for an expert to often make informed decisions based on chemical intuition, rather than heuristics that are readily translatable to logic gates and Boolean operators.

[0005] Thus, there remains a need for automated systems and methods for analyzing complex collections of molecules.SUMMARY

[0006] This Summary introduces a selection of concepts in simplified form that are described further below in the Detailed Description. This Summary neither identifies key or essential features, nor limits the scope, of the claimed subject matter.

[0007] To achieve these and other advantages and in accordance with the purpose of the disclosed subject matter, as embodied and broadly described, the disclosed subject matter includes a method including reading spectroscopic data of a source, the spectroscopic data corresponding to at least one known chemical constituent and at least one unknown chemical constituent; generating a chemical space surface having a distribution of weights corresponding to each of the at least one known chemical constituent therein; sampling a plurality of weighted points of the chemical space surface; providing a first set of vectors corresponding to the sampled weighted points to a trained ANN, the ANN being trained to produce a vector encoding the presence of one or more of a plurality of molecular subunits corresponding to the sampled weighted points in chemical space; for the first set of vectors, receiving from the ANN a second set of vectors, each of the second set of vectors corresponding one of the first set of vectors, each of the second set of vectors indicating the presence of one or more molecular subunits; for each of the second set of vectors, generating a plurality of strings of molecular subunits by concatenating random selections from its one or more molecular subunits; filtering the plurality of strings of molecular subunits for chemical validity; mapping each of the plurality of strings to the weighted chemical vector space; and outputting a plurality of molecular representations corresponding to the at least one unknown chemical constituent.2FH13297226.4MTV-25725

[0008] The herein disclosed subject matter is also directed to a method including receiving a plurality of tokens, each token representing a molecular subunit; generating a first plurality of molecular representations by concatenating selections from the plurality of tokens; combining the first plurality of molecular representations with a second plurality of molecular representations, thereby forming a training data set; producing, for each molecular representation of the training dataset, a corresponding feature vector and a one-hot vector, the one-hot vector encoding the presence of one or more of the plurality of molecular subunits; and training the model to map from the feature vector to the one-hot vector.

[0009] In some embodiments, the trained ANN is trained according to the above method.

[0010] In some embodiments, the ANN has an output layer comprising an activation function.

[0011] In some embodiments, the activation function may include a sigmoid function.

[0012] In some embodiments, each molecular representation of the first plurality of molecular representations contains a maximum of fifteen tokens.

[0013] In some embodiments, the second plurality of molecular representations is distinct from the molecular representations present in the first set.

[0014] In some embodiments, the at least one known chemical constituent is provided by user input.

[0015] In some embodiments, the weighting corresponds to chemicals having targeted properties of a molecule.

[0016] In some embodiments, the weighting corresponds to one or more known precursor molecules.

[0017] In some embodiments, the weight is calculated by using the value of the probability distribution function derived from the combined distributions at the targeted vector space.

[0018] In some embodiments, the outputted plurality of molecular representations further comprises an associated weight value.

[0019] In some embodiments, the distribution of weights comprises a plurality of Gaussian peaks.

[0020] In some embodiments, the method may include prior to generating the chemical space surface, automatically determining source parameters of the spectroscopic data, wherein automatically determining the source parameters comprises: querying, for each of a plurality of strong spectral signatures, a database of signatures for candidate molecules having signatures within a threshold range of each of the plurality of strong spectral signatures;3FH13297226.4MTV-25725calculating, for each candidate molecule, applicable transformations of the signatures; generating a histogram of the calculated Doppler velocities; fitting a Gaussian distribution to the histogram to determine a most-likely Doppler velocity; determining a linewidth by fitting a Gaussian profile to each of the plurality of strong spectral lines and calculating a median full- width half maximum; and determining an excitation temperature by optimizing simulated spectral intensities of a reference molecule to match observed intensities.

[0021] In some embodiments, the signatures may include rotational transitions.

[0022] In some embodiments, the applicable transformations may include a Doppler velocity required to shift the rotational transition to the observed frequency.

[0023] In some embodiments, the method may include for each spectral peak in the spectroscopic data, querying one or more spectroscopic databases for candidate molecules having rotational transitions within a threshold frequency of the spectral peak; simulating, for each candidate molecule, a rotational spectrum at the determined excitation temperature, linewidth, and Doppler velocity; calculating a spectroscopic match score for each candidate molecule based on a frequency match between the simulated rotational transition and the observed spectral peak a relative intensity match between simulated and observed spectral intensities; assigning, responsive to the spectroscopic match score exceeding a threshold, the candidate molecule as a molecular carrier of the spectral peak; and updating the chemical space surface to incorporate a Gaussian peak corresponding to the assigned candidate molecule.

[0024] In some embodiments, the methods may include calculating, for each candidate molecule, a structural relevance score by evaluating a weight of the candidate molecule on the chemical space surface, wherein the weight corresponds to a value of a probability distribution function derived from the combined Gaussian distributions at a vector representation of the candidate molecule; combining the spectroscopic match score and the structural relevance score to produce a global assignment score for the candidate molecule; applying a softmax function to the global assignment scores of all candidate molecules for a spectral peak to produce a local score for each candidate molecule; and assigning the candidate molecule to the spectral peak responsive to both the global assignment score and the local score exceeding respective threshold values.

[0025] In some embodiments, the methods may include restricting the plurality of strings of molecular subunits to include only atoms present in the at least one known chemical constituent; down-weighting molecular representations corresponding to chemically unstable4FH13297226.4MTV-25725molecules; and analyzing chemical properties of the at least one known chemical constituent, comprising a proportion of radical species and ionic species, and customizing the generated molecular representations to exhibit similar chemical properties.

[0026] In some embodiments, the methods may include following assignment of each candidate molecule to a spectral peak, re-evaluating prior assignments of candidate molecules to spectral peaks in view of an updated chemical space surface incorporating the newly assigned candidate molecule.

[0027] In some embodiments, the methods may include the trained ANN comprises a multilayer perceptron trained to map 32-dimensional dense vectors produced by a Variational Inference for Chemical Graph Autoencoder (VICGAE) embedding model to a 632-dimensional one-hot encoded vector representing the presence of SELFIES tokens; the multilayer perceptron comprises three hidden layers, each hidden layer having 256 nodes and a hyperbolic tangent activation function between hidden layers; and the output layer comprises a sigmoid activation function and a binary cross-entropy loss function.

[0028] The herein disclosed subject matter is also directed to a computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to perform any of the herein described methods.

[0029] The herein disclosed subject matter is also directed to a system including: a spectrometer; and a computing node operatively coupled to the optical spectrometer and configured to receive spectroscopic data therefrom and to perform any methods described herein.

[0030] The herein disclosed subject matter is also directed to an artificial neural network trained according to any method described herein, to, in response to input of spectroscopic data of a source, outputting a plurality of molecular representations of chemical constituents thereof.

[0031] The following Detailed Description references the accompanying drawings which form a part this application, and which show, by way of illustration, specific example implementations. Other implementations may be made without departing from the scope of the disclosure.5FH13297226.4MTV-25725BRIEF DESCRIPTION OF THE DRAWINGS

[0032] FIG. 1A is a schematic representation of the molecule generation process where the chemical vector space is represented in two dimensions for illustration purposes, according to aspects of the present disclosure. However, this method is applicable to any higher dimension embedding space.

[0033] FIG. IB is a flowchart of an exemplary method for machine learned chemical intuition for the Al generation of molecules within specific chemical spaces according to aspects of the present disclosure.

[0034] FIG. 2 is a schematic block diagram of a general-purpose computer which processes computer programs using a processing system according to aspects of the present disclosure. Computer programs on a general-purpose computer generally include an operating system and applications.

[0035] FIG. 3A is an exemplary illustration of a histogram produce by all candidate transitions within 100 MHz of the analyzed lines according to aspects of the present disclosure.

[0036] FIG. 3B is an exemplary illustration of a histogram after narrowing around the strongest peak of FIG. 3A and the Gaussian fit to the histogram in order to determine an optimal viSr value, according to aspects of the present disclosure.

[0037] FIG. 4 is an exemplary representation of a chemical space surface with Gaussian peaks centered on five abundant molecules in the IRA 16293-2422B star forming region, with the scatter points on the x-y axis are the UMAP 2-dimensional representations of all molecules detected in this source, according to aspects of the present disclosure.

[0038] FIG. 5 depicts simulated rotational spectrum of H13CCCCCN with the temperature, linewidth, visr, and column density determined in the automated fitting process (orange) overlaid on the GOTHAM observations of TMC-1, according to aspects of the present disclosure.

[0039] FIG. 6 depicts Table 1, listing the collection of assigned molecules, including 47 unique molecular species, according to aspects of the present disclosure.

[0040] FIG. 7 depicts Table 2, listing average chemical composition of the molecules assigned in the TMC-1 observations along with the 100 top-ranked generated molecular candidates according to aspects of the present disclosure.

[0041] FIGS. 8A-8B are exemplary illustrations of molecules assigned by the herein disclosed systems and methods in the GOTHAM observations of TMC-1, 11 of which6FH13297226.4MTV-25725including at least one assigned isotopologue, resulting in overall assignment of 76 molecular species, according to aspects of the present disclosure.

[0042] FIGS. 9A-9D are exemplary illustrations of the top 100 ranked candidate molecules for TMC-1 generated according to aspects of the present disclosure.DETAILED DESCRIPTION

[0043] As a preliminary matter, it will readily be understood by one having ordinary skill in the relevant art that the present disclosure has broad utility and application. As should be understood, any embodiment may incorporate only one or a plurality of the above-disclosed aspects of the disclosure and may further incorporate only one or a plurality of the abovedisclosed features. Furthermore, any embodiment discussed and identified as being “preferred” is considered to be part of a best mode contemplated for carrying out the embodiments of the present disclosure. Other embodiments also may be discussed for additional illustrative purposes in providing a full and enabling disclosure. Moreover, many embodiments, such as adaptations, variations, modifications, and equivalent arrangements, will be implicitly disclosed by the embodiments described herein and fall within the scope of the present disclosure.

[0044] Accordingly, while embodiments are described herein in detail in relation to one or more embodiments, it is to be understood that this disclosure is illustrative and exemplary of the present disclosure, and are made merely for the purposes of providing a full and enabling disclosure.

[0045] Thus, for example, any sequence(s) and / or temporal order of steps of various processes or methods that are described herein are illustrative and not restrictive.Accordingly, it should be understood that, although steps of various processes or methods may be shown and described as being in a sequence or temporal order, the steps of any such processes or methods are not limited to being carried out in any particular sequence or order, absent an indication otherwise. Indeed, the steps in such processes or methods generally may be carried out in various different sequences and orders while still falling within the scope of the present invention.

[0046] In an aspect, the present disclosure provides for an approach to use machine learning and artificial intelligence to predict new chemical species (molecules) that may be present in a chemical environment. The most important advancement is that the predicted molecules need not be previously known (by science): they are generated entirely from scratch by our7FH13297226.4MTV-25725approach. The application is to the analysis of any chemical environment of interest, e.g., ranging from pharmaceutical mixtures, to industrial waste, to environmental samples. It can also be used to generate new species similar to known ones (with direct applications in drug discovery, catalyst design, etc.). Briefly, we use a machine-learning-based chemical embedding model to generate vector representations of molecules of interest, which may include known components of a mixture or molecules with specific desired properties (physical, catalytic, therapeutic, etc.). These vectors are used to construct a chemical space surface, with weight concentrated around the input molecules, highlighting the regions of chemical space they occupy. An additional model is trained to extract structural information from the embedding method. We then sample the highly weighted points on the surface and identify the likely chemical components of molecules in these regions of chemical space. Based on this information, new molecular candidates are generated. Finally, the generated species are ranked by their weights on the chemical space surface, producing a rank-ordered list according to their similarity to the input species.

[0047] In another aspect, the present disclosure may provide for an automated method for assigning molecules in interstellar line surveys. For example, one or more algorithms may operate in two main stages to this end. First, the herein disclosed systems and methods may automatically determine key parameters of the data, including excitation temperature (Tex), line width (AV), and Doppler velocity ( isr). Next it may assign the observed spectral peaks by evaluating the spectroscopic match of the molecular candidates along with analyzing their chemical relevance to the interstellar source. The chemical relevance may be determined by leveraging machine-learning-based chemical embedding techniques to analyze the regions of chemical space occupied by the observed species. Following the line assignment, this information may then be used to generate new molecular candidates that occupy the same regions of chemical space. These newly generated species may serve as promising targets for further investigation in the observational data. In various embodiments, the herein disclosed algorithms may be validated using GOTHAM observations of TMC-1, successfully assigning the vast majority of the -400 observed transitions to the correct molecular species in approximately 15 minutes.

[0048] Referring now to FIGS. 1A-1B, a schematic representation 150 and a flowchart of a method 100 is shown. In various embodiments, method 100 may be a method for generating molecular representations corresponding to unknown chemical constituents within a chemical environment is described. The method may leverage machine learning techniques8FH13297226.4MTV-25725to predict chemical species that may be present in a chemical mixture or environment by constructing and sampling a weighted chemical space surface and employing a trained artificial neural network (ANN) to extract structural information from dense vector representations. The method may include the following steps, as further described throughout the disclosure.

[0049] With continued reference to FIG. IB, at step 105, method 100 may include reading spectroscopic data of a source. In various embodiments, the spectroscopic data may correspond to at least one known chemical constituent and at least one unknown chemical constituent. In various embodiments, the spectroscopic data may be obtained from a spectrometer operatively coupled to one or more systems as described herein. In various embodiments, the spectroscopic data may be obtained from any suitable spectroscopic technique, including but not limited to rotational spectroscopy, mass spectrometry, nuclear magnetic resonance (NMR), or other analytical techniques capable of producing data indicative of molecular species present in a source. In various embodiments, the source may be an interstellar region observed by radio telescopes, a pharmaceutical mixture, an industrial sample, an environmental sample, or any other chemical environment of interest. The spectroscopic data may contain numerous spectral features corresponding to molecular emission, absorption, or other signatures, and may include both identified and unidentified features requiring further analysis. In some embodiments, the at least one known chemical constituent may be provided by user input, such as when a scientist and / or other user has prior knowledge of certain molecules that may be present in the source.

[0050] With continued reference to FIG. IB, method 100, at step 110, may include generating a chemical space surface having a distribution of weights corresponding to each of the at least one known chemical constituent therein. In various embodiments, the chemical space surface may be constructed by creating weight distributions centered on vector representations of the known chemical constituents in a chemical vector space. In various embodiments, the vector representations may be generated using a machine-leaming-based chemical embedding model, such as a Variational Inference for Chemical Graph Autoencoder (VICGAE), configured to produce dense vector representations (e.g., 32-dimensional vectors) that encode molecular structural information including substructures, functional groups, and geometry. The distribution of weights may include a plurality of Gaussian peaks, each centered at the location in chemical vector space corresponding to a known chemical constituent. In various embodiments, the resulting distributions for each9FH13297226.4MTV-25725known constituent may then summed to generate a complete continuous surface representation of chemical space with weight concentrated around the known mixture components. In various embodiments, the weight of any point on the surface can be calculated using the value of the probability distribution function derived from the combined distributions at that point, thereby enabling the assessment of the chemical and structural relevance of any molecular candidate by evaluating its corresponding weight on the surface. In some embodiments, the weighting may correspond to chemicals having targeted properties of a molecule, or may correspond to one or more known precursor molecules.

[0051] With continued reference to FIG. IB, method 100, at step 115, may include sampling a plurality of weighted points of the chemical space surface. In various embodiments, the sampling targets highly weighted regions of the chemical space surface, which represent areas of chemical vector space that are most closely associated with the known chemical constituents and thus most likely to contain molecules with similar or related properties. By sampling these highly weighted points, the method focuses its subsequent analysis on the regions of chemical space that are most relevant to the chemical environment being studied.

[0052] With continued reference to FIG. IB, method 100 at step 120, may include providing a first set of vectors corresponding to the sampled weighted points to a trained ANN. In various embodiments, the trained ANN may be trained or pretrained to produce a vector encoding the presence of one or more of a plurality of molecular subunits corresponding to the sampled weighted points in chemical space. In various embodiments, the ANN may comprise a multilayer perceptron (MLP) that has been trained to map dense vector embeddings to interpretable molecular features. For example, the MLP may be trained to map 32-dimensional dense vectors produced by the VICGAE embedding model to a 632-dimensional one-hot encoded vector representing the presence of molecular subunits, representing the presence of SELFIES tokens. The ANN thus serves as a key to interpret which specific chemical substructures are likely present in the molecules around each sampled location of chemical space. The output layer of the ANN may have an activation function, such as sigmoid activation function for example, resulting in each dimension of the output having a value between 0 and 1, which can be interpreted as the probability that a molecule represented by a vector in a specific region of chemical space contains the chemical substructure associated with that dimension.

[0053] In various embodiments, the ANN may be trained, for example, by receiving a plurality of tokens, each token representing a molecular subunit; generating a first plurality of10FH13297226.4MTV-25725molecular representations by concatenating selections from the plurality of tokens; combining the first plurality of molecular representations with a second plurality of molecular representations, thereby forming a training data set; producing, for each molecular representation of the training dataset, a corresponding feature vector and a one-hot vector, the one-hot vector encoding the presence of one or more of the plurality of molecular subunits; and training the model to map from the feature vector to the one-hot vector.

[0054] With continued reference to FIG. IB, method 100, at step 125, may include for the first set of vectors, receiving from the ANN a second set of vectors. In various embodiments, each of the second set of vectors may correspond to one of the first set of vectors, and each of the second set of vectors may indicate the presence of one or more molecular subunits. In operation, the trained ANN predicts the chemical makeup of a molecule in the sampled region of chemical space by outputting a vector that provides insight into the likely chemical constituents. Each output vector encodes the likelihood that particular molecular subunits, such as specific SELFIES tokens representing molecular fragments (e.g., a carbon atom with a double bond, a nitrogen atom with a triple bond, an alcohol — OH group), are present in molecules occupying that region of chemical space.

[0055] With continued reference to FIG. IB, method 100, at step 130, may include for each of the second set of vectors, generating a plurality of strings of molecular subunits concatenating random selections from its one or more molecular subunits. The predicted substructures from each output vector are probabilistically sampled based on the ANN output probabilities and concatenated into full molecular representations. For example, for each sampled point, several hundred molecular SELFIES strings may be randomly generated by sampling the SELFIES tokens in a probabilistic manner based on the MLP output and concatenating the tokens into a complete molecule. In this manner, a large number of candidate molecules may be generated (e.g., greater than 106molecules) that reflect the structural characteristics predicted by the ANN for each region of chemical space.

[0056] With continued reference to FIG. IB, method 100, at step 135, may include filtering the plurality of strings of molecular subunits for chemical validity. In various embodiments, the filtering may employ built-in functionality of molecular representation packages, such as the SELFIES Python package, to ensure molecular validity. Additionally, domain knowledge may be incorporated at this stage. For instance, the predicted molecules may be restricted to include only the atoms present in the known chemical constituents, and chemically unstable molecules such as peroxides may be down-weighted. The algorithm may also analyze the11FH13297226.4MTV-25725chemical properties of the known chemical constituents, such as the proportion of radical or ionic species, and customize the generated molecules to exhibit similar characteristics. After filtering, duplicates and invalid molecules may be removed based on the applicable chemical criteria.

[0057] With continued reference to FIG. IB, method 100, at step 140, may include mapping each of the plurality of strings to the weighted chemical vector space. In various embodiments, dense vector representations of the filtered set of candidate molecules may be produced using the same embedding model (e.g., the VICGAE model) or others, that was employed to generate the chemical space surface. In various embodiments, the weight of each candidate molecule's vector on the chemical space surface is then determined. In various embodiments, this weight may represent the closeness of each generated molecule in the embedding space to the original known chemical constituents with the desired properties, and thus the highly weighted molecules may be strong candidates to exhibit similar properties. The candidate molecules may be ranked based on their weight on the chemical space surface, producing a rank-ordered list according to their similarity to the input species.

[0058] With continued reference to FIG. IB, method 100, at step 145, may include outputting a plurality of molecular representations corresponding to the at least one unknown chemical constituent. In various embodiments, the top-ranked generated molecules, having the greatest weight on the chemical space surface, are therefore the most chemically relevant to the system being analyzed and represent the most likely candidates for the unknown chemical constituents present in the source, as shown in FIG. 1A. In some embodiments, the outputted plurality of molecular representations may further include an associated weight value indicating the degree of similarity to the known chemical constituents. These topranked molecules can ultimately serve as starting points for follow-up experiments, calculations, or further investigation in the observational data to confirm their presence.

[0059] The various steps of the above method 100 will be described in greater detail below.Description of Model

[0060] In various embodiments, the process may begin with a chemical space surface that has weight centered on the coordinates corresponding to known molecules with desired properties. Thus, by generating new molecules that also occupy these highly weighted regions of chemical space, previously unknown candidates that have a strong likelihood of exhibiting these wanted properties can be created. The main challenge, however, is that the chemical space surface is typically generated based on an embedding method that creates12FH13297226.4MTV-25725dense vector representations of molecules. While these dense vector representations efficiently encode information about the molecular structure, they are not directly interpretable (i.e., each dimension does not correspond to a specific molecular substructure). As a result, in various embodiments, a model can be trained that can gain chemical insight from these vectors.

[0061] For example, and without limitation, to do this, a multilayer perceptron (MLP) can be trained to map a dense molecular feature vector to its likely chemical constituents. Firstly, the molecular substructures that will be predicted by the model may be determined. For example, this can be a list of chemical functional group s / fragments, specific sequences of DNA nucleotides, amino acids, a string representation of chemical features (e.g., SELFIES tokens), or many other embodiments of chemical descriptors. In various embodiments, the herein disclosed system and or method can randomly produce millions of molecules and subsequently create dense vector representations of these species using the embedding model that was employed to generate the chemical space surface. In various embodiments, the specific molecular substructures in each molecule may be stored as a one-hot encoded vector. For example, if a molecule has a ketone (C=O) group, the output vector will have a one at the dimension corresponding to this substructure. The model may then trained to map each dense vector to its one-hot encoded vector. The output layer of the model may have a Sigmoid activation function that results in each dimension having a value between 0 and 1. Thus, this can be interpreted as the probability that a molecule, represented by a vector in a specific region of chemical space, contains the chemical substructure associated with that dimension in the output vector. In principal, the model is learning the most-likely chemical constituents that make up the molecules in each region of chemical space and thus providing a key to interpret the dense embeddings. In various embodiments, this MLP model only needs to be trained once for a specific combination of embedding model and output features.

[0062] Once this model is trained, many highly weighted points on the chemical space surface can be sampled. The previously trained MLP then acts as a key to interpret which specific chemical substructures are likely present in the molecules around each of these locations of chemical space. By inputting each sampled point into the MLP, a vector that provides insight into the likely chemical makeup can be created. From this output vector, the predicted substructures can be probabilistically sampled and combine them into single molecules.13FH13297226.4MTV-25725

[0063] At this point, domain knowledge may be incorporated. For instance, the predicted molecules can be restricted to include only the atoms present in the input molecules.Additionally, in application to astronomical mixtures, it is known generally that unstable molecules like peroxides are very rarely identified in radio astronomical observations and can downweigh such species. The algorithm also analyzes the chemical properties of the detected molecules, such as the proportion of radical or ionic species, and customizes the generated molecules to exhibit similar characteristics.

[0064] At this point, a list of candidate molecules can be generated. The size of this list can be customized by a user, for example. For each species, the embedding model can then be implemented to produce a dense vector representation, and the weight of each vector on the chemical space surface will be determined. This weight represents the closeness of each generated molecule in the embedding space to the original inputted species with the desired properties, and thus the highly weighted molecules are strong candidates to exhibit similar properties. These top-ranked molecules can ultimately be starting points for follow-up experiments or calculations.

[0065] For the process disclosed herein, a chemical space surface conditioned on the molecules of interest can be first established. These could be species that have certain desired physical properties (e.g., melting point, vapor pressure, water solubility, size, and / or mass), functional or reactive properties (e.g., catalytic efficacy, and / or therapeutic activity), or those that are present in a chemical mixture that is being analyzed. This provides an informed understanding of the regions of chemical space occupied by the desired molecules, enabling the prediction of similar species.

[0066] To achieve this, the disclosure provides for a generative method to identify the most likely molecular structures that occupy the desired region of chemical space. Given a mathematical description of chemical space, with weights given to areas of interest, a model can be trained to sample chemical space and predict the likelihood that any given molecular subunit (e.g., a carbon atom with a double bond, a nitrogen atom with a triple bond, an alcohol -OH group) will be present in molecules that occupy that region of chemical space.

[0067] In various embodiments, possible molecules (in some implementations, millions) may be randomly generated that can be built from those subunits, weighted by their determined likelihoods. These molecules may then be filtered for chemical validity.

[0068] In various embodiments, then, the candidate molecules may be re-encoded as chemical vectors and compared to the original chemical space description to determine which14FH13297226.4MTV-25725occupy the highest weighted regions, thus representing the most likely candidates for similarity to the desired properties.Computer Modelling

[0069] Referring now to FIG. 2 a block diagram of a general-purpose computer which processes computer programs using a processing system is shown. Computer programs on a general-purpose computer generally include an operating system and applications. The operating system may be a computer program running on the computer that manages access to resources of the computer by the applications and the operating system. The resources generally include memory, storage, communication interfaces, input devices and output devices.

[0070] Examples of such general-purpose computers include, but are not limited to, larger computer systems such as server computers, database computers, desktop computers, laptop and notebook computers, as well as mobile or handheld computing devices, such as a tablet computer, handheld computer, smart phone, media player, personal data assistant, audio and / or video recorder, or wearable computing device.

[0071] With continued reference to FIG. 2, an example computer 500 comprises a processing system including at least one processing unit 502 and a memory 504. The computer can have multiple processing units 502 and multiple devices implementing the memory 504. A processing unit 502 can include one or more processing cores (not shown) that operate independently of each other. Additional co-processing units, such as graphics processing unit 520, also can be present in the computer. The memory 504 may include volatile devices (such as dynamic random-access memory (DRAM) or other random-access memory device), and non-volatile devices (such as a read-only memory, flash memory, and the like) or some combination of the two, and optionally including any memory available in a processing device. Other memory such as dedicated memory or registers also can reside in a processing unit. Such a memory configures is delineated by the dashed line 504 in Figure 1. The computer 500 may include additional storage (removable and / or non-removable) including, but not limited to, solid state devices, or magnetically recorded or optically recorded disks or tape. Such additional storage is illustrated in Figure 2 by removable storage 508 and nonremovable storage 510. The various components in Figure 1 are generally interconnected by an interconnection mechanism, such as one or more buses 530.

[0072] A computer storage medium is any medium in which data can be stored in and retrieved from addressable physical storage locations by the computer. Computer storage15FH13297226.4MTV-25725media includes volatile and nonvolatile memory devices, and removable and non-removable storage devices. Memory 504 , removable storage 508 and non-removable storage 510 are all examples of computer storage media. Some examples of computer storage media are RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optically or magneto-optically recorded storage device, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices. Computer storage media and communication media are mutually exclusive categories of media.

[0073] The computer 500 may also include communications connection(s) 512 that allow the computer to communicate with other devices over a communication medium.Communication media typically transmit computer program code, data structures, program modules or other data over a wired or wireless substance by propagating a modulated data signal such as a carrier wave or other transport mechanism over the substance. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal, thereby changing the configuration or state of the receiving device of the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media include any non-wired communication media that allows propagation of signals, such as acoustic, electromagnetic, electrical, optical, infrared, radio frequency and other signals. Communications connections 512 are devices, such as a network interface or radio transmitter, that interface with the communication media to transmit data over and receive data from signals propagated through communication media.

[0074] The communications connections can include one or more radio transmitters for telephonic communications over cellular telephone networks, and / or a wireless communication interface for wireless connection to a computer network. For example, a cellular connection, a Wi-Fi connection, a Bluetooth connection, and other connections may be present in the computer. Such connections support communication with other devices, such as to support voice or data communications.

[0075] The computer 500 may have various input device(s) 514 such as a various pointer (whether single pointer or multi-pointer) devices, such as a mouse, tablet and pen, touchpad and other touch-based input devices, stylus, image input devices, such as still and motion cameras, audio input devices, such as a microphone. The compute may have various output device(s) 516 such as a display, speakers, printers, and so on, also may be included. These devices are well known in the art and need not be discussed at length here.16FH13297226.4MTV-25725

[0076] The various storage 510, communication connections 512, output devices 516 and input devices 514 can be integrated within a housing of the computer, or can be connected through various input / output interface devices on the computer, in which case the reference numbers 510, 512, 514 and 516 can indicate either the interface for connection to a device or the device itself as the case may be.

[0077] An operating system of the computer typically includes computer programs, commonly called drivers, which manage access to the various storage 510, communication connections 512, output devices 516 and input devices 514. Such access generally includes managing inputs from and outputs to these devices. In the case of communication connections, the operating system also may include one or more computer programs for implementing communication protocols used to communicate information between computers and devices through the communication connections 512.

[0078] Any of the foregoing aspects may be embodied as a computer system, as any individual component of such a computer system, as a process performed by such a computer system or any individual component of such a computer system, or as an article of manufacture including computer storage in which computer program code is stored and which, when processed by the processing system(s) of one or more computers, configures the processing system(s) of the one or more computers to provide such a computer system or individual component of such a computer system.

[0079] Each component (which also may be called a “module” or “engine” or “computational model” or the like), of a computer system such as described herein, and which operates on one or more computers, can be implemented as computer program code processed by the processing system(s) of one or more computers. Computer program code includes computerexecutable instructions and / or computer-interpreted instructions, such as program modules, which instructions are processed by a processing system of a computer. Generally, such instructions define routines, programs, objects, components, data structures, and so on, that, when processed by a processing system, instruct the processing system to perform operations on data or configure the processor or computer to implement various components or data structures in computer storage. A data structure is defined in a computer program and specifies how data is organized in computer storage, such as in a memory device or a storage device, so that the data can accessed, manipulated, and stored by a processing system of a computer.Summary List of Algorithm Components17FH13297226.4MTV-25725

[0080] In various embodiments, the disclosure may provide for systems, methods and / or algorithms, including some or all of the following:

[0081] 1. The starting point is a chemical space surface weighted on known molecules with a desired property.

[0082] 2. A MLP model is trained to map the dense vector embeddings to interpretable molecular features, such as chemical fragments, sequences of DNA nucleotides, etc.

[0083] 3. Highly weighted points on the chemical space surface are sampled. These are inputted into the trained MLP to predict the likely chemical makeup of the molecules at the sampled points in chemical space.

[0084] 4. A large number of candidates are generated by probabilistically sampling the chemical features in the output vector of the MLP and combining them into molecules.

[0085] 5. Dense vector representations of these generated molecules are produced with the embedding model and they are ranked based on their weight on the chemical space surface. EXAMPLE

[0086] According to an exemplary implementation, as a proof of concept, we have demonstrated this using molecular representations (SELFIES) and embedding methods for vector generation (VICGAE) combined with surface generation methods, for example, as described in Int. Pub. No. WO 2025 / 178911, filed on February 19, 2025, which is incorporated by reference in its entirety herein. Below is described in detail according to various aspects of the present disclosure.

[0087] In an exemplary specific implementation for astronomical mixtures, the VICGAE embedder can be used and a model trained to predict SELFIES tokens. The VICGAE embedding model may be trained using a dictionary of 632 molecular SELFIES tokens, for example. Therefore, approximately one million molecules can be created by randomly combining any of these SELFIES tokens into a molecular SELFIES string representation. Randomly sampling SELFIES tokens to generate valid molecules led to the training set having an insufficient number of terminal heteroatoms in their unsaturated forms (e.g., double- bonded oxygen or triple-bonded nitrogen). To address this, additional molecules containing these specific SELFIES tokens were incorporated into the training data. The length of each string was limited to 15 tokens since the datasets predominately contain small molecules. In various embodiments, VICGAE vector representations may then be created of each randomly generated molecule. For every molecule, a one-hot encoded vector of length 632 can be produced that denotes which SELFIES tokens are present in the molecule. For18FH13297226.4MTV-25725example, and without limitation, if the token “[=C]” is contained in the compound, there is a value of one at the dimension corresponding to this token, otherwise the dimension has zero value. With these two vectors for each molecule, a MLP can be trained to map the 32-dimensional dense VICGAE vector to the 632 dimensional one-hot encoded vector. This MLP can now be applied to virtually any task that uses the VICGAE model and requires SELFIES token prediction, without needing retraining. The model architecture may consist of three hidden layers, each with 256 nodes, and a tanh activation function between the hidden layers. As mentioned previously, the output layer may contain a Sigmoid activation function, and a binary cross entropy loss function can be employed. The specific actions taken here (limiting the size, enhanced training for certain heteroatoms) may not be necessary for every embodiment of the technique; adjustments can be made to provide the desired output for the specific situation under consideration.

[0088] For example, using the MLP that predicts SELFIES tokens from VICGAE vectors, we will sample the tokens sequentially based on their outputted probabilities, concatenate them into a molecular SELFIES string and then ensure its molecular validity using the built in functionality of the SELFIES Python package.

[0089] In various embodiments, first, a multilayer perceptron (MLP) can be trained to predict SELFIES tokens, which represent molecular subunits, from a provided numerical vector representation of the molecule - in an illustrative case, a VICGAE vector. More specifically, the VICGAE embedding model may be trained using a dictionary of 632 molecular SELFIES tokens. Therefore, approximately one million molecules can be created by randomly combining any of these SELFIES tokens into a molecular SELFIES string representation. Randomly sampling SELFIES tokens to generate valid molecules led to the training set having an insufficient number of terminal heteroatoms in their unsaturated forms (e.g., double-bonded oxygen or triple-bonded nitrogen). To address this, additional molecules containing these specific SELFIES tokens were incorporated into the training data. The length of each string was limited to 15 tokens, as for this demonstration we wanted to limit our explorations to small molecules. We then created VICGAE vector representations of each randomly generated molecule. For every molecule, a one-hot encoded vector of length 632 can be produced that denotes which SELFIES tokens are present in the molecule. For example, if the token “[=C]” is contained in the compound, there is a value of one at the dimension corresponding to this token, otherwise the dimension has zero value.19FH13297226.4MTV-25725

[0090] With these two vectors for each molecule, a MLP may be trained to map the 32-dimensional feature vector to the 632-dimensional one-hot encoded vector. Thus, for any inputted feature vector, the model predicts which SELFIES tokens are present in the molecular species. The output layer of the model may have a sigmoid activation function, which results in every dimension of the output vector having a value between 0 and 1. This can therefore be interpreted as the probability that each SELFIES token is present in a molecule given its VICGAE vector representation. In various embodiments, the primary objective of this process is to identify the SELFIES tokens, representing molecular constituents, that are prevalent among molecules within a specific region of chemical vector space. A similar model can also be trained to predict molecular features beyond SELFIES tokens, such as distinct molecular fragments or substructures.

[0091] With this trained MLP, many of the highly weighted points on the chemical space surface can be sampled and the corresponding vectors can be input into the network. The MLP may then predict the chemical makeup of a molecule in the sampled region of chemical space. For each point several hundred molecular SELFIES strings may be randomly generated by sampling the SELFIES tokens in a probabilistic manner based on the MLP output and concatenating the tokens into a full molecule. At this point, domain knowledge of the system may be incorporated. For instance, the predicted molecules can be restricted to include only the atoms present in the input molecules. The algorithm may also analyze the chemical properties of the detected molecules, such as the proportion of radical or ionic species, and customizes the generated molecules to exhibit similar characteristics.

[0092] In various embodiments, VICGAE vector representations of this new set of molecular candidates can be produced, and finally determine which of these molecules have the greatest weight on the chemical space surface. The top ranked generated molecules are therefore the most chemically relevant to the system that is being analyzed. This process is depicted in FIG. 1A.

[0093] In various embodiments, the present disclosure may also provide for an automated method for assigning spectroscopic transitions in radio astronomical observations. In addition to analyzing the spectroscopic signals of molecular candidates, this approach may additionally or alternatively be configured to evaluate each candidate’s structural and chemical relevance within the context of the known molecular inventory. This process enables the algorithm to mimic aspects of human chemical intuition in a machine-based framework.20FH13297226.4MTV-25725

[0094] Overall, and according to various embodiments, the herein disclosed method may first automatically determine the Doppler velocity, molecular excitation temperature, and line width of the observational data and then may subsequently assign molecular carriers of the spectroscopic peaks. Finally, by investigating the regions of chemical space occupied by the assigned molecules (800), a rank-ordered list of new molecular candidates (900) may be generated based on their similarity to the detected inventory. The generated species may serve as additional candidates for potential identification in the observational data.FACTORS CONSIDERED IN LINE ASSIGNMENT

[0095] The line assignment algorithm disclosed herein was designed to closely emulate the manual assignment process. Therefore, it is first useful to discuss the factors that would be considered if a scientist were manually assigning the spectroscopic features from online databases.

[0096] The first consideration is typically the spectroscopic match of the simulated rotational catalog to the observational data. This involves simulating the rotational spectrum of the molecular candidate at the proper excitation temperature and investigating whether the frequencies and relative intensities of the rotational transitions match the observed peaks, after adjusting for any Doppler shift (i.e., visr). It is also common to check if there are any strong simulated peaks that are missing in the observational data.

[0097] However, it is possible that the spectroscopic signal can be a feasible match for a molecule that is not present in the interstellar source, especially if that molecule has a small number of strong transitions in the observed frequency region. Thus, if there are several remaining molecular candidates for a certain transition following the spectroscopic investigation, the scientist often relies on their “chemical intuition” to determine the most likely assignment. For astrochemical investigation, this reliance on chemical intuition is feasible, since the molecules detected in interstellar sources often occupy quite well-defined and homogeneous regions of chemical space. For example, the abundant detected species in warm star-forming regions such as IRAS 16293-2422B are generally quite saturated and oxygenated molecules (i.e., CH3OH, H2CO, CH3OCH3, CH3OCHO, etc.) that were likely formed on grain surfaces and subsequently sublimated into the gas phase upon the warming of the ices by the protostellar radiation. On the other hand, a cold prestellar source like TMC-1 has a molecular inventory dominated by highly unsaturated species like cyanopolyynes (HC3N, HC5N, HC9N, etc.) along with radicals and cyclic species. Thus, if one were21FH13297226.4MTV-25725assigning a spectroscopic transition observed toward TMC-1, one may be inclined to assign a certain transition to HC3N as opposed to a hot-core species such as CH3OCH2OH.AUTOMATED LINE ASSIGNMENT PROCESS

[0098] The following section will outline each step of the automated line assignment process according to various embodiments. This method may begin with determining the physical source parameters of the astronomical object, followed by assigning the spectroscopic peaks. Source Parameter Determination

[0099] In order to accurately assign the data, the Doppler velocity (yisr), excitation temperature (Tex), and linewidth (AV ) may first be determined. This is vital because if an incorrect visr is assumed, the simulated spectra will be shifted in frequency, which would make the correct line assignment nearly impossible. Furthermore, employing an incorrect Texwould alter the relative intensities of the simulated spectra, which once again would complicate the analysis. It is important to note that while the algorithm incorporates methods to automatically determine these source parameters, in various other embodiments they can also be manually provided by the user if known a priori.

[0100] For the visr determination, the method may begin with a limited list of around 100 astronomically common molecules, such as CO, NH3, CH3OH, N2FF, HC3N, etc. Then, for the strongest lines in the spectrum, the algorithm queries all transitions from this list of molecules within 100 MHz of the line frequencies. The visr that would be required to shift the rotational line to the observed frequency is then calculated and stored. Following the investigation of every strong line, each of the calculated visr values may be plotted on a histogram 310, an example of which is depicted in FIG. 3A. A Gaussian 320 is subsequently fit to the resulting histogram to find the peak value corresponding to the most-likely visr, an example of which is depicted FIG.3B. This method may be robust if the spectrum contains a sufficiently large number of peaks, as there should be a molecular candidate with the correct visr for most lines. In contrast, any other visr values derived from incorrect transitions will occur by random chance. As a result, the histogram peak centered on the correct visrshould be the largest. This process of histogram plotting and Gaussian fitting is illustrated in FIGS.3A-3B.

[0101] Next, the line width is determined by fitting a Gaussian profile to each of the lines in the spectrum and determining the resulting full- width half max (FWHM). The median of the determined linewidths is used for the remainder of the analysis.22FH13297226.4MTV-25725

[0102] Finally, to find the correct excitation temperature, we utilize the results from the visridentification. When determining the visr, the algorithm stores which molecular candidates have a transition with the correct velocity shift for each of the analyzed lines. The molecule with the greatest number of these transitions is then selected for use in the temperature determination since this provides the greatest number of datapoints to optimize the temperature value. However, the selected molecule is generally a species with an especially high abundance that may encounter optical depth issues. We therefore search for the isotopologues of the molecule in the data, since these are more likely to be optically thin. For example, since the selected molecule is HC5N in the analysis of TMC-1, the13C isotopologues are searched for in the observed spectrum. Since they are present, the H13CCCCCN isotopologue is selected for the remaining analysis. Using the optimization features of Scipy (Virtanen et al. 2020), the best-fit temperature and column density for this molecule are determined (keeping the visr and line width fixed).Line Assignment

[0103] Once the source parameters are determined, the line assignment process is initiated. First, in exemplary implementations, the molsim Python package may be used to determine the peaks in the spectrum along with the noise level of the observational data. For each peak, the CDMS and JPL rotational spectroscopy databases may be queried for all transitions that are at most one linewidth away from the center frequency. For a source with extremely narrow lines such as TMC-1, the frequency threshold may be expanded past one FWHM since the cataloged rotational spectra can have uncertainties greater than the linewidth. For each of the determined molecular candidates, the rotational spectrum is simulated using molsim at the determined Tex, AV , and visr. The resulting peak frequencies and intensities are stored.

[0104] Once all candidate molecules are determined through this database querying, the algorithm attempts to assign the lines. Each line can either be uniquely assigned, labeled as unidentified, or determined to have multiple potential carriers (possibly denoting a blended line). The three factors that are considered during the assignment process are: the frequency match of the observed peak to the catalog transition, the relative intensity match between the simulated and observed spectrum, and the structural / chemical relevance of each molecular candidate.Spectroscopic Match23FH13297226.4MTV-25725

[0105] For the frequency match, the score may be determined by a simple linear scaling factor based on the difference between the velocity-shifted catalog frequency and the observed peak. Next, the relative intensity match may be computed by first simulating the rotational spectrum of the molecular candidate to match the observed peak intensity. The algorithm may then check that no unrealistically strong transitions are predicted to be present in the spectrum (e.g., 10 times more intense than the strongest observed line) and that the other strong simulated lines of the molecule are present in the observational data around their expected intensity. Fairly lenient intensity thresholds may be required to account for potential uncertainties in optical depth, molecular excitation temperature, or rotational catalogs.

[0106] If the molecular candidate is isotopically substituted or vibrationally excited, the algorithm may then perform further analysis. For example, if the source is extremely cold, a molecule will be excluded if it is vibrationally excited, as there would likely be little to no population in that state.Structural Relevance Determination

[0107] Each molecule may also be scored by its structural / chemical relevance based on the other molecules observed in the source.

[0108] As previously mentioned, when manually assigning these line surveys, scientists often have a preconceived notion of a molecule’s likelihood of being present in the source, based on its similarity to the known chemical inventory. However, this is typically based on their general “chemical intuition,” which is only honed over years of experience. Therefore, one of the overarching goals of the herein disclosed systems and methods is to teach a computer to replicate this chemical intuition in order to improve the accuracy of automated chemical mixture assignment.

[0109] For this process, various embodiments rely on machine learning-based chemical embedding methods. These techniques create numerical vector representations of molecules that encode molecular information such as substructures, functional groups, and geometry. By representing a molecule as a vector, it is mapped to a specific location in chemical vector space. Similar vector representations can be generated for molecules that are chemically or structurally comparable, and these molecules will therefore be nearby each other in this chemical vector space. By identifying the regions of chemical space occupied by known mixture components, we can assess the chemical relevance of each new molecular candidate to the mixture.24FH13297226.4MTV-25725

[0110] Previous approaches to this problem relied on a graph-based architecture. This graph consisted of approximately 300,000 molecules, with each molecule corresponding to a unique graph node. Molecules were then connected via bidirectional edges if their vector representations were sufficiently similar. The graph functioned as a ranking system where all initial weight was assigned to the known mixture components. Through an iterative process, weight was then transferred through the edge connections, with the transferred weight being diminished during each additional step. Therefore, molecules that were directly connected to the known mixture components (and there- fore close in chemical vector space) received a greater amount of weight than the molecules separated by several edge connections. The resulting weight was utilized to assess the chemical relevance of each molecule in the graph to the mixture and was subsequently applied as an additional heuristic in the line assignment process.

[0111] Although this method performed well across various chemical mixtures, there remained opportunities for further improvement. These were mainly based on the non-continuous chemical space surface that was generated by the graph-based architecture. For instance, the distance threshold for an edge connection (i.e., how similar two molecules need to be in chemical vector space to be connected by an edge) placed certain constraints and limitations on the graph. To illustrate this, if the threshold value for an edge connection was set to 12 (meaning that two molecules were connected if their vector representations had a Euclidean distance less than 12), a molecule that was nearly identical to a known mixture component would have the same local connectivity as a molecule with a distance of 11.99. Moreover, a molecule that had a distance of 12.01 from a known mixture component would be ranked notably lower than the molecule with a distance of 11.99. These non-continuous edge constraints could potentially lead to unwanted behavior in the rankings. Additionally, constructing a specific molecular graph restricted our analysis to the molecules within it. Consequently, for entirely new applications involving a significantly different set of molecules, the graph would need to be regenerated from scratch.

[0112] To improve this, instead of a graph architecture, a continuous chemical space surface with Gaussian weight centered on the vector representations of the known mixture components may be created. More specifically, for every known chemical mixture component, a Gaussian peak is centered at that location in chemical vector space. The resulting Gaussian surfaces for each known mixture component are then summed to generate a complete continuous surface representation of chemical space with weight distributed25FH13297226.4MTV-25725around the known mixture components. The weight of any point on the surface can then be calculated using the value of the probability distribution function derived from the combined Gaussian distributions at that point. Thus, the chemical / structural relevance of any molecular candidate in the mixture can be assessed by evaluating the weight of the surface at its corresponding point. The parameters of the Gaussian peaks (such as the width of each Gaussian) were determined via a hyperparameter search using two separate astronomical datasets. The optimal values were determined in conjunction with several other hyperparameters, such as the global line assignment score, which will be discussed later.

[0113] The surface no longer deals with the issues of noncontinuity, since molecules that are closer to the peak of each Gaussian will have a higher ranking. Moreover, the architecture is not constrained to a predetermined molecular graph since any point in chemical vector space can be sampled. Finally, while the graph-based ranking system required around 50 seconds to converge with optimal hyperparameters, this updated technique only takes a few seconds for the score of any molecule to be determined.

[0114] A 3-dimensional illustration 410 of this process is depicted in FIG. 4. Here, the plot 410 has been limited to three dimensions for illustration purposes; however, this method is extendable to any higher-dimensional vector space. For example, the work presented herein employs the VICGAE embedding method which produces dense 32-dimensional molecular vectors. In the exemplary representation 410 shown in FIG. 4, all molecules detected toward the star forming region IRAS 16293-2422B are displayed as points in a scatter plot on the x-y axis. A Gaussian peak has been centered on five abundant molecules in this source (namely OCS, H2CO, CH3OCH3, CH3CHO, H2CCO). As observed, most species in the scatter plot are already positioned in regions of high weight, even when conditioned on a small set of molecules.

[0115] At the beginning of the line assignment process, for example, a user can choose to input a set of precursor molecules based on prior knowledge of the source. For example, one could input methanol and formaldehyde for a star forming region. In this case, the chemical space surface is initialized with weight around these molecules. On the other hand, if they select not to input any precursor molecules, the first line is assigned solely using the scores from the spectroscopic match. The assigned molecule is then the first species which is given a peak on the chemical space surface.

[0116] In various embodiments, the algorithm may operate in a sequential manner, in which the lines are assigned from strongest to weakest. After each new molecule is assigned, the26FH13297226.4MTV-25725weight surface is updated to incorporate this new information, and the score of any molecule can be recalculated. Since new information is gleaned with each new molecular assignment (and the chemical space surface becomes more informative), the previous assignments are all re-checked once a new molecule is identified.

[0117] Furthermore, the algorithm accounts for the possibility of structurally unique molecules being present in any interstellar source by enabling an “override” of the structural relevance scoring when sufficient spectroscopic evidence is available. For example, specifically, if a molecule has at least three lines in the data with a near-perfect spectroscopic match, and the structural relevance metric is the only factor reducing its molecular score, the molecule may be assigned to those lines. The species will then be added to the list of “detected” molecules, and a peak will be marked at its position on the chemical space surface. This approach allows the algorithm to explore diverse regions of chemical space, avoiding constraints to a few narrow clusters.

[0118] For every molecular candidate of each line, a global score may be tabulated by multiplying the frequency, intensity, and structural relevance scores. A softmax function may then be applied to the global scores of all candidates for that line, generating a local score for each molecule relative to the other potential candidates. A molecule is confidently assigned to a line if its global and local scores exceed their respective optimized threshold values. As mentioned previously, these threshold values were determined in conjunction with various other hyperparameters. If no molecule meets the global score threshold for a line, the transition is marked as unidentified. Conversely, if multiple molecules exceed the global score threshold but their scores are close enough that none surpass the local score threshold, the line is noted as having multiple possible carriers.Molecular Prediction

[0119] Following the assignment of each line, the chemical space surface that incorporates all of the assigned molecules in the source can be established. This provides an informed understanding of the regions of chemical space occupied by the chemical inventory, enabling the prediction of other molecules that may be present. These molecular candidates can serve as starting points for identifying the molecular carriers of the unassigned transitions.

[0120] In the previous graph-based approach, once the line assignment was completed, the graph was conditioned on all assigned molecules and identified the highest-ranked unassigned molecular species within the graph. This unfortunately limited consideration to molecules in the graph. Thus, since the computational efficiency was inversely proportional27FH13297226.4MTV-25725to the size of the graph, a graph size was selected that contained a large enough number of molecular candidates without being prohibitively slow. With the herein disclosed approach, the graph architecture is no longer limiting and it can therefore feasibly recommend any valid molecule.

[0121] To do this, since there is no longer a predetermined list of molecules to select from, a generative method may be implemented to identify the most likely molecular structures. This requires creating an approach to extract specific structural information from dense numerical molecular vectors. To do this, we trained a multilayer perceptron (MLP) to predict SELFIES tokens from an inputted VICGAE vector. More specifically, the VICGAE embedding model was originally trained using a dictionary of 632 molecular SELFIES tokens. Therefore, approximately one million molecules can be created by randomly combining any of these SELFIES tokens into a molecular SELFIES string representation. Randomly sampling SELFIES tokens to generate valid molecules led to the training set having an insufficient number of terminal heteroatoms in their unsaturated forms (e.g., double-bonded oxygen or triple-bonded nitrogen). To address this, additional molecules containing these specific SELFIES tokens were incorporated into the training data. The length of each string was limited to 15 tokens since our datasets predominately contain small molecules. VICGAE vector representations can then be created of each randomly generated molecule. For every molecule, a one-hot encoded vector of length 632 can be produced that denotes which SELFIES tokens are present in the molecule. For example, if the token “[=C]” is contained in the compound, there is a value of one at the dimension corresponding to this token, otherwise the dimension has zero value.

[0122] With these two vectors for each molecule, an MLP can be trained to map the 32-dimensional feature vector to the 632 dimensional one-hot encoded vector. Thus, for any inputted feature vector, the model predicts which SELFIES tokens are present in the molecular species. The output layer of the model has a sigmoid activation function, which results in every dimension of the output vector having a value between 0 and 1. This can therefore be interpreted as the probability that each SELFIES token is present in a molecule given its VICGAE vector representation. The primary objective of this process is to identify the SELFIES tokens, representing molecular constituents, that are prevalent among molecules within a specific region of chemical vector space. A similar model could also be trained to predict molecular features beyond SELFIES tokens, such as distinct molecular fragments or substructures.28FH13297226.4MTV-25725

[0123] With this trained MLP, many of the highly weighted points on the chemical space surface can be sampled and input the corresponding vectors into the network. The MLP will then predict the chemical makeup of a molecule in the sampled region of chemical space. For each point several hundred molecular SELFIES strings can be randomly generated by sampling the SELFIES tokens in a probabilistic manner based on the MLP output and concatenating the tokens into a full molecule. At this point, domain knowledge of the mixture can be incorporated. For instance, the predicted molecules can be restricted to include only the atoms present in the input molecules. Additionally, since unstable molecules like peroxides are very rarely identified in radio astronomical observations and can down-weight such species. The algorithm may also analyze the chemical properties of the detected molecules, such as the proportion of radical or ionic species, and customizes the generated molecules to exhibit similar characteristics.

[0124] In various embodiments, VICGAE vector representations of this new set of molecular candidates can then be produced, and finally determine which of these molecules have the greatest weight on the chemical space surface. The top ranked generated molecules are therefore the most chemically relevant to the mixture and are used as starting points in the identification of the unassigned transitions. This process is schematically depicted in FIG. 1A.

[0125] One or more computers can be used to implement such a computational pipeline, using one or more general-purpose computers, such as client devices including mobile devices and client computers, one or more server computers, or one or more database computers, or combinations of any two or more of these, which can be programmed to implement the functionality such as described in the example implementations. An example of which is schematically depicted and described with reference to FIG. 2 herein.RESULTS

[0126] In various embodiments, this algorithm may be tested on the GOTHAM observations of the cold molecular cloud TMC-1. For this exemplary embodiment investigation, the data ranging from 18 to 36.4 GHz can be focused on. The algorithm analyzed 438 lines at 5 significance.

[0127] The determined visr, linewidth, and excitation temperature were 5.83 km s \ 0.395 km s-1, and 6.41 K, respectively. These values all closely match the accepted numbers in this source, as the excitation temperature is generally 5-10 K for most molecules with linewidths around 0.3 km s-1and a visrof approximately 5.8 km s-1. Moreover, FIG. 5 shows the29FH13297226.4MTV-25725rotational spectrum 551 of H13CCCCCN simulated at the determined parameter values overlaid on top of the GOTHAM observations. There is clearly a very high level of overlap between the simulated and observed spectrum, thus suggesting that the determined values are sufficiently accurate.

[0128] Of the 438 lines, the algorithm uniquely assigned 422 to a single molecular carrier, three were listed as having several possible carriers, and 13 were unassigned. The collection of assigned molecules is listed in Table 1 displayed for convenience in FIG. 6. As can be seen, 47 unique molecular species were identified. All of these molecules have been previously detected toward TMC- 1, thus suggesting that there were no “false positive” assignments of molecules that are not truly present in the data. For 11 of these species, one or more isotopologues were also assigned. Therefore, in all, 79 molecules / isotopologues / isotopomers were identified in the data at 5 significance.

[0129] Furthermore, in nine instances, a molecule that is not present in the data was only ruled out due to a low structural relevance score. For these molecules, the spectroscopic analysis did not provide sufficient evidence to dismiss them. This highlights the importance of incorporating structural relevance in the assignment algorithm, as the addition of these nine misassignments would significantly impair the algorithm’s performance and lead to considerably less useful results.

[0130] The algorithm took a total of 15.5 minutes to complete on a laptop (Apple M2 Pro chip, 16 GB memory, 10 core CPU). Of this, 2.91 minutes were spent on source parameter determination, 11.16 minutes on catalog scraping and spectral simulation, and 1.19 minutes on line assignment. The remaining time was dedicated to overhead tasks, such as importing packages and uploading the necessary datasets.Molecular Prediction Results

[0131] When conditioned on the list of assigned species in the TMC-1 dataset, the predictive method produced 100,000 molecules. However, after removing duplicates and invalid molecules based on the provided chemical criteria, 2119 generated species remained. The top 100 ranked molecular candidates 900 are displayed in FIGS. 9A-9D. This predictive process took approximately four minutes following the line assignment. By visually comparing FIGS. 8A-8B with FIGS. 9A-9D, it is evident that the algorithm is generating predominately unsaturated carbon-chain molecules that are quite similar to the detected species. A more quantitative analysis of the chemical characteristics of the detected and top 100 ranked generated molecules is displayed in Table 2 conveniently depicted in FIG. 7. As30FH13297226.4MTV-25725can be seen, the top-ranked predicted molecules are on average almost identical in size to the observed molecular species. They are also similar in saturation, as both datasets are dominated by alkyne-containing species. Like the observed species, the predicted molecules also have a greater proportion of nitrogen than oxygen and sulfur. An especially notable result is that the top ranked generated molecule (CH2CHCCH) has been detected in TMC-1, however it does not correspond to a 5 line in our analyzed dataset, thus preventing its assignment. Several other of the top- 100 generated molecules displayed in FIGS. 9A-9D have also been detected in TMC-1, including CH3CH2CCH , HCCO, H2CCCHCCH, HCCCHS, and HCCCH2CCH. If the analysis is expanded to the top 300 ranked generated molecules, several other detected species in TMC- 1 are generated, namely CH3OH, HCCCH2CN, H2CO, HOCN, CH2CHCH3, CH2CHCHO, CH2CCHC4H, CH3CHCO, H2CCS, H2C3S, and HCSCN. This indicates that the generative method can effectively produce molecular candidates pertinent to the interstellar source.

[0132] Out of the 100,000 molecules generated by the algorithm (including duplicates), 14,610 matched molecules identified in TMC-1 during the line assignment process. This is a remarkably large percentage, as the method is capable of feasibly generating any valid molecule. The high percentage of generated species that correspond to the detected molecules suggests that the MLP effectively predicts the correct SELFIES tokens for the sampled regions of chemical space. Furthermore, this once again indicates that the method is generating molecules that are highly relevant to the interstellar source.

[0133] In various other embodiments, one or models may be trained to predict functional groups or substructures instead of SELFIES tokens. Therefore, the MLP could predict the presence of a CN or CHO group instead of needing to sequentially generate the specific SELFIES tokens that comprise them.GENERALIZATION OF STRUCTURAL RELEVANCE ANALYSIS

[0134] In this paper, the performance of the structural relevance metric and molecule generation technique on the chemical inventory of an interstellar source can be analyzed. That being said, the approach is broadly applicable across various domains. The primary aim of this methodology is to identify regions of chemical space occupied by molecules with specific properties. By targeting these regions, we can automatically generate new molecules that share those properties.

[0135] For astrochemists, the properties of interest are the abundance and detectability of molecules within a given interstellar region. However, the same approach could be extended31FH13297226.4MTV-25725to other fields, such as pharmaceutical science, by conditioning the chemical space surface on molecules with desired chemical or pharmaceutical properties. Without any additional knowledge, the method can generate previously unexplored molecules nearby the input species in chemical vector space, making them promising candidates for exhibiting similar properties. These predicted molecules could then be further investigated through calculations or experiments to assess their efficacy.

[0136] Additionally, the VICGAE embedding method was utilized as exemplary embodiment in the present disclosure due to its success on prior astrochemical studies.However, the framework is flexible and can accommodate other embedding models optimized for different applications. Adapting the method would only require retraining the MLP and fine-tuning the parameters of the chemical space surface, such as the widths of the Gaussian peaks.CONCLUSION

[0137] As discussed herein, the present disclosure provides for methods for automatically assigning molecular species in radioastronomical observations. The approach begins by determining key source parameters, including excitation temperature, linewidth, and visr. It then assigns spectral lines by querying online spectroscopic databases for molecules with known rotational transitions near the observed peak frequencies. For each molecular candidate, spectral simulations are performed, for example, using the molsim Python package to evaluate the match between the rotational catalog and the observational data. Additionally, the chemical and structural relevance of each molecule to the interstellar source is assessed by examining its position within the regions of chemical space occupied by the other assigned species. After assigning each of the identified lines, new molecular candidates are generated from the highly weighted regions of chemical space, offering potential starting points for identifying the carriers of remaining unassigned transitions.

[0138] The performance of the algorithm was demonstrated by testing it on the GOTHAM observations of TMC-1. The determined excitation temperature, linewidth, and visrclosely aligned with accepted literature values for this source. The algorithm assigned 422 of the 438 identified lines to 47 distinct molecules (79 species including isotopologues), all of which have been previously detected in TMC-1. The entire assignment process was completed in approximately 15 minutes. Furthermore, the top-ranked generated molecules closely matched the chemical characteristics of TMC-1’ s detected inventory, even including some additional known molecules in the source.32FH13297226.4MTV-25725

[0139] In another aspect of the present disclosure, the herein disclosed subject matter may provide for a method including, providing a chemical space surface weighted on known molecules with a desired property; training a multilayer perceptron (MLP) model to map dense vector embeddings to interpretable molecular features (e.g., chemical fragments, sequences of DNA nucleotides); sampling highly weighted points on the chemical space surface, wherein the highly weighted points are inputted into the trained MLP to predict the likely chemical makeup of the molecules at the sampled points in chemical space; probabilistically sampling the chemical features in the output vector of the MLP and combining the chemical features to generate a large number of candidates molecules (e.g., >106molecules); producing dense vector representations of these candidate molecules with the MLP model, wherein these representations are ranked based on their weight on the chemical space surface; and defining an algorithm for identifying the candidate molecules having desired functional properties.

[0140] In another aspect, the present disclosure provides for a computer system for generating a database comprising candidate molecules generated by the method disclosed herein, the computer system including a processing system; computer storage accessible to the processing system, and computer program instructions encoded on the computer storage, wherein when the computer program instructions are processed by the processing system, the computer system is configured to: define data structures in the computer storage representing chemical features of the candidate molecules; and execute a training program applied to the data structures to generating a database comprising candidate molecules, wherein the candidate molecules are ranked by possessing the desired functional properties.

[0141] In another aspect, the present disclosure provides for a computer program product comprising computer storage and computer program instructions encoded on the computer storage, wherein the computer program instructions, when processed by a processing system of a computer, causes the computer to perform the methods disclosed herein or implement the computer system disclosed herein.

[0142] In another aspect, the present disclosure provides for a method comprising: receiving a plurality of tokens, each token representing a molecular subunit; generating a first plurality of molecular representations by concatenating random selections from the plurality of tokens; combining the first plurality of molecular representations with a second plurality of molecular representations, thereby forming a training data set; producing, for each molecular representation of the training dataset, a corresponding feature vector and a one-hot vector, the33FH13297226.4MTV-25725one-hot vector encoding the presence of one or more of the plurality of molecular subunits; and training the model to map from the feature vector to the one-hot vector.

[0143] In another aspect, the present disclosure provides for a method comprising: reading spectroscopic data of a source, the spectroscopic data corresponding to a plurality of known chemical constituents and at least one unknown chemical constituent; generating a chemical space surface having a plurality of Gaussian peaks corresponding to each of the plurality of known chemical constituents therein; sampling a plurality of weighted points of the chemical space surface; providing a first set of vectors corresponding to the sampled weighted points to a trained ANN, the ANN being trained to produce a vector encoding the presence of one or more of a plurality of molecular subunits corresponding to the sampled weighted points in chemical space; for the first set of vectors, receiving from the ANN a second set of vectors, each of the second set of vectors corresponding one of the first set of vectors, each of the second set of vectors indicating the presence of one or more molecular subunits; for each of the second set of vectors, generating a plurality of strings of molecular subunits by concatenating random selections from its one or more molecular subunits; filtering the plurality of strings of molecular subunits for chemical validity; mapping each of the plurality of strings to the weighted chemical vector space; and outputting a plurality of molecular representations corresponding to the at least one unknown chemical constituent.

[0144] In some embodiments, the trained ANN is trained according to the methods described herein. In some embodiments, the ANN having an output layer comprising a sigmoid activation function. In some embodiments, each molecular representation of the first plurality of molecular representations contains a maximum of fifteen tokens. In some embodiments, the second plurality of molecular representations is distinct from the molecular representations present in the first set. In some embodiments, the plurality of known chemical constituents is provided by user input. In some embodiments, the weighting corresponds to chemicals having targeted properties of a molecule. In some embodiments, the weighting corresponds to one or more known precursor molecules. In some embodiments, the weight is calculated by using the value of the probability distribution function derived from the combined Gaussian distributions at the targeted vector space. In some embodiments, the outputted plurality of molecular representations further comprises an associated weight value.34FH13297226.4MTV-25725

[0145] It should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific implementations described above. The specific implementations described above are disclosed as examples only.35FH13297226.4

Claims

MTV-25725CLAIMSWhat is claimed is:

1. A method comprising:reading spectroscopic data of a source, the spectroscopic data corresponding to at least one known chemical constituent and at least one unknown chemical constituent;generating a chemical space surface having a distribution of weights corresponding to each of the at least one known chemical constituent therein;sampling a plurality of weighted points of the chemical space surface; providing a first set of vectors corresponding to the sampled weighted points to a trained ANN, the ANN being trained to produce a vector encoding the presence of one or more of a plurality of molecular subunits corresponding to the sampled weighted points in chemical space;for the first set of vectors, receiving from the ANN a second set of vectors, each of the second set of vectors corresponding one of the first set of vectors, each of the second set of vectors indicating the presence of one or more molecular subunits;for each of the second set of vectors, generating a plurality of strings of molecular subunits by concatenating random selections from its one or more molecular subunits;filtering the plurality of strings of molecular subunits for chemical validity; mapping each of the plurality of strings to the weighted chemical vector space; and outputting a plurality of molecular representations corresponding to the at least one unknown chemical constituent.

2. A method comprising:receiving a plurality of tokens, each token representing a molecular subunit; generating a first plurality of molecular representations by concatenating selections from the plurality of tokens;combining the first plurality of molecular representations with a second plurality of molecular representations, thereby forming a training data set;36FH13297226.4MTV-25725producing, for each molecular representation of the training dataset, a corresponding feature vector and a one-hot vector, the one-hot vector encoding the presence of one or more of the plurality of molecular subunits; andtraining the model to map from the feature vector to the one-hot vector.

3. The method of claim 1, wherein the trained ANN is trained according to the method of claim 2.

4. The method of any one of claims 1-3, the ANN having an output layer comprising an activation function.

5. The method of claim 4, wherein the activation function comprises a sigmoid function.

6. The method of claim 2, wherein each molecular representation of the first plurality of molecular representations contains a maximum of fifteen tokens.

7. The method of claim 2, wherein the second plurality of molecular representations is distinct from the molecular representations present in the first set.

8. The method of any one of claims 1 or 3-7, wherein the at least one known chemical constituent is provided by user input.

9. The method of any one of claims 1 or 3-8, wherein the weighting corresponds to chemicals having targeted properties of a molecule.

10. The method of any one of claims 1 or 3-9, wherein the weighting corresponds to one or more known precursor molecules.

11. The method of any one of claims 1 or 3-10, wherein the weight is calculated by using the value of the probability distribution function derived from the combined distributions at the targeted vector space.

12. The method of claim 11, wherein the outputted plurality of molecular representations further comprises an associated weight value.

13. The method of any one of claims 1 or 3-12, wherein the distribution of weights comprises a plurality of Gaussian peaks.37FH13297226.4MTV-2572514. The method of any one of claims 1 or 3-13, further comprising:prior to generating the chemical space surface, automatically determining source parameters of the spectroscopic data, wherein automatically determining the source parameters comprises:querying, for each of a plurality of strong spectral signatures, a database of signatures for candidate molecules having signatures within a threshold range of each of the plurality of strong spectral signatures;calculating, for each candidate molecule, applicable transformations of the signatures;generating a histogram of the calculated Doppler velocities;fitting a Gaussian distribution to the histogram to determine a most-likely Doppler velocity;determining a linewidth by fitting a Gaussian profile to each of the plurality of strong spectral lines and calculating a median full-width half maximum; and determining an excitation temperature by optimizing simulated spectral intensities of a reference molecule to match observed intensities.

15. The method of claim 14, wherein the signatures comprise rotational transitions.

16. The method of claim 14, wherein the applicable transformations comprise a Doppler velocity required to shift the rotational transition to the observed frequency.

17. The method of any one of claims 1 or 3-13, further comprising:for each spectral peak in the spectroscopic data, querying one or more spectroscopic databases for candidate molecules having rotational transitions within a threshold frequency of the spectral peak;simulating, for each candidate molecule, a rotational spectrum at the determined excitation temperature, linewidth, and Doppler velocity;calculating a spectroscopic match score for each candidate molecule based on a frequency match between the simulated rotational transition and the observed spectral peak a relative intensity match between simulated and observed spectral intensities;38FH13297226.4MTV-25725assigning, responsive to the spectroscopic match score exceeding a threshold, the candidate molecule as a molecular carrier of the spectral peak; andupdating the chemical space surface to incorporate a Gaussian peak corresponding to the assigned candidate molecule.

18. The method of claim 17, further comprising:calculating, for each candidate molecule, a structural relevance score by evaluating a weight of the candidate molecule on the chemical space surface, wherein the weight corresponds to a value of a probability distribution function derived from the combined Gaussian distributions at a vector representation of the candidate molecule;combining the spectroscopic match score and the structural relevance score to produce a global assignment score for the candidate molecule;applying a softmax function to the global assignment scores of all candidate molecules for a spectral peak to produce a local score for each candidate molecule; and assigning the candidate molecule to the spectral peak responsive to both the global assignment score and the local score exceeding respective threshold values.

19. The method of claim 1, further comprising:restricting the plurality of strings of molecular subunits to include only atoms present in the at least one known chemical constituent;down- weighting molecular representations corresponding to chemically unstable molecules; andanalyzing chemical properties of the at least one known chemical constituent, comprising a proportion of radical species and ionic species, and customizing the generated molecular representations to exhibit similar chemical properties.

20. The method of claim 14, further comprising:following assignment of each candidate molecule to a spectral peak, re-evaluating prior assignments of candidate molecules to spectral peaks in view of an updated chemical space surface incorporating the newly assigned candidate molecule.

21. The method of claim 1, wherein:39FH13297226.4MTV-25725the trained ANN comprises a multilayer perceptron trained to map 32-dimensional dense vectors produced by a Variational Inference for Chemical Graph Autoencoder (VICGAE) embedding model to a 632-dimensional one-hot encoded vector representing the presence of SELFIES tokens;the multilayer perceptron comprises three hidden layers, each hidden layer having 256 nodes and a hyperbolic tangent activation function between hidden layers; andthe output layer comprises a sigmoid activation function and a binary cross-entropy loss function.

22. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to perform a method according to any one of claims 1-17.

23. A system comprising:a spectrometer; anda computing node operatively coupled to the optical spectrometer and configured to receive spectroscopic data therefrom and to perform a method according to any one of claims 1-17.

24. An artificial neural network trained according to the method of claim 2 to, in response to input of spectroscopic data of a source, outputting a plurality of molecular representations of chemical constituents thereof.40FH13297226.4