Interpretable machine learning system and method for efficient discovery of high-performance materials

An interpretable machine learning system using a multi-objective model with molecular descriptors effectively addresses the inefficiencies in identifying high-performance materials by predicting key properties, achieving a high success rate in selecting suitable candidates.

WO2025170921A1PCT designated stage Publication Date: 2025-08-14OHIO STATE INNOVATION FOUND

Patent Information

Application Number
PCT/US2025/014473
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-05
Filing Date
2025-02-04
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing methods for identifying high-performance materials, such as organic electrode materials (OEMs), are inefficient and rely heavily on trial-and-error due to a vast and complex design space, making it difficult to optimize multiple properties like solubility, specific energy, and synthetic accessibility simultaneously.

Method used

An interpretable machine learning system using a multi-objective model with molecular descriptors, employing algorithms like SISSO to identify a subset of key descriptors, predicts solubility, specific energy, and synthetic accessibility, enabling the efficient selection of high-performance materials.

Benefits of technology

The system achieves a 62.9% success rate in identifying materials with viable charge-discharge cycling profiles, significantly higher than human intuition, and efficiently narrows down suitable candidates from a large molecular dataset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025014473_14082025_PF_FP_ABST
    Figure US2025014473_14082025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are systems and methods for identifying a target molecule by receiving a molecule dataset from a library of molecules comprising chemical-structural representations of candidate molecules for a material application; determining one or more target molecular property values derived from a set of predictive molecular descriptor values determined from a candidate molecule of the molecular dataset applied to a trained machine learning model, wherein the trained ML model is configured to output the set of predictive molecular descriptor values for a given molecule data input, and wherein the trained ML model was trained for a set of training data and the set of molecular descriptors generated from a molecular descriptor modeling application; and outputting, via a report or to a data store, the one or more target molecular property values for each of the molecular dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.103361-617WO1 T2024-122 INTERPRETABLE MACHINE LEARNING SYSTEM AND METHOD FOR EFFICIENT DISCOVERY OF HIGH-PERFORMANCE MATERIALS GOVERNMENT SUPPORT CLAUSE

[0001] This invention was made with government support under Grant No. 2124604 awarded by the National Science Foundation. The Government has certain rights in the invention. CROSS REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of priority to United States Provisional Application No.63 / 549,799, filed February 5, 2024, entitled “INTERPRETABLE MACHINE LEARNING SYSTEM AND METHOD FOR EFFICIENT DISCOVERY OF HIGH- PERFORMANCE MATERIALS,” which is incorporated by reference herein in its entirety. SUMMARY

[0003] An exemplary system and method are disclosed that employ interpretable machine learning to efficiently identify high-performance materials for electrodes or other applications using interpretable molecular descriptors derived from training for an interpretable multi- objective model that an optimizable, e.g., for solubility, specific energy, synthetic accessibility, or other metrics described herein. The multi-objective model comprehensively considers the risk, reward, and cost aspects of materials development, for example, by predicting solvation energy, redox potential, and symmetry-adapted synthesis accessibility as one set of metrics, to provide high precision estimation of such metrics for a target molecule.

[0004] A study was conducted that demonstrated a 62.9% success rate of identifying materials for high-performance electrodes with a viable charge-discharge cycling profile, which was observed to be at least 20.8% higher than human intuition. The study provides a baseline for further enhancements.

[0005] In an aspect, provided is a method for identifying a target molecule, the method comprising: receiving a molecule dataset from a library of molecules, the molecule dataset comprising chemical-structural representations (e.g., a Simplified Molecular Input Line Entry System (SMILES) representations) of candidate molecules for a material application (e.g., organic electrode materials (OEMs)), determining one or more target molecular property values (e.g., associated with solubility, specific energy, and synthetic accessibility) derived from a set of predictive molecular descriptor values determined from a candidate molecule ofAttorney Docket No.103361-617WO1 T2024-122 the molecular dataset applied to a trained machine learning (ML) model, wherein the trained ML model is configured to output the set of predictive molecular descriptor values for a given molecule data input, and wherein the trained ML model was trained for a set of training data and the set of down-selected molecular descriptors generated from a modeling application (e.g., Mordred descriptor package); and outputting, via a report or to a data store, the one or more target molecular property values for each of the molecular dataset, wherein the one or more determined target molecular property is used for efficient identification of a set of molecules from library of molecules for a material application.

[0006] In some aspects, the one or more target molecular property values are used in a multi-objective model (e.g., associated with solubility, specific energy, and synthetic accessibility) that evaluates a given model for two or more target molecular properties, wherein the multi-objective model is used to select the set of molecules from library of molecules for a material application.

[0007] In some aspects, the one or more target molecular property values include a first target molecular property value, wherein the first target molecular property value is determined from a first target molecular property model (e.g., linear equation) that employs two or more molecular descriptor values from the complete set of molecular descriptor values.

[0008] In some aspects, the one or more target molecular property values include redox potential, specific energy, synthesizability (e.g., symmetric-adapted synthetic accessibility), or combinations thereof.

[0009] In some aspects, a small set of molecular descriptors is determined from a significantly larger set of molecular descriptors using a sure independence screening (SIS) algorithm.

[0010] In some aspects, the SIS algorithm is a sure independence screening and sparsifying operation (SISSO) algorithm.

[0011] In some aspects, the set of molecular descriptors comprises 5 or more (e.g., 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more) distinct molecular descriptors.

[0012] In some aspects, the molecular descriptors comprise a combination of zero- dimensional descriptors, one-dimensional descriptors, two-dimensional descriptors, and three- dimensional descriptors.

[0013] In some aspects, the set of molecular descriptors is determined using PaDEL Descriptor, Mordred, Blue Desc, ChemoPy, PyDPI, Dragon, molecular operating environment (MOE), ChemAxon, Schrodinger, KNIME, or a combination thereof.Attorney Docket No.103361-617WO1 T2024-122

[0014] In some aspects, the predictive molecular descriptor values comprise one or more molecular properties selected from the group consisting of: electronegativity, molecular weight, number of nitrogen atoms, sigma electrons, polarizability, and combinations thereof.

[0015] In some aspects, the one or more target molecular property values are represented as linear expressions of the predictive molecular descriptor values.

[0016] In some aspects, the library of molecules is generated using a fast assembly of SMILES Fragments (FASMIFRA) method (e.g., from a filtered seed dataset).

[0017] In some aspects, the method further includes synthesizing a compound based on the one or more target molecular properties for the material application.

[0018] In another aspect, provided is a compound synthesized from a molecule identified for a material application according to any one of the disclosed methods for identifying a target molecule.

[0019] In another aspect, provided is a non-transitory computer-readable medium having instructions stored thereon, wherein the instructions, when executed by one or more processors, cause the one or more processors to perform and any one of the disclosed methods for identifying a target molecule.

[0020] In another aspect, provided is a system (e.g., molecule screening system) comprising: one or more processors; and a memory having instructions stored thereon, wherein the instructions, when executed by the one or more processors, cause the one or more processors to perform any one of the disclosed methods for identifying a target molecule.

[0021] In another aspect, provided is a method of training a machine learning (ML) model, comprising: generating (e.g., via a molecular descriptor modeling application) a set of molecular descriptors for a given molecule data input; applying a screening algorithm (e.g., sure independence screening and sparsifying operation (SISSO) algorithm) to identify a sub- set of molecular descriptors associated with a target molecular property; generating (e.g., via a molecular descriptor modeling application) a training data set of molecular descriptors for a corresponding set of molecules; and applying the training data set of molecular descriptors and corresponding set of molecules.

[0022] In some aspects, the training employs features from any one of the disclosed methods for identifying a target molecule.

[0023] In another aspect, provided is an ML training system comprising: one or more processors; and a memory having instructions stored thereon, wherein the instructions, when executed by the one or more processors, cause the one or more processors to perform any one of the disclosed methods of training a machine learning model.Attorney Docket No.103361-617WO1 T2024-122

[0024] In another aspect, provided is a non-transitory computer-readable medium having instructions stored thereon, wherein the instructions, when executed by one or more processors, cause the one or more processors to perform any one of the disclosed methods of training a machine learning model.

[0025] Other systems, methods, features and / or advantages will be or may become apparent to one with skill in the art upon examination of the following drawings and detailed description. It is intended that all such additional systems, methods, features and / or advantages be included within this description and be protected by the accompanying claims. DESCRIPTION OF DRAWINGS

[0026] The skilled person in the art will understand that the drawings described below are for illustration purposes only.

[0027] FIG.1 shows an exemplary system employing an interpretable machine learning to efficiently identify high-performance materials for electrodes or other applications using interpretable molecular descriptors.

[0028] FIG.2 shows an exemplary machine-learning-based workflow for discovery of high-performance OEMs. (left to right) (1) The model starts with an initial set of OEM candidates that are fed to the proposed SPARKLE (symbolic predictive algorithm for recognizing key molecular elements) method including three parts, (2) a generative AI that expands the initial molecules to a much larger set of candidates, (3) an easy-to-calculate synthetic cost measure, and (4) a machine learning algorithm that uses small amounts of DFT data to learn simple symbolic models for specific energy (e) and solubility (ΔGsol). (5) Multiobjective optimization is used to discover OEM candidates’ trade-off between specific energy, solubility, and cost. (6) The top-performing OEM candidates are synthesized, (7) fabricated into aqueous zinc-ion batteries (AZIBs), and (8) tested under over 500 charge−discharge cycles to measure realistic device-level performance metrics as a final screening step.

[0029] FIG.3 shows an exemplary method for identifying a target molecule for a material application according to one example.

[0030] FIG.4 shows an exemplary method for training a machine learning model according to one example.

[0031] FIG.5 shows the training and testing performance of SPARKLE versus discovered and conventional machine learning models for specific energy prediction; the training performance (top) and testing performance (bottom) achieved by (Panel A)Attorney Docket No.103361-617WO1 T2024-122 SPARKLE, (Panel B) NN, (Panel C) RF, and (Panel D) LASSO linear regression models for specific energy. The training set included only 100 paraquinones, while the testing set had over 103,000 quinones. It was shown that some models achieve a good training coefficient for determination (R2) and root mean squared error; however, SPARKLE is the only one that achieved good prediction performance on the test set. Since the test set is more than 1000 times the size of the training set and includes many non-paraquinones, this demonstrates the strong generalization properties afforded by SPARKLE (much more robust to overfitting than conventional black-box machine learning methods).

[0032] FIG.6.13C NMR (101 MHz, DMSO-d6) spectra of 2,3,5,6-tetrakis-1H- pyrazol-1-yl-[1,4]benzoquinone (2). A red dot indicated the NMR solvent peak (δ 39.99 – DMSO-d6).

[0033] FIG.7A.1H NMR (400 MHz, CDCl3) spectra of 2,3-bis(benzotriazol-1-yl)- 1,4-naphthoquinone (4). The peaks of the product and NMR solvent (δ 7.26 – CDCl3) were indicated by blue and red dots, respectively.

[0034] FIG.7B.13C NMR (101 MHz, CDCl3) spectra of 2,3-bis(benzotriazol-1-yl)- 1,4-naphthoquinone (4). A red dot indicated the NMR solvent peak (δ 77.02 – CDCl3).

[0035] FIG.8A.1H NMR (400 MHz, DMSO-d6) spectra of 3,7-di-p-tolyl-1H,5H-4,8- dioxa-1,2,5,6-tetraaza-anthraquinone (8). The peaks of the product, NMR solvent (δ 2.51 – DMSO-d6), and impurities (δ 5.75 – DCM, δ 3.40 – H2O, δ 1.25-0.8 – hexane) were indicated by blue, red, and yellow dots, respectively.

[0036] FIG.8B.13C NMR (101 MHz, DMSO-d6) spectra of 3,7-di-p-tolyl-1H,5H- 4,8-dioxa-1,2,5,6-tetraaza-anthraquinone (8). NMR solvent peak (δ 39.99 – DMSO-d6) was indicated by red dot.

[0037] FIG.9.1H NMR (400 MHz, CDCl3) spectra of 1,4- bis(phenylamino)anthracene-9,10-dione (9). The peaks of the product, NMR solvent (δ 7.26 – CDCl3), and impurities (δ 1.77 – H2O) were indicated by blue, red, and yellow dots, respectively.

[0038] FIG.10A.1H NMR (400 MHz, CDCl3) spectra of 3,6-dichloro-1,10- phenanthroline-5,6-dione (15). The peaks of the product, NMR solvent (δ 7.26 – CDCl3), and impurities (δ 1.77 – H2O, δ 5.29 – DCM) were indicated by blue, red, and yellow dots, respectively.

[0039] FIG.10B.13C NMR (101 MHz, CDCl3) spectra of 3,6-dichloro-1,10- phenanthroline-5,6-dione (15). A red dot indicates the NMR solvent peak (δ 77.02 – CDCl3).Attorney Docket No.103361-617WO1 T2024-122

[0040] FIG.11A.1H NMR (400 MHz, DMSO-d6) spectra of 1,4-dihydro- benzo[g]quinoxaline-2,3,5,10-tetraone (17). The peaks of the product, NMR solvent (δ 2.51 – DMSO-d6), and impurities (δ 3.40 – H2O) were indicated by blue, red, and yellow dots, respectively.

[0041] FIG.11B.13C NMR (400 MHz, DMSO-d6) spectra of 1,4-dihydro- benzo[g]quinoxaline-2,3,5,10-tetraone (17). A red dot indicated the NMR solvent peak (δ 39.99 – DMSO-d6).

[0042] FIG.12A.1H NMR (400 MHz, DMSO-d6) spectra of 2,5-bis(pyrazol-1'-yl)- 1,4-dihydroxybenzene (18). The peaks of the product, NMR solvent (δ 2.51 – DMSO-d6), and impurities (δ 3.40 – H2O) were indicated by blue, red, and yellow dots, respectively.

[0043] FIG.12B.13C NMR (400 MHz, DMSO-d6) spectra of 2,5-bis(pyrazol-1'-yl)- 1,4-dihydroxybenzene (18). A red dot indicated the NMR solvent peak (δ 39.99 – DMSO-d6).

[0044] FIG.13A.1H NMR (400 MHz, DMSO-d6) spectra of Tris(4-(1H-imidazol-1- yl)phenyl)amine (21). The peaks of the product, NMR solvent (δ 2.51 – DMSO-d6), and impurities (δ 3.40 – H2O, δ 3.45, δ 1.01 – EtOH, δ 5.74 – DCM) were indicated by blue, red, and yellow dots, respectively.

[0045] FIG.13B.13C NMR (400 MHz, DMSO-d6) spectra of Tris(4-(1H-imidazol-1- yl)phenyl)amine (21). A red dot indicated the NMR solvent peak (δ 39.99 – DMSO-d6).

[0046] FIG.14A.1H NMR (400 MHz, C6D6) spectra of 5,10-diphenyl-5,10- dihydrophenazine (23). The peaks of the product, NMR sol-vent (δ 7.16 – C6D6), and impurities (δ 2.10 – toluene, δ 0.40 – H2O) were indicated by blue, red, and yellow dots, respectively.

[0047] FIG.14B.13C NMR (101 MHz, C6D6) spectra of 5,10-diphenyl-5,10- dihydrophenazine (23). A red dot indicates the NMR solvent peak (δ 128.4 – C6D6).

[0048] FIG.15A.1H NMR (400 MHz, C6D6) spectra of 5,10-di(4-methoxyphenyl)- 5,10-dihydrophenazine (24). The peaks of the product, NMR solvent (δ 7.16 – C6D6), and impurities (δ 4.27 – DCM, δ 0.40 – H2O) were indicated by blue, red, and yellow dots, respectively.

[0049] FIG.15B.13C NMR (101 MHz, C6D6) spectra of 5,10-di(4-methoxyphenyl)- 5,10-dihydrophenazine (24). A red dot indicates the NMR solvent peak (δ 128.4 – C6D6).

[0050] FIG.16A.1H NMR (400 MHz, CDCl3) spectra of 9-(4-methoxyphenyl)-9H- carbazole (26). The peaks of the product, NMR solvent (δ 7.26 – CDCl3), and impurities (δ 3.78 – DME, δ 1.55 – H2O) were indicated by blue, red, and yellow dots, respectively.Attorney Docket No.103361-617WO1 T2024-122

[0051] FIG.16B.13C NMR (101 MHz, CDCl3) spectra of 9-(4-methoxyphenyl)-9H- carbazole (26). A red dot indicates the NMR solvent peak (δ 77.02 – CDCl3).

[0052] FIG.17.1H NMR (400 MHz, CDCl3) spectra of 9-(4-chlorophenyl)-9H- carbazole (27). The peaks of the product, NMR solvent (δ 7.26 – CDCl3), and impurities (δ 1.55 – H2O) were indicated by blue, red, and yellow dots, respectively.

[0053] FIG.18. Training of ML model for predicting solvation energy and redox potentials and validation.

[0054] FIG.19 shows a three-dimensional visualization of OEM design space considering the specific energy (reward), solvation energy (risk), and SASA (cost). Plot shows initial 20,000 candidates.

[0055] FIG.20A-20B show deployment of SPARKLE to predict specific energy (reward), solvation energy (risk), and SASA (cost) of OEMs. (FIG.20A) The plot shows initial 60,000 candidates from the 600 k design space. Two histograms indicate the distribution of the number of OEM candidates for hard-to-synthesize molecules with SASA < 2.5 (gray) and easy-to-synthesize molecules with SASA2.5 (blue). The histogram for SASA ≥ 2.5 was magnified by ten times for visibility. (FIG.20B) A plot of twenty-seven final OEM candidates selected for experimental testing. The rectangles indicate quinone- based OEMs while diamonds show non-quinone OEMs recommended by SPARKLE prediction.

[0056] FIG.21. Selected OEMs identified by SPARKLE and corresponding galvanostatic charge−discharge (GCD) profiles at 0.5 C charge−discharge rate. (Panel A) GCD profiles of eight representative n-type OEMs (2−5, 10, 11, 12, and 15). (Panel B) GCD profiles of four representative p-type OEMs (19, 20, 23, and 26).

[0057] FIG.22 shows deployment of symbolic and interpretable models to more than 600 k small molecules to identify new OEMs. The red squares are selected molecules for experimental testing.

[0058] FIG.23 shows a comparison of OEMs identified by SPARKLE and common reported OEMs. (A) Comparison of the specific energy of OEMs identified by SPARKLE compared to existing OEMs reported in the literature. (B) Comparison of the estimated cost per specific energy of selected OEMs. The cost estimation is based on the price of synthesis reagents, solvents, and reaction yields. (C). Long-term cycling stability of new OEMs identified by SPARKLE at 2 C.Attorney Docket No.103361-617WO1 T2024-122

[0059] FIG.24. Defined chemically equivalent point of 3-phenyl-2-(1H-tetrazol-5- yl)naphtho[2,3-b]furan-4,9-dione as an example. UA and CA indicate the number of unique atoms and a number of sets of chemically equivalent atoms, respectively.

[0060] FIG.25. Verification of SASA for symmetricity determination by using 350 molecule pool and several exampled molecules. DETAILED DESCRIPTION

[0061] It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate aspects, can also be provided in combination with a single aspect. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single aspect, can also be provided separately or in any suitable subcombination. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure.

[0062] Some references are cited in a reference list and discussed in the disclosure provided herein. The citation and / or discussion of such references is provided merely to clarify the description of the disclosed technology and is not an admission that any such reference is “prior art” to any aspects of the disclosed technology described herein. In terms of notation, “[n]” or the subscript n corresponds to the nth reference in a list. All references cited and discussed in this specification are incorporated herein by reference. DEFINITIONS

[0063] In this specification and in the claims that follow, reference will be made to a number of terms, which shall be defined to have the following meanings:

[0064] Throughout the description and claims of this specification, the word “comprise” and other forms of the word, such as “comprising” and “comprises,” means including but not limited to, and are not intended to exclude, for example, other additives, segments, integers, or steps. Furthermore, it is to be understood that the terms comprise, comprising, and comprises as they relate to various aspects, elements, and features of the disclosed invention also include the more limited aspects of “consisting essentially of” and “consisting of.”

[0065] As used herein, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to a “processor”Attorney Docket No.103361-617WO1 T2024-122 includes aspects having two or more such processors unless the context clearly indicates otherwise.

[0066] Ranges can be expressed herein as from “about” one particular value and / or to “about” another particular value. When such a range is expressed, another aspect includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms another aspect. It should be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint.

[0067] As used herein, the terms “optional” or “optionally” mean that the subsequently described event or circumstance may or may not occur and that the description includes instances where said event or circumstance occurs and instances where it does not.

[0068] For the terms “for example” and “such as,” and grammatical equivalences thereof, the phrase “and without limitation” is understood to follow unless explicitly stated otherwise. EXAMPLE METHOD AND SYSTEM

[0069] FIG. 1 shows an exemplary system 100 that employs interpretable machine learning 102 to efficiently identify high-performance materials for electrodes or other applications using interpretable molecular descriptors derived from training for an interpretable multi-objective model that an optimizable, e.g., for solubility, specific energy, synthetic accessibility, or other metrics described herein. The multi-objective model comprehensively considers the risk, reward, and cost aspects of materials development, for example, by predicting solvation energy, redox potential, and symmetry-adapted synthesis accessibility as one set of metrics, to provide high precision estimation of such metrics for a target molecule.

[0070] In the example shown in Fig. 1, the system 100 includes a trained ML model 102 configured for use in one or more inference systems 101 that each receives a set of candidate molecules 104 from a library of molecules 105 to generate one or more predictive molecular descriptors 106 for that set of candidate molecules 104. The system 100 includes a target molecular property model 108 that receives as its input the predictive molecular descriptor 106 to generate a set of target molecular properties 109 to be used as input in a multi-objective model 110 to select, from the candidate molecules 106’, a selected molecule based on the target molecular properties 112 as the molecular of interest.

[0071] The trained ML model 102 (shown as 102’) is generated, via a training operation 113, from a training system 114 configured to train a set of training molecules 116 using a set of molecular descriptors 118 for the training molecules 116. In the example shown in FIG.1,Attorney Docket No.103361-617WO1 T2024-122 2-10’s of molecular descriptors 118 can be, for example, employed for the training 113 and were selected from training 119. The descriptor training 119, for descriptor down selection, employs a simulation software 120 (shown as “Molecular descriptor modeling” 120) that can generate a set of molecular descriptors 122 from a library of training molecules 124. The training 119 includes a descriptor screening module 126 configured to receive the generated / simulated set of molecular descriptors 122 to identify / down select the molecular descriptors 118 to be used for the inference model training 113.

[0072] The multi-objective model 110 is configured to evaluate a given model 108 for two or more target molecular properties. Examples include solubility, specific energy, and synthetic accessibility that provides a definition for a high-performance material (e.g., for selection of OEMs).

[0073] The target molecular property model includes a transfer function (e.g., a linear equation) that employs two or more predictive molecular descriptor values 106 to generate the target molecular property values 109. Examples of the target molecular property values 109 include redox potential, specific energy, synthesizability (e.g., symmetric-adapted synthetic accessibility), or combinations thereof.

[0074] The molecular descriptor screening model 126 is configured to down-select a set of molecular descriptors 122 from the larger set of molecular descriptors. In some embodiments, the sure independence screening (SIS) algorithm or the sure independence screening and sparsifying operation (SISSO) algorithm can be employed. The down-selected set of molecular descriptors 118, from the screen operation 119, can include 5 or more distinct molecular descriptors (e.g., 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more) for use in the generation of the inference ML model 102. The molecular descriptors 118 may include a combination of zero-dimensional descriptors, one-dimensional descriptors, two- dimensional descriptors, and three-dimensional descriptors. Exemplary molecular descriptors include, for example, molecular weight, number of nitrogen atoms, sigma electrons, and polarizability.

[0075] The simulation software 120 can provide molecular descriptor modeling. Examples include PaDEL Descriptor, Mordred, Blue Desc, ChemoPy, PyDPI, Dragon, molecular operating environment (MOE), ChemAxon, Schrodinger, KNIME, or a combination thereof.

[0076] Indeed, the method of generating the inference model 114 may include generating (e.g., via a molecular descriptor modeling application) a set of molecular descriptors for a given molecule data input; applying a screening algorithm (e.g., sure independence screening and sparsifying operation (SISSO) algorithm) to identify a subset of molecular descriptorsAttorney Docket No.103361-617WO1 T2024-122 associated with a target molecular property; generating (e.g., via a molecular descriptor modeling application) a training data set of molecular descriptors for a corresponding set of molecules; and applying the training data set of molecular descriptors and corresponding set of molecules.

[0077] FIG.3 shows an exemplary method 300 for identifying a target molecule. Method 300 includes receiving a molecule dataset (302) from a library of molecules. In various aspects, the molecule dataset includes chemical-structural representations, such as Simplified Molecular Input Line Entry System (SMILES) representations, of candidate molecules for a material application (e.g., organic electrode materials (OEMs)). In various aspects, the library of molecules is generated using a fast assembly of SMILES Fragments (FASMIFRA) method (e.g., from a filtered seed dataset).

[0078] Method 300 further includes determining one or more target molecular property values (304) derived from a set of predictive molecular descriptor values determined from a candidate molecule of the molecular dataset applied to a trained machine learning (ML) model. The predictive molecular descriptor values can include properties of a prospective material that may be most relevant to a particular material application, e.g., solubility, specific energy, and synthetic accessibility. The trained ML model is configured to output the set of predictive molecular descriptor values for a given molecule data input and is trained from a set of training data and the set of molecular descriptors (e.g., down selected set of descriptors) generated from a molecular descriptor modeling application (e.g., Mordred descriptor package). In some aspects, the set of molecular descriptors comprises 5 or more (e.g., 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more) distinct molecular descriptors. In some aspects, the molecular descriptors comprise a combination of zero-dimensional descriptors, one-dimensional descriptors, two-dimensional descriptors, and three-dimensional descriptors. In some aspects, the set of molecular descriptors is determined using PaDEL Descriptor, Mordred, Blue Desc, ChemoPy, PyDPI, Dragon, molecular operating environment (MOE), ChemAxon, Schrodinger, KNIME, or a combination thereof. In some aspects, the predictive molecular descriptor values comprise one or more molecular properties selected from the group consisting of: electronegativity, molecular weight, number of nitrogen atoms, sigma electrons, polarizability, and combinations thereof.

[0079] Method 300 also includes outputting (306), via a report or to a data store, the one or more target molecular property values for each of the molecular dataset. The one or more determined target molecular property can be used for efficient identification of a set of molecules from the library of molecules for the material application.Attorney Docket No.103361-617WO1 T2024-122

[0080] The target molecular property values can then be used to evaluate (308) prospective candidate molecules for a particular material application. In some aspects, the one or more target molecular property values are used in a multi-objective model (e.g., associated with solubility, specific energy, and synthetic accessibility) that evaluates a given model for two or more target molecular properties, wherein the multi-objective model is used to select the set of molecules from a library of molecules for a material application. In some aspects, the one or more target molecular property values include a first target molecular property value, wherein the first target molecular property value is determined from a first target molecular property model (e.g., linear equation) that employs two or more predictive molecular descriptor values of the set of predictive molecular descriptor values. In some aspects, the one or more target molecular property values include redox potential, specific energy, synthesizability (e.g., symmetric-adapted synthetic accessibility), or combinations thereof. In some aspects, the one or more target molecular property values are represented as linear expressions of the predictive molecular descriptor values. After a candidate molecule has been evaluated for a material application, the compound can be synthesized. In some aspects, disclosed are compounds synthesized from a molecule identified for a material application according to any of the described methods.

[0081] In some aspects, the set of molecular descriptors (e.g., a down-selected set of descriptors) is determined from a larger set of molecular descriptors using a sure independence screening (SIS) algorithm. In some aspects, the SIS algorithm is a sure independence screening and sparsifying operation (SISSO) algorithm.

[0082] Referring now to FIG. 4, a method 400 is shown for training a machine learning model according to one aspect of the present disclosure.

[0083] Method 400 includes generating (402) a set of molecular descriptors for a given molecule data input. In some aspects, the set of molecular descriptors is generated via a molecular descriptor modeling application.

[0084] Method 400 further includes applying (404) a screening algorithm, such as a sure independence screening and sparsifying operation (SISSO) algorithm, to identify a subset of molecular descriptors associated with a target molecular property.

[0085] Method 400 also includes generating (406) a training data set of molecular descriptors for a corresponding set of molecules. The training data set of molecular descriptors can, in some aspects, be generated via a molecular descriptor modeling application.

[0086] Method 400 also includes training (408) a machine learning model using the training data set of molecular descriptors and corresponding sets of molecules.Attorney Docket No.103361-617WO1 T2024-122

[0087] Machine Learning. The exemplary system and method can be implemented using one or more artificial intelligence and machine learning operations. The term “artificial intelligence” can include any technique that enables one or more computing devices or computing systems (i.e., a machine) to mimic human intelligence. Artificial intelligence (AI) includes but is not limited to knowledge bases, machine learning, representation learning, and deep learning. The term “machine learning” is defined herein to be a subset of AI that enables a machine to acquire knowledge by extracting patterns from raw data. Machine learning techniques include, but are not limited to, logistic regression, support vector machines (SVMs), decision trees, Naïve Bayes classifiers, and artificial neural networks. The term “representation learning” is defined herein to be a subset of machine learning that enables a machine to automatically discover representations needed for feature detection, prediction, or classification from raw data. Representation learning techniques include, but are not limited to, autoencoders and embeddings. The term “deep learning” is defined herein to be a subset of machine learning that enables a machine to automatically discover representations needed for feature detection, prediction, classification, etc., using layers of processing. Deep learning techniques include but are not limited to artificial neural networks or multilayer perceptron (MLP).

[0088] Machine learning models include supervised, semi-supervised, and unsupervised learning models. In a supervised learning model, the model learns a function that maps an input (also known as feature or features) to an output (also known as target) during training with a labeled data set (or dataset). In an unsupervised learning model, the algorithm discovers patterns among data. In a semi-supervised model, the model learns a function that maps an input (also known as a feature or features) to an output (also known as a target) during training with both labeled and unlabeled data.

[0089] Neural Networks. An artificial neural network (ANN) is a computing system including a plurality of interconnected neurons (e.g., also referred to as “nodes”). This disclosure contemplates that the nodes can be implemented using a computing device (e.g., a processing unit and memory as described herein). The nodes can be arranged in a plurality of layers such as input layer, an output layer, and optionally one or more hidden layers with different activation functions. An ANN having hidden layers can be referred to as a deep neural network or multilayer perceptron (MLP). Each node is connected to one or more other nodes in the ANN. For example, each layer is made of a plurality of nodes, where each node is connected to all nodes in the previous layer. The nodes in a given layer are not interconnected with one another, i.e., the nodes in a given layer function independently of one another. AsAttorney Docket No.103361-617WO1 T2024-122 used herein, nodes in the input layer receive data from outside of the ANN, nodes in the hidden layer(s) modify the data between the input and output layers, and nodes in the output layer provide the results. Each node is configured to receive an input, implement an activation function (e.g., binary step, linear, sigmoid, tanh, or rectified linear unit (ReLU) function), and provide an output in accordance with the activation function. Additionally, each node is associated with a respective weight. ANNs are trained with a dataset to maximize or minimize an objective function. In some implementations, the objective function is a cost function, which is a measure of the ANN’s performance (e.g., an error such as L1 or L2 loss) during training, and the training algorithm tunes the node weights and / or bias to minimize the cost function. This disclosure contemplates that any algorithm that finds the maximum or minimum of the objective function can be used for training the ANN. Training algorithms for ANNs include but are not limited to backpropagation. It should be understood that an artificial neural network is provided only as an example machine learning model. This disclosure contemplates that the machine learning model can be any supervised learning model, semi-supervised learning model, or unsupervised learning model. Optionally, the machine learning model is a deep learning model. Machine learning models are known in the art and are therefore not described in further detail herein.

[0090] A convolutional neural network (CNN) is a type of deep neural network that has been applied, for example, to image analysis applications. Unlike traditional neural networks, each layer in a CNN has a plurality of nodes arranged in three dimensions (width, height, and depth). CNNs can include different types of layers, e.g., convolutional, pooling, and fully- connected (also referred to herein as “dense”) layers. A convolutional layer includes a set of filters and performs the bulk of the computations. A pooling layer is optionally inserted between convolutional layers to reduce the computational power and / or control overfitting (e.g., by down sampling). A fully-connected layer includes neurons, where each neuron is connected to all of the neurons in the previous layer. The layers are stacked similarly to traditional neural networks. GCNNs are CNNs that have been adapted to work on structured datasets such as graphs.

[0091] As used herein, the terms “loss function” or “loss model” refer to a function that indicates loss errors. As mentioned above, in some embodiments, a machine-learning algorithm can repetitively train to minimize overall loss. In some embodiments, the personalized fashion generation system employs multiple loss functions and minimizes overall loss between multiple networks and models. Examples of loss functions include a softmax classifier function (with cross-entropy loss), a hinge loss function, and a least squares loss function.Attorney Docket No.103361-617WO1 T2024-122

[0092] Other Supervised Learning Models. A logistic regression (LR) classifier is a supervised classification model that uses the logistic function to predict the probability of a target, which can be used for classification. LR classifiers are trained with a data set (also referred to herein as a “dataset”) to maximize or minimize an objective function, for example, a measure of the LR classifier’s performance (e.g., an error such as L1 or L2 loss), during training. This disclosure contemplates that any algorithm that finds the minimum of the cost function can be used. LR classifiers are known in the art and are therefore not described in further detail herein.

[0093] A Naïve Bayes’ (NB) classifier is a supervised classification model that is based on Bayes’ Theorem, which assumes independence among features (i.e., the presence of one feature in a class is unrelated to the presence of any other features). NB classifiers are trained with a data set by computing the conditional probability distribution of each feature given a label and applying Bayes’ Theorem to compute the conditional probability distribution of a label given an observation. NB classifiers are known in the art and are therefore not described in further detail herein.

[0094] A k-NN classifier is an unsupervised classification model that classifies new data points based on similarity measures (e.g., distance functions). The k-NN classifiers are trained with a data set (also referred to herein as a “dataset”) to maximize or minimize a measure of the k-NN classifier’s performance during training. This disclosure contemplates any algorithm that finds the maximum or minimum. The k-NN classifiers are known in the art and are therefore not described in further detail herein.

[0095] A majority voting ensemble is a meta-classifier that combines a plurality of machine learning classifiers for classification via majority voting. In other words, the majority voting ensemble’s final prediction (e.g., class label) is the one predicted most frequently by the member classification models. The majority voting ensembles are known in the art and are therefore not described in further detail herein. EXPERIMENTAL RESULTS AND ADDITIONAL EXAMPLES Example 1: Symbolic Predictive Algorithm for Recognizing Key Molecular Elements (SPARKLE)

[0096] Organic electrode materials (OEMs) are promising energy storage materials in rechargeable batteries due to their high structural diversity and synthetic tunability.1–3In practice, however, the virtually unlimited design space of OEM makes it challenging to identify suitable candidates with optimal stability, solubility, and redox potentials. Modifications of functional group identity and substitution patterns can result in substantial changes in multipleAttorney Docket No.103361-617WO1 T2024-122 correlated factors, including intermolecular interactions, electronic structures, and solid-state packing, which drastically impact the performance metrics of OEMs in non-intuitive fashions. As a result, the identification of new OEMs still largely depends on trial-and-error synthesis and electrochemical testing, which are inherently costly and time-intensive.

[0097] To address this issue, quantum chemical calculation and cheminformatic approaches have been developed to predict physicochemical properties of OEM candidates, such as redox potentials,6–9solubility,6,9–12, and stability.13–16Despite this progress, the feedback loop of model development-prediction-testing has yet to be realized due to the multi- objective nature of OEM discovery. Namely, the optimization of one property is often at the tradeoff of another. As a result, identifying candidates with viable traits requires extensive density functional theory (DFT) calculations, which are too computationally expensive to screen the entire design space.

[0098] Recently, there has been a shift in interest from traditional first principles models to computationally cheaper data-driven modeling paradigms.4,5Such methods can find complex nonlinear relationships among a relatively large number of variables, establishing non-intuitive mappings between molecular features and material properties.17Many model structures, such as neural networks,18,19random forests,20,21Gaussian processes,22and support vector machines23have been investigated in the context of modeling molecular properties. They eliminate the need for domain expertise, and Edisonian approaches to fitting models to data. As a tradeoff, however, they typically provide limited insights into the underlying physics and require vast amounts of data for training.

[0099] Herein, a study was conducted to delineate a new model designed with human interpretability in mind, which is referred to as Symbolic Predictive Algorithm for Recognizing Key Molecular Elements (SPARKLE). From around 1800 molecular descriptors,25the study used a sure independence screening and sparsifying operator (SISSO)24to identify a simple molecular property equation with just eight descriptors to define the property of unknown OEMs. With just 115 molecules, the study learned an accurate model of reduction potential, which generalizes beyond the class of molecules used in training to predict the reduction potential of over 100,000 held-out quinone molecules. Applying this method, the study built a multi-objective model that evaluated the risk (solubility), reward (specific energy), and cost (symmetry adapted synthetic accessibility) of unknown OEMs (FIG. 2). Deployment of this model toward a design space with over 600,000 organic molecules allowed the study to identify ~5000 synthesizable OEMs, of which 27 OEMs with exceptional tradeoff among risk, reward, and cost were selected for experimental testing. This workflow not only saves time and costAttorney Docket No.103361-617WO1 T2024-122 but has a higher success rate of 62.9% (17 of 27) at identifying OEMs with a viable galvanostatic charge-discharge profile. It was observed that the success rate is three times higher than the collective success rate of OEMs designed by human intuition over a six-year period in a laboratory (20.8%). Using this model, many stable OEMs that have comparable performance to state-of-the-art cathodes in AZIB at a much lower synthetic cost were identified.

[0100] Inverting the ML model revealed new hypotheses regarding the structure required for successful OEMs. Combining the ML model and domain expertise, an OEM that features the high energy density in AZIB was developed. The success of these new OEMs also revealed previously unknown insights into the design of stable OEMs. Specifically, the incorporation of halogen facilitates stability, potentially via halogen-pi interactions.

[0101] Study Generation of 600k design space:

[0102] Filtering available data: The study used the 140k molecules from Tabor, D. P.; Gómez-Bombarelli, R.; Tong, L.; Gordon, R. G.; Aziz, M. J.; Aspuru-Guzik, A. Mapping the Frontiers of Quinone Stability in Aqueous Media: Implications for Organic Aqueous Redox Flow Batteries. J. Mater. Chem. A 2019, 7, 12833–12841 as the initial molecular search space, which is herein incorporated by reference in its entirety. The model was designed to predict the two-proton reduction of quinones. The original design space was filtered to identify a set of unique molecules that contain the fully oxidized para- and meta-quinones, totaling 103k quinones.

[0103] Generating new molecules: The 103k filtered quinones were expanded upon using the FASMIFRA molecular generator from Berenger, F., Tsuda, K. Molecular generation by Fast Assembly of (Deep)SMILES fragments. J Cheminform 13, 88 (2021) to build a dataset including over 670,000 molecules. The FASMIFRA generator can produce out-of-sample molecules by creating sets of substructures from a given dataset and identifying valid linking points between the substructures. As a result, generated molecules are structurally similar to the 103k “seeding” quinones fed to the generator. The resulting new design space also contains molecules outside of the quinone classes due to being constructed from common substructures.

[0104] Regarding SASA: SCScore is a widely used metric for assessing synthesizability. However, it does not explicitly account for symmetry, a crucial consideration in the pursuit of new OEMs. Consequently, the SCScore was modified by incorporating a component that specifically addresses molecular symmetry. This modification was introduced as the Symmetry-Adapted Synthetic Accessibility score (SASA).Attorney Docket No.103361-617WO1 T2024-122

[0105] The number of chemically unique atoms (UA) over chemically equivalent atoms (CA) is an easy and quantifiable gauge of asymmetry. For example, a perfectly symmetric molecule does not have unique atoms, and thus, UA / CA = 0. In the contrary, a high UA / CA value indicates a high degree of molecular asymmetry, suggesting that the molecule may pose challenges during preparation. This is due to the need for precise and sequential synthetic steps to incorporate multiple distinct functional groups.

[0106] Similarly, the SCScore evaluated how the complexity of the synthetic pathway by evaluating what aspect. However, without proper normalization, the UA / CA value, which spans from 0 to infinity, is likely to dominate the SCScore, which falls within the range of 1 to 5. Consequently, the study employed the reciprocal of the SCScore (1 / SCScore, ranging from 0 to infinity) was employed. This modification allows larger scores to indicate greater ease of synthesis while ensuring that the SASA metric remains within the bounds of 0 to A where A represents the total number of atoms in the molecule. In this dataset, the smallest and largest observed SASA scores were 1.33 × 10-9and 13.3, respectively. The user-defined threshold, SASA=1, was selected to yield a manageable search space.

[0107] Discuss CA / UA Calculation

[0108] The effectiveness of SASA was examined by using a small, randomly selected set of organic molecules. The study assembled a pool of 350 molecules, and each one underwent SASA calculation. The plot of entry versus SASA value indicates that symmetric and easy-to- synthesized molecules are effectively distinguished from non-symmetric ones, using a threshold SASA value of one. For example, 149 and 247 both show a SASA value close to 1. However, the entry 149 (SASA = 1.13) shows a symmetric structure, while the entry 247 (SASA = 0.766) is not. Additionally, 26 (SASA = 4.55) and 260 (SASA = 3.87 × 10-9) have similar structures but have significantly different SASA values due to the asymmetrical structure of 260 introduced by the unique methyl group. Extending this inquiry, compounds 185 (SASA = 3.07) and 140 (SASA = 3.65 × 10-9) yielded another comparable outcome primarily influenced by the singular nitrogen atom in 140, while the even distribution of four nitrogen atoms in 185 enhances its symmetricity. It is noteworthy that the heightened SASA of OEM does not guarantee a viable synthetic pathway. Nonetheless, the SASA serves as a pivotal tool, significantly decreasing the screening time of vast OEM spaces.

[0109] Interpretation of the SPARKLE equation:

[0110] Specific energy Predicted specific energy = 1717.16 × ( ^^^^1) + 0.027413Attorney Docket No.103361-617WO1 T2024-122

[0111] ^^^^1: Barysz matrix weighted by Allred-Rocow electronegativity

[0112] ^^^^2: Barysz matrix weighted by polarizability

[0113] ^^^^3: Moreau-broto autocorrelation of lag 1, weighted by atomic number

[0114] ^^^^4: Moreau-broto autocorrelation of lag 2, weighted by atomic number

[0115] One can attempt to interpret the learned equation as specific energy being a function electronegativity divided by a function of polarizability and molecular weights. First, note that the specific energy equation is a function of the reduction potential over molecular weight. Given that the learned equation has a MW term in the denominator, it appears that the reduction potential may be governed by a competition between electronegativity, and polarizability. Although the Barysz matrix and Moreau-broto autocorrelation obfuscate the contributions of these driving physical phenomena, they such models may be useful in guiding experimentation to better understand the physical relationship between electronegativity, polarizability, and the molecular structure’s effect on reduction potential.

[0116] The present model shows a significant advantage over alternative blackbox machine learning methods [NN; RF; and linear regression with L1 regularization (LASSO)] in terms of generalization performance (Figure 5). Considering the DFT predictions of specific energy on the ≈103,000 quinones as the test set, the ML model achieves a test R2value of 0.733 while the NN, RF, and LASSO models all end up with negative R2values (implying heavily biased models, Figure 5 bottom).

[0117] It is worth noting that this level of prediction performance was obtained using only ≈100 data points (less than 0.1% of the test data set), suggesting the noted model is robust to the overfitting of the training data. The reduced level of overfitting is well beyond what is achievable with traditional black-box machine learning, which, without wishing to be bound by theory, is likely a consequence of the sparsity enforced by SPARKLE. It was demonstrated that, given ≈100 randomly selected training points, SPARKLE identifies a structurally unique model that has very low prediction error, indicating that the results shown in Figure 5 are insensitive to the specific train / test split.

[0118] Solvation energy

[0119] A similar procedure as above was repeated for solvation energy. To ensure model accuracy, a training set of 3000 randomly selected molecules was utilized. The resulting model for predicted solvation energy is expressed as follows: Solvation energy

[0120] x5: Moreau-broto autocorrelation of lag 1, weighted by atomic numberAttorney Docket No.103361-617WO1 T2024-122

[0121] x6: Molecular ID on Nitrogen atom

[0122] x7: Geary Coefficient of lag 3, weighted by polarizability

[0123] x8: Geary Coefficient of lag 2, weighted by sigma electrons

[0124] Here, it can be seen that the solvation energy can be accurately predicted from a function of the molecular weight, number of Nitrogen atoms, polarizability and sigma electrons. It is noted that polarizability and the effect of sigma electrons have some theoretical basis for why they may contribute to the aqueous solvation energy. The molecular ID on Nitrogen atoms, on the other hand, may be a corrective term to account for the increased polarizability of nitrogen molecules, which has been observed in the literature.

[0125] Results and Discussion

[0126] Symbolic Predictive Algorithm for Recognizing Key Molecular Elements (SPARKLE) Model: The study sought to develop a predictive model that reduces the need for high throughput DFT screening and / or experimentation in material searching problems. A challenge to applying machine learning (ML) models to a search for new OEMs is the selection of appropriate descriptors that correlate with target properties, e.g., solubility, redox potential, and synthetic accessibility (FIG.18, Panel A).

[0127] SPARKLE was designed with human interpretability in mind. It defines a property by considering contributions and competitions across various physically meaningful quantities as an algebraic model; similar to dimensionless numbers. SPARKLE builds equations with only a few molecular descriptors from a large feature space, constructed using an operator set on a large number of descriptors. Combinatoric screening of a combination of descriptors and operators yields the features with the highest correlation to observations. In this context, it is emphasized that descriptors xirefer to quantities computed directly from a molecule (see Todeschini and Consonni, 2000), while features refer to both the molecular descriptors and the algebraic combinations of descriptors.

[0128] Although learning the SIMs can be challenging, the resulting models have several properties that justify the approach. In particular, when the feature space consists of physically meaningful descriptors, the resulting model can describe the underlying chemical relationship, which yields the following benefits: (i) learning a model requires fewer data points, relative to purely data-driven models; (ii) the learned models can provide physical insights into the investigated phenomena; (iii) the learned models have the ability to generalize better than black-box models; and (iv) the learned models provide an efficient "latent representation" for use in efficient non-parametric modeling / optimization frameworks.Attorney Docket No.103361-617WO1 T2024-122

[0129] Even with just a few molecular descriptor features, the total possible number of combinations is in excess of 1019and would be impossible to do manually. Instead, the study used a sure independence screening and sparsifying operator (SISSO)24to identify property descriptors from high dimensional molecular descriptor feature vectors, which provides the physically meaningful features of molecules necessary to learn novel molecular property equations (FIG. 18, Panel B). These two tools work together to learn an algebraic equation that has a physically meaningful structure, which requires less data and is generalizable (FIG. 18, Panel C).

[0130] First, the study developed models with a focus on solvation-free energy (∆^^^^^^^^^^^^^^^^^^^^)and redox potentials (^^^^° ). Solvation energy is roughly proportional to the logarithm ofsolubility (^^^^, EQ.1),12which influences the viability of organic molecules at OEMs in solid- state batteries. The predicted redox potential (^^^^°), combined with the molecular mass of OEMs (MW), was used to calculate the specific energy (^^^^) assuming a two-electron redox process (n = 2, EQ.2). ∆^^^^^^^^^^^^^^^^^^^^ ∝ log^^^^(EQUATION 1) ^^^^ = ^^^^^^^^ × (^^^^° − ^^^^^^^^^^^^^^^^^^^^^^^^) / (3.6 × Mw)(EQUATION 2)

[0131] The study used DFT-computed redox potentials of 115 paraquinones to train a model (R2= 0.987) and validate it on 103 orthoquinones (R2= 0.948) and 537 additional anthraquinones quinones (R2= 0.862) at the same level of theory. Afterward, the study performed additional validation on 103k quinones collected from the literature,26having an R2value of 0.733. For the solvation energy, the study trained a model using 3,000 randomly sampled molecules from the same 103k quinones collected from the literature and used the remaining 100k for testing; the model scored 0.832, and 0.827 R2for training and testing, respectively.

[0132] The SPARKLE framework in the present example identified eight molecular descriptors key to predicting the specific energy and solvation energy, which were correlated to atomic number, polarizability, electronegativity, and the number of sigma electrons. These molecular features provide physical meaning to the SPARKLE model, which would not be possible without the highly informative feature space.

[0133] Symmetry-adapted synthesis accessibility (SASA): To facilitate further selection of OEM candidates, the study implemented a synthesis accessibility model to prioritize organicAttorney Docket No.103361-617WO1 T2024-122 structures generated from de novo molecule design. Prediction of synthesis accessibility is still an evolving field as it is difficult to classify easy-to-synthesize vs. hard-to-synthesize molecules. The study initially examined the performance of existing synthesis accessibility schemes, e.g., SYBA,27SCScore,28and SAScore.29However, these models tend to recommend highly asymmetric molecules that require multistep synthesis from often cost-prohibitive reagents. This result may be caused by the existing synthetic accessibility models being designed with medicinal chemistry applications in mind. However, it does impede their direct application for the study, which was to identify low-cost OEM materials that can be produced readily from feedstock chemicals.

[0134] Given these challenges, the model was refined to bias molecules with high symmetry. Organic compounds with highly symmetric structures typically require fewer synthetic steps from readily available starting reagents. Moreover, highly symmetric molecules tend to self-organize into orderly arrangements in the solid state. This inherent property enhances the crystal lattice energy and is therefore anticipated to decrease solubility when exposed to an electrolyte.

[0135] The study employed SCScore to provide preliminary synthetic complexity evaluation, then fine-tuned its value by introducing a chemically equivalent point calculation designed to evaluate the degree of symmetry of the molecule. Taking account of the SCScore and the numbers of chemically equivalent atoms (CA) and unique atoms (UA) among the molecules, the symmetry-adapted synthetic accessibility score (SASA) is expressed by EQ.3: ^^^^^^^^^^^^^^^^ = ^^^^^^^^ / (^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + ^^^^^^^^)(EQUATION 3)

[0136] This equation was chosen since it applies a sufficient penalty to unsymmetric molecules and establishes a clear boundary between molecules synthesizable from 1-2 steps from the rest of the pool.

[0137] The effectiveness of SASA was evaluated using a dataset of 350 organic molecules containing a wide range of symmetric and unsymmetric molecules. The SASA, when plotted against the molecule entry, found that a threshold value of 1 is a good boundary between easy- to-synthesize and hard-to-synthesize molecules.

[0138] Deployment of SIM to identify new OEMs for AZIB: With a validated multi- objective model in hand (solvation energy, redox potential, symmetry-adapted synthetic accessibility), the study deployed this model toward a design space of 600,000 compounds generated from molecular generator FASMIFRA. These generated molecule spaces emerge as compelling candidates for deployment as OEMs in aqueous Zinc-ion batteries (AZIBs). AAttorney Docket No.103361-617WO1 T2024-122 noteworthy advantage of the ∆^^^^^^^^^^^^^^^^^^^^and redox behavior prediction under an aqueous environment is that it provides more controlled and predictable results compared to those of the organic medium, where pioneering prediction models, built upon extensive experimental data, prove effective at estimating the ∆^^^^^^^^^^^^^^^^^^^^and ^^^^° in the aqueous system. Additionally, the relatively lower solubility of OEMs in aqueous environments than organic electrolyte systems has implications for their applicability in AZIBs, which was expected to mitigate the risk of dissolution during battery cycling and thus enhance stability in the lifespan of AZIBs. The combined OEMs and aqueous Zn redox system are more appealing due to their environmentally friendly approaches with inherent safe and cost-effective systems, positioning organic AZIBs as a promising future grid-scale energy storage solution.

[0139] A three-dimensional plot was constructed using the solvation energy as an estimation of risk, specific energy density as an estimation of reward, and symmetry-adapted synthesis accessibility as an estimation of cost, as shown in FIG.19.

[0140] The entire design space with 600,000 compounds would be too large to sort through manually. Therefore, the study applied a series of conditions to prioritize hit / lead compounds. First, the predicted specific energy (x-axis) was limited to above 175 Wh kg-1, considering that this is a state-of-the-art value of inorganic / organic cathode materials.

[0141] The y-axis indicates aqueous solubility (S) predicted using solvation energies (∆^^^^^^^^^^^^^^^^^^^^). More exergonic solvation energies generally correlate with lower aqueous solubility, indicating the candidate has a higher chance of being viable toward solid-state battery applications. The solvation energy was limited to less than a predefined kcal mol-1to filter out small molecules with high predicted solubility.

[0142] Lastly, the SASA was limited to ≥ 1 to facilitate the selection of synthesizable molecules. As AZIBs is considered a grid-scale energy storage system, a low synthetic cost is a priority. This SASA value limits the final candidate pool to ca. 4000 molecules, corresponding to 1.67% of the design space. The remaining candidates were plotted in a 2D projection (FIG.20A).

[0143] The large set of molecular predictions was used to identify a Pareto front corresponding to the two battery-relevant properties, which reflects the tradeoff relationship between risk (solubility) and reward (specific energy). Among the ca. 4000 molecules examined virtually, twenty-eight (1-30) OEM candidates previously untested in AZIBs were selected. These molecules were ranked from low-risk low-reward to high-risk high-reward in FIG.20B.Attorney Docket No.103361-617WO1 T2024-122

[0144] Most of the predicted OEMs were based on the para- and orthoquinone structures due to the prevalence of quinones in the training set. However, the model also extended its predictions beyond the quinones, such as predicting high-potential p-type cathode materials, including the tertiary amine or thioester motif as a redox center (21-28). It is worth noting that compounds 26 and 28 are new organic moieties yet to be tested in the OEM research field.

[0145] Testing the performance of selected candidates OEMs: The selected candidates were evaluated as cathodes in AZIB by constructing a two-electrode coin cell containing a Zn metal as anode and an aqueous solution of zinc salt as the electrolyte. The solubility of OEMs in the aqueous electrolyte was examined before proceeding with the AZIB test. It is worth noting that most of the candidates were insoluble in the common electrolyte, such as 3 M zinc triflate (Zn(OTf)2), zinc sulfate (ZnSO4), or zinc chloride (ZnCl2), suggesting the validity of the solubility prediction.

[0146] FIG. 21 shows the structure and the first galvanostatic charge-discharge (GCD) cycle of successful OEMs at a constant current density of 0.5 C (based on two electrons). All OEMs exhibit at least a two-electron discharge process. However, some show capacity beyond two-electron per molecule, perhaps due to the co-insertion of Zn2+and H+. Notably, compound 18 shows high operating potential with extraordinary energy density over 220 Wh / kg, CE > 99% and round trip efficiency > 82%. The study also observed a successful operation of p-type molecules, where 24 and 26 show high operation potentials over 1.2 V, supporting possibilities of extending the model to the unknown OEMs beyond the quinone-based molecules. Deployment of symbolic and interpretable models to identify new OEMs is shown in FIG.22.

[0147] Notably, the SPARKLE model outperformed traditional human strategies, and successfully predicted viable OEMs with a higher success rate. Candidates recommended by the model had a 58% success rate, as defined by a successful GCD cycle of the candidate with a Coulombic efficiency higher than 95%, a specific capacity higher than 150 mAh / g for n-types and 100 mAh / g for p-types, and a voltage hysteresis less than 0.5 V. The collective data over the past six years on 68 molecules tested in the laboratory shows a success rate of 19% using human intuition. While the predictive accuracy beyond quinone is lower (37.5%, 3 / 8), this study underscores the potential for expanding the scope of the model in diversifying OEM discovery. A comparison of a traditional vs. a machine learning-based workflow is shown in FIG.23.

[0148] Long-term stability testing was performed to establish their viability as OEMs in AZIB. The current model is not capable of predicting their stability after hundreds of cycles, as the decomposition pathway varies significantly as a function of functional group identityAttorney Docket No.103361-617WO1 T2024-122 and substitution pattern. Nonetheless, the much-narrowed testing space makes long-term cycling studies viable. The candidates that exhibit reversible cycling profiles were further monitored for the long-term cycle performance at a 2°C rate, and it was found that several compounds exhibited suitable stability for long-term AZIB application with 80% capacity retained after 500 cycles. Under optimized conditions, the best-performing candidates X exhibit reversible two-electron cycling with an initial capacity of 235.9 mAh g-1, of which 76.2% was retained after 100 cycles, representing comparable performance to state-of-the-art organic cathode materials in AZIBs. Example 2: Mol-SISSO Framework for Molecule Analysis and Discovery

[0149] A study was conducted to develop a framework to unlock the potential of Mordred molecular descriptors and the SISSO method in the quest to understand and identify molecules with specific properties. This framework harmoniously combines these key elements to simplify and streamline the complex landscape of molecule analysis.

[0150] Mordred molecular descriptors are central to the framework. These descriptors are like labels that provide a clear understanding of each molecule’s distinctive characteristics. By leveraging these descriptors, patterns can be uncovered, and deeper insights into why certain molecules possess the specific properties being sought. The study’s focus on molecular descriptors facilitates the development of models that not only predict properties accurately but also offer interpretability, bridging the gap between data-driven predictions and scientific understanding.

[0151] Within the framework, the power of the Sure Independence Screening and Sparsifying Operator (SISSO) method was harnessed. This method serves as the guiding light in this study, allowing one to sift through the wealth of Mordred molecular descriptors to pinpoint the most critical ones and build a linear model that drives the predictions.

[0152] In summary, this framework represents a practical approach that maximizes the utility of molecular descriptors and the SISSO method. It streamlines the complex task of understanding and identifying molecules with specific properties.

[0153] In-depth Understanding of SISSO with Mordred

[0154] The insights into Mordred descriptors calculation process empowered the compilation of a comprehensive database of molecules, each enriched with a whopping 1800 descriptors

[0015] . These descriptors encapsulate a wealth of physical information about the molecules. However, navigating this high-dimensional descriptor space necessitates the implementation of a robust predictive method, capable of efficiently handling this multifacetedAttorney Docket No.103361-617WO1 T2024-122 data. Moreover, the identification of descriptors that exhibit high correlations with the target properties of the molecules is a crucial step.

[0155] To fully comprehend the mechanics of the framework developed for harnessing Mordred descriptors and constructing predictive models, it is essential to break down the approach into four distinct steps. These steps collectively form a systematic and powerful framework that transforms the wealth of molecular information embedded within the descriptors into actionable insights and predictions for target properties.

[0156] The study begins by examining the first step, which involves the careful preprocessing of the initial set of extensive Mordred descriptors. The primary objective at this stage is to refine and prune the vast descriptors space, ultimately selecting a subset of features that holds the highest correlation to the target property. Mathematical operator settings are configured and the method SISSO is run on each feature to assess each descriptor feature individually, ensuring a comprehensive evaluation. To achieve the pruning accurately, RMSE was considered as an evaluation metric, as it quantifies the extent to which the descriptors can faithfully predict the target property. Through a rigorous evaluation process, the subset of 100 features that exhibit high correlation with the target property were identified, thus setting the state for more accurate predictions. This careful, one-by-one assessment is vital because it ensures a comprehensive evaluation, leaving no stone unturned in the quest for the most relevant descriptors.

[0157] With the first step executed successfully, the study then proceeds into the second step of the framework, which involves feature expansion. In this next stage, the aim is to take the selected subset of descriptors from the first stage and perform feature expansion with specific mathematical operator set configurations (H), which can help in transforming the captured small descriptor feature space (Φ0) into a complex high-dimensional descriptor feature space, which can provide a role in elevating the quality of the model predictions. Operator set(EQUATION 4) Initial Feature set = Φ0(EQUATION 5) Then a huge feature space with combinations is created, which enhances the representation of underlying data. Created Feature SpaceAttorney Docket No.103361-617WO1 T2024-122 (EQUATION 6)

[0158] This strategic expansion is designed to magnify the descriptor feature space exponentially, reaching approximately 1010features. This remarkable increase in dimensionality unlocks an extensive library of molecular information, each feature contributing to a more comprehensive and intricate representation of the underlying data. The introduction of new dimensions through the mathematical operators opens doors to a richer understanding of the intricate patterns, relationships, and correlations embedded within the data. This feature expansion is pivotal in the mission to equip the models with the capacity to tackle the complexity of molecular interactions and relationships, enhancing their predictive power to unprecedented levels.

[0159] Progressing through the comprehensive framework, having successfully executed the initial preprocessing and feature expansion stages, the study now arrives at an important next stage where an efficient feature selection technique is employed for identifying the best descriptors that have a high correlation with the target property. The purpose of equipping this technique is to enhance model performance and interpretability by focusing on the most salient aspects of the data. Several feature selection methods like filter methods, wrapper methods, embedded methods, and regularization-based methods have been developed to tackle the challenges of high dimensional data. These methods provide valuable solutions but often suffer from limitations of high computational complexity, difficulty in handling highly correlated features, and lack of stability. To tackle these challenges, a powerful and efficient feature selection is proposed by ”Jianqing Fan” [6] termed as “Sure Independence Screening (SIS)” which is an integral part of the methodology, and plays a role in reducing the huge dimensional space with a specific threshold.

[0160] SIS is a non-parametric statistical method that is developed for variable selection in high dimensional space based on “The fundamental idea that is if a feature is independent of the target / response variable, it is unlikely to be a good predictor of the target variable.”[6] By screening the variables based on their independence criterion, SIS reduces the dimensionality of the problem and facilitates subsequent modeling with a small subset of features.

[0161] The SIS algorithm includes a series of steps involved in evaluating scores of the variables with the target / response variable, ranking, and, lastly, the selection of variables based on the threshold of a number of features or a score value. SIS uses a metric termed “correlation magnitude” (i.e., the absolute inner product of the target variable with the feature variable) to assign the scores and then ranks them in descending order. Features with the highest scores are more likely to have strong predictive power for the target variable.Attorney Docket No.103361-617WO1 T2024-122

[0162] In the final stage of the framework, the objective is to establish a linear model that captures the relationship between the target variable and the subset of features identified during the screening stage. When dealing with the linear interdependence of feature variables (descriptors) and the target variable (target property), the preferred method is often the Least Squares (LS) method [7]. The essence of this approach lies in minimizing the loss function, which is defined as the sum of squared residuals. The optimization problem is represented as follows: arg min ‖^^^^ − ^^ ‖2^^^^^^^^^^ 2(EQUATION 7)

[0163] Here, y denotes the vector of the target variable, X is the matrix representing descriptor variables, and c signifies the coefficients vector. The overarching goal is to determine the vector c that minimizes the sum of squared residuals, where residuals represent the differences between the actual target values and the predicted outputs.

[0164] However, the challenge that is often encountered in building linear models is the risk of overfitting, a phenomenon that can adversely impact the model’s performance on unseen data. To mitigate this issue, a regularization technique is employed, which introduces sparsity into the problem. By default, L0 regularization is employed, as it promotes exact sparsity by driving most of the feature (descriptor) coefficients to zero. This technique proves especially advantageous in high-dimensional spaces, addressing the dependence of the model only on a small set of descriptors which enhances the interpretability and performance of the model.

[0165] The augmented regularization term is introduced into the optimization problem as:(EQUATION 8) where‖^^^^‖0defines the number of non-zero components of the vector.

[0166] There are various methods to implement L0 regularization, and one simplistic yet effective approach involves considering all possible combinations and selecting the combination that minimizes the error.

[0167] Results and Discussion

[0168] Implementation of SISSO on a comprehensive dataset of 115 paraquinones demonstrates exceptional performance in capturing the intricate relationships governing redox potentials. The method’s ability to navigate through the descriptors space yields accurate predictions, providing a foundation for understanding the broader organic class molecules redox behavior.Attorney Docket No.103361-617WO1 T2024-122

[0169] Linear equation resulted on application of SISSO is as follows: 0.027413(EQUATION 9)

[0170] represents −

[0171] Overview description: VR2_Dzare molecular descriptor is a normalized Randic- like eigenvector-based index from Barysz matrix weighted by Allred-Rochow electronegativity

[0016] . It is a topological descriptor that is used to describe the shape and branching of molecules. The descriptor is calculated by constructing Barysz matrix of molecule which encodes structural connectivity and bond types of the atoms in molecule. The matrix is then weighted by Allred-Rochow electronegativity of each atom, and then the eigenvector is calculated to form the descriptor. Capturing information regarding the distribution of atomic electronegativities and the molecular structure it can provide insights into the predictions of redox potential because it depends on the electronegativity, arrangement of atoms, and bond lengths information.

[0172] ^^^^2represents − VR2_Dzp

[0173] Overview description: VR2_Dzp molecular descriptor is normalized Randic-like eigenvector-based index from Barysz matrix weighted by polarizability

[0016] . It is a topological descriptor that captures the overall shape and branching pattern of molecules along with its polarizability distribution. It is calculated by constructing Barysz matrix and then weighted by the polarizability values of each atom reflecting their ability to deform their electron clouds in response to an external electric field, and then eigen vector is calculated to form the descriptor. Capturing information regarding the distribution of atomic polarizabilities and structural information of molecule it can provide insights into the predictions of redox potential.

[0174] ^^^^3represents – ATS1Z molecular descriptor

[0175] Overview description: ATS1Z molecular descriptor is Moreau-Broto autocorrelation of lag 1 weighted by atomic number

[0016] , is a topological descriptor that captures the distribution of atomic numbers within a molecule. It is calculated by considering the atomic number of each atom and its distance from a reference atom, typically the carbon atom with the lowest atomic number in the molecule. The contributions of each atom to the descriptor are weighted by their atomic number, reflecting the influence of the atomic number on the electron distribution, spatial arrangement within the molecule, and, consequently, the redox potential of the molecule.

[0176] ^^^^4represents ATS2Z molecular descriptorAttorney Docket No.103361-617WO1 T2024-122

[0177] Overview description: ATS2Z is Moreau-Broto autocorrelation of lag 2 weighted by atomic number

[0016] , is a topological descriptor that captures the distribution of atomic numbers within a molecule to a distance of two bonds. It is calculated by considering the atomic number of each atom and its distance from a reference atom, typically the carbon atom with the lowest atomic number in the molecule. The contributions of each atom to the descriptor are weighted by their atomic number, reflecting the influence of the atomic number on the electron distribution and, consequently, the redox potential of the molecule.

[0178] The training results of the linear model on the dataset of 115 paraquinones reveal a compelling performance, as evidence by an impressive R2= 0.987 and RMSE = 0.03.

[0179] To evaluate the generalizability of the SISSO model beyond the training dataset of 115 paraquinones, rigorous testing was conducted on multiple datasets. The model’s predictive accuracy was assessed using three diverse datasets, producing compelling results that underscore its robustness and versatility.

[0180] Performance Metrics:

[0181] Dataset of Orthoquinones: The model demonstrated exceptional predictive accuracy on 103 Orthoquinones dataset, achieving an R2of 0.948 and RMSE = 0.069. This indicates a strong correlation between the predicted and observed redox potentials, with the model explaining 94% of the variance in the new dataset.

[0182] Dataset of Anthraquinones: Extending the evaluation to a Dataset of 540 Anthraquinones yielded an R2value of 0.86 and RMSE=0.045. Despite potential variations in molecular structures and properties, the model showcased its ability to generalize well to different datasets. This reflects a robust predictive capacity, demonstrating consistency in capturing the nuanced interplay of molecular descriptors across diverse chemical spaces.

[0183] Dataset of Quinone derivatives: The model’s performance on the larger Dataset of quinone derivatives, consisting of 103,000 molecules, remains noteworthy. While the R2value of 0.733 is lower and RMSE = 0.056 is higher than the aforementioned datasets, it is important to highlight the challenges posed by the sheer scale and diversity of the Dataset. The model’s capacity to maintain a high level of predictive accuracy on this extensive dataset further attests to its generalizability.

[0184] Implications for Generalization: The model’s consistent performance across varied datasets substantiates its capability to generalize well beyond the initial training set of 115 paraquinones. The robust R2values achieved on three different datasets affirm that the model effectively captures the intrinsic relationships governing redox potentials in a broad spectrum of organic molecules.Attorney Docket No.103361-617WO1 T2024-122

[0185] In conclusion, the SISSO model’s ability to generalize, as evidenced by its strong performance across multiple datasets, instills confidence in its applicability to diverse chemical spaces. This adaptability positions the model as a valuable tool for predicting redox potentials in organic molecules beyond the confines of the original training set, contributing to a deeper understanding of molecular behavior in various contexts.

[0186] With the compelling results of the model in predicting redox potentials across diverse datasets, this robust foundation was leveraged for the development of a solubility model. Solubility, being a more intricate property, posed new challenges. Recognizing the inherent complexity, the study opted for a comprehensive approach, utilizing a larger dataset to encompass a broader chemical space. The lessons learned from redox potential predictions became invaluable as the study navigated the complexities of solubility, aiming to uncover the relationships between molecular descriptors and the solubility behavior of organic compounds.

[0187] Implementation of SISSO on a comprehensive dataset of 3000 quinone derivate molecules demonstrated exceptional performance in capturing the intricate relationships governing solubility. The method’s ability to navigate through the descriptors space yielded accurate predictions, providing a foundation for understanding the broader organic class molecules solubility behavior. 0.093116(EQUATION 10)

[0188] ^^^^1represents ATS1Z molecular descriptor

[0189] Overview description: ATS1Z molecular descriptor is Moreau-Broto autocorrelation of lag 1 weighted by atomic number, is a topological descriptor that captures the distribution of atomic numbers within a molecule. It is calculated by considering the atomic number of each atom and its distance from a reference atom, typically the carbon atom with the lowest atomic number in the molecule. The contributions of each atom to the descriptor are weighted by their atomic number, reflecting the influence of the atomic number on the polarity and interactions of the molecule with its environment, which in turn affects its solubility.

[0190] ^^^^2represents GATS2d molecular descriptor

[0191] Overview description: GATS2d molecular descriptor is Geary coefficient of lag 2 weighted by sigma electrons, is a topological descriptor that captures the spatial distribution of sigma electrons within a molecule. It is calculated by considering the sigma electron density of each atom and its distance from a reference atom, typically the carbon atom with the lowest atomic number in the molecule. The contributions of each atom to the descriptor are weightedAttorney Docket No.103361-617WO1 T2024-122 by their sigma electron density, reflecting the influence of sigma electrons on intermolecular interactions, which in turn affects the solubility of the molecule.

[0192] ^^^^3represents − MID N molecular descriptor

[0193] Overview description: MID N molecular descriptor is a molecular ID (MID) on N atoms is a topological descriptor that captures the connectivity and bonding patterns around nitrogen (N) atoms within a molecule. It is calculated by considering the number of N atoms and their connections to other atoms, including the types of bonds (single, double, or triple) and the presence of heteroatoms. The MID on N atoms can provide insights into the local environment of N atoms and their potential for intermolecular interactions, which can influence the solubility of the molecule.

[0194] ^^^^4represents GATS3p molecular descriptor

[0195] Overview description: GATS3p molecular descriptor is Geary coefficient of lag 3 weighted by polarizability is a topological descriptor that captures the spatial distribution of polarizability within a molecule to a distance of three bonds. It is calculated by considering the polarizability of each atom and its distance from a reference atom, typically the carbon atom with the lowest atomic number in the molecule. The contributions of each atom to the descriptor are weighted by their polarizability, reflecting the influence of polarizability on intermolecular interactions, which in turn affects the solubility of the molecule.

[0196] The training results of the linear model on the dataset of 3000 quinone derivatives revealed a good performance, as evidenced by an R2= 0.837 and RMSE = 0.16. These metrics collectively suggest accuracy, providing a robust foundation for the predictive capabilities of the model.

[0197] When subjected to the rigorous evaluation on a larger testing dataset of 100,000 molecules, the model demonstrated a consistent R2= 0.827 and RMSE=0.175. This result signifies the model’s robustness in maintaining accuracy even when faced with a substantially larger dataset. The marginal decrease in R2from the training to the testing dataset suggests the model’s ability to generalize effectively to novel instances.

[0198] The ability of the solubility model to consistently achieve high R2values on both training and testing datasets establishes its reliability for large-scale predictions. The model’s capacity to handle a diverse array of chemical structures within a significantly expanded dataset positions it as a valuable tool for predicting solubility across a broad spectrum of organic compounds.

[0199] In conclusion, the solubility model’s performance metrics, especially its robust R2and RMSE values, affirm its proficiency in capturing and predicting solubility trends. ThisAttorney Docket No.103361-617WO1 T2024-122 adaptability to varying chemical spaces reinforces its utility for researchers and practitioners seeking accurate predictions of solubility in diverse molecular contexts.

[0200] In the pursuit of discovering new and impactful molecules, a molecule generator called Fast Assembly of (Deep) SMILES Fragments was employed for generating a large dataset of 600,000 molecules, which propelled the study toward the identification of molecules exhibiting exceptional properties and promising functionalities. Going through all molecules is not an efficient approach, so a filter technique was devised which involves considering the Synthetic Complexity Score (SCS score) combination with the symmetricity of the molecules defined by the presence of single atom count and equivalent atoms count that will help in downsizing the molecule database.

[0201] The experimental phase yielded a cohort of molecules from the downsized molecules database explored by the framework, and these molecules were subjected to rigorous evaluations and contributions. Remarkably, the application of this framework resulted in a notable enhancement in the success rate of identifying molecules. The success rate catapulted from an initial 19% to an impressive 60%. This substantial improvement underscores the efficacy of the framework in not only generating molecular candidates but also in significantly augmenting the likelihood of discovering molecules with desirable properties. Example 3: TorchSisso - Efficient Torch Implementation of Sure Independence Screening and Sparsifying Operator

[0202] A study was conducted that explored the strengths and weaknesses of SISSO, and its objective was to introduce the Python implementation of the SISSO (Sure Independence Screening and Sparsifying Operator) method using the PyTorch library. This choice of implementation not only enables efficient computations in tensors but also provides support for GPU acceleration. The incorporation of PyTorch aims to enhance the accessibility and applicability of the SISSO method, given its remarkable predictive capabilities. The rationale behind implementing the SISSO method in Python lies in its potential for widespread usage, addressing some of its limitations, fostering collaboration, and facilitating seamless integration into existing scientific workflows.

[0203] Strengths of SISSO:

[0204] (1) High dimensional efficient feature selection: It can identify the relevant features and discard irrelevant or redundant features, thus reducing the dimensionality of the problem at hand.Attorney Docket No.103361-617WO1 T2024-122

[0205] (2) Interpretability: It selects the subset of important features that contribute significantly to the predictive performance, allowing researchers and practitioners to gain insights and understanding of the underlying relationships in the data.

[0206] (3) Robust to noisy data and irrelevant features, which gives enhanced predictive performance.

[0207] (4) It does work in many cases; it was a seminal paper that resulted in many scientific discoveries.

[0208] Weaknesses of SISSO:

[0209] (1) Limited exploration of feature interactions: In some scenarios, it may not capture the complex relations that are highly crucial for predictive performance. This limitation can result in overlooking important feature combinations that collectively contribute to predictive performance.

[0210] (2) Built-in Fortran - which has challenges with platform compatibility, installation, and integration with other tools.

[0211] (3) SISSO relies on CPU-based computations, which limits scalability and performance when dealing with computationally intensive tasks or large datasets.

[0212] The implementation of feature expansion closely aligns with the original SISSO method, emphasizing capturing intricate relationships between features. However, the Python implementation improves upon feature expansion by exhaustively exploring all possible combinations, leaving no stone unturned. This approach ensures a comprehensive examination of non-linear relationships between features, contributing to a more nuanced understanding of the underlying patterns in the data. The Python implementation retains the essence of the original SISSO method by incorporating the SIS (Sure Independence Screening) and SO (L0 regularization) techniques. These techniques play a crucial role in feature selection and regularization, contributing to the model’s predictive accuracy and interpretability.

[0213] The study involves a thorough analysis of the results obtained from the Python implementation of multiple case studies, setting the groundwork for a comprehensive comparison with the original SISSO method. As the results are discussed, the aim is to not only validate the accuracy of the Python implementation but also identify potential areas of improvement or optimization. This comparative analysis serves as a crucial step toward establishing the Python implementation as a reliable and accessible tool for researchers seeking to leverage the SISSO method in their studies.

[0214] Results and DiscussionAttorney Docket No.103361-617WO1 T2024-122

[0215] This study’s efforts were dedicated to optimizing the SISSO algorithm for streamlined performance, and the results are indeed promising. The most salient improvement lies in the substantial reduction of execution time compared to the original method. This enhancement is particularly crucial for applications requiring the analysis of extensive datasets or the rapid exploration of high-dimensional descriptor spaces. To quantify the efficiency gains, a series of case studies were conducted comparing the execution times and accuracy of the implementation with the original method.

[0216] TABLE 1 illustrates notable changes in the execution time and accuracy for two case studies when employing TorchSisso compared to the original SISSO method. In Case Study 1, TorchSisso achieved a reduction in execution time (0.01 seconds) while maintaining accuracy, as demonstrated in the parity plot. Conversely, in Case Study 2, there was an increase in execution time (4.08 seconds) due to the construction of an extensive feature space comprising 1,068,584 features, in contrast to the original SISSO’s 23,330 features. Despite the prolonged execution time, TorchSisso successfully generated the framed equation, capturing the nuanced relationship in the data, whereas the original SISSO method failed in equation formulation can be seen in TABLE 1. The trade-off between increased computational time and enhanced accuracy, especially in scenarios with a substantial feature space, underscores the significance of TorchSisso’s capabilities. TABLE 1. Comparison of generated equations with framed equationsTABLE 2. Comparison of TorchSisso and Original SISSO Method Performance

[0217] While TorchSisso demonstrates clear advantages in feature expansion strategies, acknowledging its strengths also opens up avenues for further advancements in the method. One notable area for improvement lies in its scalability, particularly when handling initial feature sets exceeding 100. Although TorchSisso excels in constructing expansive featureAttorney Docket No.103361-617WO1 T2024-122 spaces, its efficiency might be constrained when faced with a substantial number of initial features. The current limitations underscore the need for future releases of the package to focus on optimizing scalability and resource utilization. In addressing this limitation, future iterations of the TorchSisso package could explore innovative approaches to handle large initial feature sets more efficiently. This might involve refining algorithms, leveraging parallel processing, or incorporating advanced optimization techniques tailored for high-dimensional datasets. By enhancing scalability, TorchSisso could be better equipped to address the demands of diverse applications and datasets with varying degrees of complexity. Example 4: Preparation and Characterization of OEM Compounds

[0218] All reagents were purchased from commercial sources and used without further purification unless stated otherwise.1H and13C NMR spectra were recorded on Bruker 400 MHz spectrometer and were referenced to the NMR residual solvent peaks. s: singlet, d: doublet, t: triplet, dd: doublet of doublet, dt: doublet of triplet, m: multiplet. OEM compounds 11, 22, 25, 28, 29, and 30 were purchased from Fischer Sci (USA) and used without purification.

[0219] Synthesis of 2,3,5,6-tetrakis-[1,2,4]triazol-1-yl-[1,4]benzoquinone (1): 2,3,5,6- tetrakis-[1,2,4]triazol-1-yl-[1,4]benzoquinone was prepared according to a modified literature procedure [1,2]. The round-bottom flask was charged with p-chloranil (200 mg, 0.813 mmol, 1 eq), 1H-1,2,4-triazole (561 mg, 8.13 mmol, 10 eq), and deoxygenated DMSO (16.0 mL) and a stir bar. The mixture was stirred at 50°C for eight hours, resulting in a black solution with yellow precipitate. After cooling to room temperature, the precipitate was filtrated, washed with ice-cold acetone (10.0 mL × 3 times), and dried in vacuo at 110°C overnight to yield the product as a yellow solid (126 mg, 41.1% yield).1H NMR (400 MHz, DMSO-d6): δ 8.68 (s, 4H), δ 8.04 (s, 4H).13C NMR (101 MHz, DMSO-d6) δ 152.58, δ 147.39, δ 143.02, δ 126.16. (FIG.24 and FIG.25)1 41.2 % SCHEME 1: Synthesis of 1.

[0220] Synthesis of 2,3,5,6-tetrakis-1H-pyrazol-1-yl-[1,4]benzoquinone (2): 2,3,5,6- tetrakis-1H-pyrazol-1-yl-[1,4]benzoquinone was prepared according to a modified literature procedure [1,2]. The round-bottom flask was charged with p-chloranil (200 mg, 0.813 mmol, 1 eq), 1H-1,2,4-triazole (0.554 g, 8.13 mmol, 10 eq), and deoxygenated 1,4-dioxane (16.0 mL)Attorney Docket No.103361-617WO1 T2024-122 and stir bar. The mixture was stirred at 80 °C for five hours, generating a red precipitate. After cooling to room temperature, the precipitate was filtrated, washed with ice-cold methanol (10.0 mL) and water (10.0 mL), and dried in vacuo at 90 °C overnight to yield the product as a red solid (147 mg, 48.5% yield).1H NMR (400 MHz, DMSO-d6) δ 7.87 (d, 4H), δ 7.58 (d, 4H), δ 6.47 (t, 4H).13C NMR (101 MHz, DMSO-d6) δ 142.31, δ 140.96, δ 133.17, δ 127.25, δ 106.98. (FIG.6)2 48.5 % SCHEME 2: Synthesis of 2.

[0221] Synthesis of 2,3-bis(benzotriazol-1-yl)-1,4-naphthoquinone (4): 2,3- bis(benzotriazole-1-yl)-1,4-naphthoquinone was synthesized according to a modified literature procedure [3]. To a stirring and boiling (40 °C) solution of 2,3-dichloro-1,4-naphthoquinone (1.00 g, 4.40 mmol, 1 eq) in acetone (35.0 mL), a solution of 1,2,3-benzotriazole (1.05 g, 8.81 mmol, 2 eq) and KOH (0.494 g, 8.81 mmol, 2 eq) in acetone (8.00 mL) and water (2.00 mL) was slowly added over 30 min. The mixed solution was stirred for one hour, cooled to 0 °C in an ice bath, and the precipitate was vacuum filtrated and washed several times with distilled water, and dried in vacuo at 90 °C overnight to yield the product as a yellow solid (1.20 g, 69.4 % yield).1H NMR (400 MHz, CDCl3) δ 8.39-8.36 (dd, 2H), δ 7.99-7.96 (m, 4H), δ 7.48-7.44 (m, 2H), δ 7.37-7.33 (m, 4H).13C NMR (101 MHz, CDCl3) δ 178.74, δ 145.55, δ 135.53, δ 135.05, δ 133.78, δ 130.90, δ 129.25, δ 127.81, δ 124.89, δ 120.39, δ 110.43. (FIGS.7A-7B).SCHEME 3: Synthesis of 4.

[0222] Synthesis of dipyrido[3,2-a:2',3'-c]-benzo[3,4]-phenazine-11,16-quinone (6): 2,3- diphthlimido-1,4-naphthoquinone (6a) and 2,3-diamino-1,4-naphthoquinone (6b) were synthesized and characterized according to the same literature procedure. [4] 6a: 80.0 % yield.Attorney Docket No.103361-617WO1 T2024-122 6b: 79.2 % yield. The1H and13C NMR spectrum of these products matched that reported in the literature.[4]

[0223] Dipyrido[3,2-a:2',3'-c]-benzo[3,4]-phenazine-11,16-quinone was synthesized according to a modified literature procedure [5]. A solution of 6b (70.0 mg, 0.371 mmol, 1 eq) and 1,10-phenanthroline-5,6-dione (78.9 mg, 0.371 mmol, 1 eq) in glacial acetic acid (6.00 mL) was heated at 75 °C for three hours under ambient atmosphere and vigorous stirring. The precipitate was vacuum filtrated and washed with ethanol and diethyl ether until the filtrate was clear and dried in vacuo at 90 °C overnight to yield the product as a bright green powder (0.100 g, 74.1 % yield).SCHEME 4: Synthesis of 6.

[0224] Synthesis of quinoxaline[2,3-a]phenazine-6,7-dione (7): 7a: Rhodizonic acid dihydrate (0.100 g, 1 eq) was dissolved into the distilled water (50.0 mL) and degassed for 30 minutes by bubbling with N2gas. The prepared solution was added to a degassed solution of o-phenylenediamine (0.127 g, 2 eq) in acetic acid (200 mL, 10 vol.% in water) under a N2 atmosphere. The mixed solution was heated at reflux for three hours under an N2 atmosphere and then cooled to room temperature. The precipitate was vacuum filtrated, washed with distilled water several times, and dried in vacuo at 90 °C overnight to yield the product as a dark green powder (0.147 g, 79.5 % yield).1H NMR (400 MHz, DMSO-d6) δ 10.19 (s, 2H), δ 8.50~8.33 (dd, 4H), δ 8.06~7.96 (dt, 4H).13C NMR (101 MHz, DMSO-d6) δ 141.60 (C-N), δ 140.05 (C-OH), δ 137.57, δ 131.84, δ 130.03, δ 128.38. The1H and13C NMR spectrum of the product matched that reported in the literature. [6]

[0225] 7b: 7a (50.0 mg) was dissolved in HNO3 (5.00 mL, conc.) and heated at 90 °C for 30 minutes. Distilled water (5 mL) was added to the mixture after cooling down to room temperature, and the resultant precipitate was filtrated and washed with water several times and dried in vacuo at 90 °C overnight to yield the product as a yellow solid (0.038 g).Attorney Docket No.103361-617WO1 T2024-122SCHEME 5: Synthesis of 7.

[0226] Synthesis of 3,7-di-p-tolyl-1H,5H-4,8-dioxa-1,2,5,6-tetraaza-anthraquinone (8): 3,7-Di-p-tolyl-1H,5H-4,8-dioxa-1,2,5,6-tetraaza-anthraquinone was synthesized. Chloranilic acid (1.00 g, 4.79 mmol, 1 eq) was dissolved in ethanol (20 mL), and the 4- methylbenzohydrazide (1.44 g, 9.57 mmol, 2 eq) was slowly added to a solution over 10 min under vigorous stirring. The solution was stirred for fifteen minutes under an ambient atmosphere, and the reaction mixture was refluxed for twelve hours. After cooling down to room temperature, the brown-yellow precipitate was vacuum filtered, washed with DCM and hexane, and recrystallized in hexane to yield the product as a yellow solid (0.880 g, 45.9 % yield).1H NMR (400 MHz, DMSO-d6) δ 10.38 (s, 2H), δ 7.83-7.81 (m, 4H), δ 7.33-7.31 (d, 4H), δ 2.38 (s, 6H).13C NMR (101 MHz, DMSO-d6) δ 166.19, δ 142.29, δ 130.28, δ 129.49, δ 127.93, δ 21.50 (FIGS.8A-8B).6 45.9 % yield SCHEME 6: Synthesis of 8.

[0227] Synthesis of 1,4-bis(phenylamino) anthracene-9,10-dione (9): 1,4- bis(phenylamino) anthraquinone was synthesized according to a literature procedure. 1,4- difluoroanthraquinone (0.125 g, 0.512 mmol, 1 eq) was dissolved in DMSO (2 mL) under 50 °C with stirring for 30 min. To a solution, distilled aniline (200 μL, 2.20 mmol, 4.3 eq) was added, and the mixture was heated at 85 °C for twenty-four hours. After cooling down to room temperature, the reaction mixture was added to distilled water (100 mL), and the precipitate was vacuum filtered, washed with water (100 mL), and dried under reduced pressure. The crude product was recrystallized in hexane and ethanol (95 : 5 in volume), vacuum filtered, and driedAttorney Docket No.103361-617WO1 T2024-122 in vacuo at 80 °C overnight to yield the product as a dark purple solid (0.107 g, 53.5 % yield).1H NMR (400 MHz, CDCl3) δ 12.25 (s, 2H), 8.38 (m, 2H), 7.81-7.74 (m, 4H), 7.52 (s, 2H), 7.41-7.37 (m, 4H), 7.28 (m, 2H). (FIG.9)SCHEME 7: Synthesis of 9

[0228] Synthesis of 2,3,10,11-quinoxaline[2,3-a]phenazine-6,7-dione (10): 2,3,10,11- quinoxaline[2,3-a]phenazine-6,7-dione was prepared by following the same procedure as compound 7 by using the rhodizonic acid dihydrate and 4,5-dichloro-1,2,-phenylenediamine as reagents, yielding the product as a yellow solid.SCHEME 8: Synthesis of 10

[0229] Synthesis of tetraethyltetra-aza-anthraquinone (12):

[0230] Synthesis of 2,3,5,6-tetraphthalimido-1,4-benzoquinone (12a): 2,3,5,6- tetraphthalimido-1,4-benzoquinone was prepared according to a modified literature procedure.[5] Potassium phthalimide and p-chloranil were added to anhydrous MeCN and refluxed overnight under an N2 atmosphere. The reaction mixture was filtered while hot, and the precipitate was washed with DMF and water until the filtrate was clear and dried in vacuo at 90 °C overnight to yield the product as a dark yellow solid.

[0231] Synthesis of 2,3,5,6-tetraamino-1,4-benzoquinone (12b): 2,3,5,6-tetraamino-1,4- benzoquinone was synthesized according to a modified literature procedure.[5] A round- bottom flask was charged with 12a, and the hydrazine was added dropwise under smooth stirring. The mixture was heated at 65 °C for two hours. After cooling to room temperature, the precipitate was vacuum-filtered and washed with distilled water and tetrahydrofuran untilAttorney Docket No.103361-617WO1 T2024-122 the filtrate was clear and dried in vacuo at 90 °C overnight to yield the product as a deep dark purple solid.12a 12b 80.0 % yield 79.2 % yield SCHEME 9: Synthesis of 12a and 12b.

[0232] Tetraethyltetra-aza-anthraquinone was synthesized according to a modified literature procedure. [7] To a suspended 12b in acetic acid (10 vol.% in water) 3,4-hexadione was added and stirred at 80 °C for four hours. After cooling to room temperature, water (20.0 mL) was added, and the organic layer was extracted three times using CHCl3. The organic layer was combined, dried with MgSO4, and filtered, and the solvent was removed under reduced pressure. The residue brown solid was recrystallized in THF three times and dried in vacuo to yield the product as a brown-yellow solid (86.1 mg).12 12b xx % yield SCHEME 10: Synthesis of 12.

[0233] Synthesis of 2,5-difluoro-3,6-dianilino-1,4-benzoquinone (13): 2,5-difluoro-3,6- dianilino-1,4-benzoquinone was prepared according to a modified literature procedure [8]. NaOAc (90.1 mg, 1.10 mmol, 1.83 eq) was first dissolved in a mixed solution of ethanol (10.0 mL), acetic acid (10.0 mL), and water (5.00 mL). To a solution, 2,3,5,6-tetrafluoro-1,4- benzoquinone (108 mg, 0.599 mmol, 1 eq) was slowly added and stirred for half an hour under ambient atmosphere to give a purple solution. Aniline (0.220 mL, 2.40 mmol, 4 eq) was then added to the solution, and the resulting dark brown mixture was refluxed at 110 °C for five hours, filtered while hot, and washed with water, DCM, EtOAc, acetone, and CHCl3, and dried in vacuo to yield the product as a brown solid (121 mg, 61.8 % yield).Attorney Docket No.103361-617WO1 T2024-12213 61.8 % yield SCHEME 11: Synthesis of 13.

[0234] Synthesis of 3,6-dichloro-1,10-phenanthroline-5,6-dione (15): 3,6-dichloro-1,10- phenanthroline-5,6-dione was prepared according to a modified literature procedure [9]. 2,9- dichloro-1,10-phenanthroline (100 mg, 0.401 mmol, 1 eq) was dissolved in a 60 vol.% of H2SO4 (1mL) with proper stirring, and then NaBrO3 (66.7 mg, 0.441 mmol, 1.1 eq) was added carefully over the 20 min course. The orange color solution was stirred for twenty-four hours and poured on the ice, forming the yellow precipitate. After neutralizing with saturated NaOH, the precipitate was extracted with DCM, washed with saturated NaHCO3 (10.0 mL), dried with MgSO4, and dried in vacuo to yield the product as a yellow solid (92.0 mg, 82.1% yield).1H NMR (400 MHz, CDCl3) δ 8.43 (d, 2H), δ 7.62 (d, 2H).13C NMR (101 MHz, CDCl3) δ 177.19, δ 159.12, δ 152.22, δ 139.64, δ 127.29, δ 126.94 (FIGS.10A-10B).SCHEME 12: Synthesis of 15. Synthesis of 1,4-dihydro-benzo[g]quinoxaline-2,3,5,10-tetraone (17): 1,4-dihydro- benzo[g]quinoxaline-2,3,5,10-tetraone was prepared according to a modified literature procedure

[0010] . The reaction flask was charged with 6b (1.00 g, 3.58 mmol, 1 eq) and placed in the ice bath. Oxalyl dichloride (COCl2, 2.00 mL, 23.3 mmol, 6.5 eq) was added dropwise to a reaction flask under mechanical stirring and gradually increased the temperature to room temperature. The reaction solution was stirred for twenty-four hours. The crude product was triturated with water (10 mL × 3 times), vacuum filtrated, and dried in vacuo at 90 °C overnight to yield the product as a bright orange solid (0.705 g, 81.2 % yield).1H NMR (400 MHz, DMSO-d6) δ 12.03 (s, 2H, N-H), δ 8.05 (dd, 2H), δ 7.88 (dd, 2H).13C NMR (101 MHz, DMSO- d6) δ 177.37, δ 156.35, δ 134.95, δ 130.75, δ 126.50, δ 123.63 (FIGS.11A-11B).Attorney Docket No.103361-617WO1 T2024-1226b17 81.2 % yieldSCHEME 13: Synthesis of 17.

[0235] Synthesis of 2,5-bis(pyrazol-1'-yl)-1,4-dihydroxybenzene (18): 2,5-bis(pyrazol-1'- yl)-1,4-dihydroxybenzene was prepared according to a modified literature procedure

[0011] . The re-action flask was connected to a condenser and charged with p-benzoquinone (1.00 g, 9.25 mmol, 1 eq) and 1H-pyrazole (0.630 g, 9.25 mmol, 1 eq). After the reaction flask was vacuumed and filled with N2 gas three times, deoxygenated ethanol (10 mL) was added to a reaction flask and refluxed for twelve hours under an N2atmosphere. After cooling to room temperature, the precipitate was vacuum-filtered and washed with ethanol (10 mL × 3 times) and dried in vacuo at 80 °C overnight to yield the product as a dark brown solid (0.218 g, 9.73 % yield).1H NMR (400 MHz, DMSO-d6) δ 10.17 (s, 2H), δ 8.41 (dd, 2H), δ 7.73 (dd, 2H), δ 7.45 (s, 2H), δ 6.49 (dd, 2H).13C NMR (101 MHz, DMSO-d6) δ 141.22, δ 140.06, δ 131.45, δ 125.92, δ 111.56, δ 107.07 (FIGS.12A-12B).18 9.73 % yield SCHEME 14: Synthesis of 18.

[0236] Synthesis of Tris(4-(1H-imidazol-1-yl)phenyl)amine (21): Tris(4-(1H-imidazzol-1- yl)phenyl)amine was synthesized according to a modified literature procedure. Oven-dried Schlenk tube (10.0 mL) was charged with tris(4-bromophenyl)amine (250 mg, 0.519 mmol, 1 eq), imidazole (212 mg, 3.11 mmol, 6 eq), CuSO4·5H2O, (25.6 mg, 0.103 mmol, 0.2 eq) and K2CO3(358 mg, 2.59 mmol, 5 eq). The reaction flask was evacuate and filled with N2gas for three times, and sealed and stirred at 160 °C for twenty-four hours. After cooling down to room temperature, distilled water (5 mL) was added to the reaction mixture, stirred for two hours, and vacuum filtered. The remaining solid was extracted with methanol (20.0 mL), filtered through a Celite pad, and dried under reduced pressure. The remaining powder was triturated with ethanol (20.0 mL) for thirty minutes, filtered again through a Celite pad, and dried under reduced pressure to yield the product as a white solid (0.111 g, 48.3 % yield).1H NMR (400Attorney Docket No.103361-617WO1 T2024-122 MHz, DMSO-d6) δ 8.21 (s,3 H), δ 7.70 (s, 3H), δ 7.61 (d, 6H), δ 7.19 (d, 6H), δ 7.11 (s,3 H).13C NMR (101 MHz, DMSO-d6) δ 145.37, δ 135.60, δ 132.35, δ 129.80, δ 124.80, δ 121.94, δ 118.20. (FIG.13A and FIG.13B)SCHEME 15: Synthesis of 21.

[0237] Synthesis of 5,10-diphenyl-5,10-dihydrophenazine (23):

[0238] Synthesis of 5,10-dihydrophenazine (23a): 5,10-dihydrophenazine was prepared according to a modified literature procedure. Oven-dried three-necked reaction flask equipped with a condenser was charged with 5,10-phenazine (200 mg, 1.11 mmol, 1 eq), sodium dithionite (2.13 g, 12.2 mmol, 11 eq), and a stir bar. After the reaction flask was vacuumed and filled with N2 gas three times, de-oxygenated ethanol (15.0 mL) and deoxygenated distilled water (20.0 mL) were added to a reaction flask and refluxed for twelve hours under an N2 atmosphere. After cooling to room temperature, the precipitate was vacuum-filtered and washed with deoxygenated water (20.0 mL) and dried in vacuo at 70 °C to yield the product as a bright green solid (0.163 g, 88.0 % yield). The product was stored under vacuum until use.

[0239] Synthesis of 5,10-diphenyl-5,10-dihydrophenazine (23): 5,10-diphenyl-5,10- dihydrophenazine was synthesized according to a modified literature procedure. The reaction flask equipped with condenser was charged with 23a (100 mg, 0.549 mmol, 1 eq), 4- bromobenzene (145 μL, 1.37 mmol, 2.50 eq), tris(dibenzylideneacetone)dipalladium(0)chloroform adduct (35.2 mg, 38.4 μmol, 0.07eq.), tri- tert-butylphosphine tetrafluoroborate (15.9 mg, 54.8 μmol, 0.1 eq), sodium tert-butoxide (316 mg, 3.29 mmol, 6 eq), and stir bar. After the reaction flask was vacuumed and filled with N2 gas three times, deoxygenated toluene (4.00 mL) was added and refluxed for eight hours under an N2atmosphere. After cooling down to room temperature, distilled water (5.00 mL) and DCM (5.00 mL) were added to the reaction mixture, the organic layer was extracted, and aqueous phase was extracted three times with DCM (5.00 mL × 3 times). The combined organic layer was dried with MgSO4, filtered, and dried under reduced pressure. The residue brown-Attorney Docket No.103361-617WO1 T2024-122 yellow product was recrystallized in toluene and methanol (50:50 in volume), vacuum filtered, and dried in vacuo to yield the product as a pale yellow needle crystal (0.168 g, 61.0 % yield).1H NMR (400 MHz, C6D6) δ 7.16 (could not be resolved from NMR solvent peak), δ 7.07-7.01 (m, 2H), δ 6.28-6.25 (m, 4H), δ 5.82-5.80 (m, 4H).13C NMR (101 MHz, C6D6) δ 140.50, δ 136.82, δ 131.28, δ 131.07, δ 121.08, δ 112.77. (FIGS.14A-14B).23a 23 88.0 % yield 61.0 % yield SCHEME 16: Synthesis of 23a and 23.

[0240] Synthesis of 5,10-di(4-methoxyphenyl)-5,10-dihydrophenazine (24): 5,10-di(4- methoxyphenyl)-5,10-dihydrophenazine was synthesized according to a same procedure with 23 by using 23a (100 mg, 0.549 mmol, 1 eq) and 4-bromoanisole (172 μL, 1.37 mmol, 2.50 eq) as reagents, and refluxed for twenty-four hours under N2 atmosphere. The crude product was recrystallized by dissolving in minimal DCM, pipetting hexane until precipitate was formed, stored at –20 °C overnight, vacuum filtered, washed with ice-cold DCM, and dried in vacuo to yield a pale yellow needle crystal (0.128 g, 59.1 % yield).1H NMR (400 MHz, C6D6) δ 7.10 (d, 4H), δ 6.77 (d, 4H), δ 6.37-6.35 (dd, 4H), δ 5.93-5.90 (dd, 4H), δ 3.26 (s, 6H).13C NMR (101 MHz, C6D6) δ 159.13, δ 137.31, δ 132.79, δ 132.25, δ 121.01, δ 116.40, δ 112.67, δ 54.59. (FIGS.15A-15B).23a 24 59.1 % yield SCHEME 17: Synthesis of 24.

[0241] Synthesis of 9-(4-methoxyphenyl)-9H-carbazole (26): 9-(4-methoxyphenyl)-9H- carbazole was synthesized by following the previous literature procedure

[0012] . Oven dried Schlenk tube (10 mL) was charged with 9H-carbazole (0.100 g, 0.598 mmol, 1 eq), 4- iodoaniosole (0.174 g, 0.742 mmol, 1.24 eq), CuI (5.69 mg, 0.030 mmol, 5 mol.%), 1,10- phenanthroline (10.8 mg, 0.06 mmol, 10 mol.%), and KOH (0.068 g, 1.20 mmol, 2 eq). TheAttorney Docket No.103361-617WO1 T2024-122 reaction flask was evacuated and filled with N2 gas three times, the degassed DME (0.3 mL) and distilled water (0.7 mL) were added, sealed, and stirred at 100 °C for twenty hours. After cooling down to room temperature, ethyl acetate (10 mL) was added and washed with distilled water (10 mL × 3 times). The organic layer was separated and filtrated through the Celite pad, dried with MgSO4, filtered, and concentrated under reduced pressure. Column chromatography (hexane / ethyl acetate = 20 / 1 in volume as eluent) of crude product and the solvent was removed under reduced pressure to yield the product as a white solid (0.108 g, 66.1% yield).1H NMR (400 MHz, CDCl3) δ 8.16 (dt, 2H), δ 7.46 (d, 2H), δ 7.44-7.40 (m, 2H), δ 7.36-7.33 (dt, 2H), δ 7.31-7.27 (ddd, 2H), δ 7.26 (CDCl3), δ 7.14-7.11 (m, 2H), δ 3.93 (s, 3H).13C NMR (101 MHz, CDCl3) δ 159.00, δ 141.51, δ 130.44, δ 128.71, δ 125.97, δ 123.24, δ 120.45, δ 120.38, δ 119.77, δ 115.20, δ 109.82, δ 55.73 (FIGS.16A-16B).26 66.1 % yield SCHEME 18: Synthesis of 26.

[0242] Synthesis of 9-(4-chlorophenyl)-9H-carbazole (27): 9-(4-chlorophenyl)-9H- carbazole was synthesized by following the same method with 26 by using 9H-carbazole (0.100 g, 0.598 mmol, 1 eq) and 1-chloro-4-iodobenzene (0.177 g, 0.742 mmol, 1.24 eq) as reactants. Column chromatography (hexane / ethyl acetate = 10 / 1 in volume as eluent) of crude product yields the product as a white solid (0.101 g, 60.8 % yield).1H NMR (400 MHz, CDCl3) δ 8.15 (d, 2H), δ 7.58 (d, 2H), δ 7.51 (d, 2H), δ 7.44-7.36 (m, 4H), δ 7.30 (t, 2H) (FIG.17).27 60.8 % yield SCHEME 19: Synthesis of 27.

[0243] The following patents, applications and publications as listed below and throughout this document are hereby incorporated by reference in their entirety herein.Attorney Docket No.103361-617WO1 T2024-122 Reference list for Example 1 [1] Liang, Y.; Yao, Y. Positioning Organic Electrode Materials in the Battery Landscape. Joule 2018, 2, 1690–1706. [2] Luo, J.; Hu, B.; Hu, M.; Zhao, Y.; Liu, T. L. Status and Prospects of Organic Redox Flow Batteries toward Sustainable Energy Storage. ACS Energy Lett.2019, 4, 2220–2240. [3] Schon, T. B.; McAllister, B. T.; Li, P.-F.; Seferos, D. S. The Rise of Organic Electrode Materials for Energy Storage. Chem. Soc. Rev.2016, 45, 6345–6404. [4] Cheng, L.; Assary, R. S.; Qu, X.; Jain, A.; Ong, S. P.; Rajput, N. N.; Persson, K.; Curtiss, L. A. Accelerating Electrolyte Discovery for Energy Storage with High-Throughput Screening. J. Phys. Chem. Lett.2015, 6, 283–291. [5] Agarwal, G.; Doan, H. A.; Robertson, L. A.; Zhang, L.; Assary, R. S. Discovery of Energy Storage Molecular Materials Using Quantum Chemistry-Guided Multiobjective Bayesian Optimization. Chem. Mater. [6] Pineda Flores, S. D.; Martin-Noble, G. C.; Phillips, R. L.; Schrier, J. Bio-Inspired Electroactive Organic Molecules for Aqueous Redox Flow Batteries. 1. Thiophenoquinones. J. Phys. Chem. C 2015, 119, 21800–21809. [7] Allam, O.; Cho, B. W.; Kim, K. C.; Jang, S. S. Application of DFT-Based Machine Learning for Developing Molecular Electrode Materials in Li-Ion Batteries. RSC Adv.2018, 8, 39414–39420. [8] Xu, S.; Liang, J.; Yu, Y.; Liu, R.; Xu, Y.; Zhu, X.; Zhao, Y. Machine Learning-Assisted Discovery of High-Voltage Organic Materials for Rechargeable Batteries. J. Phys. Chem. C 2021, 125, 21352–21358. [9] Kristensen, S. B.; van Mourik, T.; Pedersen, T. B.; Sørensen, J. L.; Muff, J. Simulation of Electrochemical Properties of Naturally Occurring Quinones. Sci. Rep.2020, 10, 13571.

[0010] Robinson, S. G.; Yan, Y.; Hendriks, K. H.; Sanford, M. S.; Sigman, M. S. Developing a Predictive Solubility Model for Monomeric and Oligomeric Cyclopropenium-Based Flow Battery Catholytes. J. Am. Chem. Soc.2019, 141, 10171–10176.

[0011] Er, S.; Suh, C.; Marshak, M. P.; Aspuru-Guzik, A. Computational Design of Molecules for an All-Quinone Redox Flow Battery. Chem. Sci.2015, 6, 885–893.

[0012] Tuttle, M. R.; Brackman, E. M.; Sorourifar, F.; Paulson, J.; Zhang, S. Predicting the Solubility of Organic Energy Storage Materials Based on Functional Group Identity and Substitution Pattern. J. Phys. Chem. Lett.2023, 14, 1318–1325.

[0013] Sevov, C. S.; Hickey, D. P.; Cook, M. E.; Robinson, S. G.; Barnett, S.; Minteer, S. D.; Sigman, M. S.; Sanford, M. S. Physical Organic Approach to Persistent, Cyclable, Low-Attorney Docket No.103361-617WO1 T2024-122 Potential Electrolytes for Flow Battery Applications. J. Am. Chem. Soc. 2017, 139, 2924– 2927.

[0014] Lee, B.; Yoo, J.; Kang, K. Predicting the Chemical Reactivity of Organic Materials Using a Machine-Learning Approach. Chem. Sci.2020, 11, 7813–7822.

[0015] Sowndarya S. V., S.; St. John, P. C.; Paton, R. S. A Quantitative Metric for Organic Radical Stability and Persistence Using Thermodynamic and Kinetic Features. Chem. Sci. 2021, 12, 13158–13166.

[0016] Assary, R. S.; Zhang, L.; Huang, J.; Curtiss, L. A. Molecular Level Understanding of the Factors Affecting the Stability of Dimethoxy Benzene Catholyte Candidates from First- Principles Investigations. J. Phys. Chem. C 2016, 120, 14531–14538.

[0017] Lombardo, T.; Duquesnoy, M.; El-Bouysidy, H.; Årén, F.; Gallo-Bueno, A.; Jørgensen, P. B.; Bhowmik, A.; Demortière, A.; Ayerbe, E.; Alcaide, F.; Reynaud, M.; Carrasco, J.; Grimaud, A.; Zhang, C.; Vegge, T.; Johansson, P.; Franco, A. A. Artificial Intelligence Applied to Battery Research: Hype or Reality? Chem. Rev.2022, 122, 10899–10969.

[0018] Vermeire, F. H.; Chung, Y.; Green, W. H. Predicting Solubility Limits of Organic Solutes for a Wide Range of Solvents and Temperatures. J. Am. Chem. Soc.2022, 144, 10785– 10797.

[0019] Jin, W.; Barzilay, R.; Jaakkola, T. Junction Tree Variational Autoencoder for Molecular Graph Generation. In International conference on machine learning; PMLR, 2018; pp 2323– 2332.

[0020] Svetnik, V.; Liaw, A.; Tong, C.; Culberson, J. C.; Sheridan, R. P.; Feuston, B. P. Random Forest:  A Classification and Regression Tool for Compound Classification and QSAR Modeling. J. Chem. Inf. Comput. Sci.2003, 43, 1947–1958.

[0021] Chen, C.; Tanaka, K.; Funatsu, K. Random Forest Model with Combined Features: A Practical Approach to Predict Liquid-crystalline Property. Mol. Inform.2019, 38, 1800095.

[0022] Deringer, V. L.; Bartók, A. P.; Bernstein, N.; Wilkins, D. M.; Ceriotti, M.; Csányi, G. Gaussian Process Regression for Materials and Molecules. Chem. Rev. 2021, 121, 10073– 10141.

[0023] Jorissen, R. N.; Gilson, M. K. Virtual Screening of Molecular Databases Using a Support Vector Machine. J. Chem. Inf. Model.2005, 45, 549–561.

[0024] Ouyang, R.; Curtarolo, S.; Ahmetcik, E.; Scheffler, M.; Ghiringhelli, L. M. SISSO: A Compressed-Sensing Method for Identifying the Best Low-Dimensional Descriptor in an Immensity of Offered Candidates. Phys. Rev. Mater.2018, 2, 83802.Attorney Docket No.103361-617WO1 T2024-122

[0025] Moriwaki, H.; Tian, Y.-S.; Kawashita, N.; Takagi, T. Mordred: A Molecular Descriptor Calculator. J. Cheminform.2018, 10, 1–14.

[0026] Tabor, D. P.; Gómez-Bombarelli, R.; Tong, L.; Gordon, R. G.; Aziz, M. J.; Aspuru- Guzik, A. Mapping the Frontiers of Quinone Stability in Aqueous Media: Implications for Organic Aqueous Redox Flow Batteries. J. Mater. Chem. A 2019, 7, 12833–12841.

[0027] Voršilák, M.; Kolář, M.; Čmelo, I.; Svozil, D. SYBA: Bayesian Estimation of Synthetic Accessibility of Organic Compounds. J. Cheminform.2020, 12, 35.

[0028] Coley, C. W.; Rogers, L.; Green, W. H.; Jensen, K. F. SCScore: Synthetic Complexity Learned from a Reaction Corpus. J. Chem. Inf. Model.2018, 58, 252–261. (29] Ertl, P.; Schuffenhauer, A. Estimation of Synthetic Accessibility Score of Drug-like Molecules Based on Molecular Complexity and Fragment Contributions. J. Cheminform.2009, 1, 8. [A1] Guo, Z.; Ma, Y.; Dong, X.; Huang, J.; Wang Y.; Xia, Y. An Environmentally Friendly and Flexible Aqueous Zinc Battery Using an Organic Cathode. Angew. Chem., Int. Ed.2018, 57, 11737–11741. [A2] Nam, K. W.; Kim, H.; Beldjoudi, Y.; Kwon, T. W.; Kim, D. J.; Stoddart, J. F. Redox- Active Phenanthrenequinone Triangles in Aqueous Rechargeable Zinc Batteries. J. Am. Chem. Soc.2020, 142, 2541–2548. [A3] Wang, Q.; Liu, Y.; Chen, P. Phenazine-Based Organic Cathode for Aqueous Zinc Secondary Batteries. J. Power Sources 2020, 468, 228401. [A4] Tie, Z.; Liu, L.; Deng, S.; Zhao, D.; Niu, Z. Proton Insertion Chemistry of a Zinc–Organic Battery. Angew. Chem. Int. Ed.2020, 59, 4920–4924. [A5] Gao, Y.; Li, G.; Wang, F.; Chu, J.; Yu, P.; Wang, B.; Zhan, H.; Song, Z. A High- Performance Aqueous Rechargeable Zinc Battery Based on Organic Cathode Integrating Quinone and Pyrazine. Energy Storage Mater.2021, 40, 31–40. Reference list for Examples 2-3 [1] V. S. Bagotsky. “Fundamentals of Electrochemistry”. John Wiley Sons, Ltd, 2006. [2] Dong-Sheng Cao, Qing-Song Xu, Qian-Nan Hu, and Yi-Zeng Liang. ChemoPy: freely available python package for computational biology and chemoinformatics. Bioinformatics, 29(8):1092–1094, 2013. [3] Klaus Capelle. A bird’s-eye view of density-functional theory. Brazilian Journal of Physics, 36:1318–1343, 2002.Attorney Docket No.103361-617WO1 T2024-122 [4] Jun Yan Gui-Shan Tan Qing-Song Xu Dong-Sheng Cao, Yi-Zeng Liang and Shao Liu. PyDPI: Freely Available Python Package for Chemoinformatics, Bioin-formatics, and Chemogenomics Studies. Journal of Chemical Information and Modeling, 53(11):3086–3096, 2013. [5] Finale Doshi-Velez and Been Kim. Towards a rigorous science of interpretable machine learning, 2017. [6] Jianqing Fan and Jinchi Lv. Sure Independence Screening for Ultrahigh Dimen-sional Feature Space. Journal of the Royal Statistical Society Series B: Statistical Methodology, 70(5):849–911, 2008. [7] Luca M Ghiringhelli, Jan Vybiral, Emre Ahmetcik, Runhai Ouyang, Sergey V Levchenko, Claudia Draxl, and Matthias Scheffler. Learning physical descrip-tors for materials science by compressed sensing. New Journal of Physics, 19(2):023017, feb 2017. [8] Francesca Grisoni, Davide Ballabio, Roberto Todeschini, and Viviana Consonni. Molecular Descriptors for Structure–Activity Applications: A Hands-On Ap-proach, pages 3– 53. Springer New York, 2018. [9] P. Hohenberg and W. Kohn. Inhomogeneous electron gas. Phys. Rev., 136:B864–B871, Nov 1964.

[0010] W. Kohn and L. J. Sham. Self-consistent equations including exchange and correlation effects. Phys. Rev., 140:A1133–A1138, Nov 1965.

[0011] David R. Lide. Crc handbook of chemistry and physics. Journal of Americal Chemical Society, 2006.

[0012] Andrea Mauri, Viviana Consonni, Manuela Pavan, and Roberto Todeschini. Dragon software: An easy approach to molecular descriptor calculations. MATCH Communications in Mathematical and in Computer Chemistry, 56:237–248, 012006.

[0013] Andrea Mauri, Viviana Consonni, and Roberto Todeschini. Molecular Descrip-tors, pages 2065–2093. Springer International Publishing, Cham, 2017.

[0014] Christoph Molnar, Giuseppe Casalicchio, and Bernd Bischl. Interpretable machine learning – a brief history, state-of-the-art and challenges. In ECML PKDD 2020 Workshops, pages 417–431. Springer International Publishing, 2020.

[0015] Kawashita Norihito Takagi Tatsuya Moriwaki Hirotomo, Tian Yu-Shi. Mordred: a molecular descriptor calculator. Journal of Cheminformatics, 10, 2018.

[0016] Kawashita Norihito Takagi Tatsuya Moriwaki Hirotomo, Tian Yu-Shi. Mordred: a molecular descriptor calculator, 2018.Attorney Docket No.103361-617WO1 T2024-122

[0017] W. James Murdoch, Chandan Singh, Karl Kumbier, Reza Abbasi-Asl, and Bin Yu. Definitions, methods, and applications in interpretable machine learning. Proceedings of the National Academy of Sciences, 116(44):22071–22080, 2019.

[0018] Frank Neese. “Prediction of molecular properties and molecular spectroscopy with density functional theory: From fundamental theory to exchange-coupling”. Coordination Chemistry Reviews, 253(5):526–563, May.2009.

[0019] Runhai Ouyang, Stefano Curtarolo, Emre Ahmetcik, Matthias Scheffler, and Luca M. Ghiringhelli. Sisso: A compressed-sensing method for identifying the best low-dimensional descriptor in an immensity of offered candidates. Phys. Rev. Mater., 2:083802, Aug 2018.

[0020] Viviana Consonni Roberto Todeschini. “Handbook of Molecular Descriptors”. WILEY-VCH, 2000.

[0021] Jorge M. Seminario. An introduction to density functional theory in chemistry. In J.M. Seminario and P. Politzer, editors, Modern Density Functional Theory, volume 2 of Theoretical and Computational Chemistry, pages 1–27. Elsevier, 1995.

[0022] University of Tübingen. BlueDesc. ra.cs.uni-tuebingen.de / software / bluedesc / , 2017. Accessed 17 Aug 2017.

[0023] Yap Chun Wei. Padel-descriptor: an open source software to calculate molecular descriptors and fingerprints. Journal of computational chemistry, 32:1466–1474, 2011.

Claims

Attorney Docket No.103361-617WO1 T2024-122 CLAIMS What is claimed is:

1. A method for identifying a target molecule, the method comprising: receiving a molecule dataset from a library of molecules, the molecule dataset comprising chemical-structural representations of candidate molecules for a material application, determining one or more target molecular property values derived from a set of predictive molecular descriptor values determined from a candidate molecule of the molecular dataset applied to a trained machine learning (ML) model, wherein the trained ML model is configured to output the set of predictive molecular descriptor values for a given molecule data input, and wherein the trained ML model was trained from a set of training data and the set of molecular descriptors generated from a molecular descriptor modeling application; and outputting, via a report or to a data store, the one or more target molecular property values for each of the molecular dataset, wherein the one or more determined target molecular property is used for efficient identification of a set of molecules from the library of molecules for the material application.

2. The method of claim 1, wherein the one or more target molecular property values are used in a multi-objective model that evaluates a given model for two or more target molecular properties, wherein the multi-objective model is used to select the set of molecules from a library of molecules for a material application.

3. The method of claim 1 or 2, wherein the one or more target molecular property values include a first target molecular property value, wherein the first target molecular property value is determined from a first target molecular property model that employs two or more predictive molecular descriptor values of the set of predictive molecular descriptor values.

4. The method of any one of claims 1-3, wherein the one or more target molecular property values include redox potential, specific energy, synthesizability, or combinations thereof.

5. The method of any one of claims 1-4, wherein the set of molecular descriptors is determined from a larger set of molecular descriptors using a sure independence screening (SIS) algorithm.Attorney Docket No.103361-617WO1 T2024-122 6. The method of claim 5, wherein the SIS algorithm is a sure independence screening and sparsifying operation (SISSO) algorithm.

7. The method of any one of claims 1-6, wherein the set of molecular descriptors comprises 5 or more distinct molecular descriptors.

8. The method of any one of claims 1-7, wherein the molecular descriptors comprise a combination of zero-dimensional descriptors, one-dimensional descriptors, two-dimensional descriptors, and three-dimensional descriptors.

9. The method of any one of claims 1-8 wherein the set of molecular descriptors is determined using PaDEL Descriptor, Mordred, Blue Desc, ChemoPy, PyDPI, Dragon, molecular operating environment (MOE), ChemAxon, Schrodinger, KNIME, or a combination thereof.

10. The method of any one of claims 1-9, wherein the predictive molecular descriptor values comprise one or more molecular properties selected from the group consisting of: electronegativity, molecular weight, number of nitrogen atoms, sigma electrons, polarizability, and combinations thereof.

11. The method of any one of claims 1-10, wherein the one or more target molecular property values are represented as linear expressions of the predictive molecular descriptor values.

12. The method of any one of claims 1-11, wherein the library of molecules is generated using a fast assembly of SMILES Fragments (FASMIFRA) method.

13. The method of any one of claims 1-12, further comprising synthesizing a compound based on the one or more target molecular properties for the material application.

14. A compound synthesized from a molecule identified for a material application according to any one of the methods of claims 1-13.Attorney Docket No.103361-617WO1 T2024-122 15. A non-transitory computer readable medium having instructions stored thereon, wherein the instructions, when executed by one or more processors, cause the one or more processors to perform and any one of the methods of claims 1-13.

16. A system comprising: one or more processors; and a memory having instructions stored thereon, wherein the instructions, when executed by the one or more processors, cause the one or more processors to perform any one of the methods of claims 1-13.

17. A method of training a machine learning model, comprising: generating a set of molecular descriptors for a given molecule data input; applying a screening algorithm to identify a sub-set of molecular descriptors associated with a target molecular property; generating a training data set of molecular descriptors for a corresponding set of molecules; and training a machine learning model using the training data set of molecular descriptors and corresponding set of molecules.

18. The method of claim 17, wherein the training employs features from any one of claims 5-9.

19. An ML training system comprising: one or more processors; and a memory having instructions stored thereon, wherein the instructions, when executed by the one or more processors, cause the one or more processors to perform any one of the method of claims 17 or 18.

20. A non-transitory computer readable medium having instructions stored thereon, wherein the instructions, when executed by one or more processors, cause the one or more processors to perform any one of the methods claims 17 or 18.

Citation Information

Patent Citations

  • Systems and Methods for Predicting the Olfactory Properties of Molecules Using Machine Learning

    US20220139504A1

  • Machine learning based methods of analysing drug-like molecules

    US20220383992A1

  • Prediction of chemical compounds with desired properties

    US20230335231A1

  • Machine learning for predicting the properties of chemical formulations

    WO2022203734A1

  • Self-supervised learning-based method, apparatus and device for predicting properties of drug small molecules

    WO2023029351A1

Cited By

  • Intelligent driving material hysteresis modeling method based on knowledge-driven neural network

    CN121215099A