Generating predictions related to chemical processes using machine learning models

By encoding chemicals and processes and using machine learning models to generate latent representations, the problem of insufficient chemical workflow prediction in existing technologies is solved, and efficient prediction and generation of chemical workflows are achieved.

CN122029609APending Publication Date: 2026-05-12ALBERT INVENT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ALBERT INVENT CORP
Filing Date
2024-10-11
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing machine learning models are underutilized in areas beyond text processing, particularly in the prediction and generation of chemical workflows, where effective techniques are lacking.

Method used

By employing machine learning models to encode chemicals and processes and generate latent representations, the output properties of chemical workflows are predicted through an encoder-decoder architecture, reducing the need for physical experiments and generating novel chemical workflows to achieve predetermined goals.

Benefits of technology

By training machine learning models, it is possible to predict the properties of chemicals and processes, reducing the need for physical experiments and improving the accuracy and efficiency of chemical workflow predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122029609A_ABST
    Figure CN122029609A_ABST
Patent Text Reader

Abstract

In some embodiments, a computer-implemented method of generating predictions related to a chemical workflow including a process performed on at least one chemical is provided. A computing system encodes a representation of at least one chemical using a chemical encoder to create at least one potential chemical representation. The computing system creates a potential experimental representation based on the at least one potential chemical representation. The computing system decodes the potential experimental representation using an experimental decoder to predict one or more properties of the output of the chemical workflow, to generate an updated representation of the at least one chemical, or to generate an updated representation of the process.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications This application claims the benefit of Provisional Application No. 63 / 590341, filed on October 13, 2023, the entire disclosure of which is incorporated herein by reference for all purposes. Background Technology

[0002] Machine learning is a rapidly evolving field. Recently, new techniques for processing text have been created that allow for the prediction of novel text output based on input cues. These techniques are typically based on artificial neural networks and use architectures such as transformers. At a high level, transformers are typically trained on a text corpus and the corpus is encoded into a latent representation. A decoder is then typically used to generate novel text from the latent representation.

[0003] While the ability to generate novel output based on human understanding of training corpora has significant utility, these tools and techniques are underutilized in domains beyond text processing. What is desired are technologies for using machine learning models like these to understand and generate content in domains beyond language processing. Summary of the Invention

[0004] This summary is provided to introduce, in a simplified form, some concepts that will be further described in the following detailed description. This summary is not intended to identify key features of the subject matter protected by the claims, nor is it intended to help determine the scope of the subject matter protected by the claims.

[0005] In some embodiments, a computer-implemented method is provided for generating predictions associated with a chemical workflow, which includes a process performed on at least one chemical. The computational system encodes a representation of the at least one chemical using a chemical encoder to create at least one potential chemical representation. The computational system creates a potential experimental representation based on the at least one potential chemical representation. The computational system decodes the potential experimental representation using an experimental decoder to predict one or more properties of the output of the chemical workflow, to generate an updated representation of the at least one chemical, or to generate an updated representation of the process.

[0006] In some embodiments, a non-transitory computer-readable medium is provided on which computer-executable instructions are stored. These instructions, in response to execution by one or more processors of a computing system, cause the computing system to perform the methods described above.

[0007] In some implementations, a computing system is provided that is configured to perform the methods described above. Attached Figure Description

[0008] The foregoing aspects and numerous accompanying advantages of the invention will become more readily explained, as these aspects and advantages will become better understood when taken in conjunction with the accompanying drawings and the following detailed description, wherein: Figure 1 This is a schematic diagram of a non-limiting example implementation of a system for using machine learning to predict the properties of a chemical process, according to various aspects of this disclosure.

[0009] Figure 2 This is a block diagram illustrating various aspects of a non-limiting example implementation of a material property prediction calculation system according to various aspects of this disclosure.

[0010] Figure 3 This is a schematic diagram of a non-limiting example implementation of an encoder / decoder for processing workflow information according to various aspects of this disclosure.

[0011] Figures 4A-4C This is a schematic diagram of a technique for predicting the properties of the outputs of a chemical process, based on various aspects of this disclosure.

[0012] Figure 5 These are examples of non-limiting exemplary embodiments of the structured chemical representation according to various aspects of this disclosure.

[0013] Figure 6A and Figure 6B These are examples of non-limiting exemplary implementations of a structured process representation based on various aspects of this disclosure.

[0014] Figure 7A and Figure 7B This is a flowchart illustrating a non-limiting example implementation of a method for predicting the properties of the output of a chemical process according to various aspects of this disclosure.

[0015] Figure 8 This is a schematic diagram of a non-limiting example embodiment of an encoder / decoder according to various aspects of this disclosure, wherein context information is optionally encoded and injected into the encoding / decoding process at various points.

[0016] Figures 9A-9C Three non-limiting example techniques that can be used for potential spatial fusion according to various aspects of this disclosure are illustrated.

[0017] Figure 10 Non-limiting example implementations of agent networks according to various aspects of this disclosure are illustrated. Detailed Implementation

[0018] By storing a large amount of chemical experimental data generated from research and development work, embodiments of this disclosure are capable of training machine learning models to generate potential representations of chemicals and processes. These potential representations encode fundamental aspects of the chemicals and processes, enabling the generation of predictions of the chemical properties of the outputs of chemicals and novel experiments. In other words, embodiments of this disclosure are capable of predicting the properties of chemicals and chemicals generated from experiments that have not been previously conducted. This can reduce or eliminate the need for physical experiments on every potential combination of chemicals, combination of reaction process parameters, or chemical product to achieve a given chemical product or chemical process objective.

[0019] In some embodiments, this disclosure provides one or more machine learning models designed to consume structured and / or unstructured data that describes a complete end-to-end formulation and synthetic chemistry workflow (e.g., combinations of raw materials, parameters describing how raw materials are combined / processed into intermediates, any stages / steps for transforming intermediates into the final material, and / or any final process for transforming the final material into a form that can be characterized to understand material properties), and predict the properties of the final material or chemical composition. In some embodiments, similar models can be designed to generate novel workflows (or adapt existing workflows) to achieve predetermined objectives simply by providing a structured or unstructured description of the desired results.

[0020] This disclosure provides a machine learning architecture that creates latent representations encoded in a chemical workflow to generate predictions of the material properties of chemical products and / or to make recommendations and discoveries of novel workflows and / or chemical products based on those predictions. In some embodiments, the general technique involves transforming a description of the chemical workflow into a series of intermediate representations (latent representations) that encode relevant information about each component or process variable in the workflow in a form that can be used to predict final material properties. The latent representations themselves are learned from stored experimental data.

[0021] Depending on the type of parameters being encoded, different techniques can be used to derive latent representations. As a non-limiting example, when considering chemical compounds present in a workflow, if the molecular structure is known, a machine learning model can be used that encodes the chemical structure, or aspects thereof, as a latent representation of a vector with a certain dimension D, where D is a hyperparameter of the model determined during training. This model can be designed to encode any possible chemical compound, or any chemical compound containing the necessary chemical components, into the same latent space that can be used in synthesis or chemical processes, such that all chemical compounds, regardless of their initial molecular structure, can be analyzed by subsequent models within the same space of chemical representations. In the case of chemical process parameters, a similar technique can be employed, transforming different processes and parameters and / or groups of processes and parameters into latent representations that can be used in subsequent machine learning models to infer the interactions of those processes and parameters in the learned latent space.

[0022] Figure 1 This is a schematic diagram of a non-limiting example embodiment of a system for predicting the properties of chemical processes using machine learning, according to various aspects of this disclosure. In system 100, the material property prediction calculation system or MPP calculation system 102 is configured to receive information about experiments from one or more laboratories, including a first laboratory 104, a second laboratory 106, and a third laboratory 108. The illustration of three laboratories is merely a non-limiting example, and in some embodiments, more or fewer laboratories may provide information to the MPP calculation system 102.

[0023] Each of laboratories 104, 106, 108, or a combination thereof, may be associated with one or more computing systems configured to record aspects of experiments performed by the laboratories (e.g., chemicals used, process steps applied, parameters of the process steps, measurements / experimental results of the properties of chemical compound products), and report these aspects to MPP computing system 102 using structured, semi-structured, or unstructured representations. In some embodiments, instead of having a separate computing system at each of laboratories 104, 106, 108, one or more of laboratories 104, 106, 108 may directly record aspects of the experiments to MPP computing system 102 via a user interface provided by MPP computing system 102. By collecting information about the experiments from multiple laboratories, MPP computing system 102 can generate data that can be used to train its machine learning model, and thereby create more accurate predictions of chemical product properties and chemical products based on novel experiments.

[0024] In some implementations, the first laboratory 104, the second laboratory 106, the third laboratory 108, or a combination thereof, may be controlled by a single entity, thus allowing the MPP computing system 102 to aggregate information from any or all of these laboratories to form a single training dataset. This enables the MPP computing system 102 to provide significant benefits to a single entity. Individual laboratories controlled by a single entity can benefit from predictions generated by the MPP computing system 102 using collective information generated by any or all of these laboratories controlled by the single entity. This can be particularly significant in large organizations, where multiple laboratories throughout the organization may not be otherwise equipped to share data or experimental results in this manner, and each of the multiple laboratories throughout the organization may not individually generate enough information to allow machine learning models to be effectively trained.

[0025] In some implementations, at least one of the first laboratory 104, the second laboratory 106, and the third laboratory 108 may be controlled by different entities. In some such implementations, the MPP computing system 102 may partition the information received from laboratories controlled by different entities and train a separate machine learning model for each entity to prevent the identification of proprietary information among entities. In other such implementations, due to the low probability of proprietary information transfer, or because entities prefer to balance the risks of proprietary information transfer in order to access a larger amount of training data, the MPP computing system 102 may train at least some of its machine learning models on all received information, regardless of the source.

[0026] In some implementations, the MPP computing system 102 can also communicate with other components within system 100 to further enhance its functionality. For example, the MPP computing system 102 can send instructions to one or more automated laboratories 112 and receive experimental results from them. These instructions may include the identity and quantity of one or more chemicals to be used, one or more chemical reaction / process steps, and / or parameters associated with each chemical process. Experimental results may include measurements of one or more properties of the chemical product / chemical process outcome.

[0027] In some implementations, the MPP computing system 102 may also communicate with a reference data storage 114. The reference data storage 114 may include further information about chemicals, which may otherwise be identified by chemical identifiers such as a Chemical Abstracts Service (CAS) number or other unique identifiers. The reference data storage 114 may also include information about one or more machines used to perform process steps, such as capabilities, configurable settings, etc., which process steps can be located using identifiers such as model numbers or serial numbers.

[0028] In some implementations, system 100 may also include one or more end-user computing devices 110. End-user computing devices 110 may be used to access the user interface provided by MPP computing system 102 for inputting information, viewing results, or for any other purpose.

[0029] Figure 2 This is a block diagram illustrating various aspects of a non-limiting example embodiment of a material property prediction computational system (MPP computational system) according to various aspects of this disclosure. The illustrated MPP computational system 102 can be implemented by any computing device or collection of computing devices, including but not limited to one or more computing devices such as desktop computing devices, laptop computing devices, mobile computing devices, server computing devices, cloud computing systems, and / or combinations thereof. As described above, the MPP computational system 102 is configured to receive information from one or more laboratories regarding an executed chemical process, use that information to train one or more machine learning models to predict the properties of the output of the chemical process, and use those machine learning models to predict novel chemical processes.

[0030] As shown, the MPP computing system 102 includes one or more processors 202, one or more communication interfaces 204, model data storage 218, chemical data storage 208, process data storage 212, experimental data storage 214, result data storage 216, and computer-readable medium 206.

[0031] As used herein, “data storage” means any suitable device configured to store data for access by a computing device. One example of a data storage device is a highly reliable, high-speed relational database management system (DBMS) that runs on one or more computing devices and is accessible via a high-speed network. Another example of a data storage device is a key-value store. However, any other suitable storage technology and / or device capable of providing stored data quickly and reliably in response to queries may be used, and the computing device may be locally accessible rather than network-accessible, or may be provided as a cloud-based service. A data storage device may also include data stored in an organized manner on a computer-readable storage medium such as a hard disk drive, flash memory, RAM, ROM, or any other type of computer-readable storage medium. Those skilled in the art will recognize that the individual data storage devices described herein may be combined into a single data storage device, and / or the single data storage device described herein may be divided into multiple data storage devices, without departing from the scope of this disclosure.

[0032] In some embodiments, model data storage 218 is configured to store machine learning models that have been trained to predict the properties of the outputs of a chemical process. In some embodiments, chemical data storage 208 is configured to store structured chemical representations of chemicals to be used in one or more chemical processes. In some embodiments, process data storage 212 is configured to store structured process representations of the steps and parameters of a chemical process. In some embodiments, experimental data storage 214 is configured to store representations of experiments, which include a combination of representations of chemicals and chemical processes to be performed on the chemicals. In some embodiments, results data storage 216 is configured to store measured properties of experimental chemical products for use in training the machine learning models discussed herein.

[0033] As used herein, “computer-readable medium” means any removable or non-removable device that implements any technology capable of storing information in a volatile or non-volatile manner for reading by a processor of a computing device, including, but not limited to, the following: hard disk drives; flash memory; solid-state drives; random access memory (RAM); read-only memory (ROM); CD-ROM, DVD or other disk storage; cassette tape; magnetic tape; disk storage; or combinations thereof.

[0034] In some embodiments, processor 202 may include any suitable type of general-purpose computer processor. In some embodiments, processor 202 may include one or more dedicated computer processors or AI accelerators optimized for a specific computing task, including but not limited to graphics processing units (GPUs), vision processing units (VPUs), and tensor processing units (TPUs).

[0035] In some implementations, communication interface 204 includes one or more hardware and / or software interfaces suitable for providing a communication link between components of system 100. Communication interface 204 may support one or more wired communication technologies (including but not limited to Ethernet, FireWire, and USB), one or more wireless communication technologies (including but not limited to Wi-Fi, WiMAX, Bluetooth, 2G, 3G, 4G, 5G, and LTE), and / or combinations thereof.

[0036] As shown in the figure, logic is stored on computer-readable medium 206, which, in response to execution by one or more processors 202, enables MPP computing system 102 to provide encoder execution engine 210, property prediction engine 220, and model training engine 222. As used herein, "engine" refers to logic embodied in hardware or software instructions that can be written in one or more programming languages, including but not limited to C, C++, C#, COBOL, JAVA™, PHP, Perl, HTML, CSS, JavaScript, VBScript, ASPX, Go, and Python. Engines can be compiled into executable programs or written in interpreted programming languages. Software engines can be invoked from other engines or from themselves. Generally, the engine described herein refers to a logic module that can be combined with other engines or can be divided into sub-engines. Engines can be implemented by logic stored in any type of computer-readable medium or computer storage device and can be stored on and executed by one or more general-purpose computers, thereby creating a special-purpose computer configured to provide the engine or its functionality. Engines can be implemented by logic programmed into application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or another hardware device.

[0037] In some embodiments, encoder execution engine 210 is configured to retrieve an encoder from model data memory 218 and use the encoder to encode structured representations of chemicals and chemical processes into corresponding latent spaces, and to transform combined experimental representations from the latent space into latent experimental representations. In some embodiments, property prediction engine 220 is configured to retrieve a decoder from model data memory 218 and use the decoder to generate predicted properties based on latent experimental representations. In some embodiments, model training engine 222 is configured to improve the performance of the encoder and decoder by using the difference between measured and predicted properties of a given experiment.

[0038] The following provides a further description of the configuration of each of these components.

[0039] Figure 3This is a schematic diagram of a non-limiting example embodiment of an encoder / decoder for processing workflow information according to various aspects of this disclosure. The workflow to be processed is described by input workflow information 302, which includes chemical data 304 and process data 306. Chemical data 304 may contain information including, but not limited to, molecular structure, CAS number, bulk material properties, formulation components (e.g., weight, ratio, etc.), and / or combinations thereof, describing the composition and properties of a natural substance that is incorporated into or manipulated as part of a chemical workflow to transform and create new chemicals or other forms of substances. Process data 306 (which may be optional) may describe the processes, methods, and / or techniques to be used to process the components described by chemical data 304; and / or the setup and / or conditions of the environment and equipment used to process the components described by chemical data 304.

[0040] Chemical data 304 and process data 306 can be provided in any suitable format. In some embodiments, at least some of the data in chemical data 304 and / or process data 306 may be provided in a structured format, including but not limited to tabular, JSON, or BSON formats. In some embodiments, at least some of the data in chemical data 304 and / or process data 306 may be provided as unstructured data. Unstructured data may be provided in any suitable form, including but not limited to natural language descriptions of chemical components and / or process steps (e.g., laboratory notes, patent publications, white papers, etc.); captured images, audio, or video of chemical components and / or process steps; or other forms.

[0041] Input workflow information 302 is provided to encoder 308, which transforms the input workflow information 302 into an intermediate latent representation 310. The intermediate latent representation 310 is consumed by decoder 312, which generates predictions associated with the input workflow information 302. Encoder 308 may also share intermediate state information from various stages of encoder 308 with decoder 312, wherein optional transformation 320 is applied to the intermediate state information before being provided to decoder 312.

[0042] One example of a prediction that can be generated by decoder 312 is workflow prediction 314, which can represent a prediction of the chemical output generated when the steps described by process data 306 are applied to the components described by chemical data 304. Another example of a prediction that can be generated by decoder 312 is workflow latent representation 316, which can represent the features of the workflow extracted by encoder / decoder 300. Yet another example of a prediction that can be generated by decoder 312 is generated workflow information 318. The generated workflow information 318 may include chemical data and (optionally) process data (similar to input workflow information 302), but has generated differences between the generated workflow information 318 and the input workflow information 302, which are predicted to achieve the desired result (e.g., desired changes in the resulting chemical properties, desired changes in process steps, etc.). Any suitable type of encoder 308 and decoder 312 can be used, which are typically complementary machine learning models with appropriate architectures.

[0043] In some implementations, the latent representations (i.e., intermediate latent representation 310 and workflow latent representation 316) may be human-readable encodings of the input workflow information 302, or they may be high-dimensional or low-dimensional arrays of values ​​that encode the input workflow information 302 in a manner similar to a fingerprint or biometric feature (not necessarily human-readable). In some implementations, the latent representation may be the explicit output of encoder 308, which may be stored in isolation as a separate entity from encoder 308 and decoder 312. In some implementations, the latent representation may be implicit, as information combined within the encoder 308 / decoder 312 architecture, rather than as separate entities. That is, intermediate latent representation 310 and / or workflow latent representation 316 can be extracted from any part of the architecture of encoder 308 and / or decoder 312.

[0044] In some embodiments, chemical data 304 and process data 306 may be encoded by a single encoder 308 into a combined intermediate latent representation 310. In some embodiments, chemical data 304 and process data 306 may be encoded separately, wherein the individual latent representations are combined into a new latent representation that represents the entire workflow. In some embodiments, the encoding of chemical data 304 and / or process data 306 may be further subdivided into multiple sub-components using different sub-encoders trained to process specific types of data (e.g., chemical structures, audio, video, material property measurements / characterizations, etc.) or data from specific applications (e.g., paints / coatings, polymers, adhesives, etc.), wherein the latent representations generated by the sub-encoders are combined at later points in the processing.

[0045] In some implementations, a transformer architecture can be used to implement encoder 308 and / or decoder 312. Many types of transformers can be used, depending on the modality of the input workflow information 302. Some transformer architectures can process multiple modalities simultaneously, while others can be used to encode a specific modality (e.g., text, audio, or video) into a latent representation that can be injected as context data. In this way, encoder / decoder 300 can combine modalities of input workflow information 302 or other data, which can then be injected as context data (as described in further detail below). This enables encoder / decoder 300 to combine modalities into a form suitable for latent representation fusion, as described in further detail below. Some types of transformer architectures can be adapted to be used as encoder 308, others can be adapted to be used as decoder 312, and still others can be adapted to be used as both encoder 308 and decoder 312.

[0046] Table 1 provides a non-restrictive and non-exhaustive list of suitable transformer architectures, along with indications of whether these architectures are suitable for use as encoder 308 or decoder 312. In cases where a transformer is suitable as encoder 308 but not as decoder 312, the resulting latent representation generated by encoder 308 can be decoded by combining it with a suitable transformer decoder 312 or any other suitable type of machine learning model (e.g., deep neural models (multilayer perceptron (MLP), recurrent neural network (RNN), etc.)) or statistical models (e.g., random forest, support vector machine (SVM), linear / logistic regression, etc.). Essentially, a transformer used as encoder 308 can generate rich latent representations that can then be used with various different types of models used as decoder 312 to transform those latent representations into some kind of prediction or generation. Figures 4A-4C This is a schematic diagram of a technique for predicting the properties of the outputs of a chemical process, based on various aspects of this disclosure. Figures 4A-4C The techniques illustrated are a non-limiting example of the processing of encoder / decoder 300, in which chemical data 304 and process data 306 are encoded separately, and the underlying representations are later combined in the analysis.

[0047] Figure 4ANon-limiting example embodiments of processes for encoding chemicals according to various aspects of this disclosure are illustrated. As used herein, chemicals are inputs to chemical processes and may include information relating to one or more aspects or portions thereof. For example, a chemical may include a name, identifier, description, concentration, or other information relating to one or more aspects or portions thereof. As an example, the relevant chemical aspect or portion may include an electrophilic agent, such as a carbonyl carbon or other positively or partially positively charged atoms; or a nucleophilic agent, such as an amine, an oxygen atom, or other negatively or partially negatively charged atoms.

[0048] First, the structured chemical representation 402 is received. The structured chemical representation 402 provides information relating to each aspect or part of the process for each chemical molecule or component in a processing-friendly format, and is typically provided in a specially formatted, human-readable text format. As a non-limiting example, the information may be provided as a set of key-value pairs in JavaScript Object Notation (JSON) format.

[0049] In other implementations, other intermediate formats, including but not limited to XML, CSV, or YAML, may be used to functionally describe the data in an equivalent manner, since the chemical encoder 406 is capable of tagging and encoding any data type as plain text. In some implementations, chemical information (and / or process information described below) may be represented explicitly or implicitly in a graphical structure. For example, chemical information may be provided as an array of JSON records, where each record in the array includes a key named “step” that implies a temporal order of the data, and the data may be represented as a graph. Similarly, the SMILES string representation implies an equivalent graphical representation of molecules, where the nodes of the graph are atoms with certain properties (e.g., mass, charge, valence, etc.), and the edges of the graph are bonds with bond properties (e.g., degeneracy, orbital configuration, etc.).

[0050] In some implementations, unstructured text may also be provided as a chemical representation, or together with structured chemical representation 402, as a value in the key-value pairs of structured chemical representation 402, or may completely replace structured chemical representation 402. For example, a general description of a chemical, relevant aspects or parts of a chemical, the properties of a chemical, the use of a chemical, or the reason why a chemical is included in a process may be provided.

[0051] Figure 5 These are non-limiting example embodiments of structured chemical representations according to various aspects of this disclosure. The illustrated structured chemical representations are in JSON format, but as noted above, any other structured format can be used for structured chemical representations, and some information can also be provided in unstructured formats.

[0052] The illustrated structured chemical representation includes information on two materials, “Raw Material A” and “Raw Material B”, each represented by a set of key-value pairs. As illustrated, the key-value pairs include the name, the American Chemical Abstracts Service (CAS) registry number, a linear symbol string identifying the material using the Simplified Molecular Input Linear Input System (SMILES) format, and the concentration. These key-value pairs are shown only as non-limiting examples, and in other embodiments, more, fewer, or different key-value pairs may exist. Furthermore, more or fewer materials may be included within the structured chemical representation. Other non-limiting examples of information that can be provided for one or more materials in a structured chemical representation include other chemical identifiers, other types of linear symbols (e.g., polymer SMILES (PSMILES) strings, International Chemical Identifier (InChI) strings, Wiswesser Linear Symbol (WLN) strings, Smiles Arbitrary Target Specification (SMARTS) strings, SYBYL Linear Symbol (SLN) strings, etc.), empirical properties (e.g., boiling point, melting point, glass transition temperature, etc.), molecular descriptors (e.g., specific aspects of a molecule, specific parts within a molecule, specific molecular functionality, charge distribution, etc.), or any other relevant information used to describe a chemical compound.

[0053] return Figure 4A After obtaining the structured chemical representation 402, the structured chemical representation 402 subsequently undergoes one or more optional preprocessing steps 404. These one or more preprocessing steps 404 may include various actions to prepare the contents of the structured chemical representation 402 for further processing, including but not limited to flattening the structured chemical representation 402 by removing newlines, converting the string to a tokenized format (e.g., breaking the string down into individual words, punctuation marks, and / or other tokens, etc.), or any other suitable techniques for preprocessing. In some embodiments, preprocessing 404 may include retrieving additional information about one or more chemicals described in the structured chemical representation 402. For example, if a given chemical description in the structured chemical representation 402 includes a CAS registry number value, reference data storage 114 may be queried to retrieve additional information associated with the CAS registry number that is not otherwise provided in the structured chemical representation 402.

[0054] Then, after an optional preprocessing step 404, the structured chemical representation 402 is provided to a chemical encoder 406, which creates a latent chemical representation 408. The latent chemical representation 408 is an encoded version of the structured chemical representation 402 in a latent space, such that similar values ​​in the latent space indicate chemicals having similar properties, aspects, or parts, and different values ​​in the latent space indicate chemicals having different properties, aspects, or parts.

[0055] Any suitable architecture can be used for the chemical encoder 406, and the characteristics of the latent space (e.g., a mapping from the structured chemical representation 402 to features of the latent space) can be learned during the training of the chemical encoder 406. In some embodiments, encoder architectures commonly used for text processing can be used. It has been found that text processing encoder / decoder architectures are efficient in encoding such information into a latent space for predicting properties if the chemical information is provided in a human-readable format, and further efficiency can be obtained if the format is constructed as discussed above. Because text processing encoder / decoder architectures are able to compensate for differences in precise terminology, context, spelling, and other aspects of the input text, there is no need to impose strict format requirements on the structured chemical representation 402 (i.e., a list of key names that must be used, a list of key-value pairs that must exist, etc.), thereby improving the flexibility and predictive power of the embodiments of this disclosure.

[0056] In some implementations, an encoder architecture utilizing self-attention can be used. For example, in some implementations, a bidirectional encoder representation (BERT) from a transformer can be used for the chemical encoder 406. The BERT encoder is described in Devlin et al., “BERT: Pre-training a Deep Bidirectional Transformer for Language Understanding,” arXiv: 1810.04805v2, May 24, 2019 (the entire contents of which are hereby incorporated by reference for all purposes). To provide the structured chemical representation 402 to the chemical encoder 406, the structured chemical representation 402 is first labeled (this labeling can occur during preprocessing 404). Each label is transformed into a unique identifier within a vocabulary that serves as the basis for the chemical encoder 406, and words not present in the vocabulary are decomposed into subwords mapped to labels. A one-dimensional vector is created to represent the embedding of each label. The input is then encoded into a latent space using the embeddings to create a latent chemical representation 408.

[0057] Figure 4BNon-limiting example embodiments of processes for encoding chemical processes according to various aspects of this disclosure are illustrated. As used herein, a chemical process is a set of operations performed on a chemical or a combination of two or more chemicals. For example, a chemical process may include performing one or more steps such as: adding chemicals in a specified order; using a solvent; heating; cooling; stirring; bubbling gas; purging air or other gases; placing chemicals under a specific atmosphere such as an inert atmosphere (e.g., nitrogen, argon) or a reactive atmosphere (e.g., oxygen, hydrogen, other reactive chemical gases); placing chemicals under a vacuum (zero or near-zero pressure), increasing or decreasing pressure; freezing-pumping-thawing; introducing light of a specified wavelength or wavelength range; removing light; removing components (e.g., water, gases, solvents, byproducts); separating chemicals; purifying chemicals; drying chemicals; or any other chemical process used to perform a desired chemical reaction process and known to those skilled in the art.

[0058] Chemical processes can utilize a variety of equipment used to perform them. For example, chemical processes may utilize non-reactive glassware or other containers; stirring plates; heating elements; temperature regulators; cooling baths (e.g., ice, salt, liquid nitrogen, dry ice containing inorganic or organic solvents (e.g., ammonia, sulfur hexafluoride, fluorosulfonyl chloride, acetone, isopropanol), saline water (e.g., calcium chloride, sodium chloride)); pressure regulators; gas cylinders; inert atmosphere equipment; light sources (e.g., light emission, UV, visible); light shields; vacuum pumps; valves; desiccants; specialized glassware (e.g., dehydration, condensation); piping; connectors; evaporators; condensers; ovens; computers or machines (e.g., for the controlled introduction of chemicals, gases, solvents, etc.); equipment for removing chemicals, gases, solvents, etc.; for purifying chemicals; for characterizing chemicals, etc.); chromatographic equipment; and / or any other equipment known and used by those skilled in the art for performing chemical processes.

[0059] Any chemical process in one or more chemical processes can be configured using certain parameters for performing the chemical process. For example, parameters may include one or more of the following: a specified stirring frequency; a specified temperature for heating and / or cooling; a specified gas pressure or the absence of pressure; a specified wavelength or wavelength range of light; a specified time period (e.g., for heating, cooling, stirring, introducing light, evacuating, drying, adding chemicals or solvents, etc.); a specified rate (e.g., for stirring, introducing chemicals, removing chemicals, adding solvents, removing solvents, chromatographic purification of chemicals, etc.); or any other parameters known to those skilled in the art for effectively performing a chemical process.

[0060] Chemical processes produce chemical products, which can be intended products, final products, or intermediate products to which one or more additional chemical processes are performed, or chemical products as outputs for which one or more properties can be tested.

[0061] To encode the chemical process, a structured process representation 410 is first received. Similar to the structured chemical representation 402, the structured process representation 410 provides information related to the configuration of the chemical process and may be provided in a processing-friendly format, such as a specially formatted human-readable text format. Like the structured chemical representation 402, the structured process representation 410 may be provided in formats such as JSON, XML, CSV, YAML, or any other format that can be marked and encoded as plain text by the process encoder 414, and the information may be provided as a set of key-value pairs. In some embodiments, unstructured text may also be provided along with, or alongside, the structured process representation 410 as values ​​in the key-value pairs of the structured process representation 410, or in lieu of the structured process representation 410. For example, a general description of the chemical process, the steps of the chemical process, the benefits and / or disadvantages of the chemical process, the output of the chemical process, or any other information related to the chemical process may be provided.

[0062] Figure 6A and Figure 6B These are non-limiting example embodiments of a structured process representation based on various aspects of this disclosure, wherein Figure 6B Examples are shown in Figure 6A This is a continuation of the structured procedure representation that began in [the previous section]. The illustrated structured procedure representation uses JSON format, but as noted above, any other structured format can be used for structured procedure representation, and some information can also be provided in unstructured formats.

[0063] The illustrated structured process representation includes two types of information: (1) the batch or general process to be executed (e.g. Figure 6A (as illustrated), and (2) specific parameters associated with the process (such as...) Figure 6B (As illustrated). In general, the batch information and parameter information detail the steps of the chemical process to be performed by a chemist, an automated workbench, or any other suitable device for performing the chemical processes disclosed herein.

[0064] Batch information and parameter information may also specify one or more of the following: the order in which steps should be performed and / or the order in which chemicals should be added, specific operating conditions for each step, and / or properties to be measured under different conditions of interest in order to characterize the output of the chemical process. In some embodiments, batch information and parameter information may also include condition values ​​for one or more steps, which, based on parameters, the results of another step, or any other suitable conditions, indicate the conditions under which a given step may or may not be performed.

[0065] Figure 6A and Figure 6B The illustrated structured process representation shows information about the mixing step (“mixing component”) and the second curing step (“curing stage”), where each step is represented by a set of key-value pairs. Figure 6A The set of key-value pairs for the mixing step, named “Mixing Component”, includes multiple pairs that identify the step sequence (“1”), the equipment used to perform the action (“Mixer-E001”), the frequency of the mixing action (“200”), the unit for the mixing setting (“Hz”), the chemicals to be mixed (“Raw Material B”), the identity of the chemicals (identified by “SMILES”), and the amount of chemicals to be used (with a concentration of “0.3”).

[0066] Figure 6B The parameters illustrated herein represent information including two parameters. The first parameter, named “PG509”, provides information related to the mixer, the specific equipment to be used (“mixer 1”), the specific temperature at which mixing will occur (“35°C”), the mixing speed (“200 RPM”), the volume (“1 mL”), etc.

[0067] The second parameter, named “PG4234”, provides information related to the “curing stage”. This information includes multiple pairs that identify the step sequence (“2”), the equipment used to perform the action (“curing oven-E111”), and the specific settings for that step, including the introduction of light and the identification of the type of light (“UV”), the wavelength of the light to be used (“280”), the unit of wavelength (“nm”), the duration of curing (“curing time”) (“60”), and the unit of curing duration (“s”).

[0068] Despite Figure 6A or Figure 6BNot illustrated herein, but in some embodiments, the structured process representation may include information about the property to be measured and may include one or more steps for measuring that property. For example, the property may include any property of interest, such as boiling point, melting point, ion exchange capacity, solubility, stability under certain conditions (e.g., exposure to air or sunlight), one or more acoustic properties, one or more chemical properties, one or more electrical properties, one or more magnetic properties, one or more mechanical properties, one or more optical properties, one or more thermal properties, any other property of particular interest to the end user, or combinations thereof.

[0069] return Figure 4B After obtaining the structured process representation 410, in some embodiments, the structured process representation 410 subsequently undergoes one or more optional preprocessing steps 412. Figure 4A Similar to the optional preprocessing step 404, the one or more preprocessing steps 412 may include various actions to prepare the contents of the structured process representation 410 for further processing, including but not limited to flattening the structured process representation 410 by removing newline characters, converting strings to a tokenized format (e.g., breaking down strings into individual words, punctuation marks, and / or other tokens), or any other suitable techniques for preprocessing. In some embodiments, preprocessing 404 may include retrieving additional information about one or more machines described in the structured process representation 410. For example, if a given machine referenced in the structured process representation 410 is identified by a model or serial number, reference data memory 114 may be queried to retrieve additional information associated with the machine that is not otherwise provided in the structured process representation 410.

[0070] Then, after undergoing optional preprocessing 412, the structured process representation 410 is provided to the process encoder 414, which creates a latent process representation 416. The latent process representation 416 is an encoded version of the structured process representation 410 in a latent space, such that similar values ​​in the latent space indicate chemical processes and parameters with similar characteristics, and different values ​​in the latent space indicate chemical processes and parameters with different characteristics.

[0071] Any suitable architecture can be used for the process encoder 414, and during the training of the process encoder 414, the properties of the latent space (e.g., a mapping from the structured process representation 410 to features of the latent space) can be learned. Similar to the chemical encoder 406, encoder architectures typically used for text processing can be used for the process encoder 414, thus providing the same benefits discussed relative to the chemical encoder 406. In some embodiments, encoder architectures using self-attention (including, but not limited to, the BERT encoder) can be used for the process encoder 414, wherein the structured process representation 410 is provided to the process encoder 414 using techniques similar to those discussed above relative to providing the structured chemical representation 402 to the chemical encoder 406 (including tokenization, conversion to unique identifiers, etc.).

[0072] Figure 4C Non-limiting example embodiments of processes for combining and encoding chemicals and chemical processes, encoding combination representations to create coded experiments, and predicting properties based on coded experiments are illustrated according to various aspects of this disclosure.

[0073] First, the potential chemical representation is 408 (see...) Figure 4A ) and potential process representation 416 (see Figure 4B The two components are combined to create a combined experimental representation 418. The combined experimental representation 418 is intended to represent a process in which a chemical process and parameters represented by the potential process representation 416 are applied to a chemical substance represented by the potential chemical representation 408. The potential chemical representation 408 and the potential process representation 416 can be combined in any suitable manner. In some embodiments, the potential chemical representation 408 may simply be cascaded with the potential process representation 416 in any order.

[0074] The combined experimental representation 418 is provided to the experimental encoder 420, which creates a latent experimental representation 422. Any suitable architecture can be used for the experimental encoder 420, including text processing architectures using self-attention, including but not limited to the BERT encoder. Similarly, during the training of the experimental encoder 420, properties of the latent space (e.g., a mapping from the combined experimental representation 418 to features of the latent space) are learned, and the latent space is such that similar values ​​within the latent space indicate experiments with similar results, and different values ​​within the latent space indicate experiments with different results. It will be noted that by using an architecture utilizing self-attention, the chemical encoder 406, the process encoder 414, and the experimental encoder 420 encode both local and global contextual information into representations in their respective latent spaces. Therefore, the same chemical or parameter repeatedly used in two separate experiments may have different representations in the latent space, given the global context in which the chemical or parameter is in the separate experiment.

[0075] Finally, the latent experimental representation 422 is provided to the experimental decoder 424, which generates the predicted property 426 based on the latent experimental representation 422. Any suitable architecture can be used for the experimental decoder 424. In some embodiments, an experimental decoder 424 architecture complementary to the architectures of the chemical encoder 406, process encoder 414, and experimental encoder 420 can be used, such as a decoder associated with the BERT encoder. In some embodiments, the architecture can use a stochastic process (including but not limited to a Gaussian process) as the experimental decoder 424. In such architectures, any suitable type of kernel can be used for the Gaussian process, including but not limited to radial basis kernels. In some embodiments, an artificial neural network can be used as the experimental decoder 424. The experimental decoder 424 treats changes in the latent experimental representation 422 as representing changes in the predicted property 426 and determines the mapping between the latent experimental representation 422 and the predicted property 426 during training.

[0076] Figure 7A and Figure 7B This is a flowchart illustrating a non-limiting example embodiment of a method for predicting the properties of the outputs of a chemical process according to various aspects of this disclosure. In method 700, the MPP computing system 102 uses a machine learning model to generate coded representations of chemicals and chemical processes, and predicts properties based on these coded representations. As illustrated, method 700 also shows techniques for training the machine learning model to initially train the model or retrain a previously trained model to improve its accuracy. By training the encoder to encode chemicals and chemical processes into representations in the latent space, accurate predictions can be generated for novel chemicals and / or chemical processes (i.e., chemicals and / or chemical processes that the MPP computing system 102 has not previously considered and / or used to train the encoder or decoder). Method 700 describes, for example, Figures 4A-4C The training and use of the splitting model illustrated herein will be recognized, but similar techniques can be applied to... Figure 3 It is used in conjunction with the encoder / decoder 300, with structured or unstructured input data, and / or any other model architecture or information described herein.

[0077] From the start box, method 700 proceeds to box 702, where the MPP calculation system 102 receives a structured chemical representation 402. In some embodiments, the MPP calculation system 102 may receive the structured chemical representation 402 by providing a user interface that allows an end user to input or upload the structured chemical representation 402. In some embodiments, the MPP calculation system 102 may receive the structured chemical representation 402 by retrieving it from a chemical data storage 208 associated with the MPP calculation system 102, wherein the structured chemical representation 402 has been previously stored in the chemical data storage, such as during a previous execution of method 700, or as part of building a library of structured chemical representations 402 intended for use in experiments. In some embodiments, the MPP calculation system 102 may present a user interface that displays information about the composition of chemicals in the previously stored structured chemical representation 402 and allows the user to select one or more previously stored chemicals to be included in the new structured chemical representation 402.

[0078] At box 704, the encoder execution engine 210 of the MPP computation system 102 uses the chemical encoder 406 to generate a potential chemical representation 408 based on the structured chemical representation 402. In some embodiments, the MPP computation system 102 may retrieve the chemical encoder 406 from the model data storage 218 associated with the MPP computation system 102 and use the retrieved chemical encoder 406 to generate the potential chemical representation 408 based on the structured chemical representation 402. In some embodiments, the chemical encoder 406 may have been previously partially or fully trained using method 700.

[0079] At block 706, MPP computing system 102 receives structured process representation 410. In some embodiments, MPP computing system 102 may receive structured process representation 410 by providing a user interface that allows an end user to input or upload structured process representation 410. In some embodiments, MPP computing system 102 may receive structured process representation 410 by retrieving it from process data storage 212 associated with MPP computing system 102, wherein structured process representation 410 has been previously stored in process data storage, such as during a previous execution of method 700, or as part of building a library of structured process representations 410 intended for use in experiments.

[0080] In some implementations, the structured process representation 410 can be constructed by combining elements from a previously created structured process representation 410. For example, the user interface can present information about previously stored process steps and / or parameters, and can allow the end user to select one or more previously stored steps and / or one or more previously stored parameters to help create a structured process representation 410 for a new chemical process.

[0081] At box 708, encoder execution engine 210 uses process encoder 414 to generate latent process representation 416 based on structured process representation 410. In some embodiments, MPP computing system 102 may retrieve process encoder 414 from model data storage 218 associated with MPP computing system 102 and use the retrieved process encoder 414 to generate latent process representation 416 based on structured process representation 410. In some embodiments, process encoder 414 may have been previously partially or fully trained using method 700.

[0082] At box 710, encoder execution engine 210 combines latent chemical representation 408 and latent process representation 416 to create combined experimental representation 418. As described above, any suitable technique can be used to combine latent chemical representation 408 and latent process representation 416, including but not limited to simply cascading the two representations in the order expected by experimental encoder 420.

[0083] Method 700 then proceeds to the continuing terminal (“Terminal A”). From Terminal A ( Figure 7B Method 700 proceeds to block 712, where encoder execution engine 210 uses experimental encoder 420 to generate potential experimental representation 422 based on combined experimental representation 418. In some embodiments, MPP computing system 102 may retrieve experimental encoder 420 from model data storage 218 associated with MPP computing system 102 and use the retrieved experimental encoder 420 to generate potential experimental representation 422. In some embodiments, experimental encoder 420 may have been previously partially or fully trained using method 700.

[0084] At box 714, the property prediction engine 220 of the MPP computing system 102 uses the experimental decoder 424 to generate one or more predicted properties 426 of the output of a chemical process based on the latent experimental representation 422. In some embodiments, the property prediction engine 220 may utilize one or more types of properties pre-configured to make predictions for all latent experimental representations 422 provided to the property prediction engine. In some embodiments, the latent experimental representation 422 may suggest properties to be predicted for associated experiments. Once one or more predicted properties 426 have been generated, the property prediction engine 220 may store the predicted properties 426 in the results data storage 216 associated with the MPP computing system 102, and / or may generate a user interface to present the predicted properties 426 to an end user.

[0085] Method 700 then proceeds to decision box 716, where it is determined whether method 700 is in training mode or normal prediction mode. In some embodiments, this determination may be based on user configuration (i.e., whether the user has already placed method 700 in normal prediction mode or training mode). In some embodiments, this determination may be based on whether a predetermined number of initial training iterations have been performed or whether the model's performance has converged to an acceptable level, and if no iterations have been performed or the performance is not yet acceptable, training mode is determined. In some embodiments, this determination may be based on how long it has been since one or more of the chemical encoder 406, process encoder 414, experimental encoder 420, or experimental decoder 424 were previously retrained.

[0086] If it is determined that method 700 is in normal prediction mode, the result of decision box 716 is negative, and method 700 proceeds to the end box and terminates. Otherwise, if it is determined that method 700 is in training mode, the result of decision box 716 is positive, and method 700 proceeds to box 718. At box 718, the model training engine 222 of the MPP computing system 102 receives one or more measured properties from the laboratory performing the experiment. The measured properties are considered as the true value properties of the experiment. In some embodiments, the measured properties may be sent by the laboratory along with the identifier of the experiment that led to the measured properties to be stored in the result data storage 216 associated with the MPP computing system 102. In some embodiments, instructions may be given to the laboratory to perform experiments as defined in the structured chemical representation 402 and the structured process representation 410, and the measured properties may be linked to the associated predicted properties 426 using a unique identifier of the instructions. In some implementations, the identifier of the experiment may be a copy of the potential experimental representation 422 or any portion of the data used to create the potential experimental representation 422 (e.g., a structured chemical representation 402, a structured process representation 410, a combined experimental representation 418, and / or its latent space representation). The model training engine 222 can then obtain the measured properties from the results data storage 216 by querying the measured properties that match the predicted experiment.

[0087] At box 720, model training engine 222 compares one or more predicted properties 426 with one or more measured properties to determine the value of the loss function. In some implementations, the value of each predicted property 426 is compared with the corresponding measured property, and the difference between the two is used as the loss value to be fed into an appropriate loss function, such as mean squared error.

[0088] At box 722, model training engine 222 determines the gradient of the loss function and uses this gradient to update at least one of the experimental decoder 424, experimental encoder 420, process encoder 414, or chemical encoder 406. By updating the experimental decoder 424, experimental encoder 420, process encoder 414, and / or chemical encoder 406 using the gradient of the loss function, the performance of method 700 is iteratively improved. In some implementations, the gradient can be used to update all these models during each iteration. In some implementations, some models can remain constant while others are updated for one or more iterations.

[0089] Method 700 then proceeds to decision box 724, where it determines whether method 700 has completed training the model or whether further training iterations should be performed. In some implementations, this determination may be based on whether a predetermined number of iterations have been performed. In some implementations, this determination may be based on whether the model's performance has converged to an acceptable level. If it is determined that training is not yet complete, decision box 724 results in no, and method 700 proceeds to the continuation terminal ("terminal B") to return to box 702 for one or more subsequent training iterations. Otherwise, if training is complete, decision box 724 results in yes, and method 700 proceeds to box 726.

[0090] At box 726, model training engine 222 stores the trained experimental decoder 424, experimental encoder 420, process encoder 414, and / or chemical encoder 406 in model data storage 218 of MPP computing system 102 for subsequent execution of method 700. Method 700 then proceeds to the end box and terminates.

[0091] For ease of discussion, Figure 7A and Figure 7B The example of method 700 provided illustrates the creation of a single training example (an input structured chemical representation 402 and a structured process representation 410 paired with a set of measured properties). In some implementations, training the model may involve collecting a large number of training examples with different structured chemical representations 402, different structured process representations 410, and different measured properties before putting method 700 into training mode. In some implementations of system 100, by combining experimental results from multiple laboratories, such as all laboratories throughout a large research institution, all of which use the MPP computing system 102 to store records of experiments and results, a sufficient number of training examples can be collected fairly quickly to train the model to a high performance level.

[0092] By providing the ability to predict the outcomes of chemical processes, the MPP computational system 102 greatly enhances the ability of chemists and other researchers to explore previously untested combinations of chemicals when searching for new chemicals with desired properties. The predictions provided by the MPP computational system 102 allow researchers to avoid physically performing experiments that are unlikely to produce the desired results and focus on physically performing experiments that are more likely to have successful outputs. This reduces waste of materials, time, and other laboratory resources.

[0093] The combination of the MPP computing system 102 and the automated laboratory 112 can provide further benefits. For example, if there are regions in the potential space of the potential chemical representation 408, the potential process representation 416, or the potential experimental representation 422 that are not well covered by the training data already present in the result data storage 216, the MPP computing system 102 can automatically create experiments that will cover these missing regions, send the experiments to the automated laboratory 112, and add the automatically generated results to the result data storage 216 to improve the coverage of the training data.

[0094] Furthermore, once the chemical encoder 406, process encoder 414, experimental encoder 420, and experimental decoder 424 are fully trained, the steps of the non-training mode of method 700 can be reverse-engineered to perform inverse design of chemicals and / or chemical processes. For example, a set of starting chemicals and starting chemical processes, along with one or more desired properties, can be provided. Method 700 can be executed to generate one or more predicted properties, which are then compared with one or more desired properties to determine a loss. This loss can then be backpropagated through the experimental decoder 424 and experimental encoder 420 to determine gradients for updating the potential chemical representation and / or potential process representation. These updated potential representations can then be used to update the input chemicals and / or input chemical processes in an attempt to bring the predicted properties closer to the desired properties, and the updated information can be simulated again. By iterating this prediction / comparison / backpropagation / update loop, the MPP computational system 102 can comprehensively explore the search space of potential chemical and chemical process combinations without actually performing physical experiments, thus significantly reducing the research effort and resources required to obtain the desired properties.

[0095] The above description envisions a single chemical encoder 406 associated with a structured chemical representation 402 and a single process encoder 414 associated with a structured process representation 410. In some embodiments, multiple encoders or multiple decoders may be present. For example, encoders and / or decoders can be added to provide additional levels of detail. For example, since the structured process representation 410 includes both batch and parameter portions, separate batch encoders and parameter encoders may be provided in some embodiments. In some embodiments, value-specific encoders may be provided for well-structured key-value pairs (such as information provided in chemical linear notations) for which no format changes are expected. For example, if the chemical encoder 406 detects the presence of the string SMILES in the structured chemical representation 402, the string can be encoded using a SMILES-specific encoder before the rest of the structured chemical representation 402 is encoded.

[0096] As another example, in some implementations, multiple different encoders and / or decoders can be trained to have encoders and / or decoders considered to be domain-specific experts, thereby improving domain-specific results. For example, a first chemical encoder 406, a first process encoder 414, a first experimental encoder 420, and / or a first experimental decoder 424 can be trained using data from a first domain (e.g., experiments related to adhesives); a second chemical encoder 406, a second process encoder 414, a second experimental encoder 420, and / or a second experimental decoder 424 can be trained using data from a second domain (e.g., experiments related to inks), and so on for multiple domains (e.g., coatings, ceramics, food, polymers, etc.). In some simple implementations, the user can select a domain and use appropriate encoders and / or decoders trained using data from the selected domain. In more complex implementations, multiple encoders and / or multiple decoders can be combined in an agent network, where additional machine learning models (e.g., transformer models) are trained to determine the appropriate domain-expert encoder and / or decoder model to be used for a given experiment directly based on the structured chemical representation 402 and / or the structured process representation 410. When used in reverse engineering or in the context of “what experiment should be performed next”, implementations using proxy networks can be particularly powerful because proxy networks can leverage context learned from multiple domains to suggest additional experiments while still providing accurate domain-specific predictions.

[0097] Figure 3 The encoder / decoder 300 and illustrated in the figure Figures 4A-4C The processes illustrated describe relatively simple encoder and decoder pairs; in some implementations, more complex architectures can be used. For example, a given encoder or decoder may consist of a stack of encoders and / or decoders, such that the given encoder or decoder includes additional embedded latent representations that represent intermediate stages of information encoding and can encode multi-scale features of the information. As a non-limiting example, a given encoder may be configured to encode images of laboratory apparatus involved in a chemical workflow. Intermediate layers or sub-encoders of such an encoder may be configured to detect or analyze specific types of visual data. These layers may extract different features related to relevant chemicals, processes, or environmental factors within the workflow, including but not limited to apparatus configuration, reactant states, and / or fluid dynamics. These features detected by specific layers within the encoder can provide insights that other encoders or decoders might not capture, thus providing a deeper understanding of chemical processes at more stages that may influence analysis or decision-making.

[0098] In some implementations, a transducer architecture can be used to implement at least one encoder in an encoder. Encoders implemented using a transducer architecture may be particularly effective at processing and correlating various forms of input data due to attention mechanisms. Transducers can be configured to process various data types, including but not limited to structured / tabular data representing sensor readings, time-series data, or chemical measurements, as well as unstructured data, including but not limited to images, audio, video, or natural language text such as lab notes. For example, a transducer can ingest structured datasets (e.g., temperature readings, pressure measurements, reagent concentrations, etc.) in tabular or other formats and can learn temporal patterns or correlations associated with stages or outcomes of a workflow. As another example, a transducer can ingest video or audio input (e.g., a live stream of workflow execution or a sound recording of device operation, etc.) and can capture temporal dependencies to identify process stages, such as stirring, boiling, or other physical changes in a system performing a workflow. As yet another example, a transducer can encode textual descriptions (e.g., lab notes or procedural instructions) into meaningful representations, thereby capturing high-level context, procedural steps, and operational parameters that describe the ongoing workflow in a more abstract way. As the encoder processes these different data streams, the attention mechanism within the transducer allows it to focus on the most relevant parts of each input modality. This means the transducer can simultaneously attend to key time points (e.g., temperature spikes in sensor data), key visual elements (e.g., color changes in reactants), and / or important process cues (e.g., instructions to begin heating). The multi-scale, multi-modal feature extraction provided by the transducer architecture enables the encoder to create a comprehensive representation of the workflow, which helps in better predicting, adapting, and / or gaining insight into chemical processes.

[0099] More sophisticated encoders / decoders can optionally provide additional manipulation of the encoding / decoding process of the encoder / decoder by injecting additional context-related information at different stages of the encoding / decoding process. Figure 8 This is a schematic diagram of a non-limiting example embodiment of an encoder / decoder according to various aspects of this disclosure, wherein contextual information is optionally encoded and injected into the encoding / decoding process at various points. Encoder / decoder 800 may represent the entire encoder / decoder 300, or it may represent a sub-encoder / decoder serving as a component of a larger encoder / decoder 300. By injecting additional information at various points within encoder / decoder 800, a richer latent representation can be generated, which can better encode the complexity of the chemical workflow by considering the context at each point.

[0100] In encoder / decoder 800, input data 802 is provided to first encoder / decoder 804, which generates intermediate latent representation 806. Input data 802 can be as follows: Figure 3 The input workflow information 302 illustrated herein may also be raw data or signals input from a chemical workflow. Raw data or signal inputs may include one or more of sensor readings, video, images, time-series data, and / or other information representing an ongoing chemical process. By generating an intermediate latent representation 806 based on the input data 802, the encoder / decoder 800 can not only compress meaningful features from the input data 802 but also facilitate the reconstruction or transformation of the workflow representation for downstream tasks. Once the intermediate latent representation 806 is generated, the encoder / decoder 800 may optionally perform latent space fusion 808 to combine the intermediate latent representation 806 with other information, and the combined latent representation is provided to a second encoder / decoder 810, which generates a workflow prediction 820 (similar to...). Figure 3 Workflow prediction 314), latent representation 822 (similar to...) Figure 3 The workflow potential representation 316 in the middle) or the generated chemical workflow 824 (similar to Figure 3 One or more of the workflow information generated in (318).

[0101] At various points in the encoder / decoder 800, information from contextual data 818 can be extracted and injected into the encoding / decoding process, or directly into the latent representation. Contextual data 818 may include information representing the context in which the workflow is taking place, including but not limited to procedural information, chemical descriptor information, previous experimental results, laboratory notes, environmental factors (e.g., ambient temperature, humidity, reagent properties, etc.), or other information about the context. Contextual data 818 can represent factors influencing experimental results, allowing the encoder / decoder 800 to more accurately adjust predictions or comparisons of experiments. The richer latent representation provided by injecting context improves generalization across different environments or tasks by allowing flexible integration of both intrinsic and extrinsic data. The fusion of context and general knowledge allows the encoder / decoder 800 to be adaptable, and this fusion helps support diverse tasks and domains. In some implementations, attention mechanisms may be included as part of context encoding, ensuring that the encoder / decoder 800 emphasizes relevant context based on a given task objective.

[0102] Context data 818 can be passed to one or more optional encoders, and latent context representations output from those encoders can be provided to various points in encoder / decoder 800. As shown, a first context encoder 812 can encode context data 818 to be injected into a first encoder / decoder 804, a second context encoder 814 can encode context data 818 to be fused with an intermediate latent representation 806, and a third context encoder 816 can encode context data 818 to be injected into a second encoder / decoder 810. Each of the first context encoder 812, the second context encoder 814, and the third context encoder 816 can be trained independently and can encode different aspects of the context data 818 suitable for injection into various points of encoder / decoder 800.

[0103] Although a single encoder / decoder and context encoder block are illustrated at each location in encoder / decoder 800 for clarity, in some implementations, multiple encoders / decoders or context encoders may be provided at each point. Using multiple context encoders and / or multiple encoders / decoders at each level of the overall encoder / decoder 800 can facilitate the processing of multimodal input data 802 and / or multimodal context data 818, thereby allowing the processing of different forms of context (e.g., structured data, lab notes, audio / video, etc.).

[0104] For latent space fusion 808, any suitable technique can be used to connect the intermediate latent representation 806 and the latent representation (or other output) generated by the second context encoder 814. Figures 9A-9C Three non-limiting example techniques that can be used for latent space fusion according to various aspects of this disclosure are illustrated. These three examples demonstrate how contextual information from a molecular descriptor can be fused with an intermediate latent representation 806 of the structure of the input data to produce a fused latent representation that is related to the molecular context residing in a given environment. This type of latent space fusion improves the richness of the latent representation relative to encoder-only representations, as measured by accuracy metrics of predictions based solely on molecular descriptors, latent representations learned solely by a simple encoder, and fused latent representations involving both the encoder and context-related information injected via descriptors. In other embodiments, different types of information can be fused with the intermediate latent representation 806.

[0105] exist Figure 9A In the input field, enter data 902 (for example, such as...). Figure 8 The input data 802 described herein is provided to the encoder 906 to generate an intermediate latent representation 926. Figure 8The intermediate latent representation 806 is exemplified here. Meanwhile, the input data 902 is also associated with context data 904 (e.g., such as...). Figure 8 The context data 902 (as described in the previous section) is provided together with the context data 918 to the molecular descriptor calculator 908. In some embodiments, the molecular descriptor calculator 908 acts as a context encoder that takes input data 902 and context data 904 as input (e.g., a description of a chemical, including but not limited to a string of SMILES representing a molecule, or other structured or unstructured description of the chemical), but does not compute a learned fingerprint or other intermediate latent representation 926 (as done by encoder 906). The molecular descriptor calculator 908 computes and outputs one or more descriptors 930 about the chemical, which may represent the number of molecules specifically computed. The descriptors 930 are then fused with the intermediate latent representations 926 to create a complete latent representation 928. Because the complete latent representation 928 already incorporates context-related information, it becomes more useful in making predictions and generating results related to a given workflow.

[0106] As a non-limiting example, if the entire encoder / decoder 800 finds a potential representation that can be used to predict the reactivity of a molecule / chemical in the presence of an electromagnetic field, then descriptor 930 may include reactivity-related properties. For example, descriptor 930 may include the number of nitrogen-carbon double bonds in the molecule, empirical properties that can be used to look up the molecule's electromagnetic susceptibility at the frequency of the given electromagnetic field to be applied in a lookup table, or other reactivity-related information. These descriptors 930 can then be used as contextual information to be fused with intermediate potential representations 926, thereby allowing the complete potential representation 928 to be used for more accurate prediction and generation within a context that includes the electromagnetic field context injected as part of contextual data 904.

[0107] Figure 9A A first non-limiting example technique for fusing descriptor 930 with intermediate latent representation 926 is illustrated. As shown, intermediate latent representation 926 and intermediate latent representation 926 simply undergo cascading 910. In other words, if intermediate latent representation 926 is composed of a size of N array [a 0 a 1 , . . , a N-1 The descriptor 930 is represented by a size of ], and the descriptor 930 is represented by a size M array [d 0 d 1 , ..., d M-1 If we use ] to represent it, then cascading 910 will produce a size of N + M array [a 0 a1 , ..., a N-1 d 0 d 1 , ..., d M-1 ].

[0108] After cascading 910, the cascaded latent representation undergoes normalization 912. Any suitable normalization technique can be used. In some implementations, it can be... L p Normalization. In a general sense, L p The norm can be defined as: Once determined L p Norms can be obtained through L p The norm value is used to scale the elements of the cascaded latent representation so that the complete latent representation 928 has a norm of 1. L p Norm. In some implementations, p This can be set to 2, which represents the Euclidean norm. This transforms all the complete latent representations 928 into unit-length vectors, meaning that these unit-length vectors represent points on the surface of a multidimensional sphere. The use of normalization 912 provides numerical stability because the encoder / decoder 800 does not need to handle infinitely large or infinitely small numbers, and this encoder / decoder simplifies the distance calculations between the complete latent representations 928. These distance calculations are techniques used to train various encoders and decoders of the encoder / decoder 300 to encode similar things with similar latent representations (e.g., by including minimizing the angle between latent representations of known similar things as a training objective), and similarly maximizing the distance between latent representations of known dissimilar things.

[0109] Figure 9A The data flow illustrated herein is a non-limiting example implementation of how contextual data can be integrated with intermediate potential representations 926. Figure 9B Another non-limiting example implementation of such a fusion is illustrated. Similar to... Figure 9A ,exist Figure 9B In this process, input data 902 is provided to encoder 906 to generate intermediate latent representation 926, and is also provided, along with context data 904, to molecular descriptor calculator 908 to generate descriptor 930. Instead of... Figure 9A The example illustrates the direct fusion of descriptor 930 and intermediate latent representation 926, in Figure 9BIn this process, descriptor 930 undergoes embedding generation 914, which transforms descriptor 930 into a latent space. In some implementations, embedding generation 914 may be performed by a multilayer perceptron or other suitable model.

[0110] Normalization 918 is applied to the intermediate latent representation 926, and normalization 916 is applied to the output of the embedding generation 914. The resulting normalized vector undergoes a weighted combination 920 (e.g., αZ d + (1 – α ) Z m ,in Z d It is the normalized intermediate latent representation 926, and Z m (This is the normalized output of the embedded generation 914). The output of the weighted combination 920 undergoes an additional normalization 922 to generate the complete latent representation 928. Each of the normalizations 918, 916, and 922 can be similar to... Figure 9A The cascade 910 and normalization 912 described in the text.

[0111] Figure 9C Another non-limiting example implementation of how to achieve the fusion of context and intermediate latent representation 926 is illustrated. Figure 9C Similar to Figure 9B Normalization 918 is applied to the intermediate latent representation 926, and normalization 916 is applied to the output of the embedding generation 914. However, in Figure 9C Instead of applying weighted combination 920 to the normalized vector, the normalized vector simply undergoes cascading 924 to create the complete latent representation 928.

[0112] While the above discussion primarily describes the use of a single model to generate workflow predictions, latent representations, or generated workflow information, in some implementations, multiple models can be combined in a network of agents that collaborate together. Figure 10Non-limiting example implementations of proxy networks according to various aspects of this disclosure are illustrated. In proxy network 1000, two or more encoders / decoders 300 already trained on domain-specific datasets may be provided. For example, a first encoder / decoder 300 may be trained on paint-related data, a second encoder / decoder 300 on coating-related data, a third encoder / decoder 300 on polymer-related data, a fourth encoder / decoder 300 on adhesive-related data, and so on. As another example, one or more encoders / decoders 300 may be provided, which are trained or constructed in a manner that addresses specific workflow complexities (e.g., injecting specific types of contextual information into specific latent representations or encoders / decoders in the model).

[0113] As shown in the figure, the agent network 1000 includes a chemical supervision agent 1002 and two or more domain-specific agents (domain-specific agents 1004 to 1006). Each of the chemical supervision agent 1002 and the domain-specific agents 1004-1006 can be implemented as an encoder / decoder 300, or can be implemented using any other model architecture described herein. The chemical supervision agent 1002 represents the entire workflow and coordinates the execution of the domain-specific agents 1004-1006. The chemical supervision agent 1002 can be configured to integrate information from the domain-specific agents 1004-1006 to manage communication with the domain-specific agents 1004-1006 and ensure that the tasks delegated to the various domain-specific agents 1004-1006 are consistent with the goals of the overall workflow.

[0114] In some implementations, the chemical supervision agent 1002 may receive updates from domain-specific agents 1004-1006, adjust parameters, and / or reallocate tasks among the domain-specific agents 1004-1006 as needed. As a non-limiting example, in the context of complex workflows, the chemical supervision agent 1002 may use the results from given domain-specific agents 1004-1006 to help determine a better way to predict workflow outcomes, or to generate further refined workflows to better align with the expected goals / nature of the workflow outcomes.

[0115] Each of the domain-specific agents 1004-1006 can be dedicated to processing a specific domain or type of chemical workflow (e.g., paints / coatings, polymers, adhesives, etc.). Each domain-specific agent 1004-1006 can use its encoder / decoder architecture to process data specific to its domain, generate relevant latent representations, and decode the latent representations into actionable information that can be used by the chemical supervision agent 1002 (e.g., predicted results, another latent representation, generated workflow, etc.).

[0116] When a workflow is represented by agent network 1000, the workflow can be divided into subtasks, each of which is assigned to an appropriate domain-specific agent 1004-1006. For example, if a chemical process involves coating and adhesive, relevant data can be passed to a first domain-specific agent trained to process information related to coating, and a second domain-specific agent trained to process information related to adhesive. Each of these domain-specific agents can generate a latent representation that encodes features specific to its domain, as output and / or as an intermediate result. A chemical supervision agent 1002 can be configured to access these latent representations and can be configured to inject context or additional knowledge (as discussed above) to improve the performance of domain-specific agents 1004-1006. Each domain-specific agent 1004-1006 can continue to refine its understanding of the workflow through an iterative process of encoding and decoding. As new data enters (e.g., additional contextual data captured during the execution or simulation of a chemical workflow), domain-specific agents 1004-1006 can adapt by updating their underlying representation and transmitting relevant feedback to chemical supervisory agent 1002, which in turn coordinates updates throughout the agent network 1000.

[0117] Various actions that can be performed by the encoder / decoder 300, such as predicting experimental results and generating updates to the chemical workflow, can be extended by using a proxy network 1000 instead of a single encoder / decoder 300. For example, to predict experimental results, the proxy network 1000 can use domain-specific agents 1004-1006 to contribute additional latent representation features to be decoded for result prediction. By fusing latent representations from different domains, the proxy network 1000 can predict experimental results for complex workflows involving multiple types of materials. As another example, to generate updates to the chemical workflow, once a result is predicted, the proxy network 1000 can generate updated workflow parameters, formulation components, etc., to help optimize performance. For example, if the predicted experimental result based on polymer sensitivity to the current reaction conditions described in the workflow fails and is detected by a domain-specific agent trained on the polymer, the proxy network 1000 can generate workflow parameters to adjust temperature, pressure, or other parameters to make the process more likely to succeed.

[0118] By distributing these tasks among domain-specific agents 1004-1006, agent network 1000 can easily scale across more complex workflows without overburdening any single agent, and can increase the likelihood that agents will be able to train successfully on subdomains, while allowing the domain-specific knowledge of the agent network to be combined for the entire workflow. In some implementations, a single domain-specific agent may itself consist of a multi-agent network, enabling even domain-specific agents to be specialized in multiple domains. For example, within a domain-specific agent trained to process polymer information, there may be sub-specific agents trained relative to polymer chemical structures, another sub-specific agent trained relative to polymer synthesis processes / mechanisms, yet another sub-specific agent trained relative to the use of polymers in formulations, and so on. These sub-specific agents together will allow the polymer agent's overall analysis to be applied in a multi-scale manner to various aspects of the polymer's chemical workflow.

[0119] The embodiments of this disclosure have been tested in predicting a wide range of physical properties of chemicals with promising results. Implementations using the Uni-MOL converter model have been tested and compared with other model architectures and the TEST tool provided by the EPA, which uses a quantitative structure-activity relationship (QSAR) model and is the current gold standard for predicting boiling point, melting point, toxicology, and other properties. Using a trainable converter model will outperform a QSAR model because QSAR models are difficult to update and often overfit to the training data.

[0120] Uni-MOL is a transformer architecture that takes both atomic and pair representations as input. Pre-training is performed using two strategies: masking with random coordinates and reassignment with added noise. As described above, the internals of Uni-MOL are also manipulated to inject additional information into the latent space. The performance of Uni-MOL compared to TEST (i.e., QSAR), and the performance of a graph neural network (GNN) with context injection, are illustrated in Tables 2A-2B, 3A-3D, and 4A-4H. A GNN is a graph neural network machine learning model where data is represented in the form of a graph consisting of nodes and edges. Nodes represent entities (e.g., atoms in the case of chemical molecules, or states / stages / steps of a workflow in the case of processes), and edges represent connections between entities (e.g., bonds in the case of chemical molecules, or temporal / state dependencies in the case of processes). GNN is an artificial neural network model that takes a graph as its input and performs operations on the graph (e.g., graph convolution) to produce a latent representation of the graph. This latent representation can be used to compare different graphs or can be decoded to predict the data / structure of the graph.

[0121] The tests performed were expected to demonstrate that including context injection typically improves the model, whether using a GNN as the encoder or a transformer. However, in almost all cases, the results showed... 7 The converter is superior to other technologies. While exemplary embodiments have been exemplified and described, it should be understood that various changes may be made therein without departing from the spirit and scope of the invention.

[0122] Example The following numbered paragraphs describe non-limiting example implementations of the subject matter disclosed herein.

[0123] Example 1. A computer-implemented method for generating predictions related to a chemical workflow, the chemical workflow including a process performed on at least one chemical, the method comprising: encoding a representation of the at least one chemical by a computing system using a chemical encoder to create at least one potential chemical representation; creating a potential experimental representation based on the at least one potential chemical representation; and decoding the potential experimental representation by the computing system using an experimental decoder to predict one or more properties of the output of the chemical workflow to generate an updated representation of the at least one chemical, or to generate an updated representation of the process.

[0124] Example 2. The computer-implemented method according to Example 1, wherein predicting one or more properties of the output of the chemical workflow includes predicting one or more acoustic properties, chemical properties, electrical properties, magnetic properties, mechanical properties, optical properties, or thermal properties.

[0125] Example 3. A computer-implemented method according to any one of Examples 1-2, the method further comprising: encoding the representation of the process by the computing system using a process encoder to create a potential process representation; and wherein creating the potential experimental representation based on the at least one potential chemical representation comprises: encoding the at least one potential chemical representation and the potential process representation by the computing system using an experimental encoder to create the potential experimental representation.

[0126] Example 4. The computer-implemented method according to Example 3, the method further comprising: encoding context data related to the chemical workflow by the computing system to create a potential context representation.

[0127] Example 5. The computer-implemented method according to Example 4, the method further comprising: fusing the latent context representation with at least one of the latent chemical representation, the latent process representation, or the latent experimental representation by the computing system.

[0128] Example 6. A computer-implemented method according to any one of Examples 4-5, the method further comprising: using the latent context representation by the computing system when creating at least one of the latent chemical representation, the latent process representation, or the latent experimental representation.

[0129] Example 7. A computer-implemented method according to any one of Examples 3-6, wherein at least one of the chemical encoder, the process encoder, and the experimental encoder is a converter model.

[0130] Example 8. The computer-implemented method according to Example 7, wherein the converter model is a bidirectional encoder representation (BERT) model from the converter.

[0131] Example 9. The computer-implemented method according to Example 7, wherein the converter model is a Uni-MOL model.

[0132] Example 10. A computer-implemented method according to any one of Examples 7-9, wherein at least one of the chemical encoder, the process encoder, and the experimental encoder comprises a plurality of domain-specific encoders arranged as a proxy network.

[0133] Example 11. A computer-implemented method according to any one of Examples 3-10, wherein the experimental decoder is a Gaussian process.

[0134] Example 12. The computer-implemented method according to Example 11, wherein the Gaussian process includes a radial basis kernel.

[0135] Example 13. A computer-implemented method according to any one of Examples 3-12, wherein the representation of the process includes a structured representation of the process.

[0136] Example 14. The computer-implemented method according to Example 13, wherein the structured representation of the process is a structured text string having one or more key-value pairs, wherein the one or more key-value pairs include one or more key-value pairs representing one or more process steps and one or more key-value pairs representing one or more parameters.

[0137] Example 15. A computer-implemented method according to any one of Examples 13-14, wherein the one or more key-value pairs representing one or more process steps include representations of at least one of the following: step name; step type; step sequence; equipment identifier of the machine for performing the process step; or setting value of the settings of the machine for performing the process step.

[0138] Example 16. A computer-implemented method according to any one of Examples 13-15, wherein the one or more key-value pairs representing one or more parameters include representations of at least one of the following: parameter name; step sequence; equipment identifier of the machine associated with the parameter; attribute of the machine associated with the parameter; setting value of the machine associated with the parameter; or step condition value.

[0139] Example 17. A computer-implemented method according to any one of Examples 3-16, the method further comprising training at least one of the experimental decoder, the experimental encoder, the process encoder, or the chemical encoder by: the computing system comparing the one or more predicted properties with one or more true value properties measured from an execution instance of the chemical workflow to determine the value of a loss function; the computing system determining the gradient of the loss function; and the computing system using the gradient to update at least one of the experimental decoder, the experimental encoder, the process encoder, or the chemical encoder.

[0140] Example 18. A computer-implemented method according to any one of Examples 1-17, wherein the representation of the at least one chemical includes a structured representation of the at least one chemical.

[0141] Example 19. A computer-implemented method according to any one of Examples 1-18, wherein the structured representation of the at least one chemical is a structured text string having one or more key-value pairs, wherein the one or more key-value pairs represent one or more of the following: chemical name; chemical identifier; chemical linear symbol; concentration; property; unit; chemical property; or unstructured text associated with the at least one chemical.

[0142] Example 20. The computer-implemented method according to Example 19, wherein the chemical linear symbols include Simplified Molecular Input Linear Input System (SMILES) strings, Polymer SMILES (PSMILES) strings, International Chemical Identifier (InChI) strings, Wiswesser Linear Symbol (WLN) strings, Smiles Arbitrary Target Specification (SMARTS) strings, or SYBYL Linear Symbol (SLN) strings.

[0143] Example 21. A computer-implemented method according to any one of Examples 19-20, wherein the chemical identifier includes a CAS identifier.

[0144] Example 22. A non-transitory computer-readable medium storing computer-executable instructions that, in response to execution by one or more processors of a computing system, cause the computing system to perform the method according to any one of Examples 1-21.

[0145] Example 23. A computing system configured to perform the method according to any one of Examples 1-21.

Claims

1. A computer-implemented method for generating predictions related to a chemical workflow, the chemical workflow comprising a process performed on at least one chemical substance, the method comprising: The representation of the at least one chemical substance is encoded by a computing system using a chemical encoder to create at least one potential chemical representation; A potential experimental representation is created based on the at least one potential chemical representation; as well as The computing system uses an experimental decoder to decode the potential experimental representation to predict one or more properties of the output of the chemical workflow, in order to generate an updated representation of the at least one chemical, or to generate an updated representation of the process.

2. The computer-implemented method of claim 1, wherein predicting one or more properties of the output of the chemical workflow comprises: Predict one or more acoustic, chemical, electrical, magnetic, mechanical, optical, or thermal properties.

3. The computer-implemented method according to claim 1, further comprising: The computing system uses a process encoder to encode the representation of the process to create a potential process representation; and Creating the potential experimental representation based on the at least one potential chemical representation includes: The computing system uses an experimental encoder to encode the at least one potential chemical representation and the potential process representation to create the potential experimental representation.

4. The computer-implemented method according to claim 3, further comprising: The computing system encodes contextual data related to the chemical workflow to create a potential contextual representation.

5. The computer-implemented method according to claim 4, further comprising: The computing system fuses the latent context representation with at least one of the latent chemical representation, the latent process representation, or the latent experimental representation.

6. The computer-implemented method according to claim 4, further comprising: The latent context representation is used by the computing system when creating at least one of the latent chemical representation, the latent process representation, or the latent experimental representation.

7. The computer-implemented method according to claim 3, wherein at least one of the chemical encoder, the process encoder, and the experimental encoder is a converter model.

8. The computer-implemented method of claim 7, wherein the converter model is a bidirectional encoder representation (BERT) model from the converter.

9. The computer-implemented method according to claim 7, wherein the converter model is a Uni-MOL model.

10. The computer-implemented method of claim 7, wherein at least one of the chemical encoder, the process encoder, and the experimental encoder comprises a plurality of domain-specific encoders arranged as a proxy network.

11. The computer-implemented method of claim 3, wherein the experimental decoder is a Gaussian process.

12. The computer-implemented method of claim 11, wherein the Gaussian process comprises a radial basis kernel.

13. The computer-implemented method of claim 3, wherein the representation of the process includes a structured representation of the process.

14. The computer-implemented method of claim 13, wherein the structured representation of the process is a structured text string having one or more key-value pairs, wherein the one or more key-value pairs include one or more key-value pairs representing one or more process steps and one or more key-value pairs representing one or more parameters.

15. The computer-implemented method of claim 13, wherein the one or more key-value pairs representing one or more process steps include representations of at least one of the following: Step name; Step type; Step sequence; Equipment identifier for the machine used to perform the process steps; or The settings values ​​for the machine used to perform the process steps.

16. The computer-implemented method of claim 13, wherein the one or more key-value pairs representing one or more parameters include representations of at least one of the following: Parameter name; Step sequence; The equipment identifier of the machine associated with the parameter; The attributes of the machine associated with the parameter; The machine's setting value associated with the parameter; or Step condition value.

17. The computer-implemented method of claim 3, further comprising training at least one of the experimental decoder, the experimental encoder, the process encoder, or the chemical encoder in the following manner: The computational system compares the one or more predicted properties with one or more true value properties measured from execution instances of the chemical workflow to determine the value of the loss function; The gradient of the loss function is determined by the computational system; as well as The computational system uses the gradient to update at least one of the experimental decoder, the experimental encoder, the process encoder, or the chemical encoder.

18. The computer-implemented method of claim 1, wherein the representation of the at least one chemical comprises a structured representation of the at least one chemical.

19. The computer-implemented method of claim 1, wherein the structured representation of the at least one chemical is a structured text string having one or more key-value pairs, wherein the one or more key-value pairs represent one or more of the following: Chemical name; Chemical identifier; Chemical linear symbols; concentration; nature; unit; Chemical properties; or Unstructured text relating to the at least one of the chemicals.

20. The computer-implemented method of claim 19, wherein the chemical linear symbol comprises a Simplified Molecular Input Linear Input System (SMILES) string, a Polymer SMILES (PSMILES) string, an International Chemical Identifier (InChI) string, a Wiswesser Linear Symbol (WLN) string, a Smiles Arbitrary Target Specification (SMARTS) string, or a SYBYL Linear Symbol (SLN) string.

21. The computer-implemented method of claim 19, wherein the chemical identifier includes a CAS identifier.

22. A non-transitory computer-readable medium storing computer-executable instructions that, in response to execution by one or more processors of a computing system, cause the computing system to perform the method according to any one of claims 1 to 21.

23. A computing system configured to perform the method according to any one of claims 1 to 21.