Method and apparatus for characterizing chemicals, measuring physicochemical properties, and generating control data for synthesizing chemicals - Patents.com
Data-driven models for generating digital representations of polymers enhance synthesis efficiency and reduce waste by encoding sensor data for accurate polymer production.
Patent Information
- Application Number
- JP2025537966
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-27
- Filing Date
- 2023-12-27
- Publication Date
- 2026-01-21
Smart Images

Figure 2026502210000001_ABST
Abstract
Description
[Technical Field]
[0001] Technical Field The present disclosure relates to methods for generating digital representations of chemical substances, which may involve aspects of training neural networks to represent the chemical substances. The present disclosure also relates to computer program products, database search engines for identifying chemical substances, devices for generating measurement data related to chemical substances, and applications of the digital representations, including control data related to chemical substance synthesis specifications. The disclosed methods and aspects relate to characterizing chemical substances in terms of their physicochemical properties, control data, measurement data, composition, and / or identifiers. [Background technology]
[0002] background Chemical substances, such as polymers, come in multiple shapes, sizes, and compositions. Small molecules are often represented by their structure (chemical composition). Other basic chemical substances may be represented by their recipe and a detailed description of their manufacturing process. However, the properties of chemical substances, such as polymers, are often too complex to be represented by their recipe or their structure. In particular, polymers and / or oligomers involve a statistical distribution of repeating units and therefore require more elaborate representations. It would be desirable to provide a way to represent chemical substances in a more versatile manner. Summary of the Invention [Means for solving the problem]
[0003] overview In particular, in the chemical industry, chemicals such as new polymers are increasingly being prepared to customer requirements. This requires synthesizing the new polymer and then performing measurements of various properties. This is very expensive, and because of the low success rate of synthesizing a polymer that meets customer requirements, polymer synthesis often generates unnecessary waste. Additionally, performing the measurements is time-consuming and expensive. Therefore, there is a need to reduce the number of measurements required to fully analyze a polymer. Digital representations of chemicals can be a means to reduce the effort in determining or measuring the physical or chemical properties of chemicals, creating specifications for chemical synthesis, and / or identifying chemicals for specific purposes.
[0004] One object of the present invention is to provide a method and apparatus for generating digital representations of chemical substances. Further objects include improved uses and applications of digital representations in the context of chemical synthesis and substance characterization.
[0005] The above mentioned object is achieved by a method and a device according to the independent claims. In one aspect, a method is disclosed for generating control data indicative of synthesis specifications for chemicals, particularly polymers, the method comprising: receiving sensor data indicative of one or more measurable or measured physicochemical properties of a chemical; encoding the received sensor data using a data-driven model, the data-driven model being trained to map multimodal input data comprising sensor data and control data as modalities to encoded output data, the multimodal input data being a multimodal representation comprising sensor data and control data as modalities, and the encoded output data being a latent space representation of the multimodal input data; generating control data indicative of a synthetic specification of a chemical substance by decoding encoded multimodal input data based on or including received sensor data using a data-driven model, wherein the data-driven model is trained to map the encoded output data to multimodal output data including the sensor data and the control data as modalities, and the multimodal output data comprises a multimodal representation including the sensor data and the control data as modalities; Optionally, providing control data, e.g., for the synthesis of chemicals; Includes:
[0006] In another aspect, an apparatus for generating control data indicative of a synthesis specification for a chemical substance is disclosed, the apparatus comprising: an input interface configured to receive sensor data indicative of one or more measurable or measured physicochemical properties of the chemical; a model engine configured to encode received sensor data using a data-driven model, the data-driven model being trained to map multimodal input data comprising sensor data and control data as modalities to encoded output data, the multimodal input data being a multimodal representation comprising the sensor data and the control data as modalities, and the encoded output data being a latent space representation of the multimodal input data; a model engine configured to generate control data indicative of a synthetic specification of a chemical substance by decoding encoded multimodal input data based on received sensor data using a data-driven model, the data-driven model being trained to map the encoded output data to multimodal output data comprising the sensor data and the control data as modalities, the multimodal output data comprising a multimodal representation comprising the sensor data and the control data as modalities; and Optionally, an output interface for providing control data, e.g., for chemical synthesis; Equipped with.
[0007] In another aspect, a method for measuring a physicochemical property of a chemical substance is disclosed, the method comprising: receiving sensor data indicative of a first measurable physicochemical property of the chemical; encoding sensor data using a data-driven model, the data-driven model being trained to map multimodal input data comprising sensor data and / or measurement data as modalities to encoded output data, the multimodal input data being a multimodal representation comprising sensor data and / or measurement data as modalities, and the encoded output data being a latent space representation of the multimodal input data; generating measurement data indicative of a second measurable physicochemical property of the chemical substance by decoding encoded multimodal input data that is based on or includes the received sensor data using a data-driven model, wherein the data-driven model is trained to map the encoded output data to multimodal output data that includes the sensor data and / or the measurement data as modalities, and the multimodal output data includes a multimodal representation that includes the sensor data and the measurement data as modalities; Optionally, providing control data, e.g., for chemical characterization and / or synthesis. Includes:
[0008] In another aspect, an apparatus for measuring a physicochemical property of a chemical substance is disclosed, the method comprising the steps of: an input interface configured to receive sensor data indicative of a first measurable physicochemical property of the chemical; a model engine configured to encode sensor data using a data-driven model, the data-driven model being trained to map multimodal input data comprising sensor data and / or measurement data as modalities to encoded output data, the multimodal input data being a multimodal representation comprising sensor data and / or measurement data as modalities, and the encoded output data being a latent space representation of the multimodal input data; a model engine configured to generate measurement data indicative of a second measurable physicochemical property of the chemical substance by decoding encoded multimodal input data that includes or is based on the received sensor data using a data-driven model, the data-driven model being trained to map the encoded output data to multimodal output data that includes the sensor data and / or the measurement data as modalities, the multimodal output data including a multimodal representation that includes the sensor data and / or the measurement data as modalities; and Optionally, an output interface configured to provide measurement data, e.g., for chemical characterization and / or synthesis. Equipped with.
[0009] In another aspect, a method for generating control data indicative of synthesis specifications for chemicals, particularly polymers, is disclosed, the method comprising: providing a first synthetic specification of a reference chemical, e.g., as at least one modality of a first set; encoding the first synthesis specification into a digital representation of a reference chemical substance using a data-driven compression model; providing a database including a plurality of historical digital representations of historical chemical substances, or in other words, providing a plurality of historical digital representations of historical chemical substances stored in a database; determining a similarity score of the past digital representation with respect to a digital representation of a reference chemical substance; selecting at least one past expression based on the similarity score and generating, by decoding, a synthesis specification associated with the at least one selected past expression; Optionally, generating and / or providing control data indicative of the generated synthesis specification; Includes:
[0010] In another aspect, an apparatus for generating control data indicative of synthesis specifications for chemicals, particularly polymers, is disclosed, the apparatus comprising: an input interface configured to provide a first synthetic specification of a reference chemical, e.g., as at least one modality of a first set; a model engine configured to encode the first synthetic specification into a digital representation of the reference chemical using a data-driven compressed model; a database configured to provide a plurality of historical digital representations of historical chemical substances, or in other words to provide a plurality of historical digital representations of historical chemical substances stored in the database, and configured to generate by decoding a synthesis specification associated with at least one selected historical representation, and / or optionally configured to generate and / or provide control data indicative of the generated synthesis specification; configured to determine a similarity score of the past digital representation with respect to the digital representation of the reference chemical substance; a selection engine configured to select at least one past representation based on the similarity score; Equipped with.
[0011] In an embodiment, an apparatus for generating control data according to any one of the above aspects is implemented to perform steps according to a method for generating control data indicative of a synthesis specification of a chemical substance according to the above aspects, an embodiment of the method as described above or below with reference to the appended claims and / or drawings.
[0012] The method and apparatus may include encoding a first composite specification by using a data-driven model, the data-driven model being trained to map multimodal input data including the first composite specification as a modality to encoded output data, the multimodal input data being a multimodal representation including the first composite specification as a modality, and the encoded output data being a latent space representation of the multimodal input data; decoding a digital representation of a reference chemical by using the data-driven model, the data-driven model being trained to map the encoded output data to multimodal output data including the first composite specification as a modality, the multimodal output data including the multimodal representation of the first composite specification as a modality; optionally providing the composite specification associated with at least one selected past representation; and further optionally generating and / or providing control data indicative of the generated composite specification.
[0013] According to one aspect, a method is presented for characterizing a chemical substance in a predetermined multimodal representation, the predetermined multimodal representation having a predetermined plurality of modalities, the method comprising: receiving multimodal substance data including a first set of chemical substance modalities; encoding the multimodal data using a data-driven (especially compressed) model of chemical substances to generate encoded substance data; generating a predetermined multimodal representation comprising a second set of modalities of the chemical substance by decoding the encoded substance data using a data-driven (especially compressed) model, wherein preferably the multimodal substance data is indicative of physicochemical properties of the chemical substance, a composition of the chemical substance, and / or an identifier of the chemical substance; Including, a data-driven (particularly compressive) model is implemented to map input data to encoded output data, the input data being a multimodal representation of chemical substances and the encoded output data being a latent space representation of the input data; Preferably, the first set of modalities is different from the second set of modalities, the first and second sets of modalities being composed of a predetermined plurality of modalities, in particular at least one modality of the second set not included in the first set.
[0014] Use of measurement and / or control data generated according to the methods disclosed herein to synthesize chemicals, particularly polymers containing chemicals, having target properties.
[0015] Embodiment The data-driven model may include a compression model configured to map multimodal input data to a latent space representation for encoding and / or map the latent space representation to multimodal output data. The modalities present in the multimodal input data and the multimodal output data may be the same in training the model and / or may differ in use of the model. The model may be based on or include an autoencoder architecture. In particular, the model may be trained on multimodal input data that may include multimodal representations of sensor data and control data as modalities. The encoded output data may include a latent space representation of the multimodal input data. The mapping between the multimodal input data and the encoded output data may be trained according to or depend on multiple predetermined modalities. More particularly, the model may be trained to map the encoded output data to multimodal output data that includes sensor data and control data as modalities. The multimodal output data may include multimodal representations that include sensor data and control data as modalities. The mapping between the encoded output data and the multimodal output data may be trained on or depend on multiple predetermined modalities. When generating measurement data and / or control data using a model, input data of one or more input modalities may be provided as monomodal or multimodal input data. The trained model may map the input data to output data of one or more output modalities. The input and output modalities may be at least partially different. Thus, the trained model may provide data generation in a cross-modality mode.
[0016] A modality can be considered a dataset that represents a specific type of information related to a chemical substance, particularly a physicochemical property. In multimodal data, many different types of data are included in each multimodal dataset. The multimodal data can include multiple physicochemical properties as different modalities. The multimodal data can include one or more physicochemical properties, control data, sensor data related to measured properties, and / or measurement data related to measurable synthetically generated properties as different modalities. A modality can be represented by a dedicated data structure or a data type with a specific dimensionality. Scalars, multidimensional vectors, or tensors can be considered as data types. For example, melting temperature can be considered a physicochemical property that is a modality of a chemical substance. Another modality can be associated with a SMILES (Simplified Molecular Input Line Entry System) representation of a substance. Identifiers such as names (strings) or chemical formulas can serve as modal data. In an embodiment, each modality is represented by a specific data structure for computer-implemented processing of the modal data. The data structure can include a set of parameters that characterize the substance.
[0017] In one aspect, the method allows for generating output multimodal material data from "incomplete" input material data for a given multimodal representation. For example, the input data includes, among other modalities, a first modality but not a second modality. Using a data-driven compression model, the generated output material data may include the second modality. The method can be understood as providing measurement data from the input data.
[0018] In another embodiment, substance data including sensor data is received, the substance data including the sensor data as part of a first set of modalities, and the multimodal output data including control data as part of a second set of modalities. The first and second sets of modalities may be at least partially different. The first and second sets of modalities may include at least some of a predetermined plurality of modalities, for example, for which a data-driven model is trained. The first set may not include control data indicating chemical composition specifications.
[0019] In another embodiment, the multimodal input data relates to one or more measurable or measured physicochemical properties, one or more synthesis specifications, control data indicative of one or more synthesis specifications, chemical composition and / or chemical identifiers. The model may be trained on the multimodal input data.
[0020] In another embodiment, sensor data is received for one or more compositions of a plurality of chemicals, and the sensor data and the one or more compositions for each chemical are provided to a data-driven model to generate control data and / or associated measurement data for each chemical.
[0021] In another embodiment, measurement data indicative of one or more measurable or measured physicochemical properties of the control data and the chemical produced in accordance with the control data are generated by using a data-driven model. The one or more physicochemical properties of the measurement data may differ from the one or more physicochemical properties of the received sensor data. In this manner, not only the control data but also additional property data that the chemical produced in accordance with the control data may follow may be generated, thereby enabling reliable control data selection and chemical production.
[0022] In another embodiment, the data-driven model includes at least one multimodal variational autoencoder including at least one multimodal encoder and at least one multimodal decoder. The data-driven model may include multiple individual encoders, where a modality is assigned to each of the individual encoders of a predetermined plurality of modalities, or each individual encoder is assigned to a modality of the predetermined plurality of modalities, and the individual encoders are trained to map input data from the modalities to which the individual encoder is assigned to a common latent space representation. The data-driven model may include multiple individual decoders, where a modality is assigned to each of the individual encoders of a predetermined plurality of modalities, or each individual decoder is assigned to a modality of the predetermined plurality of modalities, and each individual decoder is trained to decode the latent space representation of the encoded input data into modality data of the generated multimodal input data, where the modality data is the modal data of the modality to which the individual decoder is assigned. The common latent space representation may include a learned probability distribution that may depend on the predetermined plurality of modalities. At least one multimodal encoder may be configured to map the individual latent space representations from the individual encoders to a common latent space, and at least one multimodal decoder may be configured to map the common latent space representation to the individual latent space representations of the individual decoders.
[0023] In another embodiment, the synthesis specification relates to the production of a chemical, particularly raw materials and operating conditions of a chemical plant for producing the chemical. The synthesis specification may relate to a production process type, auxiliary material type, and / or recipe, including, for example, reaction mixture ratios. The control data may relate to raw materials and operating conditions of a chemical plant for producing the chemical. The synthesis specification may include one or more instructions regarding how a chemical may be synthesized. In particular, the synthesis specification may include starting materials and respective production or production operating conditions for synthesis from the starting materials. The control signal may include instructions in a format that allows for automatic control of a respective chemical plant, industrial system, or work equipment for producing the chemical. The control signal may represent a machine-executable synthesis specification for the chemical.
[0024] In one embodiment, the chemical substance includes a polymer made from multiple monomers by polymerization. Generally, a macromolecule can include or be a polymer and / or oligomer. A polymer or oligomer can include one or more subgroups, all of which together form the polymer or oligomer. For example, a subgroup can refer to a portion of a polymer or oligomer, where the subgroups are linked together sequentially along a chain or network to form the polymer. Preferably, a subgroup of a polymer refers to a repeat unit that describes a portion of the polymer that, when repeated, creates a polymer chain. However, in some cases, a subgroup can also refer to a single portion of a polymer or oligomer that is not repeated. A subgroup can include repeating portions; for example, a subgroup of a polymer can include a repeating core that is also present in other subgroups, and further additional portions not present in other subgroups. A subgroup can include at least one polymerized monomer or oligomer fragment. A subgroup can include a polymerized monomer. In this context, polymerized monomers refer to the monomers after their polymerization and are sometimes referred to as "mer units" or "mers." In particular, polymerized monomer does not refer to the monomer, i.e., raw material, present in the reaction mixture prior to polymerization, but rather to repeat units derived from the monomer that have been modified during or after polymerization.
[0025] In another embodiment, the control data and / or measurement data are provided for synthesizing a chemical substance. The method may further comprise synthesizing the chemical substance according to the provided control data.
[0026] In an embodiment, the majority first set is a subset of the majority second set. In an aspect, the method allows for a process to be performed using a latent space representation of the input data. The latent space representation can be viewed as a vector in the latent space. The process in the latent space can involve searching for a vector, comparing vectors in the latent space associated with different chemicals, defining regions in the latent space, and calculating a similarity index or score. The similarity score can be based on a distance metric in the latent space, such as cosine distance, Tanimoto distance, a kernel-based distance index, or other common distance index.
[0027] In an embodiment, the data-driven compression model includes at least one trained neural network implemented to receive multimodal input data comprising a predetermined plurality of modalities, encode the input data into a latent space representation of the input data, and decode the encoded input data into multimodal output data comprising the predetermined plurality of modalities.
[0028] In embodiments, the neural network is trained based on training data that includes multimodal training data that includes a predetermined plurality of modalities, e.g., the neural network is trained using all available modalities of the predetermined plurality.
[0029] In an embodiment, the neural network includes a plurality of individual encoders, a modality is assigned to each individual encoder of a given plurality of modalities, and each individual encoder is implemented to transform input data from the modality to which the individual encoder is assigned into the same dimensionality of the latent space.
[0030] In an embodiment, the neural network includes a plurality of individual decoders, a modality being assigned to each of the individual encoders of a given plurality of modalities, and each individual decoder being implemented to decode a latent space representation of the encoded input data into modality data of the generated multi-modal substance data, the modality being the modal data of the modality to which the individual decoder is assigned.
[0031] In an embodiment, the characterizing comprises measuring a physicochemical property of the chemical substance, the substance data comprising sensor data, and the characterizing comprises: receiving sensor data indicative of a first measurable physicochemical property of the chemical substance, the sensor data being associated with at least one modality included in a first set; encoding the sensor data using a data-driven compression model of the chemicals to generate encoded sensor data; generating measurement data indicative of a second measurable physicochemical property of the chemical substance by decoding the sensor data encoded using the data-driven compression model, the measurement data being associated with at least one modality included in the second set; Includes:
[0032] In an embodiment, the chemical is a polymer. The method may be viewed as a method for measuring observable values of chemicals for which the observable values are not directly available, but which may be derived or reproduced through a data-driven model, which has inherently acquired knowledge of the desired output measurement data through pre-training with the "complete" substance training data.
[0033] In an embodiment, at least one of the group consisting of generated measurement data, recipe data indicative of chemical substances, and identification data indicative of chemical substances is output.
[0034] In an embodiment, the method comprises: For a plurality of sample chemicals, generating a latent space representation of multi-modal sample substance data associated with the sample chemicals using the data-driven compression model, the multi-modal sample substance data having a predetermined plurality of modalities; and / or storing the generated latent space representation of the multimodal sample substance data associated with the sample chemical in a database; Includes:
[0035] In an embodiment, the method comprises: receiving search input data indicative of predetermined physicochemical properties of chemical substances to be searched, the search input data being provided as a multimodal representation of the substances to be searched contained in the first set; encoding the received search input data to generate a latent space representation of the search input data; comparing the generated latent space representation of the search input data with the latent space representation of the sample chemicals to obtain a comparison result; Includes:
[0036] In an embodiment, the method comprises: Selecting at least one sample chemical depending on the comparison result. Includes:
[0037] In an embodiment, the comparing comprises: calculating a similarity score of the latent space representation of the search input data with respect to the latent space representation of the sample chemicals; and / or Determining similarity ranges in the latent space for a latent space representation of the search input data. Includes:
[0038] In an embodiment, at least one modality of the first set and / or the second set comprises a chemical synthesis specification and / or control data indicative of a chemical synthesis specification.
[0039] In an embodiment, the characterizing comprises generating control data indicative of a synthesis specification for a chemical substance, particularly a polymer, and the characterizing comprises: providing a first synthetic specification of a reference chemical as at least one modality of the first set; encoding the first synthesis specification into a digital representation of the reference chemical substance using a data-driven compression model; providing a database comprising a plurality of historical digital representations of historical chemical substances, or in other words, providing a plurality of historical digital representations of historical chemical substances stored in a database, preferably wherein the historical digital representations can be generated by a data-driven model based on multi-modal input data relating to one or more modalities of the plurality of chemical substances; determining a similarity score of the past digital representation with respect to the digital representation of a reference chemical substance; selecting at least one past expression based on the similarity score, and decoding and generating a synthesis specification associated with the at least one selected past expression; generating control data indicative of the generated synthesis specification; Includes:
[0040] In embodiments, the historical chemical is a sample chemical. The historical digital representation may include a latent space representation generated by a trained data-driven model, particularly a trained multimodal encoder. The trained model may be configured to map multimodal input data to encoded output data. The multimodal input data may include a multimodal representation, and the encoded output data may include a latent space representation of the multimodal input data. The historical digital representation may be generated by providing multimodal input data of the chemical to a trained data-driven model, particularly a trained multimodal encoder. The encoded output data or latent space representation thus generated may be stored in a database.
[0041] In an embodiment, the characterizing comprises generating control data indicative of a synthesis specification for a chemical substance, particularly a polymer, and the characterizing comprises: receiving sensor data indicative of a measurable physicochemical property of a chemical substance; encoding the sensor data using a data-driven compression model of the chemicals to generate encoded sensor data; generating control data indicative of a chemical composition specification by decoding the sensor data encoded using the data-driven compression model; Includes:
[0042] In an alternative embodiment, the first set of modalities is equal to the second set of modalities. Thus, the input may be the same as the output modality. However, the output data may refer to a different polymer or material than the input search data or material property data.
[0043] In an embodiment, using a data-driven model includes a process for generating a condensed digital representation of a chemical substance, particularly a polymer, the process comprising: receiving input data that is a multimodal representation of physicochemical properties of a chemical substance and that is indicative of measurable physicochemical properties of the chemical substance; encoding input data using a data-driven compression model of chemical substances to generate encoded substance data as a function of the received input data; generating chemical substance data indicative of measurable physicochemical properties of the chemical substance by decoding the substance data encoded using the data-driven compression model; Includes:
[0044] In an embodiment, this process is implemented as a method according to the second aspect disclosed below.
[0045] According to a second aspect, there is provided a method for generating a representation of a chemical substance, in particular a polymer, the method comprising: receiving input data, the input data being a multimodal representation of physicochemical properties of a chemical substance, the input data indicating measurable physicochemical properties of the chemical substance; encoding input data using a data-driven compression model of chemical substances to generate encoded substance data as a function of the received input data; generating chemical substance data indicative of measurable physicochemical properties of the chemical substance by decoding the substance data encoded using the data-driven compression model; Includes:
[0046] The method of the second aspect is specifically a process configured using a data-driven model.
[0047] The method of the second aspect allows for efficient processing of data referencing properties of chemical substances. The use of a data-driven model in an encoded or compressed form by using multimodal data may improve processing speed and reduce the amount of resources required in terms of energy, material, and / or computational power.
[0048] In an embodiment, a data-driven compression model maps multiple input modalities to a latent space vector. The multimodal data including the modalities requires a first amount of data, and the latent space representation of the multimodal data requires a second amount of data. Preferably, the second amount of data is less than the first amount. Thus, by using a data-driven model, the amount of data processed in the latent space can be reduced without substantially losing information about the material associated with the input multimodal data.
[0049] Multimodal data related to a chemical substance, for example, includes various data representing different aspects of the chemical substance. A first mode or modality may be contemplated as a physical observation, such as melting temperature, hardness, acidity, etc., and a second mode or modality may be contemplated as a structural aspect, such as the chirality of an enantiomer. Spectroscopic aspects may also be considered modalities.
[0050] In an embodiment, the data-driven compression model is implemented as a trained neural network, for example, the encoding step may include providing input data that is a multimodal representation as input to the trained neural network.
[0051] According to one aspect, a method for providing a data-driven compression model is disclosed, the method comprising: providing multimodal input data providing a multimodal representation of chemical substances in a multimodal initial space as input to a neural network to be trained, the neural network comprising a multimodal variational autoencoder including a multimodal encoder and a multimodal decoder, the multimodal encoder including a multimodal encoder layer having multimodal encoder weights that define how the multimodal encoder layer transforms the data, and the multimodal decoder including a multimodal decoder layer having multimodal decoder weights that define how the multimodal decoder layer transforms the data; adjusting the dimensionality of the multimodal input data using a multimodal encoder layer to obtain multimodal latent data in a latent space; decoding the multimodal latent data using a multimodal decoder layer to obtain multimodal reconstruction data in a multimodal initial space; calculating a loss function of a multimodal variational autoencoder for the current set of multimodal encoder and decoder weights based on the multimodal input data and the multimodal reconstruction data; Providing a neural network with a set of current multimodal encoder and decoder weights as a data-driven compression model; Includes:
[0052] According to one aspect, there is provided a further method for generating a representation of a chemical substance, particularly a polymer, in a latent space representation. The method may involve training a neural network: providing multimodal input data providing a multimodal representation of chemical substances in a multimodal initial space as input to a neural network to be trained, the neural network comprising a multimodal variational autoencoder including a multimodal encoder and a multimodal decoder, the multimodal encoder including a multimodal encoder layer having multimodal encoder weights that define how the multimodal encoder layer transforms the data, and the multimodal decoder including a multimodal decoder layer having multimodal decoder weights that define how the multimodal decoder layer transforms the data; adjusting the dimensionality of the multimodal input data using a multimodal encoder layer to obtain multimodal latent data in a latent space; decoding the multimodal latent data using a multimodal decoder layer to obtain multimodal reconstruction data in a multimodal initial space; Computing a loss function of a multimodal variational autoencoder for the current set of multimodal encoder and decoder weights based on the multimodal input data and the multimodal reconstruction data; Includes:
[0053] In an embodiment, the method includes outputting the trained neural network and / or configuration data indicative of a latent space representation of the chemical as a digital representation of the chemical.
[0054] The digital representation may further include an identifier for the chemical, multiple modalities associated with the chemical in terms of measurable data for the chemical, and / or specification data for synthesizing the chemical in terms of recipe data.
[0055] Using a multimodal variational autoencoder on multimodal input data allows us to process data from multiple modalities at once, combining modalities for the same chemical to represent the chemical in a single latent space representation. As a result, a neural network is trained to provide a reliable representation of the chemical in the latent space.
[0056] The generated digital representation requires less data for a complete description that includes all measurable properties of a chemical substance, thus reducing the amount of computational and memory resources required to process the data representing each chemical substance. Thus, the trained neural network can be considered a data-driven compressed model.
[0057] The combination of a multimodal decoder and an encoder may be considered an embodiment of a data-driven compression model suitable for generating digital representations of chemical substances. In particular, the autoencoder disclosed herein may be considered an embodiment of a data-driven compression model.
[0058] A chemical may be a form of matter that has a certain chemical composition and properties. Examples of chemicals include polymers and molecules. A polymer may be made up of linked monomers.
[0059] A neural network trained according to the method of the first aspect can provide a representation of chemical substances in a latent space representation. The latent space can be an abstract multidimensional space into which the neural network maps what it has learned from its training data. The latent space representation can be a mathematical representation of the training data with an adjusted (often reduced) dimensionality. "Adjusting" the dimensionality can mean increasing, decreasing, or maintaining the dimensionality, particularly to reach a desired (predetermined) dimensionality.
[0060] As used herein, a modality is information about a chemical from a particular source and / or sensor. Different modalities may be information about a chemical from different sources and / or sensors. Different modalities may be images of a chemical obtained by a camera, spectroscopic images of a chemical, recipes for a chemical, simulation data for a chemical, test data from tests on a chemical, etc. Information (data) from multiple modalities may be represented as multimodal input data.
[0061] The dimensionality of a multimodal input is defined in particular by the characteristics of the measurement: for example, in spectroscopy, the response of a chemical is measured at different wavelengths, and the wavelength range is fixed to the region where one expects to see the chemical response.
[0062] The multimodal initial space may correspond to a space in which multimodal input data is provided to the neural network. The multimodal initial space may be defined by its dimensionality. The multimodal initial space may be a space in which data from different sources and / or sensors are provided directly from the sources and / or sensors, or may be a space that has undergone some modification, such as conditioning and / or preprocessing.
[0063] A multimodal variational autoencoder (multimodal VAE) can be a variational autoencoder that combines and / or considers data from different modalities. The structure of a multimodal variational autoencoder can be similar to the structure of a standard variational autoencoder.
[0064] The multimodal encoder may be configured to receive as input multimodal input data and adapt its dimensionality to obtain multimodal latent data in the latent space, particularly where the dimensionality of the multimodal latent data is less than the dimensionality of the initial multimodal input data.
[0065] A multimodal decoder may be used to decode back data that was encoded by a multimodal encoder. In other words, a multimodal decoder receives multimodal latent data as input and outputs multimodal reconstructed data in a multimodal initial space (i.e., the same space as the initial multimodal input data).
[0066] During training, a loss function is calculated for the multimodal variational autoencoder. The loss function may indicate how good the multimodal variational autoencoder is. The loss function may correspond to the difference between the multimodal input data and the multimodal reconstruction data. The loss function may be determined by a mixture of expert techniques or a multiplication of expert techniques. In particular, the smaller the loss function, the better the multimodal variational autoencoder.
[0067] The current set of multimodal encoder and decoder weights refers to the set of multimodal encoder and decoder weights of the current run (iteration) of training the neural network.
[0068] A trained neural network can be used to receive training data related to known chemicals or other multimodal data as input and provide a latent space representation thereof as output.
[0069] Applications of the latent space representations of chemicals provided by trained neural networks are described in more detail below. Examples include providing a search engine for searching similarities between chemicals in latent space. Other examples include chemical redesign and / or the design of new chemicals.
[0070] Furthermore, the latent space representation of a chemical substance can be used, for example, in an automated manufacturing process to produce that chemical substance. Indeed, automated manufacturing machines (robots) often require very rich and compact information about the chemical substances being manufactured, which a latent space representation can provide.
[0071] According to a further embodiment, the method comprises: Updating the weights of the multimodal encoder and / or the weights of the multimodal decoder based on the loss function Further includes:
[0072] The results of the multimodal variational autoencoder, i.e., the multimodal latent data and the multimodal reconstructed data, can be changed by modifying the multimodal encoder weights and the multimodal decoder weights.
[0073] Thus, updating the weights of the multimodal encoder and / or the weights of the multimodal decoder based on the loss function allows the results (outputs) of the multimodal variational autoencoder to be modified in order to reduce the loss function and / or improve the neural network during its training, in particular to provide as accurate and / or convenient a latent space representation of the chemical-related input data as possible.
[0074] According to a further embodiment, the method comprises: repeating the steps of providing multimodal input data, adjusting the dimensionality, decoding the multimodal latent data, calculating a loss function, and / or updating weights to reduce the loss function. Further includes:
[0075] By repeating these steps, the loss function can be increasingly reduced until a satisfactory and reliable neural network is obtained. Training of the neural network can be terminated once the loss function falls below a predetermined threshold after a predetermined number of iterations (executions) of the steps of providing multimodal input data, adjusting the dimensionality, decoding the multimodal latent data, calculating the loss function, and / or updating the weights to reduce the loss function.
[0076] According to a further embodiment, the neural network further comprises a respective variational autoencoder assigned to each modality, the respective variational autoencoders comprising a respective encoder and a respective decoder, the respective encoder comprising a respective encoder layer with respective encoder weights defining how the respective encoder layer transforms the data, and the respective decoder comprising a respective decoder layer with respective decoder weights defining how the respective decoder layer transforms the data, and the method further comprises: Train each individual variational autoencoder by: inputting separate input data representing only the assigned modalities in each separate initial space to each variational autoencoder; adjusting the dimensionality of the respective input data using respective encoder layers to obtain respective latent data having a predetermined dimensionality; decoding the respective latent data using respective decoder layers to obtain respective reconstruction data in respective initial spaces; comparing the respective input data with the respective reconstruction data to obtain a comparison result; updating the individual encoder weights and / or the individual decoder weights based on the comparison results; Repeating the steps of inputting individual input data, adjusting the dimensionality of the individual input data, decoding the individual latent data, comparing the data, and updating the individual encoder weights and / or the individual decoder weights to reduce the comparison results. By training each individual variational autoencoder The method further comprises: Using individual latent data from multiple individual variational autoencoders as multimodal input data for a multimodal variational autoencoder Further includes:
[0077] Individual variational autoencoders can be used to preprocess data from different modalities before inputting them into a multimodal variational autoencoder. For example, individual variational autoencoders perform separate preprocessing or conditioning of the individual data representing the individual modalities. The individual variational autoencoders may convert data from all modalities into the same format and / or the same number of dimensions (corresponding to the "predetermined number of dimensions") in the individual latent data. This may facilitate processing by the multimodal variational autoencoder by receiving the individual latent data and / or information about the individual variational autoencoders (e.g., hyperparameters determined during training of the individual variational autoencoders) as multimodal input data.
[0078] Each individual variational autoencoder (individual VAE) may be a variational autoencoder that considers only data from a single modality (its assigned modality). Thus, data from each modality is processed separately by a single individual variational autoencoder. The structure of the individual variational autoencoders may be similar to that of a standard variational autoencoder.
[0079] The individual encoders may be configured to receive individual input data of the assigned modality as input and adjust (increase, decrease, or maintain) their dimensionality to obtain individual latent data in individual latent spaces.
[0080] The individual decoders may be used to decode back the data encoded by the individual encoders, in other words, they receive individual potential data as input and output individual reconstructed data in individual initial spaces (i.e., the same space as the individual initial input data).
[0081] During training, an individual loss function is calculated for each individual variational autoencoder. This individual loss function can indicate the degree of goodness of the individual variational autoencoder. The individual loss function corresponds to the difference between the individual input data and the individual reconstruction data and can be expressed as a comparison result thereof. In particular, the smaller the comparison result, the better the individual variational autoencoder.
[0082] Thus, by updating the individual encoder weights and / or the individual decoder weights based on the loss function, it becomes possible to modify the results (outputs) of the individual variational autoencoders in order to reduce the loss function and / or improve the neural network during training.
[0083] According to a further embodiment, the method comprises: Repeating training each individual variational autoencoder for different predetermined dimensionalities of the individual latent data; Selecting the predetermined number of dimensions for the final trained neural network that minimizes the loss function; Further includes:
[0084] In particular, the loss function of the multimodal variational autoencoder is calculated for different sets of multimodal input data corresponding to the respective latent data obtained from the respective variational autoencoders for different predetermined dimensionalities. By varying the predetermined dimensionality, the loss function can be reduced. The final trained neural network is a trained neural network. Preferably, the final trained neural network has hyperparameters and / or weights that minimize the loss function. The predetermined dimensionality can be considered a hyperparameter.
[0085] According to a further embodiment, the loss function of the multimodal variational coder is determined through a mixture of experts or a multiplication of expert techniques.
[0086] Blending expert techniques involves breaking down the predictive modeling task into subtasks, training an expert model for each, developing a gating model that learns which experts to trust based on predicted inputs, and combining the predictions. Multiplying expert techniques model a probability distribution by combining outputs from several simpler distributions.
[0087] According to a further embodiment, one of the modalities for describing a chemical substance provides a spectroscopic representation, a rheological representation, a thermal representation, a chemical representation, a structural representation, a solubility representation, a dispersibility representation, a viscosity representation, and / or a surface tension representation of the chemical substance.
[0088] In particular, one modality may correspond to any analytical and / or characterization method of chemical substances, including (i) all types of spectroscopic data, such as Fourier transform infrared spectroscopy (FTIR), Raman spectroscopy, Raman microscopy, ultraviolet spectroscopy, nuclear magnetic resonance spectroscopy (NMR), mass spectrometry (MS), gas chromatography-mass spectrometry (GCMS), etc., (ii) thermal analysis data, such as dynamic mechanical analysis (DMTA), differential scanning calorimetry (DSC), thermogravimetric analysis (TGA), dynamic mechanical analysis (DMA), etc., (iii) structural property data, such as X-ray diffraction, and (iv) data from proprietary analytical techniques that measure physical properties of polymers, such as viscosities, surface tension, etc.
[0089] For example, FTIR always contains errors arising from domain-specific sources, so neural networks are trained to deal with such errors and be able to deal with chemical data in a more general way.
[0090] According to a further embodiment, at least two modalities provided in the individual or multimodal representation of the multimodal input data comprise data in different formats.
[0091] Data in different formats may include images, tables, text, etc. Data in different formats may also include data with different dimensions and / or data with different structures or representations.
[0092] According to a further embodiment, at least a portion of the multimodal input data and / or the individual input data is sequential data.
[0093] To deal with sequential data, the neural network may be a convolutional neural network.
[0094] According to a further embodiment, the individual input data and / or the multimodal input data comprises augmented data.
[0095] Augmented data may designate artificially generated data, for example by generating a slightly modified copy of the initial data, thereby allowing for an increase in the amount of training data (individual input data and / or multimodal input data) and thereby improving the training of the neural network.
[0096] According to a further embodiment, adjusting the dimensionality of the individual input data using the individual encoder layers comprises increasing the dimensionality of the individual input data for at least one modality and decreasing the dimensionality of the individual output data for at least another modality.
[0097] Depending on the assigned modality, the individual encoder layer increases or decreases the dimensionality of the individual input data to obtain individual data in a predetermined dimensionality. For example, data representing some modalities may be scalar (optionally with error bars), and the dimensionality may be increased to achieve the predetermined dimensionality.
[0098] According to a further aspect, a method for measuring a physicochemical property of a chemical substance, in particular a desired, pre-set and / or predetermined physicochemical property of a chemical substance, comprises the steps of: receiving sensor data indicative of a first measurable physicochemical property of the chemical; encoding the sensor data using a data-driven compression model of the chemicals to generate encoded sensor data; Decoding the encoded sensor data using the data-driven compression model to generate measurement data indicative of a second measurable physicochemical property of the chemical substance.
[0099] Preferably, the encoding and decoding steps are performed according to the respective steps of the first aspect of the present disclosure or an embodiment thereof.
[0100] In an embodiment, the data-driven compression model is implemented to map input data to encoded output data according to the method of the first aspect for generating a representation of a chemical substance, wherein the input data is a multimodal representation of physicochemical properties of the chemical substance and the encoded output data is a latent space representation of the input data.
[0101] Determining the properties of chemical substances in a digital latent space representation can avoid measurements and experiments that need to be performed with samples of each substance. Furthermore, generated data indicating physicochemical properties using the presented data-driven methods and compression models can be used as training data for artificial intelligence purposes. Therefore, one aspect of the present disclosure is also the use of generated measurement data as training data. While sensor data can be obtained by hardware measurements, generated measurement data can be considered synthetic measurement data. Measurement data can include sensor data, e.g., properties measured by hardware sensors, and generated measurement data, e.g., properties synthetically generated by a data-driven model.
[0102] In an embodiment, the provided or received sensor data relates to a first type of physical measurement and the generated measurement data relates to a second type of physical measurement. For example, it may be contemplated to provide X-ray diffraction data as the received sensor data and near-infrared spectra as the generated measurement data. The digital representation allows for the partial omission of energy- and resource-intensive technical processes for measuring a sample of material.
[0103] According to a still further aspect, a measuring device for measuring physicochemical properties of chemical substances, in particular desired, pre-set and / or predetermined physicochemical properties of chemical substances, comprising: an interface device implemented to receive sensor data indicative of a first measurable physicochemical property of the chemical; an encoder device implemented to encode the received sensor data and to generate and output encoded sensor data; a decoder device implemented to generate measurement data indicative of a second measurable physicochemical property of the chemical substance and to decode the encoded sensor data; Equipped with.
[0104] The encoder device is preferably implemented to map input data to encoded output data according to the method of the first aspect for generating a representation of a chemical substance, the input data being a multimodal representation of physicochemical properties of the chemical substance in a multimodal initial space and the encoded output data being a latent space representation of the input data.
[0105] The decoder device is preferably implemented to map input data to decoded output data according to the method of the first aspect for generating a representation of a chemical substance, the input data being a multimodal latent space representation of physicochemical properties of the chemical substance and the decoded output data being multimodal reconstruction data in the multimodal initial space.
[0106] According to one aspect, there is provided a computer program product comprising instructions which, when executed by a computer, cause the computer to perform the method according to the first aspect or a method according to an embodiment thereof.
[0107] Computer program means, such as a computer program product, may be embodied on a memory card, a USB stick, a CD-ROM, a DVD or as files that can be downloaded from a server in a network, for example such files may be provided by transferring the files that make up the computer program product over a wireless communication network.
[0108] The embodiments and features described with reference to the training method of the first aspect or embodiments thereof apply, as appropriate, to the computer program product of the second aspect.
[0109] According to one aspect, there is provided a database search device or engine that is specifically implemented to identify chemical substances having predetermined physicochemical properties. a storage unit for storing a database and a trained neural network, the database comprising representations of a plurality of chemical substances obtained using the trained neural network, in particular in a latent space, and the trained neural network comprising an encoder and a decoder; an input unit for receiving search input data indicative of chemical substances to be searched; 1. A processor, comprising: adjusting the dimensionality of the search input data using an encoder to obtain encoded search data in the encoded space; comparing the encoded search data to a representation of a plurality of chemical substances in a database; selecting at least one chemical substance from the plurality of chemical substances represented in the database based on a comparison between the encoded search data and the representations of the plurality of chemical substances; a processor configured to an output unit for outputting the identifiers of the selected chemical substances; Equipped with.
[0110] The trained neural network may comprise a variational autoencoder that includes a multimodal encoder and a multimodal decoder.
[0111] The database search device can be part of a computer, particularly a personal computer or an industrial computer. A trained neural network can be used to provide latent space representations of chemical substances. In particular, the database includes latent space representations of multiple chemical substances obtained using the trained neural network. For example, the database can be updated periodically and / or continuously as new data about chemical substances is acquired.
[0112] The storage unit storing the database and the trained neural network may be any type of temporary or permanent storage device (memory). The processor may be a central processing unit (CPU) or the like configured to access the database and / or execute the neural network stored therein. The input unit may include a user interface for receiving search input data from a user, or may be a unit that can access search input data stored in the storage unit or the like.
[0113] The search input data is data that has not yet been input to a neural network and / or for which no latent representations have been stored. The search input data can be in the same format as the multimodal input data described above. The search input data can also be incomplete data representing chemical substances, for example, including only partial representations of the chemical substance's modalities.
[0114] A neural network can be used to transform the search input data into the same latent space representation as the data in the database, particularly using a neural network multimodal encoder that can combine multiple modalities of the search input data.
[0115] Preferably, the potential search data is of the same representation (and dimensionality) as the data in the database. Comparison of the potential search data with the data in the database (i.e., representations of the multiple chemical substances) can be performed by directly comparing the potential search data with the data in the database. For example, the numerical values of the potential search data assigned to each of its dimensionality can be directly compared with the numerical values of each data in the database assigned to the same dimensionality. This comparison makes it possible to determine the similarity between the potential search data and each representation of the multiple chemical substances in the latent space. A comparison score proportional to the similarity can be assigned to each representation of the multiple chemical substances in the latent space.
[0116] The at least one selected chemical may be the chemical whose latent space representation in the database is closest to the latent search data (e.g., has the highest comparison score). The plurality of selected chemicals are the N chemicals represented in the database that are closest to the latent search data. This relies on the fact that similar chemicals have similar latent space representations.
[0117] The output unit may be a user interface such as a display, a touch screen, etc. The identifier of the selected chemical may include the name, chemical composition, reference number, or other identifying information of the selected chemical. The output unit may output (display) the identifier to a user, store it in a storage unit, etc.
[0118] A database search device can be used to identify chemical substances based on their multimodal representations (search input data) by performing comparisons in latent space.
[0119] The database search device can be used to perform redesign, i.e., to find a representation of a chemical without knowing its recipe, in which case, once an unknown chemical search is found, its recipe can be derived from the recipe of the selected chemical.
[0120] Furthermore, the latent space representation provided by neural networks can mitigate few-shot learning problems (the problem of making predictions based on a limited number of samples) by reducing the dimensionality of the input data and by providing a rich feature space trained on very large datasets.
[0121] According to one embodiment, the neural network is trained according to the method of the first aspect or any embodiment thereof.
[0122] The database searching device may be further configured to perform training of the neural network according to the method of the first aspect or any embodiment thereof.
[0123] A further aspect of the present disclosure includes a method for generating control data indicative of synthesis specifications for chemicals, particularly polymers, the method comprising: providing a first synthetic specification of a reference chemical; encoding the first synthesis specification into a digital representation of the reference chemical substance using a data-driven (particularly compressed) model; providing a database comprising a plurality of historical digital representations of historical chemical substances, or in other words, providing a plurality of historical digital representations of historical chemical substances stored in a database, preferably wherein the historical digital representations can be generated by a data-driven model based on multi-modal input data relating to one or more modalities of the plurality of chemical substances; determining a similarity score of the past digital representation with respect to the digital representation of a reference chemical substance; selecting at least one past representation based on the similarity score and decoding a synthesis specification associated with the at least one selected past representation; generating control data indicative of the generated synthesis specification; Includes:
[0124] The historical digital representation of the historical chemicals may be generated by a trained data-driven model (particularly a compressed model) configured to map the multimodal input data to a latent space representation. The trained data-driven model may include an encoder configured to reduce the dimensionality of the multimodal input data.
[0125] The chemical synthesis specification preferably includes all process and recipe data required to manufacture each chemical. The synthesis specification data may include control data for operating a chemical plant in machine-readable form.
[0126] As a result of the above-described aspects, a synthesis specification is obtained that can result in a chemical similar to the reference chemical when the control data is deployed in a chemical plant, i.e., a chemical manufacturing system, according to the control data. The method also provides an alternative synthesis specification for the reference chemical.
[0127] In an embodiment, a data-driven (especially compressed) model is implemented to map input data to encoded output data, e.g., according to the methods disclosed herein for generating representations of chemical substances, where the input data is a multimodal representation of the chemical substance and the encoded output data is a latent space representation of the input data, e.g., the multimodal input data.
[0128] The database may be configured in a database search device or engine according to the aforementioned aspects.
[0129] In an embodiment, the search input data indicated a component of a chemical substance to be replaced by an alternative component, and the selected chemical substance included the alternative component in place of the component to be replaced.
[0130] In an embodiment, the search input data indicates qualitative characteristics of a chemical substance, which may refer to a classification based on predetermined regulations, such as the German Hazardous Substances Regulation (Gefahrstoffverordnung-GefStoffV).
[0131] In particular, the identifiers indicative of the components and / or qualitative characteristics to be replaced may be modalities in terms of the encoder and decoder of the neural network.
[0132] Another aspect relates to a method for generating control data indicative of synthesis specifications for chemicals, particularly polymers, the method comprising: receiving sensor data indicative of a measurable physicochemical property of a chemical substance; encoding the sensor data using a data-driven (especially compressed) model of the chemicals to generate encoded sensor data; generating control data indicative of a chemical composition specification by decoding the sensor data encoded using a data-driven (particularly compressed) model; Including, A data-driven (particularly compressed) model is implemented to map input data to encoded output data according to the method of the first aspect for generating a representation of a chemical substance, the input data being a multimodal representation of physicochemical properties of the chemical substance and the encoded output data being, in particular, a latent space representation of the input data.
[0133] The presented aspects allow for the generation of control data for synthesizing chemicals according to desired physicochemical properties without conducting experiments or test runs of real-world chemical plants, thus facilitating the creation and operation of plants for the production of such chemicals through digital representation in terms of latent space representation.
[0134] In an embodiment, a method for generating control data comprises: applying constraints indicating process or material requirements, in particular biodegradability requirements of chemicals and / or ingredients for producing chemicals, biomass-based requirements, exclusion of toxic ingredients, etc.; A step to verify whether the generated synthesis specification satisfies the constraints, especially before generating the control data. It may include at least one of:
[0135] It will be appreciated that embodiments of the method aspect are at least configured to, in response to receiving sensor data and / or receiving a first synthetic specification, e.g., a reference chemical, generate, provide, and / or output control data indicative of the generated synthetic specification of a past chemical used in the training process of the data-driven compression model.
[0136] The embodiments and features described with reference to the training method aspect of the first aspect or embodiments thereof apply mutatis mutandis to the database device of the method for determining physicochemical properties and generating control data and / or measurement data of the third aspect and other aspects.
[0137] According to some aspects, the processor may further decode the latent representation of the at least one selected chemical substance using a multimodal decoder to obtain reconstructed representation data in the multimodal initial space, and the output unit may be configured to output the reconstructed representation data, thereby providing a representation of the selected chemical substance in the initial space that is, for example, understandable and analyzable by a user.
[0138] The disclosed embodiments, among other things, allow for the substitution of known polymers with molecules having similar performance, including the creation of appropriate synthetic specifications for the substitution polymer / molecule.
[0139] All disclosed methods are preferably computer-implemented. In all embodiments, the data-driven model may be implemented in a computerized manner, for example, in terms of computer-readable functions or routines that cause a processing device to perform calculations to implement the model. The computer-readable form may include, for example, source code, pre-compiled code, and / or machine code. A data-driven model may also be considered a computerized device that receives input data and outputs output data in a desired format.
[0140] Furthermore, possible implementations or alternative solutions of the present invention also encompass combinations of features described above or below with respect to the present embodiments (not explicitly mentioned herein).Those skilled in the art can also add individual or isolated aspects and features to the most basic form of the present invention.
[0141] Further embodiments, features, and advantages of the present specification will become apparent from the following description and dependent claims, and from consideration in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0142] [Figure 1] FIG. 1 is a diagram illustrating a first example of a neural network. [Figure 2] FIG. 2 illustrates a first embodiment of a method for training the neural network of FIG. 1. [Figure 3] FIG. 2 illustrates a second embodiment of a method for training the neural network of FIG. [Figure 4] FIG. 10 is a diagram illustrating a second example of a neural network. [Figure 5] FIG. 5 illustrates a first embodiment of a method for training the neural network of FIG. 4. [Figure 6] 6A-6C illustrate different representations of the training method of FIG. 5. [Figure 7] FIG. 1 illustrates a database search device. [Figure 8]FIG. 8 illustrates a method of operation of the database search device of FIG. 7. [Figure 9] FIG. 8 illustrates a user interface for using the database search device of FIG. 7. [Figure 10] FIG. 1 illustrates an embodiment of a chemical production system. [Figure 11] FIG. 1 illustrates a user interface for using a measurement device / service. [Figure 12a] FIG. 1 illustrates an example of a model for generating control data and / or measurement data. [Figure 12b] FIG. 1 illustrates an example of a model for generating control data and / or measurement data. [Figure 13a] FIG. 10 illustrates another example of a model for generating control data and / or measurement data. [Figure 13b] FIG. 10 illustrates another example of a model for generating control data and / or measurement data. DETAILED DESCRIPTION OF THE INVENTION
[0143] In the drawings, like reference numbers designate similar or functionally equivalent elements unless otherwise specified.
[0144] The encoders, decoders, and autoencoders presented in this specification may be implemented in accordance with M. Wu, N. Goodman: “Multimodal Generative Models for Scalable Weakly-Supervised Learning”, arXiv:1802.05335, and references therein, which are incorporated herein by reference.
[0145] Figure 1 shows an example of a neural network 1 comprising a multimodal variational autoencoder 3 having a multimodal encoder 4 and a multimodal decoder 5. The neural network 1 is trained according to the method of Figure 2, and therefore Figures 1 and 2 will be discussed together below.
[0146] With respect to the following embodiments, it will be understood that the presented neural network embodies a framework for the digital representation of chemical substances. The deployed artificial neural network can be characterized in terms of parameters such as the number and properties of the implemented neurons, weights, nodes, connections, and other configuration parameters. The expression "latent space representation" in the context of this application refers to a digital representation of a chemical substance, such as a polymer, as follows: "Modality" describing a chemical substance refers to the physicochemical properties of a chemical substance that are observable by measurement and can be represented in a digital or computer-processable manner, for example, the spectroscopic, rheological, thermal, chemical, structural, solubility, dispersibility, viscosity, and / or surface tension representation of a chemical substance.
[0147] The latent space representation of a chemical substance is compressed with respect to the data volume required by the multimodal data. For example, a characterization of a chemical substance in terms of a raw parameter set describing multiple physicochemical properties and a name (e.g., a CAS (Chemical Abstracts Service) number) can be considered a multimodal representation requiring multiple data structures. After generating the latent space representation, a latent space data structure that represents the properties of the same substance becomes available, and the latent space representation requires fewer and / or smaller data structures. For example, the dimensionality of the latent space representation is less than the dimensionality of the initial multimodal representation. Because the encoder and decoder are trained on the multimodal substance data, potential information loss due to encoding is reduced or becomes negligible.
[0148] The multimodal encoder 4 and multimodal decoder 5 form an interface to the latent space representation 6 and are therefore computer-implemented embodiments of the data-driven compression model.
[0149] To train neural network 1, neural network 1 receives multimodal input data 2 as input (step S1 in FIG. 2). In the example of FIG. 1, multimodal input data 2 includes data representing seven modalities 2a-2g of the same chemical, here a polymer. Data 2 can be understood as a predetermined multimodal representation of the polymer. Reference numeral 2 represents the predetermined multimodal representation including a plurality of seven modalities 2a-2g. Modality 2a includes spectroscopic data from spectroscopic measurements, modality 2b includes rheological data, modality 2c includes X-ray diffraction data, modality 2d includes solubility data, modality 2e includes dispersed clay data (indicating the interaction of the chemical with the layered structure of the clay), modality 2f includes surface tension data, and modality 2g includes viscosity data of the polymer. Data from all dimensions was obtained by performing corresponding measurements on the polymer using sensors. Neural network 1 receives multimodal input data 2 referencing multiple polymers.
[0150] Multimodal input data 2 is provided in an initial space. In the initial space, data from each modality 2a-2g has its own dimensionality, which here corresponds to the dimensionality of the data sensed by the sensors. Thus, modalities 2a-2c have a higher dimensionality (5-100) than modalities 2d-2g, which are scalar (have only one dimension). Alternatively, in the initial space, data from each modality 2a-2g have the same dimensionality (e.g., 50).
[0151] In an alternative embodiment, the encoder 4 is replaced by separate encoders, each associated with one of the input modalities 2a to 2g, and the decoder 5 is replaced by separate decoders, each associated with one of the output modalities 7a to 7g.
[0152] In step S2 of Figure 2, the dimensionality of the multimodal input data 2 is modified using a multimodal encoder 4. The multimodal encoder 4 includes multiple multimodal encoder layers with multimodal encoder weights that define the mathematical operations by which the multimodal encoder 4 transforms the multimodal input data 2. The multimodal encoder weights are some of the parameters that are modified and optimized during training of the neural network 1, as explained further below.
[0153] In step S2, the multimodal encoder 4 reduces the dimensionality of the multimodal input data 2 to obtain multimodal latent data in a latent space 6. The latent space representation of the multimodal input data 2, i.e., the multimodal latent data, comprises 16 dimensions in this example.
[0154] In step S2, the multimodal encoder 4 combines the data from all modalities 2a-2g to form a single dataset that describes the polymer in latent space 6.
[0155] In step S3, a multimodal decoder 5 is used to decode the multimodal latent data to obtain multimodal reconstruction data 7, 7a-7g in the initial space. This involves modifying the dimensionality of the multimodal latent data to return it to the dimensionality of the initial multimodal input data 2. The multimodal decoder 5 includes multiple multimodal decoder layers, each having multimodal decoder weights that define the mathematical operation of the multimodal decoder 5 on the multimodal latent data. The multimodal decoder weights are some of the parameters that are modified and optimized during training of the neural network 1, as described further below.
[0156] In step S4 of the training method of Figure 2, a loss function for the multimodal variational autoencoder 3 is calculated. In its simplest form, the loss function indicates the degree of similarity between the multimodal input data 2 and the multimodal reconstruction data 7. Alternative methods for calculating the loss function for the multimodal variational autoencoder 3 include a mixture of experts, a mixture of Gaussians, and / or a multiplication of expert techniques.
[0157] The calculated loss function indicates how well Neural Network 1 is performing during the current run (iteration). The smaller the loss function, the better Neural Network 1 is.
[0158] Figure 3 shows a further embodiment of a method for training the neural network 1 of Figure 1. Method steps S1 to S4 of Figure 3 are identical to those of Figure 2. Depending on the calculated loss function, the neural network may update all or some of the multimodal encoder weights and / or all or some of the multimodal decoder weights in optional step S5 of Figure 3. The multimodal encoder weights and / or the multimodal decoder weights are updated by backpropagation.
[0159] As shown in Figure 3, in step S6, all method steps S1-S5 may be repeated to reduce the loss function, thereby improving neural network 1. Steps S1-S5 may be repeated a predetermined number of runs, or until the calculated loss function is less than a predetermined loss function threshold. When training stops, the multimodal encoder weights and multimodal decoder weights of the run providing the lowest loss function are retained as the weights leading to the best neural network 1. The trained neural network 1 corresponds to this best run and has its multimodal encoder weights and decoder weights.
[0160] Figure 4 shows a second example of a neural network 1. Figure 5 shows an embodiment of a method for training the neural network 1 of Figure 4. Many elements of the neural network 1 of Figure 4 and the method of Figure 5 are identical to the neural network 1 and training method of Figures 1-3, and the descriptions of Figures 4 and 5 apply equally.
[0161] The difference with the neural network 1 of FIG. 1 is that the neural network of FIG. 4 includes seven individual variational autoencoders 10, each including an individual encoder 8 and an individual decoder 9. Specifically, individual encoders 8a-8g and individual decoders 9a-9g correspond to modalities 2a-2g, respectively. Modalities 2a-2g correspond to the previously described modalities 2a-2g, but their characterization data form individual input data 12 instead of multimodal input data 2. The difference between the individual input data 12 and the multimodal input data 2 is that the individual input data 12 is input to the individual variational autoencoder 10, while the multimodal input data 2 is input to the multimodal variational autoencoder 3. Furthermore, the individual input data 12 may include data of different dimensions for different modalities, while the multimodal input data 2 may include data of the same dimensions for all modalities 2a-2g.
[0162] The individual variational autoencoders 10 are for conditioning the data 12 before inputting it into the multimodal variational autoencoder 3. The individual encoders 8a-8g transform the input data 12 from each modality 2a-2g into the same predetermined dimensionality, which may be the dimensionality of the latent space 6 (e.g., dimensionality 16).
[0163] In one embodiment, autoencoder 3 is an optional element, and individual encoders 8a-8g each convert input data 12 from each modality 2a-2g into the same predetermined dimensionality of latent space 6. Similarly, individual decoders 9a-9g map latent space vectors into respective modalities 17a-17g having particular individual dimensionality.
[0164] Due to the pre-training, various dimensions and modalities are intertwined, so that the individual decoders / encoders 8, 9 interact with latent space vectors with a given dimension. Missing input modalities can be repaired by an autoencoder structure.
[0165] In particular, in step S6 of Figure 5, individual input data 12 representing each single modality 2a-2g is input into a corresponding individual encoder 8. This means that individual input data 12 representing modality 2a is input into corresponding individual encoder 8a, individual input data 12 representing modality 2b is input into corresponding individual encoder 8b, and so on.
[0166] 5, each individual encoder 8 modifies the dimensionality of the received individual input data 12 to obtain data having a predetermined dimensionality (e.g., 16). The resulting data having a predetermined dimensionality is referred to as "individual latent data" and may correspond to the multimodal input data 2 described with respect to FIG.
[0167] In step S8, individual decoders 9a-9g are used to reconstruct the individual latent data to obtain individual reconstructed data 17 in the individual initial space (i.e., the same space as the individual input data 12). The individual reconstructed data 17 includes individual data 17a-17g for each modality 2a-2g. The individual reconstructed data 17 may be in the same space as the multimodal reconstructed data 7 of FIG. 1, and may be identical, or may be in a different space (individual latent space).
[0168] In step S9, the individual input data 12 from each modality 2a-2g is compared with the corresponding individual reconstruction data 17a-17g to obtain a comparison result. The better each individual variational autoencoder 10 is, the more similar its input data 12 and reconstruction data 17 are. The comparison result may be a loss function.
[0169] 5, the weights of the individual variational autoencoders 10 are updated as a function of the respective comparison results. In particular, based on the comparison results obtained by comparing the input data 12 of modality 2a with the respective reconstruction data 17a, the individual encoder weights of the individual encoder 8a and the individual decoder weights of the individual decoders are updated by backpropagation. This is performed similarly for each individual variational autoencoder 10.
[0170] 5, in step S21, the steps of training the individual variational autoencoder 10 (steps S6 to S10) are repeated to reduce the comparison result and thus improve the individual variational autoencoder 10. Steps S6 to S10 may be repeated until a desired comparison result is obtained or until a predetermined number of runs have been performed.
[0171] In step S11 of Figure 5, the individual latent data of the trained variational autoencoder 10 are used as multimodal input data 2 for the multimodal variational autoencoder 3 described in terms of Figures 1-3. Following step S11, the method of Figure 5 performs method steps S1-S4 using the individual latent data of the trained variational autoencoder 10 as multimodal input data 2 for the variational autoencoder 3.
[0172] Figure 6 shows another representation of the training procedure for neural network 1. In Figure 6, boxes 13, 14, and 15 represent model selection 13, individual optimization 14, and hyperparameter optimization 15, respectively.
[0173] In step S22, individual input data 12 for modalities 2a to 2g are collected. Steps S23 to S25 are part of the individual optimization, which includes steps S6 to S11 described with reference to FIG. 5. In step S24, a search space for hyperparameters of one individual variational autoencoder 10 is defined (including weights, number of layers, activation functions, channel size, etc.). In step S25, the architecture of the individual variational autoencoder 10 is optimized, particularly along steps S6 to S11. Step S23 indicates that steps S24 and S25 are performed for each modality 2a to 2g. The result of steps S23 to S25, i.e., the output of the individual optimization 14, is an optimized variational autoencoder 10 for each modality 2a to 2g.
[0174] This output is used as input to step S26, where a multimodal variational autoencoder 3 is trained against the fixed model architecture defined in steps S23-S25. Step S26 may include steps S1-S4 defined previously. Step S26 may include optimizing hyperparameters in the latent space, resulting in a joint representation of all modalities 2a-2g in the latent space 6. The optimization in steps S25 and S26 is a hyperparameter Bayesian optimization.
[0175] Arrow 16 indicates that steps S23-S26 are repeated for different values of the predetermined number of dimensions in order to optimize the loss function of the multimodal variational autoencoder 3 and achieve the best latent space representation of the chemicals.
[0176] The hyperparameters for which the loss function is minimized are saved in step S27. In particular, all information related to the trained and optimized neural network is saved. This includes the latent space variables for each dataset, information about modalities 2a-2g, and any further available information. In step S28 of FIG. 6, application testing is performed using the trained neural network 1.
[0177] The training method described with reference to Figures 1-5 provides a neural network 1 capable of representing polymers in a latent space representation. Specifically, training data and additional data representing polymers can be input to the trained neural network. The trained neural network generates a latent representation of the input data that can be stored in a database. This enables several applications, which are described in more detail below.
[0178] One example of an application of the trained neural network 1 is a database search device 20 (search engine). An example of such a database search device 20 is shown in FIG.
[0179] The retrieval device may implement a variety of functions and support a variety of methods for generating, for example, control data indicating the synthesis specifications of a desired chemical, or synthetic measurement data.
[0180] The database search device 20 of FIG. 7 includes a storage unit 21 which is a random access memory (RAM), an input unit 23, a processor 24 which is a CPU, an output unit 25, and a connection cable 26 which connects the components of the database search device 20.
[0181] The database search device 20 is part of a personal computer (PC). The storage unit 21 has a database 22 and a trained neural network 1 stored thereon. The database 22 includes latent space representations of multiple chemical substances (e.g., polymers) obtained from the trained neural network 1. In particular, to obtain the latent space representations stored in the database 22, the trained neural network 1 receives individual and / or multimodal input data 2, 12 previously used as training data and generates latent space representations in the latent space 6 using multimodal and / or individual variational autoencoders 3, 10.
[0182] Figure 8 illustrates the use of the database search device 20, and Figures 7 and 8 are discussed together below. The database search device 20 is used to search the database 22 for polymers that are the same as or similar to the searched polymer. Figure 9 illustrates the user interface 31 of the database search device 20.
[0183] In step S12 of the method of Figure 8, the input unit 23 receives search input data providing a multimodal representation of the polymer to be searched for. The search input data is provided in a multimodal initial space. The search input data is in the same format as the multimodal input data 2 described above, and is accompanied by data describing multiple modalities 2a-2g of the polymer. Optionally, the search input data includes data describing only some of the modalities 2a-2g.
[0184] The input section 32 of the user interface 31 has drop-down menus 34 and input fields 35 where the user can insert multimodal data 2. FIG. 9 shows the following potential modalities: CAS number, density, pH value, specific NMR data that may be uploaded, and viscosity. For example, a replacement for C12-15-branched linear alcohol is desired. In the exemplary diagram of FIG. 9, an ethoxylated propoxyl corresponding to CAS 1755111905-53-4 has been entered along with accessible physicochemical properties (density, pH value, viscosity, NMR file).
[0185] As explained above, a latent space representation of the multimodal substance data 2 input via interface 32 is generated by processor 24 according to the methods described above. Similar chemicals are searched for within the latent space representation by finding latent space vectors in similar regions relative to the latent space vector corresponding to the input substance, for example, ethoxylated propoxyl.
[0186] The right side of Figure 9 shows the search results. Butoxylated ethoxylic acid has been proposed as a replacement for ethoxylated propoxylic acid, a C13-15-branched linear alcohol corresponding to CAS 120313-48-6, with the physicochemical properties shown.
[0187] The interface can also output other modalities of the desired input material, such as recipes or control data for a chemical reactor.
[0188] In another example, the search input data includes physicochemical properties of a desired chemical, such as a specific thermal conductivity. As a result, the method implemented using the database search device 20 outputs control data indicating a synthesis specification. This control data specifies the components required for a chemical plant and is suitable for controlling them to produce a chemical, which in the described example is a polymer. The control data may include a digital version of a recipe for producing a chemical with the desired properties.
[0189] In step S13, processor 24 is used to convert the search input data into its latent space representation. In particular, multimodal encoder 4 of neural network 1 is used to adjust the dimensionality of the search input data to obtain multimodal latent search data in latent space 6. In this way, by developing the data-driven compression model implemented by encoder 4 and decoder 5, a digital representation of a chemical substance, e.g., a polymer, is obtained.
[0190] In step S14, the processor retrieves latent space representations of previously known polymers from a database 22 stored in the storage unit 21. The database 22 may contain latent space representations of past or known polymers.
[0191] In step S15 of Figure 8, processor 24 compares the potential search data from step S13 to the representations of polymers retrieved from database 22. This may involve calculating a similarity score.
[0192] 8, processor 24 selects at least one polymer from the plurality of polymers represented in database 22 based on the comparison results of step S15. In step S17, the scores are ranked so that a list of similar or close polymers is available in the latent space for further selection.
[0193] Here, processor 24 selects the closest polymer in latent space 6 (eg, the closest Euclidean distance between the points representing the polymer in latent space 6).
[0194] In optional step S17, the selected closest polymers are ranked by distance, ie, according to their similarity to the potential search data.
[0195] In step S18 of FIG. 8, an identifier is obtained from database 22, which includes information regarding analytical data, polymer name, synthesis specifications, etc., associated with the selected polymer.
[0196] In step S19, the output unit 25, which is a display, outputs the identifier of the selected polymer. The identifier is also stored in the database 22. The output identifier and / or its associated synthesis specifications are used to control the synthesis of the new (searched) polymer in step S20. The identifier allows the predetermined synthesis specifications associated with the identified polymer to be retrieved from the specification database 530 (see Figure 11). Step S20 may involve running an application test.
[0197] 10 illustrates a system 500 for generating chemicals based on a synthesis specification generated according to the above-described aspects and embodiments of the method and apparatus for generating control data. In this example, the system includes a user interface 510 and a processor 520 associated with a control unit 540, the control unit 540 configured to receive the control data generated according to the present disclosure. In this example, the control data is provided from a database 530, although in other examples, the control data may be provided from a server. For example, an identifier for a particular set of control data is obtained according to step S18, the identifier referencing the associated synthesis specification and respective control data set.
[0198] Vessels 550, 552 each contain a respective chemical component. Generally, there may be more than two vessels. For illustrative purposes, only two vessels are shown in the example. Valves 560, 562 are associated with vessels 550, 552. Valves 550 and 552 may be controlled to dispense the appropriate amount of each component as a component for synthesizing a selected polymer (step S17) in reactor 570 according to a synthesis specification. A motor 600 of mixer 580 may also be controlled by control unit 540 as a function of control data / synthesis specification. An optional heater 590 may also be controlled according to the synthesis specification. Finally, an outlet valve 610 in fluid communication with the reactor may be controlled by the control unit to provide the chemical to a vessel or a test system 620.
[0199] FIG. 11 illustrates another example of a user interface 41 that can be used to access a computer-implemented method for measuring physicochemical properties of chemical substances. In this example, density measurements of CAS 1755111905-53-4 C13-15 alcohol are desired, but only information about pH, viscosity, and NMR data is available and entered into sections 44 and 45. The input multimodal data is received (see S1 in FIG. 1) and encoded (S2) using a data-driven compression model by a processing device, such as processor 24 in FIG. 7. The trained neural network is deployed to generate encoded latent space data, as described above.
[0200] The processor generates measurement data indicative of the desired measurable physicochemical property (density) of the chemical substance (C13-15-branched and linear, butoxylated ethoxylated alcohols) by decoding the encoded sensor data using the data-driven compression model, which is output at output section 43. As a result, the data-driven model reconstructs the missing modality as input (density) based on the input. Thus, the measurement data can be obtained indirectly via a trained model / neural network using, among other things, a variational autoencoder device as detailed in this disclosure.
[0201] While the present invention has been described according to preferred embodiments, it will be apparent to those skilled in the art that modifications are possible in all embodiments. For example, modalities 2a-2g may be other modalities than those described above. The neural network 1 may also be used for other applications besides the database search device 20 described above. Such applications include, for example, polymer redesign based on output identifiers, polymer synthesis based on output identifiers, novel polymer design based on output identifiers, reduced shot learning, etc.
[0202] In alternative embodiments and applications of the trained autoencoder or database search device, synthetic measurement data of chemical substances is obtained based on available sensor data and an underlying data-driven compression model. It is also contemplated to generate data indicative of a second physicochemical property based on a second physicochemical property, where the first and second properties are associated with different modalities in terms of the multimodal latent space representation.
[0203] 12a and 12b show an example of a model for generating control data and / or measurement data.
[0204] FIG. 12a illustrates the training process of an exemplary model architecture based on an autoencoder architecture.
[0205] The multimodal input data may include multiple physicochemical properties as different modalities. For example, the properties may relate to measurement data recorded by a sensor. Examples include FTIR spectra, rheology, solubility, surface tension, application properties such as Shore hardness, glass transition temperature, or other measured or measurable properties of a chemical substance. The multimodal input data may include synthesis specifications and / or control data related to the synthesis specifications. The synthesis specifications may relate to raw materials, auxiliary materials such as solvents or catalysts, operating conditions of a chemical plant for producing the chemical substance, such as temperature. The multimodal input data may include identifiers related to the chemical substance, such as SMILES strings, chemical structure representations, etc.
[0206] The training dataset may include multimodal input data including measured properties P1, P2 of the chemical substance for each modality. The training dataset may include multimodal input data including a synthesis specification (SynSpec) and / or control data related to a synthesis specification used to manufacture the chemical substance. Such synthesis specifications are well known for manufacturing chemical substances, such as polymerization for polymers or oligomerization for oligomers. The training dataset may include multimodal input data including at least one identifier ID of the chemical substance.
[0207] The autoencoder may include at least one multimodal encoder and at least one multimodal decoder. The encoder may be configured to encode multimodal input data from a training dataset. The encoding may include generating a multimodal probability distribution dependent on multiple modalities provided by the multimodal input data. The probability distribution may include a Gaussian mixture model. The decoder may be configured to decode a latent space representation provided by the Gaussian mixture model into the multimodal input data.
[0208] For training, a loss function may be defined that measures the difference between the multimodal input data from the training dataset and the decoded multimodal output data, which may be minimized or maximized to learn a multimodal probability distribution by adjusting the encoder and decoder weights.
[0209] As shown in Figure 12b, the trained model can be used to generate measurement data and / or control data. In the upper example, the model is used to generate a synthetic specification Syn Spec and / or control data related to the synthetic specification from sensor data for characteristic P2. In the lower example, the model is used to generate sensor data for characteristic P2 from sensor data for characteristic P1.
[0210] 13a and 13b show another example of a model for generating control data and / or measurement data.
[0211] Similar to FIG. 12, FIG. 13 illustrates a model architecture. In this example, an autoencoder architecture is shown that includes a multimodal encoder and decoder in addition to individual encoders and decoders. The individual encoders may map multimodal input data for each modality to individual latent space representations, such as vectors or tensors. The individual representations, such as vectors or tensors, may be concatenated and provided as input to a multimodal encoder configured to encode the concatenated representations for all of a plurality of predetermined modalities into a common latent space. The multimodal decoder may be configured to decode the common latent space representation into individual latent space representations. The individual decoders may be configured to map the individual latent space representations to multimodal output data. A training dataset and training process such as those illustrated in FIG. 12 may be used for training.
[0212] As shown in Figure 13b, the trained model can be used to generate measurement data and / or control data. In the upper example, the model is used to generate a synthetic specification Syn Spec and / or control data related to the synthetic specification from sensor data for characteristic P1. In the lower example, the model is used to generate sensor data for characteristic P2 from sensor data for characteristic P1.
[0213] The autoencoder architectures shown in Figures 12 and 13 are merely examples, and other generative model architectures may be suitable for generating control data and / or measurement data. One further example may be based on or include a Generative Adversarial Network (GAN) architecture. A GAN architecture includes at least two models: a generator model and a discriminator model. The generator may take points from a latent space as input and generate new control data and / or monitoring data, while the discriminator may take control data and / or monitoring data as input and predict whether it is real (from the training dataset) or fake (synthetically generated). Both models may be trained based on a loss function that minimizes and / or maximizes the difference between the generated decisions of the discriminator. In this way, the generator is trained to generate control data and / or measurement data that are as close as possible to the real data. For unsupervised learning, a CycleGAN architecture, which is an extension of the GAN architecture and involves the simultaneous training of two generator models and two discriminator models, may be employed. The weights of the model can be adapted based on an additional consistency loss function. Further details of CycleGANs are described in J.-Y. Zhu, T. Park, P. Isola, and AAEfros, “Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks,” 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 2017, pp. 2242-2251, doi:10.1109 / ICCV.2017.244, and the model architecture described therein can be adopted for the models described herein.
[0214] Although the disclosure has been described in conjunction with several preferred embodiments and examples, other variations can be understood and practiced by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the claims.
[0215] Any steps presented herein may be performed in any order. The methods disclosed herein are not limited to any particular order of these steps. Also, it is not necessary that different steps be performed at any particular location or at any particular node of a distributed system; i.e., each of the steps may be performed at a different computing node using different equipment / data processing.
[0216] As used herein, "identifying" also includes "initiating identifying or causing to be identified," "generating" also includes "initiating generating and / or causing to be generated," and "providing" also includes "initiating identifying, generating, selecting, transmitting, and / or receiving, or causing to be identified, generating, selecting, transmitting, and / or receiving." "Initiating or causing to be performed an action" includes any processing signal that triggers a computing node or device to perform the respective action.
[0217] In the claims as well as in this specification, the word "comprising" or "including" or similar expressions does not exclude other elements or steps and should not be interpreted as limiting the described elements or steps. The indefinite article "a" or "an" does not exclude a plurality. A single element or other unit may fulfill the functions of several entities or items referred to in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation or that further elements cannot be included.
[0218] Providing within the scope of this disclosure may include any interface configured to provide data, which may include an application programming interface, a human-machine interface such as a display, and / or a software module interface. Providing may include communicating or transmitting data to an interface, particularly for display to a user or use of the data by a receiving entity.
[0219] Any disclosures and embodiments described herein relate to the methods, systems, apparatus, devices, chemicals, materials, services, uses, and computer program elements outlined above, and vice versa. Advantageously, benefits provided by any of the embodiments and examples apply equally to all other embodiments and examples, and vice versa.
[0220] All terms and definitions used herein are to be broadly understood and have their ordinary meaning. It is understood that all disclosed methods may be implemented as computer-implemented methods. In all methods involving the generation of control data, the optional step of manufacturing a chemical in accordance with a synthesis specification and / or the control data may be performed using a chemical manufacturing system having a control unit. [Explanation of symbols]
[0221] Explanation of symbols 1. Neural Networks 2. Multimodal Input Data 2a-2g Modalities 3. Multimodal Variational Autoencoder 4 Multimodal Encoder 5 Multimodal Decoder 6 Latent space 7. Multimodal Reconstruction Data 7a~7g Multimodal reconstruction data 8 Individual Encoders 8a~8g Individual Encoders 9 Individual Decoders 9a~9g Individual decoders 10 Individual Variational Autoencoders 12 Individual Input Data 13 Model Selection 14 Individual optimization 15 Hyperparameter Optimization 16 Arrow 17 individual reconstruction data 17a~17g Individual reconstruction data 20 Database Search Device 21 Memory Unit 22 Databases 23 Input Unit 24 processors 25 output units 26 Connection cable 31 User Interface 32 Input Section 33 Output Section 34 Drop-Down Menu 35 input fields 41 User Interface 42 Input Section 43 Output Section 44 Drop-Down Menu 45 input fields 500 Chemical manufacturing systems / chemical plants 510 Interface 520 processor 530 databases 540 Control Unit 550, 552 container 560, 562 valves 570 Components / Ingredients 580 Mixer 590 Heater 600 motor 610 Outlet Valve S1 Receiving multimodal input data S2 Adjusting / reducing the dimensionality of multimodal input data by encoding based on a data-driven compression model S3 Data-driven compression model based decoding S4 Calculation of loss function S5 Encoder / Decoder Weight Update S6 Repeat steps S1 to S5 S7 Adjusting the number of dimensions S8 Decoding of individual latent data based on data-driven compression model S9 Comparison of individual latent data with reconstructed data S10 Encoder / Decoder Weight Update S11 Using individual latent data as multimodal input data S12 Receiving multimodal representations / measurement data S13 Encoding search input data / conversion to latent space S14 Obtaining latent space representations from known / past chemicals S15 Comparing searched data with known data in latent space / Determining distance in latent space S16 Determine the closest point between search results and known chemicals in the latent space / select substance S17 Ranking / selection of chemical sets that are closest to the latent space representation of the search input data according to latent space distance S18 Get selected / closest chemical identifier S19 Acquisition / generation of control data showing synthesis specifications of closest chemical substances S20 Perform control / test applications for the synthesis of selected chemicals according to synthesis specifications S21 Repeat steps S6 to S10 S22 Receiving multimodal input data S23 Execution of steps for each modality S24 Setting the search space for individual autoencoders S25 Optimizing the architecture of variational autoencoders S26 Training a variational autoencoder S27 Hyperparameter Memory Running the S28 Application Test
Claims
1. 1. A method for generating control data indicative of a chemical synthesis specification, comprising: receiving sensor data indicative of one or more measurable or measured physicochemical properties of the chemical; encoding the received sensor data using a data-driven model, the data-driven model being trained to map multimodal input data, the multimodal input data including sensor data and control data as modalities, to encoded output data, the multimodal input data being a multimodal representation, the multimodal input data including sensor data and control data as modalities, and the encoded output data being a latent space representation of the multimodal input data; generating control data indicative of a synthetic specification of the chemical substance by decoding encoded multimodal input data based on the received sensor data using the data-driven model, wherein the data-driven model is trained to map encoded output data to multimodal output data comprising sensor data and control data as modalities, and wherein the multimodal output data comprises a multimodal representation comprising sensor data and control data as modalities; A method comprising:
2. 10. The method of claim 1, wherein substance data including sensor data is received, the substance data including the sensor data as part of a first set of modalities, and the multimodal output data including control data as part of a second set of modalities.
3. The method of claim 2 , wherein the first set of modalities and the second set of modalities comprise at least a portion of a predetermined plurality of modalities.
4. 10. The method of any one of the preceding claims, wherein the multimodal input data relates to at least one or more measurable or measured physicochemical properties, one or more synthesis specifications, control data indicative of one or more synthesis specifications, composition of the chemical substance and / or identifier of the chemical substance.
5. 10. The method of any one of the preceding claims, wherein control data and measurement data indicative of one or more measurable or measured physicochemical properties of the chemical produced in accordance with the control data are generated, the one or more physicochemical properties of the measurement data being different from the one or more physicochemical properties of the received sensor data.
6. 10. The method of any one of the preceding claims, wherein the sensor data is received for one or more compositions of a plurality of chemicals, and the sensor data and the one or more compositions for each chemical are provided to the data-driven model to generate control data and / or associated measurement data for each chemical.
7. 10. The method of any one of the preceding claims, wherein the data-driven model comprises at least one multimodal variational autoencoder comprising at least one multimodal encoder and at least one multimodal decoder.
8. 10. The method of any one of the preceding claims, wherein the data-driven model comprises a plurality of individual encoders, each individual encoder assigned to a modality of the predetermined plurality of modalities, and each individual encoder is trained to map the input data from the modality to which it is assigned to a common latent space representation.
9. 10. The method of any one of the preceding claims, wherein the data-driven model comprises a plurality of individual decoders, each individual decoder assigned to a modality of the predetermined plurality of modalities, and each individual decoder is trained to decode the common latent space representation of the encoded input data into modality data of the generated multi-modal input data, the modality data being modal data of the modality to which the individual decoder is assigned.
10. 10. The method of any one of the preceding claims, wherein the synthesis specification relates to the production of the chemical and the control data relates to raw materials and / or operating conditions of the chemical plant for producing the chemical.
11. 10. The method of any one of the preceding claims, wherein the chemical comprises a polymer made from multiple monomers by polymerization.
12. 10. A method according to any one of the preceding claims, wherein said control data and / or measurement data are provided for synthesizing said chemical substance.
13. 10. A method according to any one of the preceding claims, further comprising the step of synthesizing said chemical substance in accordance with said provided control data.
14. 1. An apparatus for generating control data indicating chemical synthesis specifications, comprising: an input interface configured to receive sensor data indicative of one or more measurable or measured physicochemical properties of the chemical; a model engine configured to encode the received sensor data using a data-driven model, the data-driven model being trained to map multimodal input data comprising sensor data and control data as modalities to encoded output data, the multimodal input data being a multimodal representation comprising sensor data and control data as modalities, and the encoded output data being a latent space representation of the multimodal input data; a model engine configured to generate control data indicative of a synthetic specification of the chemical substance by decoding encoded multimodal input data based on the received sensor data using the data-driven model, the data-driven model being trained to map encoded output data to multimodal output data comprising sensor data and control data as modalities, the multimodal output data comprising a multimodal representation comprising sensor data and control data as modalities; and An apparatus comprising:
15. Use of measurement and / or control data generated according to the methods disclosed herein to synthesize chemicals, particularly polymers containing chemicals, having target properties.