Method and apparatus for characterizing chemicals, measuring physicochemical properties, and generating control data for synthesizing chemicals - Patents.com
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BASF SE
- Filing Date
- 2023-04-14
- Publication Date
- 2026-04-21
AI Technical Summary
The prior art is difficult to effectively represent and analyze the multi-dimensional properties of complex chemical substances such as polymers, resulting in time-consuming, costly and wasteful resources in the chemical synthesis process.
Using a data-driven compression model, multimodal chemical data is encoded and decoded by training a neural network to generate multimodal representations to represent the multidimensional properties of chemical substances.
It realizes efficient representation and analysis of multi-dimensional properties of complex chemical substances, reduces the number of measurements and resource consumption during the synthesis process, and improves synthesis efficiency and accuracy.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to methods for generating digital representations of chemical substances, which may involve aspects of training neural networks to represent chemical substances. The present disclosure further relates to computer program products and database search devices for identifying chemical substances. Applications of the digital representations involve the generation of measurement data related to the chemical substances, and the generation of control data related to the synthesis specifications of the chemical substances. The disclosed methods and aspects relate to characterizing chemical substances in terms of their physicochemical properties, composition, and / or identifiers. [Background technology]
[0002] Chemical substances, such as polymers, come in multiple shapes, sizes, and compositions. Small molecules are often represented by their structure (chemical composition). Other basic chemical substances may be represented by their recipes, and detailed descriptions of their manufacturing processes. However, the properties of chemical substances, such as polymers, are often too complex to be represented by their recipes or their structures. It would be desirable to provide a way to represent chemical substances in a more multifaceted way.
[0003] In particular, in the chemical industry, chemicals such as new polymers are increasingly being prepared to customer requirements. This requires the synthesis of new polymers and then performing measurements of various properties. This is very expensive, and in addition, the synthesis of polymers often generates unnecessary waste due to the low success rate of synthesizing polymers that meet customer requirements. In addition, performing the measurements is time consuming and expensive. Thus, there is a need to reduce the number of measurements required to fully analyze a polymer. Digital representations of chemicals can be a means to reduce the effort in determining or measuring the physical or chemical properties of chemicals, creating specifications for the synthesis of chemicals, and / or identifying chemicals for specific purposes.
[0004] One object of the present invention is to provide a method and apparatus for generating digital representations of chemical substances. Further objects include improved uses and applications of the digital representations in the context of chemical synthesis and substance characterization.
[0005] The above mentioned objects are achieved by the methods and devices according to the independent claims. Summary of the Invention [Means for solving the problem]
[0006] According to one aspect, a method is presented for characterizing a chemical entity in a predetermined multimodal representation, the predetermined multimodal representation having a predetermined plurality of modalities, the method comprising: receiving multi-modal substance data including a first set of chemical substance modalities; encoding said multi-modal data using a data-driven (especially compressed) model of chemical substances to generate encoded substance data; generating a predetermined multimodal representation comprising a second set of chemical modalities by decoding the encoded substance data using a data-driven compression model, preferably the multimodal substance data being indicative of chemical physicochemical properties, chemical composition, and / or chemical identifiers; a data-driven compression model is implemented to map input data to encoded output data, the input data being a multi-modal representation of chemical substances, the encoded output data being a latent space representation of the input data; Preferably, the first set of modalities is different from the second set of modalities, the first and second sets of modalities being composed of a predetermined plurality of modalities, in particular at least one modality of the second set being not included in the first set.
[0007] A modality may be considered as a dataset that indicates a certain type of information related to a chemical substance, in particular a physicochemical property. In multimodal data, many different types of data are included in each multimodal dataset. A modality may be represented by a dedicated data structure or a type of data with a certain number of dimensions. The type of data may be a scalar, a multidimensional vector, or a tensor. For example, a melting temperature may be considered as a physicochemical property that is a modality of a chemical substance. Another modality may be associated with a SMILES (Simplified Molecular Input Line Entry System) representation of a substance. An identifier such as a name (string) or a chemical formula may serve as modal data. In an embodiment, each modality is represented by a specific data structure for computer-implemented processing of the modal data. The data structure may include a set of parameters that characterize the substance.
[0008] In one aspect, the method allows for generating output multimodal material data from "incomplete" input material data for a given multimodal representation. For example, the input data includes, among other modalities, a first modality but not a second modality. With the data-driven compression model, the generated output material data may include the second modality. The method may be understood as providing measurement data from the input data.
[0009] In an embodiment, the first set of majority is a subset of the second set of majority. In an aspect, the method allows the process to be performed using a latent space representation of the input data. The latent space representation can be considered as a vector in the latent space. The process in the latent space can involve searching for vectors, comparing vectors in the latent space associated with different chemicals, defining regions in the latent space, and calculating a similarity index or score.
[0010] In an embodiment, the data-driven compression model includes at least one trained neural network implemented to receive multi-modal input data comprising a predetermined number of modalities, encode the input data into a latent space representation of the input data, and decode the encoded input data into multi-modal output data comprising the predetermined number of modalities.
[0011] In an embodiment, the neural network is trained based on training data that includes multi-modal training data that includes the predetermined plurality of modalities, e.g., the neural network is trained using all available modalities of the predetermined plurality.
[0012] In an embodiment, the neural network includes a plurality of individual encoders, each individual encoder assigned to a modality of a given plurality of modalities, and each individual encoder implemented to transform input data from the modality to which the individual encoder is assigned into the same dimensionality of the latent space.
[0013] In an embodiment, the neural network includes a plurality of individual decoders, each individual decoder being assigned to a modality of a predetermined plurality of modalities, and each individual decoder being implemented to decode a latent space representation of the encoded input data into modal data of the generated multi-modal substance data, the modality being the modal data of the modality to which the individual decoder is assigned.
[0014] In an embodiment, the characterizing comprises measuring a physicochemical property of the chemical substance, the substance data comprising sensor data, and the characterizing comprises: receiving sensor data indicative of a first measurable physicochemical property of the chemical substance, the sensor data being associated with at least one modality included in a first set; encoding the sensor data using a data-driven compression model of the chemicals to generate encoded sensor data; generating measurement data indicative of a second measurable physicochemical property of the chemical by decoding the encoded sensor data using a data-driven compression model, the measurement data being associated with at least one modality included in a second set.
[0015] In an embodiment, the chemical is a polymer. The method may be viewed as a method for measuring observable values of chemicals for which the observable values are not directly available, but may be derived or regenerated through a data-driven model, which has inherently gained knowledge of the desired output measurement data through pre-training with the "full" substance training data.
[0016] In an embodiment, at least one of the group consisting of generated measurement data, recipe data indicative of chemicals, and identification data indicative of chemicals is output.
[0017] In an embodiment, the method comprises: and / or for a plurality of sample chemicals, using the data-driven compression model to generate a latent space representation of multi-modal sample substance data associated with the sample chemicals, the multi-modal sample substance data having a predetermined plurality of modalities; storing the generated latent space representation of the multi-modal sample substance data associated with the sample chemical in a database.
[0018] In an embodiment, the method comprises: receiving search input data indicative of predetermined physicochemical properties of the chemical substances to be searched, the search input data being provided as a multimodal representation of the chemical substances to be searched for in the first set; encoding the received search input data to generate a latent space representation of the search input data; and comparing the generated latent space representation of the search input data with the latent space representation of the sample chemicals to obtain a comparison result.
[0019] In an embodiment, the method comprises: selecting at least one sample chemical depending on the comparison.
[0020] In an embodiment, the comparing comprises: calculating a similarity score of the latent space representation of the search input data with respect to the latent space representation of the sample chemicals; and / or Determining a similarity range in the latent space for the latent space representation of the search input.
[0021] In an embodiment, at least one modality of the first set and / or the second set comprises a chemical compound specification and / or control data indicative of a chemical compound specification.
[0022] In an embodiment, the characterizing comprises generating control data indicative of a synthesis specification for a chemical substance, particularly a polymer, and the characterizing comprises: providing a first synthetic specification of a reference chemical as at least one modality of the first set; encoding the first synthesis specification into a digital representation of the reference chemical substance using the data-driven compression model; providing a database including a plurality of historical digital representations of historical chemical substances; determining a similarity score of the past digital representation with respect to a digital representation of a reference chemical substance; selecting at least one past representation based on the similarity score and decoding and generating a synthesis specification associated with the at least one selected past representation; and generating control data indicative of the generated synthesis specification.
[0023] In an embodiment, the past chemical is a sample chemical. In an embodiment, the characterizing comprises generating control data indicative of a synthesis specification for a chemical substance, particularly a polymer, and the characterizing comprises: receiving sensor data indicative of a measurable physicochemical property of a chemical; encoding the sensor data using a data-driven compression model of the chemicals to generate encoded sensor data; and generating control data indicative of a chemical composition specification by decoding the encoded sensor data using a data-driven compression model.
[0024] In an alternative embodiment, the first set is equal to the second set of modalities. Thus, the input may be the same as the output modality. However, the output data may refer to a different polymer or material than the input search data or material property data.
[0025] In an embodiment, using the data-driven model includes a process for generating a condensed digital representation of a chemical substance, particularly a polymer, the process comprising: receiving input data that is a multimodal representation of a physicochemical property of a chemical substance and that is indicative of a measurable physicochemical property of the chemical substance; encoding received input data using a data-driven compression model of chemical substances to generate encoded substance data as a function of said input data; and decoding the encoded substance data using a data-driven compression model to generate chemical substance data indicative of measurable physicochemical properties of the chemical substance.
[0026] In an embodiment, this process is implemented as a method according to the second aspect disclosed below.
[0027] According to a second aspect, there is provided a method for generating a representation of a chemical substance, in particular a polymer, the method comprising: receiving input data that is a multimodal representation of a physicochemical property of a chemical substance and that is indicative of a measurable physicochemical property of the chemical substance; encoding received input data using a data-driven compression model of chemical substances to generate encoded substance data as a function of said input data; and decoding the encoded substance data using a data-driven compression model to generate chemical substance data indicative of measurable physicochemical properties of the chemical substance.
[0028] The method of the second aspect is specifically a process configured using a data-driven model.
[0029] The method of the second aspect allows for efficient processing of data referencing properties of chemical substances: the use of a data-driven model in an encoded or compressed form by using multi-modal data may improve processing speed and reduce the amount of resources required in terms of energy, materials and / or computational power.
[0030] In an embodiment, the data-driven compression model maps multiple input modalities to a latent space vector. The multi-modal data including the modalities requires a first amount of data, and the latent space representation of the multi-modal data requires a second amount of data. Preferably, the second amount of data is less than the first amount. Thus, by using the data-driven model, the amount of data processed in the latent space may be reduced without substantially losing information about the material associated with the input multi-modal data.
[0031] Multimodal data relating to a chemical substance includes, for example, various data that represent different aspects of the chemical substance. A first mode or modality may be contemplated as a physical observation, such as melting temperature, hardness, acidity, etc., and a second mode or modality may be contemplated as a structural aspect, such as the chirality of an enantiomer. Spectroscopic aspects may also be considered as modalities.
[0032] In an embodiment, the data-driven compression model is implemented as a trained neural network, for example, the encoding step may include providing the input data, which is a multi-modal representation, as input to a trained neural network.
[0033] According to one aspect, a method for providing a data-driven compression model is disclosed, comprising providing multimodal input data providing a multimodal representation of chemical substances in a multimodal initial space as an input to a neural network to be trained, the neural network comprising a multimodal variational autoencoder including a multimodal encoder and a multimodal decoder, the multimodal encoder including a multimodal encoder layer with multimodal encoder weights defining how the multimodal encoder layer transforms the data, and the multimodal decoder including a multimodal decoder layer with multimodal decoder weights defining how the multimodal decoder layer transforms the data; adjusting the dimensionality of the multimodal input data using a multimodal encoder layer to obtain multimodal latent data in a latent space; decoding the multimodal latent data using a multimodal decoder layer to obtain multimodal reconstruction data in a multimodal initial space; Computing a loss function of a multimodal variational autoencoder for the current set of multimodal encoder and decoder weights based on the multimodal input data and the multimodal reconstruction data; providing a neural network having a current set of multi-modal encoder and decoder weights as a data-driven compression model.
[0034] According to one aspect, there is provided a further method for generating a representation of a chemical substance, particularly a polymer, in a latent space representation, the method comprising training a neural network; and providing multimodal input data providing a multimodal representation of chemicals in a multimodal initial space as input to a neural network to be trained, the neural network comprising a multimodal variational autoencoder including a multimodal encoder and a multimodal decoder, the multimodal encoder including a multimodal encoder layer having multimodal encoder weights that define how the multimodal encoder layer transforms the data, and the multimodal decoder including a multimodal decoder layer having multimodal decoder weights that define how the multimodal decoder layer transforms the data; adjusting the dimensionality of the multimodal input data using a multimodal encoder layer to obtain multimodal latent data in a latent space; decoding the multimodal latent data using a multimodal decoder layer to obtain multimodal reconstruction data in a multimodal initial space; and computing a loss function of a multimodal variational autoencoder for the current set of multimodal encoder and decoder weights based on the multimodal input data and the multimodal reconstruction data.
[0035] In an embodiment, the method includes a step of outputting the trained neural network and / or configuration data indicative of a latent space representation of the chemical as a digital representation of the chemical.
[0036] The digital representation may further include an identifier for the chemical, multiple modalities associated with the chemical in terms of measurable data for the chemical, and / or specification data for synthesizing the chemical in terms of recipe data.
[0037] Using a multimodal variational autoencoder on multimodal input data allows us to process data from multiple modalities at once, combining modalities of the same chemical together to represent the chemical in a single latent space representation. As a result, a neural network is trained to provide a reliable representation of the chemical in the latent space.
[0038] The generated digital representation requires less data for a complete description including all measurable properties of a chemical substance, thus reducing the amount of computational and memory resources required to process the data representing each chemical substance. The trained neural network can therefore be considered as a data-driven compressed model.
[0039] The combination of the multimodal decoder and encoder may be considered as an embodiment of a data-driven compression model suitable for generating digital representations of chemical substances. In particular, the autoencoder disclosed herein may be considered as an embodiment of a data-driven compression model.
[0040] A chemical may be a form of matter that has certain chemical composition and properties. Examples of chemicals include polymers and molecules. A polymer may be composed of linked monomers.
[0041] The neural network trained according to the method of the first aspect can provide a representation of chemical substances in a latent space representation. The latent space can be an abstract multidimensional space into which the neural network maps what it has learned from its training data. The latent space representation can be a mathematical representation of training data with an adjusted (often reduced) dimensionality. "Adjusting" the dimensionality can mean increasing, decreasing or maintaining the dimensionality, in particular to reach a desired (predetermined) dimensionality.
[0042] As used herein, a modality is information about a chemical from a particular source and / or sensor. Different modalities can be information about a chemical from different sources and / or sensors. Different modalities can be images of a chemical obtained by a camera, spectroscopic images of a chemical, recipes for a chemical, simulation data for a chemical, test data from tests on a chemical, etc. Information (data) from multiple modalities can be represented as multimodal input data.
[0043] The dimensionality of a multimodal input is defined in particular by the characteristics of the measurement: for example, in spectroscopy, the response of a chemical is measured at different wavelengths, and the wavelength range is fixed to the region where one expects to see the chemical response.
[0044] The multimodal initial space may correspond to a space in which multimodal input data is provided to the neural network. The multimodal initial space may be defined by its dimensionality. The multimodal initial space may be a space in which data from different sources and / or sensors are provided directly from the sources and / or sensors or may have been subjected to some modification, such as conditioning and / or preprocessing.
[0045] A multimodal variational autoencoder (multimodal VAE) may be a variational autoencoder that combines and / or takes into account data from different modalities. The structure of a multimodal variational autoencoder may be similar to the structure of a standard variational autoencoder.
[0046] The multimodal encoder may be configured to receive as input multimodal input data and adapt its dimensionality to obtain multimodal latent data in the latent space, in particular, the dimensionality of the multimodal latent data is smaller than the dimensionality of the initial multimodal input data.
[0047] A multimodal decoder may be used to decode back the data encoded by the multimodal encoder. In other words, the multimodal decoder receives multimodal latent data as input and outputs multimodal reconstructed data in a multimodal initial space (i.e., the same space as the initial multimodal input data).
[0048] During training, a loss function is calculated for the multimodal variational autoencoder. The loss function may indicate how good the multimodal variational autoencoder is. The loss function may correspond to the difference between the multimodal input data and the multimodal reconstruction data. The loss function may be determined by a mixture of expert techniques or a multiplication of expert techniques. In particular, the smaller the loss function, the better the multimodal variational autoencoder is.
[0049] The current set of multimodal encoder and decoder weights refers to the set of multimodal encoder and decoder weights of the current run (iteration) of training of the neural network.
[0050] A trained neural network can be used to receive training data related to known chemicals or other multi-modal data as input and provide a latent space representation thereof as output.
[0051] Applications of the latent space representations of chemicals provided by trained neural networks are described in more detail below. Examples include providing a search engine for searching similarities between chemicals in the latent space. Other examples include chemical redesign and / or design of new chemicals.
[0052] Furthermore, the latent space representation of a chemical may be used, for example, in an automated manufacturing process to produce that chemical: in fact, automated manufacturing machines (robots) often require very rich and compact information about the chemical being manufactured, which a latent space representation can provide.
[0053] According to a further embodiment, the method comprises: The method further includes updating the multi-modal encoder weights and / or the multi-modal decoder weights based on the loss function.
[0054] The results of the multimodal variational autoencoder, i.e. the multimodal latent data and the multimodal reconstructed data, can be varied by modifying the multimodal encoder weights and the multimodal decoder weights.
[0055] Thus, updating the weights of the multimodal encoder and / or the weights of the multimodal decoder based on the loss function makes it possible to modify the results (outputs) of the multimodal variational autoencoder in order to reduce the loss function and / or improve the neural network during its training, in particular to provide a latent space representation of the chemical-related input data as accurate and / or convenient as possible.
[0056] According to a further embodiment, the method comprises: The method further includes repeating the steps of providing multi-modal input data, adjusting the dimensionality, decoding the multi-modal latent data, calculating a loss function, and / or updating the weights to reduce the loss function.
[0057] By repeating these steps, the loss function may be reduced more and more until a satisfactory and reliable neural network is obtained. Training of the neural network may be terminated once the loss function falls below a predetermined threshold after a predetermined number of iterations (executions) of the steps of providing multimodal input data, adjusting the dimensionality, decoding multimodal latent data, calculating the loss function, and / or updating the weights to reduce the loss function.
[0058] According to a further embodiment, the neural network further comprises a separate variational autoencoder assigned to each modality, the separate variational autoencoder comprising a separate encoder and a separate decoder, the separate encoder comprising a separate encoder layer with separate encoder weights defining how the separate encoder layer transforms the data, the separate decoder comprising a separate decoder layer with separate decoder weights defining how the separate decoder layer transforms the data, and the method further comprises: Each variational autoencoder is fed with separate input data representing only the assigned modality in each separate initial space; Adjusting the dimensionality of the respective input data using respective encoder layers to obtain respective latent data having a predetermined dimensionality; Decoding the respective latent data using respective decoder layers to obtain respective reconstruction data in respective initial spaces; comparing the respective input data with the respective reconstruction data to obtain a comparison result; updating the respective encoder weights and / or the respective decoder weights based on the comparison results; training each individual variational autoencoder by repeating the steps of: inputting individual input data, adjusting the dimensionality of the individual input data, decoding the individual latent data, comparing the data, and updating the individual encoder weights and / or the individual decoder weights to reduce the comparison result; The method is: The method further includes using the respective latent data from the plurality of respective variational autoencoders as multimodal input data for the multimodal variational autoencoder.
[0059] The individual variational autoencoders may be used to pre-process data from different modalities before inputting them into the multimodal variational autoencoder. For example, the individual variational autoencoders perform pre-processing or conditioning of the individual data representing the individual modalities separately. The individual variational autoencoders may convert data from all modalities into the same format and / or the same number of dimensions (corresponding to the "predetermined number of dimensions") in the individual latent data. This may facilitate processing by the multimodal variational autoencoder by receiving the individual latent data and / or information about the individual variational autoencoders (such as hyperparameters determined during training of the individual variational autoencoders) as multimodal input data.
[0060] Each individual variational autoencoder (individual VAE) may be a variational autoencoder that considers only data from a single modality (the assigned modality). Thus, data from each modality is processed separately by a single individual variational autoencoder. The structure of the individual variational autoencoders may be similar to the structure of a standard variational autoencoder.
[0061] The individual encoders may be configured to receive as input the individual input data of the assigned modality and adjust (increase, decrease, or maintain) their dimensionality to obtain the individual latent data in the individual latent spaces.
[0062] The individual decoders may be used to decode back the data encoded by the individual encoders, in other words, they receive the individual latent data as input and output the individual reconstructed data in the individual initial space (i.e., the same space as the individual initial input data).
[0063] During training, an individual loss function is calculated for each individual variational autoencoder. This individual loss function may indicate the degree of goodness of the individual variational autoencoder. The individual loss function corresponds to the difference between the individual input data and the individual reconstruction data and may be expressed as a comparison result thereof. In particular, the smaller the comparison result, the better the individual variational autoencoder.
[0064] Thus, updating the individual encoder weights and / or the individual decoder weights based on the loss function makes it possible to modify the results (outputs) of the individual variational autoencoders in order to reduce the loss function and / or improve the neural network during training.
[0065] According to a further embodiment, the method comprises: repeating training of each individual variational autoencoder for different predetermined dimensionalities of the individual latent data; The method further includes selecting the predetermined number of dimensions for which the loss function is minimized as the predetermined number of dimensions for the final trained neural network.
[0066] In particular, the loss function of the multimodal variational autoencoder is calculated for different sets of multimodal input data, corresponding to the respective latent data obtained from the respective variational autoencoders for different predetermined dimensionalities. By varying the predetermined dimensionality, the loss function can be reduced. The final trained neural network is a trained neural network. Preferably, the final trained neural network has hyper-parameters and / or weights that minimize the loss function. The predetermined dimensionality can be considered as a hyper-parameter.
[0067] According to a further embodiment, the loss function of the multi-modal variational coder is determined through a mixture of experts or a multiplication of expert techniques.
[0068] Mixing expert techniques involves breaking down the predictive modeling task into subtasks, training an expert model for each, developing a gating model that learns which experts to trust based on predicted inputs, and combining the predictions. Mixing expert techniques models a probability distribution by combining the outputs from several simpler distributions.
[0069] According to further embodiments, one of the modalities for describing a chemical substance provides a spectroscopic, a rheological, a thermal, a chemical, a structural, a solubility, a dispersibility, a viscosity, and / or a surface tension representation of the chemical substance.
[0070] In particular, one modality may correspond to any analytical and / or characterization method of chemical substances, including (i) all types of spectroscopic data, such as Fourier transform infrared spectroscopy (FTIR), Raman spectroscopy, Raman microscopy, UV spectroscopy, nuclear magnetic resonance spectroscopy (NMR), mass spectrometry (MS), gas chromatography-mass spectrometry (GCMS), (ii) thermal analysis data, such as dynamic mechanical analysis (DMTA), differential scanning calorimetry (DSC), thermogravimetric analysis (TGA), dynamic mechanical analysis (DMA), (iii) structural property data, such as X-ray diffraction, and (iv) data from proprietary analytical techniques that measure physical properties of polymers, such as viscosimetry, surface tension, etc.
[0071] For example, FTIR always contains errors arising from domain-specific sources, and therefore neural networks are trained to be able to deal with such errors and deal with chemical data in a more general way.
[0072] According to a further embodiment, at least two modalities provided in the separate or multimodal representation of the multimodal input data comprise data in different formats.
[0073] Data in different formats may include images, tables, text, etc. Data in different formats may also include data having different dimensionality and / or data of different structures or representations.
[0074] According to a further embodiment, at least a portion of the multi-modal input data and / or the individual input data is sequential data.
[0075] To deal with sequential data, the neural network can be a convolutional neural network.
[0076] According to a further embodiment, the individual input data and / or the multi-modal input data comprises augmented data.
[0077] Augmented data may designate artificially generated data, for example by generating a slightly modified copy of the initial data, which allows to increase the amount of training data (individual input data and / or multimodal input data) and thereby improve the training of the neural network.
[0078] According to a further embodiment, the step of adjusting the dimensionality of the respective input data using the respective encoder layers comprises increasing the dimensionality of the respective input data for at least one modality and reducing the dimensionality of the respective output data for at least another modality.
[0079] Depending on the assigned modality, the individual encoder layers increase or decrease the dimensionality of the individual input data to obtain individual data in a predetermined dimensionality. For example, data representing some modalities may be scalar (optionally with error bars) and its dimensionality may be increased to achieve the predetermined dimensionality.
[0080] According to a further aspect, a method for measuring a physicochemical property of a chemical substance, in particular a desired, pre-set and / or predetermined physicochemical property of a chemical substance, comprises the steps of: receiving sensor data indicative of a first measurable physicochemical property of the chemical; encoding the sensor data using a data-driven compression model of the chemicals to generate encoded sensor data; and decoding the encoded sensor data using a data-driven compression model to generate measurement data indicative of a second measurable physicochemical property of the chemical.
[0081] Preferably, the encoding and decoding steps are implemented according to the respective steps of the first aspect of the present disclosure or an embodiment thereof.
[0082] In an embodiment, a data-driven compression model is implemented to map input data to encoded output data according to the method of the first aspect for generating a representation of a chemical substance, said input data being a multi-modal representation of physico-chemical properties of the chemical substance and said encoded output data being a latent space representation of the input data.
[0083] By determining the properties of chemical substances in the digital latent space representation, it is possible to avoid the need to perform measurements and experiments with samples of each substance. Furthermore, the generated data showing physicochemical properties using the presented data-driven method and compression model can be used as training data for artificial intelligence purposes. Therefore, an aspect of the present disclosure is also the use of generated measurement data as training data. While sensor data can be obtained by hardware measurement, generated measurement data can be considered as synthetic measurement data.
[0084] In an embodiment, the provided or received sensor data relates to a first type of physical measurement and the generated measurement data relates to a second type of physical measurement. For example, it may be contemplated to provide X-ray diffraction data as received sensor data and near-infrared spectra as generated measurement data. The digital representation allows to partially omit the energy- and resource-intensive technological process of measuring a sample of a substance.
[0085] According to a still further aspect, a measuring device for measuring physicochemical properties of a chemical substance, in particular a desired, pre-set and / or predetermined physicochemical property of a chemical substance, comprising: an interface device implemented to receive sensor data indicative of a first measurable physicochemical property of the chemical; an encoder device implemented to encode the received sensor data and to generate and output encoded sensor data; and a decoder device implemented to generate measurement data indicative of a second measurable physicochemical property of the chemical and to decode the encoded sensor data.
[0086] The encoder device is preferably implemented to map input data to encoded output data according to the method of the first aspect for generating a representation of a chemical substance, said input data being a multimodal representation of physicochemical properties of the chemical substance in a multimodal initial space and said encoded output data being a latent space representation of the input data.
[0087] The decoder device is preferably implemented to map input data to decoded output data according to the method of the first aspect for generating a representation of a chemical substance, said input data being a multimodal latent space representation of physicochemical properties of the chemical substance and said decoded output data being multimodal reconstruction data in a multimodal initial space.
[0088] According to one aspect, there is provided a computer program product comprising instructions which, when executed by a computer, cause the computer to perform the method according to the first aspect or a method according to an embodiment thereof.
[0089] Computer program means, such as a computer program product, may be embodied on a memory card, a USB stick, a CD-ROM, a DVD or as files that can be downloaded from a server in a network, for example such files may be provided by transferring the files constituting the computer program product over a wireless communication network.
[0090] The embodiments and features described with reference to the training method of the first aspect or embodiments thereof apply, as appropriate, to the computer program product of the second aspect.
[0091] According to one aspect, there is provided a database search device specifically implemented to identify chemical substances having predefined physicochemical properties. The database search device comprises: a storage unit for storing a database and a trained neural network, the database comprising representations of a plurality of chemical substances obtained using the trained neural network, in particular in a latent space, and the trained neural network comprising an encoder and a decoder; an input unit for receiving search input data indicative of chemical substances to be searched for; 1. A processor comprising: adjusting the dimensionality of the search input data using an encoder to obtain encoded search data in the encoded space; comparing the encoded search data to a representation of a plurality of chemical substances in a database; a processor configured to select at least one chemical substance from the plurality of chemical substances represented in the database based on a comparison between the encoded search data and the representations of the plurality of chemical substances; and an output unit for outputting an identifier of the selected chemical substance.
[0092] The trained neural network may comprise a variational autoencoder that includes a multimodal encoder and a multimodal decoder.
[0093] The database search device may be part of a computer, particularly a personal computer or an industrial computer. A trained neural network may be used to provide a latent space representation of chemical substances. In particular, the database includes latent space representations of a plurality of chemical substances obtained using the trained neural network. For example, the database may be updated periodically and / or continuously as new data regarding chemical substances is obtained.
[0094] The storage unit storing the database and the trained neural network may be any type of temporary or permanent storage device (memory). The processor may be a central processing unit (CPU) or the like configured to access the database and / or execute the neural network stored therein. The input unit may include a user interface for receiving search input data from a user, or may be any unit that can access search input data stored in the storage unit or the like.
[0095] The search input data is data that has not yet been input to a neural network and / or has no stored latent representations. The search input data may be in the same format as the multimodal input data described above. The search input data may also be incomplete data representing chemical substances, for example including only partial representations of the chemical substance's modalities.
[0096] A neural network can be used to transform the search input data into the same latent space representation as the data in the database, in particular using a neural network multi-modal encoder that can combine multiple modalities of the search input data.
[0097] Preferably, the potential search data is of the same representation (and dimensionality) as the data in the database. Comparison of the potential search data with the data in the database (i.e., the representations of the multiple chemicals) may be performed by directly comparing the potential search data with the data in the database. For example, the numerical values of the potential search data assigned to each of its dimensionalities may be directly compared with the numerical values of each data in the database assigned to the same dimensionality. This comparison allows for determining the similarity between the potential search data and each of the representations of the multiple chemicals in the latent space. A comparison score proportional to the similarity may be assigned to each of the representations of the multiple chemicals in the latent space.
[0098] The at least one selected chemical may be the chemical whose latent space representation in the database is closest to the latent search data (e.g., has the highest comparison score). The plurality of selected chemicals are the N chemicals represented in the database that are closest to the latent search data. This relies on the fact that similar chemicals have similar latent space representations.
[0099] The output unit may be a user interface such as a display, a touch screen, etc. The identifier of the selected chemical may include the name, chemical composition, reference number, or other identifying information of the selected chemical. The output unit may output (display) the identifier to a user, by storing it in a storage unit, etc.
[0100] A database search device can be used to identify chemicals based on their multi-modal representations (search input data) by performing comparisons in latent space.
[0101] The database search device can be used to perform re-engineering, i.e., to find a representation of a chemical without knowing its recipe, in which case, once an unknown chemical search is found, its recipe can be derived from the recipe of the selected chemical.
[0102] Furthermore, the latent space representation provided by neural networks can alleviate few-shot learning problems (the problem of making predictions based on a limited number of samples) by reducing the dimensionality of the input data and by providing a rich feature space trained on very large datasets.
[0103] According to an embodiment, the neural network is trained according to the method of the first aspect or any embodiment thereof.
[0104] The database searching device may be further configured to perform training of the neural network according to the method of the first aspect or any embodiment thereof.
[0105] A further aspect of the present disclosure includes a method for generating control data indicative of synthesis specifications for chemicals, particularly polymers, the method comprising: providing a first synthetic specification of a reference chemical; encoding the first synthesis specification into a digital representation of the reference chemical substance using the data-driven compression model; providing a database including a plurality of historical digital representations of historical chemical substances; determining a similarity score of the past digital representation with respect to a digital representation of a reference chemical substance; selecting at least one past representation based on the similarity score and decoding a composition specification associated with the at least one selected past representation; and generating control data indicative of the generated synthesis specification.
[0106] The chemical synthesis specification preferably includes all process and recipe data required to manufacture the respective chemical. The synthesis specification data may include control data for operating a chemical plant in machine readable form.
[0107] As a result of the above-mentioned aspects, when the control data is deployed in a chemical plant, i.e. a chemical manufacturing system according to the control data, a synthesis specification is obtained which may result in chemicals similar to the reference chemical. The method also provides an alternative synthesis specification for the reference chemical.
[0108] In an embodiment, a data-driven compression model is implemented to map input data to encoded output data according to the method of the first aspect for generating a representation of a chemical substance, said input data being a multi-modal representation of a chemical substance and said encoded output data being a latent space representation of the input data.
[0109] The database may be configured in a database search device according to the above aspects.
[0110] In an embodiment, the search input data indicated a component of a chemical substance to be replaced by an alternative component, and the selected chemical substance included the alternative component in place of the component to be replaced.
[0111] In an embodiment, the search input data indicates qualitative characteristics of a chemical substance, which may refer to a classification based on predefined regulations, such as the German Hazardous Substances Regulation (Gefahrstoffverordnung-GefStoffV).
[0112] In particular, the identifiers indicative of the components and / or qualitative characteristics to be replaced may be modalities in terms of the encoder and decoder of the neural network.
[0113] Another aspect relates to a method for generating control data indicative of synthesis specifications for chemicals, particularly polymers, the method comprising: receiving sensor data indicative of a measurable physicochemical property of a chemical; encoding the sensor data using a data-driven compression model of the chemicals to generate encoded sensor data; generating control data indicative of a chemical composition specification by decoding the encoded sensor data using a data-driven compression model; A data-driven compression model is implemented to map input data to encoded output data according to the method of the first aspect for generating a representation of a chemical substance, said input data being a multi-modal representation of physico-chemical properties of the chemical substance and said encoded output data being in particular a latent space representation of the input data.
[0114] The presented aspects allow generating control data for synthesizing chemicals according to desired physicochemical properties without performing experiments or test runs of real-world chemical plants, thus facilitating the creation and operation of plants for the production of such chemicals by digital representation in terms of latent space representation.
[0115] In an embodiment, a method for generating control data comprises: applying constraints indicating process or material requirements, in particular biodegradability requirements of chemicals and / or ingredients for producing chemicals, biomass-based requirements, exclusion of toxic ingredients, etc.; and verifying whether the generated synthesis specification satisfies the constraints, particularly before generating the control data.
[0116] The embodiments and features described with reference to the training method aspect of the first aspect or embodiments thereof apply mutatis mutandis to the database device of the method for determining physicochemical properties and generating control data and / or measurement data of the third aspect and other aspects.
[0117] According to some aspects, the processor may further decode the latent representation of the at least one selected chemical substance using a multimodal decoder to obtain reconstructed representation data in the multimodal initial space, and the output unit is configured to output the reconstructed representation data, thereby providing a representation of the selected chemical substance in the initial space that is, for example, understandable and analyzable by a user.
[0118] The disclosed embodiments, inter alia, allow for the replacement of known polymers with molecules having similar performance, including the creation of appropriate synthetic specifications for the replacement polymer / molecule.
[0119] All disclosed methods are preferably computer-implemented. Moreover, possible implementations or alternative solutions of the invention also encompass combinations of features described above or below with respect to the present embodiments (not explicitly mentioned herein). A person skilled in the art may also add individual or separate aspects and features to the most basic form of the invention.
[0120] Further embodiments, features, and advantages of the present specification will become apparent from the following description and dependent claims, when considered in conjunction with the accompanying drawings. [Brief description of the drawings]
[0121] [Figure 1] FIG. 1 is a diagram illustrating a first example of a neural network. [Diagram 2] FIG. 2 illustrates a first embodiment of a method for training the neural network of FIG. [Diagram 3] FIG. 2 illustrates a second embodiment of a method for training the neural network of FIG. [Figure 4] FIG. 13 is a diagram illustrating a second example of a neural network. [Diagram 5] FIG. 5 illustrates a first embodiment of a method for training the neural network of FIG. [Figure 6]6A-6C show different representations of the training method of FIG. 5. [Figure 7] FIG. 2 illustrates a database search device. [Figure 8] FIG. 8 illustrates a method of operation of the database search device of FIG. [Figure 9] FIG. 8 illustrates a user interface for using the database search device of FIG. [Figure 10] FIG. 1 illustrates an embodiment of a chemical production system. [Figure 11] FIG. 2 illustrates a user interface for using a measurement device / service. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0122] In the drawings, like reference numbers designate like or functionally equivalent elements unless otherwise specified.
[0123] The encoders, decoders, and autoencoders presented in this specification may be implemented along the lines of M. Wu, N. Goodman, “Multimodal Generative Models for Scalable Weakly-Supervised Learning”, arXiv:1802.05335, and references therein, which are incorporated herein by reference.
[0124] Figure 1 shows an example of a neural network 1 comprising a multimodal variational autoencoder 3 having a multimodal encoder 4 and a multimodal decoder 5. The neural network 1 is trained according to the method of Figure 2, so that in the following Figures 1 and 2 are described together.
[0125] With respect to the following embodiments, it is understood that the presented neural network embodies a framework for the digital representation of chemical substances. The deployed artificial neural network can be characterized in terms of parameters such as the number and characteristics of the implemented neurons, weights, nodes, connections, and other configuration parameters. The expression "latent space representation" in the context of this application refers to a digital representation of a chemical substance such as a polymer, as follows: "Modality" describing a chemical substance refers to the physicochemical properties of a chemical substance that are observable by measurement and can be represented in a digital or computationally processable manner, for example referring to the spectroscopic, rheological, thermal, chemical, structural, solubility, dispersibility, viscosity, and / or surface tension representation of a chemical substance.
[0126] The latent space representation of a chemical substance is compressed with respect to the data volume required by the multimodal data including modalities. For example, the characterization of a chemical substance in terms of a raw parameter set describing multiple physicochemical properties and a name (e.g., CAS (Chemical Abstracts Service) number) can be considered as a multimodal representation requiring multiple data structures. After generating the latent space representation, latent space data structures that indicate the properties of the same substance are available, and the latent space representation requires fewer and / or smaller data structures. For example, the dimensionality of the latent space representation is less than the dimensionality of the initial multimodal representation. Because the encoder and decoder are trained on the multimodal substance data, the potential information loss due to encoding is reduced or becomes negligible.
[0127] The multimodal encoder 4 and the multimodal decoder 5 form an interface to the latent space representation 6 and are therefore computer-implemented embodiments of the data-driven compression model.
[0128] To train the neural network 1, the neural network 1 receives as input multimodal input data 2 (step S1 in FIG. 2). In the example of FIG. 1, the multimodal input data 2 comprises data representing seven modalities 2a-2g of the same chemical, here a polymer. The data 2 can be understood as a predefined multimodal representation of the polymer. The reference numeral 2 represents the predefined multimodal representation comprising a plurality of the seven modalities 2a-2g. The modality 2a comprises spectroscopic data from a spectroscopic measurement, the modality 2b comprises rheological data, the modality 2c comprises X-ray diffraction data, the modality 2d comprises solubility data, the modality 2e comprises dispersed clay data (indicative of the interaction of the chemical with the layered structure of the clay), the modality 2f comprises surface tension data and the modality 2g comprises viscosity data of the polymer. The data from all dimensions was obtained by performing corresponding measurements on the polymer using a sensor. The neural network 1 receives multimodal input data 2 referring to a plurality of polymers.
[0129] The multimodal input data 2 is provided in an initial space. In the initial space, the data from each modality 2a-2g has its own dimensionality, which here corresponds to the dimensionality of the data sensed by the sensor. Thus, the modalities 2a-2c have a higher dimensionality (5-100) than the modalities 2d-2g, which are scalar (have only one dimension). Alternatively, in the initial space, the data from each modality 2a-2g have the same dimensionality (e.g., 50).
[0130] In an alternative embodiment, the encoder 4 is replaced by separate encoders, each associated with one of the input modalities 2a-2g, and the decoder 5 is replaced by separate decoders, each associated with one of the output modalities 7a-7g.
[0131] In step S2 of Figure 2, the dimensionality of the multimodal input data 2 is modified using a multimodal encoder 4. The multimodal encoder 4 comprises multiple multimodal encoder layers with multimodal encoder weights that define the mathematical operations with which the multimodal encoder 4 transforms the multimodal input data 2. The multimodal encoder weights are some of the parameters that are modified and optimized during training of the neural network 1, as will be explained further below.
[0132] In step S2, the multimodal encoder 4 reduces the dimensionality of the multimodal input data 2 to obtain multimodal latent data in a latent space 6. The latent space representation of the multimodal input data 2, i.e. the multimodal latent data, comprises 16 dimensions in this example.
[0133] In step S2, a multi-modal encoder 4 combines the data from all modalities 2a-2g to form a single dataset that describes the polymer in a latent space 6.
[0134] In step S3, a multimodal decoder 5 is used to decode the multimodal latent data to obtain multimodal reconstruction data 7, 7a-7g in the initial space. This involves modifying the dimensionality of the multimodal latent data to return it to the dimensionality of the initial multimodal input data 2. The multimodal decoder 5 includes multiple multimodal decoder layers, each having multimodal decoder weights that define the mathematical operation of the multimodal decoder 5 on the multimodal latent data. The multimodal decoder weights are some of the parameters that are modified and optimized during training of the neural network 1, as will be explained further below.
[0135] In step S4 of the training method of Fig. 2, a loss function for the multimodal variational autoencoder 3 is calculated. In its simplest form, the loss function indicates the similarity between the multimodal input data 2 and the multimodal reconstruction data 7. Alternative methods for calculating the loss function for the multimodal variational autoencoder 3 include a mixture of experts, a mixture of Gaussians, and / or a multiplication of expert techniques.
[0136] The calculated loss function indicates how well Neural Network 1 is performing during the current run (iteration). The smaller the loss function, the better Neural Network 1 is.
[0137] Figure 3 shows a further embodiment of a method for training the neural network 1 of figure 1. Method steps S1 to S4 of figure 3 are identical to those of figure 2. Depending on the calculated loss function, the neural network may update all or part of the multimodal encoder weights and / or all or part of the multimodal decoder weights in optional step S5 of figure 3. The multimodal encoder weights and / or the multimodal decoder weights are updated by backpropagation.
[0138] As shown in Fig. 3, in step S6, all method steps S1-S5 can be repeated to reduce the loss function and thereby improve the neural network 1. Steps S1-S5 may be repeated for a predefined number of runs or until the calculated loss function is smaller than a predefined loss function threshold. When training stops, the multimodal encoder weights and multimodal decoder weights of the run providing the lowest loss function are retained as the weights leading to the best neural network 1. The trained neural network 1 corresponds to this best run and has its multimodal encoder weights and decoder weights.
[0139] Figure 4 shows a second example of a neural network 1. Figure 5 shows an embodiment of a method for training the neural network 1 of Figure 4. Many elements of the neural network 1 of Figure 4 and the method of Figure 5 are identical to the neural network 1 and training method of Figures 1-3 and apply equally to the description of Figures 4 and 5.
[0140] The difference with the neural network 1 of FIG. 1 is that the neural network of FIG. 4 comprises seven individual variational autoencoders 10, each including an individual encoder 8 and an individual decoder 9. In detail, the individual encoders 8a-8g and the individual decoders 9a-9g correspond to the modalities 2a-2g, respectively. The modalities 2a-2g correspond to the previously described modalities 2a-2g, but their characterization data form the individual input data 12 instead of the multimodal input data 2. The difference between the individual input data 12 and the multimodal input data 2 is that the individual input data 12 is input to the individual variational autoencoder 10, whereas the multimodal input data 2 is input to the multimodal variational autoencoder 3. Furthermore, the individual input data 12 may include data of different dimensions for the different modalities, whereas the multimodal input data 2 may include data of the same dimensions for all the modalities 2a-2g.
[0141] The individual variational autoencoders 10 are for conditioning the data 12 before inputting it to the multimodal variational autoencoder 3. The individual encoders 8a-8g convert the input data 12 from each modality 2a-2g into the same predefined dimensionality, which may be the dimensionality of the latent space 6 (e.g., dimensionality 16).
[0142] In one embodiment, the autoencoder 3 is an optional element, and the respective encoders 8a-8g each convert the input data 12 from each modality 2a-2g into the same predefined dimensionality of the latent space 6. Similarly, the respective decoders 9a-9g map the latent space vectors into respective modalities 17a-17g having particular respective dimensionality.
[0143] Due to the pre-training, different dimensions and modalities are intertwined, so that the separate decoders / encoders 8, 9 interact with latent space vectors with a given dimension. Missing input modalities can be inpainted by an autoencoder structure.
[0144] In particular, in step S6 of Figure 5, individual input data 12 representing each single modality 2a-2g is input into a corresponding individual encoder 8. This means that the individual input data 12 representing modality 2a is input into the corresponding individual encoder 8a, the individual input data 12 representing modality 2b is input into the corresponding individual encoder 8b, and so on.
[0145] In step S7 of Figure 5, each individual encoder 8 modifies the dimensionality of the received individual input data 12 to obtain data having a predetermined dimensionality (e.g., 16). The resulting data having the predetermined dimensionality is referred to as "individual latent data" and may correspond to the multi-modal input data 2 described in conjunction with Figure 1.
[0146] In step S8, individual decoders 9a-9g are used to reconstruct the individual latent data to obtain individual reconstructed data 17 in the individual initial space (i.e. the same space as the individual input data 12). The individual reconstructed data 17 includes individual data 17a-17g for each modality 2a-2g. The individual reconstructed data 17 may be in the same space as the multimodal reconstructed data 7 of FIG. 1, and may be identical, or may be in a different space (individual latent space).
[0147] In step S9, the individual input data 12 from each modality 2a-2g is compared with the corresponding individual reconstruction data 17a-17g to obtain a comparison result. The better each individual variational autoencoder 10 is, the more similar its input data 12 and the reconstruction data 17 are. The comparison result may be a loss function.
[0148] Thus, in step S10 of Fig. 5, the weights of the individual variational autoencoders 10 are updated as a function of the respective comparison results. In particular, based on the comparison results obtained by comparing the input data 12 of modality 2a with the individual reconstruction data 17a, the individual encoder weights of the individual encoders 8a and the individual decoder weights of the individual decoders are updated by backpropagation. This is performed for each individual variational autoencoder 10 in the same way.
[0149] 5, in step S21, the steps of training the individual variational autoencoder 10 (steps S6 to S10) are repeated in order to reduce the comparison result and thus improve the individual variational autoencoder 10. Steps S6 to S10 may be repeated until a desired comparison result is obtained or until a predefined number of runs have been performed.
[0150] In step S11 of Figure 5, the individual latent data of the trained variational autoencoder 10 are used as multimodal input data 2 of the multimodal variational autoencoder 3 described in terms of Figures 1 to 3. Following step S11, the method of Figure 5 performs method steps S1 to S4 using the individual latent data of the trained variational autoencoder 10 as multimodal input data 2 of the variational autoencoder 3.
[0151] Figure 6 shows another representation of the training procedure of the neural network 1. In Figure 6, boxes 13, 14 and 15 represent model selection 13, individual optimization 14 and hyperparameter optimization 15, respectively.
[0152] In step S22, the individual input data 12 of the modalities 2a-2g are collected. Steps S23-S25 are part of the individual optimization, which includes steps S6-S11 described with reference to FIG. 5. In step S24, a search space of hyperparameters of one individual variational autoencoder 10 is defined (this includes weights, number of layers, activation functions, channel size, etc.). In step S25, the architecture of the individual variational autoencoder 10 is optimized, in particular along steps S6-S11. Step S23 indicates that steps S24 and S25 are performed for each modality 2a-2g. The result of steps S23-S25, i.e. the output of the individual optimization 14, is an optimized variational autoencoder 10 for each modality 2a-2g.
[0153] This output is used as input to step S26, where a multimodal variational autoencoder 3 is trained against the fixed model architecture defined in steps S23-S25. Step S26 may include steps S1-S4 defined previously. Step S26 may include optimization of hyperparameters in the latent space, so that a joint representation of all modalities 2a-2g in the latent space 6 is obtained. The optimization in steps S25 and S26 is a hyperparameter Bayesian optimization.
[0154] Arrow 16 indicates that steps S23-S26 are repeated for different values of a given number of dimensions in order to optimize the loss function of the multimodal variational autoencoder 3 and achieve the best latent space representation of the chemicals.
[0155] The hyperparameters for which the loss function is minimized are saved in step S27. In particular, all information related to the trained and optimized neural network is saved. This includes the latent space variables for each dataset, information about the modalities 2a-2g, and all further available information. In step S28 of FIG. 6, application testing is performed using the trained neural network 1.
[0156] The training method described with reference to Figures 1 to 5 provides a neural network 1 capable of representing polymers in a latent space representation. In particular, training data and further data representing polymers may be input to the trained neural network. The trained neural network generates a latent representation of the input data, which may be stored in a database. This enables a number of applications, which are described in more detail below.
[0157] One example of an application of the trained neural network 1 is a database search device 20 (search engine). An example of such a database search device 20 is shown in FIG.
[0158] The retrieval device may implement various functions and support various methods for generating, for example, control data indicating the synthesis specifications of a desired chemical, or synthetic measurement data.
[0159] The database search device 20 of FIG. 7 includes a storage unit 21 which is a random access memory (RAM), an input unit 23, a processor 24 which is a CPU, an output unit 25, and a connection cable 26 which connects the various components of the database search device 20.
[0160] The database search device 20 is part of a personal computer (PC). The storage unit 21 has a database 22 and a trained neural network 1 stored thereon. The database 22 includes latent space representations of a number of chemical substances (such as polymers) obtained from the trained neural network 1. In detail, to obtain the latent space representations stored in the database 22, the trained neural network 1 receives individual and / or multimodal input data 2, 12 previously used as training data and generates latent space representations in the latent space 6 using multimodal and / or individual variational autoencoders 3, 10.
[0161] Figure 8 illustrates the use of the database search device 20, and Figures 7 and 8 are discussed together below. The database search device 20 is used to search the database 22 for the same or similar polymers as the searched polymer. Figure 9 illustrates the user interface 31 of the database search device 20.
[0162] In step S12 of the method of Fig. 8, the input unit 23 receives search input data providing a multimodal representation of the polymer to be searched for. The search input data is provided in a multimodal initial space. The search input data is in the same format as the multimodal input data 2 described above, and involves data describing multiple modalities 2a-2g of the polymer. Optionally, the search input data includes only data describing some of the modalities 2a-2g.
[0163] The input section 32 of the user interface 31 has drop-down menus 34 and input fields 35 where the user can insert the multimodal data 2. FIG. 9 shows the following potential modalities: CAS number, density, pH value, specific NMR data that can be uploaded, and viscosity. For example, a replacement for C12-15-branched linear alcohols is desired. In the exemplary diagram of FIG. 9, an ethoxylated propoxyl corresponding to CAS 1755111905-53-4 has been entered with accessible physicochemical properties (density, pH value, viscosity, NMR file).
[0164] As explained above, a latent space representation of the multimodal substance data 2 input via interface 32 is generated by processor 24 according to the methods described above. Similar chemicals are searched for within the latent space representation, for example, by finding latent space vectors in similar regions with respect to the latent space vector corresponding to the input substance, ethoxylated propoxyl.
[0165] The right side of Figure 9 shows the search results. As an alternative to ethoxylated propoxylic acid, butoxylated ethoxylic acid has been proposed, a C13-15-branched linear alcohol corresponding to CAS 120313-48-6, with the indicated physicochemical properties.
[0166] The interface can also output other modalities of the desired input material, such as recipes or control data for a chemical reactor.
[0167] In another example, the search input data includes physicochemical properties of a desired chemical substance, for example a specific thermal conductivity. As a result, the method implemented with the database search device 20 outputs control data indicative of a synthesis specification. This control data specifies the elements required for a chemical plant and is suitable for controlling them to produce a chemical substance, which in the described example is a polymer. The control data may include a digital version of a recipe for producing a chemical substance with the desired properties.
[0168] In step S13, the processor 24 is used to convert the search input data into its latent space representation. In particular, the multimodal encoder 4 of the neural network 1 is used to adjust the dimensionality of the search input data to obtain multimodal latent search data in the latent space 6. Thus, by unfolding the data-driven compression model implemented by the encoder 4 and the decoder 5, a digital representation of a chemical substance, e.g. a polymer, is obtained.
[0169] In step S14, the processor retrieves latent space representations of previously known polymers from a database 22 stored in the storage unit 21. The database 22 may include latent space representations of past or known polymers.
[0170] In step S15 of Figure 8, the processor 24 compares the potential search data from step S13 to the representations of polymers retrieved from the database 22. This may involve calculating a similarity score.
[0171] 8, processor 24 selects at least one polymer from the plurality of polymers represented in database 22 based on the comparison results of step S15. In step S17, the scores are ranked so that a list of similar or close polymers is made available in the latent space for further selection.
[0172] Here, processor 24 selects the closest polymer in latent space 6 (eg, the closest Euclidean distance between the points representing the polymer in latent space 6).
[0173] In optional step S17, the selected closest polymers are ranked by distance, ie, according to their similarity to the potential search data.
[0174] In step S18 of FIG. 8, an identifier is obtained from database 22, including information regarding analytical data, polymer name, synthesis specifications, etc., associated with the selected polymer.
[0175] In step S19, the output unit 25, which is a display, outputs an identifier of the selected polymer. The identifier is also stored in the database 22. The output identifier and / or its associated synthesis specifications are used to control the synthesis of the new (searched) polymer in step S20. The identifier allows to retrieve predefined synthesis specifications associated with the identified polymer from the specification database 530 (see FIG. 11). Step S20 may involve running an application test.
[0176] 10 shows a system 500 for generating chemicals based on a synthesis specification generated according to the above aspects and embodiments of the method and apparatus for generating control data. In this example, the system comprises a user interface 510 and a processor 520, associated with a control unit 540, configured to receive the control data generated according to the present disclosure. In this example, the control data is provided from a database 530, but in other examples, the control data may be provided from a server. For example, an identifier for a particular set of control data is obtained according to step S18, the identifier referring to the associated synthesis specification and the respective control data set.
[0177] Each of the vessels 550, 552 contains a component of the chemical product. In general, there may be more than two vessels. For illustrative purposes, only two vessels are shown in the example. Valves 560, 562 are associated with the vessels 550, 552. The valves 550 and 552 may be controlled to dispense the appropriate amount of each component as a component for synthesizing the selected polymer (step S17) in the reactor 570 according to the synthesis specification. The motor 600 of the mixer 580 may also be controlled by the control unit 540 as a function of the control data / synthesis specification. The optional heater 590 may also be controlled according to the synthesis specification. Finally, an outlet valve 610 in fluid communication with the reactor may be controlled by the control unit to provide the chemical product to the vessel or to the test system 620.
[0178] In Fig. 11, another example of a user interface 41 is shown that may be used to access a computer-implemented method for measuring physicochemical properties of chemical substances. In this example, measurements of the density of CAS 1755111905-53-4 C13-15 alcohol are desired, but only information about pH value, viscosity, and NMR data is available and entered in sections 44, 45. The entered multimodal data is received (see S1 in Fig. 1) and encoded (S2) using a data-driven compression model by a processing device such as the processor 24 in Fig. 7. The trained neural network is deployed to generate encoded latent space data, as described above.
[0179] The processor generates measurement data indicative of the desired measurable physicochemical property (density) of the chemical (C13-15-branched and linear, butoxylated ethoxylated alcohols) by decoding the encoded sensor data using the data-driven compression model, which is output at the output section 43. As a result, the data-driven model reconstructs the missing modality as input (density) based on the input. Thus, the variational autoencoder device as detailed in this disclosure may be used in particular to obtain the measurement data indirectly via a trained model / neural network.
[0180] Although the present invention has been described according to preferred embodiments, it is clear to those skilled in the art that modifications are possible in all embodiments. For example, the modalities 2a-2g may be other modalities than those described above. The neural network 1 may also be used for other applications than the database search device 20 described above. Such applications include, for example, polymer redesign based on output identifiers, polymer synthesis based on output identifiers, novel polymer design based on output identifiers, reduced shot learning, etc.
[0181] In alternative embodiments and applications of the trained autoencoder or database search device, synthetic measurement data of chemical substances is obtained based on available sensor data and an underlying data-driven compression model. It is also contemplated to generate data indicative of a second physicochemical property based on a second physicochemical property, the first and second properties being associated with different modalities in terms of the multimodal latent space representation.
[0182] It is understood that all disclosed methods may be implemented as computer implemented methods. In all methods involving the generation of control data, the optional step of manufacturing a chemical according to a synthesis specification and / or the control data may be performed using a chemical manufacturing system having a control unit. [Explanation of symbols]
[0183] Reference sign 1. Neural Networks 2. Multimodal Input Data 2a-2g Modalities 3. Multimodal Variational Autoencoder 4. Multimodal Encoder 5 Multimodal Decoder 6 Latent space 7 Multimodal reconstruction data 7a-7g Multimodal reconstruction data 8 Separate Encoders 8a~8g Individual Encoders 9 Separate Decoders 9a~9g Individual decoders 10 Individual variational autoencoders 12 Individual input data 13 Model Selection 14 Individual optimization 15 Hyperparameter Optimization 16 Arrow 17 Individual reconstruction data 17a~17g Individual reconstruction data 20 Database Search Device 21 Storage Unit 22 Database 23 Input Unit 24 processors 25 Output Unit 26 Connection cable 31 User Interface 32 Input Section 33 Output Section 34 Drop-Down Menu 35 Input Fields 41 User Interface 42 Input Section 43 Output Section 44 Drop Down Menu 45 Input Fields 500 Chemical manufacturing systems / chemical plants 510 Interface 520 Processor 530 Database 540 Control Unit 550, 552 container 560, 562 Valve 570 Components / Ingredients 580 Mixer 590 Heater 600 Motor 610 Outlet valve S1 Receiving multimodal input data S2 Adjusting / reducing the dimensionality of multimodal input data by encoding based on a data-driven compression model S3 Decoding based on data-driven compression model S4 Loss function calculation S5 Encoder / Decoder Weight Update S6 Repeat steps S1 to S5 S7 Adjusting the number of dimensions S8 Decoding of individual latent data based on data-driven compression model S9 Comparison of individual latent data with reconstructed data S10 Encoder / Decoder Weight Update S11 Using individual latent data as multimodal input data S12 Receiving multimodal representations / measurement data S13 Encoding search input data / conversion to latent space S14 Obtaining latent space representations from known / historic chemicals S15 Comparing searched data with known data in latent space / determining distance in latent space S16 Determining the closest point between search results and known chemicals in latent space / selecting substances S17 Ranking / selecting the set of chemicals that are closest to the latent space representation of the search input data according to latent space distance S18 Get selected / closest chemical identifier S19 Acquire / generate control data showing the synthesis specifications of the closest chemicals S20 Run a control / test application for the synthesis of selected chemicals according to synthesis specifications S21 Repeat steps S6 to S10 S22 Receiving multimodal input data S23 Execution of steps for each modality S24 Setting the search space for individual autoencoders S25 Optimizing the architecture of variational autoencoders S26 Training a variational autoencoder S27 Hyperparameter memorization Running the S28 Application Test
Claims
1. A method for characterizing a chemical substance in a predetermined multimodal representation having a predetermined number of modalities, Receiving multimodal data (2, 12) of a substance including a first set of modalities of the chemical substance (S12), To generate encoded substance data, the multimodal data (2, 12) is encoded using a data-driven model (3, 10) of the chemical substance (S13), The process involves generating multimodal material data, which includes a second set of modalities of the chemical substance, by decoding the encoded material data using the data-driven model (3, 10), wherein the multimodal material data indicates the physicochemical properties of the chemical substance, the composition of the chemical substance, and / or an identifier of the chemical substance. The data-driven model (3, 10) is implemented to map input data (2, 12) to encoded output data (6), wherein the input data is a multimodal representation (2, 12) of the chemical substance, and the encoded output data is a latent spatial representation (6) of the input data (2, 12). A method wherein the first set of modalities differs from the second set of modalities, and the first set of modalities and the second set of modalities are composed of a predetermined number of modalities, in particular, at least one modality of the second set is not included in the first set.
2. The method according to claim 1, wherein the plurality of first sets are subsets of the plurality of second sets.
3. The method according to claim 1, wherein the data-driven model includes at least one trained neural network (1) implemented to receive multimodal input data (2, 12) including a predetermined number of modalities, encode the input data into a latent space representation (6) of the input data (2, 12), and decode the encoded input data into multimodal output data including a predetermined number of modalities.
4. The method according to claim 3, wherein the neural network (1) is trained on training data including multimodal training data that includes a predetermined plurality of modalities.
5. The method according to claim 3, wherein the neural network (1) includes a plurality of individual encoders (8a to 8g), each individual encoder (8a to 8g) is assigned to a modality (2a to 2g) among the predetermined plurality of modalities, and each individual encoder (8a to 8g) is implemented to convert the input data (2, 12) from the modality (2a to 2g) to which the individual encoder (8a to 8g) is assigned into the same number of dimensions in the latent space (6).
6. The method according to claim 3, wherein the neural network (1) includes a plurality of individual decoders (9a to 9g), each individual decoder (9a to 9g) is assigned to a modality (7a to 7g) among a predetermined plurality of modalities, each individual decoder (9a to 9g) is implemented to decode the latent spatial representation (6) of the encoded input data (2, 12) into modality data of the generated multimodal material data, and the modality is the modal data of the modality to which the individual decoders (8a to 8g) are assigned.
7. Characterization includes measuring the physicochemical properties of a chemical substance, and the substance data includes sensor data, and the characterization includes, Receiving sensor data (2, 12) indicating a first measurable physicochemical property of the chemical substance (S12), wherein the sensor data is associated with at least one modality included in the first set, To generate encoded sensor data, the sensor data (2, 12) is encoded using the data-driven model (3, 10) of the chemical substance (S13), The method according to claim 1, comprising: decoding the encoded sensor data using the data-driven model (3, 10) to generate measurement data representing a second measurable physicochemical property of the chemical substance, wherein the measurement data is associated with at least one modality included in the second set.
8. The method according to claim 7, wherein at least one of a group consisting of the generated measurement data, recipe data indicating the chemical substance, and identification data indicating the chemical substance is output.
9. For multiple sample chemical substances, the data-driven model is used to generate a latent spatial representation (6) of multimodal sample substance data related to the sample chemical substances, wherein the multimodal sample substance data has a predetermined number of modalities, and / or The method according to claim 1, further comprising storing the generated latent spatial representation (6) of the multimodal sample substance data relating to the sample chemical substance in a database.
10. Receiving search input data that indicates predetermined physicochemical properties of a chemical substance to be searched, wherein the search input data is provided as a multimodal representation of the substance to be searched included in the first set, To generate a latent space representation of the search input data, the received search input data is encoded, The method according to claim 1, further comprising comparing the generated latent spatial representation of the search input data with the latent spatial representation of a sample chemical substance in order to obtain a comparison result.
11. The method according to claim 10, further comprising selecting at least one sample chemical substance according to the comparison result.
12. To compare, With respect to the latent spatial representation of the sample chemical substance, the similarity score of the latent spatial representation of the search input data is calculated, and / or The method according to claim 10, further comprising determining the range of similarity within the latent space with respect to the latent space representation of the search input data.
13. The method according to claim 1, wherein at least one modality from the first set and / or the second set includes a synthesis specification for the chemical substance and / or control data indicating the synthesis specification for the chemical substance.
14. Characterization includes generating control data that indicates the synthesis specifications of the chemical substance, which is in particular a polymer, and the characterization is To provide a first synthesis specification of a reference chemical as at least one modality of the first set, Using the data-driven model (3, 10), the first synthesis specification is encoded into a digital representation (6) of the reference chemical substance (S13), To provide a database (22) containing multiple historical digital representations of past chemical substances, With respect to the digital representation of the reference chemical substance, the similarity score of the past digital representation is determined (S16), Based on the similarity score, select at least one past representation (S17), and decode and generate a synthesis specification associated with the selected at least one past representation. The method according to claim 1, comprising generating control data indicating the generated synthesis specifications.
15. The method according to claim 9, wherein the past chemical substance is a sample chemical substance.
16. Characterization includes generating control data that indicates the synthesis specifications of a chemical substance, particularly a polymer, and the characterization is Receiving sensor data (2, 12) that shows the measurable physicochemical properties of the chemical substance (S12), To generate encoded sensor data, the sensor data (2, 12) is encoded using the data-driven model (3, 10) of the chemical substance (S13), The method according to claim 1, comprising generating control data indicating the synthesis specifications of the chemical substance by decoding the encoded sensor data using the data-driven model (3, 10).
17. The method according to any one of claims 9, wherein the first set is equal to the second set of modalities.
18. Using the aforementioned data-driven model includes a process for generating a compressed digital representation of a chemical substance, particularly a polymer, wherein the process is The system receives input data (S12) which is a multimodal representation (2, 12) of the physicochemical properties of the chemical substance and which indicates the measurable physicochemical properties of the chemical substance. To generate substance data encoded as a function of the received input data (2, 12), the input data (2, 12) is encoded using a data-driven model of the chemical substance (3, 10) (S13), The method according to claim 1, comprising generating chemical data representing the measurable physicochemical properties of the chemical substance by decoding the encoded substance data using the data-driven model (3, 10).
19. The process for generating the compressed digital representation is (S1) provides multimodal input data (2) that provides a multimodal representation of the chemical substance in a multimodal initial space as input to a neural network (1) to be trained, wherein the neural network (1) comprises a multimodal variational autoencoder (3) including a multimodal encoder (4) and a multimodal decoder (5), the multimodal encoder (4) includes a multimodal encoder layer having multimodal encoder weights that define how the multimodal encoder layer transforms the data, and the multimodal decoder (5) includes a multimodal decoder layer having multimodal decoder weights that define how the multimodal decoder layer transforms the data, In order to obtain multimodal latent data in the latent space (6), the multimodal encoder layer is used to adjust the number of dimensions of the multimodal input data (2) (S2), To obtain the multimodal reconstruction data in the multimodal initial space, the multimodal latent data is decoded using the multimodal decoder layer (S3), Based on the multimodal input data (2) and the multimodal reconstruction data, the loss function of the multimodal variational autoencoder (3) for the current set of weights of the multimodal encoder and decoder is calculated (S4), The output of configuration data representing the trained neural network (1) and / or the latent spatial representation of the chemical substance is a digital representation of the chemical substance. The method according to claim 18, further comprising providing the trained neural network as the data-driven model.
20. The process for generating the compressed digital representation is Based on the loss function, update the weights of the multimodal encoder and / or the weights of the multimodal decoder (S5), and / or The method according to claim 19, further comprising repeating the steps (S6) of providing multimodal input data (2) (S1), adjusting the number of dimensions (S2), decoding the multimodal latent data (S3), calculating a loss function (S4), and / or updating weights to reduce the loss function (S5).
21. The neural network (1) further includes individual variational autoencoders (10) assigned to each modality (2a to 2g), the individual variational autoencoders (10) each include individual encoders (8) and individual decoders (9), the individual encoder (8) includes an individual encoder layer having individual encoder weights that define how the individual encoder layer transforms the data, the individual decoder (9) includes an individual decoder layer having individual decoder weights that define how the individual decoder layer transforms the data, and the method is Individual input data (12) representing only the assigned modality in the individual initial space are input to the individual variational autoencoder (10) (S6), In order to obtain individual latent data having a predetermined number of dimensions, the number of dimensions of the individual input data (12) is adjusted using the individual encoder layers (S7), To obtain the individual reconstruction data (17) in the individual initial space, the individual latent data is decoded using the individual decoder layer (S8), To obtain the comparison result, the individual input data (12) is compared with the individual reconstructed data (17) (S9), Based on the comparison results, update the weights of the individual encoders and / or the weights of the individual decoders (S10). By repeating the steps (S21) of inputting individual input data (12) (S6), adjusting the number of dimensions of the individual input data (12) (S7), decoding the individual latent data (S8), comparing the data (S9), and updating the weights of the individual encoder and / or the individual decoder to reduce the comparison result (S10), This further includes training each individual variational autoencoder (10), The method described above is The method according to claim 19, further comprising using the individual latent data from a plurality of individual variational autoencoders (10) as the multimodal input data (2) of the multimodal variational autoencoder (3) (S11).
22. The training of each individual variational autoencoder (10) is repeated for each individual latent data with a different predetermined number of dimensions, The method according to claim 21, further comprising selecting the predetermined number of dimensions of the final trained neural network (1) that minimizes the loss function.
23. The method according to claim 19, wherein the loss function of the multimodal variational autoencoder (3) is determined through a mixture of experts or a combination of expert techniques.
24. The method according to claim 19, wherein one of the modalities (2a to 2g) describing the chemical substance provides a spectroscopic, rheological, thermal, chemical, structural, solubility, dispersibility, viscosity, and / or surface tension representation of the chemical substance.
25. A measuring device for measuring the physicochemical properties of the chemical substance described in Claim 7, An interface device implemented to receive sensor data (2, 12) indicating a first measurable physicochemical property of the chemical substance, An encoder device (4) is implemented to encode the received sensor data (2, 12) and generate and output the encoded sensor data (6), The system includes a decoder device (5) that generates measurement data (7, 17) indicating a second measurable physicochemical property of the chemical substance and is implemented to decode the encoded sensor data (6), The encoder device (4) is implemented to map input data (2, 12) to encoded output data (6) according to the method of any one of claims 1 to 24, wherein the input data (2, 12) is a multimodal representation of the physicochemical properties of the chemical substance in a multimodal initial space, and the encoded output data (6) is a latent space representation of the input data (2, 12). A measuring device in which the decoder device (5) is implemented to map input data (6) to decoded output data (7, 17) in accordance with the method of any one of claims 1 to 7, wherein the input data (6) is a multimodal latent space representation of the physicochemical properties of the chemical substance, and the decoded output data (7, 17) is multimodal reconstruction data in the multimodal initial space.
26. A database search device (20) for identifying a chemical substance having predetermined physicochemical properties as described in Claim 10, A storage unit (21) for storing a database (22) and a trained neural network (1), wherein the database (22) includes representations of a plurality of chemical substances having physicochemical properties obtained using the trained neural network (1), and the trained neural network (1) includes an encoder (3) and a decoder (5), and the storage unit (21) An input unit (23) for receiving search input data that provides an expression indicating the predetermined physicochemical properties of the chemical substance to be searched, A processor (24), In order to obtain encoded search data, the number of dimensions of the search input data is adjusted using the encoder (4), The encoded search data is compared with the representations of the multiple chemical substances in the database (22). Based on the comparison result between the encoded search data and the representation of the plurality of chemical substances, at least one chemical substance is selected from the plurality of chemical substances represented in the database (22). A processor (24) configured as follows, A database search device (20) comprising: an output unit (25) for outputting an identifier indicating the selected chemical substance and / or a synthesis specification associated with the identifier of the selected chemical substance.