Method for selecting a polymer from a polymer database
By employing digital representations in polymer databases, similarity searches are enabled, addressing the limitations of existing systems and improving the efficiency of polymer identification and development.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BASF SE
- Filing Date
- 2025-12-17
- Publication Date
- 2026-06-25
AI Technical Summary
Existing polymer databases lack the ability to perform similarity searches based on chemical structure, hindering the identification of polymers with similar properties or substructures, which is crucial for new product development and alternative material selection.
Implementing digital representations of polymers in databases using graph, string, and feature representations allows for similarity searches, enabling the comparison and selection of polymers based on their chemical structures.
Facilitates efficient identification of polymers with similar properties, reducing the need for resource-intensive research by allowing accurate and computationally feasible searches, enhancing polymer development processes.
Smart Images

Figure EP2025087627_25062026_PF_FP_ABST
Abstract
Description
[0001] 240427W001
[0002] 1
[0003] METHOD FOR SELECTING A POLYMER FROM A POLYMER DATABASE
[0004] FIELD OF THE INVENTION
[0005] The invention relates to a method, a computer program and a system for selecting a polymer from a polymer database. Further, the invention relates to a polymer database that can be utilized in the method, the apparatus and the system.
[0006] BACKGROUND OF THE INVENTION
[0007] Generally, polymers are widely used in industrial and / or daily use products due to their broad range of application properties. The use of polymer encompasses amongst others coatings, furniture, automotive applications, lubricants, packages and foams for insulation. Thus, extensive databases for polymers exist in which, for instance, polymers, their synthesis specification and associated application properties are stored. It is then possible to search such databases of known polymers for respective application properties, ingredients, process parameters, etc. However, even today it is generally not possible to search such polymer databases for polymers that are similar to a given polymer structure, due to the complex nature of the polymer structure. Providing such a search possibility could be advantageous in many application contexts.
[0008] SUMMARY OF THE INVENTION
[0009] In nowadays existing polymer databases tend to contain information on ingredients, polymer building blocks, process condition and measurement results, for instance, referring to application properties. Thus, such databases can easily be searched for used ingredients, polymer building blocks or property ranges. However, searches for similar polymer structures is normally not possible. As a result, it is not possible to, for instance, identify polymers containing, for instance, chemical substructures, which might be important for a given application property or to perform similarity searches to identify polymers which are similar in chemical structure to a given polymer. However, such searches can be important in the development of new chemical products. For example, if a specific substructure of a polymer is associated with predetermined properties that are desired for a product, it would be helpful to search polymer databases for already existing polymers with such a substructure instead of performing an extensive and resourceintensive research program for finding such a polymer. Also, if a polymer structure is already known it could be helpful if similar polymers and their respective properties could be found in a database in order to determine alternatives or to determine the properties of the polymer.
[0010] In this context, the inventors have found that providing polymer databases with digital representations that are associated with the structure of a respective polymer, in particular, comprise a graph representation and / or a feature representation would allow for a similarity search of the polymer database that is effective and performable with
[0011] Internal 240427W001
[0012] 2 reasonable computational resources. Thus, the development processes of chemical products comprising polymers could be improved.
[0013] In a first aspect, a method, in particular, a computer implemented method, is presented for selecting a polymer from a polymer database, wherein the method comprises a) providing a polymer database comprising digital representations of a plurality of polymers, wherein a digital representation is associated with a chemical structure of a respective polymer, b) providing a digital representation of at least a part of a target polymer, wherein the digital representation is associated with a chemical structure of the at least a part of the target polymer, c) comparing the digital representation of the at least a part of the target polymer with the digital representations of polymers provided by the polymer database, wherein the compared digital representation comprise a graph representation, a string representation and / or a feature representation, d) selecting at least one polymer from the polymer database based on the comparison, and e) providing the selected at least one polymer.
[0014] In a further aspect, a polymer database is presented configured to be usable in any of the preceding claims, wherein the polymer database comprises a database comprising digital representations of a plurality of polymers, wherein a digital representation is associated with a structure of a respective polymer, wherein the digital representation comprises a graph representation, a string representation and / or a feature representation.
[0015] In a further aspect, an apparatus is presented for selecting a polymer from a polymer database, wherein the apparatus comprises computer means configured for performing method steps comprising a) providing a polymer database comprising digital representations of a plurality of polymers, wherein a digital representation is associated with a chemical structure of a respective polymer, b) providing a digital representation of at least a part of a target polymer, wherein the digital representation is associated with a chemical structure of the at least a part of the target polymer, c) comparing the digital representation of the at least a part of the target polymer with the digital representations of polymers provided by the polymer database, wherein the digital representation comprises a graph representation, a string representation and / or a feature representation, d) selecting at least one polymer from the polymer database based on the comparison, and e) providing the selected at least one polymer.
[0016] In a further aspect, a system is presented for selecting a polymer from a polymer database, wherein the system comprises a) a user interface configured for receiving a digital representation of a polymer and / or a synthesis specification of a polymer, b) an apparatus as described above for selecting a polymer from a polymer database based on the received digital representation of a polymer and / or a synthesis specification of the polymer, wherein the user interface is further configured to provide a synthesis specification of the selected polymer to the user for synthesizing the polymer.
[0017] In an aspect disclosed is a computer-implemented method for providing a polymer database, in particular a polymer database according to an aspect or embodiment, the method comprising:
[0018] Internal 240427W001
[0019] 3 providing a plurality of structure information associated with a plurality of chemical structures, wherein a respective chemical structure of the plurality of chemical structures is associated with a respective polymer of a plurality of polymers; generating, per polymer of the plurality of polymers, a digital representation based on the respective structure information associated with the respective polymer, wherein the digital representation is associated with the chemical structure of the respective polymer, and includes at least one of a graph representation, a string representation, and / or a feature representation; storing, per polymer, the respective digital representation in the polymer database; and providing the polymer database.
[0020] In a further aspect, a computer program product is presented for selecting a polymer from a polymer database, wherein the computer program product comprises program code means for causing the apparatus as described above to execute the method as described above.
[0021] The method refers to a computer-implemented method and thus can be performed by a general or dedicated computer adapted to perform the method, for instance, by executing a respective computer program. Generally, the polymer can be any polymer. Preferably, the polymer is a synthetic polymer. In an embodiment, a synthetic polymer may be a chemical compound which is produced by a chemical production from one or more starting material(s), such as monomers, and which comprises at least two polymer building blocks, e.g. monomer units. A subunit of a polymer, can comprise one or more polymer building blocks of the polymer. The polymer may be prepared from the monomers by commonly known polymerization reactions. The polymer may be produced from a single type of monomers or from different monomers. The monomer units may be distributed randomly or may be present as blocks within the polymer. The polymer may be a linear polymer. The polymer may be a branched polymer. The polymer may be a cross-linked polymer.
[0022] The method comprises providing a polymer database. A database can be a data structure storing predetermined data entities that can be accessed by database external data access processes. The providing of a polymer database can comprise providing an access to a respective polymer database. The providing of a polymer database can also refer to receiving a user input indicating which polymer database to access. Further, the providing of a polymer database can also refer to generating the polymer database, if it does not already exist, based on previously existing data on at least two polymers. The providing of the polymer database can also comprise generating for an existing polymer database digital representations of the polymers in the existing polymer database.
[0023] The polymer database comprises digital representations of a plurality of polymers. Generally, a digital representation refers to a representation of a real-world object in the digital world. The digital representation represents one or more aspects of the real-world object such that they can be processed in the digital domain and are thus accessible for computational methods. A digital representation is associated with a data structure that is configured to be processable by computational means. The digital representation of a respective polymer is associated with a
[0024] Internal 240427W001
[0025] 4 chemical structure of the respective polymer. The digital representation of a polymer can be a data structure that represents the chemical structure of a polymer in the digital domain. The digital representation comprises a graph representation, a string representation and / or a feature representation of the respective polymer.
[0026] A graph representation can comprise a representation of the structural formula of the polymer, optionally, including the arrangement of the atoms in three-dimensional space, the binding of the atoms and / or geometry of the polymer. Moreover, information about isomers or other deviations in the structure of the polymer that fall within the same chemical formula of the polymer can be expressed in a graph representation. The graph representation can be associated with polymer connectivities for instance, with the connectivities between building blocks of the polymer, wherein a connectivity can quantify a stochastic bond between building blocks of the polymer. Preferably, the graph representation represents atoms and bonds of the polymer as nodes and edges in the graph representation, wherein edges are associated with a connectivity. More preferably, the connectivity is represented by a weight defining a probability and / or frequency of the bond represented by the edge being present in or between building blocks of the polymer. A respective description of a preferred graph representation can be found in the article "A graph representation of molecular ensembles for polymer property prediction”, by M. Aldeghi and C. W. Coley, Chem. Sci. , 2022, 13, 10486-10498. The nodes and edges of the graph representation can be associated with further features of the atoms and bonds of the polymer, e.g., nuclear charge as node feature and chemical bond order as edge feature. This preferred graph representation allows to capture the stochastic nature of polymers and make it available for processing. Thus, also the stochastics of a polymer can be taken into account in the respective search. This includes both the bonding probabilities between polymer building blocks, the molar mass distribution of the polymer, byproduct formation, and the use of ingredients that are present as mixtures.
[0027] A feature representation quantifies features of the polymer that represent the structure of the polymer. The feature representation can comprise a feature vector comprising features describing the structure of the polymer. The features can be explanatory variables that represent the structure of the polymer. The features can be associated with or comprise polymer connectivity, for instance, of a stochastic quantification of a binding probability between building blocks of the polymer. The feature vector can be a counted fingerprint, for example, a MACCSkeys or Morgan fingerprint. Preferably, the feature vector is a Morgan fingerprint, wherein a feature of the fingerprint represents the number of occurrences of a certain molecular substructure in the molecule. The feature vector can be a molecular fingerprint encoding the presence or absence of chemical substructures in a binary vector. The feature vector can be an Extended-Connectivity Fingerprint. The feature vector can be obtained by aggregating features from a neural network or can be derived from representation learning, as is described below.
[0028] A string representation is a typographical notation system using, for example, ASCII characters, for representing a structure of a molecule. A string representation can describe a polymer in a language-based representation of chemical structures. A string representation can be generated by traversing a graph representation of a molecule. Preferably, the string representation is a SMILES, SMARTS, BigSM ILES, Big SMARTS, G-BigSMILES, PSMILES, CurlySMILES representation, respectively. More preferably, the digital representation is a G-BigSMILES or
[0029] Internal 240427W001
[0030] 5
[0031] BigSMARTS representation. For example, for a part of a polymer a SMILES representation can be utilized, whereas for a polymer a G-BigSMILES representation can be utilized. BigSMILES allows to address the stochastic nature of polymer structures, BigSMILES can represent macromolecules with an ensemble of possible configurations within linear, branched, random, block, alternating, and graft polymer morphologies. BigSMILES can also represent non- covalent bonding interactions through a generalizable donor-acceptor index. G-BigSMILES is an optional enhancement of the BigSMILES notation, wherein G-BigSMILES can incorporate further information of at least one of a molecular weight, a molecular weight distribution, and a probability of repeat unit connections. A graph representation and a feature representation of a polymer can be derived from a string representation.
[0032] The plurality of polymers stored on the polymer database refers to at least two polymers but preferably refers to more than two polymers. The polymer database can additionally comprise further information on the polymers, like polymer recipes, ingredients, process parameters, feed profiles, temperature profiles, pressure profiles, production facility descriptions, application properties, already performed measurements, etc. Further, the method comprises providing a digital representation of at least a part of a target polymer. The providing can refer to receiving the digital representation from an input of a user using, for instance, respective input unit. Moreover, the providing can also refer to accessing a storage unit on which the digital representation is already stored. The target polymer is a polymer for which a respective similarity search should be performed. If only a part of a target polymer is provided, then a search can be performed for polymers in the polymer database that comprise this structure. If the complete structure of the target polymer is provided one or more polymers of the polymer database similar to the complete structure can be searched. The at least a part of the polymer can comprise a building block of the polymer. The at least a part of the polymer can comprise at least two building blocks of the polymer, for example, at least two monomers. The at least a part of the polymer can comprise more than two building blocks. If only a part of the polymer is searched this part can also be referred to as substructure and the search can be referred to as substructure search. The digital representation is associated with a structure of the at least a part of the target polymer. For example, the digital representation can be associated with subunits, like one or more polymer building blocks, of the target polymer. The digital representation can be further associated with connectivities between subunits, e.g. between one or more polymer building blocks, of the polymer. The digital representation can be further associated with a group of atoms and bonds between one or more subunits of the target polymer. The digital representation can be associated with the complete structure of the target polymer by including connectivity probabilities between the subunits of the polymer. The digital representation of the at least a part of the target polymer also comprises a graph representation, a string representation and / or a feature representation.
[0033] Further, the method comprises comparing the digital representation of the at least a part of the target polymer with the digital representations of polymers provided by the polymer database. The comparing comprises comparing the digital representation of the at least a part of the target polymer with the digital representation of polymers provided by the polymer database if the digital representations are of the same kind, for instance, if both digital representations are graph representations, string representations or feature representations, respectively. If the respective digital representations comprise more than one representation kind, for example, a graph representation
[0034] Internal 240427W001
[0035] 6 and a feature representation, also both can be compared with each other, respectively. If the digital representation of the target polymer, for instance, comprises a graph representation, the graph representation is compared with the graph representation of the polymers provided by the polymer database. If the digital representation of the target polymer, for instance, comprises a string representation, the graph representation is compared with the string representation of the polymers provided by the polymer database. If the digital representation of the target polymer comprises a feature representation, the feature representation is compared with a feature representation of the polymers provided by the polymer database. Preferably, the digital representation of the target polymer is directly provided such that it comprises the same kind of digital representation as is also provided by the polymer database. However, it is generally possible to also generate for the polymers in the polymer database or at least for a part of the polymers in the polymer database a respective digital representation kind from the information provided by the polymer database.
[0036] The comparing can refer to any method that provides some kind of comparison result between the digital representation of the at least a part of the target polymer and the digital representation of the polymers provided by the polymer database. A comparison refers to determine similarities and / or differences between two entities. Thus, a result of the comparison can refer to determining differences and / or similarities between digital representations of the at least a part of the target polymer with the digital representations of polymers provided by the polymer database.
[0037] In a next step, at least one polymer is selected from the polymer database based on the comparison, in particular, based on the result of the comparison. Since the comparison generally determines similarities and / or differences between the digital representations of the at least a part of the target polymer and the digital representations in the polymer database, the selecting can be based on rules that refer to these differences and / or similarities. For example, the selecting can be based on determining the most similar polymers from the polymer database with respect to the target polymer. The selecting can also refer to selecting the most different polymers from the polymer database based on the comparison. The rules for selecting at least one polymer from the polymer database based on the comparison can depend on the respective application. For example, in some applications it can be requested to find polymers that comprise a specific structure of the target polymer, for instance, one or more specific polymer building blocks. In this case the most similar polymers from the polymer database can be selected based on comparing parts of the polymers of the polymer database with the respective part of the target polymer and select polymers of the polymer database with a similar part. However, in some applications it can also be requested that polymers are found that do not comprise one or more predetermined polymer building blocks, wherein in this case, the selecting can refer to selecting polymers from the polymer database that are dissimilar as possible to the one or more polymer building blocks and in particular do not comprise the one or more polymer building blocks.
[0038] The thus selected at least one polymer is then provided. For example, the selected at least one polymer, in particular, its digital representation, can be provided to a user utilizing a user interface together with the information on the at least one polymer that are stored on the polymer database. The polymer can be provided for further testing. The further testing can comprise producing the polymer and subjecting it to one or more further tests with respect to
[0039] Internal 240427W001
[0040] 7 one or more application goals. The polymer can be provided for producing the polymer. Preferably, the providing of the selected at least one polymer comprises providing a synthesis specification of the at least one polymer. The synthesis specification can be utilized for controlling and / or monitoring a synthesis of the at least one polymer. Preferably, the synthesis specification is configured for controlling and / or monitoring a synthesis of the at least one polymer, for instance, utilizing a respective polymer synthesis system. A polymer synthesis system can be a hardware system that is configured to perform the steps included by a synthesis specification for synthesizing a respective polymer. The providing can then comprise implementing the synthesis specification in a synthesis system that is configured for synthesizing a polymer in order to control and / or monitor the synthesis of the at least one polymer. In an embodiment, the synthesis specification of the at least one polymer is provided as part of control data configured for controlling and / or monitoring a synthesis of the selected at least one polymer. The providing thus comprises providing the control data comprising the synthesis specification to control and / or monitor a synthesis of the selected at least one polymer. The controlling and / or monitoring of the synthesis can be performed automatically but can also be performed in a user machine interaction process in which, for instance, a user can confirm a synthesis of a polymer, can correct one or more aspects of the synthesis specification, or can even deny a synthesis of a polymer.
[0041] In an embodiment, the method comprises receiving structure information associated with a chemical structure of the at least a part of the polymer and generating the digital representation of the at least a part of the target polymer based on the structure information. The structure information in the training data set can comprise connectivity information of polymer building blocks, wherein the connectivity information quantifies the stochastic connectivity probability between polymer building block, and / or a molar weight distribution of the polymer. The structure information can further comprise at least one of an amount of polymer building blocks, chemical bond statistics between polymer building blocks, polymer block assignments, average molecular weight, molecular weight distribution, information on composition drifts of polymer building blocks, reactivity ratios between polymerizable groups, branching degree, mixtures of ingredients, polymer blends, polymer tacticity, post-polymerization modifications, crosslinking densities, process data from polymerization and polymer type information. Preferably, the structure information comprises at least one of chemical bond statistics between polymer building blocks, average molecular weight, molecular weight distribution. The structure information can be provided as SMILES, SMARTS, BIG-SMILES, Big SMARTS, G-BigSM ILES, PSMILES, CurlySMILES representation of the at least a part of the target polymer. It has been found by the inventors that in particular SMILES, SMARTS and G-BigSM ILES and BigSMARTS representations are suitable to provide the structure information but at the same time are easy to handle for an operator and do not take much of storage space. However, of course the structure information can also be provided in another form, for instance, as information based on structure measurements, or utilizing, for instance, InChi. Preferably, the receiving of the structure information comprises deriving the structure information from a synthesis specification of the polymer corresponding to the at least a part of the target polymer. Since synthesis specifications are readily available for most polymers, they provide an easy starting point for searching for respective similar polymers or parts of polymers. Since the synthesis specification is also unambiguously associated with the polymer and clearly defines the result of the synthesis, the respective structure information can be derived from the synthesis
[0042] Internal 240427W001
[0043] 8 specification. For example, kinetic models of polymerization can be utilized to determine the binding probability between polymer building blocks. A synthesis specification for a polymer specifies the starting materials, ingredients and utilized processes that allow to synthesize the polymer. For example, the starting materials can refer to one or more monomers, catalysts or other additives in the respective amount, and specify the steps to be performed and in which order they are to be performed for the synthesis, parameters like temperatures and pressures that have to be applied for the reaction to take place, etc.
[0044] The digital representation of the at least a part of the target polymer can then be derived based on the structure information depending on how the structure information is provided respectively known solutions for the step can be utilized. For example, for deriving a feature vector and / or a graph representation from SMILES, SMARTS, BigSMILES, or BigSMARTS respective methods can be utilized. Preferably, the method further comprises deriving a graph structure representation for the at least a part of the target polymer as digital representation based on the structure information and wherein the digital representation of the plurality of polymers of the polymer database comprises a graph structure representation. For example, the graph representations for a polymer can be derived from a SMILES representation of polymer subunits as structure information. Further, the bonds between repeat units and their probability can be provided by the structure information. The graph representation can then be generated by representing atoms and bonds of the polymer derived from the structure information as nodes and edges of the graph representation. The edges can then be associated with the connectivity probability between respective nodes representing the atoms, in particular, between atoms of two or more different polymer building blocks. The nodes and edges of the graph representation can be associated with features, e.g., nuclear charge as node feature and chemical bond order as well as possibility as edge feature. An example, for generating a respective graph representation can be found in in the article "A graph representation of molecular ensembles for polymer property prediction”, by M. Aldeghi and C. W. Coley, Chem. Sci., 2022,13, 10486-10498.
[0045] Additionally or alternatively, the method comprises deriving a feature vector representation for the at least a part of the target polymer as digital representation based on the structure information and wherein the digital representation of the plurality of polymers of the polymer database comprises a feature vector representation. The feature vector representation can be generated directly based on the structure information provided, for instance, based on a BigSMILES representation but can also be generated, for instance, based on an already generated graph structure representation if available. Preferably, the method comprises generating a feature vector as digital representation for the at least a part of the target polymer based on the structure information utilizing a machine learning based feature model configured to generate a feature vector based on the structure information. A machine learning based feature model can be trained based on a polymer data set comprising structure information for a plurality of polymers. Preferably, the training refers to an unsupervised training with an unlabeled data set. The machine learning based feature model can be trained to determine features that are associated with the similarity between polymers in the polymer training data set during the training and construct a feature vector that is indicative of the similarity between polymers without being provided with an already determined feature vector for each polymer in the training data set, hence, without labeled information. Preferably, the structure information in the training data set comprises at least
[0046] Internal 240427W001
[0047] 9 connectivity information of polymer repeat units. Optionally, the structure information can additionally comprise at least one of an amount of polymer building blocks, chemical bond statistics between polymer building blocks, polymer block assignments, average molecular weight, molecular weight distribution, information on distribution shifts of repeat units, compositional drifts of polymer building blocks, reactivity ratios between polymerizable groups, branching degree, mixtures of ingredients, polymer blends, polymer tacticity, post-polymerization modifications, crosslinking densities, process data from polymerization and polymer type information. Preferably, representation learning is utilized based on the unlabeled training data set to train the machine learning based feature model to determine numerical representations of polymers that position the polymer in a numerical representation space in a way that similar polymers are close to each other and dissimilar polymers are far away from each other. Similar in this context can be defined based on a respective similarity measure, for instance, based on a similarity measure threshold that can be predefined, for instance, based on experience.
[0048] Preferably, the machine learning based feature model comprises a neural network algorithm, in particular, a transformer neural network algorithm. The transformer neural network algorithm can for instance be a large language model. A transformer neural network algorithm comprises a) a tokenizer, which convert text into tokens, b) an embedding layer, which converts tokens and positions of the tokens into vector representations, c) one or more transformer layers, which carry out repeated transformations on the vector representations, extracting more and more linguistic information and d) an Un-embedding layer, which converts the final vector representations back to a probability distribution over the tokens. The one or more transformer layers can comprise alternating attention and feedforward layers. The one or more transformer layers can comprise encoder layers and / or decoder layers, or further variants of these layers. Tokens are numerical representations of the respective input text elements. In this embodiment, the structure information, for instance in form of a G-BigSM ILES representation, can be provided in a letter-by-letter format as input to a large language model, wherein the large language model can then be trained based on the database to generate the respective feature vector of the polymer. The large language model can be configured to check automatically that the provided structure information is associated with a polymer that can be synthesized. Further, the large language model can be configured to provide a feature representation that is configured to provide synthesis information. Preferably, the machine learning based feature model is trained by providing a pre-trained transformer neural network algorithm and training the pre-trained transformer model based on the training data set. The pre-trained transformer neural network can be a transformer neural network algorithm that has already been trained for one or more arbitrary purposes. Preferably, the pre-trained transformer neural network is pre-trained for a similar application. For instance, the pre-trained transformer neural network can be trained based on a SMILES representation, when in the final application a G-BigSM ILES is utilized.
[0049] In an embodiment, per polymer of the plurality of polymers, the respective digital representation is associated with connectivity probabilities between subunits of the respective polymer. The probabilistic connectivity data may represent the ensemble nature of polymers, which are stochastic products composed of multiple possible chain configurations rather than a single defined structure. A configuration may include variations in chemical connectivity, such as the identity and order of monomeric units, the presence or absence of branching, and stereochemical
[0050] Internal 240427W001
[0051] 10 arrangements. A configuration may further include variations in the three-dimensional spatial arrangement of atoms, which may result from rotations about single bonds or other degrees of freedom. Such spatial variations may be referred to as conformations, and a configuration may encompass one or more conformations associated with a given connectivity. A configuration may also include variations in end groups, crosslinking sites, or pendant substituents. In certain embodiments, a configuration may be associated with a particular stereochemical arrangement, such as isotactic, syndiotactic, or atactic forms. A configuration of a polymer may include conformational states that may arise without changes in covalent connectivity. Such conformational states may include extended, coiled, folded, or aggregated arrangements in three-dimensional space, which may depend on environmental conditions such as temperature, solvent, or applied stress. The connectivity probabilities may be associated with polymerization kinetics, statistical models, or molar mass distribution data, and may relate to the likelihood of specific bonds forming between repeat units or blocks. Hence, for example, the respective digital representation may, for the connectivity probabilities, be associated with bonding probabilities between polymer building blocks and / or molar mass distribution. Probabilistic connectivity data may be encoded in different ways, for example, as weighted edges in a graph representation, as probability fields in a structured vector, or as aggregated statistics within a fingerprint computed across an ensemble of polymer chains. The ensemble may include multiple sampled configurations of the polymer, and the probabilistic connectivity values may be averaged or otherwise combined to reflect the overall distribution of structural possibilities. Incorporating such probabilistic connectivity data may allow the representation to capture variability in polymer architecture, such as branching, relative alignment of building blocks of a polymer, e.g. blockiness, or compositional drift. The particular ensemble of configurations may be decisive for technical application properties like fire-resistance, biodegradability, stiffness, interfacial tension or melting-point. Taking into account connectivity probabilities may facilitate representing the, typically, ensemble-nature of polymers. For example, a polymer may have a multitude of different configurations, and the connectivity probabilities may represent the likelihood of different configurations, these configurations having different connectivity, e.g. chemical bonds, between respective subunits. Hence, in particular by basing the comparing on digital representations that also associated with the typical ensemble nature of polymers, e.g. by the connectivity probabilities, selecting the polymer from the polymer database may be further enhanced.
[0052] In an embodiment, generating the digital representation of the at least a part of the target polymer includes determining connectivity probabilities between subunits of the target polymer based on the structure information, the connectivity probabilities at least relating to a likelihood of a bond being present between the subunits.
[0053] The embodiment may enable enhanced performance in polymer search by generating a digital representation of the target polymer that includes connectivity probabilities associated with the subunits based on structural information, thereby providing a detailed depiction of the molecular ensemble. This may facilitate robust and reliable monitoring and controlling of polymer synthesis and manufacturing processes by allowing secure comparison with digital representations in a polymer database and potentially contributing to environmental impact reduction through more efficient production management.
[0054] Internal 240427W001
[0055] 11
[0056] In an embodiment, generating the digital representation of the at least a part of the target polymer includes sampling an ensemble of possible configurations of the polymer, in particular based on the connectivity probabilities and / or a molar mass distribution of the target polymer, and determining the digital representation from the ensemble. The embodiment may facilitate stable monitoring and controlling of manufacturing and synthesis processes, as the ensemble-based digital representation is associated with physical production parameters and may contribute to environmental impact reduction.
[0057] In an embodiment, the least a part of the target polymer includes one or more subunits and / or one or more building blocks of the target polymer. The embodiment may enable robust polymer search by representing an ensemble of molecules comprising one or more subunits and / or building blocks associated with the target polymer, thereby facilitating reliable digital comparisons with polymer database entries related to chemical structures and connectivity probabilities. Additionally, this approach may enable secure monitoring and controlling of production, synthesis, and manufacturing processes through stable digital representations related to physical polymer entities, which may contribute to environmental impact reduction.
[0058] In an embodiment, the at least a part of the target polymer includes two of more subunits and the digital representation of the at least a part of the target polymer includes connectivity probabilities between the subunits. This may further improve selecting the polymer by also taking into account an ensemble nature of the target polymer or of at least a part of it.
[0059] In an embodiment, the comparison comprises generating a similarity measure between the digital representation of the at least a part of the target polymer and a respective digital representation of a polymer from the polymer database and wherein the selecting comprises selecting the polymer based on the similarity measure. A similarity measure can be defined as a real-valued function that quantifies the similarity between two objects. A similarity measure can be defined based on a distance metric in a predetermined space. A similarity measure can be defined for a feature representation, a string representation and a graph representation. For example, for a feature representation a distance between the feature vectors of the two digital representations in feature space can be defined as a similarity measure utilizing a predetermined distance metric. For graph representations, similarity measures like graph added distance can be utilized. For a string representation a similarity measure can be defined based on a similarity between different string parts of the string representation. For example, a similarity measure can be defined by a) comparing a string representations of subunits, e.g. building blocks, of respective polymers, preferably, ignoring the bond information in the string representation, b) generating a subunit similarity measures based on the respective comparisons and c) generating a similarity measure for the respective polymers based on the subunit similarity measure, for example, based on predetermined rules. A comparison between the string representations of subunits can for example be based on a comparison between regular SMILES of the subunits of the polymers. For such comparisons respective techniques exist as described for example in the article "A comparative study of SMILES-based compound similarity functions for drug-target interaction prediction.”, Ozturk, H. , et. al., BMC Bioinformatics 17, 128 (2016). Preferably, the comparison can also include the bond positions and
[0060] Internal 240427W001
[0061] 12 connectivity, for instance, by weighting respective subunit similarity scores, into the similarity measure. In an embodiment, the digital representation comprises a string representation and the comparison between the string representations of at least a part of the polymer and a string representation of a polymer of the database is based on a machine learning search model, wherein the machine learning search mode is trained to generate a similarity measure between the respective string representations, based on a historical training data set comprising respectively labeled pairs of string representations. For example, the search model can be based on a Large Language Model, based on the Transformer architecture. The labels of the training data set can be given by a respective expert that manually compares a set of pairs. For example, the search model can then be configured to map a pair of, for example, BigSM ILES representations, onto a similarity score.
[0062] The selecting of the polymers based on the similarity measure can then comprise, for instance, providing a predefined threshold similarity measure. The threshold similarity measure can define a threshold above or below, depending on the application case, which a polymer is selected. Also similarity measure ranges can be provided to select polymers that fall within or without, depending on the application case, the respective similarity measure ranges. This allows for an easy selection of polymers tailored to the respective application case.
[0063] In an embodiment, the at least a part of the target polymer is a part of the target polymer, referred to as substructure of the target polymer. For example, the substructure can refer to a subunit of the polymer. The substructure can comprise at least two polymer building blocks. In this case, it is preferred that the digital representation comprises a graph representation and / or a string representation. The comparison can then be performed based on the string representation and / or the graph representation. The string representation preferably is a SMARTS, SMILES and / or BigSMILES representation.
[0064] In an embodiment, the digital representation is associated with polymer bond connectivities and / or bond positions, wherein a polymer bond connectivity quantifies a possible bond between one or more subunits of the polymer and a bond position a possible position of one or more subunits in the polymer. The possibility of a bond or position can be quantified by a stochastic processing of the chemical possibilities. In a graph representation these information can be provided associated with an edge or node of the graph. In a feature representation these information can be provided as feature or as part of one or more features. As string representation a translation scheme can be used that allows to incorporate such information, for example, a G-BigSM ILES representation. The comparison is then also based on the polymer connectivities and / or bond positions. This allows for a more reliable search taking the stochastic nature of polymers into account.
[0065] In an embodiment, in particular of the polymer database, the digital representation is, per polymer of the plurality of polymers, associated with the structure of the respective polymer and connectivity probabilities between subunits of the respective polymer.
[0066] Internal 240427W001
[0067] 13
[0068] In an embodiment, in particular of the polymer database, the polymer database further comprises digital representations of a plurality of small molecules, the digital representation may, per polymer of the plurality of small molecules, be associated with a chemical structure of the respective small molecule. This may facilitate searching through a more diverse set of molecules, e.g. including polymers and small molecules and hence may improve manufacturing molecules with beneficial properties.
[0069] In an embodiment, the digital representations of the plurality of polymers and the digital representations of the plurality of small molecules may include common digital representations. The common digital representations may have equal dimensionality for a small molecule of the plurality of small molecules and for a polymer of the plurality of polymers. The common digital representation may comprise feature vectors of fixed length, for example molecular fingerprints, where zero-padding may be applied to small molecules to match the dimensionality of polymer fingerprints. This harmonization may enable the database to process, e.g. store, search and / or retrieve, both datasets. The common digital representation may include additional fields for polymer-specific attributes, such as molar mass distribution or connectivity statistics, which may be set to zero for small molecules. This may facilitate merging the polymer and small molecule datasets into a single data set. By aligning the representation format, the polymer database may enable searching, in a generalized way, both chemical domains, leveraging the broader chemical diversity of small molecules while maintaining relevance for polymers. For example, fingerprints may be computed for ensembles of polymer chains and aggregated into an ensemble-level fingerprint, whereas small molecules may be represented by a single fingerprint of identical dimensionality.
[0070] In an embodiment, generating the digital representation per polymer includes determining connectivity probabilities between subunits of the respective polymer based on structure information of the respective polymer, the connectivity probabilities at least relating to a likelihood of a bond being present between the subunits. This may enable robust polymer search performance by generating digital representations comprising connectivity probabilities determined based on structure information, wherein these connectivity probabilities are associated with interactions between polymer subunits and related to the likelihood of bond formation, thereby representing an ensemble of molecules of a polymer.
[0071] In an embodiment, generating the digital representation per polymer includes sampling an ensemble of possible configurations of the respective polymer, and determining the digital representation from the ensemble. The embodiment may enable a robust polymer search by sampling an ensemble of possible polymer configurations to determine a comprehensive digital representation, including examples such as graph, string, or feature representations associated with the polymer's chemical structure.
[0072] The described techniques, aspects and embodiments may enable organizations to transform how they manage and leverage polymer knowledge across industries. By introducing digital representations that capture structural details and stochastic characteristics of polymers, these approaches can support faster and more accurate identification of materials with desired properties. Some embodiments may allow similarity-based searches and substructure queries
[0073] Internal 240427W001
[0074] 14 that go beyond conventional ingredient or property filters, opening opportunities for innovation in sectors such as packaging, automotive, coatings, and specialty chemicals. These capabilities may reduce development cycles by helping teams find existing polymers that meet performance targets, potentially avoiding redundant synthesis and minimizing resource use. In addition, incorporating connectivity probabilities and ensemble-level representations may improve prediction accuracy for complex polymers, which is critical for applications where durability, recyclability, or biodegradability are key differentiators. For markets increasingly driven by sustainability and regulatory compliance, such techniques may support the design of greener products by enabling rapid screening for polymers with favourable environmental profiles. Furthermore, integration with automated synthesis systems may streamline scale- up processes, reducing time-to-market for new formulations. Overall, these approaches may deliver significant business benefits by lowering R&D costs, accelerating innovation, and enhancing competitiveness in high-value polymer markets.
[0075] It shall be understood that the methods as described above, the apparatuses as described above, the systems as described above and the computer program products as described above have similar and / or identical preferred embodiments, in particular, as defined in the dependent claims.
[0076] It shall be understood that a preferred embodiment of the present invention can also be any combination of the dependent claims or above embodiments with the respective independent claim.
[0077] These and other aspects of the present invention will be apparent from and elucidated with reference to the embodiments described hereinafter.
[0078] BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Fig. 1 shows schematically and exemplarily a method for assisting in searching a target polymer,
[0080] Figs. 2 and 3 show schematically and exemplary flowcharts of exemplary embodiments of the invention, and
[0081] Figs. 4 show schematically and exemplarily a block diagram of an exemplary system architecture of a system utilizing the invention.
[0082] DETAILED DESCRIPTION OF THE DRAWINGS
[0083] Fig. 1 shows schematically an exemplarily a method for assisting in a polymer search, in particular, in a search for at least a part of a target polymer. The method comprises providing a polymer database comprising digital representations of a plurality of polymers. In particular, the polymer database can comprise a few thousand polymers optionally together with known characteristics and / or properties of these polymers. The polymers are stored in the database in form of respective digital representations of the polymer. The digital representation can comprise a representation of the polymer based on a respective representation scheme or even based on more than one representation scheme. For example, the digital representation can comprise a string representation like G-
[0084] Internal 240427W001
[0085] 15
[0086] BigSM ILES or polymer chemprop, a graph based representation and / or a feature representation. Generally, the digital representation representing the polymer in the polymer database is indicative of a structure of a respective polymer.
[0087] Further, the method comprises providing a digital representation of at least a part of a target polymer. The at least a part of a target polymer can comprise a complete target polymer or only parts of the target polymer, for instance, specific monomers, polymer building blocks or any other form of subunit of a target polymer. For example, the subunit of the target polymer can be represented as a small molecule or a chemical functional group. The digital representation of a target polymer can be provided as a starting point for searching the polymer database for polymers that are similar to the target polymer. The digital representation of a part of the target polymer, e.g. a subunit, can be provided for searching for polymers in the polymer database that comprise the part of the target polymer. For example, the search allows to search a polymer database for polymers comprising a respective small molecule, for instance, a specific monomer, or other subunit. Searching a database for a target polymer or for parts of a target polymer allows to find respective similar polymers. This can be helpful, for instance, to find polymers that also have similar properties and / or characteristics as the respective target polymer. This allows to find new polymers for similar application for which already a polymer is known. This allows to save resources because the synthesis of a known polymer can be skipped. If the known polymer is not available anymore, its synthesis description can be used for a new synthesis, which avoids potential pitfalls during polymer synthesis. Further, if it is known for a specific chemical functional group that it leads to a specific property of the polymer it can be helpful to search a polymer database for polymers comprising these functional groups as target so that the polymer for an application can be selected from a list of polymers that very likely fulfills the respective property.
[0088] The method further comprises comparing the digital representation of the at least the part of target polymer with the digital representation of polymers provided by the polymer database. In particular, comparing the digital representations, for instance, a graph structure representation, string representation or feature vector allows for a very accurate and computational resource effective search for polymers. Thus, already providing the database with respective digital representations of polymers and searching the database by comparing the respective digital representations of a plurality of polymers of the polymer database with respective digital representations of at least a part of a target polymers allows to search huge polymers databases in a reasonable computational timeframe. Details of the comparison are provided for instance with respect to Fig. 2 and 3. If with respect to a predetermined similarity measure, one or more similar polymers are found in the polymer database at least one of these polymers can be selected. However, all of the polymers that fulfill the respective similarity criteria can be selected. The selected one or more polymers can then be provided, for instance, as a list of polymers. The providing comprises, preferably, that a synthesis specification of the respective selected polymers is provided and that the synthesis specification is configured to allow to synthesize the respective polymer by monitoring and / controlling a respective synthesis system. Moreover, respective control data can be provided comprising the synthesis specification and being configured to control a synthesis system to perform the synthesis of the respective selected polymers.
[0089] Internal 240427W001
[0090] 16
[0091] If the comparison indicates that no polymer is found on the polymers database that fulfills the respective similarity criteria the method can comprise providing a new digital representation of a new part of the target polymer or of a complete new target polymer. For example, a user can be prompted to input a new digital representation of a new part of a target polymer or of a part of a new target polymer. However, the system can also automatically switch to a new digital representation of at least a part of a target polymer, for instance, based on a list of digital representations that are to be searched. Based on the new digital representation a search can be performed.
[0092] It is noted that although in the above example a similarity criteria is utilized that indicates a target polymer is similar to a polymer of the database, in other examples also a criterion can be defined that indicates that a specific part of the target polymer should not be present in the respective polymer that is selected. For example, if it is known that a monomer or other subunit of a polymer leads to an undesired property the comparing of the digital representations can allow to select polymers from the polymer database that do not contain this monomer or subunit in the respective selected polymer.
[0093] In the following some embodiments are described in more detail with respect to Fig. 2 and 3. In the following an example is provided for utilizing graph representations for searching for polymers. An exemplary workflow for this embodiment is illustrated by Fig. 2. In this example first a recipe of a polymer is provided comprising a synthesis specification of the polymer. Structure information can be derived from the synthesis specification provided by the recipe based on chemical bonds. For example, kinetic modelling, known approximations or respective analysis tools can be utilized. For example, BIG-SMILES representations or SMILES representations of polymer repeat units as structure information can be derived. A molecular graph representation can be constructed based on bonding information of polymer provided by the structure information, for instance, from the SMILES representations of polymer repeat units as structure information. Further, the bonds between repeat units and their probability can be utilized for generating the graph representation. Within the polymer graph representation, atoms and bonds can be represented as the nodes and edges of the graph. Both can be further initiated with features, e.g., nuclear charge as node feature and chemical bond order as well as possibility as edge feature. Next, a Graph neural network is provided with the polymer graph representation to generate a feature vector. The Graph neural network can utilize message passing based on the graph representation, wherein in a layer of the neural network features associated with a graph node are transformed based on the features of the node itself and of its neighboring nodes and provided to the next layer. Thus, with each layer of the network information associated with a node are passed through the graph to neighboring nodes. A pooling layer can then generate the feature vector based on the result of the message passing in the previous layers. For example edge-centered messages of outgoing edges can be updated based on incoming edges. Updated atom features can then be obtained by a weighted sum over the features of all incoming edges. The polymer feature vector is then obtained by aggregating all final atomic features.
[0094] An alternative technique is the use of representation learning tools in the context of machine learning algorithms. The representation learning can be based on a training data set comprising an unlabeled dataset of polymers. It is possible to generate such unlabeled datasets for polymers in a first step based on generating artificial synthesis
[0095] Internal 240427W001
[0096] 17 specifications of polymer. Such a generated training dataset can comprise millions of unlabeled polymer structures. The unlabeled training datasets include preferably, the connectivity information of polymer repeat units, for example, by providing SMILES representations. Optionally, additional data can be added as part of structure information or as part of a digital representation such as the amount of polymer building blocks, chemical bond statistics between polymer building blocks, polymer block assignment, average molecular weight, molecular weight distribution, information on distribution shifts of polymer building blocks, reactivity ratios between polymerizable groups, branching degree, mixtures of ingredients, polymer blends, polymer tacticity, post-polymerization modifications, crosslinking densities, process data from polymerization or polymer type information. Part of the additional data might be derived from kinetic models, e.g., chemical bond statistics between repeat units, molecular weight distribution.
[0097] In a pretraining step, the unlabeled training dataset can be used for representation learning. The aim is to gain a feature representation of polymers. Such a representation should position polymers in a feature representation space in a way that similar polymers are close to and dissimilar ones are far away from each other. "Similar” in this context can be evaluated with respect to a predefined similarity measure. In this pretraining step, a transformer neural network, like a large language model, can be used. A large language model can receive a polymer in a letter-by-letter format. For example a SMILES representation of the polymer building blocks can be used as input and the feature model can learn to generate the feature representation of the polymer. This is analogous to large language models that learn to construct subsequent words of a sentence using not letters, but previous words in the sentence. The architecture of such models has proven to be well generalizable so that the inventors have found that these models can also be applied to polymers.
[0098] The polymer feature representation can be a vector of numbers and can be used for similarity searches in a provided polymer data set. The provided polymer data set can be provided as a polymer database comprising feature representations for a plurality of polymers and optionally additional information, like polymer properties, synthesis specification, list of ingredients, etc., associated with the respective polymer. Polymers with a similar chemical structure, are expressed in similar graph representations, which are aggregated in similar polymer feature representations. Consequently, the similarity of two polymers can be expressed as the distance of their feature representations. Thus, the search can be performed by comparing the feature representations, in this example, feature vectors, of the polymers from the polymer database with the feature representation of the provided polymer by determining a distance between the respective feature representations in feature space as similarity measure. The similarity measure can be a cosine or a Tanimoto similarity between the feature representations in a vector space. However, also other similarity measures can be utilized. Utilizing a feature vector derived from a graph representation allows to condense the most relevant aspects of the polymer into a respective digital representation. The most relevant aspects can be based on the intended application or can be based generally on the intrinsic molecular properties. Further, a feature representation allows for a much faster search than, for instance, a search based on a graph representation.
[0099] Internal 240427W001
[0100] 18
[0101] A similarity measure threshold can then be provided that defines a threshold distance. Polymers from the polymer data set that in the comparison have a distance below the threshold can then be selected as similar polymers. The selected polymers can then be provided as a list to a user together with the additionally stored information. Preferably a respective synthesis specification is provided that is configured to control and / or monitor a respective production of the respective selected polymer.
[0102] Fig. 3 shows a further exemplary embodiment. In this embodiment, only a part of a target polymer, for instance a subunit of the target polymer, is provided. This refers to a substructure search. The respective subunit can be provided as a SMART pattern representation to provide connectivity structure information on the part of the subunit. Further, a SMILES or Big-Smiles representation can be provided for the polymers of the polymer data set. These representation provide information on bond probabilities and molar mass distributions. The SMART and respective SMILES or Big-SMILES representations can then be transformed into graph representations. A respective graph search can then be performed in the polymer data set, for instance, by checked if a certain connectivity pattern of a subunit exists in a graph representation of a polymer from the data set. For example, a RDkit toolkit can be utilized for this search. Preferably, besides the existence of a subunit in a polymer of the data set, the procedure can provide information on how often the subunit is included in a polymer on average. Polymers of the data set can then be select based on if, and optionally how often, they contain the subunit. For example, a threshold or a range can be provided for the amount of a subunit in a polymer. The selected polymers can then be provided as list to a user together with the additionally stored information. Preferably a respective synthesis specification is provided that is configured to control and / or monitor a respective production of the respective selected polymer.
[0103] Alternatively to generating a graph representation from a provided string representation, the string representation can also be directly utilized in the search for substructure and complete polymer searches. For this, a similarity score for a similarity between two given polymers or parts of a polymer can be determined. This can be done, for instance, by comparing a subunit string representations between both polymers or parts of the polymer, optionally ignoring the bond symbols. This yields merely the comparison between regular SMILES of the repeating units, for which there are existing techniques. Based on this several similarity scores for pairs of repeating units can be generated. These scores can be used to derive a similarity score for the two polymers or parts of the polymer in various ways. For example, respective rules can be defined. This approach can be extended by including the bond positions and bond probabilities into the similarity measure, for example, be weighting the contribution of subunit similarity scores to the overall similarity score based on the bond positions and the connectivity. Further, a similarity measure can be determined by a trained Language Model, based on the Transformer architecture. The training can be based on a labelled set of pairs of Big-SMILES. The labels can be given by a user that manually compares a set of pairs that is sufficient to train the model. This model then learns to map a pair of BigSMILES to a similarity score.
[0104] The above invention can be applied, for instance, in situation where a user has identified a polymer with favorable application properties like a fast biodegradability. Then, the above described similarity search can be used to identify known polymers with a similar chemical structure, which hopefully are biodegradable as well. Another application of
[0105] Internal 240427W001
[0106] 19 similarity searches is that a polymer data set can be checked, during synthesis planning, if a similar polymer was synthesized before. In best case, the polymer is still physically stored, and the polymer synthesis can be skipped. In a second-best case, an old synthesis description can still be used, if that is saved in the dataset. Then, it can be learned how the old polymer was synthesized and which pitfalls can be avoid during synthesis, e.g., adjusted pH value or used catalyst to prevent side product formation.
[0107] A further example of the use of the described search method is the search of polymers according to known structure- property-relations. For example, if it is known that polymeric dye transfer inhibitors must contain certain nitrogenbased substructures, existing datasets for polymers can be searched for polymers containing such a nitrogen-based substructure. Another example would be that it is known that polymers should contain certain functional groups to achieve a certain application property. This could be due to a favorable interaction of these functional groups to certain chemicals or particles. Then, it could be searched for polymers containing such functional groups and sort the polymers ascending regarding the amount of these functional groups within the polymer. A third example is that it is known that polyesters are particularly stable against hydrolysis, if a certain functional group is in direct vicinity of the ester bond. Then, hydrolysis-stable polyesters could be identified by a substructure search for this functional group in vicinity of ester bonds.
[0108] Fig. 4 illustrates a block diagram of an exemplarily system architecture of an automated laboratory system 1000 for synthesizing a polymer with a laboratory equipment control device 1102, a network 1150 and the synthesis specification, i.e. recipe, module 1100 / 1110, and a client device 1108. The automated laboratory system includes a laboratory equipment control device layer 1152 as part of the laboratory equipment control device 1102 as well as a synthesis specification module layer 1154 associated with the synthesis specification module and a remote control or client layer 1156 associated with the client device 1108. The laboratory equipment control device layer can be split into several hierarchical layers: the hardware, the middleware and the interface layer. The hardware layer relates to hardware resources such as sensors and actuators, in particular for controlling a synthesis of a polymer. The middleware relates to any of the known middleware for laboratory or plant synthesis operations. One example is LABS / QM, providing different abstractions to hardware, network and operating system such as low-level device control and message passing. The communication layer relates to communication protocols, wherein the protocol may be REST, which may be implemented over different transport protocols (i.e. UDP, TCP, Telemetry) that allow the exchange of messages between the laboratory equipment control device and laboratory equipment devices. Such software architecture allows to control and monitor laboratory equipment without having to interact with the hardware.
[0109] The synthesis specification module layer 1154 may include: a mass storage layer, the computing layer, the interface layer. The storage layer is configured to provide mass storage for the data-driven determination model for providing a synthesis specification of a polymer based on a technical application property, as described in detail above. In particular, the functions performed by the apparatus, as described above, can be provided as program code means stored on the mass storage. Furthermore, digital representations for a plurality of polymers can be stored as polymer
[0110] Internal 240427W001
[0111] 20 database in the mass storage. Such data may be stored in structured databases such as SQL databases or in a distributed file system such as HDFS, NoSQL databases such as HBase, MongoDB. The computing layer may include an application layer that allows to customize the functionalities provided by standard cloud services to perform computing processes based on target properties. Such functionalities can include providing a digital representation of a target polymer, searching the polymer data base based on the digital representation of the target polymer, and providing a synthesis specification of a found polymer as control data, i.e. control signal, to the laboratory equipment control device.
[0112] The interface layer may implement web services, network interfaces as UDP or TCP or Websocket interfaces. For communication with the laboratory equipment control device a REST API is implemented.
[0113] The client layer 1156 provides interfaces for end-users. For end-users, the client layer 1156 can run client side Web applications, which provide interfaces to the synthesis specification module layer 1154 or the laboratory equipment control device layer 1152. Users may be provided with a Ul for selecting a digital representation of a polymer. In other examples, the users may be provided with a Ul for selecting more than one digital representation. The applications may be configured for users to monitor and control the laboratory equipment control device and the operation remotely. In other examples, the client device layer and the synthesis specification module layer may be integrated into one device. The alternatives described here are only for illustration purposes and should not be considered limiting.
[0114] Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims.
[0115] For the processes and methods disclosed herein, the operations performed in the processes and methods may be implemented in differing order. Furthermore, the outlined operations are only provided as examples, and some of the operations may be optional, combined into fewer steps and operations, supplemented with further operations, or expanded into additional operations without detracting from the essence of the disclosed embodiments.
[0116] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality.
[0117] A single unit or device may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
[0118] Procedures like the providing of the digital representation, the comparing of the digital representation, etc. performed by one or several units or devices can be performed by any other number of units or devices. These procedures can be implemented as program code means of a computer program and / or as dedicated hardware.
[0119] Internal 240427W001
[0120] 21
[0121] A computer program product may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.
[0122] Any units described herein may be processing units that are part of a classical computing system. Processing units may include a general-purpose processor and may also include a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or any other specialized circuit. Any memory may be a physical system memory, which may be volatile, non-volatile, or some combination of the two. The term "memory” may include any computer-readable storage media such as a non-volatile mass storage. If the computing system is distributed, the processing and / or memory capability may be distributed as well. The computing system may include multiple structures as "executable components”. The term "executable component” is a structure well understood in the field of computing as being a structure that can be software, hardware, or a combination thereof. For instance, when implemented in software, one of ordinary skill in the art would understand that the structure of an executable component may include software objects, routines, methods, and so forth, that may be executed on the computing system. This may include both an executable component in the heap of a computing system, or on computer- readable storage media. The structure of the executable component may exist on a computer-readable medium such that, when interpreted by one or more processors of a computing system, e.g., by a processor thread, the computing system is caused to perform a function. Such structure may be computer readable directly by the processors, for instance, as is the case if the executable component were binary, or it may be structured to be interpretable and / or compiled, for instance, whether in a single stage or in multiple stages, so as to generate such binary that is directly interpretable by the processors. In other instances, structures may be hard coded or hard wired logic gates, that are implemented exclusively or near-exclusively in hardware, such as within a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or any other specialized circuit. Accordingly, the term "executable component” is a term for a structure that is well understood by those of ordinary skill in the art of computing, whether implemented in software, hardware, or a combination. Any embodiments herein are described with reference to acts that are performed by one or more processing units of the computing system. If such acts are implemented in software, one or more processors direct the operation of the computing system in response to having executed computer-executable instructions that constitute an executable component. Computing system may also contain communication channels that allow the computing system to communicate with other computing systems over, for example, network. A "network” is defined as one or more data links that enable the transport of electronic data between computing systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection, for example, either hardwired, wireless, or a combination of hardwired or wireless, to a computing system, the computing system properly views the connection as a transmission medium. Transmission media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general-purpose or special-purpose computing system or combinations. While not all computing
[0123] Internal 240427W001
[0124] 22 systems require a user interface, in some embodiments, the computing system includes a user interface system for use in interfacing with a user. User interfaces act as input or output mechanism to users for instance via displays.
[0125] Those skilled in the art will appreciate that at least parts of the invention may be practiced in network computing environments with many types of computing system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems, microprocessorbased or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, pagers, routers, switches, datacenters, wearables, such as glasses, and the like. The invention may also be practiced in distributed system environments where local and remote computing system, which are linked, for example, either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links, through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.
[0126] Those skilled in the art will also appreciate that at least parts of the invention may be practiced in a cloud computing environment. Cloud computing environments may be distributed, although this is not required. When distributed, cloud computing environments may be distributed internationally within an organization and / or have components possessed across multiple organizations. In this description and the following claims, "cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources, e.g., networks, servers, storage, applications, and services. The definition of "cloud computing” is not limited to any of the other numerous advantages that can be obtained from such a model when deployed. The computing systems of the figures include various components or functional blocks that may implement the various embodiments disclosed herein as explained. The various components or functional blocks may be implemented on a local computing system or may be implemented on a distributed computing system that includes elements resident in the cloud or that implement aspects of cloud computing. The various components or functional blocks may be implemented as software, hardware, or a combination of software and hardware. The computing systems shown in the figures may include more or less than the components illustrated in the figures and some of the components may be combined as circumstances warrant.
[0127] Any reference signs in the claims should not be construed as limiting the scope
[0128] Internal
Claims
240427W00123CLAIMS1 . A method, in particular, a computer implemented method, for selecting a polymer from a polymer database, wherein the method comprises: providing a polymer database comprising digital representations of a plurality of polymers, wherein a digital representation is associated with a chemical structure of a respective polymer, providing a digital representation of at least a part of a target polymer, wherein the digital representation is associated with a chemical structure of the at least a part of the target polymer, comparing the digital representation of the at least a part of the target polymer with the digital representations of polymers provided by the polymer database, wherein the compared digital representation comprise a graph representation, a string representation and / or a feature representation, selecting at least one polymer from the polymer database based on the comparison, and providing the selected at least one polymer.
2. The method according to claim 1, wherein, per polymer of the plurality of polymers, the respective digital representation is associated with connectivity probabilities between subunits of the respective polymer.
3. The method according to claim 1 or 2, wherein the method comprises receiving structure information associated with a chemical structure of the at least a part of the polymer and generating the digital representation of the at least a part of the target polymer based on the structure information.
4. The method according to claim 3, wherein generating the digital representation of the at least a part of the target polymer includes determining connectivity probabilities between subunits of the target polymer based on the structure information, the connectivity probabilities at least relating to a likelihood of a bond being present between the subunits.
5. The method according to claim 3 or 4, wherein generating the digital representation of the at least a part of the target polymer includes sampling an ensemble of possible configurations of the polymer, in particular based on the connectivity probabilities and / or a molar mass distribution of the target polymer, and determining the digital representation from the ensemble.
6. The method according to any of claims 3 to 5, wherein the structure information can be provided as SMILES, SMARTS, BIG-SMILES Big SMARTS, G-BigSMILES, PSMILES, CurlySMILES representation of the at least a part of the target polymer.
7. The method according to any of claims 3 to 6, wherein the digital representation is associated with polymer bond connectivities and / or bond positions, wherein a polymer bond connectivity quantifies a possible bondInternal240427W00124 between one or more subunits of the polymer and a bond position a possible position of subunits in the polymer.
8. The method according to any of the claims 3 to 7, wherein the receiving of the structure information comprises deriving the structure information from a synthesis specification of the polymer corresponding to the at least a part of the target polymer.
9. The method according to any of claims 3 to 8, wherein the method further comprises deriving a graph structure representation for the at least a part of the target polymer as digital representation based on the structure information and wherein the digital representation of the plurality of polymers of the polymer database comprises a graph structure representation.
10. The method according to any of claims 3 to 9, wherein the method further comprises deriving a feature vector representation for the at least a part of the target polymer based on the structure information and wherein the digital representation of the plurality of polymers on the polymer database comprises a respective feature vector.11 . The method according to any of claims 3 to 10, wherein the method comprises generating a feature vector as digital representation for the at least a part of the target polymer based on the structure information utilizing a machine learning based feature model configured to generate a feature vector based on the structure information.
12. The method according to claim 11, wherein the machine learning based feature model comprises a neural network algorithm, in particular, a transformer neural network algorithm.
13. The method according to any of the preceding claims, wherein the at least a part of the target polymer includes one or more subunits and / or one or more building blocks of the target polymer.
14. The method according to any of the preceding claims, wherein the least a part of the target polymer includes two of more subunits and the digital representation of the at least a part of the target polymer includes connectivity probabilities between the subunits.
15. The method according to any of the preceding claims, wherein the comparison comprises generating a similarity measure between the digital representation of the at least a part of the target polymer and a respective digital representation of a polymer from the polymer database and wherein the selecting comprises selecting the polymer based on the similarity measure.Internal240427W0012516. The method according to claim 15, wherein the digital representation comprises a string representation and wherein the comparison between the string representations of at least a part of the polymer and a string representation of a polymer of the polymer database is based on a machine learning based search model, wherein the search model is trained to generate a similarity measure between the respective string representations, based on a historical training data set comprising respectively labeled pairs of string representations.
17. The method according to any of the preceding claims, wherein providing the selected at least one polymer comprises providing a synthesis specification of the at least one polymer configured for controlling and / or monitoring a synthesis of the at least one polymer.
18. A polymer database configured to be usable in any of the preceding claims, wherein the polymer database comprises digital representations of a plurality of polymers, wherein a digital representation is associated with a structure of a respective polymer, wherein the digital representation comprises a graph representation, a string representation and / or a feature representation.
19. The polymer database according to claim 18, wherein the digital representation is, per polymer of the plurality of polymers, associated with the structure of the respective polymer and connectivity probabilities between subunits of the respective polymer.
20. A computer-implemented method for providing a polymer database, in particular a polymer database according to claim 16, the method comprising: providing a plurality of structure information associated with a plurality of chemical structures, wherein a respective chemical structure of the plurality of chemical structures is associated with a respective polymer of a plurality of polymers; generating, per polymer of the plurality of polymers, a digital representation based on the respective structure information associated with the respective polymer, wherein the digital representation is associated with the chemical structure of the respective polymer, and includes at least one of a graph representation, a string representation, and / or a feature representation; storing, per polymer, the respective digital representation in the polymer database; and providing the polymer database.21 . The method according to claim 20, wherein generating the digital representation per polymer includes determining connectivity probabilities between subunits of the respective polymer based on structure information of the respective polymer, the connectivity probabilities at least relating to a likelihood of a bond being present between the subunits.Internal240427W0012622. The method according to any of claims 20 or 21, wherein generating the digital representation per polymer includes sampling an ensemble of possible configurations of the respective polymer, and determining the digital representation from the ensemble.
23. An apparatus for selecting a polymer from a polymer database, wherein the apparatus comprises computer means configured for performing method steps comprising: providing a polymer database structure comprising digital representations of a plurality of polymers, wherein a digital representation is associated with a chemical structure of a respective polymer, providing a digital representation of at least a part of a target polymer, wherein the digital representation is associated with a chemical structure of the at least a part of the target polymer, comparing the digital representation of the at least a part of the target polymer with the digital representations of polymers provided by the polymer database, wherein the digital representation comprises a graph representation, a string representation and / or a feature representation, selecting at least one polymer from the polymer database based on the comparison, and providing the selected at least one polymer.
24. A computer program product for selecting a polymer from a polymer database, wherein the computer program product comprises program code means for causing the apparatus of claim 23 to execute the method according to any of claims 1 to 17.Internal