Molecular graph structure based modeling for predicting properties of polymers using local clusters
Patent Information
- Application Number
- CN202480086809.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-01
- Publication Date
- 2026-09-11
AI Technical Summary
这些方法通常需要大量的人力和物力,从而导致通过试误法开发新材料的高成本和降低的效率
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This disclosure relates to molecular graph-based modeling to predict polymer properties by utilizing local oligomerization structures. Such techniques are particularly useful for predicting the density and / or dielectric constant of a given polymer to alter the behavior of a given polymer formulation. Background Technology
[0002] Polymer materials can be used in a variety of products. One such material used in flexible organic light-emitting diodes (OLEDs) and other products is low dielectric constant (Dk) thin-film encapsulation (TFE) material. Using TFE materials with low Dk values offers several advantages over other materials, including reduced parasitic capacitance, improved response time, reduced power consumption, and enhanced insulation properties when used with flexible OLEDs. Therefore, the use of TFEs with low Dk values in OLED panels is highly desirable. However, the Dk values of TFE materials currently used in OLED touch panels are greater than 3.0, limiting customers' ability to develop new high-performance flexible displays. For next-generation flexible OLED displays, TFE materials with even lower Dk values (e.g., less than 2.4) are required.
[0003] Currently, the study of polymer properties requires extensive experimentation, involving different polymer monomers and various polymerization conditions. These methods typically demand significant human and material resources, leading to high costs and reduced efficiency in developing new materials through trial and error. The cost of traditional high-precision theoretical calculations for polymer systems remains prohibitive due to the exponential increase in computational complexity. Summary of the Invention
[0004] This disclosure relates to improvements in predicting the properties of polymer samples using machine learning techniques. Predictions can be based on a set of descriptors that have been categorized based on local cluster properties of multiple polymers and / or molecular representations from a graph neural network. The descriptors can be computed via a DFT (density functional theory) on the local oligomeric structures (i.e., local cluster properties of multiple polymers, such as dielectric constants, etc.). Molecular representations can be generated by inputting a molecular graph into a pre-trained graph neural network. Compared to previous systems or methods, the combination of descriptors on oligomeric structures and representations on molecular structures can predict polymer properties with significantly fewer computational resources, potentially outperforming these systems and methods.
[0005] The foregoing description of this invention is not intended to describe every disclosed embodiment or every implementation thereof. The following description illustrates exemplary embodiments in more detail. Throughout this application, guidance is provided by a list of examples, which may be used in various combinations. In each case, the enumerated list serves only as a representative group and should not be construed as an exclusive list. Attached Figure Description
[0006] Figure 1 An example flowchart is shown for a method of using molecular graph structure-based modeling to predict polymer properties by leveraging local structure.
[0007] Figure 2 An example flowchart is shown for a method of using molecular graph structure-based modeling to predict polymer properties by leveraging local structure.
[0008] Figure 3 An example flowchart is shown for a method of using molecular graph structure-based modeling to predict polymer properties by leveraging local structure.
[0009] Figure 4 An example of a machine-readable medium is shown for modeling based on molecular diagram structures to utilize local structures to predict polymer properties.
[0010] Figure 5 An example of a device for modeling molecular graph structures to use local structures to predict polymer properties is illustrated. Detailed Implementation
[0011] This disclosure relates to methods and apparatus for predicting polymer properties based on molecular graph structure modeling, which can utilize machine learning models to predict the dielectric constant of a given polymer. The machine learning model can be a function or formula used to identify patterns in data. A machine learning module can be multiple machine learning models used together to identify patterns in data. In a specific example, the machine learning module can be organized as a neural network. The neural network can include a set of instructions executable to identify patterns in data. Some neural networks can be used to identify underlying relationships in a set of data in a manner that mimics the operation of the human brain. The neural network can adapt to changing or altered inputs, allowing it to generate the best possible results without redesigning the output criteria.
[0012] A neural network can include multiple neurons, which can be represented by one or more formulas or functions. In the context of a neural network, a neuron can receive a certain number of numbers or vectors as input and produce an output based on the properties of the neural network. For example, a neuron can receive X... k There are *k* inputs, where *k* corresponds to the index of the input. For each input, the neuron can assign a weight vector W. kThe weight vector is assigned to the input. In some implementations, the weight vector (e.g., weight values, etc.) can distinguish neurons in a neural network from one or more different neurons in that network. In some neural networks, the corresponding input vector can be multiplied by the corresponding weight vector to produce a value, as shown in Equation 1, which illustrates an example of a linear combination of the input vector and the weight vector.
[0013]
[0014] In some neural networks, nonlinear functions (e.g., activation functions) can be applied to the value f(x1, x2) produced by Equation 1. An example of a nonlinear function that can be applied to the value produced by Equation 1 is the Rectified Linear Unit (ReLU) function. Applying the ReLU function shown in Equation 2 produces the value input to the function if the value input to the function is greater than zero, or produces zero if the value input to the function is less than zero. The ReLU function is used here only as an illustrative example of an activation function and is not intended to be restrictive. Other non-restrictive examples of activation functions that can be applied in the context of neural networks can include sigmoid functions, binary step functions, linear activation functions, hyperbolic functions, leaky ReLU functions, parameterized ReLU functions, softmax functions, and / or swish functions, etc.
[0015]
[0016] During the training of a neural network, the weight vector can be changed to "tune" the network. In at least one example, the neural network can be initialized with random weights. Over time, the weights can be adjusted to improve the accuracy of the neural network. This can result in a neural network with high accuracy over time. This disclosure utilizes machine learning, such as neural networks, to predict product properties by modeling input data. In these embodiments, the weights can be tuned based on a number of factors. For example, the weights can be tuned using multiple local clusters (e.g., two-body structures, four-body structures) and corresponding density values to determine the dielectric constant of a particular polymer. In this example, the machine learning model can use monomer data and the output values of a graph neural network to determine the unknown dielectric constant of a particular polymer based on the identified local clusters of that polymer.
[0017] This disclosure relates to neural networks that combine property calculations based on microstructure properties (e.g., local cluster properties, etc.) with molecular representations from graph neural networks. In some embodiments, monomer structure files of polymers of interest (e.g., designated polymers, etc.) can be obtained and converted into oligomeric structures or local clusters. In some embodiments, an oligomeric structure can refer to a complex formed from several monomer units. In some embodiments, oligomeric structures can be synthesized through controlled chemical reactions.
[0018] To mimic the structure in actual polymers, methyl groups or other auxiliary groups are added to the terminal bonds of the repeating units. As used herein, terminal bonds of repeating units (e.g., terminal linkages, etc.) refer to chemical bonds or groups present at the ends of polymer chains that connect the last repeating unit of the polymer to the rest of the molecule or to some other entity. These terminal bonds can affect the properties of the polymer, including but not limited to: stability, reactivity, and / or how the polymer interacts with other substances.
[0019] For example, the microstructure of a polymer can be approximated by local clusters (e.g., dimeric oligomers, tetrameric oligomers, etc.), and supplementary groups (such as methyl groups) can be added to the terminal bonds. For dimeric and tetrameric molecules within local clusters, the free energies of local clusters of polymers (e.g., monomers, dimeric molecules, and tetrameric molecules) and the properties to be predicted can be simulated and calculated using quantum chemical methods (such as parametric method 6 (PM6)) or density function methods (such as Becke's three-parameter exchange function (B3LYP) combined with Lee, Yang, and Parr correlation functions).
[0020] Machine learning models (e.g., trained neural networks for predicting polymer properties) can utilize calculated density and / or dielectric constant information, along with graph structure information of monomer molecules previously obtained from graph neural networks, as combined inputs. This combined input is used to predict specific properties of a target polymer. For example, the combined input can be used to predict the dielectric constant of a specified polymer comprising a specified number of specific local clusters.
[0021] As used herein, unless the content expressly specifies otherwise, the singular forms “a / an” and “the” include both singular and plural indicators. Furthermore, throughout this application, the word “may” is used in a permissible sense (e.g., possibly, able to) rather than in a mandatory sense (e.g., must). The term “comprising” and its derivatives mean “including but not limited to”.
[0022] As will be understood, elements shown in the various embodiments herein may be added, interchanged, and / or eliminated to provide multiple additional embodiments of this disclosure. Furthermore, as will be understood, the proportions and relative scales of the elements provided in the figures are intended to illustrate certain embodiments of the invention and should not be construed as limiting.
[0023] Figure 1An exemplary flowchart illustrates a method 100 for modeling based on molecular graph structures to predict polymer properties by leveraging local structures (e.g., local oligomerization structures, etc.). Method 100 can be performed by a computing device as described herein. Method 100 can be used to train and / or utilize neural networks or other types of machine learning models to determine polymer properties by utilizing possible local structures.
[0024] In some embodiments, method 100 can be used to determine the density and / or dielectric constant (Dk) of the polymer. As described herein, the dielectric constant is a measure of a material's ability to store electrical energy in an electric field. It is a dimensionless quantity representing the ratio of the material's permittivity to the permittivity of vacuum. The higher the dielectric constant of a material, the more charge it can store. As described herein, polymers with relatively low dielectric constants can be beneficial to specific devices by reducing parasitic capacitance, improving response time, reducing power consumption, and enhancing insulation properties. For example, polymers with relatively low dielectric constants can possess properties useful when used with flexible OLEDs.
[0025] Method 100 may include generating a molecular diagram for each of the multiple polymers using a Simplified Molecular Linear Input Specification (SMILES) representation 102. SMILES refers to a symbolic system used to represent chemical structures in a concise and human-readable format. SMILES is designed to be both machine-readable and relatively easy to write and understand. SMILES uses Simple American Standard Code for Information Interchange (ASCII) characters to represent atoms, bonds, and atomic connectivity within a molecule.
[0026] Method 100 can utilize the SMILES designation to identify complementary monomers 104 of various polymers. As used herein, complementary monomers refer to monomer pairs having chemical structures or properties that enable them to react together in a specific manner to form a polymer. These monomer pairs are typically strategically selected to produce polymer chains with specific properties or functions. For example, complementary monomers can have properties or functions such as: polycondensation, addition polymerization, copolymerization, and / or crosslinking.
[0027] In some embodiments, method 100 may classify complementary monomer 104 into multiple categories. For example, complementary monomer 104 may be classified as two-body local clusters 106-1 and / or tetra-body local clusters 106-2. In other embodiments, a greater number of categories may exist. For example, method 100 may include additional categories such as monomer clusters. As used herein, a cluster or structure of polymer molecules may refer to groups of polymer molecules that interact or associate in a particular manner. The size and structure of these clusters may vary, and they play a role in determining the physical properties of polymer materials.
[0028] Dimeric local clusters 106-1 can refer to polymer dimers. In this case, two polymer chains associate with each other through various types of interactions, such as van der Waals forces, hydrogen bonding, or hydrophobic interactions. This dimer association can occur between two identical polymer chains or between two different polymer chains. Dimeric association can affect the properties of the polymer, including its solubility, viscosity, and mechanical behavior. Tetrameric local clusters 106-2 can refer to tetrameric associations of polymer molecules. In this arrangement, four polymer chains aggregate together through various interactions.
[0029] Tetrameric association can occur in more complex polymer systems and can involve multiple intermolecular forces and entanglements. These clusters can be important for understanding the behavior of polymers in solution, melt, or solid configurations. In the case of associative polymers (such as certain types of polyelectrolytes), four polymer chains may form tetrameric associations through a combination of electrostatic interactions, van der Waals forces, and hydrogen bonding. Such associations can influence the polymer's behavior in solution and its response to changes in pH and / or ionic strength.
[0030] In some implementations, conformational sampling via molecular dynamics (MD) simulations can be used to model two-body local clusters 106-1 and / or four-body local clusters 106-2. As used herein, conformational sampling via MD simulations refers to computational methods used to study the physical movement and conformational changes of molecules, particularly large biomolecules such as proteins, nucleic acids, and lipids. In some implementations, MD simulations may include dynamic simulations to calculate the time-dependent behavior of molecular systems (e.g., molecular interactions evolving over time under the influence of physical laws, etc.). Additionally, MD simulations may include force field interactions within molecules and force field interactions between molecules and their environment. In these simulations, the interactions are governed by mathematical functions called force fields. Force fields can refer to bonds, angles, dihedral angles, van der Waals forces, and / or electrostatic forces.
[0031] In some implementations, MD simulations may include predicting molecular movement using Newton's laws of motion. Additionally, conformational sampling refers to the process of exploring the various spatial arrangements that a particular molecule can adapt to. In this way, conformational sampling can utilize MD simulations to explore various spatial relationships between two-body local clusters 106-1 and / or four-body local clusters 106-2.
[0032] Two-body local clusters 106-1 can be analyzed to determine two-body properties 108-1 (e.g., density, dielectric constant, electronic structure, etc.), and four-body local clusters 106-2 can be analyzed to determine four-body properties 108-2 (e.g., density, dielectric constant, electronic structure, etc.). In some embodiments, two-body properties 108-1 and four-body properties 108-2 can be based on specific properties to be predicted by method 100. For example, two-body properties 108-1 and four-body properties 108-2 can be density values and / or dielectric constant values associated with two-body local clusters 106-1 and four-body local clusters 106-2, respectively.
[0033] In some embodiments, the two-body property 108-1 may include a calculated two-body density value 110-1. Additionally, the tetrabody property 108-2 may include a calculated tetrabody density value 110-2. In some embodiments, the two-body density value 110-1 and the tetrabody density value 110-2 may be provided as input values to a first machine learning model 114. Alternatively, graph neural network data 112 may be provided as input to the first machine learning model 114. As described herein, the graph neural network data 112 may be data associated with a graph neural network used to form monomer molecules of a specified polymer. As used herein, a graph neural network (GNN) is a type of neural network architecture designed to process and analyze data represented as a graph. GNNs can be used to learn and analyze the structure and chemical properties of molecules, where the atoms of the molecule and their connections form a graph-like structure. The output of a GNN applied to a monomer molecule can provide different data associated with the properties and / or structure of the monomer molecule. For example, the output or GNN data 112 may include, but is not limited to: node embeddings, graph embeddings, property predictions, chemical activity, visualizations, quantitative descriptors, graph-based analysis, and / or polymer interactions.
[0034] As described herein, input data can be provided to the GNN. In these embodiments, the input data provided to the GNN includes monomer molecules identified after the substitution of functional groups. As described herein, functional groups can be positioned at terminal junctions to generate different stable monomer molecules.
[0035] As described herein, GNN data 112 can include node embeddings, which can refer to vector representations of each atom in a monomer molecule. These embeddings encode information about the local environment of each atom, including its connectivity, neighboring atoms, and chemical properties. Node embeddings can be used for various downstream tasks, such as property prediction, classification, or clustering of monomer molecules. As described herein, GNN data 112 can also include graph embeddings, which can refer to graph-level embeddings that summarize the entire monomer molecule. Such embeddings capture the overall structural and chemical properties of the monomer molecule. Such embeddings can be used for tasks such as polymer classification, similarity comparison, or property prediction at the molecular level.
[0036] As described herein, GNN data 112 can include property predictions of properties such as mechanical properties (e.g., tensile strength, elasticity), thermal properties (e.g., melting point, glass transition temperature), and / or chemical properties (e.g., reactivity, solubility). Predicted properties can be scalar values or multidimensional vectors. As described herein, GNN data 112 can include the chemical reactivity or activity of specific functional groups or bonds within a polymer. This information can be valuable for understanding how polymers may react in different chemical environments or for designing polymers with specific reactivity profiles. As described herein, GNN data 112 can include visualizations, which can refer to visualizing the structure and properties of polymers. GNN data 112 can produce two-dimensional (2D) or three-dimensional (3D) representations of polymer molecules, highlighting key structural features or regions of interest. Visualization can help researchers understand the behavior of polymers.
[0037] As described herein, GNN data112 can include quantitative descriptors, such as molecular fingerprints or topological indexes. These descriptors can be used for similarity searches, virtual screening, or other cheminformatics tasks. GNN data112 can analyze the graph structure of polymers to identify important substructures (e.g., functional groups, repeating units) or detect anomalies or defects in polymer chains. As described herein, GNN data112 can predict interactions between polymer molecules and other molecules, such as solvents, additives, or nanoparticles. This information is crucial for studying the behavior of polymers in a variety of applications.
[0038] The combined input of the two-body density value 110-1, the four-body density value 110-2, and the GNN data 112 can be provided to a first machine learning model 114. The first machine learning model 114 is trained to determine the predicted density 116 of a given polymer based on those inputs. In some embodiments, the two-body dielectric constant data 118-1 can be calculated from the two-body property 108-1. Similarly, the four-body dielectric constant data 118-2 can be calculated from the four-body property 108-2. The predicted density 116 of the given polymer can be used as a combined input of the two-body dielectric constant data 118-1 and the four-body dielectric constant data 118-2 to a second machine learning model 122. The second machine learning model 122 is trained to determine the output dielectric constant 124 of the given polymer based on those inputs. In this way, method 100 can be used to calculate the dielectric constant of large polymer molecules with relatively fewer computational resources compared to previous systems and methods.
[0039] Figure 2 An exemplary flowchart illustrates a method 220 for modeling based on molecular graph structures to predict polymer properties using local structures (e.g., local oligomerization structures, etc.). In some examples, method 220 may be performed by a computing device as described herein. Method 220 may be used to train neural networks or other types of machine learning models to determine polymer properties of a given or desired polymer based on local structure data and GNN output data. In some embodiments, the trained neural network may be used to reverse engineer polymers that, when used with flexible OLEDs, can reduce parasitic capacitance, improve response time, reduce power consumption, and enhance insulation properties.
[0040] At step 244, method 220 may include generating a 3D molecular structure from the SMILES structure of a polymer monomer. In some embodiments, the polymer monomer may be a monomer structure of a variety of different polymers or possible polymers. In this way, multiple SMILES representations may reflect the properties of parts that may be different polymers. SMILES representations may be one-dimensional symbols focusing on chemical connectivity, while the 3D structure provides a detailed and three-dimensional depiction of the spatial arrangement of the polymer, thereby enabling a more comprehensive understanding of its structure, conformation, and properties. In some embodiments, multiple 3D structures may be provided for each of the multiple SMILES representations. For example, a particular SMILES representation may be used to generate a “cis” 3D structural orientation and a “trans” 3D structural orientation.
[0041] At step 246, method 220 may include replacing a methyl group or other functional group at the terminal junction of the repeating unit to form a stable monomer structure. In some embodiments, replacing the methyl group or other functional group can be used to determine a plurality of stable monomer structures that can be used to generate different polymer structures. For example, different combinations of methyl groups or other functional groups can generate stable monomer structures, while other combinations can generate unstable monomer structures. In this way, unstable monomer structures can be filtered out or not utilized. In some embodiments, when different functional groups are added to the terminal junctions, the orientation of the 3D structure can be altered, which can produce more stable or less stable molecules.
[0042] At step 248, method 220 may include: linking monomers to generate multiple possible two-body and tetra-body local clusters to form a stable structure. Different combinations of methyl groups or other functional groups can produce multiple two-body and tetra-body local structures that are stable or substantially stable to be utilized by a particular polymer. In this way, unstable two-body clusters and / or unstable tetra-body clusters can be ignored or removed from future calculations.
[0043] At step 252, method 220 may include simulating multiple monomeric local clusters, two-body local clusters, and four-body local clusters (e.g., stable structures, etc.) to determine different conformations and corresponding partition functions. In some embodiments, multiple simulated monomeric local clusters, two-body local clusters, and / or four-body local clusters can be simulated by determining the stable structure of step 248. At step 254, method 220 may include using quantum chemical methods to perform property calculations. In some examples, quantum chemical methods may include, but are not limited to, PM6, extended tight-binding (XTB), B3LYP, and / or M06-2X to determine the properties of two-body and four-body local structures.
[0044] As used herein, XTB methods can include semi-empirical (using both theoretical approximations and empirical data) quantum chemical methods for approximating the electronic structure of molecules. As used herein, the M06-2X method refers to methods used for density functional theory (DFT) calculations. M06-2X is a hybrid elementary GGA (generalized gradient approximation) functional. Compared to many other functionals, it includes a higher amount of Hartree-Fock exchange (approximately 54%), which helps to accurately model non-covalent interactions and transition states. In some implementations, the M06-2X method can accurately describe the kinetics of non-covalent interactions and chemical reactions.
[0045] As used herein, DFT calculations can refer to quantum mechanical modeling methods used to study the electronic structure of many-body systems, particularly atoms, molecules, and condensed phases. For example, DFT calculations can include computational quantum mechanical modeling methods that focus on electron density rather than wavefunctions to calculate the properties of matter. In some implementations, DFT calculations can be based on electron density as the primary quantity. They can utilize theorems that allow the properties of a system to be determined by the spatial distribution of electron density. In some implementations, DFT calculations can utilize the Hohenberg-Kohn theorem and / or the Kohn-Sham formula. DFT calculations can utilize energy calculations, material properties, and / or functional approximations.
[0046] For example, DFT calculations can calculate the total energy of a system (e.g., two-body local clusters, four-body local clusters, etc.), which can include contributions from kinetic energy, potential energy, and / or electron-electron interaction energy. Additionally, DFT calculations can predict the physical and / or chemical properties of materials and molecules, such as molecular structure, reaction energies, electronic properties, and / or other properties. In some implementations, the exact functional form may not be known, and therefore approximations can be used. Approximations such as the local density approximation (LDA) and / or the generalized gradient approximation (GGA) can be used.
[0047] At step 256, method 220 may include: calculating a statistical average of the corresponding properties calculated by the partition function. Calculating the statistical average of the corresponding properties may be based on the number of two-body local structure properties and four-body local structure properties within a specified polymer. In some embodiments, the statistical average may refer to mathematical calculations, such as the mean, mode, or other mathematical averages. In this way, the number of two-body local structures and four-body local structures with a particular property can be used to determine the statistical average of each property. In some embodiments, the statistical average may be a measure of the trend of the set of values for that property. The statistical average may summarize the dataset using a single number representing the center point of the data distribution.
[0048] At step 258, method 220 may include: predicting polymer properties using the calculated properties. As described herein, combined properties of two-body local structures and four-body local structures can be provided to generate predicted properties of the polymer.
[0049] Figure 3An exemplary flowchart of a method 330 for modeling based on molecular graph structures to predict polymer properties using local structures (e.g., local oligomerization structures, etc.) is illustrated. In some examples, method 330 may be performed by a computing device as described herein. Method 330 may be used to train neural networks or other types of machine learning models to determine polymer properties of a given or desired polymer based on local structure data and GNN output data. In some embodiments, the trained neural network may be used to reverse engineer polymers that, when used with flexible OLEDs, can reduce parasitic capacitance, improve response time, reduce power consumption, and enhance insulation properties.
[0050] At step 361, method 330 may be performed to generate a dataset comprising the corresponding structure and properties of each of a plurality of monomers corresponding to a specified polymer. In some embodiments, the corresponding structure and properties corresponding to the specified polymer may be specific structures and / or properties desired for a particular function. For example, the corresponding structures and properties may be used for a specific purpose, such as, but not limited to, OLED display materials. In some embodiments, the generated dataset of a plurality of monomers may be selected based on the stability of the plurality of monomers. In some embodiments, the dataset may include a plurality of monomers simulated based on known monomers. That is, the plurality of monomers may be virtual representations of monomers including specific structural features and / or specific properties that correspond to the desired properties of monomers having desired properties, such as a specific dielectric constant or a dielectric constant within a specific range.
[0051] In some embodiments, method 330 may be performed to generate structures of multiple polymers using SMILES representations of multiple polymers. As described herein, the structures of multiple polymers may be SMILES representations that can be used to represent the chemical linkages of the multiple polymers. In some embodiments, the multiple polymers may be converted from SMILES representations to 3D representations to represent the orientation of the multiple polymers.
[0052] In some embodiments, method 330 can be performed to model multiple polymers in a dataset by substituting functional groups at the terminal junctions of repeating monomer units in multiple polymers. In some embodiments, substituting functional groups can be used to determine multiple stable monomer structures that can be used to generate different polymer structures. For example, different combinations of methyl groups or other functional groups can generate stable monomer structures, while other combinations can generate unstable monomer structures. In this way, unstable monomer structures can be filtered out or not utilized. In some embodiments, when different functional groups are added to the terminal junctions, the orientation of the 3D structure can be altered, which can produce more stable or less stable molecules.
[0053] At step 362, method 330 can be performed to convert the corresponding structure into a graph-connected structure. In some embodiments, multiple local clusters and / or polymers identified as stable structures can be converted or utilized to generate a graph-connected structure. The graph-connected structure can be a GNN representation of multiple local clusters and / or a GNN representation of the polymer.
[0054] At step 363, method 330 may be performed to determine multiple local clusters based on multiple monomers. Determining multiple local clusters may include identifying two-body clusters and / or three-body clusters from a plurality of monomers determined to be stable. In some embodiments, determining multiple local clusters may include identifying multiple local clusters of a specified polymer. Identifying multiple local clusters may include identifying the specified polymer as a combination of multiple local clusters, and the polymer properties may be approximately represented by a statistical average of the properties of the local clusters. In some embodiments, the topological connectivity and interactions (e.g., dipole-dipole interactions) between local clusters may remain substantially unchanged during the optimization process.
[0055] In principle, as long as the conformational distribution of local clusters and their interaction energies are identified, the overall polymer properties remain unchanged. Local cluster models focus on the local microstructure of the polymer, thus expecting that the properties corresponding to these structures can form the basis for predicting the overall polymer properties. In principle, if the size of the local clusters increases, the scope of study can continue to expand until it approaches the entire polymer. Unfortunately, with the increase in local cluster size and the exponential increase in the number of conformations, the computational cost increases exponentially and may eventually exceed computational capabilities.
[0056] Considering the balance between computational accuracy and efficiency, a threshold size for the monomer can be chosen (e.g., molecular weight, number of atoms of a specific size, etc.). For example, a threshold of 20 heavy atoms can be chosen, where each of the multiple local clusters has no more than 20 heavy atoms. For some cases with cross-linking, single chains may need to satisfy a basic T-shaped cross or cross-crossing structure. As used herein, cross-linking or cross-linking refers to a bond that links one polymer chain to another. For example, cross-linking can be covalent, ionic, and / or physical. As used herein, T-shaped cross-linking and cross-crossing can refer to specific types of cross-linking or cross-linking. For example, T-shaped cross-linking can occur when one polymer chain forms a “T” shape with another polymer chain. For example, a single polymer chain acting as a branch can connect at points along the length of another polymer chain, thus forming a “T” shape. Cross-crossing can occur when multiple cross-crossings occur between the same two polymer chains, or when one polymer chain connects to many other polymer chains at several points. Some specific cross-linking structures may result in all local clusters connecting into a single entity. Meanwhile, representative molecular conformations may affect properties such as charge mobility, molecular shape, and cavity size.
[0057] In some embodiments, method 330 can be performed to simulate monomer-stable, two-body, and four-body stable structures from multiple polymers to determine multiple local clusters. In these embodiments, method 330 can be performed to calculate the density of the multiple local clusters using density functional theory and to calculate the dielectric constant of the multiple local clusters using quantum chemical methods. As described herein, each of the monomer structure, two-body structure, and four-body structure can be used to determine the corresponding density, dielectric constant, or other properties. In this way, the statistical average of the monomer structure, two-body structure, and / or four-body structure within the polymer can be used to determine the corresponding properties of the entire polymer.
[0058] At step 364, method 330 can be performed to calculate the corresponding polymer properties of each of the plurality of local clusters based on density functional theory (DFT). In some embodiments, the density of the plurality of local clusters can be used to determine the dielectric constant of the plurality of local clusters. In a similar manner to density, the statistical average of the determined dielectric constants of the plurality of local clusters can be used to determine the dielectric constant of the entire polymer or the polymer as a whole.
[0059] At step 365, method 330 may be performed to input the corresponding polymer properties and the output values of the graph neural network into a machine learning model, which is trained to determine unknown polymer properties of the specified polymer. In some embodiments, monomer data or polymer properties of monomers may be provided as input to the machine learning model, which is trained to determine unknown polymer properties of the specified polymer. Input monomer data and / or local cluster data may include the corresponding polymer properties of the local clusters. As described herein, local cluster data may be provided to the machine learning model. In some examples, input monomer data and / or local cluster data may include density data associated with the monomer data and / or local cluster data.
[0060] As described herein, density data for two-body and four-body clusters can be provided along with the output of the GNN to determine the density of the entire specified polymer. In these examples, the density data of the specified polymer can be used as input along with the calculated two-body and four-body dielectric constant data to determine the dielectric constant of the entire specified polymer.
[0061] In some implementations, method 330 can be performed to provide inputs to the GNN, including the nuclear charge number of the specified polymer, the net atomic charge of each monomer molecule in a plurality of monomer molecules, the number of heavy atom connections, and the number of hydrogen atom connections. Using this type of data, the GNN can be used to learn and analyze the structural and chemical properties of the specified polymer. As described herein, the output of a GNN applied to polymer molecules can provide diverse data associated with the properties and / or structure of the polymer. For example, the output or GNN data can include, but is not limited to: node embedding, graph embedding, property prediction, chemical activity, visualization, quantitative descriptors, graph-based analysis, and / or polymer interactions.
[0062] At step 366, method 330 may be performed to receive predictions of unknown polymer properties of a specified polymer from a machine learning model. As described herein, the polymer property may be a predicted unknown dielectric constant. The predicted unknown dielectric constant may be based on a statistical average of the dielectric constants of multiple local clusters associated with the specified polymer. As described herein, one or more machine learning models may be trained to predict the unknown dielectric constant of the specified polymer based on monomer data and / or local cluster data. In some embodiments, method 330 may be performed to determine the dielectric constant of the specified polymer using a statistical average of the corresponding properties of multiple local clusters.
[0063] Figure 4An example of a machine-readable medium 440 is illustrated for modeling based on molecular graph structures to predict polymer properties using local structures (e.g., local oligomerization structures, etc.). The machine-readable medium 440 can be communicatively connected to a processor resource 471 via a communication path 472. In some examples, the communication path 472 may include a wired or wireless connection that allows communication between devices and / or components within a single device. As used herein, the processor resource 471 may include, but is not limited to: a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a metal programmable cell array (MPCA), a semiconductor-based microprocessor, or other combinations of circuitry and / or logic for coordinating the execution of instructions 473, 474, 475, 476, 477, 478. In a particular example, the processor resource 471 utilizes a non-transitory computer-readable medium storing instructions 473, 474, 475, 476, 477, 578, which, when executed, cause the processor resource 471 to perform the corresponding function.
[0064] Machine-readable storage medium 440 can be an electronic, magnetic, optical, or other physical storage device that stores executable instructions. Therefore, a non-transitory machine-readable medium (MRM) (e.g., machine-readable medium 440) can be, for example, a non-transitory MRM including random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), a storage drive, an optical disk, etc. Machine-readable medium 440 can be disposed within a controller and / or computing device. In this example, executable instructions 473, 475, 476, 477, and 478 can be "installed" on the device. Additionally and / or alternatively, machine-readable medium 440 can be a portable, external, or remote storage medium, for example, allowing a computing system to download instructions 473, 475, 476, 477, and 478 from a portable / external / remote storage medium. In this case, the executable instructions can be part of an "installation package".
[0065] Machine-readable medium 440 includes instructions 473 for generating a dataset that includes the structures and properties of multiple known polymers. In some embodiments, the dataset includes structures and properties as desired structures and properties of a specified polymer. In this way, the dataset can include information about the structures and properties of known polymers, including properties of the specified polymer.
[0066] In some embodiments, machine-readable medium 440 may include instructions for using a machine learning model to predict multiple polymer structures within defined ranges of dielectric constant and density. As described herein, the multiple polymer structures may be simulated representations of real polymer structures including desired properties. In this way, the defined ranges of dielectric constant and / or defined ranges of density may be desired properties of a specified polymer.
[0067] Machine-readable medium 440 includes instructions 474 for generating a molecular map of each of the plurality of known polymers from SMILES representations of the plurality of known polymers. Machine-readable medium 440 includes instructions 475 for determining a plurality of local clusters for each of the plurality of known polymers, wherein the plurality of local clusters include simulated monomer-stable structures, two-body-stable structures, and four-body-stable structures from the plurality of known polymers.
[0068] Machine-readable medium 440 includes instructions 476 for calculating the density and dielectric constant of each of a plurality of local clusters. Machine-readable medium 440 includes instructions 477 for calculating a statistical average of the density and dielectric constant of the plurality of local clusters for each of a plurality of known polymers. In some embodiments, the statistical average is based on the number of each of the plurality of local clusters.
[0069] Machine-readable medium 440 includes instructions 478 for training a machine learning model to determine an unknown dielectric constant of a specified polymer from the output of a GNN for the specified polymer and density data determined from the molecular graph structure of the specified polymer. Training the machine learning model may include changing and / or updating the weights or weight vectors of the machine learning model to improve the accuracy of the machine learning model's ability to accurately predict the properties of the polymer based on the input. As described herein, the input may include the output of the GNN and density data and / or dielectric constant data of two-body local clusters and four-body local clusters.
[0070] In some embodiments, the machine-readable medium 440 includes instructions for providing the nuclear charge number of a specified polymer, the net atomic charge of each monomer molecule in the monomer molecule, the number of heavy atom connections, and the number of hydrogen atom connections as inputs to the GNN. In these embodiments, the machine-readable medium 440 may include instructions for providing the output of the GNN as inputs to a machine learning model.
[0071] In some implementations, machine-readable medium 440 may include instructions for training a machine learning model to utilize density data of two-body local clusters and four-body local clusters as descriptors. Two-body local clusters and four-body local clusters can be used as inputs or descriptors to the machine learning model. As used herein, a descriptor refers to a feature or attribute used to represent and understand data within a machine learning model. Descriptors can serve as fundamental inputs that a machine learning model can use to make predictions or classifications.
[0072] Figure 5 Examples of devices 550 for modeling molecular graph structures to predict polymer properties using local structures (e.g., local oligomerization structures, etc.) are illustrated. In some examples, device 550 is a computing device that includes processor resources 571 and machine-readable medium 540 for storing instructions 582, 583, 584, 585, 586, 587, 588 that are executed by processor resources 571 to perform specific functions. Figure 5 This illustrates how a computing device can execute instructions to perform the functions described herein.
[0073] Device 550 includes instructions 582 stored in machine-readable medium 650, which are executed by processor resource 571 to generate a dataset containing the structure, density, and dielectric constant of multiple polymers. As described herein, the dataset may include information relating to the simulated polymer, which may include properties that are the same as or similar to the desired properties of a specified polymer.
[0074] The device 550 includes instructions 583 stored by a machine-readable medium 650, which are executed by a processor resource 571 to generate a molecular map of each of the multiple polymers using SMILES representations of the multiple polymers.
[0075] Device 550 includes instructions 584 stored by machine-readable medium 650, which are executed by processor resource 571 to determine multiple local clusters for each of a plurality of polymers. In some embodiments, device 550 may include instructions for simulating the free energy and corresponding properties of the multiple local clusters. In some embodiments, the multiple local clusters include monomer interactions, two-body interactions, and four-body interactions of the multiple polymers.
[0076] Device 550 includes instructions 585 stored by machine-readable medium 540, which are executed by processor resource 571 to calculate the local density and local dielectric constant of each of a plurality of local clusters. In some embodiments, device 550 may include instructions for calculating the density of the plurality of local clusters using Becke’s three-parameter exchange function method (B3LYP) in combination with Lee, Yang, and Parr correlation function methods, and for calculating the unknown dielectric constant using a parametric method (PM6).
[0077] The device 550 includes instructions 586 stored in a machine-readable medium 540, which are executed by a processor resource 571 to calculate a statistical average of the density and dielectric constant of multiple local clusters of each of the multiple polymers.
[0078] In some embodiments, device 550 may include instructions for calculating the polymerization mean and maximum values of elements of a plurality of polymers, and for calculating the polymerization mean and maximum value of each of the plurality of polymers as a whole molecule.
[0079] Device 550 includes instructions 587 stored by machine-readable medium 650, which are executed by processor resource 571 to input monomer data into a machine learning model trained to determine the unknown dielectric constant of a specified polymer from the output values of the GNN and density values of local clusters from a plurality of identifiers of the specified polymer. In some embodiments, device 550 may include instructions for providing polymer property data as input to the GNN.
[0080] In some embodiments, device 550 may include instructions for determining the topological connectivity patterns of an unknown polymer. As used herein, a topological connectivity pattern is a representation of how the atoms or segments of the polymer are connected to each other, including arrangements of polymer chains and any possible branches or crosslinks. That is, a topological connectivity pattern is a spatial map of the polymer molecular structure and how each component of its constituent parts is connected.
[0081] The device 550 includes instructions 588 stored by a machine-readable medium 650, which are executed by a processor resource 571 to receive a prediction of an unknown dielectric constant for a specified polymer from a machine learning model. As described herein, a first machine learning model can be used to determine the density of multiple local clusters, and a second machine learning model can be used to determine the dielectric constant based on the determined density. In this way, a desired dielectric constant for a specified polymer can be achieved.
[0082] Although specific embodiments have been described above, these embodiments are not intended to limit the scope of this disclosure, even where only a single embodiment is described with respect to a particular feature. Unless otherwise stated, the examples of features provided in this disclosure are intended to be illustrative and not restrictive. The above description is intended to cover such alternatives, modifications, and equivalents that will be obvious to those skilled in the art who benefit from this disclosure.
[0083] The scope of this disclosure includes any feature or combination of features disclosed herein (express or implicit), or any generalization thereof, whether or not it alleviates any or all of the problems addressed herein. Various advantages of this disclosure have been described herein; however, embodiments may provide some, all, or none of these advantages, or may provide other advantages.
[0084] In the foregoing specific embodiments, for the purpose of simplifying this disclosure, some features are combined in a single embodiment. This approach of the disclosure should not be construed as reflecting an intention that the disclosed embodiments of the disclosure must use more features than are expressly recited in each claim. Rather, as reflected in the following claims, the subject matter of the invention does not consist of all features of a single disclosed embodiment. Therefore, the following claims are hereby incorporated into the specific embodiments, wherein each claim exists independently as a separate embodiment.
Claims
1. A method, the method comprising: Generate a dataset that includes the structure and properties of each of a plurality of monomers corresponding to a specified polymer; Convert the corresponding structure into a graph-connected structure; Multiple local clusters are determined based on the aforementioned multiple monomers; The corresponding polymer properties of each of the multiple local clusters are calculated based on density functional theory (DFT). The corresponding polymer properties and the output values of the graph neural network are input into a machine learning model, which is then trained to determine the unknown polymer properties of a given polymer. as well as The machine learning model receives predictions of the unknown polymer properties of the specified polymer.
2. The method according to claim 1, further comprising: The graph neural network is provided with inputs including the nuclear charge of the specified polymer, the net atomic charge of each monomer molecule in the plurality of monomer molecules, the number of heavy atom connections, and the number of hydrogen atom connections.
3. The method according to claim 1, further comprising: The structures of the plurality of polymers are generated using a simplified molecular linear input specification.
4. The method according to claim 1, further comprising: The plurality of polymers in the dataset are simulated by substituting functional groups at the end junctions of repeating monomer units of the plurality of polymers.
5. The method according to claim 1, further comprising: The plurality of local clusters are determined by simulating monomer-stable, two-body-stable, and four-body-stable structures of the plurality of polymers.
6. The method according to claim 5, further comprising: The density of the plurality of local clusters is calculated using the DFT, and the dielectric constant of the plurality of local clusters is calculated using quantum chemical methods.
7. The method according to claim 1, further comprising: The dielectric constant of the specified polymer is determined by using the statistical average of the corresponding properties of the plurality of local clusters.
8. A machine-readable medium storing machine-readable instructions that, when executed by processor resources of a device, cause the processor to: Generate a dataset that includes the structures and properties of multiple known polymers; Generate a molecular graph for each of the known polymers from the simplified molecular linear input specification (SMILES) representation of the plurality of known polymers; Determine multiple local clusters for each of the plurality of known polymers, wherein the plurality of local clusters include simulated monomer-stable structures, two-body-stable structures and four-body-stable structures from the plurality of known polymers; Calculate the density and dielectric constant of each of the plurality of local clusters; Calculate the statistical average of the density and dielectric constant of the plurality of local clusters for each of the plurality of known polymers; and A machine learning model is trained to determine the unknown dielectric constant of a given polymer from the output of a graph neural network of the given polymer and density data determined from the molecular graph structure of the given polymer.
9. The machine-readable medium of claim 8, wherein the machine-readable medium includes instructions for using the machine learning model to identify a plurality of polymer structures within a defined range of dielectric constants and a defined range of density.
10. The machine-readable medium of claim 8, wherein the machine-readable medium includes instructions for providing the nuclear charge number of the specified polymer, the net atomic charge of each of the plurality of monomer molecules, the number of heavy atom connections, and the number of hydrogen atom connections as inputs to the graph neural network.
11. The machine-readable medium of claim 10, wherein the machine-readable medium includes instructions for providing the output of the graph neural network as input to the machine learning model.
12. The machine-readable medium of claim 8, wherein the machine-readable medium includes instructions for training the machine learning model to utilize the density data of two-body local clusters and four-body local clusters as descriptors.
13. The machine-readable medium of claim 8, wherein the statistical average is based on the number of each of the plurality of local clusters.
14. An apparatus, the apparatus comprising: processor; and Non-transitory memory resources, on which machine-readable instructions are stored, which, when executed, cause the processor resources to: Generate a dataset that includes the structure, density, and dielectric constant of multiple polymers; A molecular graph for each of the plurality of polymers is generated using the simplified molecular linear input specification (SMILES) representation of the plurality of polymers; Identify multiple local clusters for each of the plurality of polymers; Calculate the local density and local dielectric constant of each of the plurality of local clusters; Calculate the statistical average of the density and dielectric constant of the plurality of local clusters for each of the plurality of polymers; Monomer data is input into a machine learning model, which is trained to determine the unknown dielectric constant of a given polymer from the output values of a graph neural network and the density values of local clusters of multiple identifiers of that polymer; and The machine learning model receives a prediction of the unknown dielectric constant of the specified polymer.
15. The device of claim 14, wherein the processor is configured to calculate the density of the plurality of local clusters using Becke’s three-parameter exchange function method (B3LYP) in combination with Lee, Yang and Parr correlation function methods, and to calculate the unknown dielectric constant using a parametric method (PM6).
16. The device of claim 14, wherein the processor is configured to determine the topological connectivity pattern of the unknown polymer.
17. The apparatus of claim 14, wherein the processor is configured to provide polymer property data as input to the graph neural network.
18. The device of claim 14, wherein the processor is configured to simulate the free energy and corresponding properties of the plurality of local clusters.
19. The device of claim 18, wherein the plurality of local clusters comprises monomer interactions, two-body interactions, and four-body interactions of the plurality of polymers.
20. The apparatus of claim 14, wherein the processor is configured to calculate the average and maximum polymerization values of the elements of the plurality of polymers, and to calculate the average and maximum polymerization values of each of the plurality of polymers as a whole molecule.