Polymer screening method based on machine learning
By employing a machine learning-based polymer screening method and automated synthesis, the problems of low efficiency and high cost in screening antimicrobial polymer materials in traditional methods have been solved, enabling the screening and design of efficient and low-cost multifunctional synergistic antimicrobial polymer materials.
Patent Information
- Application Number
- CN202510841848.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-11-14
AI Technical Summary
Traditional methods are inefficient and costly in high-throughput screening of antimicrobial polymer materials, cannot effectively guide the design of new materials, and lack sufficient understanding of multi-scale structure-property relationships.
A machine learning-based polymer screening method is adopted, combined with automated synthesis. A graph-based deep learning model is used to represent the polymer molecular structure. Through active learning strategy and uncertainty sampling, polymer properties are predicted, and potential high-quality materials are screened out through automated synthesis.
It achieves efficient and low-cost high-throughput screening, breaks through the bottleneck of multi-scale parameter combination optimization of multifunctional synergistic antibacterial polymers, and improves the efficiency and accuracy of material design.
Smart Images

Figure CN120954559A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of chemistry and information technology, and more specifically to a polymer screening method based on machine learning. Background Technology
[0002] In recent years, the problem of multidrug-resistant bacteria has intensified globally, becoming a global challenge. However, the development of new antibiotics is extremely difficult; in the past 50 years, only a few have been approved by the FDA, and none are effective against Gram-negative bacteria. Humanity is gradually entering a "post-antibiotic era." Currently, various sectors are striving to develop new antibacterial strategies. Among these, polymer materials possess advantages such as controllable structure, diverse functions, and good processing performance. Through rational structural design, multiple functions, including antibacterial properties and biocompatibility, can be integrated to prevent and treat microbial infections. However, due to the numerous construction parameters of polymer materials, there is a lack of systematic understanding of the structure-activity relationship between multi-scale parameter combinations and multifunctional synergistic antibacterial effects in antibacterial polymer materials. Traditional trial-and-error methods are typically low-throughput, time-consuming, and costly, failing to achieve efficient combination optimization. Therefore, the development of new and effective methods is urgently needed.
[0003] The advent of photoinitiation technology has broken the stringent reaction conditions of traditional thermal initiation, making high-throughput synthesis possible. Among them, photoinduced electron / energy transfer reversible addition-fragmentation chain transfer polymerization (PET-RAFT) is the most commonly used. The photoredox catalyst enters an excited state under visible light irradiation, and then reduces the chain transfer agent through the PET process to form free radicals. The ability of ground-state molecular oxygen to transform into a singlet state through triplet-triplet affinity (TTA) endows PET-RAFT with excellent oxygen resistance, simplifying the polymerization process and reducing production costs.
[0004] If high-throughput strategies have shifted biomaterials research from traditional trial-and-error methods to parallel experiments, the introduction and development of automation technologies can further reduce human and material costs, and improve experimental throughput, speed, accuracy, and reproducibility. Automation technologies include automation of synthesis and automation of characterization. Automated synthesis is mainly achieved through liquid or solid workstations; and currently, many instruments, such as microplate readers, dynamic light scattering systems, and chromatographs, can be linked with upstream synthesis or post-processing via robotic arms and dual-arm collaborative robots to achieve automated characterization.
[0005] While the use of automated robots in research has made high-throughput scanning possible, brute-force scanning cannot grasp structure-activity relationships or guide the design of new materials, thus remaining inefficient and costly. Therefore, the introduction of machine learning plays a crucial role in the field of biomedical macromolecules and has seen significant progress in recent years. As a branch of artificial intelligence, machine learning allows computer programs to simulate human learning behavior—that is, to discover patterns from existing data, continuously improve themselves, and make predictions and judgments about unknown data.
[0006] Therefore, linking automation and machine learning is expected to be an effective means of high-throughput deciphering the structure-property relationships of materials at multiple scales, and it remains a technological gap that urgently needs to be filled in the current market. Summary of the Invention
[0007] This invention provides a polymer screening method based on machine learning, which can be combined with automated synthesis to achieve high-throughput screening and can be applied to antimicrobial polymer screening, etc.
[0008] The specific technical solution is as follows: A polymer screening method based on machine learning, wherein the raw material composition of the polymer contains two or more monomers and each monomer independently carries one or more active groups, and the active groups are related to the properties of the polymer to be screened. The machine learning-based polymer screening method uses a graph-based deep learning model to represent polymer molecular structures, with one polymer molecular structure corresponding to one polymer molecular structure graph. Each polymer molecular structure diagram contains a global virtual node and multiple monomer diagrams. Each monomer diagram corresponds to each monomer used to synthesize the polymer molecular structure. For any monomer diagram, non-hydrogen atoms (e.g., C, N, O, etc.) and half-unsaturated bonds (e.g., carbon-carbon double bonds, etc.) used to connect with another monomer (which can be the same or a different monomer) are used as nodes. Each monomer diagram is connected to a monomer virtual node. The global virtual node connects all nodes in all monomer diagrams of the polymer molecular structure diagram in which it is located, as well as each monomer virtual node. The proportion of each monomer in all monomers that make up the polymer molecular structure is reflected by the connection length between the monomer virtual node corresponding to that monomer and the global virtual node. A graph-based deep learning model is trained using polymers with known structures and properties as seed datasets. The model employs an active learning strategy, with inputs including polymer molecular structures, specifically monomer types and the proportion of each monomer in the feed, and outputs including polymer properties. A trained graph-based deep learning model is used to predict and screen the properties of polymers with known structures but unknown properties.
[0009] In some embodiments, the active learning strategy in the machine learning-based polymer screening method includes uncertainty sampling.
[0010] In some embodiments, the machine learning-based polymer screening method scores polymers with known structures when predicting polymer performance using a graph-based deep learning model. This scoring integrates polymer performance and uncertainty, where the uncertainty score is determined based on the upper limit of the confidence interval. Further optionally, the machine learning-based polymer screening method sets a scoring threshold, clusters polymers with predicted scores greater than or equal to the threshold, and selects polymers from different clusters for performance verification and to train and optimize the graph-based deep learning model, avoiding overfitting. Specifically, the clustering method of this invention can be K-Means clustering, etc.
[0011] In some embodiments, the machine learning-based polymer screening method, when training the graph-based deep learning model, sets a molecular structure similarity threshold, wherein the molecular structure similarity between any polymer training samples is greater than or equal to the set threshold. Further optionally, in the machine learning-based polymer screening method, the graph-based deep learning model represents the polymer on a two-dimensional reduced-dimensional graph based on the polymer molecular structure using t-distributed random embedding (t-SNE), and uses the distance between polymer training samples on the two-dimensional reduced-dimensional graph to reflect the molecular structure similarity between polymer training samples.
[0012] In some embodiments, the polymer screening method based on machine learning uses random copolymers synthesized using PET-RAFT. PET-RAFT, as a living radical polymerization method, is particularly suitable for this invention because the polymer molecular structure can be considered determined when the types and proportions of monomers used in the synthesis are fixed.
[0013] In some embodiments, the monomers in the machine learning-based polymer screening method include one or more of vinyl monomers, acrylate monomers, and acrylamide monomers.
[0014] This invention can be used to screen biomimetic antimicrobial polymers with antimicrobial peptides.
[0015] In some embodiments, the active groups in the machine learning-based polymer screening method include one or more of cationic groups, hydrophobic groups, and hydrophilic groups.
[0016] In some embodiments, the machine learning-based polymer screening method includes polymer properties such as antibacterial properties and biocompatibility. Further optionally, biocompatibility can be reflected by hemolytic toxicity.
[0017] It is understood that this invention is universal, and the active groups therein can also be other functional groups, which can endow the polymer with other properties.
[0018] In some embodiments, the polymer screening method based on machine learning involves obtaining the polymer through PET-RAFT of monomers in the presence of a photosensitizer and a chain transfer agent under visible light irradiation. Further optionally, the PET-RAFT is performed in solution. Even more optionally, the solvent in the solution includes at least one selected from water, ethanol, methanol, tetrahydrofuran, N,N-dimethylformamide, and dimethyl sulfoxide. The visible light irradiation time can be adjusted according to the solution volume and the power of the light source.
[0019] In some embodiments, the machine learning-based polymer screening method further includes automated polymer synthesis and performance verification. Optionally, the automated synthesis employs a Hamilton MLSTARlet pipetting robot, which can convert information such as the material design composition, degree of polymerization, and viscosity into liquid sequence, concentration, volume, and dispensing speed.
[0020] In some embodiments, the machine learning-based polymer screening method uses a graph-based deep learning model to recommend polymer molecular structures based on current training results (the recommended polymer molecular structures are generally expected to have target performance). Polymers are then automatically synthesized according to the recommended polymer molecular structures, and performance experiments are conducted. The automatically synthesized polymer molecular structures and their performance are used to train the graph-based deep learning model to optimize the polymer molecular structure recommendations.
[0021] Herein, the present invention provides a preferred polymer screening method based on machine learning, including an automated and machine learning-assisted high-throughput screening strategy and its application in the screening of antimicrobial polymers.
[0022] An automated and machine learning-assisted high-throughput screening strategy combines the design, preparation, and screening of antimicrobial peptide biomimetic molecules to obtain antimicrobial polymers; the synthesis method is PET-RAFT; the automation technology is implemented by the Hamilton MLSTARlet pipetting robot; the machine learning strategy includes a graph-based deep learning model and a multi-objective Bayesian optimization problem based on a sample pool.
[0023] Specifically, it includes: (1) Synthesis of random copolymer molecular library: The monomers are premixed and then subjected to living radical polymerization under the action of photosensitizer and chain transfer agent and irradiation with visible light to obtain random copolymers; (2) Hamilton MLSTARlet pipetting robot-assisted automated synthesis: The control program of the pipetting robot is programmed in Python, which can realize the conversion of information such as the composition, degree of polymerization, and viscosity of the material design into liquid sequence, concentration, volume and the speed of pipetting liquid; (3) Machine learning-assisted recommendation of the structure of synthetic polymers to achieve screening: a graph-based deep learning model is used to represent the molecular structure, and a multi-objective Bayesian optimization problem based on the sample pool is used to describe the screening target.
[0024] The preferred synthesis method of this invention is PET-RAFT, which can be photoinitiated at room temperature and pressure, and does not require dehydration or deoxygenation, but requires the addition of photosensitizers and chain transfer agents.
[0025] The Hamilton MLSTARlet pipetting robot, through calibration of pipetting parameters and the writing of built-in programs, can convert information such as the composition, degree of polymerization, and viscosity of the material design into liquid sequence, concentration, volume, and the speed of suction and dispensing.
[0026] In this invention, the polymer molecular structure is represented using a graph-based deep learning model. Each monomer component of the polymer is represented as a separate graph, with a global virtual node connecting all nodes in each monomer component. Simultaneously, monomer virtual nodes, corresponding one-to-one with the types of monomer components, are introduced to connect the monomer component graph and the global virtual node, specifying the feed ratio of each monomer component. Then, a pre-trained graph transformer is used to extract polymer features, followed by a readout operator. Finally, a projection head is used to map the layer representation to the attribute space (polymer properties).
[0027] In this invention, the screening objective can be categorized as a multi-objective Bayesian optimization problem based on a sample pool. A query strategy is employed to select the next sample to be labeled from the sample pool. Uncertainty and diversity are considered simultaneously during the strategy's implementation. An ensemble method is used to estimate the polymer's performance and the model's uncertainty, and an upper confidence level is applied as the retrieval function to score the remaining samples in the sample pool. Then, samples with higher scores are selected for clustering.
[0028] Compared with the prior art, the beneficial effects of this invention are as follows: 1) This invention preferably employs living radical polymerization to synthesize the target molecule, achieving a reaction efficiency close to 100%, greatly simplifying post-processing and enabling increased experimental throughput. Furthermore, the polymerization process is photoinitiated, conducted under ambient temperature and pressure, and requires no dehydration or deoxygenation, simplifying the polymerization process and improving reaction efficiency. The monomers used in the reaction are acrylates or acrylamides, and their function is achieved by introducing active groups at the ends, which is beneficial for constructing large molecular libraries. Based on this synthetic route, this work can employ antimicrobial peptide biomimetic molecular design, introducing cationic, hydrophobic, and hydrophilic groups at the monomer ends in different proportions to construct antimicrobial polymer libraries. It is worth noting that this invention focuses on applications in the antimicrobial field, but this technical route can be extended to multiple fields, such as protein conjugates and antifouling materials. Any polymer molecular libraries designed and prepared using this technical route for applications in other fields are within the scope of protection of this invention.
[0029] 2) This invention preferably combines automation technology with machine learning to systematically study the combined effect of the molecular structure of antibacterial polymers on antibacterial properties and biocompatibility, establishes a new “global optimization” model for multi-factor combined effects, breaks through the bottleneck of multi-scale parameter combination optimization of multifunctional synergistic antibacterial polymers, and is an innovation in the design and development mode of antibacterial polymer material systems. Attached Figure Description
[0030] Figure 1 This is a schematic diagram illustrating an automated and machine learning-assisted high-throughput screening strategy of the present invention and its application in the screening of antimicrobial polymers.
[0031] Figure 2 This is a schematic diagram of the Hamilton MLSTARlet pipetting robot of the present invention.
[0032] Figure 3 This is a schematic diagram illustrating the construction of a polymer graph, which is an example of the present invention.
[0033] Figure 4 This is a schematic diagram of the model training and prediction process of the present invention.
[0034] Figure 5 This is a t-SNE visualization of the screening process in Embodiment 1 of the present invention. Detailed Implementation
[0035] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Operating methods not specifically specified in the following embodiments are generally performed under conventional conditions or as recommended by the manufacturer.
[0036] See Figure 1An automated and machine learning-assisted high-throughput screening strategy includes: (I) Synthesis of random copolymer molecular libraries.
[0037] Monomers are premixed at a certain concentration, and under the action of photosensitizers and chain transfer agents, and under the irradiation of visible light, the monomers undergo free radical polymerization to obtain random polymers.
[0038] The photosensitizer was zinc tetraphenylporphyrin, the chain transfer agent was 2-(dodecylthiocarbonylthiothiothio)propionic acid, and the light source was green light (530 nm, 5 W).
[0039] The monomer is one or more of vinyl, acrylate, or acrylamide substances, specifically selected from cationic, hydrophobic, and hydrophilic monomers. Cationic monomers include N-(2-aminoethyl)acrylamide, N-(3-aminopropyl)acrylamide, N-(4-aminobutyl)acrylamide, N-(5-aminopentyl)acrylamide, N-(6-aminohexyl)acrylamide, N-(8-aminooctyl)acrylamide, N-(2-guanidinylethyl)acrylamide, N-(3-guanidinylpropyl)acrylamide, N-(4-guanidinylbutyl)acrylamide, N-(6-guanidinylhexyl)acrylamide, 1-(4-(aminomethyl)piperidin-1-yl)propyl-2-en-1-one, 2-(dimethylamino)ethyl methacrylate, 3-acrylamido-N,N,N-trimethylpropyl-1-amine, N-(2-(1H-indol-3-yl)ethyl)acrylamide, N-(4-) N-(2-(1H-imidazol-1-yl)acrylamide, N-(2-(1H-imidazol-1-yl)ethyl)acrylamide, hydrophobic monomers including methyl methacrylate, ethyl methacrylate, isopropyl methacrylate, butyl methacrylate, N-pentylacrylamide, N-isopentylacrylamide, N-neopentylacrylamide, N,N-diethylacrylamide, phenyl methacrylate, N-benzylacrylamide, cyclohexyl methacrylate, N-heptylacrylamide, N-octylacrylamide, hydrophilic monomers including N-hydroxymethylacrylamide, N-(2-hydroxyethyl)acrylamide, oligo(ethylene glycol) methyl ether acrylate, 2-methacryloyloxyethyl phosphocholine, N-[2-(3,4-dihydroxyphenyl)ethyl]-2-methylacrylamide, 4-acryloylmorpholine, β-(acryloyloxy)propionic acid, for example, the synthesis process can controllably use one each of cationic monomer, hydrophobic monomer, and hydrophilic monomer.
[0040] The synthesis of random copolymer molecular libraries was carried out in solution.
[0041] The visible light irradiation time can be adjusted according to the solution volume and the power of the light source. For example, in some embodiments, the total solution volume is 300 μL, the light source power is 5 W, and the visible light irradiation time is 16 h.
[0042] The solvent in the solution may include at least one of water, ethanol, methanol, tetrahydrofuran, N,N-dimethylformamide, and dimethyl sulfoxide. In some embodiments, dimethyl sulfoxide is selected. If other solvents are selected, an additional reactive oxygen species scavenger needs to be added.
[0043] (II) See Figure 2 Random copolymers were synthesized using an automated method assisted by a Hamilton MLSTARlet pipetting robot.
[0044] The control program for the pipetting robot is programmed in Python, which can convert information such as the composition, degree of polymerization, and viscosity of the material design into liquid sequence, concentration, volume, and the speed of liquid aspiration.
[0045] In some embodiments, the monomer, chain transfer agent, and photosensitizer are pre-prepared into solutions with concentrations of 0.25 mmol / mL, 0.05 mmol / mL, and 0.5 mg / mL, respectively. Then, a programmed pipetting robot is used to add total amounts of the monomer, chain transfer agent, and photosensitizer of 240 μL, 30 μL, and 30 μL, respectively. After the samples are added, the mixture is irradiated with green light for 16 h to obtain a polymer with a degree of polymerization of 40.
[0046] (III) Machine learning is used to recommend the structure of synthetic polymers and then screen them.
[0047] The polymer molecular structure is represented using a graph-based deep learning model. Each monomer component of the polymer is represented as a separate graph, with a global virtual node Vg connecting all nodes in each monomer component. Simultaneously, monomer virtual nodes (V1, V2, V3, etc.) are introduced in a one-to-one correspondence with the monomer component types, connecting the monomer component graph and the global virtual node Vg respectively to specify the feed ratio of each monomer component. An exemplary polymer graph construction diagram is shown below. Figure 3 As shown. Then, as Figure 4 As shown, a pre-trained graph transformer is used to extract polymer features, followed by a readout operator. Finally, a projection head is used to map the layer representation to the attribute space, and the performance data is output.
[0048] The selection objective can be categorized as a multi-objective Bayesian optimization problem based on a sample pool. A query strategy is employed to select the next sample to be labeled from the pool. Uncertainty and diversity are considered simultaneously during the strategy's implementation. An ensemble approach is used to estimate the polymer's performance and the model's uncertainty, and a confidence cap is applied as the retrieval function to score the remaining samples in the pool. The samples with higher scores are then selected for K-Means clustering. In some embodiments, n_clusters = 50 for K-Means clustering, and the top 3 samples from each cluster are selected as the final samples.
[0049] Example 1: Thirty different acrylates or acrylamides were selected as polymer monomers, tetraphenylporphyrin zinc was used as a photosensitizer, 2-(dodecylthiocarbonylthiothiothio)propionic acid was used as a chain transfer agent, green light (530nm, 5W) was used as the light source, and Escherichia coli was used as the target bacterial strain for screening.
[0050] (1) Synthesis of random copolymer molecular library: 11 cationic monomers, 13 hydrophobic monomers, 6 hydrophilic monomers and 16 different ratios were selected. A combination of 1 cationic monomer, 1 hydrophobic monomer and 1 hydrophilic monomer was selected to obtain an anti-E. coli polymer library with an expected 13,728 combinations.
[0051] (2) Hamilton MLSTARlet pipetting robot-assisted automated synthesis: The monomer, chain transfer agent and photosensitizer were prepared in advance into solutions with concentrations of 0.25 mmol / mL, 0.05 mmol / mL and 0.5 mg / mL. Then, the pipetting robot was programmed to add 240 μL, 30 μL and 30 μL of monomer, chain transfer agent and photosensitizer respectively. After the samples were added, the samples were irradiated with green light for 16 h to obtain a polymer with a degree of polymerization of 40.
[0052] (3) Machine learning-assisted recommendation of the structure of synthetic polymers for screening: Each monomer component of the polymer is represented as a separate graph, with one global virtual node connecting all nodes in each monomer component. Three virtual nodes are introduced to connect the graphs of the three monomer components and the global virtual node, specifying the feed ratio of each monomer component. Then, a pre-trained graph transformer is used to extract polymer features, followed by a readout operator. Finally, a projection head is used to map the layer representation to the attribute space. A query strategy is adopted to select the next sample to be labeled from the sample pool. An ensemble method is used to estimate the polymer performance and model uncertainty, and a confidence cap is applied as the acquisition function to score the remaining samples in the sample pool. Then, samples with higher scores are selected for K-Means clustering (n_clusters=50), and the top three samples from each cluster are selected as the final samples. Figure 5 This is a t-SNE visualization of the screening process in Embodiment 1 of the present invention.
[0053] Example 2: Thirty-three different acrylates or acrylamides were selected as polymer monomers, tetraphenylporphyrin zinc was used as a photosensitizer, 2-(dodecylthiocarbonylthiothiothio)propionic acid was used as a chain transfer agent, green light (530nm, 5W) was used as the light source, and Candida albicans was used as the target bacterial species for screening.
[0054] (1) Synthesis of random copolymer molecular library: 14 cationic monomers, 13 hydrophobic monomers, 6 hydrophilic monomers, 7 different cationic:hydrophobic:hydrophilic ratios and 4 different cationic 1:cationic 2 ratios were selected. Combinations of 2 cationic monomers, 1 hydrophobic monomer and 1 hydrophilic monomer were selected to obtain an anti-Candida albicans polymer library with an estimated 198,744 combinations.
[0055] (2) Hamilton MLSTARlet pipetting robot-assisted automated synthesis: The monomer, chain transfer agent and photosensitizer were prepared in advance into solutions with concentrations of 0.25 mmol / mL, 0.05 mmol / mL and 0.5 mg / mL. Then, the pipetting robot was programmed to add 240 μL, 30 μL and 30 μL of monomer, chain transfer agent and photosensitizer respectively. After the samples were added, the samples were irradiated with green light for 16 h to obtain a polymer with a degree of polymerization of 40.
[0056] (3) Machine learning-assisted recommendation of the structure of the synthetic polymer for screening: Each monomer component of the polymer is represented as a separate graph, with one global virtual node connecting all nodes in each monomer component. Four virtual nodes are introduced to connect the graphs of the four monomer components and the global virtual node, specifying the feed ratio of each monomer component. Then, a pre-trained graph transformer is used to extract polymer features, followed by a readout operator. Finally, a projection head is used to map the layer representation to the attribute space. A query strategy is adopted to select the next sample to be labeled from the sample pool. An ensemble method is used to estimate the polymer performance and model uncertainty, and a confidence cap is applied as the acquisition function to score the remaining samples in the sample pool. Then, samples with higher scores are selected for K-Means clustering (n_clusters=50), and the top 3 samples from each cluster are selected as the final samples.
[0057] Example 3: Thirty-six different acrylates or acrylamides were selected as polymer monomers, tetraphenylporphyrin zinc was used as a photosensitizer, 2-(dodecylthiocarbonylthiothiothio)propionic acid was used as a chain transfer agent, green light (530nm, 5W) was used as the light source, and methicillin-resistant Staphylococcus aureus was used as the target bacterial strain for screening.
[0058] (1) Synthesis of random copolymer molecular library: 16 cationic monomers, 13 hydrophobic monomers, 7 hydrophilic monomers, 16 different cationic:hydrophobic:hydrophilic ratios, 4 different cationic 1:cationic 2 ratios, and 4 different hydrophilic 1:hydrophilic 2 ratios were selected. Combinations of 2 cationic monomers, 1 hydrophobic monomer, and 2 hydrophilic monomers were selected to obtain a polymer library of 8,386,560 possible combinations for resistance to methicillin-resistant Staphylococcus aureus.
[0059] (2) Hamilton MLSTARlet pipetting robot-assisted automated synthesis: The monomer, chain transfer agent and photosensitizer were prepared in advance into solutions with concentrations of 0.25 mmol / mL, 0.05 mmol / mL and 0.5 mg / mL. Then, the pipetting robot was programmed to add 240 μL, 30 μL and 30 μL of monomer, chain transfer agent and photosensitizer respectively. After the samples were added, the samples were irradiated with green light for 16 h to obtain a polymer with a degree of polymerization of 40.
[0060] (3) Machine learning-assisted recommendation of the structure of the synthetic polymer for screening: Each monomer component of the polymer is represented as a separate graph, with one global virtual node connecting all nodes in each monomer component. Five virtual nodes are introduced to connect the graphs of the five monomer components and the global virtual node, specifying the feed ratio of each monomer component. Then, a pre-trained graph transformer is used to extract polymer features, followed by a readout operator. Finally, a projection head is used to map the layer representation to the attribute space. A query strategy is adopted to select the next sample to be labeled from the sample pool. An ensemble method is used to estimate the polymer performance and model uncertainty, and a confidence cap is applied as the acquisition function to score the remaining samples in the sample pool. Then, samples with higher scores are selected for K-Means clustering (n_clusters=50), and the top three samples from each cluster are selected as the final samples.
[0061] Example 4: Thirty-six different acrylates or acrylamides were selected as polymer monomers, tetraphenylporphyrin zinc was used as a photosensitizer, 2-(dodecylthiocarbonylthiothiothio)propionic acid was used as a chain transfer agent, green light (530nm, 5W) was used as the light source, and Pseudomonas aeruginosa was used as the target bacterial species for screening.
[0062] (1) Synthesis of random copolymer molecular library: 16 cationic monomers, 13 hydrophobic monomers, 7 hydrophilic monomers, 16 different cationic:hydrophobic:hydrophilic ratios, 4 different cationic 1:cationic 2 ratios, and 4 different hydrophobic 1:hydrophobic 2 ratios were selected. Combinations of 2 cationic monomers, 2 hydrophobic monomers, and 1 hydrophilic monomer were selected to obtain an anti-Pseudomonas aeruginosa polymer library with an expected 16,773,120 combinations.
[0063] (2) Hamilton MLSTARlet pipetting robot-assisted automated synthesis: The monomer, chain transfer agent and photosensitizer were prepared in advance into solutions with concentrations of 0.25 mmol / mL, 0.05 mmol / mL and 0.5 mg / mL. Then, the pipetting robot was programmed to add 240 μL, 30 μL and 30 μL of monomer, chain transfer agent and photosensitizer respectively. After the samples were added, the samples were irradiated with green light for 16 h to obtain a polymer with a degree of polymerization of 40.
[0064] (3) Machine learning-assisted recommendation of the structure of the synthetic polymer for screening: Each monomer component of the polymer is represented as a separate graph, with one global virtual node connecting all nodes in each monomer component. Five virtual nodes are introduced to connect the graphs of the five monomer components and the global virtual node, specifying the feed ratio of each monomer component. Then, a pre-trained graph transformer is used to extract polymer features, followed by a readout operator. Finally, a projection head is used to map the layer representation to the attribute space. A query strategy is adopted to select the next sample to be labeled from the sample pool. An ensemble method is used to estimate the polymer performance and model uncertainty, and a confidence cap is applied as the acquisition function to score the remaining samples in the sample pool. Then, samples with higher scores are selected for K-Means clustering (n_clusters=50), and the top three samples from each cluster are selected as the final samples.
[0065] Example 5: Thirty-six different acrylates or acrylamides were selected as polymer monomers, tetraphenylporphyrin zinc was used as a photosensitizer, 2-(dodecylthiocarbonylthiothiothio)propionic acid was used as a chain transfer agent, green light (530nm, 5W) was used as the light source, and Acinetobacter baumannii was used as the target bacterial species for screening.
[0066] (1) Synthesis of random copolymer molecular library: 16 cationic monomers, 13 hydrophobic monomers, 7 hydrophilic monomers, 16 different cationic:hydrophobic:hydrophilic ratios, 4 different cationic 1:cationic 2 ratios, 4 different hydrophobic 1:hydrophobic 2 ratios, and 4 different hydrophilic 1:hydrophilic 2 ratios were selected. Combinations of 2 cationic monomers, 2 hydrophobic monomers, and 2 hydrophilic monomers were selected to obtain an estimated 201,277,440 combinations of anti-Acinetobacter baumannii polymer libraries.
[0067] (2) Hamilton MLSTARlet pipetting robot-assisted automated synthesis: The monomer, chain transfer agent and photosensitizer were prepared in advance into solutions with concentrations of 0.25 mmol / mL, 0.05 mmol / mL and 0.5 mg / mL. Then, the pipetting robot was programmed to add 240 μL, 30 μL and 30 μL of monomer, chain transfer agent and photosensitizer respectively. After the samples were added, the samples were irradiated with green light for 16 h to obtain a polymer with a degree of polymerization of 40.
[0068] (3) Machine learning-assisted recommendation of the structure of the synthetic polymer for screening: Each monomer component of the polymer is represented as a separate graph, with one global virtual node connecting all nodes in each monomer component. Six virtual nodes are introduced to connect the graphs of the six monomer components and the global virtual node, specifying the feed ratio of each monomer component. Then, a pre-trained graph transformer is used to extract polymer features, followed by a readout operator. Finally, a projection head is used to map the layer representation to the attribute space. A query strategy is adopted to select the next sample to be labeled from the sample pool. An ensemble method is used to estimate the polymer performance and model uncertainty, and a confidence cap is applied as the acquisition function to score the remaining samples in the sample pool. Then, samples with higher scores are selected for K-Means clustering (n_clusters=50), and the top three samples from each cluster are selected as the final samples.
[0069] Furthermore, it should be understood that after reading the above description of the present invention, those skilled in the art can make various alterations or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims.
Claims
1. A polymer screening method based on machine learning, characterized in that, The raw material composition of the polymer contains two or more monomers, and each monomer independently carries one or more active groups. The active groups are related to the properties of the polymer to be screened. The machine learning-based polymer screening method uses a graph-based deep learning model to represent polymer molecular structures, with one polymer molecular structure corresponding to one polymer molecular structure graph. Each polymer molecular structure diagram contains a global virtual node and multiple monomer diagrams. Each monomer diagram corresponds to each monomer used to synthesize the polymer molecular structure. For any monomer diagram, non-hydrogen atoms and half-unsaturated bonds used to connect with another monomer are used as nodes. Each monomer diagram is connected to a monomer virtual node. The global virtual node connects all nodes in all monomer diagrams in the polymer molecular structure diagram and each monomer virtual node. The proportion of each monomer in all monomers that make up the polymer molecular structure is reflected by the connection length between the monomer virtual node corresponding to that monomer and the global virtual node. A graph-based deep learning model is trained using polymers with known structures and properties as seed datasets. The model employs an active learning strategy, with inputs including polymer molecular structures, specifically monomer types and the proportion of each monomer in the feed, and outputs including polymer properties. A trained graph-based deep learning model is used to predict and screen the properties of polymers with known structures but unknown properties.
2. The polymer screening method based on machine learning according to claim 1, characterized in that, The active learning strategy includes uncertainty sampling; When predicting polymer properties, graph-based deep learning models score polymers with known structures. This scoring integrates polymer properties and uncertainties, with the uncertainty score determined based on the upper limit of the confidence interval.
3. The polymer screening method based on machine learning according to claim 2, characterized in that, A scoring threshold is set, and polymers with predicted scores greater than or equal to the threshold are clustered. Polymers from different clusters are selected for performance verification and used to train and optimize the graph-based deep learning model to avoid overfitting.
4. The polymer screening method based on machine learning according to claim 1, characterized in that, When training a graph-based deep learning model, a molecular structure similarity threshold is set, and the molecular structure similarity between any polymer training samples is greater than or equal to the set threshold.
5. The polymer screening method based on machine learning according to claim 4, characterized in that, Graph-based deep learning models represent polymers on a two-dimensional reduced-dimensional graph using t-distribution random embedding based on the polymer molecular structure, and use the distance between polymer training samples on the two-dimensional reduced-dimensional graph to reflect the molecular structural similarity between polymer training samples.
6. The polymer screening method based on machine learning according to claim 1, characterized in that, The polymer is a random copolymer synthesized using photoinduced electron / energy transfer reversible addition-fragmentation chain transfer polymerization (PET-RAFT); The monomers include one or more of vinyl monomers, acrylate monomers, and acrylamide monomers; The active group includes one or more of cationic groups, hydrophobic groups, and hydrophilic groups; The polymer properties include antibacterial properties and biocompatibility.
7. The polymer screening method based on machine learning according to claim 1 or 6, characterized in that, The polymer is obtained by PET-RAFT of monomers in the presence of photosensitizers and chain transfer agents and under visible light irradiation.
8. The polymer screening method based on machine learning according to claim 7, characterized in that, The PET-RAFT is performed in solution; The solvent in the solution includes at least one of water, ethanol, methanol, tetrahydrofuran, N,N-dimethylformamide, and dimethyl sulfoxide.
9. The polymer screening method based on machine learning according to claim 1, characterized in that, The machine learning-based polymer screening method also includes automated polymer synthesis and performance verification; A graph-based deep learning model recommends polymer molecular structures based on the current training results. Polymers are then automatically synthesized according to the recommended polymer molecular structures, and performance experiments are conducted. The automatically synthesized polymer molecular structures and their performance are used to train the graph-based deep learning model to optimize the polymer molecular structure recommendations.
10. The polymer screening method based on machine learning according to claim 9, characterized in that, The automated synthesis was performed using a Hamilton MLSTARlet pipetting robot.