A method for designing temperature-resistant water-soluble monomers based on generative artificial intelligence
By constructing datasets and predictive models, and combining generative artificial intelligence strategies, the problems of limited types and low screening efficiency in the design of heat-resistant and water-soluble monomers were solved. Efficient multi-objective screening within the macromolecular structure space was achieved, and monomers with high heat resistance and water solubility were obtained.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN UNIV
- Filing Date
- 2026-03-20
- Publication Date
- 2026-06-23
AI Technical Summary
The types of existing heat-resistant water-soluble monomers are limited, and research and development mainly rely on experience and trial and error, resulting in low screening efficiency. There is a lack of unified and calculable evaluation indicators for water solubility, making it difficult to achieve synergistic screening of heat resistance and water solubility in artificial intelligence-assisted design.
A dataset of polymer temperature resistance and water solubility was constructed, a predictive model of thermal decomposition temperature and monomer solubility was established, a quantifiable water solubility evaluation index was introduced, a generative artificial intelligence strategy was used to generate virtual monomers, and candidate monomers were screened through multi-objective optimization.
The automatic generation and performance prediction of virtual monomers within a larger molecular structural space significantly shortens the R&D cycle, improves screening efficiency, and yields candidate monomers with novel structures and excellent theoretical performance.
Smart Images

Figure CN122266543A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence-enabled new material discovery, specifically involving a method for designing high-temperature resistant, water-soluble monomers using machine learning and generative algorithms. Background Technology
[0002] Copolymer modification with heat-resistant functional monomers is a key strategy for improving the structural stability of water-soluble polymers in high-temperature and high-salt environments. By introducing functional units such as sulfonic acid groups, amide groups, heterocyclic rigid structures, or hydrophobic segments into the polymer molecular chain, the thermal stability, salt resistance, and viscosity retention of the polymer can be improved to a certain extent. Therefore, it is widely used in fields such as high-temperature oil displacement and fracturing.
[0003] However, the variety of existing heat-resistant functional monomers is relatively limited. For a long time, the design of high-temperature water-soluble polymers has mainly focused on the structural modification and copolymerization ratio adjustment of a few commercially available monomers. As oil and gas exploration and development continue to advance into deeper and ultra-deep formations, the thermal stability and salt resistance provided by traditional monomer structures are gradually approaching their limits, making it difficult to meet the requirements for use under higher temperatures and more complex formation conditions. On the other hand, the development of new functional monomers usually relies on empirical design and repeated trial and error in experiments, resulting in long research and development cycles, low screening efficiency, and difficulty in quickly locating candidate structures that combine heat resistance and water solubility in the vast chemical structure space.
[0004] The gradual application of machine learning and generative artificial intelligence technologies in materials design has revolutionized traditional materials research and development models, significantly shortening the development cycle and reducing costs of new materials. However, in the design of temperature-resistant water-soluble monomers, the lack of unified and calculable evaluation indicators for the water solubility of their polymers makes it difficult to introduce them as effective constraints into the model optimization process, resulting in the difficulty of achieving synergistic screening of temperature resistance and water solubility using artificial intelligence methods.
[0005] Given the above challenges, there is an urgent need to develop a heat-resistant water-soluble monomer design method based on generative artificial intelligence. This method would establish quantifiable water-soluble evaluation indicators and introduce them into monomer design, enabling the automatic generation, performance prediction, and multi-objective collaborative screening of virtual monomers within a larger molecular structure space. This would allow for the acquisition of potential monomers with novel structures and excellent theoretical performance, providing new clues and directions for subsequent experimental verification. Summary of the Invention
[0006] This invention addresses the problems of limited types of existing heat-resistant water-soluble monomers, reliance on experience-based trial and error in research and development, and low screening efficiency. In particular, the lack of a unified and calculable evaluation index for water solubility makes it difficult to optimize screening by synergistically using heat resistance and water solubility as constraints in the AI-assisted design process. This invention provides a heat-resistant water-soluble monomer design method based on generative artificial intelligence, establishes a quantifiable water solubility evaluation index and introduces it into monomer design, and realizes the automatic generation of virtual monomers, prediction of heat resistance and water solubility performance, and multi-objective high-throughput screening in a larger molecular structure space.
[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: This invention provides a method for designing heat-resistant water-soluble monomers based on generative artificial intelligence, comprising the following steps: S1: Construct a dataset of the polymer's temperature resistance and water solubility properties.
[0008] The chemical structures and thermal decomposition temperatures of polymer repeating units (or corresponding monomers) were collected from existing polymer databases and related literature. T d We collected monomer solubility (log S) data from small molecule databases and literature; we standardized, deduplicated, and removed outliers from the collected data to form a structure-performance dataset for model training and validation.
[0009] S2: Construct a predictive model for thermal decomposition temperature and a predictive model for monomer dissolution performance.
[0010] The collected polymer chemical structures are represented as SMILES strings and converted into feature vectors that can be processed by machine learning. The structures are then encoded using molecular fingerprints or molecular descriptors. (The text then repeats the process, so the translation will only include the first instance.) T d A quantitative structure-performance (QSPR) prediction model was constructed using log S. The prediction performance of various machine learning algorithms was compared through cross-validation and other methods to determine the coefficient of determination. R 2 Using mean squared error (MSE) as the evaluation metric, the best-performing prediction model was selected, and the selected model was optimized to improve its generalization ability and prediction accuracy.
[0011] S3: Establish a comprehensive evaluation index for water solubility.
[0012] Taking into account monomer solubility (log S), Hildebrand solubility parameter (δ), oil-water partition index (log P), and the sum of hydrogen bond donor and acceptor numbers (HBD+HBA), a comprehensive evaluation function for characterizing polymer water solubility is constructed, and a logistic regression model is trained to learn the weight coefficients of each feature variable, thereby outputting a comprehensive score for polymer water solubility.
[0013] in, S total The comprehensive score for the water solubility of the polymer is given by w1 to w4, which are the weight coefficients of the corresponding terms automatically learned during the logistic regression training process; log S comes from the output of the monomer solubility prediction model constructed in step S2; δ is calculated using the group contribution method combined with SMARTS matching, that is, firstly, SMARTS automatically identifies each structural group in the molecule, and then the solubility parameter of the molecule is obtained by summing the contribution values corresponding to each group; logP, HBD, and HBA can all be automatically calculated using RDKit based on the molecular structure.
[0014] S4: Build a generative virtual monolithic database.
[0015] Based on a generative artificial intelligence strategy, the structure of existing heat-resistant polymers and water-soluble polymers (or the corresponding monomer / repeating unit structures of polymers) is deconstructed to extract the core skeleton and functional group fragments of monomers. The existing water-soluble polymers and heat-resistant polymers are divided into reactive molecular fragments using the Reaction Rule-Based Molecular Fractionation (BRICS) method to obtain a fragment library. Then, the fragments are spliced, assembled and recombined according to the BRICS fragment connection rules to generate a massive number of virtual monomer structures and construct a virtual monomer database.
[0016] S5: Screening for heat-resistant water-soluble polymer monomers.
[0017] The thermal decomposition temperature prediction model obtained in step S2 and the water solubility evaluation index obtained in step S3 are applied to the virtual monomer database in step S4. High-throughput performance calculation and multi-objective optimization screening are performed on the virtual monomers to obtain candidate monomers with both high temperature resistance and good water solubility.
[0018] Furthermore, step S1 also includes: performing statistical processing (e.g., taking the median or average) on the case of multiple sets of performance data corresponding to the same structure; and dividing the data into training set and test set according to a preset ratio (e.g., 4:1).
[0019] Furthermore, the molecular fingerprint mentioned in step S2 refers to Morgan's fingerprint.
[0020] Furthermore, the machine learning algorithms described in step S2 include, but are not limited to, feedforward neural networks (FNN), random forests, gradient boosting trees, support vector regression, etc.
[0021] Furthermore, step S2 can optimize the selected model by methods such as reducing feature dimensions, adding L2 regularization, setting an early stopping mechanism, and enhancing SMILES.
[0022] Further, in step S3, the logistic regression training process is used to learn the weight coefficients and map the linear combination to the final water solubility score and dissolution probability; soluble and insoluble are determined according to the dissolution probability, preferably with a determination threshold of 0.5, that is, when the dissolution probability is greater than or equal to 0.5, it is determined to be soluble, otherwise it is determined to be insoluble.
[0023] Further, in step S4: the fragment library is divided into three parts, namely fragments containing polymerizable groups or structural fragments with polymerizable sites, heat-resistant fragments obtained from the resolution of heat-resistant polymers, and hydrophilic fragments obtained from the resolution of water-soluble polymers; under the premise of following the BRICS linker site matching rules, the above fragments are strategically spliced and assembled so that the generated structure expands the chemical space while maintaining chemical rationality, and ensures that the generated virtual molecule has polymerizable sites to participate in the polymerization reaction as a monomer.
[0024] Furthermore, the specific screening method in step S5 is as follows: In step S2, the virtual monomer structure generated in step S4 is input into the thermal decomposition temperature prediction model and the logistic regression model respectively to obtain the corresponding predicted thermal decomposition temperature value T. d and prediction S total value.
[0025] Furthermore, step S5 also includes constraining the syntheticity of candidate monomers by combining the synthetic accessibility score (SA Score), conducting multi-objective trade-offs and screening among temperature resistance, water solubility and syntheticity, and based on the Pareto front principle, prioritizing the retention of candidate monomers that have advantages in multiple key indicators and are difficult for other structures to surpass as a whole, outputting the monomer structure in the optimal solution set; comparing the output monomer structure with the collected real structures to obtain a novel temperature-resistant and water-soluble monomer with excellent overall performance.
[0026] Furthermore, step S5 may also include structural rationality verification and risk constraint screening of candidate monomers, eliminating structural units that are difficult to control in polymerization reactions, as well as structural units with potential environmental risks, toxicity, or highly irritating functional groups, thereby improving the feasibility and safety of output candidate monomers in actual synthesis and application.
[0027] The beneficial effects of this invention are: 1. This invention establishes a quantifiable comprehensive evaluation index for polymer water solubility, transforming water solubility from an empirical judgment into a calculable and learnable model constraint, thus making the synergistic screening of temperature resistance and water solubility operable and repeatable.
[0028] 2. This invention utilizes a machine learning QSPR model to rapidly predict polymer thermal decomposition temperature and monomer solubility, enabling high-throughput evaluation of a vast number of candidate structures without extensive experimental synthesis, thereby significantly shortening the R&D cycle and reducing screening costs.
[0029] 3. This invention employs a reaction-rule-based molecular fragmentation and recombination generative strategy to construct a virtual monomer library, which effectively expands the chemical space of candidate structures while maintaining chemical rationality and polymerizability, thereby increasing the probability of obtaining novel structural monomers.
[0030] 4. This invention introduces syntheticity constraints such as SA Score and combines them with multi-objective optimization strategies to achieve a comprehensive trade-off between temperature resistance, water solubility and syntheticity, thereby improving the synthetic availability and reliability of the output candidate monomers.
[0031] 5. The process of this invention has a high degree of modularity. By replacing or expanding the dataset and prediction model, the method can be transferred to other solvent systems, other target properties, or other categories of polymerizable monomers for design and screening tasks, thus exhibiting good versatility and scalability. Attached Figure Description
[0032] Figure 1 This is a data distribution diagram in an embodiment of the present invention; Figure 2 This is a diagram showing the selection of machine learning models in an embodiment of the present invention; Figure 3 This is a comparison chart of experimental and predicted values of the thermal decomposition temperature prediction model in an embodiment of the present invention; Figure 4 This is a comparison chart of experimental and predicted values of the solubility prediction model in an embodiment of the present invention; Figure 5 This is a graph showing the relationship between water solubility comprehensive score and solubility probability in an embodiment of the present invention; Figure 6 This is a schematic diagram of the process for constructing a virtual monomer library based on the BRICS break-recombination rule in an embodiment of the present invention; Figure 7 This is a multi-objective optimization diagram based on the Pareto front in an embodiment of the present invention. Detailed Implementation
[0033] The technical solution of the present invention will be further described in detail below with reference to embodiments, but the scope of protection of the present invention is not limited thereto. For those skilled in the art, several changes and improvements can be made without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention.
[0034] Example S1: Construct a dataset of the polymer's temperature resistance and water solubility properties, and perform data distribution analysis (corresponding to...) Figure 1 ).
[0035] Construction of temperature-resistant dataset Obtain the structural information and corresponding thermal decomposition temperature data of the polymer (its repeating units / equivalent structural units) (denoted as ). T d In this embodiment, structural information is represented by the computer-processable string SMILES, and the structure is standardized by removing irrelevant ions / solvents, normalizing the structure, removing duplicate samples, and handling outliers.
[0036] Water-soluble dataset construction Obtain the structural information of the monomer and its solubility data in water (denoted as log S), and perform structural standardization, deduplication, and outlier processing consistent with the heat resistance dataset.
[0037] Data partitioning and distribution visualization The dataset was divided into training and testing sets in a 4:1 ratio.
[0038] like Figure 1 The image shown is a data distribution diagram in this embodiment, illustrating the temperature resistance data used for model construction. T d The distribution of solubility data (log S) within the numerical range is used to assess the data coverage and sample distribution characteristics.
[0039] S2: Construct a predictive model for thermal decomposition temperature and a predictive model for monomer dissolution performance.
[0040] Machine learning model selection To achieve rapid performance prediction of candidate monomers, this embodiment establishes quantitative structure-property (QSPR) models for both temperature resistance and water solubility. Since different algorithms have varying adaptability to different datasets, this embodiment first performs model screening.
[0041] Structural feature construction Transform the SMILES structure into input features that can be used for machine learning. Features may include molecular fingerprints and / or molecular descriptors.
[0042] Candidate Model Set and Evaluation Metrics Multiple candidate machine learning models were constructed, and their predictive performance was compared. Candidate models include, but are not limited to: random forest regression, gradient boosting tree regression, support vector regression, and feedforward neural network regression.
[0043] Evaluation metrics can include the coefficient of determination (R²) and mean squared error (MSE), combined with cross-validation to assess model stability.
[0044] Model screening results like Figure 2 The diagram shown illustrates the selection of machine learning models in this embodiment. FNN performed best on the dataset, therefore it was chosen as the optimal algorithm for the subsequent temperature resistance prediction model and solubility prediction model.
[0045] Construction of thermal decomposition temperature prediction model (corresponding) Figure 3 ) After completing the model selection, a thermal decomposition temperature prediction model was established for the temperature resistance index.
[0046] Model training Using the optimal model determined in step 2, with the structural feature vector as input and the thermal decomposition temperature... T d The system is trained to output values. During training, overfitting risk can be reduced and generalization ability improved through hyperparameter optimization, regularization, and early stopping strategies.
[0047] Model Validation and Results Demonstration The model is evaluated on the test set, and the results of the comparison between the predicted values and the experimental values are output.
[0048] like Figure 3 The figure shown is a comparison between the experimental and predicted values of the thermal decomposition temperature prediction model in this embodiment. The correspondence between the experimental and predicted values demonstrates the model's prediction accuracy and fit. This model predicts R0. 2 The score reached 0.98, indicating that the trained machine learning model can accurately predict the thermal decomposition temperature.
[0049] Construction of monomer solubility prediction model (corresponding) Figure 4 ) To address water solubility-related properties, this embodiment establishes a small molecule solubility prediction model to output the predicted log S value of candidate monomers.
[0050] Model training Using the optimal model determined in step 2, we train and optimize the model with the Morgan fingerprint of the monomer structure as the input vector and the solubility log S as the output value.
[0051] Model Validation and Results Demonstration Output the comparison between predicted and experimental values on the test set.
[0052] like Figure 4 The figure shown is a comparison of the experimental and predicted values of the solubility prediction model in this embodiment, used to characterize the model's ability to predict solubility. This model predicts R... 2 A score of 0.99 indicates that the trained machine learning model can accurately predict the solubility of small molecules.
[0053] S3: Establish a comprehensive evaluation index for water solubility.
[0054] To quantify the water solubility of polymers, this embodiment uses monomer solubility (log S), Hildebrand solubility parameter (δ), oil-water partition index (log P), and the sum of hydrogen bond donor and acceptor numbers (HBD+HBA) as inputs to characterize water solubility. A logistic regression model is used to learn the weights of each factor to obtain a comprehensive water solubility score S. total Wherein, log S is output by the monomer solubility prediction model trained in step S2, δ is calculated using the group contribution method combined with SMARTS identification and summation, and log P, HBD, and HBA can be automatically calculated based on molecular structure using tools such as RDKit. The logistic regression training process is used to learn the weight coefficients and map the linear combination to the final water solubility score and dissolution probability. Solubility / insolubility is determined according to the dissolution probability. In this embodiment, a threshold of 0.5 is used, that is, when the dissolution probability is greater than or equal to 0.5, it is judged as soluble; otherwise, it is judged as insoluble. Figure 5 As shown in this embodiment, S total The graph showing the relationship between Statal and solubility probability demonstrates a strong correlation between Statal and the soluble / insoluble determination.
[0055] S4: Generative Virtual Monolith Library Construction: Virtual Monolith Library Construction Based on BRICS Fragmentation-Recombination Rules (corresponding to...) Figure 5 ).
[0056] To expand the monomer structure space and improve the discovery efficiency of novel monomers, this embodiment uses the BRICS break-recombination rule to construct a virtual monomer library.
[0057] Source of the clip A set of known structures with heat resistance characteristics (derived from heat-resistant polymers) and a set of known structures with hydrophilic characteristics (derived from water-soluble polymers) were selected as the matrix structure sources for fragment generation.
[0058] BRICS Fracture and Fragment Library Establishment The parent structure is broken according to BRICS rules to obtain connectable fragments, and the connection site types and connection compatibility information of the fragments are recorded. The fragment library contains: fragments for providing temperature resistance characteristics; fragments for providing hydrophilic / water-soluble characteristics; and fragments containing polymerizable groups or structural fragments with polymerizable sites to ensure that the generated molecules are polymerizable.
[0059] Monolithic structure generation and structural verification The fragments are combined and reconstructed according to the BRICS connection rules to generate candidate monomer structures, and the structure validity is verified and filtered, including but not limited to: valence state verification, structure legality verification, removal of duplicate structures, removal of missing aggregateable sites or obviously unreasonable structures, etc.
[0060] like Figure 6 The diagram shown illustrates the process of constructing a virtual monomer library based on the BRICS break-recombination rule in this embodiment. This process yields a large-scale virtual monomer library, providing a set of candidate structures for subsequent high-throughput prediction and screening.
[0061] S5: Screening for heat-resistant, water-soluble monomers Multi-objective optimization screening based on Pareto front (corresponding) Figure 7 ) For each candidate monomer in the virtual monomer library, this embodiment performs rapid prediction of its temperature resistance and water solubility related properties, and conducts multi-objective optimization screening to obtain candidate monomers that combine high temperature resistance and good water solubility. The thermal decomposition temperature prediction model is then applied to each candidate monomer to obtain... T d Predicted values, logistic regression model, yield S total Predicted values. Additionally, a syntheticity index (SA Score) is calculated to improve the feasibility of the screening results.
[0062] Multi-objective optimization strategy Using temperature resistance and water solubility-related indicators as optimization objectives, for example: maximizing T d Predicted value and maximizing S total Predicted values and performed multi-objective optimization under constraints of syntheticity or structural rationality.
[0063] like Figure 7 The diagram shown is a multi-objective optimization graph based on the Pareto front in an embodiment of the present invention. The Pareto optimality principle is used to retain the set of structures that cannot be simultaneously surpassed by other candidate entities on all objectives, thus forming a candidate entity set.
[0064] Output It outputs candidate monomer structures on the Pareto front and their corresponding prediction metrics, and can further perform candidate ranking, clustering to remove redundancy, and subsequent verification according to application requirements.
[0065] In summary, this invention addresses the structural design problem of temperature-resistant, water-soluble monomers by constructing an integrated intelligent design process encompassing data acquisition, performance prediction, generative structure construction, and multi-objective optimization screening. By establishing thermal decomposition temperature and solubility prediction models, and combining a BRICS-based virtual monomer library construction method with a Pareto front multi-objective optimization strategy, synergistic optimization and high-throughput screening of candidate monomers' temperature resistance and water solubility are achieved. This method effectively overcomes the problems of low efficiency, long cycle, and high cost associated with traditional trial-and-error monomer design, improves the specificity of structural design and the reliability of prediction results, and provides a systematic and scalable technical path for the research and development of temperature-resistant, water-soluble functional monomers.
[0066] It should be noted that, without departing from the core ideas of this invention, those skilled in the art can reasonably adjust or equivalently replace the data scale, model type, feature selection method, optimization algorithm, multi-objective weight settings, and screening thresholds according to specific application needs, all of which can achieve the technical effects described in this invention. The specific algorithm models, parameter ranges, and structural combinations involved in the above embodiments are merely illustrative examples of preferred embodiments and are not intended to limit the scope of protection of this invention.
[0067] Therefore, any technical means that are the same as or equivalent to the technical solution of this invention, or any equivalent transformations, substitutions or improvements made based on the technical concept of this invention, shall fall within the protection scope of this invention.
Claims
1. A method for designing heat-resistant water-soluble monomers based on generative artificial intelligence, characterized in that, It includes the following steps: S1: Constructing a dataset of polymer temperature resistance and water solubility properties Collect the chemical structure and thermal decomposition temperature (Td) data of existing polymer repeating units or corresponding monomers, and collect monomer solubility (logS) data; standardize the collected data and form a structure-performance dataset for model training and validation; S2: Constructing a thermal decomposition temperature prediction model and a monomer dissolution performance prediction model The collected polymer chemical structures are represented as SMILES strings and converted into feature vectors that can be processed by machine learning. The structures are then encoded using molecular fingerprints or molecular descriptors. (The text then repeats the process, so the translation will only include the first instance.) T d A quantitative structure-performance prediction model was constructed using log S. The prediction performance of various machine learning algorithms was compared using methods such as cross-validation to determine the coefficient of determination. R 2 Using mean squared error (MSE) as the evaluation metric, the best-performing prediction model was selected, and the selected model was optimized to improve generalization ability and prediction accuracy. S3: Establish a comprehensive evaluation index for water solubility Taking into account monomer solubility log S, Hildebrand solubility parameter δ, oil-water partition index log P, and the sum of hydrogen bond donors and acceptors HBD+HBA, a comprehensive evaluation function for characterizing polymer water solubility is constructed, and a logistic regression model is trained to learn the weight coefficients of each feature variable, thereby outputting a comprehensive score for polymer water solubility. , in, S total For the comprehensive water solubility score, w1 to w4 are the weight coefficients of the corresponding terms automatically learned during the logistic regression training process; log S comes from the output of the monomer solubility prediction model constructed in step S2; δ is calculated using the group contribution method combined with SMARTS matching, that is, firstly, SMARTS automatically identifies each structural group in the molecule, and then the solubility parameter of the molecule is obtained by summing the contribution values corresponding to each group; logP, HBD, and HBA can all be automatically calculated based on the molecular structure using RDKit; S4: Building a Generative Virtual Monolithic Database Based on a generative artificial intelligence strategy, the structure of existing heat-resistant polymers and water-soluble polymers, or the corresponding monomers and repeating units of polymers, is deconstructed to extract the core skeleton and functional group fragments of monomers. Using a molecular fragmentation method based on reaction rules, the existing water-soluble polymers and heat-resistant polymers are divided into reactive molecular fragments to obtain a fragment library. Then, according to the connection rules, the fragments are spliced, assembled and recombined to generate a large number of virtual monomer structures and construct a virtual monomer database. S5: Screening for heat-resistant water-soluble polymer monomers The thermal decomposition temperature prediction model obtained in step S2 and the water solubility evaluation index obtained in step S3 are applied to the virtual monomer database in step S4. High-throughput performance calculation and multi-objective optimization screening are performed on the virtual monomers to obtain candidate monomers with both high temperature resistance and good water solubility.
2. The method according to claim 1, characterized in that, Step S1 also includes: performing statistical processing on the case of multiple sets of performance data corresponding to the same structure; and dividing the data into training set and test set according to a preset ratio.
3. The method according to claim 1, characterized in that, The molecular fingerprint mentioned in step S2 refers to Morgan's fingerprint.
4. The method according to claim 1, characterized in that, The machine learning algorithm mentioned in step S2 includes, but is not limited to, one of the following: feedforward neural network (FNN), random forest, gradient boosting tree, and support vector regression.
5. The method according to claim 1, characterized in that, Step S2 optimizes the selected model by reducing feature dimensions, adding L2 regularization, setting an early stopping mechanism, or using SMILES enhancement.
6. The method according to claim 1, characterized in that, In step S3, the logistic regression training process is used to learn the weight coefficients and map the linear combination to the final water solubility score and dissolution probability. Solubility and insolubility are determined according to the dissolution probability. The determination threshold is set to 0.5, that is, when the dissolution probability is greater than or equal to 0.5, it is determined to be soluble, otherwise it is determined to be insoluble.
7. The method according to claim 1, characterized in that, In step S4: the fragment library is divided into three parts: fragments containing polymerizable groups or structural fragments with polymerizable sites, heat-resistant fragments obtained from the resolution of heat-resistant polymers, and hydrophilic fragments obtained from the resolution of water-soluble polymers. Under the premise of following the BRICS linker site matching rules, the above fragments are strategically spliced and assembled so that the generated structure expands the chemical space while maintaining chemical rationality, and ensures that the generated virtual molecule has polymerizable sites to participate in the polymerization reaction as a monomer.
8. The method according to claim 1, characterized in that, The specific screening method for step S5 is as follows: In step S2, the virtual monomer structure generated in step S4 is input into the thermal decomposition temperature prediction model and the logistic regression model respectively to obtain the corresponding predicted thermal decomposition temperature value. T d and prediction S total value.
9. The method according to claim 1, characterized in that, Step S5 also includes constraining the syntheticity of candidate monomers by combining the synthetic accessibility score, conducting multi-objective trade-off screening between temperature resistance, water solubility and syntheticity, and based on the Pareto front principle, prioritizing the retention of candidate monomers that have advantages in multiple key indicators and are difficult to be surpassed by other structures as a whole, and outputting the monomer structure in the optimal solution set. By comparing the output monomer structure with the collected real structures, a novel, temperature-resistant, water-soluble monomer with excellent overall performance was obtained.
10. The method according to claim 1, characterized in that, Step S5 may also include verifying the structural rationality and screening for risk constraints of candidate monomers, eliminating structural units that are difficult to control in the polymerization reaction, as well as structural units with potential environmental risks, toxicity, or strong irritant functional groups.