Method for predicting performance of high polymer material based on machine learning of small sample
The data is expanded through molecular dynamics simulation and improved interpolation method, combined with the XGBoost model, the problem of insufficient data in the performance prediction of polymer materials is solved, and the glass transition temperature of dissolved polystyrene butadiene rubber is achieved quickly and accurately predicted, and the performance prediction of other polymer materials is extended to the performance prediction of other polymer materials.
Patent Information
- Application Number
- CN202410951701.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-16
- Publication Date
- 2025-08-08
AI Technical Summary
When predicting the performance of polymer materials, especially the glass transition temperature of dissolved polystyrene butadiene rubber, the prior art has problems such as insufficient data volume and difficulty in providing targeted explanations in generalized data sets, resulting in limited development of accurate prediction models.
The original data was obtained through molecular dynamics simulation, data augmentation was performed using improved nearest neighbor interpolation method, and quantitative structure-effect relationship was constructed in combination with clustering analysis and XGBoost model to achieve efficient prediction of the performance of dissolved styrene butadiene rubber.
The rapid and accurate prediction of the glass transition temperature of dissolved polystyrene butadiene rubber was achieved under limited samples, solving the problems of long experimental cycles and high cost, and expanding to the performance prediction of other polymer materials.
Smart Images

Figure CN120452582A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of application of machine learning technology, and specifically relates to a method for predicting polymer material properties based on small sample machine learning. Background Art
[0002] As the most widely produced and consumed synthetic rubber worldwide, solution-polymerized styrene-butadiene rubber (SBR) exhibits excellent wear resistance, heat resistance, and aging resistance, and its structural composition is closely related to these properties. The glass transition temperature (GTP), one of its key processing parameters, is typically measured experimentally using differential scanning calorimetry (DSC) and dynamic mechanical analysis (DMA). However, achieving the desired GTP by varying the structural unit content requires a lengthy trial-and-error process, significantly hindering R&D progress.
[0003] With the advent of the big data era, the enormous potential of combining artificial intelligence with traditional scientific research has given rise to the concept of "AI for Science." As a key branch of artificial intelligence, machine learning has been rapidly applied to materials science research within the context of the Materials Genome Project and the new paradigm of "data-driven innovation." With the assistance of machine learning, materials design has become more convenient, enabling more efficient development of quantitative structure-activity relationships (QSPRs) for predicting material properties.
[0004] While many previous studies have explored the application of machine learning to predict material properties, most of these studies rely on relatively generalized datasets from databases or literature to generate large amounts of raw data. For example, researchers collected polymer data from the PoLyInfo database and used various feature representations to predict glass transition temperatures. (J. Chem. Inf. Model. 2021, 61, 5395–5413) Similarly, a research team developed a predictive model for the dielectric constant by aggregating values for 738 polymers from existing literature. (npj Comput. Mater. 2020, 6, 1–9) However, due to the generalization of the datasets, these works often struggle to provide targeted insights when studying specific materials. Furthermore, using experiments or simulations to provide raw data often results in insufficient data, and these limitations hinder the development of accurate predictive models. Obtaining highly targeted and sufficiently large datasets and using them to accurately predict polymer properties remains an urgent challenge. Summary of the Invention
[0005] The purpose of the present invention is to solve the difficulties existing in the above-mentioned prior art and provide a method for predicting the properties of polymer materials based on small sample machine learning. Small sample data of the glass transition temperature of solution-polymerized styrene-butadiene rubber (SSBR) is obtained through molecular dynamics simulation, and an improved nearest neighbor interpolation method is used for data enhancement. The contents of the four structural units of solution-polymerized styrene-butadiene rubber are used as input, and a machine learning model is used to achieve efficient prediction of target performance.
[0006] The present invention is achieved through the following technical solutions:
[0007] A method for predicting polymer material properties based on small sample machine learning, the method comprising:
[0008] Obtain raw data through molecular dynamics simulation;
[0009] The original data is expanded by interpolation to obtain an expanded data set;
[0010] Perform cluster analysis on the data set to obtain a class-balanced data set;
[0011] The dataset is fed into the machine learning prediction model for training and prediction.
[0012] Furthermore, the operation of providing raw data through molecular dynamics simulation includes:
[0013] By adjusting the content of structural units, multiple solution-polymerized styrene-butadiene rubber system models with different structural unit contents were obtained.
[0014] The temperature-volume method is used to calculate each system model to obtain the corresponding original data.
[0015] Furthermore, the solution-polymerized styrene-butadiene rubber system model includes 10 identical solution-polymerized styrene-butadiene rubber single chains, each chain includes 60 repeating structural units; and each piece of the original data includes four characteristic variables and one label value.
[0016] Furthermore, the step of expanding the original data by interpolation method includes:
[0017] Establish an interpolation grid and use a random number generator to insert a set number of new data points near the original data points;
[0018] Calculate the label value of the new data point to obtain the expanded data set.
[0019] Furthermore, calculating the label value of the new data point includes:
[0020] The four-dimensional features of the original data are reduced to two dimensions as its spatial coordinates through principal component analysis;
[0021] The spatial relationship between the original data is calculated through the semivariogram function;
[0022] For new data points, the weight is determined by combining the spatial data between them and the original data;
[0023] According to the weight, the weighted sum is used to obtain the label value of the new data point.
[0024] Furthermore, the label value of the new data point is obtained using the following formula:
[0025]
[0026] where λ i (i=1,…,n) is the weight, f(X i ) is the label value of the original data point in the dataset.
[0027] Furthermore, the weight is obtained by the following formula:
[0028]
[0029] Where μ is the Lagrange multiplier, γ(h) is the semivariogram equation, and λ k (k=1,…,n) is the weight, X i is the original data point, X k is another original data point, X j is a new data point, γ(x i -x k ) represents the semivariance equation of each known point and other known points; γ(x i -x j ) represents the semivariance equation between each known point and a new point to be estimated.
[0030] Furthermore, the K-means method is used to obtain the number of data points in each category; and the ratio of the number of data points in each category is compared to see whether it exceeds 3:1;
[0031] The ratio of the number of data points of each category does not exceed 3:1, and there is no category imbalance in the dataset;
[0032] The ratio of the number of data points in each category exceeds 3:1, and the dataset is class-imbalanced. The SMOTE method is used for quadratic interpolation to obtain a class-balanced dataset.
[0033] Furthermore, the operations of inputting the data set into the machine learning prediction model for training and performing prediction include:
[0034] The class-balanced dataset is divided into training set and test set according to the ten-fold cross-validation method;
[0035] The training set is input into the machine learning prediction model for training, and the test set is used for verification, and finally a trained model is obtained;
[0036] The contents of the four structural units of solution-polymerized styrene-butadiene rubber are input into the trained prediction model, and the trained prediction model outputs the prediction results of the polymer material properties.
[0037] Furthermore, the machine learning prediction model adopts the XGBoost model, and its setting parameters include: a learning rate of 0.08, a maximum tree depth of 13, and a number of trees of 120.
[0038] Compared with prior art, the present invention has the beneficial effects as follows: in the case of a finite sample of less than 50 sample size, original training data is provided for machine learning by molecular dynamics simulation, not only the specific properties of specific polymer materials can be studied specifically, the problems such as the longer cycle of actual experiment, higher cost can also be solved. In addition, the present invention improves the method for calculating label values in the nearest neighbor interpolation method, effectively overcomes the situation that the label value large area that is easy to occur when using the nearest neighbor interpolation method is identical, lacks continuity between data points, and balances between computational efficiency and interpolation accuracy, so as to provide sufficient and accurate data sets for subsequent prediction. Finally, the trial and error method in place of the experiment is realized, the rapid and accurate prediction of the rubber material target properties such as solution-polymerized butadiene styrene rubber (i.e. SSBR) is completed, and other polymer materials or other performances can be extended to. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 Technical flow chart of the present invention.
[0040] Figure 2 The Elbow method was used to determine the number of clusters.
[0041] Figure 3 Determine XGBoost model parameters through learning curve.
[0042] Figure 4 Comparison chart of experimental simulation values and XGBoost model prediction values.
[0043] Figure 5 R in 900 repeated experiments 2 changes in the situation.
[0044] Figure 6 Comparison chart of machine learning prediction results and simulation calculation results after selecting new structural unit content. DETAILED DESCRIPTION
[0045] The present invention is further described in detail below with reference to the accompanying drawings.
[0046] The technical solution of the present invention is a method for predicting the properties of polymer materials based on small sample machine learning. The method provides raw data through molecular dynamics simulation, expands the data through improved nearest neighbor interpolation method, and trains the XGBoost model to construct a quantitative structure-activity relationship between the content of the four structural units of SSBR and the glass transition temperature, thereby achieving efficient prediction of the glass transition temperature.
[0047] like Figure 1 As shown, the specific steps of the method of the present invention are as follows:
[0048] Step 1: Molecular dynamics simulation to obtain raw data
[0049] By regulating the content of structural units, namely the content of four structural units: styrene, 1,2-butadiene, cis-1,4-butadiene and trans-1,4-butadiene, a solution-polymerized styrene-butadiene rubber (SSBR) system model was constructed using all-atom molecular dynamics simulation, and the temperature-volume method was used to calculate each system model separately, and finally the original data corresponding to each system model was obtained.
[0050] Specifically, 40 different SSBR system models were created by adjusting the structural unit content. The molar ratio of each structural unit in the model varied, with the styrene molar content ranging from 5% to 50%, ensuring that the total molar content of the four structural units equaled 100%. All SSBR system models contained 10 identical SSBR chains, each containing 60 repeating structural units.
[0051] The temperature-volume method uses the volume change of a system at different temperatures to determine the glass transition temperature. Specifically, volume data is collected at each temperature point, and the point with the maximum volume change is the glass transition temperature. Each raw data entry contains four characteristic variables (i.e., the contents of the four structural units) and a label value (i.e., the corresponding glass transition temperature).
[0052] Step 2: Expand the original data by interpolation to obtain the data set
[0053] An interpolation grid is established, and a random number generator is used to insert new data points that are no less than 5 times the original data points near the original 40 original data points. Furthermore, the label values of the new data points are calculated using an improved nearest neighbor interpolation method to obtain an expanded data set.
[0054] Specifically, during the interpolation grid implementation, a random number generator is created for each feature. The generator follows a normal distribution with the specific value of each feature as the mean and the mean of each feature multiplied by a constant factor as the variance. The number of data generated by the random number generator is the interpolation multiplier in the final interpolation grid.
[0055] Furthermore, the nearest neighbor interpolation method is improved to calculate the label value. The specific process is as follows:
[0056] The four-dimensional features of the original data are reduced to two dimensions using principal component analysis, serving as their spatial coordinates. The spatial relationships between the original data points are calculated using the semivariogram function. For new data points, the weights are determined based on the spatial relationship between them and the original data points, and a weighted sum is taken to obtain a revised label value. Specifically, the label value of the new data point is obtained by weighted summing the label values of all currently known original data points. The specific correction process is as follows:
[0057] There are n original data points X i (i=1,…,n), X j (j=0,…,m) is the new data point to be corrected, f(X j ) is the label value of the new data point, then the correction formula is:
[0058]
[0059] where λ i (i=1,…,n) is the weight, f(X i ) is the label value corresponding to the original data point in the dataset.
[0060] Furthermore, the weight information is calculated by the Lagrange multiplier method, which is calculated as follows:
[0061]
[0062] Where μ is the Lagrange multiplier, γ(h) is the semivariogram equation, and λ k (k=1,…,n) is the weight, X i is the original data point, X k is another original data point, X j is a new data point, γ(x i -x k ) represents the semivariance equation of each known point and other known points; γ(x i -x j ) represents the semivariance equation between each known point and a new point to be estimated.
[0063] After solving, we can get the weight information. Substitute the obtained weight information into Formula 1 to obtain the corrected label value.
[0064] Ultimately, this step will expand the simulated 40 data sets to 800 sets of data.
[0065] Step 3: Perform cluster analysis on the data in the dataset
[0066] When there is a significant difference in the number of different categories in a dataset, known as imbalance, machine learning prediction models tend to prioritize learning features from the more numerous categories and may ignore samples from the minority categories, thus affecting model performance. Therefore, before inputting the data into the model for training, cluster analysis is performed on the dataset to detect whether there is a class imbalance problem. The cluster analysis method uses the K-means method to obtain the data volume for each category, that is, the number of data points in each cluster. The number of data points in different clusters is compared to see if there is a significant difference. When the ratio of the data volume between categories exceeds 3:1, it can be determined that the data has class imbalance.
[0067] Specifically, the Elbow method is first used to determine the number of clusters to be 2. The Elbow method gradually increases the number of clusters K and calculates the mean square error (i.e., Distortion Score) within each cluster under each K value, and plots the relationship between the mean square error within the cluster and the K value. Figure 2 As shown in the figure, observe the trend of the intra-cluster mean square error and find the location where the curve shows a clear bend, known as the "elbow." The K value corresponding to this elbow is the optimal number of clusters. Then, use the K-means method to cluster and see if the ratio of the data volume between the two clusters is less than 3:1. If there is class imbalance, use the SMOTE method for quadratic interpolation. If there is no class imbalance, directly input the data into the machine learning prediction model.
[0068] Step 4: Input the data set into the machine learning prediction model for training and prediction
[0069] Specifically, the XGBoost model was used for prediction: the expanded 800 data sets were divided into training sets and test sets for training according to the ten-fold cross-validation method; the contents of the four structural units of solution-polymerized styrene-butadiene rubber (SSBR) were used as the input for the XGBoost model prediction, and the glass transition temperature was finally output.
[0070] S4-1. Specifically, the 10-fold cross validation method is applied: the data is randomly divided into 10 parts, and 1 part is selected as the test set each time, and the remaining 9 parts are used as the training set. The cycle is repeated, and the average value of the ten experimental results is used as the final evaluation standard. The evaluation indicators are generally RMSE and R 2 .
[0071] S4-2, select R 2 Plot a learning curve for an indicator, such as Figure 3As shown in Figure 2, the specific parameters of the XGBoost model are set: the learning rate is 0.08, the maximum tree depth is 13, and the number of trees is set to 120. 2 It reflects the goodness of fit of the model, with values close to 1 indicating a better fit.
[0072] Among them, the training process of the XGBoost model is as follows:
[0073] A random seed is selected, and the random seed is obtained by a random number generator. Specifically, the present invention selects a random number in the range of 1 to 3000.
[0074] The model evaluates the prediction error by calculating the current loss.
[0075] For each data point, the gradient and second-order derivative of the loss function are calculated to construct a new decision tree. The maximum tree depth is set to 13 to prevent each decision tree from becoming too complex and causing the model to overfit. Each time a new tree is constructed, the model attempts to split across all possible values of each feature, selecting the split point that minimizes the loss function.
[0076] New trees are added to the model, their contribution to the model is controlled via the learning rate, and the model's predictions are updated.
[0077] The above process will be iterated multiple times, gradually adding more trees to improve the model performance until the predetermined stopping condition is met, that is, the set number of trees is reached 120 or the training error no longer decreases significantly.
[0078] S4-3. Input the contents of the four structural units of the solution-polymerized styrene-butadiene rubber into the above-trained prediction model, and the final output result is the glass transition temperature predicted by the model.
[0079] Step 5: Verify the robustness and universality of the prediction results
[0080] S5-1, select different random seeds from the random number generator and repeat the model prediction to ensure the robustness of the prediction method of the present invention. Specifically, 900 repeated prediction experiments were performed to obtain 900 R 2 , all R 2 All of them are kept in a relatively high range, that is, the present invention has robustness.
[0081] S5-2. Design verification data, i.e., multiple sets of different structural unit content combinations, and simultaneously perform machine learning model predictions and molecular dynamics simulations to compare the glass transition temperature results. The specific verification process involves inputting the verification data into the XGBoost model to obtain a prediction result; then reusing the verification data using molecular dynamics simulations to obtain a calculation result. Comparing the two results separately, the comparison results for multiple sets of verification data are very close, indicating that the present invention has high universality.
[0082] Below in conjunction with accompanying drawing and embodiment, the present invention is described in further detail.In the present embodiment, using solution-polymerized styrene-butadiene rubber as the polymer material object, its four kinds of structural units (styrene, 1,2-butadiene, cis-1,4-butadiene and trans-1,4-butadiene) content are predicted to the influence of its glass transition temperature.
[0083] 1. An example of obtaining data through molecular dynamics simulation is as follows:
[0084] According to step 1, specifically, Materials Studio software is used to construct a solution-polymerized styrene-butadiene rubber (SSBR) system model by random copolymerization. By adjusting the content of structural units (i.e., the content of styrene, 1,2-butadiene, cis-1,4-butadiene, and trans-1,4-butadiene), a total of 40 solution-polymerized styrene-butadiene rubber (SSBR) models with different styrene and butadiene contents are obtained, and the molar styrene content ranges from 5% to 50%. All models contain 10 identical solution-polymerized styrene-butadiene rubber (SSBR) single chains, each containing 60 repeating units. The molecular weight of each SSBR single chain is between 3000 and 5000 g·mol -1 between.
[0085] After obtaining the all-atom model of solution-polymerized styrene-butadiene rubber (SSBR), a constant temperature and pressure (NPT) ensemble was used to equilibrate all systems, and periodic boundary conditions were applied in the x, y, and z directions. The equations of motion were integrated using the Velocity-Verlet algorithm, and an annealing process was implemented with a time step of 1 fs. Specifically, after an initial short equilibration period, the system temperature was raised to 400 K while maintaining the pressure at 0.1 MPa, and an NPT equilibration period of 5 ns was performed to ensure sufficient relaxation of the molecular chains. Subsequently, the temperature was lowered to 298 K, and an NPT equilibration period of another 5 ns was performed. The annealing process was repeated twice to ensure complete relaxation of the solution-polymerized styrene-butadiene rubber (SSBR) system. Subsequently, the system was maintained at 298 K and 0.1 MPa under the NPT ensemble for 20 ns to stabilize parameters such as density and non-bonded interactions. Ultimately, 40 simulation data points, namely the original data, were obtained. Each piece of raw data includes four characteristic variables (styrene unit content, 1,2-butadiene unit content, cis-1,4-butadiene unit content, and trans-1,4-butadiene unit content) and a label value (glass transition temperature).
[0086] 2. The data is expanded by interpolation to obtain a data set as follows:
[0087] The nearest neighbor interpolation method is used to establish an interpolation grid near the original 40 sets of original data points. If the interpolation factor is too large, the correlation between the expanded data set and the original data set will decrease, while if it is too small, the purpose of expansion will not be achieved. In this embodiment, the interpolation factor is determined to be 20.
[0088] According to step 2, the improved nearest neighbor interpolation method is used to calculate the label values of all inserted new data points, and finally 800 groups of data points are obtained, that is, the expanded data set is obtained.
[0089] Specifically, the four-dimensional features of the original data are reduced to two dimensions through principal component analysis, which serve as its spatial coordinates. The spatial relationships between the original data points are calculated using methods such as the semivariogram. For each new data point in the dataset, a weight is determined based on the spatial relationship between it and the original data point. A new label value is then obtained through weighted summation. The specific correction process is as follows:
[0090] Assume there are n known points X i (i=1,…,n), X j (j=0,…,m) is the new data point to be corrected, f(X j ) is the label value of the new data point, and the correction formula is:
[0091]
[0092] Among them, λ i (i=1,…,n) is the weight, f(X i ) is the label value corresponding to the new data point in the dataset.
[0093] The weights are calculated using the Lagrange multiplier method, and the specific process is as follows:
[0094] According to the Lagrange multiplier method, that is:
[0095]
[0096] Where μ is the Lagrange multiplier, γ(h) is the semivariogram equation, and λ k (k=1,…,n) is the weight, X i is the original data point, X k is another original data point, X j is a new data point, γ(x i -x k ) represents the semivariance equation of each known point and other known points; γ(x i -x j ) represents the semivariance equation between each known point and a new point to be estimated.
[0097] Formula 2 is expanded to:
[0098]
[0099] Select its linear form, that is:
[0100] γ(h)=C j +C1h
[0101] Among them C j represents the semivariance value at distance h=j, C1 reflects the spatial autocorrelation of the data, and h represents the spatial distance vector between two data points. Substituting the above formula into the Lagrange equation yields:
[0102]
[0103] After solving, the weight information can be obtained.
[0104] Substitute the obtained weight information into Formula 1 to obtain the corrected label value.
[0105] 3. An example of cluster analysis of data in a data set is as follows:
[0106] According to step 3, specifically, different cluster numbers K are selected, and the intra-cluster mean square error (Distortion Score) under the corresponding K value is calculated, and the relationship between the intra-cluster mean square error and the K value is plotted, as shown in Figure 2 As shown in the figure, observe the trend of the intra-cluster mean squared error (MSE) and find the location where the curve shows a clear bend, i.e., the "elbow." The K value corresponding to this elbow is the optimal number of clusters. At this point, increasing the K value does not significantly improve the intra-cluster mean squared error. The figure shows that in this embodiment, the number of clusters K is set to 2.
[0107] In this embodiment, the K-means method is used to obtain two clusters containing 413 and 387 data points respectively. The ratio between the two categories of data is less than 3:1, so there is no category imbalance.
[0108] 4. An example of inputting a data set into a machine learning model for training is as follows:
[0109] According to step 4, set up the XGBoost model, such as Figure 3 As shown, the specific parameters are: learning_rate is set to 0.08, max_depth is set to 13, and number of trees is set to 120. The data set is randomly divided into 10 subsets in equal proportions using ten-fold cross validation. In this embodiment, automatic random division is achieved through code. One subset is selected as the test set and the remaining 9 subsets are selected as the training set, and the model is trained ten times. After each training, the model performance index is calculated, and finally the ten results are averaged to obtain the trained model. The R of the final model in this embodiment is 2The value is 0.9710 and the RMSE is 0.736, which shows that the model has a good prediction effect.
[0110] Furthermore, Table 1 gives the model prediction values and the simulation values obtained through experiments under 5 structural unit combinations as examples. Figure 4 , it can be observed that the predicted values predicted by the model are very close to the simulated values.
[0111] Table 1
[0112]
[0113] 5. Examples for verifying robustness and universality are as follows:
[0114] According to S5-1 in step 5, different random numbers are obtained by the random number generator. In this embodiment, random numbers in the range of 1 to 3000 are selected and the experiment is repeated 900 times. Figure 5 As shown in the final prediction results of repeated experiments, the indicator R 2 It remains within the range of 0.9785 to 0.9933, indicating that the performance of the machine learning model prediction method is not an accidental result of any specific data, proving the robustness of the prediction method proposed in this invention.
[0115] According to S5-2 of step 5, 4 new groups of four structural unit contents were selected, and machine learning model prediction and molecular dynamics simulation calculation were performed simultaneously to compare the results of glass transition temperature. Figure 6 As shown in the figure, the specific verification process is to input the new structural unit content into the XGBoost model to obtain the prediction result; at the same time, the new structural unit content is calculated using molecular dynamics simulation to obtain the calculation result; and then compare them respectively.
[0116] In this embodiment, the contents of the four new structural units and their simulated and predicted values are shown in Table 2, where T g -MD refers to the value calculated by simulation, T g -ML refers to the value predicted by the model, and the final result is similar. However, the prediction method of the present invention can predict the glass transition temperature more quickly, and can achieve efficient and accurate prediction of the glass transition temperature of solution-polymerized styrene-butadiene rubber (i.e., SSBR) based on the content of the four structural units within a certain range, which helps to accelerate the research and development process and solve the problems of long experimental cycle, high cost, and insufficient data volume. In addition, the method proposed in the present invention can use transfer learning or adjust intrinsic parameters, etc. to expand the scope of application to the target performance of other polymer materials.
[0117] Table 2
[0118]
[0119] The above technical solution is only one embodiment of the present invention. For those skilled in the art, it is easy to make various types of improvements or modifications based on the principles disclosed in the present invention, and it is not limited to the technical solution described in the above specific embodiments of the present invention. Therefore, the above description is only preferred and does not have a restrictive meaning.
Claims
1. A method for predicting polymer material properties based on small sample machine learning, characterized in that: The method comprises: Obtain raw data through molecular dynamics simulation; The original data is expanded by interpolation to obtain an expanded data set; Perform cluster analysis on the data set to obtain a class-balanced data set; The dataset is fed into the machine learning prediction model for training and prediction.
2. The method for predicting polymer material properties based on small sample machine learning according to claim 1, characterized in that: The operation of obtaining raw data through molecular dynamics simulation includes: By adjusting the content of structural units, multiple solution-polymerized styrene-butadiene rubber system models with different structural unit contents were obtained. The temperature-volume method is used to calculate each system model to obtain the corresponding original data.
3. The method for predicting polymer material properties based on small sample machine learning according to claim 2, characterized in that: The solution-polymerized styrene-butadiene rubber system model includes 10 identical solution-polymerized styrene-butadiene rubber single chains, each chain includes 60 repeating structural units; each piece of the original data includes four characteristic variables and a label value.
4. The method for predicting polymer material properties based on small sample machine learning according to claim 1, characterized in that: The operation of expanding the original data by interpolation includes: Establish an interpolation grid and use a random number generator to insert a set number of new data points near the original data points; Calculate the label value of the new data point to obtain the expanded data set.
5. The method for predicting polymer material properties based on small sample machine learning according to claim 4, characterized in that: The operation of calculating the label value of the new data point includes: The four-dimensional features of the original data are reduced to two dimensions as its spatial coordinates through principal component analysis; The spatial relationship between the original data is calculated through the semivariogram function; For new data points, the weight is determined by combining the spatial data between them and the original data; According to the weight, the weighted sum is used to obtain the label value of the new data point.
6. The method for predicting polymer material properties based on small sample machine learning according to claim 5, characterized in that: The label value of the new data point is obtained using the following formula: where λ i (i=1,…,n) is the weight, f(X i ) is the label value of the original data point in the dataset.
7. The method for predicting polymer material properties based on small sample machine learning according to claim 5 or 6, characterized in that: The weight is obtained by the following formula: Where μ is the Lagrange multiplier, γ(h) is the semivariogram equation, and λ k (k=1,…,n) is the weight, X i is the original data point, X k is another original data point, X j is a new data point, γ(x i -x k ) represents the semivariance equation of each known point and other known points; γ(x i -x j ) represents the semivariance equation between each known point and a new point to be estimated.
8. The method for predicting polymer material properties based on small sample machine learning according to claim 1, characterized in that: The cluster analysis includes: using the K-means method to obtain the number of data points in each category; comparing whether the ratio of the number of data points in each category exceeds 3:1; The ratio of the number of data points of each category does not exceed 3:1, and there is no category imbalance in the dataset; The ratio of the number of data points in each category exceeds 3:1, and the dataset is class-imbalanced. The SMOTE method is used for quadratic interpolation to obtain a class-balanced dataset.
9. The method for predicting polymer material properties based on small sample machine learning according to claim 1, characterized in that: The operations of inputting the data set into the machine learning prediction model for training and performing prediction include: The class-balanced dataset is divided into training set and test set according to the ten-fold cross-validation method; The training set is input into the machine learning prediction model for training, and the test set is used for verification, and finally a trained model is obtained; The contents of the four structural units of solution-polymerized styrene-butadiene rubber are input into a trained prediction model, and the trained prediction model outputs prediction results of polymer material properties.
10. The method for predicting polymer material properties based on small sample machine learning according to claim 9, characterized in that: The machine learning prediction model adopts the XGBoost model, and its setting parameters include: a learning rate of 0.08, a maximum tree depth of 13, and the number of trees is 120.