A surfactant formulation recommendation model training method and system
By using orthogonal experimental design and machine learning model training, the problems of low efficiency and inconsistent data management in traditional surfactant formulation optimization have been solved, enabling efficient and accurate surfactant formulation recommendations and improving formulation development efficiency and success rate.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-03-31
AI Technical Summary
Traditional surfactant system formulation optimization relies on human experience, resulting in low experimental efficiency, inconsistent data management, poor operational repeatability, inability to systematically cover complex parameter spaces, and inability to reveal nonlinear interactions and effects between components, thus limiting the global optimization potential of the formulation.
An initial experimental matrix is generated using orthogonal experimental design. Key factors are identified through analysis of variance. A supplementary experimental matrix is generated by combining central composite design. A surfactant formulation recommendation model is trained using machine learning model, including a main trend prediction sub-model and a residual correction sub-model, to achieve efficient and accurate formulation optimization.
It has enabled a shift from low-throughput, experience-driven random adjustments to high-throughput, data-driven global system optimization, significantly improving the efficiency and success rate of formulation development and building an efficient, accurate, and traceable intelligent R&D paradigm.
Smart Images

Figure CN121118698B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of chemical engineering technology, and in particular relates to a method and system for training a surfactant formulation recommendation model. Background Technology
[0002] Traditional surfactant formulation optimization has long relied on manual experience and trial-and-error adjustments, a model with significant efficiency bottlenecks and insufficient scientific rigor. Currently, most laboratories still use the traditional procedures of manually preparing solutions and measuring interfacial tension, resulting in low throughput, poor repeatability, and unavoidable systematic errors introduced by human intervention. Simultaneously, outdated data management methods mean that key information such as formulation parameters, interfacial tension values, and environmental conditions are often scattered across paper documents or various electronic devices, lacking a unified, structured, and traceable data integration platform. This makes it difficult to review and optimize experimental processes in a closed-loop manner.
[0003] Furthermore, because the entire process relies on manual labor, the experimental cycle is long and resource-intensive, researchers often cannot strictly execute scientifically designed multi-factor orthogonal experiments in practice, and can only rely on limited experience to perform local parameter tuning. This "experience-driven" optimization model is difficult to systematically cover the complex parameter space in multi-component formulations, and cannot effectively reveal the nonlinear interactions and effects between components, thus limiting the global optimization potential of the formulation.
[0004] Therefore, in the face of the above challenges, it is particularly urgent to construct a training method and system for a surfactant formulation recommendation model. Summary of the Invention
[0005] Therefore, it is necessary to provide a method and system for training a surfactant formulation recommendation model to address the aforementioned technical problems.
[0006] Firstly, this application provides a method for training a surfactant formulation recommendation model, the method comprising:
[0007] Based on the number of surfactant components in the target system, an initial experimental matrix is generated using orthogonal experimental design. The target system represents the specific scope and final goal of this experiment. Each experimental unit in the experimental matrix includes different surfactant ratios, surfactant condition parameters, and measured interfacial tension values.
[0008] A variance analysis is performed on the initial experimental matrix to identify factors that have an impact on the interfacial tension value greater than a set threshold. Based on these factors, experimental points are generated to obtain a supplementary experimental matrix generated based on these experimental points. Here, each experimental point represents an experimental scheme formed by traversing and combining the factors.
[0009] The initial experimental matrix and the supplementary experimental matrix are merged to obtain the experimental matrix training set;
[0010] Using the experimental matrix training set, the initial surfactant formulation recommendation model is trained to obtain the target surfactant formulation recommendation model. The initial surfactant formulation recommendation model includes an initial main trend prediction sub-model and an initial residual correction sub-model. The initial residual correction sub-model is configured to take the prediction result of the initial main trend prediction sub-model as input and correct the residual of the prediction result. The target surfactant formulation recommendation model is configured to receive an input vector consisting of surfactant ratio and condition parameters and output the predicted interfacial tension value.
[0011] In some feasible methods, the step of generating an initial experimental matrix using orthogonal experimental design based on the number of surfactant components in the target system includes:
[0012] Based on the composition of surfactants in the target system, the factors and the levels of each factor are determined. The factors include the ratio of various surfactants, as well as temperature, rotation speed and chemical form condition parameters. The levels represent discrete values preset by various factors.
[0013] Based on the factors and their levels, a matching orthogonal table is selected from the standard orthogonal table library. The factors and levels are then substituted into the selected orthogonal table to perform orthogonal experimental design and generate an initial experimental plan matrix.
[0014] Experiments are conducted at each experimental point in the initial experimental plan matrix. The interfacial tension value at each experimental point is measured and recorded, and the interfacial tension value is added as a new column to the initial experimental plan matrix to form the initial experimental matrix.
[0015] In some feasible methods, the step of performing variance analysis on the initial experimental matrix to identify factors whose influence on the interfacial tension value is greater than a set threshold, and generating experimental points based on the factors to obtain a supplementary experimental matrix generated based on the experimental points, includes:
[0016] Based on the initial experimental matrix, with the interfacial tension value as the dependent variable and the ratio of each surfactant and the condition parameters as independent variables, an analysis of variance was performed to calculate the p-values of each factor and the interactions between factors. Factors and interaction terms with p-values less than the preset significance level were identified as factors that have a significant impact on interfacial tension, thus obtaining a list of factors.
[0017] Based on the list of factors, the optimal response regions of the factors in the experimental space are collected using the central composite design method, and a supplementary experimental plan matrix of the experimental point set is generated.
[0018] Experiments were conducted using each experimental point in the supplementary experimental plan matrix, and the interfacial tension value at each new experimental point was measured and recorded to obtain the supplementary experimental matrix.
[0019] In some feasible methods, the step of collecting the optimal response regions of factors in the experimental space and generating a supplementary experimental plan matrix of experimental point sets based on the factor list using the central composite design method includes:
[0020] Based on the list of factors, define the upper and lower limits for each factor, and calculate the coordinate values of the corresponding center point and axial point to obtain a structured factor design space table.
[0021] Based on the structured factor design space table, and using the mathematical model of central composite design, an initial CCD point set containing the coordinates of experimental points is generated, wherein the initial CCD point set includes a set of coordinates of cubic points, axial points, and center points.
[0022] The initial CCD point set is converted into a plan table, generating a supplementary experimental plan matrix for the experimental point set. The plan table lists the running sequence of each experiment and the specific settings of each factor.
[0023] In some feasible methods, the step of merging the initial experimental matrix and the supplementary experimental matrix to obtain the experimental matrix training set includes:
[0024] Align the initial experimental matrix and the supplementary experimental matrix to obtain two structure-aligned standardized experimental matrices;
[0025] The two structure-aligned standardized experimental matrices are merged along the row direction to obtain the merged deduplicated experimental data table.
[0026] The surfactant ratio and condition parameter columns in the merged deduplication experimental data table are used as feature variables, and the interfacial tension value column is used as the target variable to obtain the experimental matrix training set.
[0027] In some feasible methods, the step of training the initial surfactant formulation recommendation model using the experimental matrix training set to obtain the target surfactant formulation recommendation model includes:
[0028] Using the data in the experimental matrix training set, the correlation between tension and factors in the initial main trend prediction sub-model is trained to obtain the target main trend prediction sub-model, wherein the target main trend prediction sub-model represents the output of the preliminary interface tension prediction value based on the input formula;
[0029] Based on the prediction results of the target main trend prediction sub-model on the experimental matrix training set, the prediction residual is calculated. Using the data in the experimental matrix training set as input and the prediction residual as the training target, the initial residual correction sub-model is trained to obtain the target residual correction sub-model.
[0030] The target main trend prediction sub-model and the target residual correction sub-model are combined to obtain the target surfactant formulation recommendation model, wherein the output of the target surfactant formulation recommendation model is the sum of the output of the target main trend prediction sub-model and the output of the target residual correction sub-model.
[0031] In some feasible methods, after the step of training the initial surfactant formulation recommendation model using the experimental matrix training set to obtain the target surfactant formulation recommendation model, the method further includes:
[0032] User preferences are used as conditions and input into the target surfactant formulation recommendation model to form multiple optimal solutions and obtain a candidate formulation list. The user preferences represent interfacial tension performance preference and cost priority preference.
[0033] Secondly, this application provides a surfactant formulation recommendation model training system, applied to the aforementioned surfactant formulation recommendation model training method, the system comprising:
[0034] The generation unit is used to generate an initial experimental matrix based on the number of surfactant components in the target system using orthogonal experimental design. The target system represents the specific scope and final goal of this experiment. Each experimental unit in the experimental matrix includes different surfactant ratios, surfactant condition parameters, and measured interfacial tension values.
[0035] The analysis unit is used to perform variance analysis on the initial experimental matrix, identify factors that have an impact on the interfacial tension value greater than a set threshold, and generate experimental points based on the factors to obtain a supplementary experimental matrix generated based on the experimental points, wherein the experimental points represent experimental schemes formed by traversing and combining the factors.
[0036] The merging unit is used to merge the initial experimental matrix and the supplementary experimental matrix to obtain an experimental matrix training set.
[0037] The result unit is used to train the initial surfactant formulation recommendation model using the experimental matrix training set to obtain the target surfactant formulation recommendation model. The initial surfactant formulation recommendation model includes an initial main trend prediction sub-model and an initial residual correction sub-model. The initial residual correction sub-model is configured to take the prediction result of the initial main trend prediction sub-model as input and correct the residual of the prediction result. The target surfactant formulation recommendation model is configured to receive an input vector consisting of surfactant ratio and condition parameters and output the predicted interfacial tension value.
[0038] Thirdly, this application provides a computer storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the aforementioned method.
[0039] Fourthly, this application provides a computer program that, when executed by a processor, implements the steps of the aforementioned method.
[0040] Beneficial Effects: This application provides a method for training a surfactant formulation recommendation model. The method includes: generating an initial experimental matrix based on the number of surfactant components in a target system using orthogonal experimental design, wherein the target system represents the specific scope and final goal of this experiment, and each experimental unit in the experimental matrix includes different surfactant ratios, surfactant condition parameters, and measured interfacial tension values; performing variance analysis on the initial experimental matrix to identify factors that have an impact on the interfacial tension value greater than a set threshold, and generating experimental points based on the factors to obtain a supplementary experimental matrix generated based on the experimental points, wherein the experimental points represent combinations based on the factors. The experimental scheme is formed; the initial experimental matrix and the supplementary experimental matrix are merged to obtain an experimental matrix training set; the initial surfactant formulation recommendation model is trained using the experimental matrix training set to obtain a target surfactant formulation recommendation model. The initial surfactant formulation recommendation model includes an initial main trend prediction sub-model and an initial residual correction sub-model. The initial residual correction sub-model is configured to take the prediction result of the initial main trend prediction sub-model as input and correct the residual of the prediction result. The target surfactant formulation recommendation model is configured to receive an input vector consisting of surfactant ratio and conditional parameters and output the predicted interfacial tension value. Through the above method, an efficient, accurate, and traceable intelligent R&D paradigm is constructed. This method effectively overcomes the core defects of traditional manual experience-based trial-and-error modes, such as low experimental efficiency, poor data quality, insufficient parameter coverage, and inability to reveal complex nonlinear relationships. It achieves a fundamental shift from low-throughput, experience-driven random adjustments to high-throughput, data-driven global system optimization, thereby significantly improving the efficiency and success rate of formulation R&D. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a flowchart of a surfactant formulation recommendation model training method in one embodiment. Detailed Implementation
[0043] To facilitate understanding of this application, a more complete description will be provided below with reference to the accompanying drawings, which illustrate embodiments of the present application. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of this application will be thorough and complete.
[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used herein includes any and all couplings of one or more of the associated listed items.
[0045] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish one element from another.
[0046] The following explanations of some terms used in this application are provided to aid in understanding the application:
[0047] Central composite design (CCD) is an experimental design method for response surface methodology. It is primarily used to efficiently explore the nonlinear relationships (e.g., quadratic relationships) between identified key influencing factors and the output (i.e., the "response"), thereby finding optimal process parameters or formulation ratios. In this application, after screening a few key factors through orthogonal experiments, CCD is used to conduct targeted supplementary experiments to obtain high-quality data sufficient to train a powerful predictive model.
[0048] Support Vector Regression (SVR) is an application of Support Vector Machines (SVMs) to regression problems. SVR allows a portion of the data points to deviate from the model, as long as this deviation is within an acceptable range. SVR's nonlinear mapping capability can effectively capture the complex main effects and interactions between formulation, process parameters, and interfacial tension, providing a robust basis for subsequent residual correction. The goal of SVR is not to minimize the error of all points, but to find a function that ensures the vast majority of data points fall within this interval, while making this interval as "flat" as possible (i.e., the function is as simple as possible to avoid overfitting).
[0049] Random Forest (RF) is an ensemble learning algorithm that builds a large number of decision trees and combines their predictions (e.g., by averaging) to obtain a more accurate and stable model.
[0050] Five-fold cross-validation is a statistical method used to evaluate the generalization performance of machine learning models and to fine-tune model parameters. Its core idea is to repeatedly and effectively utilize a limited dataset to obtain a more stable and reliable model evaluation result through multiple training and testing cycles.
[0051] Early stopping is a regularization technique used in machine learning model training to prevent overfitting. Its core idea is to continuously monitor the model's performance on the validation set during training. If performance stops improving or even begins to decline, the training process is terminated early, even if the model's loss on the training set is still decreasing.
[0052] XGBoost, or eXtreme Gradient Boosting, is a high-performance gradient boosting algorithm that progressively corrects errors by building a model sequentially. In this application, it is used as a residual correction model, leveraging its high accuracy and anti-overfitting properties to refine and correct the predictions of the main model, thereby constructing a powerful interface tension prediction system.
[0053] The Non-dominated Sorting Genetic Algorithm II (NSGA-II) is one of the most widely used multi-objective optimization algorithms. It is an evolutionary algorithm based on genetic algorithms, specifically designed to solve problems with multiple conflicting optimization objectives.
[0054] Pareto describes a state where further optimization is impossible: any adjustment that attempts to improve one objective will inevitably worsen at least one other objective. The set of these optimal solutions is called the Pareto front, which defines the best trade-off boundary in multi-objective problems.
[0055] Monte Carlo simulation is a computational method that approximates the solution of complex problems by generating a large number of random samples (i.e., probabilistic simulation). Its core idea is to use randomness to simulate uncertainty, and through tens of thousands or even millions of experiments, to finally obtain stable predictions based on the average or distribution of all possible outcomes.
[0056] like Figure 1 As shown, in a first aspect, this application provides a method for training a surfactant formulation recommendation model, the method comprising:
[0057] S100: Based on the number of surfactant components in the target system, an initial experimental matrix is generated using orthogonal experimental design.
[0058] The target system represents the specific scope and ultimate goal of this experiment, and each experimental unit in the experimental matrix includes different surfactant ratios, surfactant condition parameters, and measured interfacial tension values.
[0059] This step aims to systematically obtain a batch of high-quality initial experimental data covering the target formulation parameter space with the fewest number of experiments through scientific and efficient experimental design methods.
[0060] Specifically, generating the initial experimental matrix based on the number of surfactant components in the target system using orthogonal experimental design may include the following steps:
[0061] S101, based on the number of surfactant components in the target system, determine the factors and the level of each factor.
[0062] The factors include the ratio of various surfactants, as well as temperature, rotation speed and chemical form condition parameters, and the level represents the discrete values preset by various factors.
[0063] Specifically, the purpose of S101 is to clarify the variables and scope to be studied. Based on the target system of this optimization, all key variables that need to be adjusted are listed. These variables are divided into two categories: formulation factors and condition factors. Further, formulation factors refer to the ratio (concentration) of each surfactant component. For example, if the system contains three surfactants, the factors are "component A concentration," "component B concentration," and "component C concentration." Condition factors refer to environmental or process parameters during the experiment, such as temperature, stirring speed, and pH value.
[0064] Setting factor levels: Pre-determine several discrete value points for each factor, called "levels". Levels are typically set to two (e.g., "high", "low") or three (e.g., "low", "medium", "high"), and their specific values need to be determined based on preliminary experiments or professional knowledge to ensure coverage of the best possible range. For example, the levels for the "temperature" factor can be set to 30°C, 45°C, and 60°C.
[0065] S102, Based on the factors and their levels, select a matching orthogonal table from the standard orthogonal table library, and substitute the factors and levels into the selected orthogonal table to perform orthogonal experimental design and generate an initial experimental plan matrix.
[0066] Among them, by utilizing the characteristics of orthogonal arrays that are "uniformly dispersed and neatly comparable", the most representative small number of combinations are selected from all possible formula combinations to form an experimental scheme.
[0067] Specifically, based on the total number of factors and the number of levels for each factor determined in S101, an orthogonal array is matched from a pre-defined standard orthogonal array library. For example, for an experiment with 4 factors and 3 levels for each factor, the L9 orthogonal array is usually selected, which only requires 9 experiments.
[0068] Assign factors and levels: The factors and levels determined in step S101 are assigned to the columns of the selected orthogonal array according to certain rules.
[0069] Certain rules indicate that:
[0070] a. Examine interactions: Based on the assumptions or expertise made before the experimental design, determine whether it is necessary to examine the interaction between specific factor pairs (such as "surfactant A" and "temperature").
[0071] b. Consult the interaction table: If you need to examine interactions, consult the "interaction table" corresponding to the selected orthogonal array to find the specific column used to analyze the interaction. For example, in table L8, the interaction between columns 1 and 2 is located in column 3.
[0072] c. Assignment columns: Assign the most important factor (such as the concentration of the main surfactant) to column 1; assign another factor that interacts with it (such as temperature) to column 2; reserve column 3 for their interactions, and do not assign any actual factors to this column; randomly assign the remaining factors to the remaining empty columns.
[0073] Finally, fill in the specific level values of each factor (such as temperature: 30°C, 60°C) into their respective columns according to the original "1" and "2" level labels in the orthogonal table.
[0074] Generating the initial experimental plan matrix: Each row of the orthogonal array constitutes a combination of experimental conditions to be executed, and the set of all rows generates the "initial experimental plan matrix". This initial experimental plan matrix is a data table that clearly lists the specific settings of each factor for each experiment, but it does not yet contain the experimental results.
[0075] S103, conduct experiments according to each experimental point in the initial experimental plan matrix, measure and record the interface tension value of each experimental point, and add the interface tension value as a new column to the initial experimental plan matrix to form the initial experimental matrix.
[0076] The experiment is carried out according to each column of the initial experimental plan matrix, the target performance index (interfacial tension) is accurately measured, and the results are correlated with the experimental conditions to form a complete dataset.
[0077] For example, an "initial experimental plan matrix" can be issued to an automated solution preparation and measurement device (both automated solution preparation and measurement devices are existing equipment). The device automatically completes the sample preparation and condition control (such as constant temperature and stirring) operations at each experimental point according to the order in the initial experimental plan matrix.
[0078] For each prepared sample, an interfacial tension value can be measured using a tensiometer, ensuring that the measurement conditions are stable and consistent.
[0079] As another example, based on the number of surfactant components in the target system (e.g., up to 8), an orthogonal experimental design (L9, L10, L20, L30, L40, L50, L60, L70, L80, L90, L1 ... 27 or L 81 (Type) Generates an experimental matrix. Each experimental unit contains:
[0080] Different surfactant ratios ;
[0081] Conditional parameters such as temperature, rotation rate, and chemical properties;
[0082] The measured interfacial tension results.
[0083] Through experimental execution, the system obtains the initial training dataset:
[0084] ;
[0085] in, This represents the initial training dataset (initial experiment matrix). This represents the input feature vector (components and experimental conditions). This represents the measured value of interfacial tension.
[0086] Finally, the measured "interfacial tension values" are used as a new data column and correlated with the "initial experimental plan matrix" generated in S102. The resulting complete data table, which includes both input conditions and output results, is the "initial experimental matrix." This matrix will serve as the basis for subsequent data analysis and model training.
[0087] S200, perform variance analysis on the initial experimental matrix to identify factors that have an impact on the interfacial tension value greater than a set threshold, and generate experimental points based on the factors to obtain a supplementary experimental matrix generated based on the experimental points.
[0088] Wherein, the experimental point represents an experimental scheme formed by traversing and combining the factors.
[0089] It should be noted that the core influencing factors are selected from the initial experimental matrix, and focused experimental designs are carried out around these factors to obtain supplementary data that can reveal complex nonlinear relationships, laying the foundation for building a high-precision prediction model.
[0090] Specifically, obtaining the supplementary experimental matrix generated based on the experimental points may include the following steps:
[0091] S201. Based on the initial experimental matrix, with the interfacial tension value as the dependent variable and the ratio of each surfactant and the condition parameters as independent variables, perform an analysis of variance to calculate the p-value of each factor and the interaction between factors. Factors with p-values less than the preset significance level and interaction terms are identified as factors that have a significant impact on interfacial tension, thus obtaining a factor list.
[0092] Specifically, mathematical statistics methods are used to quantify the influence of each factor (and the interaction between factors) on the interfacial tension results, and the factors with significant influence are identified as key factors and factors that need to be used.
[0093] It should be noted that, based on the "initial experimental matrix," the measured "interfacial tension value" is taken as the effect to be explained (dependent variable), and the surfactant ratios and condition parameters (temperature, rotation speed, etc.) are taken as possible causes (independent variables). The null hypothesis of the statistical analysis is: a certain independent variable has no significant effect on interfacial tension. Next, an analysis of variance is performed to calculate the p-value for each independent variable (and its interaction). The p-value represents the probability of observing the current experimental data (or more extreme data) given that the null hypothesis is true (i.e., the factor is invalid). The smaller the p-value, the less likely the effect of the factor is to be caused by random error, i.e., the more significant the effect. The calculated p-values are compared with the preset significance level (usually set to 0.05 or 0.01). If the p-value of a factor is less than the significance level, the null hypothesis is rejected, and the factor is determined to be a "key factor" with a significant effect on interfacial tension. All the determined key factors constitute the "factor list."
[0094] S202, Based on the list of factors, the optimal response region of each factor in the experimental space is collected using the central composite design method, and a supplementary experimental plan matrix of the experimental point set is generated.
[0095] Specifically, in a high-dimensional parameter space composed of key factors, points are arranged according to specific geometric rules. These points can efficiently support the modeling of nonlinear relationships (such as parabolic relationships) between factors and performance.
[0096] It should be noted that for each factor in the list of key factors, based on its performance in the initial experiment, a numerical range requiring fine-tuning is determined (i.e., setting new upper and lower limits for the experiment). Using a central composite design method, three types of experimental points are systematically generated within the defined experimental space: cubic points, axial points, and center points. Finally, all types of experimental points and their corresponding key factor values are summarized into a structured table, namely the "Supplementary Experiment Plan Matrix." This matrix clearly lists the execution sequence of each round of supplementary experiments and the specific set values for each key factor.
[0097] It should also be noted that step S202 may include:
[0098] S2021, Based on the factor list, define the upper and lower limits of each factor, and calculate the coordinate values of the corresponding center point and axial point to obtain a structured factor design space table.
[0099] Specifically, a numerical range that needs to be explored in detail is determined for each key factor, and the core coordinate points of this space are calculated to provide a benchmark for subsequent point placement.
[0100] First, for each factor in the list of key factors, analyze its value in the "initial experiment matrix" and its corresponding effect on interface tension to determine a more targeted experimental range. The upper and lower limits of this range should cover its possible optimal response area.
[0101] Next, the coordinates of the center point and axial points are calculated. The center point refers to the midpoint of the value range of each key factor. For example, if the lower limit of a factor is 1% and the upper limit is 5%, then its center point coordinates are (1% + 5%) / 2 = 3%. The axial point refers to a point set outwards along the coordinate axis of each factor, based on the center point, by a certain distance. This distance is determined by the central composite design model. The value of the axial point ensures the specific geometric properties of the experimental design (such as rotatability) to better detect bending effects. For example, if the center point is 3% and the α (axial point) value is 1.5, then the axial point coordinates of this factor can be calculated as 3% ± (1.5 × (5%-3%) / 2).
[0102] Finally, the names, upper and lower limits, center point coordinates, and calculated axial point coordinates of all the above key factors are compiled into a structured key factor design space table.
[0103] S2022, Design a spatial table based on the structured factors, and use the mathematical model of central composite design to generate an initial CCD point set containing the coordinates of the experimental points.
[0104] The initial CCD point set includes a set of points with coordinates of cubic points, axial points, and center points.
[0105] Specifically, within the defined experimental space, three types of experimental points with specific mathematical meanings are systematically generated. The combination of these points can efficiently fit complex response surfaces.
[0106] Specifically, the three types of experimental points with specific mathematical significance include:
[0107] Cube point generation: Experimental points are set at the vertices of a "cube" composed of upper and lower limits of key factors. These points are mainly used to accurately estimate the main effects of factors and the interaction effects between factors.
[0108] Generate axial points: Set experimental points at the axial point coordinates of each factor pre-calculated in S2021. At this point, other factors remain at the center point level. These points are designed to probe and fit nonlinear relationships (bending effects).
[0109] Set a center point: Conduct several repeated experiments at the center of the space (i.e., take the center point value for all factors). These points are used to estimate experimental error and check for model curvature.
[0110] Converging point set: The coordinates of all the above cubic points, axial points and center points are gathered together to form the "initial CCD point set".
[0111] S2023 converts the initial CCD point set into a plan table and generates a supplementary experimental plan matrix for the experimental point set.
[0112] The schedule lists the order of each experiment and the specific settings for each factor.
[0113] Specifically, the initial CCD point set obtained in step S2022 is transformed into a specific and ordered list of instructions to guide the corresponding experimental equipment.
[0114] For example, firstly, all experimental points in the initial CCD point set are randomly sorted. This is to eliminate potential unknown time-related biases (such as equipment state drift) during the experiment and improve the reliability of the experimental results. Then, a planning matrix is generated, creating a well-structured table, namely the "Supplementary Experimental Planning Matrix." This matrix should contain at least the following:
[0115] Experiment run sequence number: indicates the order in which the experiments are executed.
[0116] Key Factors Columns: Each column corresponds to a key factor, and the factor value is the specific set value of the factor in this experiment (obtained from the CCD point set).
[0117] Experimental point type identifier: Indicate whether the experimental point is a cubic point, an axial point, or a center point, to facilitate subsequent data analysis.
[0118] S203. Using each experimental point in the supplementary experimental plan matrix, conduct experiments and measure and record the interfacial tension value of each new experimental point to obtain the supplementary experimental matrix.
[0119] Specifically, the supplementary experimental plan matrix instruction is sent to the corresponding experimental equipment (which is the existing experimental equipment), and the sample preparation and condition control operations for each new experimental point are completed sequentially. Next, the interfacial tension value of each newly prepared sample is automatically measured. Then, the newly measured "interfacial tension value" is used as a new column and associated with the supplementary experimental plan matrix generated by S202 to form a "supplementary experimental matrix" containing all new experimental conditions and their results. This supplementary experimental matrix, together with the previous "initial experimental matrix," constitutes the complete dataset for subsequent model training.
[0120] S300, the initial experimental matrix and the supplementary experimental matrix are merged to obtain the experimental matrix training set.
[0121] It should be noted that the S300 step involves effectively integrating and standardizing the two sets of data obtained in the early stages through different experimental designs (orthogonal experiments and central composite designs) to form a unified and standardized dataset, providing a high-quality data foundation for subsequent machine learning model training.
[0122] Specifically, obtaining the experimental matrix training set may include the following steps:
[0123] S301, Align the initial experimental matrix and the supplementary experimental matrix to obtain two structure-aligned standardized experimental matrices.
[0124] Specifically, examine the columns of the initial experimental matrix and the supplementary experimental matrix. Ensure they contain the same columns for surfactant ratios, conditional parameters (such as temperature and rotation speed), and interfacial tension values. If one matrix is missing a column from the other (e.g., a conditional parameter was not studied in the supplementary experiment), this column needs to be added to the supplementary matrix, filled with a pre-defined missing value (such as NaN) or the parameter's default value to ensure structural consistency. Next, adjust the order of the columns in both matrices to make them completely identical. For example, arrange them in the order of surfactant A concentration, surfactant B concentration, temperature, rotation speed, and interfacial tension value. This adjustment yields the standardized experimental matrices for the two structures.
[0125] S302, merge the two structure-aligned standardized experimental matrices along the row direction to obtain the merged deduplicated experimental data table.
[0126] Specifically, the two experimental matrix data sets are concatenated to form a larger dataset, and any completely duplicate experimental records are removed to ensure the uniqueness of the data.
[0127] For example, the two structure-aligned matrices obtained in S301 are concatenated along the row direction (i.e., the vertical direction). After merging, the columns of the new data table remain unchanged, while the number of rows is the sum of the number of rows in the two original matrices. The merged data table is then checked to determine if there are rows where all feature variables (surfactant ratio and condition parameters) have exactly the same values (i.e., two completely identical experiments). If such rows exist, they are considered duplicate data, and only one row is retained to avoid introducing unnecessary bias into model training.
[0128] After merging and deduplication, the merged and deduplicated experimental data table is obtained.
[0129] S303, the surfactant ratio and condition parameter columns in the merged deduplication experimental data table are used as feature variables, and the interfacial tension value column is used as the target variable to obtain the experimental matrix training set.
[0130] Specifically, in the merged deduplication experimental data table: all surfactant ratio columns and condition parameter columns are used as inputs to the model, i.e., feature variables (which can be represented by the symbol x). The interfacial tension value column is used as the output to be predicted by the model, i.e., the target variable (which can be represented by the symbol y).
[0131] The final generated "experimental matrix training set" is a structured collection containing (x, y) data pairs. Here, x is a two-dimensional array (or matrix), where each row represents an experimental formulation and its conditions, and each column represents a feature; y is a one-dimensional array (or vector), where each element is the interfacial tension measurement value of the corresponding row. This training set can be directly used for training the initial surfactant formulation recommendation model in step S400.
[0132] S400, using the experimental matrix training set, the initial surfactant formulation recommendation model is trained to obtain the target surfactant formulation recommendation model.
[0133] The initial surfactant formulation recommendation model includes an initial main trend prediction sub-model and an initial residual correction sub-model. The initial residual correction sub-model is configured to take the prediction result of the initial main trend prediction sub-model as input and correct the residual of the prediction result. The target surfactant formulation recommendation model is configured to receive an input vector consisting of surfactant ratio and condition parameters and output the predicted interfacial tension value.
[0134] It should be noted that in step S400, a hybrid machine learning model capable of predicting interfacial tension for any formulation is trained using the integrated experimental data (experimental matrix training set). This model employs a two-layer structure of "main trend prediction + residual correction," which improves prediction accuracy through cascading.
[0135] Specifically, obtaining a target surfactant formulation recommendation model may include the following steps:
[0136] S401, using the data in the training set of the experimental matrix, train the relationship between the tension and factors of the initial main trend prediction sub-model to obtain the target main trend prediction sub-model.
[0137] The target main trend prediction sub-model represents the output of preliminary interfacial tension prediction values based on the input formula.
[0138] Specifically, a target master trend prediction sub-model is trained to learn and capture the main correlations and global trends between surfactant formulations, process conditions, and interfacial tension from experimental data. This model is responsible for providing a preliminary and reasonable prediction.
[0139] For example, all columns representing surfactant ratios (e.g., component concentrations) and conditional parameters (e.g., temperature, rotation speed) in the experimental matrix training set are used as input features for the initial main trend prediction sub-model. Assuming there are p features, the input of one experiment can be represented as a p-dimensional feature vector x (feature variable). The measured interfacial tension values in the experimental matrix training set are used as the target to be learned and predicted by the initial main trend prediction sub-model. Let it be a scalar y (target variable). The entire experimental matrix training set (containing data from N experiments) is input into the initial main trend prediction sub-model. The initial main trend prediction sub-model adjusts its internal parameters using its built-in algorithms (such as Support Vector Regression (SVR) or Random Forest Regression (RF) algorithms), aiming to make the predicted value f1(x) of the initial main trend prediction sub-model as close as possible to the true interfacial tension value y. Furthermore, the mathematical essence of this learning process is minimizing a loss function. After training, the initial main trend prediction sub-model evolves into the target main trend prediction sub-model. The model can take a new recipe vector and output a preliminary interfacial tension prediction that reflects the main patterns known based on the current data.
[0140] S402, based on the prediction results of the target main trend prediction sub-model on the experimental matrix training set, calculate the prediction residual, use the data in the experimental matrix training set as input, use the prediction residual as the training target, train the initial residual correction sub-model, and obtain the target residual correction sub-model.
[0141] Specifically, the target main trend prediction sub-model may fail to capture all the complex patterns in the data, leaving systematic prediction errors. Step S402 aims to train a specialized model to learn and correct these errors, thereby improving overall prediction accuracy.
[0142] For example, the trained target main trend prediction sub-model is used to predict each sample in the experimental matrix training set to obtain an initial prediction value. For each sample, the difference between its predicted value and the actual measured value is calculated, i.e., the prediction residual. This can be calculated using difference calculation. Next, the input features used in training the initial main trend prediction sub-model are exactly the same, namely, the feature vector x composed of the surfactant ratio and conditional parameters. The prediction residual calculated in the previous step is used as the new learning target. In other words, the task of the initial residual correction sub-model is not to predict interfacial tension, but to predict the prediction error of the target main trend prediction sub-model.
[0143] It should be noted that the feature vector x and the corresponding predicted residuals are used as training data and input into the initial residual correction sub-model (such as the XGBoost model). The internal algorithm of the initial residual correction sub-model learns the complex mapping relationship between features and residuals. The goal is to make the predicted values of the initial residual correction sub-model as close as possible to the true residuals. After training, the initial residual correction sub-model evolves into the target residual correction sub-model. This model can receive a new recipe vector and output a predicted residual correction value.
[0144] S403, combine the target main trend prediction sub-model and the target residual correction sub-model to obtain the target surfactant formulation recommendation model, wherein the output of the target surfactant formulation recommendation model is the sum of the output of the target main trend prediction sub-model and the output of the target residual correction sub-model.
[0145] Specifically, two sub-models, each adept at capturing the main trend and correcting residuals respectively, are integrated to construct the final hybrid prediction model. The final prediction equals the main trend prediction plus a fine-tuning of the main trend error. The final prediction output of the target surfactant formulation recommendation model is a linear superposition of the outputs of the two sub-models.
[0146] It should be noted that the target main trend prediction sub-model is responsible for capturing the primary, global relationship between the formulation and interfacial tension, providing a robust prediction baseline. The target residual correction sub-model is responsible for capturing the subtle, nonlinear, and complex local patterns that the target main trend prediction sub-model fails to learn, and for fine-tuning the baseline prediction.
[0147] In the S400 steps, for example, a two-layer hybrid model of "main trend + residual correction" is established, with the overall functional form as follows:
[0148] ;
[0149] in, This represents the model's predictive performance index for surfactant formulations. This represents a target surfactant formulation recommendation model. This represents the target main trend prediction sub-model. This represents the target residual correction sub-model.
[0150] Furthermore, the target main trend prediction sub-model The formula is:
[0151] ;
[0152] in, This is the radial basis function kernel, used to compute new samples. With the support vectors Similarity between them For the first The weight coefficients corresponding to each support vector, where b is the bias term. This represents the total number of support vectors.
[0153] Target Residual Correction Submodel The formula is:
[0154] Using XGBoost residual ( Represented as , Learning from (representing actual observations):
[0155] ;
[0156] in, Indicates the summation index. Indicates the learning rate. Indicates the first The predicted values of the regression trees, This represents the input feature vector. This represents the total number of regression trees (weak learners).
[0157] In one embodiment, after training the initial surfactant formulation recommendation model using the experimental matrix training set to obtain the target surfactant formulation recommendation model, the method further includes:
[0158] User preferences are used as conditions and input into the target surfactant formulation recommendation model to form multiple optimal solutions and obtain a candidate formulation list. The user preferences represent interfacial tension performance preference and cost priority preference.
[0159] Specifically, after obtaining a model that can accurately predict interfacial tension, a limited number of candidate solutions that achieve the best balance between "performance" and "cost" are automatically selected from possible formulations based on the user's specific preferences, forming a list of candidate formulations that can be directly used for decision-making.
[0160] For example, after the model training is completed, a performance-cost balance is achieved through a multi-objective optimization algorithm (NSGA-II):
[0161] Minimize interface tension
[0162] Minimize raw material costs.
[0163] Constraint: The sum of the proportions of all components is 1.
[0164] Optimization results:
[0165] Generate the Pareto optimal recipe set Each solution corresponds to a balance between performance and cost.
[0166] The system automatically recommends 3–5 candidate formulations based on user preferences (such as performance priority or cost priority).
[0167] Based on the user's selection, the system automatically extracts the corresponding formulation and its predicted performance (predicted interfacial tension value, estimated cost) from the Pareto optimal solution set, forming a clear list of candidate formulations for the user to finally evaluate and select.
[0168] In one embodiment, after training the initial surfactant formulation recommendation model using the experimental matrix training set to obtain the target surfactant formulation recommendation model, the method further includes:
[0169] Using the target surfactant formulation recommendation model, Monte Carlo simulations were performed in the formulation parameter space to obtain the region with the largest prediction variance and the boundary region with the best prediction performance.
[0170] Based on the region with the largest prediction variance and the boundary region with the best prediction performance, data augmentation is performed to obtain new experimental data.
[0171] The newly added experimental data is combined with the experimental matrix training set to train the target surfactant formulation recommendation model, thus obtaining the final surfactant formulation recommendation model.
[0172] Specifically, within the formulation parameter space (i.e., the allowed range of values for all surfactant ratios and condition parameters), a large number (e.g., tens of thousands or hundreds of thousands) of virtual formulation points are randomly generated using Monte Carlo simulation. For each virtual point, a pre-trained target surfactant formulation recommendation model is used to make predictions in two aspects:
[0173] Predictive performance: i.e., the interfacial tension value predicted by the target surfactant formulation recommendation model;
[0174] Prediction uncertainty: Assess the confidence level of the target surfactant formulation recommendation model in predicting this point.
[0175] Based on prediction performance and prediction uncertainty, key regions are identified, namely, high-uncertainty regions: areas containing virtual points with the largest prediction variance are selected. These regions represent areas where the model's accuracy is relatively low due to a lack of training data. Performance boundary regions are also identified: areas containing virtual points with optimal prediction performance (e.g., the lowest predicted interface tension values). These regions are potential optimal solution boundaries (i.e., Pareto fronts) that require more accurate models to characterize. This results in two sets of target region coordinates that require focused attention: a) high-uncertainty regions; b) high-performance boundary regions.
[0176] For the target region coordinate set, select a moderate number (e.g., several dozen) of new experimental points based on strategies (such as uniform sampling, maximizing diversity, etc.). These points combine the goals of "exploration" (reducing uncertainty) and "utilization" (optimizing performance). Perform experiments to obtain real-world data: According to the generated experimental plan, actually prepare these new formulation points in the laboratory (or using experimental equipment) and measure their actual interfacial tension values. Form a new experimental data matrix containing the characteristics (ratios, conditions) of the new formulation points and their corresponding measured interfacial tension values.
[0177] Next, the newly added experimental data matrix and the experimental matrix training set will be merged to form an amplified training set. The target surfactant formulation recommendation model will then be trained to obtain a further improved final surfactant formulation recommendation model.
[0178] It should be noted that the high uncertainty region (the region with the largest prediction variance) refers to the interfacial tension predicted by the model as several random formulations, and at the same time, the confidence level (variance) of the model for each prediction is calculated. The formulation points with the highest prediction variance are identified, and the parameter space where these points are clustered together is the high variance region. The high performance boundary region refers to the same high uncertainty region, but the model prediction results are based on the predicted interfacial tension values (performance scores). The formulation points with the best prediction performance (lowest interfacial tension) are identified, and cluster analysis is performed to find the concentrated distribution area, which is the high performance boundary region.
[0179] The above method introduces a data augmentation and model iteration mechanism after training the initial prediction model, resulting in significant additional benefits. This scheme, through Monte Carlo simulation and uncertainty assessment of the model itself, can automatically and accurately identify the most ambiguous regions (high variance regions) and potentially optimal performance regions (high-performance boundary regions) in the current model's understanding. By fusing new data with existing data and retraining the model, the limitations of the initial training dataset are effectively overcome. This results in a model that not only performs well on known data points but also exhibits higher prediction accuracy and reliability in the global parameter space, especially near key optimal solution boundaries. This significantly improves the intelligence level of the entire formula recommendation system and the confidence level of the final output solution.
[0180] In summary, this application provides a method for training a surfactant formulation recommendation model. By integrating orthogonal experimental design, central composite design, and machine learning algorithms, it achieves efficient data acquisition, model training, prediction optimization, and closed-loop updates for surfactant systems. Based on experimental design theory, it utilizes an automated system to acquire experimental data in high throughput, and combines artificial intelligence algorithms to establish an interfacial tension prediction model and formulation recommendation mechanism. The model training and usage process forms an iterative closed loop of "experimental design, data acquisition, intelligent factor screening, central composite design expansion, AI optimization and formulation recommendation, model usage and dynamic updates," achieving optimal performance and maximized experimental efficiency for the surfactant system.
[0181] Secondly, this application provides a surfactant formulation recommendation model training system, applied to the aforementioned surfactant formulation recommendation model training method, the system comprising:
[0182] The generation unit is used to generate an initial experimental matrix based on the number of surfactant components in the target system using orthogonal experimental design. The target system represents the specific scope and final goal of this experiment. Each experimental unit in the experimental matrix includes different surfactant ratios, surfactant condition parameters, and measured interfacial tension values.
[0183] The analysis unit is used to perform variance analysis on the initial experimental matrix, identify factors that have an impact on the interfacial tension value greater than a set threshold, and generate experimental points based on the factors to obtain a supplementary experimental matrix generated based on the experimental points, wherein the experimental points represent experimental schemes formed by traversing and combining the factors.
[0184] The merging unit is used to merge the initial experimental matrix and the supplementary experimental matrix to obtain an experimental matrix training set.
[0185] The result unit is used to train the initial surfactant formulation recommendation model using the experimental matrix training set to obtain the target surfactant formulation recommendation model. The initial surfactant formulation recommendation model includes an initial main trend prediction sub-model and an initial residual correction sub-model. The initial residual correction sub-model is configured to take the prediction result of the initial main trend prediction sub-model as input and correct the residual of the prediction result. The target surfactant formulation recommendation model is configured to receive an input vector consisting of surfactant ratio and condition parameters and output the predicted interfacial tension value.
[0186] Thirdly, this application provides a computer storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the aforementioned method.
[0187] Fourthly, this application provides a computer program that, when executed by a processor, implements the steps of the aforementioned method.
[0188] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0189] The various embodiments in this disclosure are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0190] The scope of protection of this disclosure is not limited to the embodiments described above. Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from its scope and spirit. If such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, then the intent of this disclosure also includes such modifications and variations.
Claims
1. A method for training a surfactant formulation recommendation model, characterized in that, The method comprises: According to the number of components of the surfactant in the target system, an initial experiment matrix is generated by using orthogonal experiment design, wherein the target system represents the specific range and final target of the experiment, and each experiment unit in the experiment matrix includes different surfactant ratios, surfactant condition parameters and measured interfacial tension values; Variance analysis is performed on the initial experiment matrix to identify factors that have an impact on the interfacial tension value greater than a set threshold, and an experiment point is generated according to the factors to obtain a supplementary experiment matrix generated by experiments according to the experiment point, wherein the experiment point represents an experimental scheme formed by traversing combinations according to the factors; The initial experiment matrix and the supplementary experiment matrix are merged to obtain an experiment matrix training set; An initial surfactant formula recommendation model is trained using the experiment matrix training set to obtain a target surfactant formula recommendation model, wherein the initial surfactant formula recommendation model includes an initial main trend prediction sub-model and an initial residual correction sub-model, the initial residual correction sub-model is configured to take the prediction result of the initial main trend prediction sub-model as input and correct the residual of the prediction result, and the target surfactant formula recommendation model is configured to receive an input vector composed of surfactant ratios and condition parameters and output a predicted interfacial tension value; and The step of performing variance analysis on the initial experiment matrix to identify factors that have an impact on the interfacial tension value greater than a set threshold and generating an experiment point according to the factors to obtain a supplementary experiment matrix generated by experiments according to the experiment point comprises: According to the initial experiment matrix, variance analysis is performed on the interfacial tension value as the dependent variable and each surfactant ratio and condition parameter as the independent variable, the p values of each factor and the interaction between factors are calculated, and factors and interaction terms with p values less than a preset significance level are determined as factors that have a significant impact on the interfacial tension to obtain a factor list; According to the factor list, a central composite design method is used to collect the best response region of the factors in the experimental space to generate a supplementary experiment plan matrix of the experiment point set; Each experiment point in the supplementary experiment plan matrix is used for experiment, and the interfacial tension value of each new experiment point is measured and recorded to obtain a supplementary experiment matrix; The step of using the central composite design method to collect the best response region of the factors in the experimental space according to the factor list to generate a supplementary experiment plan matrix of the experiment point set comprises: According to the factor list, the upper and lower limits of each factor are defined, and the coordinate values of the corresponding central point and axial point are calculated to obtain a structured factor design space table; According to the structured factor design space table, an initial CCD point set containing experiment point coordinates is generated by using the mathematical model of central composite design, wherein the initial CCD point set includes a point set of cubic points, axial points and central point coordinates; The initial CCD point set is converted into a plan table to generate a supplementary experiment plan matrix of the experiment point set, wherein the running order of each experiment, the specific setting values of each factor are listed in the plan table.
2. The surfactant formulation recommendation model training method of claim 1, wherein, The step of generating the initial experiment matrix according to the number of components of the surfactant in the target system by using the orthogonal experiment design comprises: According to the number of components of the surfactant in the target system, determine the factors and levels of each factor, wherein the factors include the ratio of various surfactants, and the temperature, rotation speed and chemical form condition parameters, and the levels represent the discrete values preset for each factor; According to the factors and levels of each factor, select a matching orthogonal table from the standard orthogonal table library, and input the factors and levels into the selected orthogonal table to perform orthogonal experiment design and generate an initial experiment plan matrix; According to each experiment point in the initial experiment plan matrix, perform experiments, measure and record the interfacial tension value of each experiment point, and add the interfacial tension value as a new column to the initial experiment plan matrix to form the initial experiment matrix.
3. The surfactant formulation recommendation model training method of claim 1, wherein, The step of merging the initial experiment matrix and the supplementary experiment matrix to obtain the experiment matrix training set comprises: Align the initial experiment matrix and the supplementary experiment matrix to obtain two standardized experiment matrices with aligned structures; Merge the two standardized experiment matrices with aligned structures in the row direction to obtain a merged and deduplicated experiment data table; Take the surfactant ratio and condition parameter columns in the merged and deduplicated experiment data table as characteristic variables, and take the interfacial tension value column as a target variable to obtain the experiment matrix training set.
4. The surfactant formulation recommendation model training method of claim 1, wherein, The step of training the initial surfactant formula recommendation model using the experiment matrix training set to obtain the target surfactant formula recommendation model comprises: Train the initial main trend prediction sub-model using the data in the experiment matrix training set to obtain a target main trend prediction sub-model, wherein the target main trend prediction sub-model represents the output of the preliminary interfacial tension prediction value according to the input formula; Based on the prediction result of the target main trend prediction sub-model on the experiment matrix training set, calculate the prediction residual, and train the initial residual correction sub-model using the data in the experiment matrix training set as input and the prediction residual as a training target to obtain a target residual correction sub-model; Combine the target main trend prediction sub-model and the target residual correction sub-model to obtain a target surfactant formula recommendation model, wherein the output of the target surfactant formula recommendation model is the sum of the output of the target main trend prediction sub-model and the output of the target residual correction sub-model.
5. The surfactant formulation recommendation model training method of claim 1, wherein, After the step of training the initial surfactant formula recommendation model using the experiment matrix training set to obtain the target surfactant formula recommendation model, the method further comprises: Input the user preferences into the target surfactant formula recommendation model as conditions to form multiple optimal solutions and obtain a candidate formula list, wherein the user preferences represent the interfacial tension performance preferences and the cost priority preferences.
6. A surfactant formulation recommendation model training system characterized by, The system is applied to the surfactant formula recommendation model training method in any one of claims 1-5. The generating unit is configured to generate an initial experiment matrix according to the number of components of the surfactant in a target system by using an orthogonal test design, wherein the target system represents a specific range and a final target of the experiment, and each experiment unit in the experiment matrix includes different surfactant ratios, surfactant condition parameters, and measured interfacial tension values. The analyzing unit is configured to perform variance analysis on the initial experiment matrix, identify factors that have an impact on the interfacial tension values greater than a set threshold, and generate experiment points according to the factors to obtain a supplementary experiment matrix generated by experiments according to the experiment points, wherein the experiment points represent experiment schemes formed by traversing combinations according to the factors. The merging unit is configured to merge the initial experiment matrix and the supplementary experiment matrix to obtain an experiment matrix training set. The result unit is configured to train an initial surfactant formula recommendation model by using the experiment matrix training set to obtain a target surfactant formula recommendation model, wherein the initial surfactant formula recommendation model includes an initial main trend prediction sub-model and an initial residual correction sub-model, the initial residual correction sub-model is configured to take a prediction result of the initial main trend prediction sub-model as an input and correct a residual of the prediction result, and the target surfactant formula recommendation model is configured to receive an input vector composed of surfactant ratios and condition parameters and output a predicted interfacial tension value.
7. A computer storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 5.
8. A computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Prediction method and device, electronic equipment and storage medium
CN110867254A
Model parameter adjustment method and device based on an orthogonal experiment, and storage medium
CN112541588A