Constrained multi-objective optimization method for chemical process
By combining a dual-population multi-objective optimization algorithm with differential evolution and NSGA-II, and using an SVR surrogate model and K-means clustering, the problem of balancing convergence and diversity in chemical processes was solved, achieving efficient optimization of chemical processes and environmentally friendly production decisions.
Patent Information
- Application Number
- CN202511742452.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-02-13
AI Technical Summary
Traditional multi-objective optimization algorithms struggle to balance convergence and diversity in chemical processes, and their ability to handle infeasible solutions is insufficient, failing to meet the optimization needs of complex chemical processes.
A multi-objective optimization algorithm based on a dual-population model is adopted, which combines differential evolution and non-dominated sorting genetic algorithm (NSGA-II) with support vector regression (SVR) surrogate model. Parents are selected by reference vector and feasibility priority principle, and the auxiliary population is updated by combining K-means clustering and potential score to achieve multi-objective optimization.
It improves the convergence and versatility of the algorithm in chemical processes, enabling energy conservation and emission reduction while meeting production requirements, and providing better decision support.
Smart Images

Figure CN121525508A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of multi-objective optimization, and particularly relates to a multi-objective optimization method with constraints for a chemical process. BACKGROUND
[0002] Optimization of a chemical process is related to product quality, pollutant emission and system energy consumption, and is of great significance to the realization of the "double carbon" goal. The chemical process often has characteristics such as nonlinearity and multivariable coupling, and is related to multiple process parameters. The optimization target is often conflicting and limited by constraints, and the contradictory relationship between the targets requires the decision maker to select a satisfactory compromise solution among multiple targets under the premise of meeting the constraints, which is challenging. Traditional process optimization mainly involves equipment modification and production process improvement, both of which are costly and difficult to further optimize. In recent years, artificial intelligence has developed rapidly and has become a widely cross-cutting frontier discipline. More and more researchers combine multi-objective optimization algorithms with the chemical industry and have achieved good results. However, as the chemical process becomes more and more complex and the coupling between equipment becomes stronger, traditional multi-objective optimization algorithms cannot meet the demand. The main problems include that the convergence and diversity of the algorithm cannot be well balanced, the handling of infeasible solutions is too simple, and the performance of the algorithm is insufficient to meet the requirements. Therefore, improving the multi-objective optimization algorithm to adapt to the increasingly complex chemical process is one of the hot research contents in this field. SUMMARY
[0003] In view of the problems existing in the above-mentioned existing method, a constraint multi-objective optimization algorithm for a chemical process is provided. First, a sulfuric acid data set of sulfur production in a factory is collected, the decision variables are determined, and the SO2 conversion rate, acid mist emission and sulfuric acid production are used as optimization objectives. The data set is preprocessed, missing values are checked and removed, duplicate values are removed, abnormal values are detected and removed, and data normalization is performed. In order to balance convergence and diversity, a double-population-based multi-objective optimization algorithm is used to perform multi-objective optimization of sulfuric acid production from sulfur. The main population mainly stores feasible solutions to ensure the convergence of the algorithm, and the auxiliary population collects high-quality infeasible solutions to ensure the diversity of the algorithm. First, parent selection is performed, and reference vectors are used to promote the diversity of selection. By associating the solution with the reference vector and considering the feasibility and perpendicular distance, it is ensured that the selected parent is uniformly distributed in the target space. Second, the DE algorithm is used to generate offspring, and then the CV-based NSGA-II algorithm is used to update the main population, and the feasible solution is preferentially retained. Finally, the auxiliary population is updated using K-means clustering combined with the defined "potential score", and the "potential" infeasible solution is retained. The inverse generation distance IGD and the hyper volume HV are used as evaluation indexes. The method proposed in the application realizes smaller IGD and larger HV in multi-objective optimization, which indicates that the algorithm has good convergence and diversity, provides a reference for engineers to set operating conditions, and helps to save energy and increase efficiency while meeting production requirements and reducing environmental pollution.
[0004] To achieve the above technical purposes, the present application provides the following technical solutions:
[0005] The constraint multi-objective optimization algorithm for a chemical process provided by the application specifically comprises the following steps:
[0006] S1, collecting the steady-state data of sulfuric acid production from sulfur in a factory, including 11 input decision variables, 3 target variables and the mass fraction of produced sulfuric acid, the target variables are the content of discharged acid mist, the SO2 conversion rate and the sulfuric acid production, the mass fraction of sulfuric acid is a constraint condition, and the remaining two constraint conditions are the inlet temperatures of the converter three sections and the converter four sections in the decision variables.
[0007] S2, pre-processing the data, including checking and removing missing values, removing duplicate values, abnormal value detection and removal, and data normalization;
[0008] S3, selecting support vector regression SVR as the proxy model of the sulfuric acid production from sulfur process, dividing the training set and the test set using the pre-processed data, training the SVR model using the training set, and using the trained SVR model for the test set to test the model performance;
[0009] S4, multi-objective optimization is carried out on the agent model. The decision variables of the model and the corresponding objective function and constraint condition are regarded as each individual in the population. The population is divided into a main population and an auxiliary population and the two populations are initialized respectively, and then the main population and the auxiliary population are combined, and the mating parent is selected by combining the reference vector and the feasibility priority principle; the offspring is generated by using the mutation, crossover and selection operations of the differential evolution algorithm (DE); the main population is updated by using the NSGA-II based on CV; the potential score P is defined and the archive population is updated by combining K-means clustering. The steps of selecting mating parents, generating offspring, updating the main population and updating the auxiliary population are cycled until the maximum evolution generation is reached.
[0010] S5, a set of Pareto optimal solutions is generated, and the site operator selects appropriate input decision variables as operating conditions for production according to the actual situation in the optimal solution set.
[0011] Beneficial effects: the application uses a double-population evolution mechanism, which can balance the convergence and diversity of the algorithm. When selecting parents, the multi-resolution reference vector is combined with the feasibility priority, which ensures the quality of the solution and the uniformity of the solution in the target space. When updating the auxiliary population, the concept of "potential score" is introduced, combined with K-means clustering, which ensures the diversity of the auxiliary population solution and provides direction for population evolution. Compared with the prior art, the application has different degrees of improvement in balancing the convergence and diversity of the algorithm, processing constraints and improving the performance of the algorithm. BRIEF DESCRIPTION OF DRAWINGS
[0012] The above and / or additional aspects and advantages of the application will become apparent and more readily appreciated from the following description of the embodiments, taken in conjunction with the accompanying drawings, in which:
[0013] Figure 1 The overall flowchart of the method proposed by the application is shown in the figure;
[0014] Figure 2 The population scatter plot obtained by the algorithm on the sulfuric acid system for producing sulfuric acid from sulfur is shown in the figure;
[0015] Figure 3 The change trend of the inverse generation distance IGD in the evolution process is shown in the figure;
[0016] Figure 4 The change trend of the hyper volume HV in the evolution process is shown in the figure. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical scheme and advantages of the application more clear and obvious, the application is further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the application and not to limit the application.
[0018] It is to be understood that the term "and / or", as used herein in the specification and in the claims, is used to mean one or the other of the listed terms or both of the listed terms.
[0019] The overall flow of the method is shown in Figure 1 The specific implementation steps are as follows:
[0020] S1, collect the sulfuric acid production data of a sulfur plant in a steady state, including 11 input decision variables, 3 objective variables and the mass fraction of the produced sulfuric acid, the objective variables being the content of acid mist emission, unit kg / hr; SO2 conversion rate, unit %; sulfuric acid production, unit kg / hr. The constraint conditions are the mass fraction of finished sulfuric acid, unit %; the inlet temperature of the converter three-stage, unit ℃; the inlet temperature of the converter four-stage, unit ℃. Table 1 and Table 2 are respectively the value range of the 11 decision variables and the constraint conditions in this implementation case:
[0021] Table 1 Value range of decision variables
[0022]
[0023] Table 2 Constraint conditions
[0024] Constrained variable Range Converter three-stage inlet temperature (°C) <435 Converter four-stage inlet temperature (°C) <410 Finished sulfuric acid mass fraction (%) ≥98.5
[0025] S2, perform pre-processing operations on the data, including checking and removing missing values, removing duplicate values, detecting and removing outliers, and data normalization;
[0026] Use quartile range to detect outliers. First, arrange the data from small to large, and the values of 25%, 50% and 75% of the data set are denoted as Q1, Q2 and Q3 respectively. IQR can be expressed as IQR = Q3-Q1. For data point x, if Q1-1.5IQR < x < Q3+1.5IQR, then x is considered to be a normal point, otherwise x is considered to be an outlier. Use Min-Max to normalize the data, and the Min-Max normalization formula is:
[0027]
[0028] wherein, denotes the jth value of the ith feature after normalization; denotes the jth value of the ith feature before normalization; max(x i ) denotes the maximum value of the ith feature before normalization; min(x i ) denotes the minimum value of the ith feature before normalization.
[0029] S3, Select support vector regression (SVR) as the proxy model of the sulfuric acid production process from sulfur, use the preprocessed data, divide the first 80% of the sample set into the training set, and the last 20% of the sample set into the test set, use the training set to train the SVR model, and the trained SVR model is used for the test set to test the model performance;
[0030] SVR is selected as the proxy model. SVR maps the input data into high-dimensional space by introducing a kernel function, and determines a fitting function f(x) such that the error between the predicted value and the true value of the data point is minimized. The fitting function f(x) of the SVR model is defined as:
[0031]
[0032] where f(x) is the predicted value of the output variable, represents the kernel function, and ω and b represent the weight vector and the bias term, respectively.
[0033] The common kernel functions of SVR are radial basis function (RBF), polynomial kernel function and sigmoid kernel function. The radial basis function RBF is selected in this embodiment. The radial basis function has an important parameter τ, which controls the mapping distribution of the input data in the high-dimensional space. If the value of τ is too large, the model will be underfitting, which will reduce the prediction accuracy. If the value of τ is too small, the model training time will be longer.
[0034] SVR introduces a linear insensitive loss function ε to enhance the robustness of the model. ε represents the tolerance error, i.e. the acceptable range of the prediction error between the predicted value and the true value of the data point. When the difference between the true value y and the predicted value f(x) is |y-f(x)|>ε, it is considered that the predicted value has an error, and when |y-f(x)|≤ε, it is considered that the predicted value meets the accuracy requirement, thus the loss function L of SVR is defined as:
[0035]
[0036] For the samples with prediction error greater than ε, a penalty factor C is defined to punish it, and the size of C represents the degree of punishment for the samples that do not meet the prediction accuracy.
[0037] In this embodiment, for the three target variables, the values of the three parameters τ, ε and C of SVR are shown in Table 3:
[0038] Table 3 Values of SVR model hyperparameters
[0039] Target variable τ ε C Acid mist emission 2.868 0.023 14.388 SO2 conversion 5.901 0.0007 1042.6 Sulfuric acid production 4.77 0.0129 296.77
[0040] S4, multi-objective optimization is performed on the agent model. The decision variables and corresponding objective functions and constraint conditions of the model are regarded as each individual in the population. The population is divided into a main population and an auxiliary population and the two populations are initialized respectively, and then the main population and the auxiliary population are combined, and the mating parent is selected by combining the reference vector and the feasibility priority principle; the offspring is generated by using the mutation, crossover and selection operations of the differential evolution algorithm (DE); the main population is updated by using the NSGA-II based on CV; the potential score P is defined and the archive population is updated by combining the K-means clustering. The steps of selecting the mating parent, generating the offspring, updating the main population and updating the auxiliary population are repeated until the maximum evolution generation is reached;
[0041] For a multi-objective optimization problem, it can be abstracted as the following mathematical expression:
[0042] minF(x)=(f1(x),f2(x),...,f m (x)) T
[0043] s.t.g i (x)≤0,i=1,2,...,p
[0044] x min ≤x≤x max
[0045] Wherein, x=(x1,x2,...,x n ) represents the decision variable with a dimension of n, F(x)=(f1(x),f2(x),...,f m (x)) T represents the objective vector with a dimension of m, in the embodiment, m is 3, f1(x), f2(x) and f3(x) are acid mist emission, SO2 conversion rate and sulfuric acid production respectively. g is a constraint condition, p is the number of constraint conditions, x max and x min represent the upper limit and the lower limit of the decision variable respectively.
[0046] Step S4 specifically includes:
[0047] S4-1, the population is divided into two sub-populations, namely a main population and an auxiliary population, the main population mainly stores feasible solutions, and the auxiliary population stores high-quality infeasible solutions. The main population is initialized:
[0048]
[0049] Wherein, X represents an individual in the population, V max and V min represent the upper limit and the lower limit of the decision space respectively, and n represents the number of decision variables, which is 11 in the embodiment.
[0050] S4-2, after initialization, the reference vector method and the feasibility priority principle are combined to select the mating parents from the main population and the auxiliary population. The reference vector method refers to generating a set of uniformly distributed reference vectors, each of which represents a direction in the target space. The diversity of the parent population is maintained by dividing the target space into multiple sub-regions. The reference vector is a unit vector, that is:
[0051] v = (v1, v2,..., vm) m ), ||v||2 = 1
[0052] Where m represents the number of target variables, which is 3 in this embodiment.
[0053] After generating the reference vector, the cosine similarity between the candidate solution vector and the reference vector is calculated, and each candidate solution is assigned to the reference vector with the maximum cosine similarity. The formula for calculating the cosine similarity is:
[0054] cos(θ) = (u·v) / (||u||·||v||)
[0055] Where cos(θ) represents the cosine similarity, u and v represent the solution vector and the reference vector respectively, and ||·|| represents the Euclidean norm of the vector.
[0056] Different solution vectors are assigned to corresponding reference vectors to form several vector sets. Traverse each vector set, first select a feasible solution vector in the vector set with the smallest perpendicular distance to the reference vector as the parent individual. If there is no feasible solution in the vector set, select the infeasible solution with the smallest CV value. The formula for calculating the perpendicular distance is:
[0057] d = ||u-(u·v)v||
[0058] The formula for calculating CV is:
[0059]
[0060] CV i (x) = max(0, g i (x))
[0061] From the formula, when CV = 0, the solution x satisfies the constraint and is a feasible solution, otherwise it is an infeasible solution.
[0062] If all vector sets are traversed and the selected individuals are not enough to fill the parent population, the remaining candidate solutions are selected based on the following formula:
[0063] score = CV + 1×10 -6 d
[0064] Select individuals with small score values until the parent population is filled.
[0065] When performing mating parent selection, the number of reference vectors is not fixed during the evolution process, but changes with the evolution generation number, so a multi-resolution reference vector is innovatively proposed based on the reference vector. The multi-resolution reference vector changes the number of reference vectors by changing the resolution value of the population at different evolution stages. The resolution formula is:
[0066]
[0067] Wherein, a represents the resolution, gen represents the evolution generation number, Maxgen represents the maximum evolution generation number, a1, a2 and a3 are constants. In this embodiment, a1, a2 and a3 are respectively taken as 0.4, 0.7 and 1.
[0068] The number of reference vectors at different evolution stages is calculated by the formula:
[0069] N ref = a(gen) numref + n obj
[0070] Wherein, N ref represents the number of reference vectors, numref represents the number of reference vectors, n obj represents the number of boundary reference vectors, which is used to ensure that the boundary direction of the solution space has a special search vector to cover the extreme direction.
[0071] In the early stage of population evolution, the number of reference vectors is small, at this time the algorithm mainly performs large-scale global search to find the feasible solution as soon as possible, which is conducive to the convergence of the algorithm. With the evolution, the number of reference vectors gradually increases, covering more directions, and tends to fine search, which is conducive to uniformly covering the Pareto front and improving the diversity of solutions.
[0072] S4-3, after completing the selection of mating parents, then select offspring from the mating parents. The differential evolution algorithm DE is used to select offspring. DE improves the diversity of the population and selects offspring by performing mutation, crossover and selection operations.
[0073] Mutation: Perform mutation operation in the parent population to generate new individuals. Taking the tth evolution as an example, the formula of mutation operation is:
[0074]
[0075] Wherein, i represents the index of the ith individual in the population, r1, r2 and r3 are three random numbers, and These are the r1, r2, and r3 particles, respectively, and F is the scaling factor, which is 0.7 in this embodiment.
[0076] Crossover: After generating mutated individuals, a crossover operation is performed. The formula for the crossover operation is as follows:
[0077]
[0078] Where k∈{1,2,3,...,n} represents the k-th dimension of the i-th individual, rand(k)∈[0,1] is a random number, rand(k)∈{1,2,3,...,n} is also a random number, and CR represents the crossover probability, which is 0.7 in this embodiment.
[0079] Selection: If the mutated individual is superior, replace the current offspring; otherwise, maintain the original generation. Based on the principle of feasibility first, select feasible individuals. If both the mutated individual and the current offspring are feasible, compare their dominance relationship.
[0080] S4-4. After performing mating parent selection and offspring generation operations, update the main population. Update the main population using non-dominated sorting. Taking the minimization of the objective function as an example, for two individuals x1 and x2 in the feasible solution set, if they simultaneously satisfy (1): f i (x1)≤f i (x2); (2) f j (x1)<f j If x1 dominates x2, denoted as x1 < x2, then solution x1 is superior to x2. Otherwise, x1 and x2 are not mutually dominant, and are called mutually non-dominant solutions. Non-dominant ordering refers to dividing candidate solutions into several Pareto fronts based on their dominance. Solutions in front 1 are not dominated by other solutions and have the highest priority. Solutions in front 2 are not dominated by any other solutions besides front 1, and so on. When there are enough feasible individuals in the original main population, non-dominant ordering is performed on the feasible individuals. Solutions in front 1, front 2, ..., front j are sequentially filled into the main population. When selecting the last front, crowding ordering is used to select some individuals from it to fill the main population. The formula for calculating crowding ordering is expressed as:
[0081]
[0082] Where M is the number of target variables, and Let represent the target values of the two individuals adjacent to the candidate solutions after ascending order of the m-th target. and respectively, represent the maximum and minimum values of the front on the mth objective. Individuals with larger crowding distances are preferred to be reserved.
[0083] If there are not enough feasible solutions in the original main population, all the feasible solutions are selected as the updated main population, and then the individuals with smaller CV values are selected from the infeasible solutions to fill the main population.
[0084] S4-5, define the "potential score", and update the auxiliary population combined with K-means clustering. The auxiliary population stores infeasible solutions with "potential". The concept of "potential score" is proposed, which measures whether a candidate solution has potential from two aspects. On the one hand, the degree of constraint violation CV, the smaller the value, the greater the "potential" of the candidate solution. On the other hand, the quality of the solution. The quality of the solution is obtained by calculating the Euclidean distance between the objective function value corresponding to the candidate solution and the ideal optimal objective function value. Let the objective function value corresponding to the ith candidate solution be U i =(u1, u2,..., u m ), and the ideal optimal objective function value be O best =(o1, o2,..., o m ), then the calculation formula of the Euclidean distance is:
[0085]
[0086] The smaller d is, the higher the quality of the candidate solution is.
[0087] The potential score P is based on CV and d, and its formula is:
[0088] P = lambda (1-CV) + (1-lambda) (1-d)
[0089] Where P is the potential score, lambda is the weight coefficient, which is a constant, and measures the weight size of CV and d. In this embodiment, lambda is taken as 0.5. The larger P is, the more potential the candidate solution has.
[0090] Take the potential score P of the candidate solution as the screening index, and update the auxiliary population combined with the K-means clustering operation. The clustering operation is to divide the candidate solutions into several clusters, so that the candidate solutions in the same cluster are as close to each other as possible, and the candidate solutions in different clusters are different, so as to ensure the diversity of the updated auxiliary population. Its goal is to minimize the squared error within the cluster, which is expressed by the formula:
[0091]
[0092] Where n represents the number of candidate solutions, K represents the number of clusters, x i represents the ith candidate solution, σ j represents the centroid of the jth cluster, and r ijxj = 1 if x is assigned to the jth cluster, otherwise xj = 0. i rj = 1 if x is assigned to the jth cluster, otherwise rj = 0. ij
[0093] When performing the clustering operation, first the centroid of each cluster is initialized. After initialization, the candidate solutions are assigned to different clusters using Euclidean distance, at this time the centroid is updated, the centroid of each cluster is recalculated, the centroid calculation formula is:
[0094]
[0095] The assignment of candidate solutions and the updating of the centroid are repeatedly repeated until the centroid no longer changes, at this time the centroid is the center of the corresponding cluster.
[0096] When updating the auxiliary population using the "potential score" and K-means clustering, first determine the center of the cluster, and divide the candidate solutions into different clusters. For each cluster, select the candidate solution with the highest potential score P in the cluster as an individual in the updated auxiliary population. If the population cannot be filled, continue to select the second highest potential score solution in each cluster, and so on, until the auxiliary population is filled, and the auxiliary population updating is completed.
[0097] S5, using the multi-objective optimization algorithm described in step S4, the candidate solution in the input decision variable space is iterated to obtain the Pareto optimal solution set, and the termination condition of iteration is to reach the preset maximum number of generations. The field operator can select the appropriate operating conditions for production according to the actual needs in the Pareto optimal solution set. The population size, maximum evolution number, resolution and other parameters of the algorithm can be adjusted by the implementer according to experience.
[0098] The results obtained by applying the multi-objective optimization algorithm described in step S4 to the production of sulfuric acid from sulfur are shown in Tables 4 and Figure 2 , Figure 3 , Figure 4 In this embodiment, the size of the main population is 200, the size of the auxiliary population is 160, and the evolution number is 150. Table 4 is part of the optimal solution set in the Pareto optimal solution set and the corresponding target variable.
[0099]
[0100] Figure 2 The algorithm is shown in the figure. Each point in the figure represents a Pareto optimal solution, corresponding to a specific set of decision variables, and the three-dimensional coordinate values are the specific performance of the solution on the three objectives.
[0101] Figure 3 and Figure 4 Inverse Generational Distance (IGD) and Hypervolume (HV) in the process of population evolution, respectively. The calculation formula of IGD is expressed as:
[0102]
[0103] wherein, represents the reference Pareto front solution set, represents the approximate Pareto front solution set obtained by the algorithm, |P * represents the number of solutions in the reference Pareto front solution set, |A * represents the number of solutions in the approximate Pareto front solution set, p and a represent the solutions in the reference Pareto front solution set and the approximate Pareto front solution set, respectively, and ||·|| represents the Euclidean norm. The smaller the value of IGD, the better the convergence and distribution of the solution set obtained by the algorithm.
[0104] The calculation formula of HV is expressed as:
[0105]
[0106] wherein, z * is the reference point, X represents the solution set composed of non-dominated solutions in A * , ψ represents the volume of the hypercube surrounded by X and z * , and ∪ represents the union set. The larger the value of HV, the better the solution quality of the algorithm.
[0107] It can be seen that IGD always remains at a low level in the evolution process, and HV becomes larger and larger with the increase of the evolution generation, and finally maintains at a high level. This shows that the Pareto front calculated by the algorithm is very close to the real Pareto front of the embodiment and can uniformly cover the Pareto front surface, indicating that the algorithm has good performance in finding optimal solutions, and the Pareto optimal solution found has high reference value in actual engineering.
Claims
1. A constrained multi-objective optimization method for chemical processes, characterized in that, Specifically, the following steps are included: S1. Collect steady-state data of sulfuric acid production from sulfur in a certain factory, including 11 input decision variables, 3 target variables, and the mass fraction of sulfuric acid produced. The target variables are the content of emitted acid mist, SO2 conversion rate, and sulfuric acid yield. The mass fraction of sulfuric acid is a constraint condition. The other two constraints are the inlet temperature of the third stage and the inlet temperature of the fourth stage of the converter in the decision variables. The 11 input decision variables represent the operating conditions related to sulfuric acid production from sulfur, and are denoted as input 1, input 2, ..., input 11. S2. Perform preprocessing operations on the data, including checking and removing missing values, removing duplicate values, detecting and removing outliers, and normalizing the data. S3. Support Vector Regression (SVR) was selected as the proxy model for the sulfur-to-sulfuric acid production process. The preprocessed data was used to divide the data into training and test sets. The training set was used to train the SVR model, and the trained SVR model was used on the test set to test the model performance. S4. Perform multi-objective optimization on the surrogate model, treating the model's decision variables, corresponding objective functions, and constraints as each individual in the population; The population is divided into a primary population and a secondary population, and the two populations are initialized separately. Then, the primary and secondary populations are merged, and mating parents are selected based on reference vectors and feasibility priority principles. Offspring are generated using mutation, crossover, and selection operations of the differential evolution algorithm (DE). The primary population is updated using CV-based NSGA-II. A potential score P is defined and updated using K-means clustering. The process of selecting mating parents, generating offspring, updating the primary population, and updating the secondary population is repeated until the maximum number of generations is reached. S5. Generate the Pareto optimal solution set. On-site operators select appropriate input decision variables from the optimal solution set as operating conditions for production based on the actual situation.
2. The method according to claim 1, characterized in that... The 11 input decision variables in step S1 are represented as follows: Input 1 indicates the inlet temperature of the air feed for sulfuric acid production from sulfur. Input 2 indicates the pressure increase of the blower in the sulfur production process; Input 3 indicates the operating pressure of the sulfur reactor for sulfuric acid production from sulfur. Input 4 indicates the SO2 content at the outlet of the sulfur reactor used for sulfur production from sulfur. Inputting 5 indicates the O2 content in the air fed into the sulfuric acid production process from sulfur. Input 6 indicates the inlet temperature of the three sections of the converter; Input 7 indicates the inlet temperature of the four sections of the converter; Input 8 indicates the operating pressure of the four stages of the converter; Input 9 to indicate the outlet flow rate of the sulfur reactor; Inputting 10 indicates the outlet temperature of the sulfur reactor; Inputting 11 indicates the flow rate of sulfur in the sulfur production process.
3. The method according to claim 1, characterized in that... In step S2, outlier detection and removal specifically involve: Outlier detection is performed using the interquartile range (IQR). First, the data are arranged from smallest to largest. The 25%, 50%, and 75% values of this group of data are denoted as Q1, Q2, and Q3, respectively. IQR is expressed as IQR = Q3 - Q1. For a data point x, if Q1 - 1.5IQR < x < Q3 + 1.5IQR, then x is considered a normal point; otherwise, x is considered an outlier.
4. The method according to claim 1, characterized in that... In step S3, the fitting function f(x) of the SVR model is defined as: Where f(x) is the predicted value of the output variable. This represents the kernel function, where ω and b represent the weight vector and bias term, respectively. The radial basis function (RBF) was selected as the kernel function. In SVR, the values of the kernel function's scaling parameter τ, insensitive loss bandwidth ε, and penalty factor C are as follows:
5. The method according to claim 1, characterized in that... A multi-objective optimization problem with minimizing constraints can be abstracted as follows: minF(x)=(f1(x),f2(x),...,f m (x)) T s.t.g i (x)≤0,i=1,2,...,p x min ≤x≤x max Where x = (x1, x2, ..., x n Let f(x) represent a decision variable of dimension n, where F(x) = (f1(x), f2(x), ..., fn(x)). m (x)) T Let f1(x), f2(x), and f3(x) represent the target vector of dimension m, where f1(x), f2(x), and f3(x) are the acid mist emission, SO2 conversion rate, and sulfuric acid production, respectively; g represents the constraints, p represents the number of constraints, and x represents the target vector of dimension m. max and x min These represent the upper and lower limits of the decision variable, respectively.
6. The method according to claim 1, characterized in that... The method for selecting the mating parent in step S4 is as follows: A reference vector is introduced, and the number of reference vectors is dynamically adjusted at different stages of evolution. Candidate solutions are associated with the nearest reference vector through an inner product to form a reference vector set. The reference vector set is traversed, and for each reference vector set, in combination with the principle of feasibility priority, the feasible solution that is closest to the reference vector in vertical distance is selected first, and finally the mating pool is filled.
7. The method according to claim 6, characterized in that... Multi-resolution reference vectors change the number of reference vectors by altering the resolution values at different evolutionary stages of the population. The formula for resolution is expressed as: Where α represents the resolution, gen represents the number of generations, Maxgen represents the maximum number of generations, and α1, α2, and α3 are constants. In this embodiment, α1, α2, and α3 are taken as 0.4, 0.7, and 1, respectively.
8. The method according to claim 1, characterized in that... The method for updating the main population in step S4 is as follows: First, N feasible solutions are selected using the non-dominated sorting and crowding calculation of NSGA-II. If the number of feasible solutions is less than N, all feasible solutions are selected first, and then individuals with the smallest constraint violation are selected from the infeasible solutions to make up the N. When using NSGA-II, the feasible solutions are divided into multiple fronts using the non-dominated sorting, and the selection is carried out in the order of the fronts until the number exceeds N. When selecting the last front, some individuals are selected from it using the crowding sorting.
9. The method according to claim 1, characterized in that... The method for updating the auxiliary population in step S4 is as follows: Introducing the potential score P, the formula is as follows: P=lambda(1-CV)+(1-lambda)(1-d) Where P represents the potential score; lambda is a parameter used to measure the importance of constraint violation and solution quality; CV represents the constraint violation of each candidate solution, the smaller the value, the lower the constraint violation of the solution; d represents the distance between each solution and the ideal optimal solution, the smaller the value, the closer the solution is to the ideal optimal solution and the better the quality of the solution; the larger P is, the more "potential" the corresponding solution has. After calculating the P-value for each solution, determine the number of clusters K and assign the candidate solutions to the nearest clusters according to Euclidean distance. During the update, first select the solution with the highest P-value in each cluster. If the auxiliary population is not full, continue to select the second best solution in each cluster. If it is still not full, sort the remaining unselected solutions according to their P-values and select the solutions with higher P-values until the auxiliary population is full.
10. The method according to claim 1, characterized in that... In step S5, specifically: Based on the objective function and constraints defined in the above steps, the multi-objective optimization algorithm described in step S4 is used to iterate candidate solutions in the input decision variable space to obtain the Pareto optimal solution set. The termination condition of the iteration is reaching the preset maximum number of generations. On-site operators select appropriate operating conditions from the Pareto optimal solution set for production according to actual needs. The population size, maximum number of generations, and resolution parameters of the algorithm are adjusted by the implementers based on experience.