Machine learning based optimization method for wild canola blanching process parameters
Patent Information
- Application Number
- CN202610792993.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-03
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]为了解决现有技术中搜索边界同步收缩适应性不足、均匀采样难以聚焦优选区域以及迭代后期易陷入局部最优的问题,本发明提出一种基于机器学习的野油菜漂烫工艺参数优化方法
本发明提供了一种野油菜漂烫工艺参数优化方法,用于缓解漂烫过程中灭酶效果与营养物质保留难以兼顾的问题。通过构建涵盖过氧化物酶残留率、叶绿素保留率及维生素C保留率的多目标评价体系,能够对漂烫工艺效果进行综合量化评估。在迭代寻优过程中,引入归一化信息熵对各工艺参数维度的搜索边界进行独立调节,并结合核密度估计与拟蒙特卡罗序列映射,使生成的新样本更倾向分布于优质解密集区域,从而提高参数搜索精度与整体收敛效率。此外,本发明通过监测超体积增量引入多样性增强机制,配合重尾t分布模型进行拓展采样,缓解常规寻优过程容易陷入局部最优的问题,有助于保持搜索后期的样本多样性,并提高漂烫工艺参数优化结果的稳定性。
Smart Images

Figure CN122818892A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of parameter optimization, and in particular relates to a method for optimizing the parameters of the blanching process of wild rapeseed based on machine learning. Background Technology
[0002] Wild rapeseed is rich in nutrients such as vitamin C and chlorophyll, possessing high nutritional value and potential for food development. However, during processing and storage, wild rapeseed is susceptible to enzymatic browning caused by endogenous enzymes such as peroxidase, resulting in color deterioration and nutrient loss. Therefore, blanching is typically an important pretreatment step before processing wild rapeseed, used to reduce the activity of relevant enzymes through moderate heat treatment and delay subsequent quality deterioration. In actual production, there are significant multi-objective constraints among process parameters such as blanching temperature, blanching time, and material-to-water ratio. Higher temperatures or longer blanching times, while beneficial for reducing peroxidase residue and improving enzyme inactivation, can easily lead to vitamin C degradation and chlorophyll loss. Conversely, milder blanching conditions help maintain nutrients and a green appearance but may result in insufficient enzyme inactivation, affecting product storage stability. Therefore, achieving an optimal balance between enzyme activity inhibition, nutrient retention, and color preservation is a key issue that needs to be addressed in optimizing the blanching process for wild rapeseed.
[0003] Chinese patent application CN121835379A discloses a multi-objective parameter optimization method for hydropower units based on active learning and preference guidance, comprising: S1. Constructing a nonlinear simulation model based on the mechanism of hydropower units to build a multi-objective optimization problem; S2. Initializing the population, performing the first round of multi-objective optimization, and obtaining an initial Pareto front solution set; S3. Selecting a representative subset of solutions from the current Pareto front solution set, performing preference labeling, and generating a reference point based on the labeling results; S4. Incorporating preference distance into the crowding calculation to guide the population to evolve towards the reference point, thereby obtaining an updated Pareto front solution set; S5. Training a preference surrogate model based on historical labeled data, and automatically generating a reference point by the model when the model confidence reaches a threshold; S6. Terminating the optimization if the termination condition is met, and outputting the current Pareto optimal solution set and an optimization process report.
[0004] In the optimization of blanching process parameters for wild rapeseed, to address the multi-objective parameter optimization problem, the field of food engineering has gradually introduced multi-objective evolutionary algorithms and intelligent optimization methods. These methods construct evaluation functions to find non-dominated solution sets within the process parameter search space. However, existing techniques typically employ simultaneous compression of each parameter dimension during search boundary shrinkage, making it difficult to independently adjust based on the concentration of Pareto front solutions across different parameter dimensions. This can easily reduce search efficiency or prematurely eliminate potential preferred regions. When generating new samples within the updated search boundary, uniform random sampling is often used, failing to fully utilize the probability distribution characteristics reflected by the dense regions of current preferred solutions, resulting in insufficient local fine-grained search capabilities. As iterations progress and the search boundary continues to shrink, population diversity may decrease, and the optimization process can easily fall into local optima. When the Pareto front stagnates, failure to compensate for directional perturbations by incorporating the spatial principal component structure of the current solution set will affect the global optimization effect and result stability of the blanching process parameter combination. Summary of the Invention
[0005] To address the problems of insufficient adaptability of synchronous shrinkage of search boundaries, difficulty in focusing on the optimal region by uniform sampling, and easy getting trapped in local optima in the later stages of iteration in existing technologies, this invention proposes a machine learning-based method for optimizing the parameters of wild rapeseed blanching process.
[0006] This invention proposes a machine learning-based method for optimizing the parameters of the blanching process for wild rapeseed, comprising the following steps: S1: Construct a multi-objective evaluation function for blanched wild rapeseed; during the iterative optimization process, identify the non-dominated solution set from the current process parameter sample set to form the Pareto front of the current generation; S2: For the solutions in the Pareto front, calculate the normalized information entropy of each parameter dimension, such as blanching temperature, blanching time, and material-to-water ratio, and set an independent boundary contraction degree to update the search boundary. The more concentrated the distribution of the solutions in a certain parameter dimension, the smaller the normalized information entropy and the greater the corresponding boundary contraction degree. Within the updated search boundary, perform kernel density estimation on the numerical distribution of each parameter dimension of the Pareto front to construct a probability density function. Through inverse transformation sampling, map the uniformly distributed sample points generated by the quasi-Monte Carlo sequence to the first new sample set biased towards the dense region of the current optimal solution. S3: Calculate the hypervolume increment within a preset number of consecutive iterations of the Pareto front. If it is lower than a preset threshold, activate the diversity enhancement mechanism and generate a second new sample set based on the heavy-tailed t-distribution model aligned with the covariance matrix and the principal component analysis results of the Pareto front. Combine this second new sample set with the first new sample set to form the next generation sample population. If this mechanism is not activated, the first new sample set supplemented with random perturbation samples will form the next generation sample population. Repeat the iteration process until the search boundary range of each parameter dimension is less than the preset termination threshold, and output the Pareto front set composed of the optimal combination of process parameters.
[0007] This invention constructs a multi-objective evaluation function and introduces Pareto front identification to simultaneously optimize three mutually constraining indicators: peroxidase residue rate, chlorophyll retention rate, and vitamin C retention rate, achieving a synergistic balance between enzyme inactivation and nutrient retention in the blanching process. Based on this, the search boundary is independently adjusted using the normalized information entropy of each parameter dimension, matching the degree of boundary contraction to the distribution concentration of the current optimal solution across dimensions. This avoids premature exclusion of potential preferred regions by traditional synchronous contraction methods, improving search efficiency. Combining kernel density estimation with inverse transformation sampling of quasi-Monte Carlo sequences makes the generated new samples more likely to be distributed in the dense region of the current optimal solution, enhancing local fine-grained search capabilities. Furthermore, by monitoring the hypervolume increment to activate the diversity enhancement mechanism, a second new sample set with strong exploratory capabilities is generated using a heavy-tailed t-distribution model, effectively alleviating the problem of population diversity decline and getting trapped in local optima in the later stages of iteration. Finally, the Pareto front set is output through adaptive boundary termination conditions, providing an optimal parameter combination for the wild rapeseed blanching process that balances enzyme activity inhibition and nutrient retention.
[0008] Preferably, the multi-objective evaluation function optimizes the peroxidase residue rate, chlorophyll retention rate, and vitamin C retention rate of wild rapeseed after blanching.
[0009] Preferably, a multi-objective evaluation function for blanching wild rapeseed is constructed, including: setting blanching temperature, blanching time, and material-to-water ratio as independent variables, conducting a three-factor, three-level response surface experiment, and measuring the actual values of peroxidase residue rate, chlorophyll retention rate, and vitamin C retention rate under the corresponding set experimental parameters; using the least squares method to fit the measured actual data of each index, and constructing independent multivariate quadratic polynomial response surface regression models for peroxidase residue rate, chlorophyll retention rate, and vitamin C retention rate, respectively. The regression models include the first-order term, quadratic term, and interaction term of the independent variables; setting the unified optimization direction to find the minimum value, keeping the response surface regression model of peroxidase residue rate unchanged, and taking negative values for the response surface regression models of chlorophyll retention rate and vitamin C retention rate, thus constructing a multi-objective evaluation function transformed into a single minimization direction.
[0010] By using response surface methodology and multivariate quadratic polynomial regression fitting, a continuously differentiable mathematical model for each index was constructed, providing an accurate basis for calculating the objective function for subsequent iterative optimization. At the same time, by taking negative values for chlorophyll retention rate and vitamin C retention rate, the optimization direction was unified, which facilitates the use of standard multi-objective evolutionary algorithms for determining dominance relationships and searching for Pareto fronts.
[0011] Preferably, during the iterative optimization process, identifying the non-dominated solution set constituting the Pareto front of the current generation from the current process parameter sample set includes: for any two solutions A and B in the current process parameter sample set, determining the dominance relationship under a multi-objective evaluation function transformed into a single minimization direction; if the calculated values of peroxidase residual rate, negative chlorophyll retention rate, and negative vitamin C retention rate of solution A are all less than or equal to the calculated values of solution B, and at least one of the calculated values is strictly less than that of solution B, then solution A is determined to dominate solution B; the number of times each solution in the sample set is dominated by other solutions is counted, and solutions with a dominance count of 0 are extracted to form a non-dominated solution set, and the process parameter combination of all solutions in this non-dominated solution set is recorded as the Pareto front of the current generation.
[0012] The non-dominated sorting algorithm is used to extract the Pareto front from the current sample set. It can quickly and accurately identify the high-quality solution set in the current generation that is not dominated by any other solution. This provides high-quality basic data for subsequent search boundary updates and new sample generation, and improves the convergence direction of the optimization process.
[0013] Preferably, the normalized information entropy of each parameter dimension—blanching temperature, blanching time, and material-to-water ratio—is calculated, and an independent boundary shrinkage degree is set to update the search boundary, including: Divide the current search boundary of each parameter dimension into a preset number of intervals, count the frequency of solutions in the current Pareto front falling within each interval, and calculate the probability distribution corresponding to each interval; calculate the discrete information entropy of each parameter dimension based on the probability distribution, and divide it by the maximum possible information entropy to obtain the normalized information entropy; multiply the current search boundary width of each parameter dimension by the product of the corresponding normalized information entropy and a fixed shrinkage coefficient to obtain the updated search boundary width of that parameter dimension; centering on the median or mean of the current Pareto front solutions in that parameter dimension, redefine the upper and lower limits according to the updated search boundary width to generate the updated search boundary.
[0014] By calculating the normalized information entropy of each parameter dimension, the distribution concentration of Pareto front solutions across different process parameter dimensions is quantified, allowing for independent adjustment of the search boundary contraction magnitude for each dimension. This mechanism enables dimensions with more concentrated distributions to experience greater boundary contraction, thereby focusing search resources on regions with dense optimal solutions while avoiding the loss of potential edge solutions due to uniform contraction.
[0015] Preferably, a probability density function is constructed by kernel density estimation of the numerical distribution of each parameter dimension of the Pareto front. The uniformly distributed sample points generated by the quasi-Monte Carlo sequence are then mapped to a first new sample set biased towards the dense region of the current optimal solution through inverse transformation sampling. This includes: using a Gaussian kernel function to estimate the kernel density of the Pareto front solution distribution in each parameter dimension, and truncating it within the updated search boundary to construct a smooth, non-parametric truncated probability density function model; numerically integrating the truncated probability density function model to calculate a monotonically increasing truncated cumulative distribution function; generating a uniformly distributed quasi-Monte Carlo random sample point set within the interval 0 to 1 using the Sobol sequence; and performing an inverse transformation sampling mapping on the uniformly distributed quasi-Monte Carlo random sample points by calculating the inverse function of the truncated cumulative distribution function to generate a first new sample set falling within the updated search boundary and conforming to the truncated probability density function distribution.
[0016] By using kernel density estimation to perform nonparametric modeling of the Pareto front distribution across various parameter dimensions, the probability density characteristics of the current optimal solution set can be accurately reflected without pre-setting a specific distribution form. By combining uniformly distributed samples generated from quasi-Monte Carlo sequences with inverse transformation sampling, the sampling points are biased towards high-density regions, thereby achieving efficient and low-redundancy local fine sampling within the updated search boundary, improving search efficiency and convergence accuracy.
[0017] Preferably, the generation of a second new sample set based on a heavy-tailed t-distribution model aligned with the covariance matrix and the principal component analysis results of the Pareto front includes: standardizing the values of the current Pareto front solution across each process parameter dimension; calculating the covariance matrix of the standardized data; performing principal component analysis on the covariance matrix to extract principal eigenvectors and eigenvalues; constructing a multivariate heavy-tailed t-distribution model, aligning the scale matrix of the heavy-tailed t-distribution model with the matrix reconstructed from the principal eigenvectors and eigenvalues, and setting its degree of freedom parameters to preset constants; generating a random offset vector centered on the random solution in the current Pareto front using the multivariate heavy-tailed t-distribution model; adding the random solution to the offset vector scaled according to the standard deviation of each parameter dimension to obtain candidate sample points; and selecting candidate sample points located within the boundaries of the global initial process parameters to generate a second new sample set.
[0018] Principal component analysis was used to extract the main distribution direction and scale information of the Pareto front in the process parameter space, and a covariance matrix aligned with it was constructed. This covariance matrix was then used as the scale matrix for the heavy-tailed t-distribution model. The heavy-tailed t-distribution has a thicker tail probability, which can generate a bias vector with a larger span than the Gaussian distribution. This provides an effective global perturbation when population diversity declines in the later stages of iteration, helping the algorithm escape local optima and expand sampling along the main extension direction of the Pareto front, thus enhancing the global optimization capability of multi-objective optimization.
[0019] Preferably, the method for calculating the hypervolume increment within a predetermined number of consecutive iterations of the Pareto front is as follows: The peroxidase residue rate, chlorophyll retention rate, and vitamin C retention rate are uniformly converted into normalized minimization targets. Specifically, the peroxidase residue rate, after normalization according to its value range, is used as the first target. The chlorophyll retention rate and vitamin C retention rate are converted into the second and third targets, respectively, by subtracting the normalized retention rate from the numerical value of 1. Using the theoretical worst point in the transformed target space as a reference point, the current Pareto front hypervolume is calculated and recorded. The difference in hypervolume within the last five iterations is calculated as the hypervolume increment.
[0020] After unifying the three objectives into a normalized minimization objective, calculating the hypervolume of the Pareto front and its increment within consecutive iterations allows for quantitative monitoring of the convergence status of the optimization process. The hypervolume is a comprehensive indicator of the quality of the Pareto front, and its increment reflects the degree of improvement of the front in recent iterations. By setting a hypervolume increment threshold to trigger a diversity enhancement mechanism, an adaptive response is achieved when the algorithm stalls during convergence, avoiding the problem of being trapped in local optima and unable to recover.
[0021] Preferably, the step of constructing the next generation sample population from the first new sample set supplemented with random perturbation samples is as follows: generating a small random perturbation quantity that follows a uniform distribution, and adding the small random perturbation quantity to a portion of the samples in the first new sample set; combining the elite solution in the current Pareto front or the additionally generated uniform perturbation samples to supplement the sample number to a preset population size, which serves as the next generation sample population.
[0022] Without triggering the diversity enhancement mechanism, applying a uniform random perturbation to the first new sample set introduces appropriate randomness without disrupting the original sampling distribution, maintaining basic population diversity and preventing premature convergence. Simultaneously, combining elite solutions or additional perturbation samples to supplement the population size to a preset number ensures the stability of the sample size in each generation, facilitating stable algorithm iteration.
[0023] Preferably, the specific steps for repeatedly executing the iterative process until the search boundary range of each parameter dimension is less than a preset termination threshold, and outputting the Pareto front set composed of the optimal process parameter combination, are as follows: After each generation of the next generation sample population, calculate the difference between the upper and lower limits of the search boundary after the update of all current parameter dimensions; determine whether the difference of the blanching temperature is less than a preset temperature difference, whether the difference of the blanching time is less than a preset time difference, and whether the difference of the material-to-water ratio is less than a preset ratio difference; if all the above difference judgments are satisfied, then determine that the search boundary range of each parameter dimension is less than the preset termination threshold, and exit the iteration loop; extract the solutions with a dominance count of zero in the final generation of the next generation sample population, export the blanching temperature, blanching time, and material-to-water ratio values corresponding to the solutions, and output them as a formatted data file as the Pareto front set composed of the optimal process parameter combination.
[0024] Independent termination thresholds are set for each parameter dimension. Iteration stops only when the search boundary width of all dimensions is less than the corresponding threshold, ensuring that the optimization results achieve sufficient convergence accuracy across all process parameter dimensions. The final output is a formatted data file, facilitating subsequent production applications or process archiving, thus achieving a complete closed loop from multi-objective optimization to the derivation of actual process parameters.
[0025] The present invention has the following technical effects: This invention provides a method for optimizing the blanching process parameters of wild rapeseed to alleviate the difficulty of simultaneously achieving enzyme inactivation and nutrient retention during blanching. By constructing a multi-objective evaluation system covering peroxidase residue rate, chlorophyll retention rate, and vitamin C retention rate, the blanching process effect can be comprehensively and quantitatively evaluated. During the iterative optimization process, normalized information entropy is introduced to independently adjust the search boundaries of each process parameter dimension. Combined with kernel density estimation and quasi-Monte Carlo sequence mapping, the generated new samples are more likely to be distributed in the high-quality solution-dense region, thereby improving the parameter search accuracy and overall convergence efficiency. In addition, this invention introduces a diversity enhancement mechanism by monitoring the hypervolume increment and expands the sampling using a heavy-tailed t-distribution model, alleviating the problem of conventional optimization processes easily getting trapped in local optima. This helps maintain sample diversity in the later stages of the search and improves the stability of the blanching process parameter optimization results. Attached Figure Description
[0026] Figure 1 A flowchart of a machine learning-based method for optimizing the parameters of wild rapeseed blanching process; Figure 2 The response surface plot is for the peroxidase residual rate; Figure 3 This is a schematic diagram showing the distribution relationship between the initial Pareto solution set, the first new sample set, and the second new sample set in the three-dimensional target space. Figure 4This is a comparison chart of the experimental results. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] This invention discloses a method for optimizing the blanching process parameters of wild rapeseed based on machine learning, referring to... Figure 1 This includes the following steps: S1, construct a multi-objective evaluation function, and identify non-dominated solutions from the sample set to form the Pareto front.
[0029] A multi-objective evaluation function was constructed with the peroxidase residue rate, chlorophyll retention rate, and vitamin C retention rate of wild rapeseed after blanching as the optimization objectives. During the iterative optimization process, the non-dominated solution set was identified from the current process parameter sample set to form the Pareto front of the current generation.
[0030] In one embodiment, a mathematical model for the evaluation function is established using response surface methodology combined with the Box-Behnken three-factor, three-level response surface design principle. First, the blanching temperature, blanching time, and material-to-water ratio of wild rape are determined as independent variables through single-factor experiments. The Box-Behnken design module of DesignExpert software is then used to generate a response surface experimental table containing center point repetitions. Based on the parameters set in the experimental table, a wild rape blanching experiment is conducted. The peroxidase residue rate is determined using the guaiacol method, the chlorophyll retention rate is determined using spectrophotometry, and the vitamin C retention rate is determined using the dichlorophenolindophenol titration method. Multiple polynomial regression fitting is performed on the linear, quadratic, and interaction terms between each optimization objective and the independent variables using OLS functions, constructing three continuously differentiable multi-objective evaluation functions as the basis for calculating the fitness of the objectives in iterative optimization.
[0031] After initializing the process parameter sample set, the samples are substituted into the multi-objective evaluation function to calculate the fitness of each sample across the three objectives. Subsequently, a fast non-dominated sorting algorithm is used to identify the Pareto front. Specifically, first, each solution in the sample set is traversed, and the number of solutions dominated by that solution (i.e., the dominance count) and the set of other solutions dominated by that solution are calculated. Then, solutions with a dominance count of zero are selected to form the first layer of non-dominated solutions. Next, the solutions in the first layer are removed from the sample set, and the dominance counts of the remaining solutions are updated. This process is repeated until all solutions are stratified, and finally, the first layer of non-dominated solutions is extracted as the Pareto front for the current iteration.
[0032] In another optional embodiment, the construction of a multi-objective evaluation function, with the optimization objectives of peroxidase residue rate, chlorophyll retention rate, and vitamin C retention rate after blanching of wild rapeseed, includes: The blanching temperature, blanching time and material-to-water ratio of wild rape were set as independent variables. A three-factor, three-level response surface experiment was conducted to measure the actual values of peroxidase residue rate, chlorophyll retention rate and vitamin C retention rate under the corresponding experimental parameters. The least squares method was used to fit the real data of each measured index, and independent multivariate quadratic polynomial response surface regression models were constructed for the peroxidase retention rate, chlorophyll retention rate and vitamin C retention rate, respectively. The regression models include the first term, the quadratic term and their interaction term of the independent variables. By setting the unified optimization direction to find the minimum value, keeping the response surface regression model of peroxidase residue rate unchanged, and taking negative values for the response surface regression models of chlorophyll retention rate and vitamin C retention rate, a multi-objective evaluation function transformed into a single minimization direction is constructed.
[0033] Define the core process variables for blanching wild rapeseed and their global search boundary range: blanching temperature. Blanching temperature: 85℃ to 95℃, rinsing time The time interval was 60s to 120s, and the material-to-water ratio was denoted by R and used as the third independent variable X3. R was the ratio of water volume to rapeseed mass, in mL / g, with a value ranging from 5 to 15, corresponding to a process expression of 1:5 to 1:15. Based on Box-Behnken response surface design, a three-factor, three-level response surface experiment was conducted. The experiment included 12 sets of factor level combination experiments and 5 sets of center point replication experiments for estimating pure error, for a total of 17 experiments.
[0034] After the experiment, the peroxidase residual rate of each group was obtained. chlorophyll retention rate and vitamin C retention rate The actual measured values, including those under typical experimental conditions. The value range is generally between 2.5% and 15.0%. and The distribution is generally between 60.0% and 95.0%. Ordinary least squares method was used to perform multivariate quadratic polynomial fitting at the matrix algebra level on the three sets of multidimensional response data. The preferred approach was to first convert the blanching temperature, blanching time, and water-to-material ratio into dimensionless coded variables. , , To reduce the influence of different physical dimensions on the scale of regression coefficients, the response surface regression equations are solved separately: The coefficients for each component are: one intercept term, three linear term coefficients, three quadratic term coefficients, and three interaction term coefficients. For the intercept term, , , The coefficient of the linear term, , , The coefficient of the quadratic term, , , The coefficients represent the three interaction terms. Under conditions where the material-to-water ratio is at a central or fixed level, the response surface of the peroxidase residue rate as a function of blanching temperature and blanching time is shown below. Figure 2 As shown.
[0035] To satisfy the unified minimization dominance search in modern multi-objective evolutionary optimization algorithms, a vector-form multi-objective evaluation function is defined. Among these, for the peroxidase residual rate that needs to be inhibited and minimized, its expression is preserved. For the chlorophyll retention rate and vitamin C retention rate, which need to be maximized, the optimization direction is reversed by taking the overall negative value. and Through the aforementioned space reversal transformation, the original food engineering problem, which mixed objectives of minimization and maximization in different directions, is equivalently transformed, under the sense of dominance, into a standard multi-objective optimization problem that seeks to achieve a multidimensional Pareto minimum for the vector function F(X).
[0036] In an optional embodiment, the step of identifying the non-dominated solution set constituting the Pareto front of the current generation from the current process parameter sample set during the iterative optimization process includes: For any two solutions A and B in the current process parameter sample set, the dominance relationship is determined under the multi-objective evaluation function with a single minimization direction. If the calculated values of peroxidase residual rate, negative chlorophyll retention rate and negative vitamin C retention rate of solution A are all less than or equal to the calculated values of solution B, and at least one of the calculated values of these indicators is strictly less than that of solution B, then solution A is determined to dominate solution B. The number of times each solution in the statistical sample set is dominated by other solutions is counted. Solutions with a dominance count of 0 are extracted to form a non-dominated solution set. The process parameter combinations of all solutions in this non-dominated solution set are recorded as the Pareto front of the current generation.
[0037] Set the total population size of *Raphanus sativus* in the current generation to N=100. For any two candidate solutions A and B within the population, extract the fitness target value vectors obtained by substituting them into the response surface model. and Only when three parallel conditions are met. and and Meanwhile, Boolean logic expressions The solution is formally established and recorded in the topological relation only when the operation result is True. Successful Domination Solution Boolean state.
[0038] During the dominance census of a sample set of 100, an integer dominance counter array of length 100 is initialized, with each element... The initial number of individuals is set to 0. Two nested loops are started to perform pairwise cross-adversarial comparisons on individuals in the solution space, executing a total of 100×99 judgment logics.
[0039] In a comparison, if the solution at index i is determined to be dominant by any other valid solution, then the counter for that solution is incremented. After all the loop comparison instructions have been executed, the counter array is traversed, and condition masks are used. The process parameter sequences of all highly optimal individuals that still maintain a count value of 0 (i.e., are not dominated by any contemporaneous competitors) are isolated and extracted. Simultaneously, an adaptive size matrix is recorded, representing the specific blanching temperature, blanching time, and material-to-water ratio combination mapped onto the real-world dimension of the non-dominated solution set. For example, if a generation evolves into 24 non-repeating optimal frontier solutions, the set of these 24 solutions is taken as the effective Pareto front of the current generation.
[0040] S2 calculates the normalized information entropy to independently shrink the boundary, and combines kernel density estimation with quasi-Monte Carlo sampling to generate a first new sample set biased towards the optimal region.
[0041] For the solutions in the Pareto front, the normalized information entropy of each parameter dimension, such as blanching temperature, blanching time, and material-to-water ratio, is calculated, and an independent boundary contraction degree is set to update the search boundary. The more concentrated the distribution of solutions in a certain parameter dimension, the smaller the normalized information entropy and the greater the corresponding boundary contraction degree. Within the updated search boundary, the kernel density estimation is performed on the numerical distribution of each parameter dimension of the Pareto front to construct a probability density function. Through inverse transformation sampling, the uniformly distributed sample points generated by the quasi-Monte Carlo sequence are mapped to a first new sample set biased towards the dense region of the current optimal solution.
[0042] First, extract the process parameters for all solutions in the current Pareto front. For the three dimensions of blanching temperature, blanching time, and material-to-water ratio, divide the parameter value range of the current dimension into ten equidistant intervals. Count the frequency of the solution in each interval and divide it by the total number of solutions to obtain the probability of each interval. Next, calculate the Shannon information entropy of this dimension and divide it by the natural logarithm of the number of intervals to obtain the normalized information entropy with a value between 0 and 1. Then, multiply the current search boundary width by the normalized information entropy and the preset boundary contraction coefficient to obtain the updated search boundary width. Finally, using the median or mean of the current Pareto front in each dimension as the boundary relocation center, redetermine the upper and lower limits on both sides according to half of the updated search boundary width to obtain the updated search boundary.
[0043] For the Pareto front parameter solutions within the updated search boundary, a Gaussian kernel function is used to estimate the kernel density of the numerical distributions in the three dimensions of blanching temperature, blanching time, and material-to-water ratio. Grid search combined with cross-validation is then used to determine the optimal smoothing bandwidth, thereby constructing continuous probability density functions for each dimension. Next, the probability density functions are numerically integrated to obtain the cumulative distribution function. Subsequently, a uniformly distributed quasi-Monte Carlo sequence in the multidimensional space is generated as initial probability points. Finally, an inverse mapping relationship is constructed based on the cumulative distribution function, and these uniformly distributed probability points are mapped back to the process parameter space through one-dimensional interpolation, thus generating a new first sample set with a higher sampling density in the dense region of the current optimal solution.
[0044] In an optional embodiment, the step of calculating the normalized information entropy of each parameter dimension—blanching temperature, blanching time, and material-to-water ratio—and setting an independent boundary shrinkage degree to update the search boundary includes: Divide the current search boundary of each parameter dimension into a preset number of intervals, count the frequency of the solution in the current Pareto front falling in each interval, and calculate the probability distribution corresponding to each interval. The discrete information entropy of each parameter dimension is calculated based on the probability distribution, and then divided by the maximum possible information entropy to obtain the normalized information entropy. Multiply the current search boundary width of each parameter dimension by the product of the corresponding normalized information entropy and a fixed shrinkage coefficient to obtain the updated search boundary width of that parameter dimension. Centered on the median or mean of the current Pareto front solution in this parameter dimension, the upper and lower limits are redefined according to the updated search boundary width to generate the updated search boundary.
[0045] For the three independent one-dimensional continuous process parameters of blanching temperature, blanching time, and material-to-water ratio in multi-objective optimization, a preset partitioning parameter K=10 is used. Taking the material-to-water ratio parameter as an example, this dimension is actually calculated using the material-to-water ratio R. Assuming that after reading the current Pareto front set data, the current local boundary of the material-to-water ratio R is found to be a lower limit of 8 and an upper limit of 12, corresponding to a process expression of 1:8 to 1:12, then the total width of the current local search domain interval is 4. This is uniformly divided into 10 continuous and non-overlapping sub-intervals, each with a width of 0.4. By iterating through the material-to-water ratio attribute values of all solutions within this Pareto front, the frequency of solutions falling into the aforementioned 10 corresponding sub-intervals is counted. Then, by dividing each frequency by the total number M of effective frontier solutions in the current iteration state, the discrete probability mass distribution of the material-to-water ratio is obtained. ( The one-dimensional discrete information entropy is calculated on the extracted distribution state, and the result is divided by the theoretical maximum possible information entropy of the ten interval states. This allows the information entropy index to be mapped and calculated as normalized information entropy. When the Pareto solution set clusters around a specific material-to-water ratio, for example, when most frontier solutions cluster around 1:10, the calculated probability quality distribution becomes extremely sharp, leading to a drop in information entropy. It will move closer to 0.
[0046] Read the preset fixed boundary shrinkage coefficient hyperparameter, for example, set... Extract the original search boundary span width value of this parameter dimension before entering this iteration. The material-to-water ratio dimension mentioned earlier is 4. Perform iterative compression and update to obtain the latest boundary width range value. If a certain evolutionary algebra calculates the normalized information entropy in the feed-to-water ratio dimension... The new width obtained by the corresponding contraction in this generation is 2.04. After completing the span reduction, the median point of the current material-to-water ratio data in the frontier is obtained. Assuming that the median value of the material-to-water ratio is R=10.5, corresponding to a process expression of 1:10.5, the data is radiated outwards from this median point as the spatial anchoring center. The latest boundary width of 2.04 obtained in the previous step is divided into two equal parts. Thus, a new search range with a new positioning is applied to the space through linear translation. The lower limit of the iterative dynamic local search for the material-to-water ratio dimension is updated to R=9.48, corresponding to 1:9.48, and the upper limit is updated to R=11.52, corresponding to 1:11.52. This adaptive control process is also applicable to the blanching temperature and blanching time dimensions of wild rapeseed to realize automatic volume pruning of the multi-dimensional search superrectangular space based on the monitoring of the discrete state of the frontier space of the evolutionary population.
[0047] In an optional embodiment, the step of constructing a probability density function by kernel density estimation of the numerical distribution of each parameter dimension of the Pareto front within the updated search boundary, and mapping the uniformly distributed sample points generated by the quasi-Monte Carlo sequence to a first new sample set biased towards the dense region of the current optimal solution through inverse transformation sampling, includes: The distribution of Pareto front solutions across various parameter dimensions is estimated using a Gaussian kernel function, and then truncated within the updated search boundary to construct a smooth, non-parametric truncated probability density function model. Numerical integration is performed on the truncated probability density function model to calculate the monotonically increasing truncated cumulative distribution function; A uniformly distributed set of quasi-Monte Carlo random sample points is generated in the interval from 0 to 1 using the Sobol sequence; By calculating the inverse function of the truncated cumulative distribution function, an inverse transformation sampling mapping is performed on the uniformly distributed quasi-Monte Carlo random sample points to generate a first new sample set that falls within the updated search boundary and follows the distribution of the truncated probability density function.
[0048] Within the updated three-dimensional local boundaries of blanching temperature, blanching time, and material-to-water ratio, kernel density estimation is performed by independently introducing a one-dimensional Gaussian kernel function with automatic bandwidth adjustment for each orthogonal dimension of the random variable. For any process parameter dimension X, the optimal smoothing bandwidth is calculated according to the Silverman empirical criterion. ,in Let M be the unbiased sample standard deviation of this parameter in the current solution set, and M be the individual size of the frontier solution set. A Gaussian kernel probability density function is constructed based on this. To avoid generating invalid parameters that exceed the boundary, the continuous probability density function is truncated within the updated boundary [L, U], where L and U represent the lower and upper limits of the updated search boundary, respectively. The total probability area of this truncated interval is calculated through numerical integration, and its reciprocal is used as a normalization weight to scale the probability values within the interval to satisfy the statistical principle that the total probability is 1. For the truncated probability density function PDF, it is discretized into 1000 small grids, and the trapezoidal integral method is used to approximate the numerical cumulative integral, obtaining the mapping coordinate matrix of the monotonically increasing discrete truncated cumulative distribution function CDF.
[0049] During the random mapping phase, a low-dissimilarity sequence generator with dimension D=3 is initialized. The Sobol low-dissimilarity quasi-Monte Carlo sequence generator is invoked to generate sequences of size D=3 within the interval [0,1]. The uniformly distributed sample matrix. For any random scalar in the sequence distributed in the interval [0,1] The CDF inverse lookup table, consisting of 1000 grid points, is invoked to calculate the inverse-transformed process parameter values through linear interpolation of adjacent discrete points. Through the aforementioned inverse transformation sampling mapping, the originally uniformly distributed Monte Carlo samples are reconstructed into parameter clusters that conform to the actual boundary and statistical distribution characteristics. The resulting 80 candidate solutions containing rapeseed processing parameters are distributed within the neighborhood of the optimal feasible solution of the current Pareto front, thereby improving the efficiency of local spatial optimization and reducing the probability of repeated sampling in low-value regions.
[0050] S3 monitors the hypervolume increment to trigger diversity enhancement, generates new samples using a heavy-tailed t-distribution, and iterates until the boundary converges to output the optimal solution set.
[0051] Calculate the hypervolume increment within a preset number of consecutive iterations of the Pareto front. If it is lower than a preset threshold, activate the diversity enhancement mechanism. Generate a second new sample set based on a heavy-tailed t-distribution model aligned with the covariance matrix and the principal component analysis results of the Pareto front. Combine this second new sample set with the first new sample set to form the next generation sample population. If the mechanism is not activated, supplement the first new sample set with randomly perturbed samples to form the next generation sample population. Repeat the iteration process until the search boundary range of each parameter dimension is less than the preset termination threshold, and output the Pareto front set composed of the optimal combination of process parameters.
[0052] At the end of each iteration, the peroxidase residue rate, chlorophyll retention rate, and vitamin C retention rate are first uniformly converted into normalized minimization objectives. Specifically, the peroxidase residue rate, normalized according to its value range, is used as the first objective. The chlorophyll retention rate and vitamin C retention rate are converted into the second and third objectives, respectively, by subtracting the normalized retention rate from 1. Then, using the theoretical worst point in the transformed objective space as a reference point, the hypervolume of the current Pareto front is calculated and recorded. The difference in hypervolume over the last five iterations is calculated as the hypervolume increment. If the average of this increment is lower than a preset minimum threshold, the algorithm is considered to be trapped in a local optimum, thus activating the diversity enhancement mechanism.
[0053] Principal component analysis is performed on the process parameter matrix of the current Pareto front solution to extract eigenvalues and eigenvectors. A rotation matrix is constructed using the eigenvectors, and combined with the amplified eigenvalues to construct a covariance matrix aligned with the principal component directions. A heavy-tailed t-distribution model is constructed using this covariance matrix and the centroid of the current Pareto front as parameters. A second new sample set with strong global exploration capabilities is generated using the random variable generation function of this model. This second new sample set is merged with the first new sample set to form the next generation sample population. Candidate samples are truncated or supplemented according to a preset population size N. If the hypervolume increment is not lower than a preset threshold, this mechanism is not activated, and a small, uniformly distributed random perturbation is generated. For example, this perturbation is added to a portion of the samples in the first new sample set and combined with elite solutions or additional uniformly perturbation samples from the current Pareto front to supplement the population to the preset population size N, forming the next generation sample population.
[0054] After each generation of the next generation sample population, the process returns to the steps of fitness calculation and non-dominated solution set identification. Then, the difference between the upper and lower bounds of the search boundary after updating all parameter dimensions is calculated. If the difference in blanching temperature is less than a preset temperature difference, the difference in blanching time is less than a preset time difference, and the difference in material-to-water ratio is less than a preset ratio difference, i.e., the preset termination threshold condition is met, the iteration loop is exited. Finally, solutions with a dominance count of zero are extracted from the final generation of the next generation sample population, and the corresponding blanching temperature, blanching time, and material-to-water ratio values R are derived. Finally, formatted data processing is performed and the data is output as a CSV file, serving as the Pareto front set composed of the optimal combination of blanching process parameters for wild rapeseed.
[0055] In an optional embodiment, generating a second new sample set based on a heavy-tailed t-distribution model aligned with the Pareto front principal component analysis results using the covariance matrix includes: standardizing the values of the current Pareto front solution across each process parameter dimension; calculating the covariance matrix of the standardized data; performing principal component analysis on the covariance matrix to extract principal eigenvectors and eigenvalues; constructing a multivariate heavy-tailed t-distribution model, setting the scaling matrix of the heavy-tailed t-distribution model to be aligned with the matrix reconstructed from the principal eigenvectors and eigenvalues, and setting its degree-of-freedom parameter to a preset constant; generating a random offset vector centered on a random solution in the current Pareto front using the multivariate heavy-tailed t-distribution model; adding the random solution to the offset vector scaled according to the standard deviation of each parameter dimension to obtain candidate sample points; and selecting candidate sample points located within the boundaries of the global initial process parameters to generate the second new sample set.
[0056] Set a reference point based on the engineering extreme value, for example, in the unnormalized target space that has been unified to the minimization direction. In this context, the reference point can be set to 100% peroxidase retention, 0% negative chlorophyll retention, and 0% negative vitamin C retention. If a normalized target space is used, the normalized peroxidase retention is taken as the first target, and the chlorophyll retention and vitamin C retention are converted into the second and third targets respectively by subtracting the normalized retention rate from 1. This ensures that all three targets satisfy the condition that smaller values indicate better performance and larger values indicate worse performance. In this case, the worst-case reference point can be set to (1,1,1). The increment of the Pareto front hypervolume is continuously statistically analyzed over five consecutive generations using the Monte Carlo point estimation method. If this increment is continuously lower than the set stagnation threshold... The system is identified as potentially trapped in local extrema, triggering a diversity enhancement mechanism. Upon triggering, the non-dominated solution of the current generation, i.e., the three-dimensional process parameter set, is Z-score standardized to eliminate dimensional influences. The standardized data is then used for calculations. covariance matrix Principal component analysis or singular value decomposition (SVD) is used to extract three orthogonal eigenvectors and their corresponding eigenvalues to characterize the main distribution direction and variance distribution of the Pareto front in the process parameter space. The scale matrix is then reconstructed using the extracted eigenvectors and eigenvalues. This ensures that the spatial distribution direction of the matrix matches the distribution trend of the Pareto front in the process parameter space.
[0057] To generate perturbation variables with large spans, a multivariate heavy-tailed t-distribution model is constructed. The degrees of freedom parameter of this distribution is set to a constant. This is done to retain a thicker tail probability, increasing the probability of generating far-end mutation solutions; at the same time, the above-aligned scaling matrix is... Substitute these values into the model to control the directional weights of multidimensional mutations. Based on this, randomly select an elite solution from the current Pareto front population as the baseline center. A set of standardized offset vectors in multiple directions is generated using the aforementioned multivariate heavy-tailed t-distribution model. .because This represents the perturbation relative to the reference solution; therefore, scaling is only required according to the standard deviation of each parameter dimension to obtain the offset vector with actual physical dimensions. , where σ is the standard deviation vector of the current Pareto front across all process parameter dimensions. Calculation This process yields new candidate sample points. This avoids repeatedly stacking the sample mean as an additional offset.
[0058] Perform global boundary validity checks on all generated candidate sample points, such as checking if the blanching time is within the preset valid range of [60s, 120s]. If the parameter exceeds the limit, use a mirror bounce strategy or resampling until sample points that meet the boundary constraints are generated. Repeat this process until a sample of size [size missing] is obtained. The valid samples, as the second new sample set, are combined with a size of [missing information]. The first new sample set with a size of 80 is merged to form the next generation sample population with a size of N=100. If the number of samples after merging exceeds N, samples with higher non-dominated levels are preferentially retained and truncated to N, and then used in the next generation of iteration and evaluation processes. The distribution relationship of the initial Pareto front solution set, the first new sample set, and the second new sample set in the three-dimensional target space is as follows: Figure 3 As shown.
[0059] A baseline control group and two ablation experimental groups were set up to verify the performance of multi-objective optimization. The baseline control group adopted a conventional non-dominated sorting genetic optimization mechanism with random mutations and no adaptive search boundary. The locally enhanced ablation group, in addition to the control group, separately introduced information entropy boundary contraction and an inverse transformation sampling module based on kernel density estimation combined with quasi-Monte Carlo sequences, but removed the heavy-tailed t-distribution mechanism. The complete scheme group simultaneously introduced a local sampling module based on kernel density estimation and a global perturbation sampling module based on principal component analysis aligned with the heavy-tailed t-distribution. The unified initial conditions for the three experiments were: population size 100, maximum number of evolutionary iterations 200 generations, and hypervolume increment monitoring trigger threshold of 1×10⁻⁶. -4 .
[0060] Comparison of the three quality indicators in the Pareto front solution sets of each group. Figure 4 As shown, in the baseline control group, the mean residual peroxidase rate was 8.2%, the mean chlorophyll retention rate was 76.5%, and the mean vitamin C retention rate was 72.4%. In the locally enhanced ablation group, the residual peroxidase rate decreased to 6.1%, and the chlorophyll retention rate and vitamin C retention rate increased to 82.3% and 79.8%, respectively. In the complete protocol group, the mean residual peroxidase rate decreased to 4.5%, and the chlorophyll retention rate and vitamin C retention rate increased to 89.6% and 85.7%, respectively.
[0061] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for optimizing the parameters of wild rapeseed blanching process based on machine learning, characterized in that, Includes the following steps: S1: Construct a multi-objective evaluation function for blanched wild rapeseed; during the iterative optimization process, identify non-dominated solution sets from the current process parameter sample set to form the Pareto front of the current generation; S2: For the solutions in the Pareto front, calculate the normalized information entropy of each parameter dimension, such as blanching temperature, blanching time, and material-to-water ratio, and set an independent boundary contraction degree to update the search boundary. The more concentrated the distribution of the solutions in a certain parameter dimension, the smaller the normalized information entropy and the greater the corresponding boundary contraction degree. Within the updated search boundary, perform kernel density estimation on the numerical distribution of each parameter dimension of the Pareto front to construct a probability density function. Through inverse transformation sampling, map the uniformly distributed sample points generated by the quasi-Monte Carlo sequence to the first new sample set biased towards the dense region of the current optimal solution. S3: Calculate the hypervolume increment within a preset number of consecutive iterations of the Pareto front. If it is lower than a preset threshold, activate the diversity enhancement mechanism and generate a second new sample set based on the heavy-tailed t-distribution model aligned with the covariance matrix and the principal component analysis results of the Pareto front. Combine this second new sample set with the first new sample set to form the next generation sample population. If this mechanism is not activated, the first new sample set supplemented with random perturbation samples will form the next generation sample population. Repeat the iteration process until the search boundary range of each parameter dimension is less than the preset termination threshold, and output the Pareto front set composed of the optimal combination of process parameters.
2. The method for optimizing the blanching process parameters of wild rapeseed based on machine learning according to claim 1, characterized in that, The multi-objective evaluation function uses the residual peroxidase rate, chlorophyll retention rate, and vitamin C retention rate of wild rapeseed after blanching as optimization objectives.
3. The method for optimizing the blanching process parameters of wild rapeseed based on machine learning according to claim 2, characterized in that, A multi-objective evaluation function for blanching wild rapeseed was constructed, including: setting blanching temperature, blanching time, and material-to-water ratio as independent variables, conducting a three-factor, three-level response surface experiment, and measuring the actual values of peroxidase residue rate, chlorophyll retention rate, and vitamin C retention rate under the corresponding experimental parameters; using the least squares method to fit the measured actual data of each index, and constructing independent multivariate quadratic polynomial response surface regression models for peroxidase residue rate, chlorophyll retention rate, and vitamin C retention rate, respectively. The regression models include the first-order term, quadratic term, and interaction term of the independent variables; setting the unified optimization direction to find the minimum value, keeping the response surface regression model of peroxidase residue rate unchanged, and taking negative values for the response surface regression models of chlorophyll retention rate and vitamin C retention rate, thus constructing a multi-objective evaluation function transformed into a single minimization direction.
4. The method for optimizing the blanching process parameters of wild rapeseed based on machine learning according to claim 1, characterized in that, During the iterative optimization process, the non-dominated solution set is identified from the current process parameter sample set to form the Pareto front of the current generation. This includes: for any two solutions A and B in the current process parameter sample set, the dominance relationship is determined under a multi-objective evaluation function transformed into a single minimization direction; if the calculated values of peroxidase residual rate, negative chlorophyll retention rate, and negative vitamin C retention rate of solution A are all less than or equal to the calculated values of solution B, and at least one of the calculated values is strictly less than that of solution B, then solution A is determined to dominate solution B; the number of times each solution in the sample set is dominated by other solutions is counted, and solutions with a dominance count of 0 are extracted to form the non-dominated solution set, and the process parameter combination of all solutions in this non-dominated solution set is recorded as the Pareto front of the current generation.
5. The method for optimizing the blanching process parameters of wild rapeseed based on machine learning according to claim 1, characterized in that, Calculate the normalized information entropy of each parameter dimension, including blanching temperature, blanching time, and material-to-water ratio, and set independent boundary shrinkage degrees to update the search boundary, including: Divide the current search boundary of each parameter dimension into a preset number of intervals, count the frequency of solutions in the current Pareto front falling within each interval, and calculate the probability distribution corresponding to each interval; calculate the discrete information entropy of each parameter dimension based on the probability distribution, and divide it by the maximum possible information entropy to obtain the normalized information entropy; multiply the current search boundary width of each parameter dimension by the product of the corresponding normalized information entropy and a fixed shrinkage coefficient to obtain the updated search boundary width of that parameter dimension; centering on the median or mean of the current Pareto front solutions in that parameter dimension, redefine the upper and lower limits according to the updated search boundary width to generate the updated search boundary.
6. The method for optimizing the blanching process parameters of wild rapeseed based on machine learning according to claim 1, characterized in that, The process involves: estimating the kernel density of the numerical distribution of the Pareto front across its various parameter dimensions to construct a probability density function; and mapping uniformly distributed sample points generated by the quasi-Monte Carlo sequence to a first new sample set biased towards the dense region of the current optimal solution through inverse transformation sampling. This includes: estimating the kernel density of the Pareto front solution distribution across its various parameter dimensions using a Gaussian kernel function and truncating it within the updated search boundary to construct a smooth, non-parametric truncated probability density function model; numerically integrating the truncated probability density function model to calculate a monotonically increasing truncated cumulative distribution function; generating a uniformly distributed quasi-Monte Carlo random sample point set within the interval 0 to 1 using a Sobol sequence; and performing an inverse transformation sampling mapping on the uniformly distributed quasi-Monte Carlo random sample points by calculating the inverse function of the truncated cumulative distribution function to generate a first new sample set that falls within the updated search boundary and follows the truncated probability density function distribution.
7. The method for optimizing the blanching process parameters of wild rapeseed based on machine learning according to claim 1, characterized in that, A second new sample set is generated based on a heavy-tailed t-distribution model aligned with the covariance matrix and the principal component analysis results of the Pareto front. This process includes: standardizing the values of the current Pareto front solution across each process parameter dimension; calculating the covariance matrix of the standardized data; performing principal component analysis on the covariance matrix to extract principal eigenvectors and eigenvalues; constructing a multivariate heavy-tailed t-distribution model, aligning the scaling matrix of the heavy-tailed t-distribution model with the matrix reconstructed from the principal eigenvectors and eigenvalues, and setting its degrees of freedom parameters to preset constants; generating random offset vectors centered on random solutions in the current Pareto front using the multivariate heavy-tailed t-distribution model; adding the random solutions to the offset vectors scaled according to the standard deviation of each parameter dimension to obtain candidate sample points; and selecting candidate sample points located within the boundaries of the global initial process parameters to generate the second new sample set.
8. The method for optimizing the blanching process parameters of wild rapeseed based on machine learning according to claim 1, characterized in that, The method for calculating the hypervolume increment within a predetermined number of consecutive iterations of the Pareto front is as follows: The peroxidase residue rate, chlorophyll retention rate, and vitamin C retention rate are uniformly converted into normalized minimization targets. Specifically, the peroxidase residue rate, after normalization according to its value range, is used as the first target. The chlorophyll retention rate and vitamin C retention rate are converted into the second and third targets, respectively, by subtracting the normalized retention rate from the numerical value of 1. Using the theoretical worst point in the transformed target space as a reference point, the current Pareto front hypervolume is calculated and recorded. The difference in hypervolume within the last five iterations is calculated as the hypervolume increment.
9. The method for optimizing the blanching process parameters of wild rapeseed based on machine learning according to claim 1, characterized in that, The steps for constructing the next generation sample population from the first new sample set supplemented with random perturbation samples are as follows: generating a small random perturbation quantity that follows a uniform distribution, and adding the small random perturbation quantity to a portion of the samples in the first new sample set; combining the elite solution in the current Pareto front or the additional uniform perturbation samples, supplementing the sample number to a preset population size, which serves as the next generation sample population.
10. The method for optimizing the blanching process parameters of wild rapeseed based on machine learning according to claim 1, characterized in that, The specific steps for repeatedly executing the iterative process until the search boundary range of each parameter dimension is less than the preset termination threshold, and outputting the Pareto front set composed of the optimal combination of process parameters are as follows: After each generation of the next generation of sample population, calculate the difference between the upper and lower limits of the search boundary after the current update of all parameter dimensions; determine whether the difference of the blanching temperature is less than the preset temperature difference, whether the difference of the blanching time is less than the preset time difference, and whether the difference of the material-to-water ratio is less than the preset ratio difference. If all the above difference judgments are satisfied, the search boundary range of each parameter dimension is determined to be less than the preset termination threshold, and the iteration loop is exited; the solutions with a dominance count of zero in the next generation sample population of the final generation are extracted, and the blanching temperature, blanching time and material-to-water ratio corresponding to the solution are exported and output as a formatted data file, which serves as the Pareto front set composed of the optimal process parameter combination.
Citation Information
Patent Citations
Hydroelectric generating set multi-target parameter optimization method based on active learning and preference guidance
CN121835379A