Spider wasp optimizer algorithm- and machine learning algorithm-based method for predicting elastic modulus of glass

By optimizing the random forest model using the spider-bee optimization algorithm, the problem of inaccurate prediction of glass elastic modulus was solved, and fast and accurate prediction of glass performance was achieved.

WO2026016416A1PCT designated stage Publication Date: 2026-01-22CNBM RESEARCH INSTITUTE FOR ADVANCED GLASS MATERIALS GROUP CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/142199
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-19
Filing Date
2024-12-25
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing technologies cannot accurately and quickly predict the elastic modulus of glass, especially the influence of glass composition and manufacturing process parameters on the elastic modulus, which leads to errors in glass performance prediction.

Method used

A random forest model was constructed using the spider bee optimization algorithm and machine learning algorithm. By collecting data on glass composition and preparation process parameters, the parameters of the random forest model were optimized using the spider bee optimization algorithm to improve prediction accuracy.

Benefits of technology

It improves the accuracy and speed of glass elastic modulus prediction, reduces prediction errors, and enables rapid and accurate glass performance prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024142199_22012026_PF_FP_ABST
    Figure CN2024142199_22012026_PF_FP_ABST
Patent Text Reader

Abstract

A Spider wasp optimizer algorithm- and machine learning algorithm-based method for predicting the elastic modulus of glass. The method comprises: constructing an elastic modulus database of glass materials of different compositions; constructing feature descriptors; providing training set samples, and initializing parameters of a Spider wasp optimizer algorithm; setting optimization parameters of the Spider wasp optimizer algorithm, using the optimization parameters to optimize random forest algorithm parameters so as to obtain optimal random forest algorithm parameters, and establishing a glass elastic modulus prediction model; and inputting test sample data to predict the elastic modulus of glass. In the present application, the Spider wasp optimizer algorithm is used to optimize the random forest algorithm parameters for optimization, the structure is simple, the convergence speed and accuracy are improved, and the optimal random forest algorithm parameters obtained by optimization can significantly improve the performance of the random forest algorithm, which is of practical significance for improving the accuracy of predicting the elastic modulus of glass.
Need to check novelty before this filing date? Find Prior Art

Description

A Glass Elastic Modulus Prediction Method Based on Spider-Bee Optimization Algorithm and Machine Learning Algorithm

[0001] This application claims priority to Chinese Patent Application No. 202410981786.3, filed on July 19, 2024, entitled "Method for Predicting the Elastic Modulus of Glass Based on Spider-Bee Optimization Algorithm and Machine Learning Algorithm", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of glass elastic modulus prediction technology, and in particular to a glass elastic modulus prediction method based on spider bee optimization algorithm and machine learning algorithm. Background Technology

[0003] Glass is a non-equilibrium, amorphous material that can spontaneously relax to a supercooled liquid state. Unlike crystals, glass does not need to meet strict stoichiometric rules and can be considered a continuous solution of chemical elements. Therefore, a large number of elements can be components of glass; a change of 1 mol% in 80 chemical elements can produce 10... 52 There are several possible glass compositions. However, only 10 inorganic glasses have been reported. 6 There are approximately 100 types, which means there is still a huge space to explore new glass components with special properties.

[0004] As a fundamental mechanical property of solid materials, the accurate and rapid measurement of elastic modulus is crucial and essential. Methods for measuring the elastic modulus of solid materials fall into four categories: quasi-static methods, low-frequency methods, resonance methods, and wave propagation methods. Quasi-static methods determine the elastic modulus by measuring the relationship between force and deformation under quasi-static loading. Dynamic methods are categorized by measurement frequency from low to high: low-frequency methods, resonance methods, and wave propagation methods. Currently, the more mature methods for testing the elastic modulus of solid materials are the free beam resonance method and the pulse excitation method within the resonance method. Both methods essentially measure the elastic modulus by exciting the bending vibration of a long strip sample. The free beam resonance method typically uses two suspension wires to suspend the sample and simultaneously excite and receive vibration signals. By changing the suspension method, bending or torsional vibration can be selected for excitation. However, this method is prone to sample detachment and damage, and requires a relatively long experimental preparation time. Molecular dynamics simulations can quickly simulate the properties of glass, but they are subject to certain errors.

[0005] Currently, there is no prediction method that can effectively solve the problem of inaccurate prediction of the elastic modulus of glass. Summary of the Invention

[0006] To address the problem of inaccurate and rapid prediction of the elastic modulus of glass in existing technologies, this application proposes a method for predicting the elastic modulus of glass based on the spider-bee optimization algorithm and machine learning algorithms. The technical solution adopted in this application is as follows:

[0007] The first aspect of this application provides a method for predicting the elastic modulus of glass based on the spider-bee optimization algorithm and machine learning algorithm, comprising the following steps:

[0008] Step 1: Collect elastic modulus data of oxide glasses with different components and construct a glass elastic modulus database. The glass elastic modulus database includes a one-to-one mapping of glass components and their corresponding elastic modulus.

[0009] Step 2: Use the element molar content and preparation process parameters as input parameters for descriptors. That is, use the molar content of each component that makes up the glass as a set of descriptors, and use the glass preparation process parameters, including heating rate, melting temperature and holding time, to construct descriptors.

[0010] Step 3: Using the descriptors constructed in Step 2 as the input to the random forest model and the elastic modulus database constructed in Step 1 as the output of the random forest model, construct the training set and the test set to establish the random forest model.

[0011] Step 4: Introduce the spider-bee optimization algorithm to optimize the parameters of the selected random forest model;

[0012] Step 5: Build a top-performing random forest model based on the optimized parameters;

[0013] Step 6: For the glass composition to be predicted, use the optimal random forest model to predict the glass elastic modulus of the glass composition.

[0014] A second aspect of this application provides a method for training a random forest model for predicting the elastic modulus of glass, comprising:

[0015] Acquire characteristic data of various types of glass and the corresponding relationship between their elastic moduli; the characteristic data includes glass composition and glass preparation process parameters.

[0016] A training set is established based on the aforementioned correspondence;

[0017] Obtain the random forest model to be trained and determine the parameters to be optimized in the random forest model; the random forest model is used to output the predicted elastic modulus based on the input feature data;

[0018] Based on a preset fitness function, the parameters to be optimized in the random forest model are optimized using the spider-bee optimization algorithm to obtain a trained random forest model; the fitness function is constructed based on the accuracy of the random forest model in predicting the elastic modulus for the feature data in the training set.

[0019] A third aspect of this application provides a method for predicting the elastic modulus of glass, which obtains characteristic parameters of the glass to be predicted; the characteristic data includes glass composition and glass preparation process parameters.

[0020] The feature parameters are input into a pre-trained random forest model to obtain the elastic modulus predicted by the random forest model for the feature parameters.

[0021] The random forest model is trained based on the method described in the second aspect above.

[0022] The beneficial effects of this application are as follows:

[0023] This application uses the spider-bee optimization algorithm to optimize the parameters of the random forest algorithm. It has a simple structure, improves the convergence speed and accuracy, and the optimal random forest algorithm parameters obtained by optimization can significantly improve the performance of the random forest algorithm, which has practical significance for improving the accuracy of predicting the elastic modulus of glass. Attached Figure Description

[0024] The accompanying drawings, which are provided to further understand this application and constitute a part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application.

[0025] Figure 1 is a flowchart illustrating the glass elastic modulus prediction method based on the spider bee optimization algorithm and machine learning algorithm provided in the embodiments of this application. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this invention are within the scope of protection of this invention.

[0027] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to Figure 1. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0028] Example 1:

[0029] Step 1: Collect elastic modulus data of oxide glasses with different compositions and construct a glass elastic modulus database. Specifically, collect elastic modulus data of 800 sets of SiO2-Na2O-CaO-Al2O3 glasses with different contents and construct a glass elastic modulus database.

[0030] The glass elastic modulus database includes a one-to-one mapping between glass components and their corresponding elastic moduli. In this embodiment, it specifically includes 800 glass components and their corresponding glass elastic moduli.

[0031] Step 2: Use elemental molar content and preparation process parameters as descriptors for input parameters;

[0032] Step 2 includes the following steps 2-1 and 2-2:

[0033] Step 2-1: Use the elemental molar content of 800 SiO2-Na2O-CaO-Al2O3 glasses as a descriptor;

[0034] Step 2-2: Construct a descriptor using glass preparation process parameters (heating rate, melting temperature, holding time);

[0035] Specifically, for each of the 800 groups of SiO2-Na2O-CaO-Al2O3 glasses with different contents, a set of descriptors can be constructed based on the elemental molar content of the glass components of that glass, and another set of descriptors can be constructed based on the glass preparation process parameters. The properties of that glass can be characterized by these two sets of descriptors.

[0036] Step 3: Using the descriptors constructed in Step 2 as the input to the model and the elastic modulus database constructed in Step 1 as the output of the model, construct 640 sets of data to build a training set and 160 sets of data to build a test set, and establish a random forest model.

[0037] Specifically, for each type of glass, the two sets of descriptors constructed in steps 2-1 and 2-2 are used as input to the random forest model. The random forest model is then used to output the predicted elastic modulus for that type of glass based on these two sets of descriptors.

[0038] In this embodiment, for the 800 sets of one-to-one mapping glass components and their corresponding elastic moduli (specifically, the true elastic moduli), 640 sets are used as the training set to optimize the parameters of the initially established random forest; the other 160 sets are used as the test set to test the performance of the optimized random forest model in the future.

[0039] After completing the initial random forest model, the process of optimizing the parameters of the random forest model is described in step 4 below:

[0040] Step 4: Introduce the spider-bee optimization algorithm to optimize the parameters of the selected random forest model;

[0041] Step 4 includes the following steps 4-1 to 4-8:

[0042] Step 4-1: The hyperparameters of the random forest algorithm (number of decision trees, maximum depth of decision trees, minimum number of separated samples) are the parameters that need to be optimized. The position coordinates of the spider bee are the parameters that need to be optimized. Define the population size of the spider bee optimization algorithm as 50 and the maximum number of iterations as 1000.

[0043] In the spider-bee optimization algorithm, the position coordinates of each spider bee (all spider bees in the population are female) in the spider-bee population represent a solution to the optimization problem (i.e., a value of a parameter to be optimized in the random forest model). If there are D parameters to be optimized, the position coordinates of the spider bees are specifically represented as follows:

[0044] in, The coordinates of the i-th spider wasp in the spider wasp population are represented by SW. i,D To SW i,D For the different components of the position coordinates, SW i,1 To SW i,D These are values ​​for the parameters that need to be optimized in the D term, which represents the i-th spider wasp.

[0045] The initial population of the spider-bee optimization algorithm can be represented as follows:

[0046] Among them, SW Pop Let N be the initial population of spider wasps, and N be the population size. Each row in the matrix on the right side of Equation (1) represents the position coordinates of a spider wasp in the population.

[0047] The following formula can be used to generate any solution in the search space:

[0048] Where t is the generation index (representing the t-th iteration); i is the median index (representing the i-th spider wasp in the population), i = 1, 2, 3, ..., N; that is, The coordinates representing the position of the i-th spider wasp in the t-th generation spider wasp population; Characterize the lower bound of the search; Characterize the upper bound of the search; It is a D-dimensional vector composed of D-dimensional random initialization numbers between 0 and 1. and It can be preset according to actual needs.

[0049] In practical applications, the location coordinates of N spider bees can be randomly generated based on the above formula (2) (the initial value of N can be preset according to actual needs). The location coordinates of these N spider bees can be substituted into formula (1) to complete the establishment of the initial population.

[0050] Step 4-2: In the search phase, the female spider wasp randomly explores the search space with a constant step size, looking for spiders suitable for their offspring. This behavior is modeled as shown in Equation (3), updating the current position of each female spider wasp at each generation t with a constant motion (i.e., constant speed) to simulate the spider wasp's exploration behavior.

[0051] Here, 'a' and 'b' are two indicators (i.e., the aforementioned mesosome index) randomly selected from the population to determine the direction of exploration. In other words, and Two spider bees are randomly selected from the t-th generation spider bee population. The spider bees will follow the determined exploration direction to explore in order to find spiders. Let r be the position coordinates of the i-th spider wasp at the (t+1)-th iteration; use μ1 to determine the constant motion through the current direction, as follows: μ1=|rn|×r1 (4)

[0052] Here, r1 is a number randomly generated between 0 and 1, and rn is also a random number, but generated using a normal distribution.

[0053] Step 4-3: Sometimes, spider wasps cannot find the spider that fell from the ball; therefore, they search the entire area around the exact location where the spider fell. To address this behavior, in order to enable the optimization algorithm to explore the area around the spider's fall point with a smaller step size than Equation (3), another equation was constructed under a different exploration method. This equation randomly selects spider wasps from the population to represent the location of the fallen spider and updates the spider wasp's position based on the fact that the spider wasps have a constant movement in each generation. The equation is described as follows: μ2=B×cos(2πl) (6)

[0054] Where c is an indicator randomly selected from the population, that is... Let l be a randomly selected spider wasp from the t-th generation spider wasp population, and l be a randomly generated number between 1 and -2. B = 1 / (1+e^(-t)) l ), It is a D-dimensional vector randomly generated between 0 and 1.

[0055] The random implementation of generating the spider wasp at the next position (that is, generating the position of the spider wasp for the next iteration) based on formulas (3) and (5) is as follows:

[0056] Among them, r3 and r4 are two random numbers in the range [0,1].

[0057] Step 4-4: Follow and Escape Phase: This simulates the process of a spider wasp chasing a spider, where the distance between the prey (spider) and the spider wasp is initially small and may increase or decrease depending on the speed of both (that is, the distance between the spider wasp and the spider may increase or decrease due to the difference in speed between them). Under this trend, the process of a spider wasp chasing a spider can be simulated by the following formula:

[0058] Where 'a' is an indicator randomly selected from the population; 't' and 't' are... max These represent the current iteration number and the maximum value of the preset t (i.e., the maximum iteration number), respectively. r is a D-dimensional vector, randomly generated within the interval [0,1]; r6 is a random number within the interval [0,1]; C is a distance control factor that determines the speed of the spider bee, used to control the speed of the spider bee starting from 2 and decreasing linearly to 0.

[0059] Equations (8) and (9) can simulate two scenarios: (1) the spider-wasp is faster than the spider, i.e., C > 0.5, and (2) the prey (spider) is faster than the spider-wasp, i.e., C < 0.5. In the latter case, the spider-wasp moves faster than the prey, simulating the process of the spider escaping the spider-wasp, and the distance between the spider-wasp and the spider increases. Therefore, when the spider-wasp moves at a speed less than 0.5, its rate of change of position is so small that it cannot catch up with the spider-wasp.

[0060] As the spider flees the spider wasp, the distance between them gradually increases. This stage is the initial development phase. This behavior is simulated using the following formula as the distance increases:

[0061] in It is a vector generated based on a normal distribution between k and -k. Therefore, k is generated using equation (11), and the distance between the spider wasp and the spider is gradually increased. k = 1 - (t / t) max (11)

[0062] The trade-off between the catch-up trend represented by equation (8) and the escape trend represented by equation (9) is achieved stochastically, as shown in the following equation:

[0063] The exchange between the search phase and its subsequent mechanisms (i.e., the follow and escape phases) is adjusted according to the following formula:

[0064] Where p is a random number in [0,1].

[0065] Steps 4-5: The utilization phase. The first equation is to lure the spider to the most suitable area for it, assuming it's the best location to build a nest, so that the paralyzed spider can be placed there and lay its eggs on its abdomen. The first equation is described as follows:

[0066] in This represents the optimal solution so far. Specifically, for the latest generation of spider wasp populations, the fitness value of the solution represented by each spider wasp in the population can be determined based on a preset fitness function, and the solution represented by the spider wasp with the highest fitness value is considered the optimal solution.

[0067] Specifically, during the iterative process of the spider-bee optimization algorithm, the fitness value of each solution represented by a spider-bee (i.e., a value of the parameter to be optimized in the random forest model) is evaluated. The fitness value reflects the quality of the solution for the optimization problem and guides the search process for the optimal solution. During the algorithm's iteration, solutions to the optimization problem need to be selected, updated, and eliminated based on the fitness value. Furthermore, the fitness value can also be used as one of the conditions for stopping the algorithm's iteration (for example, stopping the iteration might be due to reaching the maximum number of iterations, the spider-bee's fitness value converging to a stable state, or the spider-bee's fitness value reaching a preset threshold, etc.).

[0068] Fitness functions are typically designed as expressions directly related to the objective of the optimization problem. In this embodiment, the fitness function specifically reflects the performance of the random forest model under the parameter configuration of the spider-bee representation to be evaluated. In one example, the fitness function can be based on the accuracy of the random forest model's predictions of the elastic modulus of descriptors in the training set, as follows:

[0069] Where f() represents the fitness function, N total Represents the total number of samples contained in the training set. Characterizing the random forest model in The number of samples accurately predicted under the given parameter configuration. For a random forest model to accurately predict a sample means that the elastic modulus predicted by the random forest model for that sample is consistent with the actual elastic modulus corresponding to that descriptor in the training set.

[0070] The second equation will build nests at locations randomly selected from the spider wasps in the population, using an additional step size to avoid building two nests at the same location. The second equation is designed as follows:

[0071] Where r3 is a random number generated in the interval [0,1]; γ is a number generated based on levy flights; and a, b, and c are indices of three solutions randomly selected from the population. It is a binary vector used to determine when to apply a step size to avoid building two nests at the same location; The allocation formula is as follows:

[0072] in, and These are two vectors representing random values ​​within the interval [0,1]. The different formulas (14) and (15) characterizing the nest-building behavior of spider wasps are randomly interchanged as follows:

[0073] After constructing the nest, the spider wasp will lure suitable spiders into the pre-prepared nest. The exchange between the follow / escape phase and the exploitation phase is adjusted according to the following formula:

[0074] Steps 4-6: Matching Phase: Each spider wasp represents a possible solution for the current generation (i.e., a possible value of the parameter to be optimized in the random forest model), while the spider wasp egg represents a newly generated potential solution for that generation. The new solution (i.e., the spider wasp egg) is generated by the following formula:

[0075] Crossover is and The unified crossover operator applied between them, where CR is the crossover rate: and These are two vectors representing the female and male spider wasps, respectively. The male spider wasp vector is generated based on the following formula:

[0076] In the formula Let be the position of the male spider wasp in the (t+1)th iteration, β and β1 are two numbers randomly generated according to a normal distribution, and e is an exponential constant. and It is generated by the following formula:

[0077] In the formula, a, b, and c are the indices of three solutions randomly selected from the population (that is, the median indices of three spider wasps randomly selected from the population), a≠i≠b≠c, and f is the fitness function. The crossover operation is used to recombine the genetic material of the parent spider wasps to produce offspring (eggs) with the characteristics of both parents.

[0078] Since a, b, and c are the exponents of three solutions randomly selected from the population, equations (21) and (22) above are essentially:

[0079] in, and These are three spider wasps randomly selected from the current generation of the spider wasp population. Let i represent the i-th spider bee in the t-th iteration.

[0080] Steps 4-7: During the iteration process, some spider wasps in the population will be terminated, providing more functional evaluations for other spider wasps. This also reduces population diversity and accelerates convergence towards the near-optimal solution. Each time the entire function is evaluated, the length of the new population (i.e., the size of the new population) will be updated using the following formula: N' = N min +(NN min )×k (23)

[0081] Where, N min N is the preset minimum population size, N is the current population size, and N' is the updated population size.

[0082] Steps 4-8: Output the location and accuracy (fitness value) of the spider wasp. The location of the spider wasp is the parameter value of the random forest.

[0083] The specific implementation of steps 4-8 will be explained in detail in steps A4-1 to A4-7 below.

[0084] Step 5: Build a top-performing random forest model based on the optimized parameters;

[0085] The above briefly explains the basic principles of the Spider-Bee optimization algorithm and the approach to applying it to parameter optimization in random forest models. In practical applications, the Spider-Bee optimization algorithm can be integrated into the parameter optimization process of random forest models based on the following workflow:

[0086] Step A1: Obtain characteristic data of various different glasses and the corresponding relationship between their elastic moduli; the characteristic data includes glass composition and glass preparation process parameters.

[0087] Step A2: Establish a training set based on the correspondence.

[0088] Step A3: Obtain the random forest model to be trained and determine the parameters to be optimized for the random forest model; the random forest model is used to output the predicted elastic modulus based on the input feature data.

[0089] For details on steps A1-A3, please refer to the previous explanation of step 4.

[0090] Step A4: Based on the preset fitness function, optimize the parameters of the random forest model using the spider-bee optimization algorithm to obtain a trained random forest model; the fitness function is constructed based on the accuracy of the random forest model in predicting the elastic modulus of the feature data in the training set.

[0091] Specifically, step A4 can be implemented based on the following detailed steps:

[0092] Step A4-1: Using equation (2), generate the initial spider bee population; the position coordinates of each spider bee in the spider bee population represent a value of the parameter to be optimized in the random forest model.

[0093] Step A4-2: Calculate the fitness value of each spider wasp in the initial spider wasp population.

[0094] After establishing the initial spider wasp population, iterative updates to the spider wasp population can be performed based on steps A4-3 to A4-6 as follows:

[0095] Step A4-3: Select at least a portion of the spider bees in the current spider bee population. The selected spider bees are used to generate the location coordinates of the next generation of spider bees.

[0096] Step A4-4: Generate r6 and determine the relationship between r6 and the preset weight ratio TR. If r6 is less than TR, then for each selected spider wasp, update the position coordinates of the spider wasp based on equation (18) to obtain a new generation of spider wasps; if r6 is not less than TR, then for each selected spider wasp, update the position coordinates of the spider wasp based on equation (19) to obtain a new generation of spider wasps.

[0097] Step A4-4 involves updating the location of the spider wasp. For details, please refer to the explanations of steps 4-2 to 4-6 above.

[0098] Step A4-5: Update the N value based on equation (23) to obtain the updated population size N'. Calculate the fitness value of each new generation of spider bees and sort each new generation of spider bees and each spider bee in the current spider bee population in descending order of fitness value. Select the top N' spider bees with the highest fitness values ​​from the sorted spider bees to form the next generation of spider bee population.

[0099] Step A4-6: Determine whether the preset termination iteration condition has been met (e.g., the preset number of iterations has been reached, or the fitness value of the individual with the highest fitness value in the population exceeds the threshold). If not met, return to step A4-3; if met, proceed to step A4-7.

[0100] Step A4-7: Identify the spider wasp with the highest fitness value in the latest generation of spider wasp population, and use the model parameter value represented by the location coordinates of this spider wasp as the target value of the parameter to be optimized, so as to obtain the trained random forest model (that is, the random forest model with optimal parameters).

[0101] Step 6: For the glass composition to be predicted, use the optimal random forest model to predict the glass elastic modulus of the glass composition.

[0102] Specifically, refer to the explanations in steps 2-1 and 2-2 above to construct a descriptor for the glass composition to be predicted using elemental molar content and preparation process parameters (that is, construct a descriptor based on the characteristic data of the glass). Input the constructed descriptor into the random forest model established in step 5 to obtain the glass elastic modulus predicted by the random forest for the glass composition to be predicted.

[0103] Example 2:

[0104] Step 1: Collect elastic modulus data of oxide glasses with different compositions and construct a glass elastic modulus database. Specifically, collect elastic modulus data of 1600 sets of SiO2-Na2O-CaO-Al2O3 glasses and construct a glass elastic modulus database.

[0105] Step 2: Use elemental molar content and preparation process parameters as descriptors for input parameters;

[0106] Step 2 includes the following steps 2-1 and 2-2:

[0107] Step 2-1: Use the elemental molar content of 1600 SiO2-Na2O-CaO-Al2O3 glasses as a descriptor;

[0108] Step 2-2: Construct a descriptor using glass preparation process parameters (heating rate, melting temperature, holding time).

[0109] Step 3: Using the descriptors constructed in Step 2 as the input to the model and the elastic modulus database constructed in Step 1 as the output of the model, construct 1280 sets of data to build a training set and 320 sets of data to build a test set, and establish a random forest model.

[0110] The process of establishing the initial random forest model described above is consistent with steps 1-3 of Example 1 above, and you can refer to the description above for details.

[0111] Step 4: Introduce the spider-bee optimization algorithm to optimize the parameters of the selected random forest model;

[0112] Step 4 includes the following steps 4-1 to 4-8:

[0113] Step 4-1: The hyperparameters of the random forest algorithm (number of decision trees, maximum depth of decision trees, minimum number of separated samples) are the parameters that need to be optimized. The position coordinates of the spider bee are the parameters that need to be optimized. Define the population size of the spider bee optimization algorithm as 100 and the maximum number of iterations as 2000.

[0114] For details on defining the spider wasp population, please refer to the explanation of step 4-1 in Example 1 above.

[0115] Step 4-2: Search Phase. The female spider wasp randomly explores the search space with a constant step size, as previously described, to find spiders suitable for her offspring. This behavior is modeled as shown in Equation (3), updating the current position of each female wasp at each generation t with a constant motion (i.e., constant speed) to simulate the female spider wasp's exploration behavior.

[0116] Here, a and b are two indicators randomly selected from the population to determine the exploration direction. Spider wasps will follow the determined exploration direction to search for spiders. μ1 is used to determine the constant movement through the current direction, as shown in the following formula: μ1=|rn|×r1 (4)

[0117] Here, r1 is a number randomly generated between 0 and 1, and rn is also a random number, but generated using a normal distribution.

[0118] Step 4-3: Sometimes, spider wasps cannot find the spider that fell from the ball; therefore, they search the entire area around the exact location where the spider fell. To address this behavior, in order to enable the optimization algorithm to explore the area around the spider's fall point with a smaller step size than Equation (3), another equation is constructed under a different exploration method. This equation randomly selects spider wasps from the population to represent the location of the fallen spider and updates the spider wasp's position based on the fact that the spider wasps have a constant movement in each generation. The equation is described as follows: μ2=B×cos(2πl) (6)

[0119] Where c is an index randomly selected from the population, l is a number randomly generated between 1 and -2, and B = 1 / (1+e)l ), It is a D-dimensional vector randomly generated between 0 and 1.

[0120] The randomization of the spider-wasp at the next position based on formulas (3) and (5) is implemented as follows:

[0121] Among them, r3 and r4 are two random numbers in the range [0,1].

[0122] Step 4-4: Follow and Escape Phase: This phase simulates the process of a spider wasp chasing a spider, where the distance between the prey (spider) and the spider wasp is initially small and may increase or decrease depending on the speed of both. Under this trend, the process of a spider wasp chasing a spider can be simulated by the following formula:

[0123] Where 'a' is the index of a solution randomly selected from the population; 't' and 't' are also present. max These represent the current generation number of the population and the maximum value of the preset t, respectively. r is a D-dimensional vector, randomly generated within the interval [0,1]; r6 is a random number within the interval [0,1]; C is a distance control factor that determines the speed of the spider bee, used to control the speed of the spider bee starting from 2 and decreasing linearly to 0.

[0124] Equations (8) and (9) can simulate two scenarios: (1) the spider wasp is faster than the spider (prey), i.e., C > 0.5, and (2) the spider is faster than the spider wasp, i.e., C < 0.5. In scenario (2), the spider wasp moves faster than the prey, simulating the process of the spider escaping the spider wasp, and the distance between the spider wasp and the spider increases. Therefore, when the spider wasp moves at a speed less than 0.5, its rate of change of position is so small that it cannot catch up with the spider wasp.

[0125] As the spider flees the spider wasp, the distance between them gradually increases. This stage is the initial development phase. This behavior is simulated using the following formula as the distance increases:

[0126] in It is a vector generated based on a normal distribution between k and -k. Therefore, k is generated using equation (11), and the distance between the spider wasp and the spider is gradually increased. k = 1 - (t / t) max (11)

[0127] The trade-off between the catch-up trend represented by equation (8) and the escape trend represented by equation (9) is achieved stochastically, as shown in the following equation:

[0128] The exchange between the search phase and its subsequent mechanisms (i.e., the follow and escape phases) is adjusted according to the following formula:

[0129] Where p is a random number in [0,1].

[0130] Steps 4-5: The utilization phase. The first equation is to lure the spider to the most suitable area for it, assuming it's the best location to build a nest, so that the paralyzed spider can be placed there and lay its eggs on its abdomen. The first equation is described as follows:

[0131] in The first equation represents the optimal solution so far. The second equation will build nests at locations randomly selected from the population of spider wasps, using an additional step size to avoid building two nests at the same location. The second equation is designed as follows:

[0132] Where r3 is a random number generated in the interval [0,1]; γ is a number generated based on the levy flight; a, b, and c are indices of three solutions randomly selected from the population; It is a binary vector used to determine when to apply a step size to avoid building two nests at the same location; The allocation formula is as follows:

[0133] in, and These are two vectors representing random values ​​within the interval [0,1]. The different formulas (14) and (15) characterizing the nest-building behavior of spider wasps are randomly interchanged as follows:

[0134] After constructing the nest, the spider wasp will lure suitable spiders into the pre-prepared nest. The switching between the follow / escape phase and the exploit phase is adjusted according to the following formula:

[0135] Steps 4-6: Matching Phase: Each spider wasp represents a possible solution for the current generation, while the spider wasp egg represents a newly generated potential solution for that generation. The new solution (i.e., the spider wasp egg) is generated by the following formula:

[0136] Crossover is and The unified crossover operator applied between them, where CR is the crossover rate; and These are two vectors representing the female and male spider wasps, respectively. The male spider wasp vector is generated based on the following formula:

[0137] In the formula Let represent the position of the male spider wasp in generation t+1, where β and β1 are two numbers randomly generated according to a normal distribution, and e is an exponential constant. and It is generated by the following formula:

[0138] In the formula, a, b, and c are the exponents of three solutions randomly selected from the population, such that a≠i≠b≠c, and f is the fitness function. Crossover is used to recombine the genetic material of the two parent spider wasps to produce offspring (eggs) with characteristics of both parents.

[0139] Steps 4-7: During the iteration process, some spider wasps in the population will be terminated, providing more functional evaluations for other spider wasps. This also reduces population diversity and accelerates convergence towards the near-optimal solution. Each time the entire function is evaluated, the length of the new population (i.e., the size of the new population) will be updated using the following formula: N' = N min +(NN min )×k (23)

[0140] Where, N min N is the preset minimum population size, N is the current population size, and N' is the updated population size.

[0141] Steps 4-8: Output the location and accuracy (fitness value) of the spider wasp. The location of the spider wasp is the parameter value of the random forest.

[0142] For details, please refer to the previous description of steps 4-8 in Example 1.

[0143] Step 5: Build a top-performing random forest model based on the optimized parameters;

[0144] Step 6: For the glass composition to be predicted, the optimal method is to use a random forest model to predict the glass elastic modulus of the glass composition.

[0145] Steps 5-6 above can be referred to the previous description of steps 5-6 in Example 1.

[0146] After testing the performance of the random forest model obtained in step 5, the model performance of Examples 1-2 is shown in the table below:

[0147] Table 1: Performance Table of Models in Examples 1 and 2

[0148] In the table, R² (R-Square, coefficient of determination) ranges from [0,1]. The larger the value, the more accurate the model prediction.

[0149] The smaller the values ​​of MSE (Mean Squared Error) and MAE (Mean Absolute Error), the smaller the prediction error of the model.

[0150] As can be seen from Table 1, after optimizing the parameters of the random forest model using the spider-bee optimization algorithm, the prediction performance of the random forest model for the elastic modulus of glass was significantly improved.

[0151] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in this application, based on the technical solution and inventive concept of this application, should be included within the scope of protection of this application.

Claims

1. A method for predicting glass elastic modulus based on a spider optimization algorithm and a machine learning algorithm, characterized in that, The method comprises the following steps: Step 1: Collect the elastic modulus data of different component oxide glasses, and construct a glass elastic modulus database comprising a one-to-one mapping of glass components and their corresponding elastic modulus; Step 2: Take the element molar content and the preparation process parameters as input parameters, i.e., take the molar content of each component of the composition glass as a group of descriptors, and take the glass preparation process parameters, including the heating rate, melting temperature, and holding time, to construct the descriptors; Step 3: Take the descriptors constructed in step 2 as the input of the random forest model, and take the elastic modulus database constructed in step 1 as the output of the random forest model, construct the training set and the test set, and establish the random forest model; Step 4: Introduce the spider optimization algorithm to optimize the parameters of the selected random forest model; Step 5: Establish the random forest model with the optimal performance based on the optimized parameters; Step 6: For the glass components to be predicted, the optimal random forest model is used to predict the glass elastic modulus of the glass components. 2.The glass modulus of elasticity prediction method based on the spider optimization algorithm and the machine learning algorithm according to claim 1, wherein, The step 4 comprises the following sub-steps: Step 4-1: Select the parameters to be optimized for the random forest algorithm, the position coordinates of the spider mites are the parameters to be optimized, define the population size N of the spider mite optimization algorithm, and search the upper bound Search lower bound where SW Pop is the initial population of spider bees, the following equation can generate any solution in the search space: wherein t is a generation index; i is a middle index, i = 1, 2, 3, …, N; It is a vector composed of a D-dimensional random initialization between 0 and 1; Step 4-2: Search phase, the wasp explores the search space randomly with a constant step size, looking for a spider that is suitable for its offspring, as described previously, this behavior is modeled as shown in equation (3), which updates the current position of each wasp with a constant motion at each generation t to simulate the exploratory behavior of the wasp: Where a and b are two indices randomly selected from the population to determine the exploration direction, the spider explores according to the exploration direction, and the constant motion of the current direction is determined by μ1, and the formula is as follows: μ1=|rn|×r1 (4) Where r1 is a number randomly generated between 0 and 1, and rn is also a random number, but generated using a normal distribution; Step 4-3: Spider mites sometimes cannot find the spider that has fallen from the ball, so they search the entire area around the exact place where the spider fell; in response to this behavior, in order to enable the proposed algorithm to explore the area around the place where the spider fell with small steps different from equation (3), another equation under different exploration methods was constructed, which randomly selects a spider mite from the population to represent the position of the fallen spider and updates the position of the spider mite based on the fact that the spider mite has a constant movement in each generation, which is described as follows: μ2=B×cos(2πl) (6) Where c is an index randomly selected from the population, and l is a number randomly generated between 1 and -2, Equations (3) and (5) generate the random implementation of the female spider mite at the next position as follows: Where r3 and r4 are two random numbers in [0,1]; Step 4-4: Chasing and escaping phase: The process of the spider wasp chasing the spider was simulated, where the distance between the spider as prey and the wasp was initially small and could increase or decrease depending on the speed of the wasp and the prey, under this trend, we proposed a mathematical model to simulate two cases: (1) the wasp is faster than the spider (prey) C>0.5, (2) the prey is faster than the wasp C<0.5, where C is a distance control factor that determines the speed of the wasp, starting from a speed of 2, linearly decreasing to 0; in the latter case, the wasp moves faster than the prey, simulating the escape of the spider as prey from the wasp, the distance between the wasp and the spider increases, so when the wasp moves at a speed less than 0.5, the rate of change of its position is particularly small, so that it cannot reach the prey, the formula is as follows: where a is an index randomly selected from the population: t and t max respectively represent the current iteration number and the maximum value of t preset: It is a vector that represents a value randomly generated in the interval [0,1]: r6 is a random number in the interval [0,1]; As the spider prey escapes from the spider wasp, the distance between the spider wasp and the spider increases gradually, this stage is the initial development stage, with the increase of the distance, the following formula is used to simulate this behavior: wherein It is a vector generated between k and -k according to the normal distribution, according to which k is generated by formula (11), gradually increasing the distance between the spider and the spider, k = 1 - (t / t max ) (11) The trade-off between these two trends is realized randomly, as shown in the following equation: The exchange between the search phase and the follow and flee phase is adjusted according to the following formula: Where p is a random number in [0,1]; Step 4-5: Utilization phase, the first equation is based on pulling the spider towards the area that is most suitable for the spider and considering it as the best place to build a nest in order to place the paralyzed spider at the place and lay eggs in its abdomen, the first equation is described as follows: wherein Representing the best solution so far, the second equation will build a nest at the location of a randomly selected spider mite from the population, using an extra step to avoid building two nests at the same location, the equation is designed as follows: wherein r3 is a random number generated in the interval [0, 1]; γ is a number generated according to the Zephyr flight: a, b, c are indices of three solutions randomly selected from the population; is a binary vector that determines when to apply a step to avoid building two nests at the same location; The allocation formula for the above is as follows: wherein and are two vectors representing random values in the interval [0,1], (14), (15) are randomly permuted according to The spider bee will drag the appropriate spider to the pre-prepared nest, Step 4-6: Matching phase: Each wasp represents a possible solution of the current generation, while the wasp egg represents a newly generated potential solution of the generation, the new solution / wasp egg is generated by the following formula: wherein Crossover is a uniform crossover operator applied between solutions, and having a cross ratio (CR); and are two vectors representing female and male wasps, respectively, and the male wasps generated in our proposed algorithm differ from the female wasps according to the following formula: where β and β1are two numbers randomly generated according to a normal distribution, e is the exponential constant, and are generated from the following formula: In the formula, a, b, and c are three indices of solutions randomly selected from the population, and a≠i≠b≠c, the crossover is the recombination of the genetic material of the parent spider, and the offspring egg has the characteristics of the parents; Step 4-7: In the iteration process, some spiders in the population will be terminated to provide more function evaluations for other spiders, while also reducing the diversity of the population and accelerating the convergence speed to the near-optimal solution. The length of the new population will be updated using the following formula at each function evaluation: N = N min + (N - N min ) x k (23) Step 4-8: Output the position and accuracy of the spider, and the position of the spider is the parameter value of the random forest.

3. A method of training a random forest model for predicting glass elastic modulus, characterized in that, It comprises: Obtaining the corresponding relationship between the characteristic data of a plurality of different glasses and the elastic modulus thereof; The characteristic data comprises glass components and glass preparation process parameters; Establishing a training set based on the corresponding relationship; Obtaining a random forest model to be trained, and determining parameters to be optimized of the random forest model; the random forest model is used to output a predicted elastic modulus based on input feature data; According to a preset fitness function, the parameters to be optimized of the random forest model are optimized by a bumblebee optimization algorithm to obtain a trained random forest model; the fitness function is constructed according to an accuracy rate of the elastic modulus predicted by the random forest model for the feature data in the training set.

4. The method of claim 3, wherein, The optimization of the parameters to be optimized of the random forest model according to the preset fitness function by the bumblebee optimization algorithm comprises: Constructing an initial bumblebee population; wherein the position coordinates of each bumblebee in the bumblebee population represent a value of the parameter to be optimized; Iteratively updating the current bumblebee population based on the bumblebee optimization algorithm to generate a next-generation bumblebee population until a preset termination iteration condition is met; Determining the fitness values of each bumblebee in the latest generation bumblebee population based on the fitness function, taking the value represented by the bumblebee with the highest fitness value as the target value of the parameter to be optimized, and obtaining the trained random forest model.

5. The method of claim 4, wherein, The iterative updating of the current bumblebee population based on the bumblebee optimization algorithm to generate a next-generation bumblebee population comprises: Selecting at least part of the bumblebees in the current bumblebee population; For each selected bumblebee, determining a target stage from each update stage based on a random decision mechanism between the update stages in the bumblebee optimization algorithm, and updating the position coordinates of the bumblebee based on the position update formula corresponding to the target stage to obtain the position coordinates of the new-generation bumblebee; the update stages include a search stage, a following and escaping stage, a utilization stage and a matching stage, and each update stage has a corresponding position update formula; Determining the fitness values of each new-generation bumblebee based on the fitness function, and sorting each new-generation bumblebee and each bumblebee in the current bumblebee population in order of fitness value from high to low; Obtaining a first quantity representing the population size of the next-generation bumblebee population, and obtaining the next-generation bumblebee population from the first quantity of bumblebees in the sorted bumblebees.

6. The method of claim 4, wherein, The position coordinates of the wasps in the initial wasp population are generated based on the following formula: wherein, characterizing the position coordinates of the ith spider wasp in the tth generation of spider wasp population; Characterizing search lower bounds characterizing search upper bounds; is a D-dimensional vector composed of a D-dimensional random initialization array between 0 and 1.

7. The method of claim 5, wherein, The determination of a target stage from each update stage based on the random decision mechanism between the update stages in the bumblebee optimization algorithm comprises: Determining whether r6 is less than a preset trade-off rate; r6 is a random number in the interval [0, 1]; When r6 is not less than the trade-off rate, the matching stage is determined as the target stage; When r6 is less than the trade-off rate, it is judged whether i is less than N x k; N is the population size of the current wasp population, k = 1 - (t / t max ), t and t max respectively represent the current iteration number and the maximum value of preset t; When i is not less than N×k, the utilization stage is determined as the target stage; when i is less than N×k, it is determined whether p is less than k; p is a random number in the interval [0, 1]; When p is less than k, the search stage is determined as the target stage, and when p is not less than k, the following and escaping stage is determined as the target stage.

8. The method according to claim 5 or 7, characterized in that, The position update formula corresponding to the search phase is: wherein a, b and c are meso indices randomly selected from the current wasp population; characterizing the position coordinates of the ith spider wasp in the tth generation of spider wasp population; Represents the position coordinates of the i-th spider wasp at the (t+1)-th iteration; μ2=B×cos(2πl), B=1 / (1+e l ), where l is a randomly generated number between 1 and -2; is a D-dimensional vector randomly generated between 0 and 1 ; r3 and r4 are two random numbers in the interval [0, 1 ]; Characterizing search lower bounds representing the search upper bound; The position update formula corresponding to the following and escaping phases is: wherein is a D-dimensional vector, randomly generated in the interval [0,1]; r6is a random number in the interval [0,1]; is a vector generated according to a normal distribution between k and -k, k = 1 - (t / t max ); t and t max represent the current iteration number and the preset maximum value of t, respectively; The position update formula corresponding to the utilization stage is: wherein, represents the best solution so far; γ is a number generated according to the penalty flight; and are two vectors representing random values in the interval [0, 1]; The position update formula corresponding to the matching stage is: wherein Crossover is and The uniform crossover operator is applied between them, CR is the crossover rate: and are two vectors representing a female spider wasp and a male spider wasp, respectively; wherein the male spider wasp is generated based on the following formula: β and β1are two numbers randomly generated according to a normal distribution, e is an exponential constant, and are generated from the formula: Wherein, f() represents the fitness function.

9. The method of claim 5, wherein, The population size of the next generation of spider wasp population is determined based on the following equation: N' = N min + (N - N min ) x k where N' is the population size of the next generation of spider mite population, N is the population size of the current spider mite population min is a preset minimum population size, and N is the population size of the current spider mite population.

10. A method of predicting the modulus of elasticity of a glass, characterized by, Comprise: Obtaining characteristic parameters of a glass to be predicted; the characteristic data comprises glass components and glass preparation process parameters; Inputting the characteristic parameters into a pre-trained random forest model to obtain an elastic modulus predicted by the random forest model for the characteristic parameters; The random forest model is obtained by training based on the method in any one of claims 3-9.

Citation Information

Patent Citations

  • Durable concrete mix proportion optimization method based on RF-NSGA-II

    CN111986737A

  • Elastic modulus prediction method based on molecular dynamics and elastic network regression model

    CN113255083A

  • Glass hardness prediction method based on squirrel optimization algorithm and machine learning algorithm

    CN114373523A

  • Oxide glass performance prediction method based on interpretable high-dimensional space prediction model

    CN117497087A

  • Glass elastic modulus prediction method based on spider bee optimization algorithm and machine learning algorithm

    CN118969144A