Performance evaluation method of sparse robust extreme learning machine based on L1 / L2

By introducing L1/L2 norm regular terms and non-convex loss functions in the extreme learning machine model, the problem of ELM model being sensitive to outliers and difficulty in optimizing NP in L0 norm is solved, achieving higher sparsity and robustness, and improving the generalization performance of the model.

CN119990223APending Publication Date: 2025-05-13XI'AN PETROLEUM UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510060959.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, the classic ELM model is sensitive to outliers and noise, and the L0 norm optimization problem NP is difficult, which is difficult to effectively solve in practical applications, resulting in insufficient model sparsity and robustness.

Method used

By introducing the L1/L2 norm regular terms and using a proximal operator to construct the analytical solution of L1/L2, combined with a non-convex exponential square loss function, the limit learning machine model is optimized to improve sparsity and robustness.

Benefits of technology

Improve the sparsity and robustness of the model and enhance the generalization performance of the model, especially in the presence of outliers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990223A_ABST
    Figure CN119990223A_ABST
Patent Text Reader

Abstract

The invention discloses a performance evaluation method for a sparse robust extreme learning machine based on L1 / L2, and the method comprises the steps: obtaining a preprocessing data set which comprises a training set and a test set, and the training set comprises a preset number of abnormal data; an extreme learning machine model is constructed according to the ELM model, and the extreme learning machine model comprises an L2 norm regular term; optimizing the extreme learning machine model to obtain an optimized extreme learning machine model, so as to convert the solution of the extreme learning machine model into an L1 / L2 analytical solution and a non-convex function solution; training the optimized extreme learning machine model according to the preprocessed data set to obtain a trained extreme learning machine model; and evaluating the trained extreme learning machine model according to sparsity and a root mean square error (RMSE). Through the method provided by the invention, the sparsity and robustness of the model are improved, and the model has better generalization performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of neural networks, and in particular relates to a performance evaluation method of a sparse robust extreme learning machine based on L1 / L2. Background Art

[0002] Sparse hidden layers can reduce the complexity of the model, reduce the risk of overfitting, and improve the generalization ability of the model. However, the number of hidden nodes in the classic ELM is determined based on experience, and some of these hidden nodes may be redundant, which is meaningless to improve the performance of the model. Later, some methods appeared to improve the sparsity of the model, mainly pruning and incremental learning. Pruning is to remove redundant nodes, and the incremental method is the opposite, which is to add nodes one by one or in groups. There is also a combination of L1 norm regularization technology and extreme learning machine to improve the sparsity of the model. But these methods are not sparse enough. Later, it was discovered that the L0 norm can theoretically provide ideal sparsity, that is, the number of non-zero elements, but its optimization problem is usually NP-hard and difficult to solve effectively in practical applications.

[0003] The choice of loss function has an important impact on the robustness of the ELM model. The classic ELM uses a squared loss function, which is very sensitive to outliers and noise, which may affect the robustness of the model in challenging scenarios. To solve this problem, someone built a unified model from the perspective of iterative reweighting, combining different loss functions and regularization terms to enhance the robustness of ELM regression. However, when there are many outliers, the unbounded loss function will still be affected. Summary of the invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a performance evaluation method for a sparse robust extreme learning machine based on L1 / L2. The technical problem to be solved by the present invention is achieved by the following technical solutions:

[0005] A performance evaluation method for a sparse robust extreme learning machine based on L1 / L2, comprising:

[0006] Acquire a preprocessed data set, wherein the preprocessed data set includes a training set and a test set, and the training set includes a preset number of abnormal data;

[0007] Constructing an extreme learning machine model according to the ELM model, wherein the extreme learning machine model includes an L2 norm regularization term;

[0008] Optimizing the extreme learning machine model to obtain an optimized extreme learning machine model, so as to convert the solution of the extreme learning machine model into an L1 / L2 analytical solution and a non-convex function solution;

[0009] Training the optimized extreme learning machine model according to the preprocessed data set to obtain a trained extreme learning machine model;

[0010] The trained extreme learning machine model is evaluated according to sparsity and root mean square error RMSE.

[0011] In a specific embodiment, after obtaining the preprocessed data set, the method further includes:

[0012] Calculate the scaling factor based on the number of samples in the preprocessed dataset;

[0013] Calculate a normalization coefficient and a normalization offset according to the scaling factor;

[0014] The preprocessed data set is normalized according to the normalization coefficient and the normalization offset.

[0015] In a specific embodiment, the preprocessed data set includes a pollution data set and a NO2 data set, the ratio of the training set to the test set is 2:1, and the preset number is 10%-40% of the training set data.

[0016] In a specific implementation, the extreme learning machine model is:

[0017]

[0018] stβ=u,

[0019] Where C is the regularization parameter, β = [β1,β2,…,β L ] T is the weight connecting the hidden layer and the output layer, z i =y i -h(x i )u,z i is the training error between the predicted value and the true value of the i-th training sample, and σ is the upper bound of the exponential square loss function.

[0020] In a specific implementation, converting the solution of the extreme learning machine model into an L1 / L2 analytical solution and a non-convex function solution includes:

[0021] Convert the solution of the extreme learning machine model into the solution of the following formula:

[0022]

[0023] Where C is the regularization parameter, β is the output weight vector, u is the auxiliary variable, v is the Lagrange multiplier, η is the step size, and η>0. k+1 represents the β,u obtained in the k+1th iteration k+1 represents u, v obtained in the k+1th iteration k+1represents v obtained at the k+1th iteration.

[0024] In one specific embodiment, the analytical solution of L1 / L2 is solved by the proximal operator;

[0025] The proximal operator form is

[0026] The analytical solution of L1 / L2 is in,

[0027] In one specific embodiment, the non-convex function solution is solved by a convex difference algorithm;

[0028] The solution of the non-convex function is

[0029] Among them, I1 is the L×L identity matrix, I2 is the N×N identity matrix, H is the N×L hidden layer output matrix, H T is the transpose of the L×N hidden layer output matrix, Y is an N×1 vector, which is the actual value of the training sample, S is an N×1 vector, and η is the step size.

[0030] In a specific embodiment, training the optimized extreme learning machine model according to the preprocessed data set includes:

[0031] Input the training set into the established extreme learning machine model;

[0032] Randomly generate weight w in the range of [-1,1] and bias b in the range of [0,1] to obtain output weight β;

[0033] Initialize β, u, v, that is, β = β0, u = β0, v = β0, substitute into the loop and use the ADMM algorithm to alternately optimize β, u and v until the loop cutoff condition is met to obtain β, where the loop cutoff condition is:

[0034] In a specific embodiment,

[0035] In a specific embodiment, Among them, y i is the ith actual value, is the i-th predicted value.

[0036] Beneficial effects of the present invention:

[0037] Although the L0 norm can provide ideal sparsity in theory, its optimization problem is usually NP-hard and difficult to solve effectively in practical applications. Because of the characteristics of L1 / L2 being non-convex, non-parameterized and scale-invariant, the zero norm can be approximated. Therefore, the present invention introduces the L1 / L2 norm regularization term and constructs the analytical solution of L1 / L2 with the proximal operator; the classic ELM uses a convex square loss function and is very sensitive to outliers and noise samples. Therefore, the present invention introduces a non-convex exponential square loss function to constrain the negative impact of outliers on the model. Through the method of the present invention, the sparsity and robustness of the model are improved, allowing the model to have better generalization performance.

[0038] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a flowchart of a performance evaluation method for a sparse robust extreme learning machine based on L1 / L2 provided by an embodiment of the present invention;

[0040] Figure 2 1 is a schematic diagram of an extreme learning machine network structure provided by an embodiment of the present invention;

[0041] Figure 3 An exponential square loss function image provided for an embodiment of the present invention. DETAILED DESCRIPTION

[0042] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0043] Embodiment 1

[0044] See also Figure 1 , Figure 1 The present invention provides a flow chart of a performance evaluation method for a sparse robust extreme learning machine based on L1 / L2, including:

[0045] Acquire a preprocessed data set, wherein the preprocessed data set includes a training set and a test set, and the training set includes a preset number of abnormal data;

[0046] Constructing an extreme learning machine model according to the ELM model, wherein the extreme learning machine model includes an L2 norm regularization term;

[0047] Optimizing the extreme learning machine model to obtain an optimized extreme learning machine model, so as to convert the solution of the extreme learning machine model into an L1 / L2 analytical solution and a non-convex function solution;

[0048] Training the optimized extreme learning machine model according to the preprocessed data set to obtain a trained extreme learning machine model;

[0049] The trained extreme learning machine model is evaluated according to sparsity and root mean square error RMSE.

[0050] Specifically, after obtaining the preprocessed data set, the following steps are also included:

[0051] Calculate the scaling factor based on the number of samples in the preprocessed dataset;

[0052] Calculate a normalization coefficient and a normalization offset according to the scaling factor;

[0053] The preprocessed data set is normalized according to the normalization coefficient and the normalization offset.

[0054] Specifically, the preprocessed data set includes a pollution data set and a NO2 data set, the ratio of the training set to the test set is 2:1, and the preset number is 10%-40% of the training set data.

[0055] In a specific implementation, the extreme learning machine model is:

[0056]

[0057] stβ=u,

[0058] Where C is the regularization parameter, β = [β1,β2,…,β L ] T is the weight connecting the hidden layer and the output layer, z i =y i -h(x i )u,z i is the training error between the predicted value and the true value of the i-th training sample, and σ is the upper bound of the exponential square loss function.

[0059] In a specific implementation, converting the solution of the extreme learning machine model into an L1 / L2 analytical solution and a non-convex function solution includes:

[0060] Convert the solution of the extreme learning machine model into the solution of the following formula:

[0061]

[0062] Where C is the regularization parameter, β is the output weight vector, u is the auxiliary variable, v is the Lagrange multiplier, η is the step size, and η>0. k+1 represents the β,u obtained in the k+1th iteration k+1 represents u, v obtained in the k+1th iteration k+1 represents v obtained at the k+1th iteration.

[0063] Specifically, the analytical solution of L1 / L2 is solved by the proximal operator;

[0064] The proximal operator form is

[0065] The analytical solution of L1 / L2 is in,

[0066] Specifically, the solution of non-convex functions is solved by the convex difference algorithm;

[0067] The solution of the non-convex function is

[0068] Among them, I1 is the L×L identity matrix, I2 is the N×N identity matrix, H is the N×L hidden layer output matrix, H T is the transpose of the L×N hidden layer output matrix, Y is an N×1 vector, which is the actual value of the training sample, S is an N×1 vector, and η is the step size.

[0069] Specifically, training the optimized extreme learning machine model according to the preprocessed data set includes:

[0070] Input the training set into the established extreme learning machine model;

[0071] Randomly generate weight w in the range of [-1,1] and bias b in the range of [0,1] to obtain output weight β;

[0072] Initialize β, u, v, that is, β = β0, u = β0, v = β0, substitute into the loop and use the ADMM algorithm to alternately optimize β, u and v until the loop cutoff condition is met to obtain β, where the loop cutoff condition is:

[0073] Specifically,

[0074] Specifically, the root mean square error Among them, y i is the ith actual value, is the i-th predicted value.

[0075] Although the L0 norm can provide ideal sparsity in theory, its optimization problem is usually NP-hard and difficult to solve effectively in practical applications. Because of the characteristics of L1 / L2 being non-convex, non-parameterized and scale-invariant, the zero norm can be approximated. Therefore, the present invention introduces the L1 / L2 norm regularization term and constructs the analytical solution of L1 / L2 with the proximal operator; the classic ELM uses a convex square loss function and is very sensitive to outliers and noise samples. Therefore, the present invention introduces a non-convex exponential square loss function to constrain the negative impact of outliers on the model. Through the method of the present invention, the sparsity and robustness of the model are improved, allowing the model to have better generalization performance.

[0076] Embodiment 2

[0077] This embodiment is described with a specific example.

[0078] Step 1: Collect and preprocess data:

[0079] Download the simulation data sets pollution and NO2 from the official website. The pollution data set is 60X15, that is, 60 samples and 15 features, and the NO2 data set is 500X7, that is, 500 samples and 7 features. The data sets are divided into training sets and test sets in a ratio of 2:1. In order to compare the performance of the algorithm in a noisy environment, different proportions of outliers are added to the training data. The number of these outliers is 10%, 20%, 30% and 40% of the original training data, and the outliers are randomly selected from [ymin, ymax]. In order to eliminate the influence of the dimensions of each indicator on the model prediction, we normalize the data set and map the data to the range of [0,1]. Specifically, obtain the number of sample features, calculate the scaling factor sca(i) = maxa(i)-mina(i); then calculate the normalization coefficient A(i) = (1-0) / sca(i); calculate the normalization offset B(i) = lb-A(i)*mina(i); ensure that the normalized data remains within a specific range. Finally, each column in the sample is normalized using the calculated coefficient A(i) and offset B(i).

[0080] Step 2: Establish an extreme learning machine model:

[0081] according to Figure 2 The network structure of the extreme learning machine shown in the figure has a training sample of {(x1,y1),…,(x i ,y i ),…,(x N ,y N )}, when the pollution data set, N = 40, when the NO2 data set, N = 300. Each input sample x i ∈Rn , the corresponding output sample y i ∈R m .h(w j ,x,b j ) is the activation function of ELM, and the output function of ELM with L hidden nodes is in the form of:

[0082]

[0083] Here the hidden layer nodes L = 600, h(x i )=[h(a1,b1,x i ),…,h(a L ,b L ,x i )] is the activation function, β=[β1,β2,…,β L ] T is the weight connecting the hidden layer and the output layer, w j is the weight vector connecting the jth hidden node and the input layer, b j is the threshold of the jth hidden node. In ELM, the error between the actual value and the predicted value is 0, so we can get the equation: Hβ=Y.

[0084] The actual output matrix is ​​Y = [y1,…,y i ,…,y N ] T , the output matrix of the hidden layer is

[0085] It is a linear system of equations, so the output weight can be solved quickly, β = H + Y.

[0086] H + is the Moore-Penrose generalized inverse matrix of the matrix H. Since the standard ELM is prone to overfitting, an L2 norm regularization term is added to the standard ELM to form a regularized extreme learning machine.

[0087]

[0088] st.h(x i )β=y i -ξ i i=1,2,…,N

[0089] Where C is the regularization parameter, ξ i is the training error.

[0090] The optimal solution of β can be obtained according to the Karush-Kuhn-Tucker (KKT) condition.

[0091]

[0092] Step 3: Optimize the extreme learning machine model:

[0093] The extreme learning machine model constructed by the present invention is:

[0094]

[0095] Where C is the regularization parameter, z i =y i -h(x i )u, σ is the parameter of the exponential square loss function, is the upper bound of the loss function, z i is the training error, which is the difference between the predicted value and the true value of the i-th training sample.

[0096] Define the Lagrangian function:

[0097]

[0098] Where C is the regularization parameter, v is the Lagrange multiplier, η is the step size, and η>0.

[0099] The Alternating Direction Method of Multipliers (ADMM) is used to solve convex optimization problems with separable structures. The main idea is to decompose the original problem model into some easily solvable sub-optimization problems by relaxing and augmenting the original problem model. The solution to the original problem is obtained by solving the sub-problems.

[0100] The present invention decomposes the original problem into three sub-problems for solution, and the process of using ADMM iteration is as follows:

[0101]

[0102] Where C is the regularization parameter, β is the output weight vector, u is the auxiliary variable, v is the Lagrange multiplier also called the dual variable, η is the step size of the ADMM optimization algorithm, and η>0. k+1 represents the β,u obtained in the k+1th iteration k+1 represents u, v obtained in the k+1th iteration k+1 represents v obtained at the k+1th iteration.

[0103] Let’s solve the above sub-problems separately:

[0104] Step 3.1: Use the proximal operator to construct the analytical solution of L1 / L2, including:

[0105] If the objective function is differentiable everywhere, then the gradient descent method can be used directly. However, in most cases, the objective function f is not differentiable. Usually, subgradients are used for optimization. However, the disadvantage of subgradients is that they are slow to solve and usually do not produce sparse solutions.

[0106] The proximal operator decomposes f(x) into a differentiable convex function and a non-differentiable convex function, that is: f(x) = g(x) + h(x). The proximal operator is an extension of the gradient. When the function f is a smooth function, the proximal operator is the gradient.

[0107] Proximal operator prox f :R n →R n , f is defined as:

[0108]

[0109] prox f (v) is called the proximal endpoint of v relative to f. As a function of v, the term on the right is also defined as prox f,ρ (v), is called the proximal operator of f with a penalty term ρ.

[0110] Request for solution We use the proximal operator to solve for β. The proximal operator is in the form of prox L1 / L2 (q) represents the continuous update of q, so that the optimal q is found. The function value is minimum. At this time, let β approximate q and get the value of β. Therefore, we can get β k+1 ∈prox L1 / L2 at this time,

[0111] Step 3.2: Solve the non-convex function using the convex difference algorithm, including:

[0112] The exponential square loss function is as follows: the image is Figure 3 As shown:

[0113]

[0114] As can be seen from the image, z represents the difference between the predicted value and the true value, and the parameter σ is used to control the smoothness of the loss function and the sensitivity to outliers. As the value of z increases, the growth rate of the loss function will gradually slow down. The exponential square loss function is a non-convex loss function that has good robustness to outliers and noise. Especially when the data distribution is uneven or has heavy tail characteristics, the exponential square loss function can reduce the sensitivity to extreme values.

[0115] However, the exponential square loss function Φ(u) is non-convex and difficult to optimize. Therefore, the DC algorithm is used to transform the non-convex exponential square loss function into the difference between two convex functions, and Φ(u) = Φ1(u)-Φ2(u). Among them,

[0116] Therefore, we can get Obviously, Φ1(u), Φ2(u) and They are all convex functions, which can be optimized using the convex difference algorithm:

[0117]

[0118] Next, we take the derivative of Φ2(u), and we have make have to therefore but Where S = [s1, s2, …, s N ] T .

[0119] final Finally, we get:

[0120]

[0121] Through Beta k+1 and v k Calculate u k+1 I1 is the L×L identity matrix, I2 is the N×N identity matrix, H is the N×L hidden layer output matrix, H T is the transpose of the L×N hidden layer output matrix, Y is an N×1 vector, which is the actual value of the training sample, S is an N×1 vector, and η is a parameter, which is the step size of the ADMM algorithm.

[0122] Step 4: Training model and prediction model:

[0123] After normalizing the data set and dividing it into training set and test set, the training set is input into the established extreme learning machine model, where the weight w is randomly generated in the range of [-1,1], the bias b is randomly generated in the range of [0,1], and the activation function is the sigmoid function, and the output weight β is obtained, which is recorded as β0. Initialize β, u, v, that is, β=β0, u=β0, v=β0, substitute it into the loop and use the ADMM algorithm to alternately optimize β, u and v until the loop cutoff condition is met to obtain the final β. Among them, the loop cutoff condition is: ∈ is a very small positive number used for numerical stability. If the output weight |β i |<10 -6, then it is recorded as a sparse node sparenodes, which is accumulated continuously. Substitute the output weight β obtained by the training model into the ELM model for prediction, and get the predicted output y_pre.

[0124] Step 5: Experimental analysis and evaluation:

[0125] In the experiment, the optimal parameters are selected by standard grid search. The search range of the regularization parameter C of each algorithm is {2 -19 ,2 -18 ,…,2 19 ,2 20}, the loss function parameter σ of L1 / L2ELM is selected in the range of (0,2), and the step range is (0,100). The sparsity of the hidden layer nodes and the root mean square error RMSE between the predicted output and the actual output are used as the standard to measure the performance and generalization ability of the extreme learning machine. Ten experiments were conducted on the training and test samples respectively, and the average performance of these 10 experiments was taken as the final result. The larger the sparsity value, the sparser the output weight of the model.

[0126]

[0127]

[0128] The present invention conducts a comparative analysis of sparsity and RMSE with ELM, L21ELM, RELM and WELM on two data sets, pollution and NO2, for different outlier ratios.

[0129] Table 1 Sparsity of algorithms under different outlier ratios

[0130]

[0131] Table 2 RMSE of algorithms under different outlier ratios

[0132]

[0133]

[0134] Table 1 shows the sparsity of the algorithms under different outlier ratios. Observing the five experiments on the two data sets, it is found that the change in algorithm sparsity is not significant under different outlier ratios. The sparsity of ELM and WELM is close to 0, while the sparsity of L21ELM is maintained at around 0.1. RRELM has better sparsity than L21ELM due to the use of L1 norm regularization, while L1 / L2ELM has the best sparsity among the five algorithms due to its approximate zero norm.

[0135] Table 2 shows the RMSE of the algorithms under different outlier ratios. By observing the RMSE under different outlier ratios of the two data sets, it is found that with the increase of the outlier ratio, the RMSE of each algorithm shows different degrees of change. The RMSE of ELM, WELM and L21ELM increased by 20 on the pollution data set, RELM increased by 50, and L1 / L2ELM only increased by 10. On the NO2 data set, the RMSE of ELM and WELM increased by 0.3, RELM and L21ELM increased by 0.2, and L1 / L2ELM only increased by 0.02. The RMSE of L1 / L2ELM is better than other algorithms in all cases, except that it is slightly inferior to ELM when there are no outliers, showing its superior robustness and stability in the presence of outliers.

[0136] In summary, RELM has good sparsity but insufficient robustness; L21ELM has good robustness but insufficient sparsity. In contrast, L1 / L2ELM performs well in both sparsity and robustness, especially in the presence of outliers.

[0137] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification.

[0138] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "one" or "an" does not exclude multiple situations. A single processor or other unit may implement several functions listed in a claim. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0139] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the scope of protection of the present invention.

Claims

1. A performance evaluation method for a sparse robust extreme learning machine based on L1 / L2, characterized in that: include: Acquire a preprocessed data set, wherein the preprocessed data set includes a training set and a test set, and the training set includes a preset number of abnormal data; Constructing an extreme learning machine model according to the ELM model, wherein the extreme learning machine model includes an L2 norm regularization term; Optimizing the extreme learning machine model to obtain an optimized extreme learning machine model, so as to convert the solution of the extreme learning machine model into an L1 / L2 analytical solution and a non-convex function solution; Training the optimized extreme learning machine model according to the preprocessed data set to obtain a trained extreme learning machine model; The trained extreme learning machine model is evaluated according to sparsity and root mean square error RMSE.

2. The performance evaluation method of the L1 / L2-based sparse robust extreme learning machine according to claim 1, characterized in that: After obtaining the preprocessed dataset, it also includes: Calculate the scaling factor based on the number of samples in the preprocessed dataset; Calculate a normalization coefficient and a normalization offset according to the scaling factor; The preprocessed data set is normalized according to the normalization coefficient and the normalization offset.

3. The performance evaluation method of the L1 / L2-based sparse robust extreme learning machine according to claim 1, characterized in that: The preprocessed data set includes a pollution data set and a NO2 data set, the ratio of the training set to the test set is 2:1, and the preset number is 10%-40% of the training set data.

4. The performance evaluation method of the L1 / L2-based sparse robust extreme learning machine according to claim 1, characterized in that: The extreme learning machine model is: stβ=u, Where C is the regularization parameter, β = [β1,β2,…,β L ] T is the weight connecting the hidden layer and the output layer, z i =y i -h(x i )u,z i is the training error between the predicted value and the true value of the i-th training sample, and σ is the upper bound of the exponential square loss function.

5. The performance evaluation method of the L1 / L2-based sparse robust extreme learning machine according to claim 4, characterized in that: Convert the solution of the extreme learning machine model into an L1 / L2 analytical solution and a non-convex function solution, including: Convert the solution of the extreme learning machine model into the solution of the following formula: Where C is the regularization parameter, β is the output weight vector, u is the auxiliary variable, v is the Lagrange multiplier, η is the step size, and η>0. k+1 represents the β,u obtained in the k+1th iteration k+1 represents u, v obtained in the k+1th iteration k+1 represents v obtained at the k+1th iteration.

6. The performance evaluation method of the L1 / L2-based sparse robust extreme learning machine according to claim 5, characterized in that: The analytical solution of L1 / L2 is solved by the proximal operator; The proximal operator form is The analytical solution of L1 / L2 is in, 7. The performance evaluation method of the L1 / L2-based sparse robust extreme learning machine according to claim 5, characterized in that: The solution of non-convex functions is solved by the convex difference algorithm; The solution of the non-convex function is Among them, I1 is the L×L identity matrix, I2 is the N×N identity matrix, H is the N×L hidden layer output matrix, H T is the transpose of the L×N hidden layer output matrix, Y is an N×1 vector, which is the actual value of the training sample, S is an N×1 vector, and η is the step size.

8. The performance evaluation method of the L1 / L2-based sparse robust extreme learning machine according to claim 5, characterized in that: Training the optimized extreme learning machine model according to the preprocessed data set includes: Input the training set into the established extreme learning machine model; Randomly generate weight w in the range of [-1,1] and bias b in the range of [0,1] to obtain output weight β; Initialize β, u, v, that is, β = β0, u = β0, v = β0, substitute into the loop and use the ADMM algorithm to alternately optimize β, u and v until the loop cutoff condition is met to obtain β, where the loop cutoff condition is:

9. The performance evaluation method of the L1 / L2-based sparse robust extreme learning machine according to claim 1, characterized in that:

10. The performance evaluation method of the L1 / L2-based sparse robust extreme learning machine according to claim 1, characterized in that: Root mean square error Among them, y i is the ith actual value, is the i-th predicted value.