Semi-supervised smooth support vector regression method based on Bessel difference function
By using the Bessel difference function approximates the loss function in the semi-supervised support vector regression method, the existing methods have solved the problem of low regression accuracy and low algorithm efficiency in large-scale data processing, and achieved more efficient regression performance.
Patent Information
- Application Number
- CN202510276751.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
The existing semi-supervised support vector regression method has low regression accuracy and low algorithm efficiency in large-scale data processing, and has little research on optimization algorithms, which is difficult to meet the processing needs of the big data era.
A class of semi-supervised smooth support vector regression method based on Bezier's difference function is adopted to approximate the loss function through Bezier's interpolation function, replacing the non-smooth part in the original model, thereby improving the regression accuracy and algorithm efficiency.
The regression accuracy and algorithm efficiency of semi-supervised support vector regression are improved, and large-scale semi-supervised data can be processed more effectively.
Smart Images

Figure CN120218288A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of artificial intelligence machine learning and data mining, and relates to a semi-supervised efficient regression method for large-scale data, in particular to a class of semi-supervised smooth support vector regression method based on Bessel interpolation function. Background Art
[0002] In the era of artificial intelligence, "big data" comes from various industrial fields and life application scenarios, and has the characteristics of huge data volume, diversity and complex structure. Therefore, it becomes particularly important to predict the trend of future data. Based on its strong statistical technical support and solid mathematical theory, support vector regression (SVR) shows broad application prospects in machine learning of large-scale data. SVR was initially proposed to solve the non-linear regression problem with the maximum margin. It predicts continuous results by finding a function such that the maximum deviation between the function and the actual observed values does not exceed ε, thus extending the support vector machine method to the regression field. The goal is to ensure that the function is as flat as possible, achieving a balance between model complexity and prediction error. Therefore, many excellent improved SVR models have been proposed to cope with different scenarios.
[0003] However, traditional fully supervised SVR highly depends on labeled data sets, which limits its application in industrial and daily life scenarios. In the real world, many soft sensors and hardware sensors generate a large number of data sets, and a large part of these data sets have no corresponding labels. To solve this problem, some scholars have proposed excellent semi-supervised learning regression methods (SSR) in the past few decades. SSR lies between supervised learning and unsupervised learning, and trains a better regression model by using labeled data and unlabeled data. The premise is that unlabeled data can provide additional structural or distribution information, which helps the better generalization of the model, especially in the case of scarce labeled data. The achievements of semi-supervised regression have been widely applied in various fields, such as digital image processing, rock mechanics, semiconductor manufacturing, wireless sensors, and traffic prediction.
[0004] The research work of SSR can be generally divided into three categories: non-parametric semi-supervised regression methods, parametric semi-supervised regression methods - hybrid methods, and semi-supervised ordinal regression methods. First of all, non-parametric semi-supervised regression methods include five branches: collaborative regression, kernel regression, graph regularization, local linear SSR, and Gaussian process regression.
[0005] 1. Collaborative regression is a semi-supervised method proposed by Blum and Mitchell in 1998. Subsequently, two improved models emerged: multi-view collaborative regression and single-view collaborative regression.
[0006] 2. Kernel regression is one of the most commonly used non-parametric methods in regression analysis. A kernel is a positive definite symmetric function that satisfies the conditions of Mercer's theorem and corresponds to the inner product in the feature space. Semi-supervised kernel regression is divided into kernel regression, support vector regression, and output kernel regression.
[0007] 3. Graph-based learning methods combine many SSL studies. Graph regularization describes the intrinsic geometric structure of the data and improves the prediction performance of the algorithm by constructing a nearest neighbor graph and using it as a regularization term in the objective function. SSR based on graph regularization can be divided into graph Laplacian regularization, Hessian regularization, parallel field regularization, and spectral regression.
[0008] 4. Local linear SSR attempts to solve the common local constant estimation problem in regression problems. Rwebangira and Lafferty published a famous study that elaborated on the efficiency of local linear SSR and proposed a method called local linear semi-supervised regression (LLSR).
[0009] 5. Gaussian process regression combines multi-task regression. Zhang et al. proposed a semi-supervised multi-task regression (SSMTR) algorithm, which assumes a common prior among all kernel parameters. Followed by the parametric semi-supervised regression method - the hybrid method. Although non-parametric methods have been studied for many years, parametric methods are less applied because it is difficult to model the relationship between independent variables to estimate the output variable. The existing parametric methods mainly include: semi-supervised linear regression and semi-supervised hybrid methods.
[0010] Finally, there is semi-supervised ordinal regression. Ordinal regression refers to a multi-class regression problem with order constraints. In practical applications, when there is a measurement distance but no clear rank among the values of finite discrete variables, an ordinal regression model can be used. Ordinal regression can be divided into finite variables and discrete variables.
[0011] In the development of support vector regression, Lee et al. proposed a smoothed support vector machine (SSVR) for insensitive regression in 2005. This model uses an insensitive loss function and a Sigmoid smoothing function to improve the regression accuracy. After that, Chapelle and Zien replaced the non-smooth term in the objective function. Xiong et al., Liu et al., and Yuan et al. proposed polynomial functions and piecewise spline polynomials respectively. However, the polynomial function is too complex and computationally intensive. Yang et al. proposed a square root function.
[0012] Currently, the existing technologies have the deficiencies of low accuracy and low algorithm efficiency for semi-supervised data regression, and there is little research on the optimization algorithm of semi-supervised support vector regression. However, the real social demand for processing massive semi-supervised data in the big data era is increasing continuously. Summary of the Invention
[0013] To overcome the deficiencies of the above-mentioned existing technologies, the object of the present invention is to provide a semi-supervised smooth support vector regression method based on Bessel interpolation functions, by proposing a new type of semi-supervised smooth support vector machine based on Bessel interpolation functions, achieving a good approximation of the loss function, replacing the non-smooth part in the objective function of the semi-supervised support vector regression model, thereby improving the regression accuracy of the semi-supervised support vector machine and the algorithm efficiency.
[0014] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0015] A semi-supervised smooth support vector regression method based on Bessel interpolation functions, characterized by comprising the following steps:
[0016] Step 1: Establish a support vector regression model based on the insensitive loss function;
[0017] Step 2: Construct a semi-supervised support vector regression model based on the ε-insensitive loss function;
[0018] Step 3: Establish a semi-supervised support vector regression machine based on a class of Bessel smooth functions;
[0019] Step 4: For the Bessel smooth semi-supervised support vector regression in the non-linear case, use the kernel function to map the original low-dimensional data to high-dimensional data;
[0020] Step 5: A fast solution method for the new type of Bessel smooth semi-supervised support vector regression model.
[0021] Furthermore, for the support vector regression model based on the insensitive loss function established in Step 1 ( ), for the given data set find the optimal regression hyperplane;
[0022] (1);
[0023] It is to find the minimum regression risk function:
[0024] (2);
[0025] Define the ε-insensitive loss function as:
[0026] (3);
[0027] Then the square of the ε-insensitive loss function can be expressed as:
[0028] (4);
[0029] Using the above definitions, the minimum regression risk function (2) can be transformed into an unconstrained support vector regression model (5) with an ε-insensitive loss function:
[0030] (5).
[0031] Furthermore, for the construction of the semi-supervised support vector regression model ( ) based on the ε-insensitive loss function described in Step 2, for semi-supervised learning, the fully supervised model needs to be transformed to handle some unlabeled corresponding data. Let the labeled training data set be denoted as , and the unlabeled data set be denoted as . is the label corresponding to the unlabeled data predicted by optimizing the model. The semi-supervised support vector regression model based on the ε-insensitive loss function can be expressed as:
[0032] (6);
[0033] In formula (6), is the slack variable for the labeled data, is the slack variable for the unlabeled data. At the same time, formula (6) can be transformed into an unconstrained model:
[0034] (7);
[0035] Let , then formula (7) can be transformed into an unconstrained semi-supervised support vector regression model (8) based on the ε-insensitive loss function:
[0036] (8);
[0038] Furthermore, for the establishment of a semi-supervised support vector regression machine based on the Bessel smooth function ( ) described in Step 3, according to the non-differentiability of the last term in formula (8) in Step 2, which results in a large amount of computational effort during the model solution process. Based on the good performance of the Bessel interpolation function in approximating the objective function, the Bessel function is used to replace the above non-differentiable term, obtaining a semi-supervised support vector regression machine based on the Bessel smooth function ( ), as shown in formula (9):
[0039] (9);
[0040] Further, for the Bessel smooth semi-supervised support vector regression for the non-linear case in Step 4, the kernel function is used to map the original low-dimensional data to high-dimensional data, and is used to replace in the objective function (9).
[0041] (10);
[0042] Further, for the fast solution of the new type of Bessel smooth semi-supervised support vector regression model in Step 5, for any convex programming in the interval satisfies then the Newton-Armijo fast solution algorithm can be used. Since the Bessel smooth semi-supervised support vector regression model belongs to a strictly convex programming, thus, the Newton-Armijo algorithm becomes the preferred method for solving the new model.
[0043] Further, the algorithm solving steps for the Newton-Armijo algorithm to solve the new model are as follows:
[0044] (1) Train BS 4 the labeled data in the VR model , and sort the unlabeled data according to the magnitude of the value;
[0045] (2) Set the parameters , initialize and ;
[0046] (3) Calculate the gradients and ;
[0047] (4) If or is satisfied, then the iteration stops and the current is accepted. Otherwise, calculate the Newton direction according to formula (11);
[0048] (11);
[0049] (5) Calculate the Armijo step size: Select the step size , such that , and let ;
[0050] (6) Go back to step (3) for the next loop;
[0051] (7) Output weights w and b;
[0052] In this process, only formula (11) needs to be solved instead of the entire second-order programming, and the global optimal solution of the algorithm can be obtained.
[0053] The beneficial effects of the present invention are:
[0054] In the specific implementation process, the present invention compares the semi-supervised support vector regression machine based on the insensitive loss function S3VR, the smooth semi-supervised support regression model based on the Sigmoid function The performance of S4VR on different data. These data include artificial data: two-dimensional Mexican hat data, Friedman function, multivariate function, plane function, fourth-order polynomial function, and 12 UCI classic regression datasets (Banknote, ENB, Bank_8fm, Bank_8fh, Bank_32nm, Energy, Parkinson_sound, Qsar_fish, Parkinsons, Naval_Plants, Electrical_Grid, DrivFace) and large-scale NDC data. Description of the Drawings
[0055] Figure 1 is the flowchart of the present invention.
[0056] Figure 2 is the flowchart of the present invention.
[0057] Figure 3 is the graph of different function data.
[0058] Figure 4 is the comparison of four regression results tested on the non-linear 2D Mexican hat data.
[0059] Figure 5 is the trend of the correlation coefficient of 12 UCI datasets with different labeling ratios.
[0060] Figure 6 is the trend of the correlation coefficient change of three regression methods.
[0061] Figure 7 is the trend of the calculation time change of three regression methods. Detailed Embodiment
[0062] The present invention will be further described below in conjunction with the drawings and embodiments.
[0063] As Figure 1 、 2, as shown in FIGS. 3, a semi-supervised smooth support vector regression method based on Bessel interpolation function includes the following steps:
[0064] Step 1: Establish a support vector regression model based on an insensitive loss function;
[0065] Step 2: Construct a semi-supervised support vector regression model based on the ε-insensitive loss function;
[0066] Step 3: Establish a semi-supervised support vector regression machine based on a Bessel smooth function;
[0067] Step 4: For the Bessel smooth semi-supervised support vector regression in the non-linear case, use a kernel function to map the original low-dimensional data to high-dimensional data;
[0068] Step 5: A fast solution method for a new type of Bessel smooth semi-supervised support vector regression model.
[0069] Example:
[0070] The experimental effect is specifically implemented on the Matlab R2011 simulation platform, and the steps are as follows:
[0071] Step 1: Experimental parameter setting: For each experiment, each experiment is conducted 10 times. The settings of the experiment have the following rules. (1) Regarding the labeled ratio. In each trial, the labeled data and the unlabeled data are randomly divided according to the following rules: (i) For BS 4 VR test, stricter conditions should be set. (ii) The ratio of the labeled data should be prevented from being very high. Therefore, we set the labeled ratio in a large range of 1%, 5%, 10%, 30%, 50% to 70%.
[0072] (2) Optimal parameter setting. The parameters have a significant impact on the learning accuracy and calculation time. The present invention preprocesses the hyperparameters and applies a parameter adaptive optimization method based on the genetic algorithm.
[0073] (3) For the non-linear case, the setting range of the Gaussian kernel function parameter is {2 -2 , 2 -1 , 1, 2 1 , 2 2}.
[0074] Step 2: Comparison of artificial data simulation experiments. First, the first group of artificial data is two-dimensional Mexican hat data, and the regression results of four algorithms are qualitatively compared through this data. As Figure 4 shown. Figure 4 It shows that all four algorithms can achieve the regression and prediction of two-dimensional Mexican hat data.
[0075] Secondly, the second set of data is the regression simulation of four groups of mathematical polynomial functions. The four polynomials are shown in Table 1, and the labeled data ratio ranges from 1% to 70%. Two comparison factors are introduced in this experiment: (1) Mean correlation error , calculated as , where , are the actual output and the predicted output, and , are the average values of the actual output and the predicted output respectively. (2) Correlation coefficient , calculated according to the rule
[0076] Table 1. Descriptions of four artificial regression functions
[0077] The specific simulation comparison results are shown in Table 2;
[0078] Table 2. Four mathematical polynomial regression tests are performed on three models; (the bold part is the best result)
[0079]
[0080] As can be seen from Table 2, the differences in the mean relative errors of the artificial data are not significant, and the labeled ratio varies from 1% to 70%. The comparison factor of the correlation coefficient increases significantly as the labeled ratio increases. This indicates that the latter can well test the performance of different models. From the "Average" row, it can be seen that the proposed BS 4 VR method performs the best.
[0081] Step 3: Comparison of simulation experiments on real-world UCI datasets;
[0082] The real data from industries and daily life is highly complex, and the regression model should be tested first. In this experiment, we used 12 UCI datasets to train and test the proposed algorithm. The 12 training data can be seen in Table 3.
[0083] Table 3. Test results of three regression methods on 12 UCI datasets. (The bold part is the best result);
[0084]
[0085]
[0086] Table 3 shows the regression results of 12 UCI datasets. It can be found that: (1) The larger the sample size and the number of attributes of the data, the smaller the average relative error and the higher the correlation coefficient. This indicates that these algorithms have better regression performance. This trend is particularly suitable for the Energy (19735*27), Naval_Plants (11934*17), and Electrical_Grid (10000*13) datasets. (2) In most cases, BS 4 VR performs better than S 4 VR and S 3 VR. To visualize the change trend of the correlation coefficient under different labeling ratios, we plotted Figure 5 .
[0087] In Figure 5 , the red line representing the BS 4 VR method is always higher than the lines of other colors, such as in data (1), (3), (4), (5), (7), (9), (11), (12). This means that from the perspective of the correlation coefficient, BS 4 VR achieves the best regression accuracy.
[0088] Table 4 more clearly lists the average rankings of these four regression algorithms. The last row of Table 4 shows that BS 4 VR ranks first, while the other regression algorithms are ranked behind.
[0089] Table 4. Average rankings of BS 4 VR, S 4 VR, and S 3 VR: Average relative error and correlation coefficient values:
[0090] To more precisely quantify whether there are significant differences among these three regression methods, we applied the Friedman test. The Friedman statistic is calculated based on the following two parameters: one represents the number of datasets , and the other represents the number of algorithms . Since we have two comparison factors, the Friedman statistic can be calculated in two ways.
[0091] (1) For the 12 UCI datasets, from the perspective of the average relative error , the Friedman statistic can be calculated according to the following process. For the experiments on the above UCI datasets, and under the original hypothesis that all algorithms are equal, the Friedman statistic can be calculated as:
[0092]
[0093] For the three algorithms and twelve datasets, the degrees of freedom of the distribution are . The significance level when the critical value or threshold is 3.443. Obviously , this means that from the perspective of comparing the relative errors, there are significant differences among these three regression methods.
[0094] (2) From the perspective of the correlation coefficient R,
[0095]
[0096] the degrees of freedom are . The significance level when the critical value or threshold is 3.443. Obviously , this indicates that from the perspective of comparing the correlation coefficients, there are also significant differences among the three regression methods.
[0097] To further evaluate which model is the best, we use the Nemenyi test for further verification. When , for this experiment , this means that when the ranking difference between two models is greater than CD, one model is better with 90% confidence.
[0098] Obviously, the average relative error ranking difference between BS 4 VR and S 3 VR (2.493 - 1.396 = 1.097) is greater than the critical difference CD 0.935. This shows that the regression performance of BS 4 VR is significantly better than that of S 3 VR. Similarly, from the perspective of comparing the correlation coefficients of factors, the performance of BS 4 VR is also significantly better than that of S 4 VR (2.389 - 1.222 = 1.167 > 0.935) and S 3 VR (2.306 - 1.222 = 1.084 > 0.935).
[0099] Step 4: Comparison of the non - linear Gaussian kernel regression models based on large - scale NDC data.
[0100] In Step 4, we used the experiment based on the NDC dataset to determine BS 4The performance of VR in terms of regression accuracy and calculation time. The large NDC dataset is designed to measure the robustness of the proposed method. In Table 5, we selected the 4th-order BS 4 VR-I, 5th-order BS 4 VR-II, and 6th-order BS 4 VR-III models.
[0101] Table 5. BS 4 Comparison of regression accuracy and calculation time of different models within the VR series;
[0102] As can be seen from Table 5: (1) From the perspective of the correlation coefficient, there is not much difference in performance among these three methods. This situation is different from our intuition that the higher the order, the higher the regression accuracy. (2) From the perspective of calculation time, there are obvious differences. The higher the order of the Bessel function, the longer the calculation time for regression. When the sample size of the training data is greater than 10,000, the time consumption will increase sharply. This reminds us that when dealing with regression problems of large-scale data, we should consider the comprehensive performance. To intuitively display the change trends of the correlation coefficients and calculation times of the three methods, we designed Figure 6 and Figure 7 .
[0103] Figure 6 shows the change trend of the correlation coefficient of each method. The red line represents BS 4 VR-I, the blue line represents BS 4 VR-II, and the black line represents BS 4 VR-III. We can find that these three lines almost coincide in most cases. This indicates that the differences between them are very small. Figure 7 shows that the lines of BS 4 VR-I and BS 4 VR-II are very close, while the black line of BS 4 VR-III is always higher than the other two lines. Sometimes, the calculation time of BS 4 VR-III is quite high. This indicates that there are a large number of iterative steps. To further quantitatively compare the performance of these three models, Table 6 lists the average rankings of each method.
[0104] Table 6. BS 4 VR-I, BS 4 VR-II, and BS 4 VR-III's average rankings regarding the correlation coefficient and calculation time:
[0105]
[0106] The Friedman analysis can also be carried out from two aspects: (1) From the perspective of the correlation coefficient,
[0107]
[0108] For the three algorithms and eight datasets, the degrees of freedom of the distribution are . When the significance level is , the critical value or threshold is 3.739. Obviously , this indicates that there are significant differences among the above methods.
[0109] (2) Distinguish the significant differences of different algorithms from the perspective of computing time.
[0110]
[0111] Obviously , it can be seen that there are also significant differences among the above three algorithms in terms of computing time.
[0112] Furthermore, we use the Nemenyi test to distinguish which algorithm has the best performance. For the NDC experiment, when the significance level is , this means that if the difference in the average rankings of two models exceeds CD, the confidence that one of these two models is better than the other in performance is 90%. In terms of the correlation coefficient, BS 4 VR-II and BS 4 the average ranking difference of VR-III is 2.667 - 1.396 = 1.271 > CD. This shows that BS 4 VR-II is better than the latter in terms of the regression prediction accuracy. On the other hand, in the comparison of computing time, since 2.771 - 1.127 = 1.644 > CD, therefore, the model BS 4 VR-I has better performance than BS 4 VR-III in terms of computing time.
Claims
1. A semi-supervised smooth support vector regression method based on Bessel difference function, characterized in that: The following steps are involved: Step 1: Establish a support vector regression model based on an insensitive loss function; Step 2: Construct a semi-supervised support vector regression model based on the ε-insensitive loss function; Step 3: Establish a semi-supervised support vector regression machine based on Bessel smooth function; Step 4: Bessel smoothed semi-supervised support vector regression for nonlinear situations uses a kernel function to map the original low-dimensional data to high-dimensional data; Step 5: Fast solution method for the new type of Bessel smooth semi-supervised support vector regression model.
2. A semi-supervised smooth support vector regression method based on Bessel difference function according to claim 1, characterized in that: The support vector regression model based on the insensitive loss function described in step 1 is established ( ), for a given data set Find the most regressive hyperplane; (1); is to find the minimum regression risk function: (2); The ε-insensitive loss function is defined as: (3); Then the square of the ε-insensitive loss function It can be expressed as: (4); Using the above definition, the minimum regression risk function (2) can be transformed into an unconstrained support vector regression model with an ε-insensitive loss function (5): (5)。 3. A semi-supervised smooth support vector regression method based on Bessel difference function according to claim 1, characterized in that: Step 2 describes the construction of a semi-supervised support vector regression model based on the ε-insensitive loss function ( ), for semi-supervised learning, fully supervised model Transformation is required to process some unlabeled corresponding data, and there is a labeled training data set express , unlabeled dataset express . is the label corresponding to the unlabeled data predicted by the optimization model. The semi-supervised support vector regression model based on the ε-insensitive loss function can be expressed as: (6); In formula (6), is a slack variable with labeled data, is the slack variable for unlabeled data. At the same time, formula (6) can be converted into an unconstrained model: (7); set up , then formula (7) can be transformed into an unconstrained semi-supervised support vector regression model based on the ε-insensitive loss function (8): (8)。 4. A semi-supervised smooth support vector regression method based on Bessel difference function according to claim 1, characterized in that: Step 3 describes the establishment of a class of semi-supervised support vector regression machines based on Bessel smooth functions ( ), according to the last term of formula (8) in step 2, which is non-differentiable, a large amount of calculation is generated in the process of solving the model. Based on the good performance of Bessel difference function in approximating the objective function, the Bessel function is used Replacing the above non-differentiable terms, we obtain a class of semi-supervised support vector regression machines based on Bessel smooth functions ( ), as shown in formula (9): (9)。 5. A semi-supervised smooth support vector regression method based on Bessel difference function according to claim 1, characterized in that: The Bessel smoothed semi-supervised support vector regression for nonlinear situations described in step 4 uses a kernel function to map the original low-dimensional data to high-dimensional data. Substituting the objective function (9) Thus, the Bessel smoothed semi-supervised support vector regression model for nonlinear data is obtained, as shown in (10); (10)。 6. A semi-supervised smooth support vector regression method based on Bessel difference function according to claim 1, characterized in that: The fast solution for the new type of Bessel smooth semi-supervised support vector regression model described in step 5 is for any convex programming In the interval satisfy The Newton-Armijo fast solution algorithm can be used, and the Bessel smooth semi-supervised support vector regression model belongs to strict convex programming. Therefore, the Newton-Armijo algorithm becomes the preferred method for solving the new model.
7. A semi-supervised smooth support vector regression method based on Bessel difference function according to claim 6, characterized in that: The Newton-Armijo algorithm becomes the algorithm for solving the new model. The solution steps are as follows: (1) Training BS 4 Labeled Data in VR Models ,according to The magnitude of the value is important for unlabeled data Sort by (2) Setting parameters ,initialization and ; (3) Calculate the gradient and ; (4) If satisfied or The iteration stops and the current Otherwise, calculate the Newton direction according to formula (11); (11); (5) Calculate the Armijo step size: Select the step size , so that , and let ; (6) Return to step (3) for the next cycle; (7) Output weight w,b; In this process, only formula (11) needs to be solved instead of the entire second-order planning, and the global optimal solution of the algorithm can be obtained.