Data mining method based on least square method and symbol change detection function evaluation

Through the combination method of least squares method and symbol change detection function evaluation, the problem of difficulty in fitting implicit function is solved, more accurate data mining is achieved, and obvious and implicit functions can be mined, which improves the fitting effect and readability.

CN120372436APending Publication Date: 2025-07-25SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510373638.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

It is difficult to effectively mine hidden function relationships in the prior art, especially in symbolic regression, where there is difficulty in evaluating the fitness of hidden functions, making it difficult for the model to accurately fit the data set.

Method used

The combination method of least squares method and symbol change detection function evaluation is used to mine multiple mathematical expressions from the data set, and complex implicit function expressions are generated through linear combinations, and the fitting effect is evaluated in combination with the fitness function.

Benefits of technology

It improves the accuracy and readability of implicit function fitting, can effectively avoid convergence to ordinary solutions, enhances the search ability in solving polynomial implicit function equations, adapts to different data distributions, and stably mines the obvious and implicit functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372436A_ABST
    Figure CN120372436A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data mining, in particular to a data mining method based on least square method and symbol change detection function evaluation, which comprises the following steps: S1, mining a plurality of mathematical expressions from a data set to be mined, a series of independent mathematical expressions are linearly combined by adopting a least square method to generate a more complex implicit function expression; and S2, evaluating the fitting effect of the generated implicit function expression on the data set by adopting a fitness function. According to the method, a learned mathematical expression can better fit a data set, and a data hiding inherent rule can be more accurately revealed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data mining, and more specifically, to a data mining method based on the least squares method and the evaluation of a sign change detection function. Background Art

[0002] With the continuous progress of hardware technologies such as multi-core processor architectures and GPUs, and machine learning, the capture and storage of data have grown rapidly in recent years, and data mining has become the research focus of more and more scholars. Massive amounts of data can be collected from various sources, such as personal electronic devices, the stock market, transportation networks, etc. Historical data can be mined to reveal internal laws or predict future trends.

[0003] Symbolic Regression (SR) is a machine learning technique that aims to automatically discover a mathematical model (such as an algebraic equation or function) from data to fit the data, and the mathematical model should describe the relationship of the data as accurately as possible. Different from traditional regression analysis methods (such as polynomial regression, LASSO regression), symbolic regression does not need to rely on a preset model form, and mathematical symbols and model structures are also part of the regression, so it can discover more flexible and more complex models. Compared with black-box regression methods such as neural networks, symbolic regression can provide relationships in the form of mathematical equations, and this white-box relationship is usually considered particularly useful because they can be meaningfully interpreted.

[0004] Many methods have been developed for symbolic regression, such as: physics-informed artificial intelligence Feynman2.0, enhanced reinforcement learning method, feature synthesis. Among many different implementations of symbolic regression, genetic programming is a traditional and mainstream method. Usually, symbolic regression refers to explicit function symbolic regression, that is, given a data set where x i ∈R n is the input, and y i ∈R is the output, then symbolic regression aims to find a function y = f(x) that can best fit the data set. However, the form of the implicit function f(x, y) = 0 has stronger expressive power and is usually used to describe various complex conic curves and surfaces, which is becoming increasingly important in geometric modeling, visualization, and computer graphics. Implicit functions are also often used to represent partial differential equations or explain physical phenomena.

[0005] Although implicit functions have more powerful expressive power than explicit functions, unfortunately, few symbolic regression methods are able to successfully capture implicit relationships, and many leading methods and benchmarks simply do not consider implicit function equations. One crucial reason is that in the mainstream genetic programming-based implementations of symbolic regression, the fitness evaluation of implicit functions faces difficulties. The earliest research on symbolic regression of implicit functions was by Michael Schmidt and Hod Lipson. In 2009, they analyzed the difficulties of symbolic regression of implicit functions, among which how to prevent the model from converging to a trivial solution such as f(x,y)=sin 2 (x)+cos 2 (x)-1 is a crucial difficulty. For this reason, they proposed a derivative-based fitness evaluation method. However, this method performs unsatisfactorily on some datasets. Subsequently, other scholars have also proposed new methods. For example, Roberts et al. proposed using the Kullback-Leibler divergence as the fitness calculation formula for implicit functions; Chen et al. proposed a comprehensive learning fitness evaluation method (CL-FEM) as the fitness evaluation method for implicit functions. However, the above methods all have their respective limitations.

[0006] Therefore, how to better reveal the hidden internal laws of data through mathematical expressions is an urgent problem for those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a data mining method based on least squares method and symbolic change detection function evaluation, so that the learned mathematical expression can better fit the dataset and more accurately reveal the hidden internal laws of the data.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] A data mining method based on least squares method and symbolic change detection function evaluation, comprising the following steps:

[0010] S1. Mine multiple mathematical expressions from the dataset to be mined, and use the least squares method to generate a more complex implicit function expression by linearly combining a series of independent mathematical expressions;

[0011] S2. Use a fitness function to evaluate the goodness of fit of the generated implicit function expression to the dataset.

[0012] Further, the dataset to be mined is physical modeling data, financial modeling data, traffic network data or signal processing data.

[0013] Further, S1 includes:

[0014] S11. Use a genetic programming algorithm to mine multiple sets of mathematical expression sets from the dataset to be mined, and each set of mathematical expressions is used as a sub-population;

[0015] S12. Randomly select a mathematical expression from each sub-population, denoted as f1, f2, …, f m ; Calculate a series of coefficients β0, β1, …, β m by the least squares method so that the following equation holds:

[0016] (β0 + 1) + β1f1 + β2f2 + … + β m f m = 1

[0017] S13. After obtaining a series of coefficients, obtain a more complex expression by linearly combining multiple implicit function expressions, denoted as:

[0018] f = β0 + β1f1 + β2f2 + … + β m f m

[0019] S14. Repeat S12 - S13 to obtain multiple implicit function expressions.

[0020] Further, in S2, for the dataset where x i ∈R n is the input, y i ∈R is the output, and the implicit function f(x, y) = 0, calculate the fitness function of f(x, y) = 0 for the dataset through a series of steps to evaluate the goodness of fit of the implicit function expression to the dataset.

[0021] Further, S2 includes:

[0022] S21. For a point (x', y') in the dataset, substitute x' into the implicit function expression f(x, y). Since x' is a constant vector, simplify f(x, y) to g(y);

[0023] S22. Generate a series of sampling points a1, a2, …, a m ;

[0024] S23. Solve the values of g(y) at a1, a2, …, a m to get g(a1), g(a2), …, g(a m );

[0025] S24. Traverse g(a1), g(a2), …, g(a m) value, solve the zero point of the function according to the following rules: If g(a i ) = 0, it means that a i is the zero point of the function; if g(a i ) × g(a i+1 ) < 0, it means that the zero point is in the interval (a i , a i+1 ). As long as the distance between the sampling points a i and a i+1 is small enough, it is considered that a i + (a i+1 - a i ) / 2 is the approximate zero point of the function; after traversing all the values, take the point closest to the supervised value y i as the solution of the function;

[0026] S25. For each point (x i , y i ) in the dataset, repeat the steps from S21 to S24 to obtain the solution of the function corresponding to the implicit function expression at each y i ;

[0027] S26. Use the root mean square error formula as the fitness function to calculate the value of the implicit function expression f(x, y) on this fitness function.

[0028] Further, in S26, the expression of the fitness function is:

[0029]

[0030] where N is the size of the dataset, y i is the supervised value corresponding to the i-th data point x i , o i is the solution of the function corresponding to the implicit function expression at each y i ; MSE represents the root mean square error. By comparing the difference between the value of the currently mined implicit function expression on the y-axis and the Euclidean distance of the supervised value, judge the goodness of fit of the implicit function expression to the dataset.

[0031] Further, the sampling points in S22 are obtained by the following method:

[0032] Uniformly generate sampling points at the same sampling interval near each supervised value y i ;

[0033] Continuously generate sampling points with different interval sizes on the number axis at a varying sampling interval near each supervised value y i .

[0034] Further, S2 further includes: performing a fitness evaluation on each implicit function expression generated by S1, and screening out the implicit function expression with the best fitting effect on the data set.

[0035] As can be seen from the above technical solutions, compared with the prior art, the present invention has the following beneficial effects:

[0036] 1. The present invention first mines multiple mathematical expressions from data sets such as physical modeling data, financial modeling data, or traffic network data, and then obtains the best linear coefficients through the least squares method, and combines the multiple mathematical expressions into a more complex implicit function expression, so that the learned mathematical expressions can better fit the data set. At the same time, not only can explicit functions be mined from the data set, but also implicit function expressions can be learned, and it will not converge to a trivial solution. Compared with the single-expression strategy, the present invention can enhance the search ability of the algorithm in solving polynomial implicit function equations, suppress redundant parts, and improve the accuracy and readability of the results.

[0037] 2. Most current data mining techniques are dedicated to mining explicit functions from data sets, while the present invention can not only mine explicit functions from data sets, but also mine implicit functions. For example, the present invention can mine implicit function equations such as lemniscate, conchoid, cissoid, and folium of Descartes. Compared with the existing methods for mining implicit functions, it also has certain advantages. For example, the present invention is not easily affected by different distributions of data in the data set compared with the existing derivative-based implicit function fitness evaluation method. In particular, when facing sparse data sets, the present invention generally maintains stable performance, while the performance of the derivative-based fitness evaluation method drops severely.

[0038] 3. Compared with the existing comprehensive learning implicit function fitness evaluation method, the present invention will not converge to a trivial solution that is approximately equal to zero but not zero (for example, f(x,y)=x 8 ×sin(y), where x belongs to (-1,1)). Compared with the existing method for solving implicit function symbolic regression problems using mixed integer linear programming (MILP), the present invention has the advantage of not requiring artificial prior setting of outliers. If the outliers are set improperly in MILP, it is easy to obtain incorrect mathematical expressions. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0040] Figure 1Flowchart of the data mining method based on least squares method and sign change detection function evaluation provided by the present invention;

[0041] Figure 2 Flowchart of S2 provided by the present invention. Detailed implementation manners

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] As Figure 1 shown, the embodiments of the present invention disclose a data mining method based on least squares method and sign change detection function evaluation, including the following steps:

[0044] S1. Mine multiple mathematical expressions from the dataset to be mined, and use the least squares method to generate a more complex implicit function expression by linearly combining a series of independent mathematical expressions;

[0045] S2. Use a fitness function to evaluate the goodness of fit of the generated implicit function expression to the dataset.

[0046] The above steps will be further described below.

[0047] S1. Mine multiple mathematical expressions from the dataset to be mined, and use the least squares method to generate a more complex implicit function expression by linearly combining a series of independent mathematical expressions.

[0048] Among them, the dataset to be mined is physical modeling data, financial modeling data, traffic network data or signal processing data.

[0049] For example, in the double pendulum model in a physical experiment system, a large amount of data of pendulum angles θ1 and θ2 is collected in real time through sensors, and the lengths l1 and l2 of the pendulum arms and the masses m1 and m2 of the pendulum bobs are given, that is, the given dataset where Denote the swing angle data obtained during the \(i\)-th measurement of the sensor. Then, using the present invention, the Lagrangian equation can be mined without any prior knowledge of physics, kinematics, etc. This method combines multiple mined sub-equations using the least squares method to form a more complex equation to improve the fitting effect on the data set. At the same time, different from the existing fitting effect evaluation methods, the symbol change detection function evaluation of the present invention can be used not only to evaluate explicit function equations but also to evaluate implicit function equations, that is, implicit functions can be mined from the data set.

[0050] Specifically, S1 includes the following steps:

[0051] S11. Use the genetic programming algorithm to mine multiple sets of mathematical expression sets for characterizing data relationships from the data set to be mined, and each set of mathematical expressions is used as a sub-population;

[0052] S12. Randomly select a mathematical expression from each sub-population, denoted as \(f_1,f_2,\cdots,f\) m ; Calculate a series of coefficients \(\beta_0,\beta_1,\cdots,\beta\) m by the least squares method such that the following equation holds:

[0053] (\(\beta_0 + 1\))+\(\beta_1f_1+\beta_2f_2+\cdots+\beta\) m \(f\) m = 1

[0054] S13. After obtaining a series of coefficients, obtain a more complex expression by linearly combining multiple implicit function expressions, denoted as:

[0055] \(f=\beta_0+\beta_1f_1+\beta_2f_2+\cdots+\beta\) m \(f\) m

[0056] S14. Repeat S12 - S13 to obtain multiple implicit function expressions.

[0057] For the convenience of description, in this embodiment, three sub-populations are adopted, and each sub-population represents a set of mathematical expressions mined by the genetic algorithm from the data set. S12 and S13 can be expressed as:

[0058] S12. Randomly select an expression from each sub-population, denoted as \(f_1,f_2,f_3\) respectively, and calculate a series of coefficients \(\beta_0,\beta_1,\beta_2,\beta_3\) by the least squares method such that the following equation holds:

[0059] (\(\beta_0 + 1\))+\(\beta_1f_1+\beta_2f_2+\beta_3f_3 = 1\)

[0060] S13. After obtaining a series of coefficients, generate a more complex expression by linearly combining multiple implicit function expressions, that is:

[0061] f(x, y) = β0 + β1f1 + β2f2 + β3f3

[0062] Repeat the operations in S12 - S13 until a total of 50 implicit function expressions are generated through linear combination.

[0063] S2. Use the fitness function to evaluate the goodness of fit of the generated implicit function expressions to the data set.

[0064] For the data set where x i ∈R n is the input, y i ∈R is the output, and the implicit function f(x, y) = 0, calculate the fitness function of f(x, y) = 0 for the data set through a series of steps to evaluate the goodness of fit of the implicit function expressions to the data set. As Figure 2 shown, S2 specifically includes:

[0065] S21. For a point (x', y') in the data set, for example, if the data set can be the data collected by sensors in the double - pendulum system mentioned above, then a data point is Substitute x' into the implicit function expression f(x, y). Since x' is a constant vector, simplify f(x, y) to an equation related only to the variable y, that is, g(y);

[0066] S22. Generate a series of sampling points a1, a2, …, a m on the y - axis; The sampling points are obtained by the following method:

[0067] 1) Uniformly generate sampling points at the same sampling interval near each supervised value y i ;

[0068] 2) Continuously generate sampling points with different interval sizes on the number line at a varying sampling interval near each supervised value y i ;

[0069] S23. Solve the values of g(y) at a1, a2, …, a m to get g(a1), g(a2), …, g(a m );

[0070] S24. Traverse the values of g(a1), g(a2), …, g(a m ) and solve for the zero points of the function according to the following rules: If g(a i ) = 0, it means that a i is the zero point of the function; If g(a i ) × g(a i+1) < 0 indicates that the zero point is within the interval (a i , a i+1 ). As long as the distance between the sampling points a i and a i+1 is small enough, then a i + (a i+1 - a i ) / 2 is considered the approximate zero point of the function; after traversing all the values, the point closest to the supervision value y i is taken as the solution of the function;

[0071] S25. For each point (x i , y i ) in the dataset, repeat the steps from S21 to S24 to obtain the solutions of the functions corresponding to the implicit function expression at each y i ;

[0072] S26. Use the root mean square error formula as the fitness function to calculate the value of the implicit function expression f(x, y) on this fitness function. The expression of the fitness function is:

[0073]

[0074] where N is the size of the dataset, y i is the supervision value corresponding to the i-th data point x i , and o i is the solution of the function corresponding to the implicit function expression at each y i ; MSE represents the root mean square error. By comparing the difference between the value of the currently mined implicit function expression on the y-axis and the Euclidean distance of the supervision value, the goodness of fit of the implicit function expression to the dataset is judged.

[0075] Fitness characterizes the goodness of fit of a mathematical expression to a dataset. For example, if the internal law in the dataset is the unit circle equation, then mining x 2 + y 2 = 1.5 will obtain a smaller fitness value (indicating a good fit), and mining x 2 + y 2 = 4 will obtain a larger fitness value (indicating a poor fit).

[0076] After that, fitness evaluation is performed on each implicit function expression generated in S1 respectively, and genetic selection operations are performed on the population using roulette or other mechanisms until the evolution terminates. The steps of evolution include selection, crossover, and mutation. Finally, the implicit function expression with the best fit to the dataset is selected.

[0077] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For related parts, reference can be made to the description in the method section.

[0078] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data mining method based on the evaluation of the least squares method and the sign change detection function, characterized in that It includes the following steps: S1. Mine multiple mathematical expressions from the dataset to be mined, and use the least squares method to generate a more complex implicit function expression through linear combination of a series of independent mathematical expressions; S2. Use a fitness function to evaluate the goodness of fit of the generated implicit function expression to the dataset.

2. The data mining method based on the evaluation of the least squares method and the sign change detection function according to claim 1, characterized in that The dataset to be mined is physical modeling data, financial modeling data, traffic network data or signal processing data.

3. The data mining method based on the evaluation of the least squares method and the sign change detection function according to claim 1, characterized in that S1 includes: S11. Use the genetic programming algorithm to mine multiple sets of mathematical expression sets from the dataset to be mined, and each set of mathematical expressions is used as a sub-population; S12. Randomly select a mathematical expression from each sub-population, denoted as f1, f2, …, f m ; Calculate a series of coefficients β0, β1, …, β m by the least squares method, such that the following equation holds: (β0 + 1) + β1f1 + β2f2 + … + β m f m = 1 S13. After obtaining a series of coefficients, obtain a more complex expression through linear combination of multiple implicit function expressions, expressed as: f = β0 + β1f1 + β2f2 + … + β m f m S14. Repeat S12 - S13 to obtain multiple implicit function expressions.

4. The data mining method based on the least squares method and the evaluation of the sign change detection function according to claim 1, wherein In S2, for the data set where x i ∈R n is the input, y i ∈R is the output, and the implicit function f(x, y) = 0, which is calculated through a series of steps The fitness function of f(x, y) = 0 for the dataset to evaluate the goodness of fit of the implicit function expression to the dataset.

5. The data mining method based on the evaluation of the least squares method and the sign change detection function according to claim 1, wherein S2 It includes: S21. For a point (x', y') in the dataset, substitute x' into the implicit function expression f(x, y). Since x' is a constant vector, simplify f(x, y) to g(y); S22. Generate a series of sampling points a1, a2, …, a on the y-axis m ; S23. Solve for the values of g(y) at a1, a2, …, a m to obtain g(a1), g(a2), …, g(a m ); S24. Traverse the values of g(a1), g(a2), …, g(a m ), and solve for the function zero according to the following rules: If g(a i ) = 0, it indicates that a i is the zero of the function; if g(a i ) × g(a i+1 ) < 0, it indicates that the zero is in the interval (a i , a i+1 ). As long as the distance between the sampling points a i and a i+1 is small enough, it is considered that a i + (a i+1 - a i ) / 2 is the approximate zero of the function; after traversing all the values, take the point closest to the supervised value y i as the solution of the function; S25. For each point (x i , y i ) in the dataset, repeat the steps from S21 to S24 to obtain the solutions of the functions corresponding to the implicit function expression at each y i ; S26. Use the root mean square error formula as the fitness function to calculate the value of the implicit function expression f(x, y) on this fitness function.

6. The data mining method based on the least squares method and the symbol change detection function evaluation according to claim 5, wherein In S26, the expression of the fitness function is: where N is the size of the data set, and y i is the supervised value corresponding to the i-th data point x i , and o i is the solution of the function corresponding to the implicit function expression at each y i ; MSE represents the root mean square error. By comparing the difference between the value of the currently mined implicit function expression on the y-axis and the Euclidean distance of the supervised value, the goodness of fit of the implicit function expression to the data set is judged.

7. The data mining method based on the least squares method and the symbol change detection function evaluation according to claim 5, characterized in that, The sampling points in S22 are obtained by the following method: By generating sampling points uniformly at the same sampling interval near each supervised value y i ; By generating, at varying sampling intervals near each supervised value y i successively on the number line sampling points of different interval sizes.

8. The data mining method based on the least squares method and the symbol change detection function evaluation according to claim 3, characterized in that S2 also includes: separately performing fitness evaluation on each implicit function expression generated by S1, and screening out the implicit function expression with the best fit to the dataset.