A Symbolic Regression Method Based on Active Learning Strategy and Library Space Optimization

Through active learning strategies and library space optimization methods, the problem of search space expansion in semantic genetic planning is solved, efficient symbol regression is achieved, training time is reduced and accuracy is maintained.

CN117290811BActive Publication Date: 2025-07-04UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311181953.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-13
Publication Date
2025-07-04
Estimated Expiration
2043-09-13

AI Technical Summary

Technical Problem

With the increase in the scale of the problem, the search space of traditional semantic genetic planning methods has expanded dramatically, resulting in a significant reduction in search efficiency and it is difficult to obtain correct results in a limited time.

Method used

Using an active learning strategy and library space optimization method, data is filtered through improved active learning strategies, combined with DBSCAN clustering algorithm to remove noise, K-center clustering is used to reduce the library space, and semantic vectors are optimized through cosine distance and linear scaling to reduce semantic space and library search space.

Benefits of technology

The accuracy of semantic genetic planning is maintained, while significantly reducing the running time and improving the problem solving speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117290811B_ABST
    Figure CN117290811B_ABST
Patent Text Reader

Abstract

The object of the present invention is to provide a symbolic regression method based on an active learning strategy and library space optimization, belonging to the technical field of data processing. The method includes: first, selecting data based on an improved active learning strategy to effectively reduce the semantic vectors of each subtree in semantic genetic programming; then initializing a library space composed of random subtrees, and performing K-means clustering on the subtrees in the library. On the basis of clustering, by calculating the entropy value of each semantic vector, the search space is further reduced; and in the library search stage, cosine distance is combined with subtree linear scaling to match the optimal subtree. The present invention reduces the dimension of the semantic vectors and shrinks the library space, while ensuring a high accuracy rate in semantic genetic programming, greatly reducing the training time and reducing the consumption of computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a symbolic regression method based on an active learning strategy and library space optimization. Background Art

[0002] Data processing technology refers to a series of technologies and methods for collecting, storing, processing, and analyzing data. Among them, it is crucial to discover the complex relationships hidden behind the data and transform them into mathematical expressions that conform to the data distribution. Symbolic Regression (SR) provides an effective solution to this problem. Specifically, the basic idea of symbolic regression is to derive a mathematical expression through a series of symbolic operations based on the known relationship between independent variables and dependent variables, and this expression can best fit the existing data. These symbolic operations include addition, subtraction, multiplication, division, squaring, square rooting, etc. Among them, it searches in the space composed of mathematical expressions, trying to find a mathematical expression that fits the given data set. Traditional methods mainly use evolutionary computing techniques, especially Genetic Programming (GP), to solve this problem. In GP-based symbolic regression, mathematical expressions are represented as symbolic trees, where leaf nodes are input variables and constants, and non-leaf nodes are operators. The GP algorithm usually initializes a population composed of many symbolic trees, and then this population evolves generation by generation through methods such as crossover and mutation. Usually, the GP algorithm uses a fitness evaluation function to evaluate the quality of each individual in the population. Through the survival of the fittest, the population searches for the optimal individual. Among them, the individual fitness only depends on the final effect of program execution, and intermediate effects, such as the values calculated by subtrees of individual trees, are ignored.

[0003] In recent years, the method of introducing geometric semantics into genetic programming has received a lot of attention. The key innovation is to guide individuals towards a better fitness direction through the semantic space. The mapping from the original space to the semantic space in semantic genetic programming provides a theoretical framework for designing semantic operators. Geometric semantic operators aim to make the uncertainty of traditional operators develop towards certainty, set boundaries through semantics, and generate subprograms with similar or better performance than their parents through iteration. Currently, this research direction is still in the exploration stage, and the solutions vary. For example, some methods propose an angle selection operator and two angle geometric search operators, which bring new geometric properties to geometric operators by using angle perception, can approximate the target semantics in each iteration, and more importantly, can resist overfitting. In addition, some methods replace the tree structure with a more complex structure with a simpler tree structure through semantics, thereby reducing the computational cost. However, these methods all require a lot of time for library search and crossover and mutation of semantic vectors.

[0004] Therefore, semantic GP remains a simple and powerful tool for solving the symbolic regression problem. However, as the problem scale increases, the search space expands rapidly, and the search efficiency of traditional semantic GP methods is greatly reduced, making it difficult to obtain correct results within a limited time. The inefficiency of semantic GP stems from the maintenance and search of the semantic library, as well as the computational complexity of semantic vectors. Therefore, starting from this factor, this proposal presents a new and efficient GP method that can solve the technical problems of high computational complexity and long training time in previous semantic GP methods. Summary of the Invention

[0005] The object of the present invention is to provide a symbolic regression method based on an active learning strategy and library space optimization to solve the technical problems existing in the above-mentioned prior art. Semantic GP remains a simple and powerful tool for solving the symbolic regression problem. However, as the problem scale increases, the search space expands rapidly, and the search efficiency of traditional semantic GP methods is greatly reduced, making it difficult to obtain correct results within a limited time. The inefficiency of semantic GP stems from the size of its semantic space and the size of the library search space.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A symbolic regression method based on an active learning strategy and library space optimization, comprising the following steps:

[0008] S1: Measure the informativeness, diversity, and representativeness of data through an improved active learning strategy, so as to screen the input data, thereby reducing the labeling cost and reducing the dimension of the semantic vector; the specific steps are as follows:

[0009] S11: Obtain a data set containing N sample data; normalize the data, and then sequentially select sample data corresponding to the geometric sequence 1, 2, 4, 16,... through iteration. Assume that the total number of selected samples is represented by k. For the remaining N - k sample data that have not been selected Calculate the distance between them and the selected samples:

[0010]

[0011] where x n represents a certain sample in the sample set to be selected, and x m represents a certain sample in the selected sample set, represents the shortest distance from x n to the k selected samples to measure the sample diversity.

[0012] S12: Select a regression model f(x), and input the remaining N - k data Input into the regression model to obtain the output Calculate the distance between the regression result and the label:

[0013]

[0014] where y n represents the label of sample x n and is used to measure the informativeness of the data. For measuring the informativeness of the data.

[0015] S13: And by performing operation calculations on and to comprehensively represent the diversity and informativeness of the data, and select data by measuring the diversity and informativeness of the data;

[0016] S14: Use the DBSCAN clustering algorithm to perform noise removal operations on the sample data and select representative data;

[0017] S2: Initialize the population and represent the mathematical expression in a tree structure;

[0018] S3: Construct the initial library space from all subtrees of all individual trees in the population, calculate the semantic vectors of each subtree in the library space; further use the clustering algorithm for all individual trees in the library space, calculate and compare the similarity of the semantic vectors of two individual trees to measure the similarity between individual trees, and remove similar subtrees through similarity comparison;

[0019] S4: Each individual tree can be decoded into a mathematical expression, and its corresponding semantic vector is calculated. By performing calculations on the semantic vector of the parent tree, a new vector is generated as the value of the offspring vector to obtain the target semantics;

[0020] S5: Measure the similarity between the semantics of the subtrees in the library and the expected semantics of the offspring, and select the optimal subtree;

[0021] S6: After finding a subtree in the library that is closest to the expected semantics, perform linear scaling on the selected subtree in the library during the replacement process to reduce the error between the semantics of the subtree in the library and the target semantics.

[0022] A symbolic regression method based on an active learning strategy and library space optimization provided by the present invention can not only maintain the original accuracy of semantic genetic programming, but also greatly reduce the semantic space and library search space, thereby reducing the running time and accelerating the problem-solving speed. Brief Description of the Drawings

[0023] Figure 1 It is a flowchart of a symbolic regression method based on an active learning strategy and library space optimization provided by an embodiment of the present invention.

[0024] Figure 2 The flowchart of data selection based on the active learning strategy provided by the embodiments of the present invention.

[0025] Figure 3 The flowchart of initializing the library space provided by the embodiments of the present invention.

[0026] Figure 4 The schematic diagram of the change results before and after the semantic GP optimization with different iteration numbers provided by the embodiments of the present invention.

[0027] Figure 5 The schematic diagram of the results of the training time before and after the semantic GP optimization provided by the embodiments of the present invention. Specific embodiments

[0028] To make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the specific implementation manners of the present invention will be further described in detail below in conjunction with specific embodiments.

[0029] Taking the creep life feature selection of nickel-based superalloys as an example in this embodiment, the creep life data of 1200 nickel-based superalloy samples and their corresponding eight related features are obtained, which are: γ' volume fraction, shear modulus, antiphase domain boundary energy, stacking fault energy, γ' melting temperature, misfit degree, initial creep rate, applied stress, and creep temperature.

[0030] Based on the above nickel-based superalloy creep life data set, this embodiment provides a symbolic regression method based on the active learning strategy and library space optimization, and its process is as Figure 1 shown, and specifically includes the following steps:

[0031] S1: Measure the informativeness, diversity, and representativeness of the data through an improved active learning strategy, so as to screen the input data, thereby reducing the labeling cost and reducing the dimension of the semantic vector. The specific steps are as follows:

[0032] S11: Obtain the nickel-based superalloy creep life data set, given 8 features and labels of a single sample, and select the data, including: performing data normalization on the features. The process is as follows: for the feature parameters of the input sample, perform data standardization based on the mean and standard deviation of the original data. After standardization, the N data satisfy that the sample average within a certain feature is 0 and the variance is 1, and the mean of each feature of all samples in the data set is used as the centroid;

[0033] First, select a sample data closest to the centroid, and then select the sample data corresponding to the number of the geometric sequence 1, 2, 4, 16,... in turn. The total number of selected samples is represented by k. For the remaining N-k sample data Calculate the distances between them and the selected samples:

[0034]

[0035] Among them, x n represents a certain sample in the set of samples to be selected, and x m represents a certain sample in the selected sample set, represents the shortest distance from x n to the k selected samples. Try to make the distances between samples larger as much as possible to measure the realization of sample diversity.

[0036] S12: Use the XGBoost model as the regression model f(x), and input the N - k unselected input data into the XGBoost model to obtain the output Calculate the distance between the regression result and the label:

[0037]

[0038] Among them, y n represents the label of the sample x n , that is, the corresponding true regression result; subtract the output result of the regression model from the true result of the sample itself. If this difference is larger, it means that the model is less likely to regress the true result of the sample, and the sample contains more information, which can be used to measure the informativeness of the data;

[0039] S13: And multiply by to comprehensively represent the diversity and informativeness of the data:

[0040]

[0041] Further sort from largest to smallest and select the top n data;

[0042] S14: Then use the DBSCAN clustering algorithm to remove noise and select representative data at the same time. The steps are as follows:

[0043] S141: Initialization: Set the radius ε and the density threshold MinPts;

[0044] S142: Randomly select an unvisited data point;

[0045] S143: Check whether the number of data points in the ε neighborhood of this data point is greater than or equal to MinPts. If so, mark this data

[0046] point as a core point, otherwise mark it as a noise point;

[0047] S144: If the data point is a core point, starting from this point, all unvisited data points within its radius ε neighborhood

[0048] are added to the current cluster;

[0049] S145: Repeat step 4 until the ε neighborhoods of all data points in the current cluster have been visited;

[0050] S146: Mark all data points in the current cluster as visited;

[0051] S147: If the number of data points in the current cluster is greater than or equal to MinPts, add this cluster to the final clustering result;

[0052] S148: Repeat steps S142 - S147 until all data points have been visited;

[0053] The final clustering result is a set of clusters, where each cluster is composed of a core point and the data points within its ε neighborhood. There will also be some data points marked as noise points, which do not belong to any cluster.

[0054] S2: Initialize the population. The steps of representing the mathematical expression in a tree structure are as follows:

[0055] Implement the symbolic regression algorithm using an evolutionary algorithm and a tree - based coding method. In the proposed method, first, we need to define a set of mathematical operators, such as addition, subtraction, multiplication, division, etc., and operands, such as integers or decimals; these operators and operands will be used to construct mathematical expressions; starting from the root node, build sub - trees recursively; for each operator node, select an appropriate number of child nodes and select the corresponding operator or operand for each child node. Build sub - trees recursively until reaching the leaf nodes.

[0056] S3: The steps of initializing the library space and optimizing are as follows:

[0057] S31: First, form the library space by all sub - trees of the individual trees in all populations, and bring the data in the dataset into the individual trees, calculate the semantic vectors of each sub - tree in the library space, and all sub - tree semantic vectors form the semantic space;

[0058] S32: Then, apply the K - means clustering algorithm to all sub - trees in the library space. The specific steps are as follows:

[0059] S321: Normalize all sub - tree semantic vectors: (original value - minimum value) / (maximum value - minimum value) to obtain the set M of normalized clustering centers;

[0060] S322: Initialize: Set the number of clusters K, and randomly select K data points as the initial cluster centers:

[0061] M = m1, m2, …, m K ;

[0062] S323: Distance calculation: For each data point x i , calculate its distance to each center point m j , where the center point m j is the sample closest to the centroid:

[0063] d(x i , m j ) = (x i - c j ) 2 ;

[0064] where c j is the sample data corresponding to the center point sample m j ;

[0065] S324: Assignment: Assign each data point x i to the center point with the closest distance, that is, find the minimum distance:

[0066] j = argmin j d(x i , c j )

[0067] j is the subscript index corresponding to the minimum distance, and then assign the data point x i to the center point m j ;

[0068] S325: Update: For each cluster, calculate the total distance of all data points assigned to it, and select one of the data points as the new center point to minimize the total distance;

[0069] Repeat steps S323 - S325 until the center points no longer change or reach the predetermined number of iterations. The final clustering result is a set of clusters, where each cluster consists of a set of data points that are closest to the same center point;

[0070] S33: Use the subtree of the cluster center in each clustering cluster to represent the entire cluster, and further calculate the entropy value of each subtree semantic vector. Compare the entropies of pairwise semantic vectors to measure the similarity between subtrees, and remove similar subtrees through a threshold to reduce the library space and improve the search efficiency.

[0071] S4: Generate offspring target semantics by performing semantic crossover and mutation in the semantic space. The specific steps are as follows:

[0072] Through the selection operator, two parent generations p1 and p2 are selected, and the semantic vectors s(p1) and s(p2) of the two parent generations p1 and p2. The crossover operation generates the target semantics of the offspring:

[0073] o1 = s(p1) * k+(1 - k) * s(p2)

[0074] o2 = s(p1) * (1 - k)+k * s(p2)

[0075] where k is a random number between 0 and 1, and o1 and o2 are the target semantics of the offspring generated by the crossover operation;

[0076] The target semantics generated by the mutation operation is

[0077] m=(s(p1)+s(p2)) / 2;

[0078] By setting a threshold between 0 and 1 and generating a random number between 0 and 1 in each round for judgment. When the generated random number is greater than the set threshold, the crossover operation is executed; otherwise, the mutation operation is executed. And according to the executed operation, the final target semantics is obtained.

[0079] S5: The random expectation operator finds the subtree and replaces it

[0080] In the random expectation operator, a node is randomly selected from the parent individual tree, and semantic backpropagation is performed to obtain the expected semantics of the node. Search all subtrees in the library, find the subtree with the most similar semantics between the expected semantics and the subtree semantics in the library, and replace the subtree corresponding to the expected semantics with this subtree, so that the semantics of the final entire tree is as close as possible to the target semantics. The cosine distance is used to measure the similarity relationship between the expected semantics and the subtree semantics in the library:

[0081]

[0082] where γ is the cosine distance, t is the expected semantics, and ct is the subtree semantics;

[0083] S6: Linearly scale the individuals matched in the library as follows:

[0084] Whenever a library search is performed, for the subtree with the most similar semantics between the expected semantics and the subtree semantics in the library, the most ideal state is that the semantics of a certain subtree in the library is exactly equal to the expected semantics, which is often impossible and even has a large difference. Therefore, the least squares method is used for the expected semantics and the subtree semantics selected from the library to reduce the gap between the expected semantics and the subtree semantics in the library. This process involves calculating the optimal a and b coefficients, and multiplying the entire subtree selected from the library by b and then adding a, which can realize the linear transformation of the subtree semantics in the library and make its semantics more tend to the expected semantics. The a and b coefficients are calculated as follows:

[0085]

[0086] where t is the expected semantic vector, and ct is the semantic vector of the subtree in the library;

[0087] Execute the above process, continuously iterate new individuals through crossover and mutation, and update the population. Finally, select the optimal individual solution using the evaluation index. Specifically, select the mean squared error as the evaluation index, and the individual with the lowest mean squared error is the optimal individual.

[0088] Figure 4 It is a schematic diagram of the change results before and after semantic GP optimization for different iteration numbers. It can be seen that after sample selection based on the improved active learning strategy and library space reduction, compared with before improvement, the accuracy remains basically unchanged.

[0089] Figure 5 It is a schematic diagram of the training time results before and after semantic GP optimization. It can be seen that after sample selection based on the improved active learning strategy and library space reduction, compared with before improvement, the training time is significantly reduced.

Claims

1. A symbolic regression method based on an active learning strategy and library space optimization, characterized in that It includes the following steps: S1: Measure the informativeness, diversity, and representativeness of the data through an improved active learning strategy, and screen the input data. Specifically, it includes the following steps: S11: Obtain a dataset, where the dataset contains N sample data; specifically, the data are the creep life data of nickel-based superalloys and their corresponding eight related characteristics, namely: γ' volume fraction, shear modulus, antiphase domain boundary energy, stacking fault energy, γ' melting temperature, misfit, initial creep rate, applied stress, and creep temperature; normalize the data, and then sequentially select sample data corresponding to the number of terms in the geometric sequence 1, 2, 4, 8, 16,... through iteration. Assume that the total number of selected samples is represented by k. For the remaining N - k sample data Calculate the distances between them and the selected samples: where x n represents a sample in the set of samples to be selected, and x m represents a sample in the set of selected samples, represents the shortest distance from x n to k selected samples, to measure the achievement of sample diversity; S12: Select a regression model f(x), and input the N - k input data that have not been selected into the regression model to obtain an output Calculate the distance between the regression result and the label: Among them, y n represents the label of the sample x n , which is used to measure the informativeness of the data; S13: By operating on and to comprehensively represent the diversity and informativeness of the data, and select data by measuring the diversity and informativeness of the data; S14: Use the clustering algorithm to perform noise removal on the sample data and select representative data; S2: Initialize the population and represent the mathematical expression in a tree structure; S3: All subtrees of all individual trees in the population constitute the initial library space. Calculate the semantic vectors of each subtree in the library space using the dataset obtained in S1. Further, use the clustering algorithm for all individual trees in the library space, calculate and compare the similarity of the semantic vectors of two individual trees to measure the similarity between individual trees, and remove similar subtrees through similarity comparison; S4: Each individual tree can be decoded into a mathematical expression, and its corresponding semantic vector is calculated. By performing a calculation operation on the semantic vector of the parent tree, a new vector is generated as the value of the offspring vector to obtain the target semantics; S5: Measure the similarity between the semantics of the subtrees in the library and the expected semantics of the offspring, and select the optimal subtree; S6: After finding a subtree in the library that is closest to the expected semantics, perform linear scaling on the selected subtree in the library during the replacement process to reduce the error between the semantics of the subtree in the library and the target semantics.

2. The symbolic regression method based on an active learning strategy and library space optimization according to claim 1, wherein In step S1, the regression model is the XGBoost model, and the clustering algorithm is the DBSCAN clustering algorithm. The specific steps for noise removal are as follows: S141: Initialization: Set the radius ε and the density threshold MinPts; S142: Randomly select an unvisited data point; S143: Check whether the number of data points within the ε-neighborhood of this data point is greater than or equal to MinPts. If so, mark this data point as a core point; otherwise, mark it as a noise point; S144: If this data point is a core point, starting from this point, add all unvisited data points within its radius ε-neighborhood to the current cluster; S145: Repeat step S144 until all ε-neighborhoods of all data points in the current cluster have been visited; S146: Mark all data points in the current cluster as visited; S147: If the number of data points in the current cluster is greater than or equal to MinPts, add this cluster to the final clustering result; S148: Repeat steps S142 - S147 until all data points have been visited; The final clustering result is a set of clusters, where each cluster consists of a core point and the data points within its ε-neighborhood. At the same time, there will also be some data points marked as noise points that do not belong to any cluster.

3. The symbolic regression method based on an active learning strategy and library space optimization according to claim 2, wherein The specific steps of step S2 are as follows: Use the evolutionary algorithm and tree encoding method to implement the symbolic regression algorithm; in the proposed method, first define a set of mathematical operators and operands, which will be used to construct mathematical expressions; starting from the root node, construct subtrees recursively; for each operator node, select an appropriate number of child nodes and select the corresponding operator or operand for each child node; Construct subtrees recursively in this way until reaching the leaf nodes.

4. A symbolic regression method based on an active learning strategy and library space optimization according to claim 3, characterized in that, The specific steps of step S3 are as follows: S31: Compose all subtrees of the individual trees in all populations into a library space, bring the data in the dataset into the individual trees, calculate the semantic vectors of each subtree in the library space, and all subtree semantic vectors form a semantic space; S32: Apply the K - center clustering algorithm to all subtrees in the library space. The specific steps are as follows: S321: Normalize all subtree semantic vectors: (original value - minimum value) / (maximum value - minimum value) to obtain the normalized clustering center set M; S322: Initialize: Set the number of clusters K and randomly select K data points as the initial clustering centers: M = m1, m2, …, m K ; S323: Distance calculation: For each data point x i , calculate its distance to each center point m j . The center point m j is the sample closest to the centroid: d(x i ,m j ) = (x i - c j ) 2 ; Among them, c j is the sample data corresponding to the center point sample m j ; S324: Assignment: Assign each data point x i to the nearest center point, i.e., find the minimum distance: j = argmin j d(x i , c j ) j is the subscript index corresponding to the minimum distance, and then the data point x i is assigned to the center point m j ; S325: Update: For each cluster, calculate the total distance of all data points assigned to it, and select one of the data points as the new center point to minimize the total distance; Repeat steps S323 - S325 until the center points no longer change or reach a predetermined number of iterations; The final clustering result is a set of clusters, where each cluster consists of a set of data points that are closest to the same center point; S33: Use the subtree at the cluster center in each clustering cluster to represent the entire cluster, and further calculate the entropy value of each subtree semantic vector. Compare the entropies of pairwise semantic vectors to measure the similarity between subtrees, and remove similar subtrees through a threshold, thereby shrinking the library space and improving the search efficiency.

5. A symbolic regression method based on an active learning strategy and library space optimization according to claim 4, characterized in that, The specific steps of step S4 are as follows: Through the selection operator, select two parent trees p1 and p2, and the semantic vectors s(p1) and s(p2) of the two parent trees p1 and p2. The crossover operation generates the target semantics of the offspring: o1 = s(p1)*k+(1 - k)*s(p2) o2 = s(p1)*(1 - k)+k*s(p2) where k is a random number between 0 and 1, and o1 and o2 are the target semantics of the offspring generated by the crossover operation; The target semantics generated by the mutation operation is m=(s(p1)+s(p2)) / 2; By setting a threshold between 0 and 1 and generating a random number between 0 and 1 in each round for judgment, when the generated random number is greater than the set threshold, perform the crossover operation, otherwise perform the mutation operation, and obtain the final target semantics according to the executed operation.

6. The symbolic regression method based on an active learning strategy and library space optimization according to claim 5, wherein The specific steps of step S5 are as follows: In the random expectation operator, randomly select a node from the parent individual tree and perform semantic backpropagation to obtain the expected semantics of the node. Search all subtrees in the library, find the subtree whose semantic is most similar to the expected semantics in the library, and replace the subtree corresponding to the expected semantics, so that the semantics of the final entire tree is as close as possible to the target semantics. The cosine distance is used to measure the similarity relationship between the expected semantics and the subtree semantics in the library: where γ is the cosine distance, t is the expected semantics, and ct is the subtree semantics.

7. A symbolic regression method based on an active learning strategy and library space optimization according to claim 6, characterized in that The specific steps of step S6 are as follows: Whenever a library search is performed, the expected semantics is the most similar to the subtree semantics in the library. The ideal state is that the semantics of a subtree in the library is exactly equal to the expected semantics, which is often impossible, and even has a large difference. Therefore, the expected semantics and the semantics of the subtree selected in the library are least squared to reduce the gap between the expected semantics and the semantics of the subtree in the library. This process involves calculating the optimal a and b coefficients, and multiplying the subtree selected in the library by b and then adding a to achieve a linear transformation of the semantics of the subtree in the library, making its semantics more inclined to the expected semantics; the a and b coefficients are calculated as follows: Where t is the expected semantic vector, ct is the semantic vector of the subtree in the library; By executing the above process, new individuals are continuously iterated through crossover mutation, and the population is updated. Finally, the optimal individual solution is selected using the evaluation index.

8. A symbolic regression method based on an active learning strategy and library space optimization according to claim 7, characterized in that In step S6, the mean square error is selected as the evaluation index, and the individual with the lowest mean square error is the optimal individual.

Citation Information

Patent Citations

  • Adaptive symbol regression method based on multi-task genetic programming algorithm

    CN115543556A

  • Alternative techniques for design of experiments

    US20190383874A1