Genetic programming symbol regression method based on deep learning driving
Through the deep learning-driven genetic programming symbol regression method, the air quality prediction model is optimized using the symbol tree component prediction model, which solves the problem of insufficient air quality prediction efficiency and accuracy in the prior art, and achieves more efficient and accurate air quality prediction.
Patent Information
- Application Number
- CN202510401754.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has problems in repeated implementation of complete evolutionary processes, heavy computational burdens, and low resource utilization efficiency in air quality prediction. Multi-task optimization lacks efficient and accurate measurement methods, which limits the solution efficiency and accuracy of air quality prediction tasks.
The genetic programming symbol regression method driven by deep learning is adopted to predict the distribution of nodes and edges through the symbol tree component prediction model, and use it as prior knowledge to optimize the genetic programming symbol regression process, build an air quality prediction model, and combine the advantages of deep learning in data feature extraction and prediction to improve the model solution efficiency and accuracy.
The solution efficiency and accuracy of the air quality prediction model are significantly improved. The prior knowledge generated by deep learning is guided by genetic programming and optimized variation operations, achieving faster inference speed and more accurate prediction results, and solving the inefficiency and low accuracy problems existing in traditional methods.
Smart Images

Figure CN120337175A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and more specifically, to a genetic programming symbolic regression method driven by deep learning. Background Art
[0002] In the field of environmental research, especially in air quality prediction, researchers need to construct a model that can accurately predict the air pollution level based on various information such as meteorological data, pollutant emissions, traffic flow, and industrial activities. These models not only require a high degree of automation and flexible expression ability to process complex environmental data, but also require strong interpretability so that policymakers and environmental scientists can understand the mechanism behind the model, thus providing strong support for air pollution control and public health decision-making.
[0003] Currently, symbolic regression, as a typical representative of interpretable tasks, shows great application potential in the field of environmental research, especially in air quality prediction. Its core goal is to automatically discover and derive a mathematical model that describes the variation law of air pollutant concentration by using the given environmental data and a preset function set.
[0004] However, traditional single-task optimization genetic programming methods have problems such as repeatedly executing the complete evolution process, heavy computational burden, and low resource utilization efficiency when dealing with different air quality prediction tasks. Although the genetic programming method based on multi-task optimization provides the possibility of reducing the computational cost, the existing technology lacks an efficient and accurate measurement method to evaluate the similarity between different air quality prediction tasks, which limits the actual application effect of multi-task optimization.
[0005] Therefore, how to improve the solution efficiency and accuracy of air quality prediction tasks has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0006] In view of this, in order to at least partially solve the above problems, the present invention provides a genetic programming symbolic regression method driven by deep learning, which is particularly suitable for air quality prediction in the field of environmental research. This method significantly improves the solution efficiency and accuracy of the air quality prediction model by optimizing the evolution process of single-task genetic programming or multi-task genetic programming.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] A genetic programming symbolic regression method driven by deep learning for air quality prediction, comprising:
[0009] Obtain the environmental data to be regressed, and use the symbolic tree component prediction model to predict the distribution of nodes and edges in the symbolic tree based on the air quality data to be regressed;
[0010] Optimize the genetic programming symbolic regression process by taking the distributions of nodes and edges as prior knowledge, so as to construct a more accurate and efficient air quality prediction model.
[0011] Preferably, the symbolic tree component prediction model sequentially includes an input layer, a feature extraction layer, and an output layer;
[0012] The input layer is a fully connected feedforward network containing two ReLU activation functions, which is used to map the dimension of the input environmental data to the dimension required by the symbolic tree component prediction model.
[0013] The feature extraction layer is a BERT network, which is used to extract the global feature representation of the data after dimension adjustment.
[0014] The output layer is two parallel fully connected branches. The first fully connected branch is used to map the global feature representation to the prediction space, and generate the distribution of nodes through the softmax layer. The second fully connected branch is used to map the global feature representation to the prediction space, and generate the distribution of edges through the softmax layer.
[0015] Preferably, an environmental data set is constructed to train the symbolic tree component prediction model. In the data set, the true distribution P nodes of nodes and the true distribution P edge of edges are obtained through the following formula:
[0016]
[0017] where P node (c i ) represents the distribution probability of nodes of category c i (such as sin, cos, etc.) in the symbolic tree. N node is the total number of nodes in the symbolic tree. The indicator function takes the value of 1 if and only if the category node j of the node is equal to the category c i , otherwise it takes the value of 0;
[0018]
[0019] In the formula, represents the probability of the edge (c i →c j ), that is, the distribution probability that the type of the parent node is c i and the type of the child node is c j (for example, cos→exp). N edge is the total number of edges in the symbolic tree. The indicator function is 1 when the category of the parent node parent l of the edge edge lEqual to category c i and the child node category child l is equal to category c j take the value of 1, otherwise take the value of 0.
[0020] Preferably, the loss function during the training of the symbol tree component prediction model is:
[0021]
[0022] where KL represents the divergence loss, which is an index measuring the difference between two probability distributions, that is, the information loss caused when using one distribution to approximate another distribution, and P node is the true node distribution, is the predicted node distribution, and P edge is the true edge distribution, is the predicted edge distribution.
[0023] Among them, KL is defined as follows:
[0024]
[0025] Among them, P and respectively represent the true distribution and the predicted distribution, and x corresponds to the specific category.
[0026] Preferably, before using the component distributions of nodes and edges as prior knowledge, convert the predicted edge distribution into a two-dimensional matrix representation and normalize each row.
[0027] Preferably, optimize the normalized edge distribution and the predicted node distribution, including:
[0028] Remove redundant outputs according to the regression target task;
[0029] And impose a minimum probability threshold constraint.
[0030] Preferably, the steps for optimizing single-task genetic programming symbolic regression include:
[0031] Initialization: Select nodes for initialization using the roulette wheel algorithm according to the node distribution in the prior knowledge;
[0032] Mutation: Randomly select the node to be mutated and perform one mutation based on the node distribution in the prior knowledge; then perform a second mutation according to the type of the parent node of the current mutated node and its corresponding edge distribution.
[0033] Preferably, the steps for optimizing multi-task genetic programming symbolic regression include:
[0034] Calculate the task similarity according to the distributions of nodes and edges in the prior knowledge to obtain the task similarity matrix;
[0035] Convert the task similarity matrix into a knowledge transfer matrix;
[0036] Filter the knowledge transfer matrix according to a preset scalar threshold to generate task groups.
[0037] Preferably, calculate the task similarity according to the distribution of nodes and edges in the prior knowledge according to the following formula;
[0038]
[0039] where, and represent the predicted values of the node distributions of tasks T1 and T2, and represent the predicted values of the edge distributions of tasks T1 and T2, KL represents the KL divergence, which is an index for measuring the difference between two probability distributions, and λ n and λ e are balance coefficients for controlling the influence of nodes and edges. An exponential function is used for transformation so that the similarity values between tasks are in the range of (0, 1].
[0040] Preferably, convert the task similarity matrix into a knowledge transfer matrix through the following formula:
[0041]
[0042] In the formula, p ij represents an element in the knowledge transfer matrix, s ij represents an element in the task similarity matrix, i and j represent the serial numbers of tasks, N represents the total number of tasks, and α represents the temperature coefficient for controlling the knowledge transfer intensity.
[0043] The present invention discloses and provides a symbolic regression solving method that integrates the advantages of genetic programming and deep learning in a specific scenario. This method gives full play to the excellent performance of deep learning in data feature extraction and prediction, and organically combines the prior knowledge generated by the deep network with the global search ability of genetic programming, providing a new and efficient symbolic regression method for air quality prediction in the field of environmental research, effectively solving the problems existing in the prior art, and having remarkable technological progress and practical application value.
[0044] Compared with the prior art, the present application has the following effects:
[0045] 1. The symbolic tree feature prediction model provided by the present application has a fast inference speed and relatively accurate prediction results. By predicting the features of nodes and edges in advance, not only the utilization rate of symbolic regression data is improved, but also more comprehensive symbolic tree prediction information is provided; in terms of feature representation, the present invention uses the component distribution information to describe the symbolic tree features, which can describe the features more accurately;
[0046] 2. Using the prediction result of the symbol tree feature prediction model as prior knowledge to guide genetic programming can solve the problem of low utilization of data features in traditional genetic programming. Specifically,
[0047] 1) To make the distribution of the initial population closer to the true value, when initializing the population, the initial solution is generated through the probability of the symbol tree node component distribution. During the population evolution process, the model controls the mutation operation through the probability of the symbol tree edge component distribution, thereby guiding the genetic programming to evolve in a specific direction and improving the mutation efficiency. The present invention accelerates the search speed and accuracy of single-task optimization genetic programming and solves the problem of low efficiency of single-task optimization genetic programming.
[0048] 2) The present invention predicts the component distributions of symbol tree nodes and edges for different tasks, calculates the distribution similarity between tasks, and divides the tasks according to the similarity, constructs an accurate task similarity matrix for multi-task optimization genetic programming, and realizes the collaborative optimization between similar tasks, solving the problem that traditional multi-task optimization genetic programming cannot efficiently and accurately measure task similarity. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0050] Figure 1 Flowchart of the genetic programming symbolic regression method driven by deep learning disclosed in the present invention;
[0051] Figure 2 Deep learning model framework diagram for predicting the component distribution of symbol tree nodes and edges disclosed in the present invention;
[0052] Figure 3 Overall framework diagram of the single-task optimization genetic programming algorithm guided by deep learning disclosed in the present invention;
[0053] Figure 4 Overall framework diagram of the multi-task optimization genetic programming algorithm guided by deep learning disclosed in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0055] In the following description, many specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0056] In one embodiment, the present application discloses a genetic programming symbolic regression method driven by deep learning. By combining deep learning with genetic programming, the evolution process of single-task genetic programming or multi-task genetic programming is optimized to improve the solution efficiency and accuracy of the air quality prediction model.
[0057] The purpose of air quality prediction is to utilize information such as meteorological data, pollutant emissions, traffic flow, industrial activities, etc., and through constructing a mathematical model or a machine learning model, to predict the changing trends of the concentrations of air pollutants (such as PM2.5, PM10, NO2).
[0058] In air quality prediction, the definitions of the dependent variable and the independent variables are as follows:
[0059] Dependent variable: Usually the concentration of air pollutants, such as PM2.5 (μg / m³), PM10 (μg / m³), NO2 (μg / m³), CO (mg / m³), etc. These variables can be continuous values, used to represent the concentration level of pollutants at a certain time point or in a certain area.
[0060] Independent variables: Include various quantifiable data related to air quality, such as:
[0061] Meteorological data: Temperature (°C), humidity (%), wind speed (m / s), air pressure (hPa), precipitation (mm).
[0062] Pollutant emissions: SO2 (μg / m 3 )、CO (mg / m 3 )、NO2 (μg / m 3 )、VOCs (μg / m 3 ).
[0063] Traffic flow: The number of vehicles per unit time on the road (vehicles / hour), total exhaust emissions (mg / m 3 ).
[0064] Industrial activities: Industrial production index (IPI), energy consumption (tons of standard coal).
[0065] This prediction can help the government and environmental protection agencies optimize pollution control strategies, reduce the impact of air pollution on public health, while improving the early warning ability of the environmental monitoring system, and achieving more precise air pollution control and decision-making support.
[0066] In one embodiment, the steps of the genetic programming symbolic regression method are referred to Figure 1 and specifically include:
[0067] Obtain the environmental data to be solved, which contains multiple observation variables (independent variables) and target variables (dependent variables). Use the symbolic tree component prediction model to predict the distribution of nodes and edges in the symbolic tree based on the environmental data;
[0068] And use the distribution of nodes and edges as prior knowledge to optimize the genetic programming symbolic regression process, so as to construct a more accurate and efficient air quality prediction model.
[0069] Specifically, the present invention can be implemented through the following embodiments:
[0070] Embodiment 1
[0071] First, construct a symbolic tree component prediction model;
[0072] In this embodiment, the model framework diagram is as Figure 2 shown; it successively includes an input layer, a feature extraction layer, and an output layer;
[0073] Input layer: A fully connected feedforward network FFN containing two ReLU activation functions, used to map the dimension of the input data points to the dimension required by the symbolic tree component prediction model;
[0074] In this embodiment, a classification token [CLS] is introduced, which is encoded as a D + 1-dimensional vector. Preferably, the value of each dimension is 10,000; this token is used for downstream distribution learning tasks;
[0075] When inputting a set of N floating-point data points (y, x) ∈ R × R D (where x represents the independent variable and y represents the dependent variable), it is mapped from D + 1 dimension to D model dimension through the input layer, and then passed into the feature extraction structure, where N is the length of the test data, D is the dimension of the variable, and D model represents the dimension required by the subsequent network layer;
[0076] Feature extraction layer: In this embodiment, a standard BERT network is selected as the backbone network of the model, and an Encoder Only architecture is adopted to extract the global feature representation of the data after dimensional adjustment. Considering the characteristics of symbolic regression data, the input set of N data points has permutation invariance, so traditional position information encoding is not required, and the permutation invariance of the data can be directly removed or learned through the learnable position encoding of BERT.
[0077] Output layer: It consists of two parallel fully connected branches. The first fully connected branch is used to map the global feature representation to the prediction space, and generate the distribution of nodes through the softmax layer. The second fully connected branch is used to map the global feature representation to the prediction space, and generate the distribution of edges through the softmax layer.
[0078] In this application, the global feature representation corresponding to [CLS] in the BERT encoding result is extracted and input into two parallel fully connected branches respectively. After each branch is mapped to the prediction space through the fully connected layer, the softmax layer is used to generate the corresponding point and edge component distribution predictions, so as to achieve accurate modeling of the target distribution.
[0079] Embodiment 2
[0080] Construct an environmental data set to train the symbolic tree component prediction model;
[0081] Among them, the data set construction process includes:
[0082] 2.1 Generate a large number of symbolic trees by pre-order traversing environmental data;
[0083] 2.2 Remove the symbolic trees whose tree height and number of nodes do not meet the requirements from the generated symbolic trees;
[0084] 2.3 Simplify the remaining symbolic trees using the simplify method of SymPy;
[0085] 2.4 Sample the symbolic trees within the required data range to obtain the corresponding variables, function values, and the component distributions of the corresponding symbolic tree points and edges.
[0086] Among them, the true distribution P nodes and the true distribution P edge are obtained through the following formula:
[0087]
[0088] Among them, P node (c i ) represents the distribution probability of the i-th node of category c (such as sin, cos, etc.) in the symbolic tree, and N nodeis the total number of nodes in the symbol tree, indicating the function takes the value of 1 if and only if the category of the node node j is equal to the category c i , otherwise it takes the value of 0;
[0089]
[0090] In the formula, represents the probability of the edge (c i →c j ), that is, the probability distribution that the parent node type is c i and the child node type is c j (for example, cos→exp), N edge is the total number of edges in the symbol tree, and the indicating function takes the value of 1 when the parent node category parent l of the edge edge l is equal to the category c i and the child node category child l is equal to the category c j , otherwise it takes the value of 0.
[0091] Furthermore, the loss function during the training of the symbol tree component prediction model of this application is:
[0092]
[0093] where KL represents the divergence loss, which is an index for measuring the difference between two probability distributions, that is, the information loss caused when using one distribution to approximate another distribution, and P node is the true node distribution, is the predicted node distribution, P edge is the true edge distribution, is the predicted edge distribution.
[0094] Among them, KL is defined as follows:
[0095]
[0096] Among them, P and respectively represent the true distribution and the predicted distribution, and x corresponds to the independent variable data category of the discrete distribution, which represents the edge category and the point category in this article.
[0097] Example Three
[0098] Obtain the data to be regressed, and use the symbol tree component prediction model to predict the distributions of nodes and edges in the symbol tree;
[0099] Assume that the set of node types in the symbol tree contains n different types of nodes, and the number of function nodes is m. After training, the output results of the deep learning model for the test task are as follows:
[0100] Predicted value of node distribution: It is a probability vector of length n, representing the component distribution of each node type.
[0101]
[0102] Among them, p i represents the component probability of the i-th type of node in the symbol tree.
[0103] Predicted value of edge distribution: It is a probability vector of length n, representing the component distribution of each node type.
[0104]
[0105] Among them, q k represents the component probability of the k-th type of edge in the symbol tree.
[0106] In one embodiment, the predicted edge distribution is converted into a two-dimensional matrix representation, where each row corresponds to the component distribution of the edges starting from a specific function node. Specifically, assume that the symbol tree contains n types of node types, and the number of function nodes is m. Then the converted matrix has dimensions of m×n, and its form is as follows:
[0107]
[0108] Among them, q i,j represents the component probability of the edge pointing from the i-th function node to the j-th type of node.
[0109] Then, each row of the matrix is normalized to obtain a new set of probability vectors Specifically, it is represented as follows:
[0110]
[0111] This process is to normalize each row as a whole.
[0112] To further optimize the above technical solution, the normalized edge distribution and the predicted node distribution are optimized, including:
[0113] 1) Remove redundant outputs according to the regression target task; Since the output size of the neural network is fixed, different tasks share the same output dimension. For example, even if the target problem only involves variable x1, the model's output will still include x2 and x3; Similarly, even if the target equation does not contain the constant C, the output may still include the constant term. Therefore, we need to remove these redundant parts according to the specific problem requirements so that the output only contains the required symbol set;
[0114] 2) And impose a minimum probability threshold constraint; Since the same mathematical expression may have multiple equivalent forms, for example, sin(x) and sin(x + x - x) have the same mathematical meaning. Without constraints, the search in the optimization process may be overly concentrated on certain specific structures, thus affecting the diversity of solutions. To avoid falling into local optima while preferentially selecting key symbol nodes, we impose a minimum probability threshold constraint on each component of the probability distribution and then normalize it to ensure that all candidate symbols have a certain sampling probability, so as to enhance the exploration ability of the search process and the diversity of solutions.
[0115] After two steps of processing, we get is the processed node probability distribution; is the set of processed edge probability distributions.
[0116] Example 4
[0117] Use the distributions of nodes and edges as prior knowledge to optimize the genetic programming symbolic regression process;
[0118] In some embodiments, the single-task optimization genetic programming is guided according to the distributions of nodes and edges. The specific flowchart can be as Figure 3 shown; including:
[0119] 1) Initialization: According to the node distribution in the prior knowledge Use the roulette wheel algorithm to select nodes for initialization;
[0120] 2) Mutation: In the mutation process of genetic programming based on deep learning prior knowledge, during mutation, it is carried out in two steps. In the first step of mutation, randomly select a node, and based on Use the roulette wheel algorithm to perform mutation operations on the node to be mutated in its optional set. In the second step of mutation, randomly select a node again, find the parent node of the node to be mutated, obtain its type i, and based on Use the roulette wheel algorithm to perform mutation operations on the node to be mutated in its optional set.
[0121] In some embodiments, multi-task optimization genetic programming is guided by the distribution of nodes and edges. For this part, the present invention proposes a statistical method based on symbolic trees to evaluate the similarity of tasks by measuring the similarity of node distribution and edge connection distribution. The node distribution similarity is used to measure whether the occurrence probabilities of various symbols (such as operators and variables) in different tasks are similar, while the edge connection distribution similarity is used to evaluate whether the connection patterns of the symbolic tree structures are similar. Based on these similarity metrics, we construct a reasonable task similarity measure and divide similar tasks and control the gene exchange between multi-task populations accordingly, thereby improving the effectiveness and efficiency of parallel optimization.
[0122] The specific optimization flowchart can be as Figure 4 shown; the steps include:
[0123] 1) Calculate the task similarity according to the distribution of nodes and edges in prior knowledge to obtain a task similarity matrix; the calculation formula is:
[0124]
[0125] where and represent the predicted values of the node distributions of tasks T1 and T2, and represent the predicted values of the edge distributions of tasks T1 and T2, KL represents the KL divergence, an index used to measure the difference between two probability distributions, and λ n and λ e are balance coefficients for controlling the influence of nodes and edges. An exponential function is used for transformation so that the similarity values between tasks are in the range of (0, 1].
[0126] Assuming there are N tasks, then the task similarity matrix can be obtained:
[0127]
[0128] 2) Convert the task similarity matrix into a knowledge transfer matrix;
[0129] First, convert the elements in the task similarity matrix into elements in the knowledge transfer matrix through the following formula:
[0130]
[0131] In the formula, p ij represents the element in the knowledge transfer matrix, s ij represents the element in the task similarity matrix, i and j represent the serial numbers of tasks, N represents the total number of tasks, and α represents the temperature coefficient for controlling the knowledge transfer intensity.
[0132] Then, the following knowledge transfer matrix P is obtained;
[0133]
[0134] 3) Filter the knowledge transfer matrix according to a preset scalar threshold to generate task groups.
[0135] In this application, by setting a similarity threshold, tasks with a similarity higher than the threshold are grouped into the same group for subsequent adoption of a more efficient collaborative optimization strategy, while tasks with a similarity lower than the threshold are grouped separately to reduce interference with the optimization processes of other tasks, thereby improving the overall optimization performance.
[0136] Further use the knowledge transfer matrix P to guide the knowledge sharing strategy between tasks. For example, in the multi-task optimization of multiple populations, the individual exchange between different task populations can be flexibly adjusted according to this matrix to optimize the efficiency of knowledge transfer. Specifically, the number of individual exchanges N between task i and task j i,j is calculated by the following formula:
[0137]
[0138] where N t represents the preset total number of exchanged individuals, and P i,j is the knowledge transfer probability from task population i to task population j.
[0139] The present invention first generates a large number of symbolic regression data to train a neural network model to extract the distribution characteristics of nodes and edges in the symbolic tree; then uses the trained model to predict the characteristic information of the symbolic tree in a specific problem and integrates it as prior knowledge into the evolutionary process of single-task genetic programming, thereby effectively improving the search efficiency and the quality of solutions. In addition, the model can also identify highly similar tasks, construct an accurate task similarity matrix for multi-task optimization genetic programming, and achieve collaborative optimization between similar tasks. This method gives full play to the excellent performance of deep learning in data feature extraction and prediction, and organically combines the prior knowledge generated by the deep network with the global search ability of genetic programming, providing a new and efficient symbolic regression method for air quality prediction in the field of environmental research, effectively solving the problems existing in the prior art, and having significant technological progress and practical application value
[0140] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for the relevant parts.
[0141] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A genetic programming symbolic regression method driven by deep learning, characterized in that: Obtain environmental data to be regressed, and use a symbolic tree component prediction model to predict the distribution of nodes and edges in the symbolic tree based on the environmental data; Use the distribution of nodes and edges as prior knowledge to optimize the genetic programming symbolic regression process to construct an air quality prediction model.
2. The symbolic regression method according to claim 1, characterized in that The symbolic tree component prediction model sequentially includes an input layer, a feature extraction layer, and an output layer; The input layer is a fully connected feedforward network containing two ReLU activation functions, which is used to map the dimension of the input environmental data to the dimension required by the symbolic tree component prediction model. The feature extraction layer is a BERT network, which is used to extract the global feature representation of the data after dimension adjustment. The output layer is two parallel fully connected branches. The first fully connected branch is used to map the global feature representation to the prediction space, and the distribution of nodes is generated through the softmax layer. The second fully connected branch is used to map the global feature representation to the prediction space, and the distribution of edges is generated through the softmax layer.
3. The symbolic regression method according to claim 1, wherein Construct an environmental data set to train the symbol tree component prediction model. In the data set, the true distribution P of nodes nodes and the true distribution P of edges edge are obtained by the following formula: Among them, P node (c i ) represents the distribution probability of nodes of category c i (such as sin, cos, etc.) in the symbol tree, N node is the total number of nodes in the symbol tree, and the indicator function takes the value 1 if and only if the category of the node node j is equal to category c i , otherwise it takes the value 0; In the formula, represents the probability of edge (c i →c j ), that is, the distribution probability that the parent node type is c i and the child node type is c j (for example, cos→exp), N edge is the total number of edges in the symbol tree, and the indicator function takes the value of 1 when the parent node category parent l of the edge edge l is equal to the category c i and the child node category child l is equal to the category c j , and takes the value of 0 otherwise.
4. The symbolic regression method according to claim 1, wherein The loss function during the training of the symbolic tree component prediction model is: where KL represents the divergence loss, and P node is the true node distribution, is the predicted node distribution, and P edge is the true edge distribution, is the predicted edge distribution.
5. The symbolic regression method according to claim 1, wherein Before using the component distributions of nodes and edges as prior knowledge, convert the predicted edge distribution into a two-dimensional matrix representation and normalize each row.
6. The symbolic regression method according to claim 5, wherein Optimize the normalized edge distribution and the predicted node distribution, including: Remove redundant outputs according to the regression target task; And impose a minimum probability threshold constraint.
7. The symbolic regression method according to claim 1, characterized in that Optimize single-task genetic programming symbolic regression. The steps include: Initialization: Select nodes for initialization according to the node distribution in the prior knowledge; Mutation: Randomly select a node to be mutated, perform a mutation based on the node distribution in the prior knowledge; then perform a secondary mutation according to the type of the parent node of the current mutated node and its corresponding edge distribution.
8. The symbolic regression method according to claim 1, characterized in that, Optimize multi-task genetic programming symbolic regression. The steps include: Calculate the task similarity according to the node and edge distributions in the prior knowledge to obtain a task similarity matrix; Convert the task similarity matrix into a knowledge transfer matrix; Filter the knowledge transfer matrix according to a preset scalar threshold to generate task groups.
9. The symbolic regression method according to claim 8, wherein, Calculate the task similarity according to the node and edge distributions in the prior knowledge according to the following formula; Among them, and represent the predicted node distribution values of tasks T1 and T2, and represent the predicted edge distribution values of tasks T1 and T2. KL represents the KL divergence, which is an index used to measure the difference between two probability distributions. λ n and λ e are the balance coefficients that control the influence of nodes and edges.
10. The symbolic regression method according to claim 8, wherein Convert the task similarity matrix into a knowledge transfer matrix through the following formula: where p ij represents an element in the knowledge transfer matrix, s ij represents an element in the task similarity matrix, i and j represent the serial numbers of tasks, N represents the total number of tasks, and α represents the temperature coefficient for controlling the knowledge transfer intensity.