Symbol regression-based time sequence analysis method
By combining Monte Carlo tree search and strategy-value network, symbol enhancement strategies are introduced, and the balance between prediction accuracy and interpretability in time series analysis is solved, and efficient and interpretable time series modeling and prediction are achieved, which is suitable for finance, medicine, meteorology and other fields.
Patent Information
- Application Number
- CN202510400843.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
When existing time series analysis methods deal with complex nonlinear and non-stationary data, it is difficult to find a balance between prediction accuracy and model interpretability. The traditional symbol regression method has low computational efficiency and limited generalization ability.
Combining Monte Carlo tree search and strategy-value network, we explore mathematical expression space through intelligent exploration, introduce symbol enhancement strategies, automatically learn composite functions, and improve search efficiency and model generalization capabilities.
It realizes efficient and interpretable time series modeling and prediction, improves computing efficiency and prediction accuracy, and is suitable for many fields such as finance, medicine, and meteorology.
Smart Images

Figure CN120337173A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of time series analysis, and particularly relates to a time series analysis method based on symbolic regression. Background Art
[0002] In the context of the rapid development of information technology, data has become a key fundamental resource for promoting scientific and technological progress and social development. As an important form of data, time series data has a wide range of applications in many fields such as finance, healthcare, meteorology, and industrial monitoring due to its ability to reflect the evolution of phenomena over time. How to effectively analyze the evolution law of time series data, extract core patterns, and achieve accurate prediction has become a hot and difficult issue in current research.
[0003] Traditional time series analysis methods are mainly divided into two categories: statistical modeling and deep learning. Statistical modeling methods, such as autoregressive model (AR), moving average model (MA), and Kalman filter, usually perform well in dealing with simple and relatively stable time series prediction tasks based on the assumptions of linearity and stationarity. However, when facing complex dynamic systems with high nonlinearity and non-stationarity, the limitations of these methods gradually emerge, and it is difficult to fully capture the deep features and internal laws in the data. In contrast, deep learning methods that have developed rapidly in recent years, including recurrent neural network (RNN), long short-term memory network (LSTM), gated recurrent unit (GRU), and Transformer model, etc., can effectively capture complex dynamic relationships and thus significantly improve the prediction performance due to their nonlinear network structure and powerful feature extraction ability. However, deep learning models usually require a large amount of data support, the training process is complex, and the demand for computing resources is high. More importantly, due to the highly complex internal computing process, the model decision-making often lacks interpretability and shows typical "black box" characteristics. This opacity greatly reduces the credibility of the decision-making process, especially in high-risk decision-making fields such as medical diagnosis, financial risk control, and autonomous driving, severely limiting the practical application of deep learning models.
[0004] Therefore, a core challenge in the field of time series modeling is how to find an effective balance between prediction accuracy and model interpretability. Achieving high interpretability while maintaining prediction accuracy is a complex and challenging task. For this reason, researchers need to develop more transparent and easy-to-interpret new machine learning frameworks, or explore hybrid methods that combine the advantages of statistical modeling and deep learning. In addition, using posterior interpretation techniques to enhance the transparency and decision-making credibility of the model has also become an important research direction.
[0005] In this context, symbolic regression, as an emerging modeling method, has attracted wide attention. Symbolic regression can automatically discover mathematical expressions from data and is particularly suitable for revealing the underlying mathematical laws behind time series data. Different from traditional parameter fitting methods (such as linear regression or polynomial regression), symbolic regression does not require a priori setting of a specific expression form. Instead, it explores the space of symbolic expressions in a data-driven manner and automatically constructs the optimal mathematical model. This method can flexibly and accurately describe the internal dynamic characteristics of time series. However, traditional symbolic regression methods mainly rely on evolutionary algorithms or genetic programming to generate and optimize expressions by simulating the natural evolution process. Such methods usually have low computational efficiency and slow convergence speed, making it difficult to meet the requirements of large-scale data and complex feature spaces. In addition, traditional symbolic regression often requires manual setting of the search space and initial conditions, and its generalization ability and adaptability are limited to a certain extent.
[0006] To address these issues, improving the computational efficiency, self-adaptability of the search strategy, and generalization ability of symbolic regression methods has become an important research direction. For example, combining reinforcement learning with symbolic regression, enabling the agent to autonomously learn the optimal search strategy to improve search efficiency and model accuracy; using the mutation and crossover operations of genetic programming and combining modern optimization techniques such as particle swarm optimization (PSO) or ant colony optimization (ACO) to accelerate convergence and avoid falling into local optima; in addition, introducing the ideas of transfer learning and meta-learning to enable the symbolic regression model to quickly adapt between different datasets and tasks, thereby enhancing the generalization ability. There are also studies attempting to combine symbolic regression with deep learning to construct a hybrid model to take into account the feature extraction ability of deep learning and the interpretability of symbolic regression. These efforts aim to promote the application of symbolic regression in time series modeling, achieve an effective balance between prediction accuracy and model interpretability, and provide new possibilities for solving key problems in current time series analysis. However, breaking through this performance bottleneck in symbolic regression, constructing an efficient, flexible, and generalizable modeling strategy, while maintaining sufficient interpretability, remains a formidable challenge. Summary of the Invention
[0007] To effectively solve the above problems, the present invention proposes a time series analysis method based on symbolic regression with high efficiency, low cost, and strong interpretability, which is specifically used for the modeling and prediction tasks of time series.
[0008] The time series analysis method based on symbolic regression proposed by the present invention combines Monte Carlo tree search (MCTS) with a policy-value network. Through the policy optimization ability of Monte Carlo tree search, it conducts efficient and intelligent exploration in the vast search space of mathematical expressions. At the same time, it uses the policy-value network to evaluate candidate expressions in real time, effectively guiding the search process and significantly improving the search efficiency and quality. In addition, in order to further enhance the generalization performance and adaptability of symbolic regression expressions, a symbolic enhancement strategy is innovatively introduced, which automatically learns composite function structures from data to more flexibly capture dynamic patterns and non-linear relationships in complex time series. Through this comprehensive strategy integrating search, reinforcement learning, and symbolic enhancement, the present invention effectively improves the computational efficiency, prediction accuracy, and model generalization ability of the symbolic regression method, thus significantly expanding the practical application potential of symbolic regression in the field of complex time series analysis and prediction.
[0009] The present invention constructs a mathematical expression of time series in the form of a composite function, rather than simply a linear combination. This approach allows the model to automatically learn multi-level mathematical structures, making the modeling process more flexible. For example, in financial time series analysis, this method can automatically discover the mathematical patterns of market fluctuations, and in medical signal analysis, it can deduce the changing rules of pathophysiological processes. The core advantage of this modeling method is that it not only ensures the interpretability of mathematical expressions but also enhances the adaptability of the model to different types of time series data. In addition, to improve computational efficiency, this method introduces a reinforcement learning mechanism and adopts an intelligent strategy in the symbolic regression search process to reduce the search space and improve the speed of expression discovery. At the same time, it combines neural networks for search optimization, enabling the model to adaptively select the optimal expression and improve the accuracy and stability of time series prediction. Compared with traditional symbolic regression methods, this method has significant improvements in computational complexity, search efficiency, and modeling accuracy.
[0010] In summary, the time series analysis method (modeling method) based on symbolic regression proposed by the present invention realizes efficient and interpretable time series modeling and prediction by combining Monte Carlo tree search optimized by neural networks and a symbolic enhancement strategy. This method can automatically discover mathematical laws from data, improve the transparency and generalization ability of the model, and provide an efficient and intelligent modeling tool for time series analysis, which is applicable to multiple fields such as finance, medicine, meteorology, and industry.
[0011] The time series analysis method based on symbolic regression proposed by the present invention specifically includes the following steps:
[0012] (1) Data preparation: Before optimizing the expression, it is first necessary to construct a complete initial environment to support the subsequent optimization process. This environment includes the following key elements:
[0013] Formula library: Used to pre-store common mathematical formulas for reference and invocation during the optimization process.
[0014] Operator and function library: Includes basic mathematical operators (such as addition +, subtraction -, multiplication *, division / ) and advanced mathematical functions (such as logarithm log, exponential exp, sine sin, cosine cos, etc.). These operators and functions are the basic components for constructing and optimizing expressions.
[0015] For example, parse the input mathematical expression and transform it into a tree structure. In this structure: Each node represents an operation or an operand. The root node represents the final calculation result of the expression. The internal nodes represent operators or functions, and the leaf nodes represent operands. For example, the expression log(cos(x)+sin(y)) can be parsed into the following tree structure:
[0016] Root node: log (logarithmic function)
[0017] Child node: + (addition operation)
[0018] Children of +:
[0019] cos(x): Further decomposed into the cos function and the variable x
[0020] sin(y): Further decomposed into the sin function and the variable y.
[0021] (2) Model construction: The core of expression optimization is to achieve automated adjustment of the expression through a policy-value network. Model construction includes the construction of a policy network and a value network;
[0022] The policy network is responsible for selecting a suitable node for operation in the current expression structure; this node can be the root node, a child node, or a leaf node, and the selection criterion is based on the priority score of the node by the policy network; the value network is used to evaluate the contribution of each operation to the overall optimization goal and provide guidance for the policy network.
[0023] Definition of the operation: On each selected node, multiple operations can be performed, including adding functions (such as logarithm, exponential, etc.), replacing operators (such as replacing + with *), deleting redundant nodes, or changing the hierarchical structure of the expression. For example, on the parent node log() or exp(), given the child node +, it can be expanded to log(+) or exp(+).
[0024] Achieving automated adjustment of the expression through a policy-value network includes representation learning and state update of the policy-value network:
[0025] In an expression tree, the state of each node is determined not only by its own characteristics but also depends on the characteristic information of its child nodes. The model explicitly determines the dependencies between nodes through a neural network, thereby effectively aggregating the characteristics of child nodes to update the state of the current node. For example, for the node "+", its child nodes are cos(x) and sin(y) respectively, and the model will aggregate the characteristics of the two according to the coefficients learned by the neural network to form a new state representation of the node "+".
[0026] Expression score calculation:
[0027] After each operation, update the state of the entire expression and calculate its optimization score through a value network. The optimization score measures the quality of the current expression and is used to guide subsequent operation selections. For example, the simplified expression log(exp(x)) may have a higher score than the original expression log(cos(x)+sin(y)).
[0028] Furthermore, the policy-value network is a set of neural network structures sharing a common backbone, which is composed of an encoding layer, a feature fusion layer, and two independent multi-layer perceptrons as a whole; the policy-value network accepts two sets of inputs: the current expression path in Monte Carlo tree search and the time series to be modeled; the backbone network first performs embedding and position encoding processing on the expression path (as a symbol sequence) and the time series to be modeled respectively, then extracts deep joint features through a multi-layer Transformer encoder, mixes the features, and then uses two independent multi-layer perceptrons as a policy selector and a value estimator to input the currently selected policy and the estimated value of the current state; the policy selector outputs the probability distribution of the operation nodes, and the value estimator outputs the scoring result of the current expression state, providing accurate and intelligent decision-making basis for the expression generation and optimization process.
[0029] (3) Iteratively optimize parameters: In the process of expression optimization, update the parameters of the policy-value network through multiple rounds of iteration to achieve gradual optimization. The specific process is as follows:
[0030] (a) Initialize parameters:
[0031] Initialize the parameters for the policy network and the value network, and set the learning rate and the optimization objective function (such as computational efficiency or expression complexity).
[0032] (b) Gradient descent update:
[0033] After each operation, calculate the gradient according to the objective function and update the network parameters through the gradient descent method. The negative gradient direction represents the optimization path, and the model gradually approaches the optimal solution through multiple iterations.
[0034] (c) Reward signal feedback:
[0035] To further enhance the model's selection ability, a reward signal is generated based on the optimization effect after each operation. This reward signal is obtained by calculating the symbolic regression expression as described above, and the value function of the expression is fed back to the policy network in the model for model optimization. This mechanism is similar to policy optimization in reinforcement learning and can effectively improve the model's performance in complex expression optimization.
[0036] (4) Complete network alignment: After multiple rounds of iteration, the expression is finally optimized under the guidance of the policy-value network. The optimization process includes the following steps:
[0037] (a) Backtracking and cumulative reward:
[0038] For the expression after each operation, accumulate the reward and update the total score to provide a more accurate reference for subsequent operations.
[0039] (b) Determine the optimal expression:
[0040] After meeting the termination conditions (such as reaching the set number of iterations or the expression score no longer improving), output the expression with the highest score as the final optimization result. For example, optimizing the initial expression log(cos(x)+sin(y)) to log(exp(x)) not only has higher computational efficiency but also a more concise structure.
[0041] (c) Expression output and verification:
[0042] After optimization, verify the generated expression to ensure its functional equivalence with the initial expression, and at the same time compare the scores before and after optimization to evaluate the optimization effect.
[0043] The method of the present invention has been experimentally verified on multiple real-world datasets, including fields such as financial forecasting, meteorological modeling, and urban computing. The experimental results show that this method is superior to existing symbolic regression methods in terms of data fitting ability, computational efficiency, and prediction stability. At the same time, compared with deep learning models, it provides higher interpretability. Especially in application scenarios that require transparency, such as medical diagnosis, intelligent monitoring, and scientific modeling, this method can provide intuitive mathematical expressions, making the prediction results more understandable and providing a reliable analysis basis for users.
[0044] The present invention realizes a fully automated process for expression optimization. In the experiment, different expression datasets were used for verification, and the results show that the optimization method based on the policy-value network has the following advantages compared with traditional optimization methods: it improves the optimization efficiency, reduces the redundancy of expressions, and enhances the computational stability of expressions, especially showing significant effects when dealing with complex nested expressions.
[0045] Key technologies and advantages in the present invention:
[0046] (1) Employ symbolic regression and Monte Carlo tree search (MCTS):
[0047] Utilize MCTS to efficiently explore the expression space, intelligently balance the exploration and exploitation of expressions, significantly improve the search efficiency and reduce the computational complexity; compared with traditional genetic programming methods, MCTS can more effectively handle complex and large-scale data sets.
[0048] (2) Construct a policy-value network;
[0049] To further optimize the search efficiency, the present invention introduces a policy-value network to combine neural networks to calculate the advantages and disadvantages of candidate mathematical expressions. Compared with traditional random search methods, the policy-value network can intelligently guide the search process, reduce ineffective searches, and improve the model convergence speed.
[0050] (3) Adopt a symbolic enhancement strategy;
[0051] The expression search space of symbolic regression is relatively large. Therefore, the present invention adopts a symbolic enhancement strategy to automatically extract high-value mathematical composite functions from data for optimizing the mathematical expression ability of time series modeling. For example, automatically learn common mathematical patterns, such as periodic functions, exponential decay functions, logarithmic growth functions, etc., to improve the accuracy and generalization ability of modeling.
[0052] (4) Enhanced interpretability
[0053] The time series mathematical expressions generated by this method have high interpretability and can clearly display the evolution pattern of data. For example, in financial market prediction, this method can generate an interpretable mathematical model to describe the internal law of price fluctuations; in infectious disease analysis, it can generate mathematical expressions of the change of infectious diseases over time, providing a reliable analysis tool for medical research. Brief description of the drawings
[0054] Figure 1 Is the neural-enhanced Monte Carlo tree search.
[0055] Figure 2 Is the schematic diagram of the policy-value network.
[0056] Figure 3 Is the qualitative analysis of the neural-enhanced Monte Carlo tree search for symbolic regression. Detailed implementation manners
[0057] The present invention will be further introduced below through simulation experiments in combination with the drawings.
[0058] The simulation experiments were conducted on a computer with an 8-core Intel CPU with 64GB of physical memory and an NVIDIA V100 GPU. The operating system was Ubuntu 18.04, and the main coding language was Python. The simulation experiments were conducted on three public datasets: Weighted Influenza-like Illness Percentage (WILI), Australian Daily Currency Exchange Rate (ACER), and Atmospheric Pressure (AP).
[0059] Both the policy network and the value network are constructed based on the Transformer attention mechanism and jointly modeled through a shared backbone structure. Among them, the symbol sequence and the time series are first embedded into a unified high-dimensional vector space. The symbols are mapped through the embedding layer, and the numerical sequence is processed by linear transformation, and a positional encoder is added to strengthen the sequential perception. The encoder consists of 3 layers, each layer including a multi-head self-attention mechanism, a feed-forward network, GELU activation, layer normalization, and residual connection to extract the deep dependencies between sequences. The model finally generates the operation selection probability distribution (policy output) and the expression score (value output) through multi-layer perceptrons with two branches respectively. The policy output is processed by softmax, and the value output is normalized by sigmoid. The training uses cross-entropy and mean squared error loss functions. The optimizer is Adam, the learning rate is 1e-4, the hidden layer dimension is 64, it contains 4 attention heads, and the feed-forward network is a two-layer neural network with a middle layer of 256 dimensions.
[0060] The experimental results show that:
[0061] Improved fitting accuracy: This method improves the prediction accuracy compared with traditional symbolic regression methods and has better interpretability than deep learning methods.
[0062] Improved computational efficiency: Monte Carlo tree search combined with the policy-value network effectively reduces the computational overhead, enabling symbolic regression to run efficiently on large-scale datasets.
[0063] Enhanced generalization ability: The symbolic enhancement strategy improves the adaptability of the model, enabling accurate modeling of different types of time series data.
[0064] The superiority of this method is demonstrated through quantitative analysis and qualitative analysis.
[0065] For quantitative analysis, on public datasets, by comparing the prediction performance with current time series modeling methods and comparing the fitting performance and efficiency with current symbolic regression algorithms, the superiority of this method in multiple quantitative metrics, including fitting error (coefficient of determination R 2 ), trend error (correlation coefficient CORR), and efficiency (average time cost per sample ATC), is demonstrated. See Table 1 for details.
[0066] The coefficient of determination is a dimensionless index used to evaluate the model fitting ability, reflecting the proportion of the model error relative to the average error. In quantitative analysis, the present invention is significantly superior to other existing methods in terms of the coefficient of determination. The increase is about 203.04%. These data fully illustrate the outstanding performance of the model proposed by the present invention in terms of fitting ability. The correlation coefficient is used to evaluate the consistency between the model predicted value and the actual observed value. Under this index, the proposed model also shows significant advantages. The increase is about 103.65%. These results further confirm the effectiveness and reliability of the proposed model in capturing the trend of actual data. At the same time, in terms of efficiency, the algorithm efficiency is measured by evaluating the average time cost of each sample. The results show that the present invention has achieved a significant reduction in the average time cost, with an increase of about 68.06%.
[0067] Table 1
[0068]
[0069]
[0070] For qualitative analysis, some sequences were randomly selected and fitted, as specifically shown in Figure 3 . Figure 3 Each graph in it is derived from different real datasets, a set of real time series with a length of 36 randomly extracted according to sliding sampling. The first 30 timestamps are used for fitting to show the fitting performance, and the last 6 timestamps are used for prediction to show the reliability of the fitted expression. The analysis shows that the present invention has successfully fitted complex real data using specific mathematical models. It is worth noting that these mathematical models can be expressed by relatively simple expressions, clearly showing the trend of the time series. These visualization results not only prove the ability of the present invention in fitting complex data, but also demonstrate its effectiveness in refining and expressing the data trend. In this way, the present invention can provide profound and intuitive insights into time series data, helping to better understand the data. In addition, it is particularly emphasized that after using the model to fit the data, a comparison was also made on the subsequent data that did not participate in the fitting, which actually constitutes a prediction behavior. The gray background part in the graph shows the comparison between the subsequent data that did not participate in the fitting and the model prediction curve. It can be observed from this that the present invention can not only accurately fit the existing observed data, but also make reliable predictions for the subsequent data. This not only shows that the present invention has good prediction ability, making the expressions mined by it more valuable, but more importantly, its reliable prediction ability stems from the fact that its expressions precisely capture the pattern of data evolution.
[0071] Generally speaking, this method effectively combines the policy network and the value network, and realizes the efficient modeling and decision-making of expression optimization through the attention mechanism, providing a new intelligent solution for the simplification and optimization of mathematical expressions.
Claims
1. A time series analysis method based on symbolic regression, characterized in that, Combine Monte Carlo Tree Search (MCTS) with a policy-value network, and through the policy optimization ability of MCTS, conduct efficient and intelligent exploration in the vast search space of mathematical expressions; Utilize the policy-value network to evaluate candidate expressions in real time to guide the search process and improve search efficiency and quality; In addition, introduce a symbolic enhancement strategy to automatically learn composite function structures from data to more flexibly capture dynamic patterns and non-linear relationships in complex time series; through a comprehensive strategy integrating search, reinforcement learning, and symbolic enhancement, effectively improve the computational efficiency, prediction accuracy, and model generalization ability of the symbolic regression method; Among them, construct the mathematical expression of the time series in the form of a composite function, allowing the model to automatically learn multi-level mathematical structures, making the modeling process more flexible, ensuring the interpretability of the mathematical expression, and enhancing the adaptability of the model to different types of time series data; in addition, introduce a reinforcement learning mechanism, adopt an intelligent strategy in the symbolic regression search process to reduce the search space and improve the speed of expression discovery; at the same time, combine neural networks for search optimization, enabling the model to adaptively select the optimal expression and improve the accuracy and stability of time series prediction.
2. The time series analysis method based on symbolic regression according to claim 1, wherein The specific steps are as follows: (1) Data preparation: Before optimizing the expression, it is first necessary to construct a complete initial environment to support the subsequent optimization process; The initial environment includes: Formula library: Used to pre-store common mathematical formulas for reference and invocation during the optimization process; Operator and function library: Contains basic mathematical operators, including addition +, subtraction -, multiplication *, division / , and advanced mathematical functions, including logarithm log, exponential exp, sine sin, cosine cos; these operators and functions are the basic components for constructing and optimizing expressions; Parse the input mathematical expression and convert it into a tree structure; in the tree structure, each node represents an operation or operand; the root node represents the final calculation result of the expression; the internal nodes represent operators or functions, and the leaf nodes represent operands; (2) Model construction: The core of expression optimization is to achieve automatic adjustment of the expression through a policy-value network; model construction includes the construction of a policy-value network; The policy network is responsible for selecting a suitable node for operation in the current expression structure; this node can be the root node, a child node, or a leaf node, and the selection criterion is based on the priority score of the node by the policy network; the value network is used to evaluate the contribution of each operation to the overall optimization goal and provide guidance for the policy network; Definition of operations: On each selected node, multiple operations can be performed, including adding functions, replacing operators, deleting redundant nodes, or changing the hierarchical structure of the expression; Achieve automatic adjustment of the expression through the policy-value network, including representation learning and state update of the policy-value network: In the expression tree, the state of each node is determined not only by its own characteristics but also by the characteristic information of its child nodes; the model explicitly determines the dependencies between nodes through a neural network, thereby effectively aggregating the characteristics of child nodes to update the state of the current node; Expression score calculation: After each operation, update the state of the entire expression and calculate its optimized score through a value network; the optimized score measures the quality of the current expression and is used to guide subsequent operation selections; (3) Iteratively optimize parameters: During the process of expression optimization, update the parameters of the policy network and the value network through multiple rounds of iteration to achieve gradual optimization; the specific process is as follows: (a) Initialize parameters: Initialize the parameters for the policy network and the value network, and set the learning rate and the optimization objective function; (b) Update by gradient descent: After each operation, calculate the gradient according to the objective function and update the network parameters by the gradient descent method; the negative gradient direction represents the optimization path, and the model gradually approaches the optimal solution through multiple iterations; (c) Feedback of reward signal: Generate a reward signal according to the optimization effect after each operation. This reward signal is obtained by calculating the symbolic regression expression as described above, and the value function of the expression is fed back to the policy network in the model for model optimization; (4) Complete network alignment: After multiple rounds of iteration, finally complete the optimization of the expression under the guidance of the policy-value network; the optimization process includes the following steps: (a) Backtracking and cumulative reward: For the expression after each operation, accumulate the reward and update the total score, thereby providing a more accurate reference for subsequent operations; (b) Determine the optimal expression: After meeting the termination condition, output the expression with the highest score as the final optimization result; (c) Expression output and verification: After optimization, verify the generated expression to ensure its functional equivalence with the initial expression, and at the same time compare the scores before and after optimization to evaluate the optimization effect.
3. The time series analysis method based on symbolic regression according to claim 2, wherein The policy-value network is a neural network structure that shares a common backbone. As a whole, it consists of an encoding layer, a feature fusion layer, and two independent multi-layer perceptrons; the policy-value network accepts two sets of inputs: the current expression path in Monte Carlo tree search and the time series to be modeled; the backbone network first performs embedding and position encoding processing on the expression path as a symbol sequence and the time series to be modeled respectively, then extracts deep joint features through a multi-layer Transformer encoder, mixes the features, and then passes through two independent multi-layer perceptrons as a policy selector and a value estimator to input the currently selected policy and the estimated value of the current state; the policy selector outputs the probability distribution of the operation nodes, and the value estimator outputs the scoring result of the current expression state, providing a precise and intelligent decision-making basis for the expression generation and optimization process.
Citation Information
Cited By
Oil and gas reservoir production decline curve modeling method, device and equipment based on prefabricated search space symbol regression algorithm, medium and product
CN120974449A
Distributed joint debugging simulation platform construction method based on symbol regression
CN121598793A