Material science formula data set construction method for symbol regression
By automatically extracting and processing latex formulas from materials science PDF books, we generate high-quality symbolic regression datasets, addressing the problems of insufficient dataset size and diversity, and enabling effective evaluation of symbolic regression algorithms and data-driven discovery of materials science models.
Patent Information
- Application Number
- CN202510936559.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-17
AI Technical Summary
Existing material science formula datasets are small in scale and single in diversity, which increases the risk of overfitting and leads to biased evaluation results.
Automatically parse Markdown text from PDF books on materials science, extract formulas in Latex format using regular expressions, and manually verify and filter formulas that are not applicable to symbolic regression data points. Classify and rename variables, convert to Python format, and generate high-quality symbolic regression datasets.
A large-scale, high-quality symbolic regression benchmark dataset was constructed, which significantly improved the efficiency and scale of data acquisition, supported the verification and optimization of complex symbolic regression algorithms, and promoted data-driven material science model discovery.
Smart Images

Figure CN120804708A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of material science formula, and particularly relates to a material science formula data set construction method for symbolic regression. BACKGROUND
[0002] With the proposal of the Materials Genome Initiative, high-throughput experiments have developed rapidly, and the amount of data has increased exponentially. How to induce scientific laws from the growing amount of data has become a problem that scientists need to solve. In recent years, the rapid development of machine learning, especially deep learning, has given birth to a variety of new methods that can effectively reproduce known data and make reasonable inferences on new inputs. However, the black box nature of these models makes it difficult for researchers to analyze the internal mechanism of their prediction decisions. This characteristic is particularly critical in cross-disciplinary application scenarios - when domain experts (such as physics and chemistry researchers) attempt to introduce machine learning tools into their discipline, the lack of model interpretability will severely limit their practical application value.
[0003] Under this background, symbolic regression, as an important branch of interpretable machine learning, focuses on mining analytical expressions that conform to mathematical norms from data. The core goal of this technology is to build mathematical models that have both prediction accuracy and formal simplicity. It is worth noting that, although theoretically we can infinitely approach the observed data by increasing complexity, overly complex and lengthy functions often lack practical value. Therefore, the essence of symbolic regression is to seek a Pareto optimal solution between data fitting fidelity and mathematical expression simplicity.
[0004] Among them, the benchmark data set makes the performance of different symbolic regression algorithms can be objectively measured and compared, which helps researchers understand the advantages and disadvantages of their own algorithms, and promotes the improvement and innovation of algorithms. However, the existing benchmark data set for evaluating symbolic regression algorithms has the problems of insufficient data quantity, single diversity, etc., which increases the risk of overfitting and leads to biased evaluation results. SUMMARY
[0005] In view of the above defects of the prior art, the technical problem to be solved by the present application is the problem of insufficient data quantity and single diversity in the existing material science formula data construction method, which increases the risk of overfitting and leads to biased evaluation results. The present application provides a material science formula data set construction method for symbolic regression, which collects data points from formulas and constructs a benchmark database for symbolic regression, solving the problems of small size of existing formula data sets for symbolic regression and lack of formula data sets specialized in the field of material science.
[0006] To achieve the above purpose, the present application provides a material science formula data set construction method for symbolic regression, comprising the following steps:
[0007] Step 1: Analyze the PDF book of the materials science major into Markdown text;
[0008] Step 2: Extract latex-formatted equations from the Markdown text using regular expressions;
[0009] Step 3: Manually compare the extracted latex equations with the equations in the original text, and verify that the accuracy of the extracted equations is 98% through manual verification;
[0010] Step 4: Filter the latex equations to remove those that cannot be used to generate symbolic regression data points;
[0011] Step 5: Remove units and unnecessary font styles from the latex equations;
[0012] Step 6: Classify latex equations according to their latex keywords and extract algebraic equations;
[0013] Step 7: Rename variables in algebraic equations to simplify variable names and highlight the operational properties of the algebraic equations;
[0014] Step 8: Convert latex equations to python equations;
[0015] Step 9: Extract explicit equations from python equations;
[0016] Step 10: Define the value range of the independent variable in the explicit equation, uniformly sample the independent variable, and calculate the dependent variable according to the explicit equation. Batch the process, and each equation corresponds to a set of data points;
[0017] Step 11: Use the data_profiling tool to generate an html report for each equation's data points, to visually see the number of variables, variable values, etc. for each set of data, making it easier for users to understand the data set.
[0018] Further, in step 1, the PDF book of the materials science major is analyzed into Markdown text, specifically using the open-source tool minerU to parse the PDF book of the materials science major into markdown format.
[0019] Further, in step 2, latex-formatted equations are extracted from the Markdown text using regular expressions, where the latex-formatted equations contain an equal sign, and both sides of the equal sign cannot be empty. According to the characteristics of latex-formatted equations, regular expressions are used to extract inline equations and inline equations.
[0020] Further, the latex formula is screened, and the formula which cannot be used to generate symbol regression data points is cleaned up, including:
[0021] In combination with the characteristics of latex formula set, the screening rules are formulated;
[0022] According to the screening rules, the regular expression is designed, and the formula which cannot be used to generate symbol regression data points is cleaned up.
[0023] Further, the screening rules include:
[0024] Screening equations using abstract functions;
[0025] Screening equations without or with only one independent variable;
[0026] Screening equations containing N-term expression summation or multiplication operation;
[0027] Screening equations with more than one equal sign;
[0028] Screening equations containing approximately equal sign;
[0029] Screening chemical equations.
[0030] Further, according to the latex keywords of different types of formulas, the formulas are classified, and the algebraic formulas are extracted, including the following steps:
[0031] The latex formula set is divided into algebraic equation, differential equation, integral equation, tensor / matrix equation;
[0032] According to the latex keywords of different types of equations, the second regular expression is designed;
[0033] The algebraic formula is extracted for subsequent processing.
[0034] Further, the variables in the algebraic formula are renamed, the variable name is simplified, and the operation property of the formula is highlighted, including:
[0035] For each formula, the number of variables n in the formula is counted;
[0036] According to the order of variables appearing from left to right in the formula, the variables are standardized and renamed, and the naming standard is {x1,x2,...,xn}, so as to weaken the intuitive influence of complex variable name on the formula and highlight the operation in the formula.
[0037] Further, the explicit formula is extracted from the python algebraic formula, and the explicit formula refers to the formula in which there is only one variable on one side of the equal sign, and the value of the variable depends on the calculation of the remaining variables on the other side of the equation.
[0038] Further, the step of determining whether to display the formula is as follows:
[0039] Check if the formula contains an equal sign =, this step filters out the formula without the equal sign;
[0040] Check if the right side is a single variable, this step is beneficial to unify the variable on the left side;
[0041] Check if the expression contains illegal symbols;
[0042] The formula that meets the above three steps is added to the display formula file.
[0043] Further, step 10, define the value range of the independent variable in the display formula, uniformly collect the independent variable, and calculate the dependent variable according to the display formula, and batch the process, each formula corresponds to a group of data points; Including the following steps:
[0044] Extract the variables in the formula as the table header of the data points;
[0045] The generated data points meet the requirements of the conditional expression;
[0046] Generate data points.
[0047] Technical effects
[0048] The material science formula data set construction method for symbolic regression provided by the application has the following technical effects:
[0049] 1. Automatically construct high-quality symbolic regression benchmark data sets: by automatically extracting, cleaning and classifying formulas from material science books, and accurately collecting their corresponding data points, the first large-scale, structured symbolic regression special data set for the material science field is automatically constructed, solving the key bottleneck problem of lack of high-quality, targeted benchmark data sets in this field.
[0050] 2. Significantly improve data acquisition efficiency and scale: overcome the low efficiency and limited scale of traditional manual collection and arrangement of material formulas and data points. This method uses an automated processing flow to efficiently mine potential formulas and data from a large number of literature, greatly improving the speed and scalability of data set construction, providing a sufficient data foundation for training and verifying complex symbolic regression models.
[0051] 3. Support verification and optimization of complex symbolic regression algorithms: the constructed data set contains typical formulas and real data points in materials science with different complexities, providing a reliable basis for evaluating and comparing the performance, robustness and generalization ability of different symbolic regression algorithms (such as genetic programming, neural networks, sparse regression, etc.) in the specific field of materials science.
[0052] 4. Promote data-driven model discovery in the field of materials science: lay the foundation for data-driven materials model discovery. Symbolic regression models trained on this dataset are expected to rediscover known physical laws directly from experimental or simulation data, or to mine potential, unexpressed quantitative relationships between material variables, accelerating the computational design and performance prediction of new materials.
[0053] The concept, specific structure and generated technical effects of the present application will be further described below in combination with the drawings, so as to fully understand the purpose, features and effects of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0054] Figure 1 is a flowchart of a material science formula dataset construction method for symbolic regression according to a preferred embodiment of the present application;
[0055] Figure 2 is a data point report example schematic diagram of a material science formula dataset construction method for symbolic regression according to a preferred embodiment of the present application. DETAILED DESCRIPTION
[0056] In order to make the technical problems to be solved by the present application, the technical solutions and beneficial effects more clear and explicit, the present application will be further described in detail below in combination with the drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0057] In the following description, specific details such as specific internal procedures, techniques, etc. are presented in order to thoroughly understand the embodiments of the present application, but it should be clear to those skilled in the art that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits and methods are omitted to avoid unnecessary details that hinder the description of the present application.
[0058] As shown in Figure 1 The present application provides a material science formula dataset construction method for symbolic regression, characterized in that it comprises the following steps:
[0059] Step 1: Analyze the PDF books of materials science into Markdown text; specifically, analyze the PDF books of materials science into markdown format using the open source tool minerU. Among them, the books of materials science include books in the fields of materials thermodynamics, dynamics, etc.
[0060] Step 2, extract latex formatted formulas from Markdown text using regular expressions; where latex formatted formulas contain equal signs, and both sides of the equal sign cannot be empty, according to the characteristics of latex formatted formulas, extract inline formulas and inline formulas using regular expressions, the following is the regular expression used to extract formulas from the text.
[0061] double_dollar_formulas = re.findall(r'\$\$(.*?)\$\$', text, re.DOTALL)
[0062] single_dollar_formulas = re.findall(r'\$(.*?)\$(?!\$)', text, re.DOTALL)
[0063] Step 3, manually compare the extracted latex formulas with the formulas in the original text, and verify that the overall accuracy of the extracted formulas is 98%. Take the book "Composites And Metamaterials" as an example, convert it to markdown format using the minerU tool, and extract 557 inline and interline formulas from the book using regular expressions. After manual verification, it is found that the number of extracted latex formulas consistent with the original text is 553, and the accuracy of formula extraction is 99.28%. Then manually delete the incorrectly extracted formulas to ensure that all formulas in the subsequent data set are extracted from the book without errors.
[0064] Step 4, screen latex formulas and clean up formulas that cannot be used to generate symbolic regression data points; including:
[0065] Step 4.1 Develop screening rules based on latex formula set characteristics; the screening rules include the following:
[0066] Screening equations using abstract functions;
[0067] Screening equations with no or only one independent variable;
[0068] Screening equations containing N-term summation or multiplication operations;
[0069] Screening equations with more than one equal sign;
[0070] Screening equations containing approximately equal signs;
[0071] Screening chemical equations.
[0072] Step 4.2 Regular expressions are designed according to the screening rules of 4.1 to clean up the formulas that cannot be used to generate symbolic regression data points. Each screening rule corresponds to a regular expression as follows:
[0073] Abstract function: r'\b\w+\([^)]*,' such as formula f(x, y) = x^2 + y^2 is screened out;
[0074] Equation with no or only one variable: r'^\s*\w+\s*=\s*[^a-zA-Z]+$' such as formula C = 3.14 is screened out;
[0075] Sum or product operation containing N-term expression: r'(\\sum|\\prod|\\bigcup|\\bigcap|sum\(' such as formula
[0076] is screened out;
[0077] Number of equal signs > 1: r'(?<![<=!])=(?!=)' such as formula x = y = z + 1 is screened out;
[0078] Contains approximately equal sign: r'\\approx|≈' such as formula π ≈ 3.14 is screened out;
[0079] Chemical equation:
[0080]
[0081] such as chemical equation
[0082] 2H2(g) + O2(g) → 2H2O(l)
[0083] CH4 + 2O2 → CO2 + 2H2O
[0084] is screened out;
[0085] Step 5, call large language model, for each latex formula, delete the units and redundant font style settings in it, such as delete before: F = m·a where (kg), a (m / s 2 ), delete after: F = m·a.
[0086] Step 6, according to the latex keywords of different types of formulas, classify the latex formulas and extract algebraic formulas; including the following steps:
[0087] Step 6.1, combine the latex formula set after step 5 with the formula category, and expect to be divided into four categories: algebraic formula, differential formula, integral formula, and tensor / matrix formula;
[0088] Step 6.2, design the second regular expression according to the latex keyword of different types of equations, where each formula type corresponds to the second regular expression as follows:
[0089] Algebraic formula:
[0090]
[0091] Differential formula:
[0092]
[0093]
[0094] Integral formula:
[0095]
[0096] Matrix / tensor formula:
[0097]
[0098] Step 6.3, extract the algebraic formula for subsequent processing.
[0099] Step 7, rename the variables in the algebraic formula, simplify the variable name, and highlight the operation properties of the algebraic formula; specifically including:
[0100] Step 7.1, for each formula, count the number of variables n in it;
[0101] Step 7.2, according to the order of the variables appearing from left to right in the formula, standardize the renaming of the variables, the naming standard is: {x1,x2,...,xn}, so as to weaken the intuitive influence of complex variable name on the formula, highlight the operation in the formula, for example
[0102]
[0103] Convert to
[0104] x0 2 = x1 2 x2 2 + x3 2 x4 2 + x5 2 x6 2 -(x1x2 2 + x3x4 2 + x5x6 2 ) 2
[0105] Step 8, convert latex formula to python formula; specifically including
[0106] Step 8.1: Exponent conversion, convert the exponential form a^{{b}} in latex format to the power-of form a**b in python format, for example x0^{{e}}→x0**e, and for multi-character exponents, add parentheses: x^{{2*k+1}}→x**(2*k+1);
[0107] Step 8.2: Fraction processing, convert the fraction form \\frac{{a}}{{b}} in latex format to the python form (a / b), and pay attention to nested fraction processing: \\frac{{e^{{-\\frac{{x}}{{2}}}}}}{{3}}→
[0108] (e**(-x / 2)) / 3;
[0109] Step 8.3: Explicit multiplication, add the omitted multiplication sign, such as 3x→3*x, adjacent variables: xy→x*y, and variables after parentheses: (a+b)c→(a+b)*c;
[0110] Step 8.4: Greek letter conversion, support recognizing common letters and removing backslashes (escape characters), for example \\alpha→alpha, \\beta→beta, \\pi→pi, etc.
[0111] Step 8.5: Special operator symbol processing, remove backslashes, for example \\sqrt{{ax}}→sqrt(a*x), \\log→log, \\sin→sin, e^x→exp(x);
[0112] Step 8.6: While converting the formula format, keep the operation priority.
[0113] The following is a conversion example:
[0114]
[0115] Step 9: Extract explicit formulas from python algebraic formulas, explicit formulas are formulas where one side of the equation has only one variable, and the value of this variable depends on the calculation of the remaining variables on the other side of the equation, for example x0=x1*(x2-x3)*x4 is an explicit formula;
[0116] Step 9.1: Check if the formula contains the equal sign =, this step filters out formulas without the equal sign;
[0117] Step 9.2: Check if the right side is a single variable, this step helps to unify the variables on the left side;
[0118] Step 9.3: Check if the expression contains illegal symbols (not variables or standard mathematical symbols);
[0119] Step 9.4: If the above three steps are satisfied, add the formula to the display formula file. The following is the algorithm to determine if it is a display formula:
[0120] def is_explicit_equation(equation):
[0121] #1. Whether it contains an equal sign
[0122] if '=' not in equation:
[0123] return False
[0124] #2. Split left and right sides
[0125] left_side, right_side = equation.split('=', 1)
[0126] left_side = left_side.strip()
[0127] #3. Check if the left side is in the form of a variable: x0, x1,...
[0128] pattern = r'^x\d+$'
[0129] #4. Check if the right side contains mathematical operators (determine if it is an expression)
[0130] if not contains_math_operators(right_side):
[0131] return False
[0132] #5. Check if it contains illegal characters
[0133] if contains_invalid_symbols(equation):
[0134] return False
[0135] #6. Satisfy the conditions: left side is a variable, right side has mathematical structure, no illegal symbols.
[0136] return bool(re.match(pattern, left_side))
[0137] Step 10. Batch generate data points for each formula. The dataset consisting of formula-data point pairs can be used as the basis for subsequent symbolic regression experiments.
[0138] Step 10.1: Extract variables in the formula as the header of data points.
[0139] Step 10.1.1: Split the formula into left side (dependent variable) and right side (expression).
[0140]
[0141] Step 10.1.2: Extract all independent variables from the right side using regular expressions (format x0, x1, x2, etc.), def extract_variables(expression):
[0142] return sorted(set(re.findall(r'x\d+', expression))).
[0143] Step 10.2: The generated data points meet the requirements of the conditional expression.
[0144] Step 10.2.1: The expression that needs to consider legality is called the conditional expression.
[0145] Step 10.2.2: Determine the square root and even root function: the value of the conditional expression must be ≥ 0.
[0146] Step 10.2.3: Determine the logarithmic function: the value of the conditional expression must be > 0.
[0147] Step 10.2.4: Determine the fraction and division: the value of the conditional expression must ≠ 0.
[0148] Step 10.2.5: Determine the inverse trigonometric function (arcsin, arccos): the value of the conditional expression must ∈ [-1, 1].
[0149] Step 10.2.6: Determine the tangent and cotangent function: the value of the conditional expression must ≠ k∈N;
[0150] Step 10.2.7: Determine the hyperbolic function: the value of the conditional expression must ∈ [-5, 5].
[0151] Step 10.2.8: Determine the inverse hyperbolic function acosh: the value of the conditional expression must be ≥ 1.
[0152] Step 10.2.9: Determine the power function: the base conditional expression must be ≥ 0, and the exponent conditional expression must ∈ Z.
[0153] The following is the algorithm of 10.2.1-10.2.9 above:
[0154]
[0155]
[0156] Step 10.3: Generate data points;
[0157] Step 10.3.1: Parse the expression and convert the formula to a function that can be used for calculation using sympy;
[0158] Step 10.3.2: Randomly generate independent variable values according to the set value range (initially (1, 5));
[0159] Step 10.3.3: Return to 10.2 to determine whether the generated set of values passes the domain limit. If it passes, it is considered a valid point. If it does not pass, discard the set of values and return to 10.3.2 to continue generating;
[0160] Step 10.3.4: Repeat the operation until the number of valid points collected reaches the target number or exceeds the maximum number of attempts;
[0161] Step 10.3.5: Form a DataFrame with valid points and reorder the columns, with the dependent variable at the beginning.
[0162] The following is the algorithm of 10.3.1-10.3.5:
[0163] The input equation is the displayed formula, num_points is the number of data points to be collected, max_attempts is the maximum number of attempts to prevent the point from being trapped in a dead loop due to difficult conditions, and var_range is the range of independent variable values.
[0164] The output is the data point df
[0165]
[0166]
[0167] Step 11, for each material science formula completed in step 10, use the data_profiling tool (ydata-profiling) to automatically analyze and visualize the characteristics of each formula's data point set, generating detailed HTML reports, significantly reducing the data usage threshold.
[0168] Figure 2are 100 data points sampled from the formula x0 = (4 / 6) * (x1 / x2) * x3, with the independent variables x1-x3 uniformly sampled from [2, 5]. Figure 2 As can be seen from the table, the minimum value of the dependent variable x0 is 1.330382, the maximum value is 3.3665268, and the Distinct value is 100%, which means that there is no repetition of the value of x0.
[0169] In this embodiment, we use the proposed method to construct a high-quality symbolic regression benchmark dataset and conduct a detailed analysis, covering typical formulas and their corresponding data points in the field of materials science. Through the automated process, we successfully overcome the low efficiency of manually collecting and organizing material formulas and data points, significantly improving the efficiency and scale of data acquisition, providing a solid data foundation for the training and verification of symbolic regression models.
[0170] The construction of the dataset not only supports the verification and optimization of complex symbolic regression algorithms, but also promotes data-driven model discovery in the field of materials science. The symbolic regression model trained using this dataset can effectively rediscover known physical laws from experimental or simulation data, or mine potential, not explicitly stated quantitative relationships between material variables. This achievement verifies the effectiveness of the proposed method, provides strong support for the computational design and performance prediction of new materials, and accelerates the research and innovation process in the field of materials science.
[0171] The above describes the preferred embodiments of the present application in detail. It should be understood that those skilled in the art can make many modifications and changes without creative labor according to the concept of the present application. Therefore, any technical solutions obtained by logical analysis, reasoning or limited experiments based on the prior art according to the concept of the present application shall be within the protection scope defined by the claims.
Claims
1. A method for constructing a materials science formula dataset for symbolic regression, characterized in that: The following steps are involved: Step 1: Parse the PDF books on materials science into Markdown text; Step 2: Use regular expressions to extract latex format formulas from Markdown text; Step 3: Manually compare the extracted latex formula with the formula in the original text. After manual verification, the accuracy of the extracted formula is 98%; Step 4: Filter the latex formulas and remove the formulas that cannot be used to generate symbolic regression data points; Step 5: Remove the units and redundant font style settings in the latex formula; Step 6: Classify the latex formulas according to the latex keywords of different types of formulas and extract the algebraic formulas; Step 7: Rename the variables in the algebraic formula to simplify the variable names and highlight the operational properties of the algebraic formula; Step 8, convert the latex formula into a python formula; Step 9, extract the explicit formula from the Python formula; Step 10: Define the value range of the independent variable in the display formula, evenly collect the independent variable, and calculate the dependent variable according to the display formula. The process is batched and each formula corresponds to a set of data points. Step 11: Use the data_profiling tool to generate an HTML report for each formula data point to clearly see the number of variables, variable values, and other information for each set of data, making it easier for users to understand the data set.
2. The method for constructing a materials science formula dataset for symbolic regression according to claim 1, wherein: Step 1: Parse the PDF books on materials science into Markdown text. Specifically, use the open source tool minerU to parse the PDF books on materials science into markdown format.
3. The method for constructing a materials science formula dataset for symbolic regression according to claim 1, wherein: Step 2: Use regular expressions to extract formulas in latex format from Markdown text. Formulas in latex format contain equal signs, and both sides of the equal sign cannot be empty. Based on the characteristics of formulas in latex format, use regular expressions to extract inter-line formulas and in-line formulas.
4. The method for constructing a materials science formula dataset for symbolic regression according to claim 1, wherein: Filter the latex formulas and remove the formulas that cannot be used to generate symbolic regression data points, including: Combine the characteristics of latex formula set to formulate screening rules; Design regular expressions based on the screening rules to eliminate formulas that cannot be used to generate symbolic regression data points.
5. The method for constructing a material science formula dataset for symbolic regression according to claim 4, wherein: The screening rules include: Filter equations using abstract functions; Screen equations with no or only one independent variable; Filter equations containing summation or multiplication operations of N expressions; Filter equations with the number of equal signs greater than 1; Filter equations containing approximately equal signs; Filter chemical equations.
6. The method for constructing a material science formula dataset for symbolic regression according to claim 1, wherein: According to the latex keywords of different types of formulas, the formulas are classified and the algebraic formulas are extracted. The specific steps include: Divide the latex formula set into algebraic equations, differential equations, integral equations, and tensor / matrix equations; Design the second regular expression based on the latex keywords of different types of equations; Extract the algebraic formula for subsequent processing.
7. The method for constructing a materials science formula dataset for symbolic regression according to claim 1, wherein: Rename the variables in algebraic formulas to simplify the variable names and highlight the operational properties of the formulas, including: For each formula, count the number of variables n; Variables are standardized and renamed according to the order in which they appear from left to right in the formula. The naming standard is: {x1,x2,...,xn}, which weakens the intuitive impact of complex variable names on the formula and highlights the operations in the formula.
8. The method for constructing a materials science formula dataset for symbolic regression according to claim 1, wherein: Extract explicit formulas from Python algebraic formulas. Explicit formulas are formulas in which there is only one variable on one side of the equal sign, and the value of this variable depends on the calculation of the remaining variables on the other side of the equation.
9. The method for constructing a material science formula dataset for symbolic regression according to claim 8, wherein: The steps to determine whether it is a display formula are as follows: Check whether the formula contains an equal sign =. This step will filter out formulas without an equal sign. Check whether the right side is a single variable. This step helps to uniformly place the variables on the left side. Check whether the expression contains illegal symbols; The formula that satisfies the above three steps is added to the display formula file.
10. The method for constructing a material science formula dataset for symbolic regression according to claim 1, wherein: Step 10, defining the value range of the independent variable in the display formula, uniformly collecting the independent variable, and calculating the dependent variable according to the display formula, and batching the process so that each formula corresponds to a set of data points; including the following steps: Extract the variables in the formula as the header of the data point; The generated data points meet the requirements of the conditional expression; Generate data points.