A Method for Optimizing and Reconstructing Bridge Scour Formula Based on Symbolic Regression

By optimizing and reconstructing the bridge scour formula based on symbolic regression, and combining prior knowledge of pier scour with physical relationships, a more accurate and universal scour depth prediction model is generated using a genetic programming algorithm. This solves the problem of insufficient accuracy and universality of existing pier scour formulas, and improves the safety and economy of bridge design.

CN118070383BActive Publication Date: 2025-12-02SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410173733.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-07
Publication Date
2025-12-02
Estimated Expiration
2044-02-07

AI Technical Summary

Technical Problem

Existing bridge pier scour formulas are insufficient in terms of accuracy and universality, leading to safety hazards and economic problems in bridge design.

Method used

A bridge scour formula optimization and reconstruction method based on symbolic regression is adopted. Combining prior knowledge of pier scour and physical relationships, a genetic programming algorithm is used to learn the complex relationship between scour depth and influencing factors from measured data, generating a formula with good predictive performance and easy understanding.

Benefits of technology

It improves the accuracy and generalization of bridge scour depth prediction, enhances the safety and economy of bridge design, reduces damage accidents caused by scour, and saves maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118070383B_ABST
    Figure CN118070383B_ABST
Patent Text Reader

Abstract

This invention discloses a method for optimizing and reconstructing bridge scour formulas based on symbolic regression, comprising the following steps: obtaining statistical data on bridge scour, which are used as training and testing sets respectively; determining the initial structure of the bridge scour formula; firstly determining candidate operators, then introducing a group of functions with significant physical meaning and statistical relationships in scour to determine candidate operational variables; using a symbolic regression method based on genetic programming to generate an optimized formula using the data from the training set; adjusting parameters to generate a group of candidate formulas; evaluating the formula performance based on the testing set, and selecting the formula form with the best prediction effect and the simplest formula structure. This invention, based on existing standard formulas, uses a symbolic regression method based on genetic programming algorithms, integrating machine learning methods with nonlinear representation capabilities to more accurately mine rich features in the data. Simultaneously, by combining prior knowledge of scour and physical relationships, it can improve the predictive performance and generalization of the scour calculation formula.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of civil engineering and water conservancy technology, and in particular relates to a method for optimizing and reconstructing bridge scour formulas based on symbolic regression. Background Technology

[0002] Localized scour of bridge piers is a phenomenon where water flow erodes the surrounding sediment due to the obstruction of the pier. It is a complex interaction process between water flow, sediment, and the structure. According to relevant literature, in recent years, approximately 50% of bridge failures have been caused by scour damage to the bridge foundation structure and related hydraulic effects, resulting in huge economic losses. Therefore, estimating the depth of localized scour is an important aspect of bridge safety assessment. Currently, scholars both domestically and internationally have proposed many formulas for calculating the depth of localized scour of bridge piers, based on various influencing factors. Most of these are empirical or semi-empirical formulas based on fitting field observations and experimental data. The calculation results are often only relatively accurate under specific conditions, exhibiting poor stability and generalization. Some formulas yield overly conservative results, which are detrimental to the economic efficiency of engineering design; others yield results that are too small, which can pose serious safety hazards to bridge scour design. Summary of the Invention

[0003] Purpose of the invention: To address the issues of insufficient accuracy and universality of existing bridge pier scour formulas, the purpose of this invention is to provide a method for optimizing and reconstructing bridge scour formulas based on symbolic regression. This method, based on existing standard formulas, combines prior knowledge of bridge pier scour and physical relationships, and uses a symbolic regression method based on genetic programming algorithms to learn from a large amount of measured data, discovering the complex relationship between scour depth and various influencing factors, and obtaining a formula with good predictive performance that is easy to understand and use.

[0004] Technical Solution: To achieve the above objectives, this invention discloses a method for optimizing and reconstructing bridge scour formulas based on symbolic regression, comprising the following steps:

[0005] (1) Obtain statistical data on bridge scour, and use it as the training set and test set respectively;

[0006] (2) Determine the initial structure for the bridge scour formula;

[0007] (3) First, determine the candidate operators, and then introduce a group of functions with significant physical meaning and statistical relationship in scouring to determine the candidate operation variables;

[0008] (4) Use the symbolic regression method based on genetic programming to generate the optimization formula using the data in the training set;

[0009] (5) Adjust parameters to generate a group of alternative formulas.

[0010] (6) Evaluate the formula effect based on the test set and select the formula form with the best prediction effect and the simplest formula structure.

[0011] In step (1), statistical data is obtained from historical literature and used as the training and test sets, respectively, ensuring that the references of the training and test sets do not overlap and that the data are independent of each other; the statistical data includes:

[0012] Data{b, L, θ, V1, y1, d 50 , σ g y s}

[0013] b is the width of the bridge pier, L is the length of the bridge, θ is the angle of attack of the water flow, V1 is the average velocity of the water flow upstream of the bridge pier, y1 is the water depth upstream of the bridge pier after typical scouring, and d 50 σ represents the particle size of the sediment. g For the uniformity of sediment, y s This represents the local scour depth of the bridge pier.

[0014] Preferably, in step (2), a new calculation formula is generated by optimizing and reconstructing the CSU equation structure adopted by the US standard:

[0015]

[0016] In the formula: y s y1 represents the local scour depth of the bridge pier, y2 represents the upstream water depth of the bridge pier after general scour, K1 is the pier shape correction factor, K2 is the flow angle of attack correction factor, K3 is the riverbed condition correction factor, b is the pier width, and F is the bridge pier width. r For the Froude number of the water flow upstream of the bridge pier, Where V1 is the average velocity of the water flow upstream of the bridge pier, and g is the acceleration due to gravity.

[0017] Take the opposite side of the formula:

[0018]

[0019] The initial structure of the optimized bridge scour formula is as follows:

[0020]

[0021] Furthermore, in step (3), common operators are used to determine the candidate operators, including:

[0022] Operator}+, -, ×, ÷, ^2, ^3, sin, cos}

[0023] Based on the bridge scour test results and correlation analysis, the following parameters were found to have a significant statistical relationship with the scour depth:

[0024] Parameter{b n d 50 , σ g ,y1,V1}

[0025] Among them, b n For the effective width, b n = bcosθ + Lsinθ, where b is the width of the pier, L is the length of the bridge, θ is the angle of attack of the water flow, and d 50 σ represents the particle size of the sediment. g To represent the uniformity of sediment, y1 is the water depth upstream of the bridge pier after typical scouring, and V1 is the average velocity of the water flow upstream of the bridge pier.

[0026] Local scour depth y of bridge pier s The following relationship exists with Parameter:

[0027]

[0028]

[0029]

[0030]

[0031]

[0032] Based on the analysis of the scouring test results, the above function set can be adopted as: g1(x, y) = x + y, g2(x) = sin(ax), g3(x) = b x (0 < b < 1), g4(x) = x c-x g5 = log d (x+1)(d>1); Parameters a, b, c, and d should be determined based on the dataset.

[0033] Candidate operands include:

[0034]

[0035] Furthermore, step (4) specifically includes the following steps:

[0036] (4.1) Based on the candidate operators and candidate operation variables obtained in step (3), generate a random initial expression population containing operators, constants, and variables. The population size i is defined as the number of individuals in the population. The initialization process needs to be repeated until i expressions are generated: Original{E1, E2, ..., E...} i};

[0037] (4.2) Based on the training set data, the fitness function is used to evaluate the population performance, and the top n expressions with the best performance are selected as the maternal population for the next generation: Parents{E1, E2, ..., E... n}, n = 0.8i;

[0038] (4.3) Genetic programming algorithm is used to perform genetic iteration on the selected population based on the training set data to generate the subpopulation with the best fitness. It is divided into three stages: crossover stage, mutation stage and replacement stage.

[0039] (4.4) Calculate the new population New{E1, E2, ..., E} based on the training set data. n Fitness, and the new population New}E1, E2, ..., E n Sort the parent populations (E1, E2, ..., E...) in ascending order of fitness, and use them as the maternal lineage for the next iteration. n}; Return to step (4.3), and after reaching the set number of iterations, extract the expression with the best fitness in the population as the target formula E. target .

[0040] Preferably, the fitness calculation formula for each expression in step (4.2) is as follows:

[0041]

[0042] Where N is the total number of samples, It is individual E i The predicted value y for the j-th sample i It is the corresponding true value.

[0043] Furthermore, in step (4.3), the crossover stage refers to randomly selecting the subtrees of two maternal populations and exchanging them to generate two new offspring trees; the mutation stage refers to randomly mutating the nodes of the individual expression tree while keeping the tree structure unchanged; and the replacement stage refers to regenerating a new tree to replace the current tree.

[0044] Furthermore, in step (4.4), the target formula E target for:

[0045]

[0046] Preferably, the adjustable parameter set P in step 5 is:

[0047] P = {num, depth, p} variation p crossover p mutation p replacement ,run}

[0048] Where num is the number of initialization expressions, depth is the depth of the tree, and p variation p represents the probability that a node is a variable when the tree is being spanned. crossover p represents the probability of intersection between different expressions. mutation p represents the probability of node mutation. replacement The probability of generating a new tree to replace the current tree is given by `run`, which is the number of genetic iterations.

[0049] Furthermore, in step (6), the alternative formula group is calculated. MSE and R 2 Rank the prediction results:

[0050]

[0051]

[0052] Where N is the total number of samples, It is individual E i The predicted value y for the j-th sample i It corresponds to the actual value. It is the average of the true values; the smaller the MSE, the better the prediction effect; R 2 The closer the result is to 1, the better the prediction effect.

[0053] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: Based on the existing standard formula, the present invention combines prior knowledge of bridge scour and physical relationships, and uses a symbolic regression method based on genetic programming algorithm to learn from a large amount of measured data, discovering the complex relationship between scour depth and various influencing factors, and obtaining a formula with good predictive performance that is easy to understand and use; The present invention overcomes the shortcomings of traditional statistical methods in accurately mining the rich features in the data and the limited amount of data, and the optimized formula is more accurate and has better generalization than the traditional formula. Attached Figure Description

[0054] Figure 1 This is a flowchart of the present invention;

[0055] Figure 2 This is a flowchart of the symbolic regression method based on genetic algorithm in this invention;

[0056] Figure 3 The scouring depth y in this invention s With sediment particle size d 50 A schematic diagram of the changing relationship;

[0057] Figure 4 The scouring depth y in this invention s With uniformity of sediment σ g A schematic diagram of the changing relationship;

[0058] Figure 5 The scouring depth y in this invention s A schematic diagram showing the relationship between the water flow depth y1 and the water flow depth.

[0059] Figure 6 The scouring depth y in this invention s A schematic diagram showing the relationship between the change in water flow velocity V1 and the flow velocity V1. Detailed Implementation

[0060] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0061] Symbolic regression is a supervised learning method that attempts to discover a hidden mathematical formula to predict a target variable using feature variables. The mainstream algorithm for solving symbolic regression problems is genetic programming. Genetic programming algorithms are a class of intelligent computational optimization methods that achieve autonomous computer programming by simulating the "survival of the fittest" evolutionary phenomenon in nature. The encoding mechanism of genetic programming algorithms is relatively flexible and suitable for representing the structure of evolutionary functions. Furthermore, the algorithm's search process is based on random heuristic search, eliminating the need for complex optimization models, thus giving it strong versatility.

[0062] By optimizing and reconstructing existing scour formulas using symbolic regression, a more accurate and universally applicable scour depth prediction model can be constructed by combining prior knowledge of scour with physical relationships, based on existing empirical formulas. This scour depth prediction model can better adapt to different river conditions, pier shapes, and hydrodynamic effects. The resulting formula not only improves predictive performance but is also easier for engineers to understand and use, thereby enhancing the safety and economy of bridge design.

[0063] In summary, applying symbolic regression to the study of local scour depth in bridge piers not only compensates for the shortcomings of existing empirical formulas, but also provides strong technical support for bridge scour protection design and maintenance. This helps reduce bridge damage accidents caused by scour, ensures traffic safety, and saves maintenance costs, thus having significant practical implications and broad application prospects in the field of bridge engineering.

[0064] like Figure 1 As shown, the present invention provides a method for optimizing and reconstructing bridge scour formulas based on symbolic regression, comprising the following steps:

[0065] (1) Obtain statistical data on bridge scour, and use it as the training set and test set respectively;

[0066] In step (1), statistical data are obtained from historical literature and used as the training set and test set, respectively, ensuring that the references of the training set and test set do not overlap and that the data are independent of each other; the statistical data includes:

[0067] Data{b, L, θ, V1, y1, d 50 , σ g y s )

[0068] b is the width of the bridge pier, L is the length of the bridge, θ is the angle of attack of the water flow, V1 is the average velocity of the water flow upstream of the bridge pier, y1 is the water depth upstream of the bridge pier after typical scouring, and d 50 σ represents the particle size of the sediment. g For the uniformity of sediment, y s This represents the local scour depth of the bridge pier.

[0069] (2) Determine the initial structure for the bridge scour formula;

[0070] In step (2), a new calculation formula is generated by optimizing and reconstructing the CSU equation structure adopted by the US standard:

[0071]

[0072] In the formula: y s y1 represents the local scour depth of the bridge pier, y2 represents the upstream water depth of the bridge pier after general scour, K1 is the pier shape correction factor, K2 is the flow angle of attack correction factor, K3 is the riverbed condition correction factor, b is the pier width, and F is the bridge pier width. r For the Froude number of the water flow upstream of the bridge pier, Where V1 is the average velocity of the water flow upstream of the bridge pier, and g is the acceleration due to gravity.

[0073] Take the opposite side of the formula:

[0074]

[0075] The initial structure of the optimized bridge scour formula is as follows:

[0076]

[0077] (3) First, determine the candidate operators, and then introduce a group of functions with significant physical meaning and statistical relationship in scouring to determine the candidate operation variables;

[0078] like Figure 2 As shown, in step (3), common operators are used to determine candidate operators, including:

[0079] Operator{+,-,×,÷,^2,^3,sin,cos}

[0080] Based on the bridge scour test results and correlation analysis, the following parameters were found to have a significant statistical relationship with the scour depth:

[0081] Parameter{b n d50 , σ g ,y1,V1}

[0082] Among them, b n For the effective width, b n = bcosθ + Lsinθ, where b is the width of the pier, L is the length of the bridge, θ is the angle of attack of the water flow, and d 50 σ represents the particle size of the sediment. g To represent the uniformity of sediment, y1 is the water depth upstream of the bridge pier after typical scouring, and V1 is the average velocity of the water flow upstream of the bridge pier.

[0083] like Figure 3 , Figure 4 , Figure 5 and Figure 6 As shown, the local scour depth y of the bridge pier s The following relationship exists with Parameter:

[0084]

[0085]

[0086]

[0087]

[0088]

[0089] Based on the analysis of the scouring test results, the above function set can be adopted as: g1(x, y) = x + y, g2(x) = sin(ax), g3(x) = b x (0 < b < 1), g4(x) = x c-x g5 = log d (x+1)(d>1); Parameters a, b, c, and d should be determined based on the dataset.

[0090] This invention proposes candidate operands including:

[0091]

[0092] (4) Using the genetic programming-based symbolic regression method, the optimization formula is generated using the training set data. The specific steps are as follows:

[0093] (4.1) Based on the candidate operators and candidate operation variables obtained in step (3), generate a random initial expression population containing operators, constants, and variables. The population size i is defined as the number of individuals in the population. The initialization process needs to be repeated until i expressions are generated: Original{E1, E2, ..., E...} i};

[0094] (4.2) Based on the training set data, the fitness function is used to evaluate the population performance, and the top n expressions with the best performance are selected as the maternal population for the next generation: Parents{E1, E2, ..., E... n}, n = 0.8i;

[0095] The fitness calculation formula for each expression is as follows:

[0096]

[0097] Where N is the total number of samples, It is individual E i The predicted value y for the j-th sample i It is the corresponding true value;

[0098] (4.3) Based on the training set data, a genetic programming algorithm is used to perform genetic iteration on the selected population to generate the subpopulation with the best fitness, which is divided into three stages:

[0099] Crossover: Randomly select the subtrees of two maternal populations and swap them to generate two new offspring trees;

[0100] Mutation: Randomly mutate the nodes of the individual expression tree while maintaining the tree structure;

[0101] Replacement: Regenerate a new tree to replace the current tree;

[0102] (4.4) Calculate the new population New{E1, E2, ..., E} based on the training set data. n Fitness, and the new population New{E1, E2, ..., E n Sort the parent populations (E1, E2, ..., E...) in ascending order of fitness, and use them as the maternal lineage for the next iteration. n Return to step (4.3), and after reaching the set number of iterations, extract the expression with the best fitness from the population as the target formula E. target ;

[0103] Target Formula E target as follows:

[0104]

[0105] (5) Adjust parameters to generate a group of alternative formulas.

[0106] In step 5, the adjustable parameter set P is:

[0107] P = {num, depth, p} variation p crossover pmutation p replacement ,run}

[0108] Where num is the number of initialization expressions, depth is the depth of the tree, and p variation p represents the probability that a node is a variable when the tree is being spanned. crossover p represents the probability of intersection between different expressions. mutation p represents the probability of node mutation. replacement The probability of generating a new tree to replace the current tree is given by `run`, which is the number of genetic iterations.

[0109] (6) Evaluate the formula effect based on the test set and select the formula form with the best prediction effect and the simplest formula structure.

[0110] In step (6), the alternative formula group is calculated. MSE and R 2 Rank the prediction results:

[0111]

[0112]

[0113] Where N is the total number of samples, It is individual E i The predicted value y for the j-th sample i It corresponds to the actual value. It is the average of the true values; the smaller the MSE, the better the prediction effect; R 2 The closer the result is to 1, the better the prediction effect.

Claims

1. A method for optimizing and reconstructing bridge scour formulas based on symbolic regression, characterized in that, Includes the following steps: (1) Obtain statistical data on bridge scour, which will be used as the training set and the test set respectively; in step (1), statistical data are obtained from historical literature and used as the training set and the test set respectively, ensuring that the references of the training set and the test set do not overlap and that the data are independent of each other; the statistical data includes: Data{b, L, θ, V1, y1, d 50 ,s g ,y s } b is the width of the bridge pier, L is the length of the bridge, θ is the angle of attack of the water flow, V1 is the average velocity of the water flow upstream of the bridge pier, y1 is the water depth upstream of the bridge pier after typical scouring, and d 50 σ represents the particle size of the sediment. g For the uniformity of sediment, y s This represents the local scour depth of the bridge pier. (2) Determine the initial structure of the bridge scour formula; in step (2), optimize and reconstruct the CSU equation structure adopted by the US standard to generate a new calculation formula: In the formula: y s y1 represents the local scour depth of the bridge pier, y2 represents the upstream water depth of the bridge pier after general scour, K1 is the pier shape correction factor, K2 is the flow angle of attack correction factor, K3 is the riverbed condition correction factor, b is the pier width, and F is the bridge pier width. r For the Froude number of the water flow upstream of the bridge pier, Where V1 is the average velocity of the water flow upstream of the bridge pier, and g is the acceleration due to gravity. Take the opposite side of the formula: The initial structure of the optimized bridge scour formula is as follows: (3) First, candidate operators are determined, and then a group of functions with significant physical meaning and statistical relationship in scouring is introduced to determine candidate operational variables; common operators are used in step (3) to determine candidate operators, including: Operator{+,-,×,÷,^2,^3,sin,cos} Based on the bridge scour test results and correlation analysis, the following parameters were found to have a significant statistical relationship with the scour depth: Parameter{b n ,d 50 ,s g ,y1,V1} Among them, b n For the effective width, b n = bcosθ + Lsinθ, where b is the width of the pier, L is the length of the bridge, θ is the angle of attack of the water flow, and d 50 σ represents the particle size of the sediment. g To represent the uniformity of sediment, y1 is the water depth upstream of the bridge pier after typical scouring, and V1 is the average velocity of the water flow upstream of the bridge pier. Local scour depth y of bridge pier s The following relationship exists with Parameter: Based on the analysis of the scouring test results, the above function set can be adopted as: g1(x, y) = x + y, g2(x) = sin(ax), g3(x) = b x (0 <b<1),g4(x)=x c-x g5 = log d (x+1)(d>1); Parameters a, b, c, and d should be determined based on the dataset. Candidate operands include: (4) Use the symbolic regression method based on genetic programming to generate the optimization formula using the data in the training set; (5) Adjust parameters to generate a group of alternative formulas. (6) Evaluate the formula effect based on the test set and select the formula form with the best prediction effect and the simplest formula structure.

2. The bridge scour formula optimization and reconstruction method based on symbolic regression according to claim 1, characterized in that: Step (4) specifically includes the following steps: (4.1) Based on the candidate operators and candidate operation variables obtained in step (3), generate a random initial expression population containing operators, constants, and variables. The population size i is defined as the number of individuals in the population. The initialization process needs to be repeated until i expressions are generated: Original{E1, E2, ..., E...} i }; (4.2) Based on the training set data, the fitness function is used to evaluate the population performance, and the top n expressions with the best performance are selected as the maternal population of the next generation: Parents{E1, E2, ..., E... n }, n = 0.8i; (4.3) Genetic programming algorithm is used to perform genetic iteration on the selected population based on the training set data to generate the subpopulation with the best fitness. It is divided into three stages: crossover stage, mutation stage and replacement stage. (4.4) Calculate the new population New{E1, E2, ..., E} based on the training set data. n Fitness, and the new population New{E1, E2, ..., E n Sort the parent populations (E1, E2, ..., E...) in ascending order of fitness, and use them as the maternal lineage for the next iteration. n }; Returning to step (4.3), after reaching the set number of iterations, extract the expression with the best fitness from the population as the target formula E. target .

3. The bridge scour formula optimization and reconstruction method based on symbolic regression according to claim 2, characterized in that: The fitness calculation formula for each expression in step (4.2) is as follows: Where N is the total number of samples, It is individual E i The predicted value y for the j-th sample i It is the corresponding true value.

4. The bridge scour formula optimization and reconstruction method based on symbolic regression according to claim 3, characterized in that: In step (4.3), the crossover stage refers to randomly selecting the subtrees of two maternal populations and exchanging them to generate two new offspring trees; the mutation stage refers to randomly mutating the nodes of the individual expression tree while keeping the tree structure unchanged; and the replacement stage refers to regenerating a new tree to replace the current tree.

5. The bridge scour formula optimization and reconstruction method based on symbolic regression according to claim 4, characterized in that: The target formula E in step (4.4) target for:

6. The bridge scour formula optimization and reconstruction method based on symbolic regression according to claim 5, characterized in that: The adjustable parameter set P in step 5 is as follows: P={num,depth,p variation ,p crossover ,p mutation ,p replacement ,run} Where num is the number of initialization expressions, depth is the depth of the tree, and p variation p represents the probability that a node is a variable when the tree is being spanned. crossover p represents the probability of intersection between different expressions. mutation p represents the probability of node mutation. replacement The probability of generating a new tree to replace the current tree is given by `run`, which is the number of genetic iterations.

7. The method for optimizing and reconstructing bridge scour formula based on symbolic regression according to claim 6, characterized in that: In step (6), the alternative formula group is calculated. MSE and R 2 Rank the prediction results: Where N is the total number of samples, It is individual E i The predicted value y for the j-th sample i It corresponds to the actual value. It is the average of the true values; the smaller the MSE, the better the prediction effect; R 2 The closer the result is to 1, the better the prediction effect.