Method for determining excavation sequence of deep and shallow foundation pits of transfer channel by using intelligent algorithm
Through the reinforcement learning algorithm to optimize the excavation sequence of deep and shallow foundation pits, the problems of low efficiency and high risk in traditional methods are solved, and safe and efficient construction of deep and shallow foundation pit excavation in Ningbo soft soil strata is achieved, which is suitable for intelligent construction decisions in complex working conditions.
Patent Information
- Application Number
- CN202510620828.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-26
AI Technical Summary
When excavating deep and shallow foundation pits in soft soil strata, traditional methods cannot effectively solve dynamic decision-making problems, resulting in low efficiency and high cost in the selection of excavation sequence, and the inability to optimize strategies based on real-time monitoring data, which poses a risk of structural coupling.
The reinforcement learning algorithm is adopted, and by establishing a three-dimensional finite element model and offline training sample set, combining the Q-learning algorithm for dynamic strategy training, optimizing the excavation order of the depth and shallow foundation pits, and iteratively updates the Q value using the ε-greedy strategy and the Bellman equation, comprehensively considering displacement, settlement and construction efficiency, and determining the optimal excavation order is achieved.
It improves the safety and efficiency of excavation of deep and shallow foundation pits, reduces the design workload, demonstrates dynamic adaptability and high-dimensional data processing capabilities, and is suitable for intelligent construction decisions in complex working conditions.
Smart Images

Figure CN120542166A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of subway transfer channel construction, and in particular to a method for determining the excavation sequence of deep and shallow foundation pits of a transfer channel by utilizing an intelligent algorithm. Background Art
[0002] With the continuous development of urban underground space, deep and shallow foundation pit excavation has become a common construction type. However, in soft soil strata such as Ningbo, due to the constraints of the special engineering properties of soft soil on the excavation sequence, deep and shallow foundation pit excavation in soft soil strata with high water content faces core challenges such as high stratum sensitivity, significant displacement differences caused by slight differences in excavation sequence due to the rheological characteristics of soft soil, and structural coupling risks. Traditional methods rely on experience to select the excavation sequence, which has problems such as low numerical simulation efficiency, high cost of exhaustive calculation of all sequences, insufficient dynamic adjustment capabilities, and inability to optimize strategies based on real-time monitoring data.
[0003] In the prior art, although random forests and genetic algorithms are used for geotechnical parameter inversion, they cannot solve the problem of dynamic decision-making. For example, Chinese patent CN119720746A discloses a fast intelligent prediction method for deep foundation pit deformation in soft soil areas. The intelligent prediction method includes the following steps: randomly generating small strain hardening (HSS) model parameters of different soft soil layers through the Monte Carlo algorithm, simulating the foundation pit excavation process using the finite element software PLAXIS, and obtaining the deformation of the deep foundation pit under different stratum conditions; summarizing the soil HSS model parameters and the maximum lateral displacement of the foundation pit retaining structure, the maximum surface settlement outside the pit, and the maximum positive bending moment of the retaining wall toward the pit into the soft soil area deep foundation pit deformation database; introducing the artificial neural network (ANN) method to establish the relationship between the HSS model parameters in the soft soil area and the deep foundation pit deformation, and realizing fast intelligent prediction of the deformation of deep foundation pits in soft soil areas. Although this method can predict the stratum deformation after each foundation pit excavation, it cannot solve the dynamic decision-making problem during foundation pit excavation and cannot obtain the optimal excavation sequence. Summary of the Invention
[0004] The present invention provides a method for determining the excavation sequence of deep and shallow foundation pits for transfer passages by using an intelligent algorithm, so as to solve the dynamic decision-making problem during subway foundation pit excavation.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A method for determining the excavation sequence of deep and shallow foundation pits for a transfer passage using an intelligent algorithm, characterized by comprising:
[0007] S1. Establish a three-dimensional finite element model
[0008] A three-dimensional finite element model was constructed based on the ground parameters, support structure parameters, and boundary conditions of the target subway foundation pit project. The alternating construction process of deep and shallow foundation pits under different excavation sequences was simulated, and the displacement and settlement corresponding to each construction stage were output.
[0009] S2. Generate multimodal training sample set
[0010] Numerical simulations of three excavation modes, namely, deep excavation followed by shallow excavation, shallow excavation followed by deep excavation, and alternating deep and shallow excavation, were performed using the three-dimensional finite element model. Excavation steps under each working condition and their corresponding displacement and settlement data were extracted to form an offline training sample set containing complete construction process data. The offline training sample set was then introduced into a reinforcement learning algorithm.
[0011] S3. Build a reinforcement learning algorithm environment
[0012] S31. Define a state space, including the current deep foundation pit depth, the number of completed deep foundation pit excavation layers, the current shallow foundation pit depth, the number of completed shallow foundation pit excavation layers, the maximum displacement, the maximum settlement, the executed excavation sequence, and a discretized state dimension consisting of two-dimensional actions of excavating the deep foundation pit and excavating the shallow foundation pit;
[0013] S32, defining an action space, including action 1 for excavating a deep foundation pit and action 2 for excavating a shallow foundation pit;
[0014] S4. Reinforcement Learning Strategy Optimization
[0015] Use offline Q-learning algorithm model for dynamic strategy training, including:
[0016] S41. Set the learning rate α, discount factor γ, number of training rounds, exploration rate ε parameter and 2^12 state space possibility; initialize the state matrix and Q value consisting of depth, displacement / settlement and excavation sequence;
[0017] S42. Select the current excavation action using the ε-greedy strategy. For excavation actions previously included in the offline training sample set, the corresponding displacement and settlement data are directly used. When encountering excavation actions outside the offline training sample set, the displacement and settlement data are obtained by learning the solution simulated by learning the offline training sample set.
[0018] S43. Execute reward calculation and strategy update based on the displacement and settlement data:
[0019] S431. Calculate the current reward value based on the safety thresholds of displacement and settlement and the construction efficiency;
[0020] S432, iteratively updating the Q value through the Bellman equation;
[0021] S433, performing penalty determination based on the presence or absence of current excavation data in the historical database and the displacement / settlement exceeding the limit;
[0022] S44, loop through S42-S43 until the training termination condition is met;
[0023] S45. After the training is terminated, a convergence test is performed. When the average reward threshold reaches the set threshold and the Q value tends to be stable, the model is determined to have converged and the strategy test is started.
[0024] S5. Result test and output:
[0025] The trained Q-table is imported into the verification environment for strategy testing. A pure greedy strategy is used to select the optimal action in the Q-table and output the optimal excavation sequence that meets safety and efficiency requirements.
[0026] Furthermore, in step S42, the current excavation action is selected by the ε-greedy strategy, and its execution strategy is:
[0027]
[0028] Where ε is the exploration probability, 0≤ε≤1, Q(s,a) is the value function of action a in state s, argmax a' Q(s,a') is the current optimal action. The ε-greedy strategy selects the action with the highest current Q value with a probability of 1-ε, maximizes the immediate reward, and selects uniformly randomly from all actions with a probability of ε to discover potentially better actions.
[0029] Furthermore, in step S431, the current reward value is calculated based on the safety thresholds of displacement and settlement and the construction efficiency. The design mechanism of the current reward value is as follows:
[0030] The expression of the reward function is:
[0031] R=w p R progress +w s R safety +w c R completion +w q R sequence (2)
[0032] Where w p is the progress weight coefficient, R progress For progress rewards, w s is the safety weight coefficient, R safety For safety punishment, w c To complete the reward weight coefficient, R completion To complete the reward, w q is the sequence penalty weight coefficient, Rsequence is the sequence penalty;
[0033] The progress bonus is calculated based on the ratio of the total number of deep and shallow foundation pit excavations to the total number of excavations. The calculation formula is:
[0034]
[0035] Where N deep is the number of times the deep foundation pit has been excavated, N shallow is the number of times the shallow foundation pit has been excavated, N total is the total number of excavations;
[0036] The safety penalty is calculated based on the difference between the displacement and the displacement safety threshold, and the difference between the settlement and the settlement safety threshold, and the formula is:
[0037]
[0038] Where, d is the maximum displacement of the ground-connected wall, d max is the displacement safety threshold of the ground-connected wall, s is the maximum surface settlement, s max is the safety threshold of surface settlement;
[0039] When the excavation times of both deep and shallow foundation pits reach the design upper limit, and the displacement and settlement do not exceed the safety threshold, a one-time high completion reward will be granted. The formula is:
[0040]
[0041] When the same foundation pit is repeatedly excavated, the sequence penalty is calculated based on the ratio of the square of the length of consecutive identical actions to the total number of excavations. The formula is:
[0042]
[0043] Where c i is the length of the i-th segment of continuous identical action, and k is the number of segments of continuous identical action in the current excavation sequence.
[0044] Furthermore, in step S432, the Q value is iteratively updated by the Bellman equation, which is as follows:
[0045] Q(s,a)←Q(s,a)+α[R+γmax a' Q(s',a')-Q(s,a)] (7)
[0046] Where s is the current state, a is the current action, Q(s,a) is the current Q value, which represents the historical experience value of executing action a in state s, and the initial value is 0; max a'Q(s',a') is the maximum expected value of the next state; R is the immediate reward, which is calculated by the reward function described in S431, α is the learning rate, γ is the discount factor, [R+γmax a' Q(s′,a′)-Q(s,a)] is the temporal difference error, which reflects the difference between the actual experience and the expected value and is used to correct the Q table.
[0047] Furthermore, in step S433, penalty determination is performed based on the set dynamic penalty mechanism according to the displacement exceeding the limit, settlement exceeding the limit or data missing. The dynamic penalty mechanism is designed as follows:
[0048] After each excavation action is executed, the Q-learning algorithm model generates a unique code based on the current excavation sequence and queries the historical database to determine the presence or absence of the unique code. The formula is:
[0049]
[0050] When the unique code does not exist in the historical database, it is determined that the unique code is missing and isDataValid = 0. At this time, the learning rate is dynamically adjusted and the learning rate is proportionally decayed to avoid noisy updates:
[0051] α adjusted =0.2α (9)
[0052] Where α is the learning rate set by S31 above;
[0053] At this time, multi-dimensional penalty calculation is performed, and the penalty value consists of two parts: security penalty and data missing penalty;
[0054] Safety penalty is calculated based on the displacement and settlement exceeding the limit: the corresponding formula is
[0055] Safety penalty = -50[ / (·)(Δ d >d max )]+[ / (·)(Δ s >s max )] (10)
[0056] Where / (·) is an indicator function, which takes 1 when the limit is exceeded and 0 otherwise;
[0057] Data missing penalty: if the current state has no historical data support, that is, when the unique code is missing, an additional penalty is imposed: the corresponding formula is
[0058]
[0059] In addition, when the unique code exists in the historical database, the maximum displacement and settlement values of the response are obtained from the historical database. When the unique code does not exist in the historical database, while dynamically adjusting the learning rate, the existing data is learned and a displacement / settlement value is generated, and a contribution penalty is applied to the displacement / settlement in proportion. The formula is:
[0060]
[0061] Combining multi-dimensional penalties and contribution penalties, the complete penalty calculation process is as follows:
[0062]
[0063] Furthermore, the training termination condition in step S44 is:
[0064] The conditions for successful termination are as follows:
[0065]
[0066] The failure termination conditions are as follows:
[0067]
[0068] If the successful termination condition is met, the result is recorded and the next step of convergence determination is entered. If the failed termination condition is met, the calculation is stopped and returned to S41 to reinitialize the calculation.
[0069] Furthermore, the conditions for convergence determination in step S45 are as follows:
[0070] Average reward threshold:
[0071]
[0072] N is the number of training rounds in the statistical window, R episode,i is the total reward of the i-th training round, τ is the preset threshold;
[0073] Q value stability:
[0074] σ(Q current -Q previous )<0.1 (17)
[0075] σ represents the standard deviation, which calculates the Q value fluctuation of the last N updates;
[0076] If the convergence judgment fails, the process returns to the S41 round initialization to continue the Q-value update calculation. If the convergence judgment succeeds, it means that the Q-learning algorithm model has learned a stable strategy, the Q-value changes slowly, and the average reward is close to the theoretical optimal value, and the calculation ends and enters the strategy test.
[0077] Compared with the prior art, the present invention has the following beneficial effects:
[0078] By introducing a reinforcement learning algorithm during the excavation of deep and shallow subway foundation pits, the present invention can more efficiently and conveniently determine the excavation sequence of deep and shallow foundation pits, while comprehensively considering the potential risks of each excavation step. This method significantly increases the safety of the deep and shallow foundation pit excavation process on the basis of reducing the workload of excavation design.
[0079] In addition, the present invention demonstrates the core advantages of dynamic adaptability, high-dimensional data processing capabilities, cross-scenario migration and real-time interactive optimization capabilities in the optimization of the excavation sequence of deep and shallow foundation pits through reinforcement learning algorithms: it is particularly suitable for intelligent construction decision-making in complex working conditions such as soft soil layers, and realizes real-time optimization of the excavation sequence through autonomous exploration and trial-and-error mechanisms. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0081] Figure 1 It is a process step diagram of the present invention;
[0082] Figure 2 It is a schematic diagram of a deep foundation pit;
[0083] Figure 3 This is a schematic diagram of a shallow foundation pit;
[0084] Figure 4 It is a schematic diagram of the soil layers where the deep and shallow foundation pits are located;
[0085] Figure 5 This is a schematic diagram of modeling using PLAXIS3D finite element software;
[0086] Figure 6 This is a schematic diagram of the monitoring point arrangement during the deep and shallow foundation pit excavation process;
[0087] Figure 7 It is the optimal excavation sequence result. DETAILED DESCRIPTION
[0088] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0089] This embodiment provides a method for determining the excavation sequence of deep and shallow foundation pits for transfer passages using an intelligent algorithm, such as Figure 1 Shown, including:
[0090] S1. Establish a three-dimensional finite element model
[0091] A three-dimensional finite element model was constructed based on the ground parameters, support structure parameters, and boundary conditions of the target subway foundation pit project. The alternating construction process of deep and shallow foundation pits under different excavation sequences was simulated, and the displacement and settlement corresponding to each construction stage were output.
[0092] S2. Generate multimodal training sample set
[0093] Numerical simulations of three excavation modes, namely, deep excavation followed by shallow excavation, shallow excavation followed by deep excavation, and alternating deep and shallow excavation, were performed using the three-dimensional finite element model. Excavation steps under each working condition and their corresponding displacement and settlement data were extracted to form an offline training sample set containing complete construction process data. The offline training sample set was then introduced into a reinforcement learning algorithm.
[0094] S3. Build a reinforcement learning algorithm environment
[0095] S31. Define a state space, including the current deep foundation pit depth, the number of completed deep foundation pit excavation layers, the current shallow foundation pit depth, the number of completed shallow foundation pit excavation layers, the maximum displacement, the maximum settlement, the executed excavation sequence, and a discretized state dimension consisting of two-dimensional actions of excavating the deep foundation pit and excavating the shallow foundation pit;
[0096] S32, defining an action space, including action 1 for excavating a deep foundation pit and action 2 for excavating a shallow foundation pit;
[0097] S4. Reinforcement Learning Strategy Optimization
[0098] Use offline Q-learning algorithm model for dynamic strategy training, including:
[0099] S41. Set the learning rate α, discount factor γ, number of training rounds, exploration rate ε parameter and 2^12 state space possibility; initialize the state matrix and Q value consisting of depth, displacement / settlement and excavation sequence;
[0100] S42. Select the current excavation action using the ε-greedy strategy. For excavation actions previously included in the offline training sample set, the corresponding displacement and settlement data are directly used. When encountering excavation actions outside the offline training sample set, the displacement and settlement data are obtained by learning the solution simulated by learning the offline training sample set.
[0101] S43. Execute reward calculation and strategy update based on the displacement and settlement data:
[0102] S431. Calculate the current reward value based on the safety thresholds of displacement and settlement and the construction efficiency;
[0103] S432, iteratively updating the Q value through the Bellman equation;
[0104] S433, performing penalty determination based on the presence or absence of current excavation data in the historical database and the displacement / settlement exceeding the limit;
[0105] S44, loop through S42-S43 until the training termination condition is met;
[0106] S45. After the training is terminated, a convergence test is performed. When the average reward threshold reaches the set threshold and the Q value tends to be stable, the model is determined to have converged and the strategy test is started.
[0107] S5. Result test and output:
[0108] The trained Q-table is imported into the verification environment for strategy testing. A pure greedy strategy is used to select the optimal action in the Q-table and output the optimal excavation sequence that meets safety and efficiency requirements.
[0109] Specifically, the three-dimensional finite element model was established. Based on the data of the deep and shallow foundation pits of the Ningbo soft soil Zemin Station and the connecting channel, the deep and shallow foundation pit three-dimensional finite element model was established using PLAXIS3D finite element software and the HSS model; the construction process of the deep foundation pit (divided into 8 excavations) and the shallow foundation pit (divided into 4 excavations) under different excavation sequences was simulated, and the displacement and settlement data corresponding to each construction stage were output;
[0110] The data for establishing the three-dimensional finite element model of the deep and shallow foundation pit include: stratum parameters (such as Figure 4 As shown in the figure, it is composed of miscellaneous fill A, clay B, silty clay C, round gravel D, silt E, and rock layer F), support structure parameters (the first and sixth layers of the deep foundation pit are concrete supports, and the rest are steel supports; the first layer of the shallow foundation pit is concrete support, and the rest are steel supports) and boundary conditions; the modeling of the three-dimensional finite element model is as follows Figure 5 As shown in the figure, 5 is a deep foundation pit model and 6 is a shallow foundation pit model; the composition of the deep foundation pit model and the shallow foundation pit model is as follows Figure 2 and Figure 3As shown, in the figure, 1 is a concrete support, 2 is a steel support, 3 is a foundation pit ground connecting wall, and 4 is a foundation pit bottom plate.
[0111] Specifically, the multi-mode training sample set was generated. Three excavation modes (deep then shallow, shallow then deep, and alternating deep and shallow) were numerically simulated using the three-dimensional finite element model. 15 sets of excavation steps and their corresponding displacement and settlement data under various working conditions were extracted. An offline training sample set containing complete construction process data was constructed, and the offline training sample set was introduced into the reinforcement learning algorithm. Specific data of the offline training sample set are shown in Table 1 below:
[0112]
[0113]
[0114]
[0115]
[0116]
[0117] Table 1
[0118] In Table 1, the locations of the various detection points are as follows: Figure 6 As shown, 5 is a deep foundation pit, 6 is a shallow foundation pit, Ax is the x-axis displacement monitoring point of the shallow foundation pit diaphragm wall, Bx is the x-axis displacement monitoring point of the deep foundation pit diaphragm wall, Cx is the x-axis displacement monitoring point of the deep foundation pit diaphragm wall at the intersection of the deep and shallow foundation pits, Dy is the y-axis displacement monitoring point of the deep foundation pit diaphragm wall at the intersection of the deep and shallow foundation pits, Ez is the surface settlement monitoring point of the shallow foundation pit, Fz is the surface settlement monitoring point of the deep foundation pit, and Gz is the surface settlement monitoring point of the shallow foundation pit at the intersection of the deep and shallow foundation pits.
[0119] Specifically, the construction of the reinforcement learning algorithm environment includes defining a state space and defining an action space;
[0120] In the definition of the state space, the total excavation steps for deep foundation pits are defined as 8 layers, the total excavation steps for shallow foundation pits are defined as 4 layers, the maximum displacement is d, and the maximum settlement is s. In addition, the current deep foundation pit depth, the number of completed deep foundation pit excavation layers, the current shallow foundation pit depth, the number of completed shallow foundation pit excavation layers, the executed excavation sequence, and the discretized state dimension including the two-dimensional actions of deep foundation pit excavation and shallow foundation pit excavation are defined.
[0121] In the definition of the action space, two discrete actions are defined: action one is to excavate a deep foundation pit, and action two is to excavate a shallow foundation pit.
[0122] Specifically, the reinforcement learning strategy optimization adopts an offline Q-Learning algorithm model to perform dynamic strategy training, and the specific steps include:
[0123] First, the reinforcement learning algorithm is set up: the learning rate α is set to 0.1, the discount factor γ is set to 0.99, the exploration rate ε is set to 0.3, the number of training rounds is set to 1000, and the state space possibility is set to 2^12; the excavation depth of deep and shallow foundation pits is reset to zero, the displacement, settlement and other deformations are reset to zero, the excavation sequence is cleared, and the initial state vector is obtained; the Q value is initialized to 0.
[0124] Second, the current excavation action is selected through the ε-greedy strategy, and its execution strategy is:
[0125]
[0126] In the Ningbo soft soil deep and shallow foundation pit excavation project, the exploration probability ε is set to 0.3, Q(s,a) is the value function of action a under state s, arg max a' Q(s,a') is the current optimal action. The ε-greedy strategy selects the action with the highest Q value with a probability of 70% to maximize the immediate reward, and selects uniformly randomly among all actions with a probability of 30% to find potentially better actions.
[0127] For excavation actions that have been included in the offline training sample set in the early stage, the corresponding stratum response data is directly used; when encountering excavation actions outside the offline training sample set, the stratum response data is obtained by learning the scheme simulated by learning the offline training sample set.
[0128] Third, reward calculation and strategy update are performed based on the formation response data, which includes three steps:
[0129] First, the current reward value is calculated based on the safety thresholds of displacement and settlement and the construction efficiency. The design mechanism of the current reward value is as follows:
[0130] The expression of the reward function is:
[0131] R=w p R progress +w s R safety +w c R completion +w q R sequence (2)
[0132] Where, the progress weight coefficient w p Set to 0.4, R progress is the progress reward, and the safety weight coefficient w s Set to 0.3, R safety For security penalty, complete the reward weight coefficient w c is 0.2, R completionTo complete the reward, the sequence penalty weight coefficient w q is 0.1, R sequence is the sequence penalty;
[0133] The progress bonus is calculated based on the ratio of the total number of shallow foundation pit excavations to the total number of excavations to encourage the excavation progress to proceed as planned. The bonus is proportional to the amount of work completed. The calculation formula is:
[0134]
[0135] Where, the number of deep foundation pit excavations N deep Set to 8, the number of shallow foundation pit excavations N shallow Set to 4, the total number of excavations is N total Set to 12;
[0136] Safety penalty is calculated based on the difference between the displacement and the displacement safety threshold, and the difference between the settlement and the settlement safety threshold, to penalize excessive ground-connected wall displacement and surface settlement, forcing the model to prioritize safety. The formula is:
[0137]
[0138] Where d is the maximum displacement of the ground-connected wall; the displacement safety threshold d is max According to the first-level foundation pit control standard in the Technical Specification for Building Foundation Pit Support (JGJ 120-2016), it is set to 50mm; s is the maximum surface settlement; the surface settlement safety threshold s max According to the Ningbo Metro Protection Regulations, the settlement limit around sensitive areas (such as operating tunnels) is set at 30mm;
[0139] Completion reward: When the number of deep and shallow foundation pit excavations reaches the design upper limit, and the displacement and settlement do not exceed the safety threshold, a one-time high completion reward is awarded to guide the model to pursue complete construction. The formula is:
[0140]
[0141] When the same layer of foundation pit is repeatedly excavated, the sequence penalty is calculated based on the ratio of the square of the length of the consecutive identical actions to the total number of excavations. In order to suppress the behavior of repeatedly excavating the same layer of foundation pit and avoid repeated calculations, the formula is:
[0142]
[0143] Where c i is the length of the i-th segment of continuous identical action, and k is the number of segments of continuous identical action in the current excavation sequence;
[0144] This reward function quantifies the comprehensive impact of the excavation sequence of deep and shallow foundation pits in Ningbo's soft soil area on construction safety, efficiency, and continuity through weighted multi-objective analysis. Its core goal is to maximize construction progress and optimize operational continuity while controlling deformation risks.
[0145] Secondly, the Q value is iteratively updated through the Bellman equation, which is as follows:
[0146] Q(s,a)←Q(s,a)+α[R+γmax a' Q(s',a')-Q(s,a)] (7)
[0147] Where s is the current state, a is the current action, Q(s,a) is the current Q value, which represents the historical experience value of executing action a in state s, and the initial value is 0; max a' Q(s',a') is the maximum expected value of the next state; R is the immediate reward, calculated by the reward function, the learning rate α is 0.1, the discount factor γ is 0.99, [R+γmax a' Q(s′,a′)-Q(s,a)] is the temporal difference error, which reflects the difference between the actual experience and the expected value and is used to correct the Q table.
[0148] Finally, according to the displacement exceeding the limit, settlement exceeding the limit or data missing, penalty judgment is made based on the set dynamic penalty mechanism. The design of the dynamic penalty mechanism is as follows:
[0149] After each excavation action is performed, the Q-Learning algorithm model generates a unique code based on the current excavation sequence and queries the historical database to determine the presence or absence of the unique code. The formula is:
[0150]
[0151] When the unique code does not exist in the historical database, it is determined that the unique code is missing and isDataValid = 0. At this time, the learning rate is dynamically adjusted and the learning rate is proportionally decayed to avoid noisy updates:
[0152] α adjusted =0.2α (9)
[0153] Where α is the learning rate set in the above reinforcement learning settings, and the learning rate α is set to 0.1;
[0154] At this time, multi-dimensional penalty calculation is performed, and the penalty value consists of two parts: security penalty and data missing penalty;
[0155] Safety penalties, calculated based on displacement and settlement exceeding limits:
[0156] Safety penalty = -50[ / (·)(Δd >d max )]+[ / (·)(Δ s >s max )] (10)
[0157] Where, / (·) is an indicative function, which takes the value 1 when the limit is exceeded and 0 otherwise; the penalty for exceeding the limit of displacement / settlement triggers (-50 points / item)
[0158] Data missing penalty: if the current state has no historical data support, that is, when the unique code is missing, an additional penalty is imposed: the corresponding formula is:
[0159]
[0160] Missing data triggers a penalty (-10 points);
[0161] In addition, when the unique code exists in the historical database, the maximum displacement and settlement values of the response are obtained from the historical database. When the unique code does not exist in the historical database, while dynamically adjusting the learning rate, the existing data is learned and a displacement / settlement value is generated, and a contribution penalty is applied to the displacement / settlement in proportion. The formula is:
[0162]
[0163] The complete penalty calculation process is as follows:
[0164]
[0165] Fourth, the conditions for terminating the training are:
[0166] The conditions for successful termination are as follows:
[0167]
[0168] The failure termination conditions are as follows:
[0169]
[0170] If the successful termination conditions are met, the results are recorded and the next step of convergence determination is entered. If the failure termination conditions are met, the calculation is stopped and reinitialized.
[0171] Fifth, the conditions for convergence determination are as follows:
[0172] Average reward threshold:
[0173]
[0174] N is the number of training rounds in the statistical window, R episode,i is the total reward of the i-th training round, τ is the preset threshold;
[0175] Q value stability:
[0176] σ(Q current -Q previous )<0.1 (17)
[0177] σ represents the standard deviation, which calculates the Q value fluctuation of the last N updates;
[0178] If the convergence fails, the initialization phase is restarted and the Q value is updated. If the convergence succeeds, it means that the model has learned a stable strategy, the Q value changes slowly, and the average reward is close to the theoretical optimal value. The calculation ends and the strategy test begins.
[0179] Specifically, the result test and output are as follows: the trained Q-table is imported into the verification environment for strategy testing, the exploration probability ε in the ε-greedy strategy is set to 0, and the optimal action in the Q-table is selected for calculation using a pure greedy strategy. The optimal excavation sequence that meets safety and efficiency requirements is output, which is the optimal excavation sequence for deep and shallow foundation pits in soft soil in Ningbo.
[0180] The best excavation sequence is Figure 7 As shown in the figure, the optimal excavation sequence is: two layers of deep foundation pit excavation - two layers of shallow foundation pit excavation - two layers of deep foundation pit excavation - two layers of shallow foundation pit excavation - four layers of deep foundation pit excavation. After the excavation is completed, the maximum displacement of the ground diaphragm wall is 41.7 mm, and the maximum surface settlement is 25.7 mm.
[0181] Compared with the prior art, the present invention has the following beneficial effects:
[0182] By introducing a reinforcement learning algorithm during the excavation of deep and shallow subway foundation pits, the present invention can more efficiently and conveniently determine the excavation sequence of deep and shallow foundation pits, while comprehensively considering the potential risks of each excavation step. This method significantly increases the safety of the deep and shallow foundation pit excavation process on the basis of reducing the workload of excavation design.
[0183] In addition, the present invention demonstrates the core advantages of dynamic adaptability, high-dimensional data processing capabilities, cross-scenario migration and real-time interactive optimization capabilities in the optimization of the excavation sequence of deep and shallow foundation pits through reinforcement learning algorithms: it is particularly suitable for intelligent construction decision-making in complex working conditions such as soft soil layers, and realizes real-time optimization of the excavation sequence through autonomous exploration and trial-and-error mechanisms.
[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for determining the excavation sequence of deep and shallow foundation pits for a transfer channel using an intelligent algorithm, characterized in that: include: S1. Establish a three-dimensional finite element model A three-dimensional finite element model was constructed based on the ground parameters, support structure parameters, and boundary conditions of the target subway foundation pit project. The alternating construction process of deep and shallow foundation pits under different excavation sequences was simulated, and the displacement and settlement corresponding to each construction stage were output. S2. Generate multimodal training sample set Numerical simulations of three excavation modes, namely, deep excavation followed by shallow excavation, shallow excavation followed by deep excavation, and alternating deep and shallow excavation, were performed using the three-dimensional finite element model. Excavation steps under each working condition and their corresponding displacement and settlement data were extracted to form an offline training sample set containing complete construction process data. The offline training sample set was then introduced into a reinforcement learning algorithm. S3. Build a reinforcement learning algorithm environment S31. Define a state space, including the current deep foundation pit depth, the number of completed deep foundation pit excavation layers, the current shallow foundation pit depth, the number of completed shallow foundation pit excavation layers, the maximum displacement, the maximum settlement, the executed excavation sequence, and a discretized state dimension consisting of two-dimensional actions of excavating the deep foundation pit and excavating the shallow foundation pit; S32, defining an action space, including action 1 for excavating a deep foundation pit and action 2 for excavating a shallow foundation pit; S4. Reinforcement Learning Strategy Optimization Use offline Q-learning algorithm model for dynamic strategy training, including: S41. Set the learning rate α, discount factor γ, number of training rounds, exploration rate ε parameter and 2^12 state space possibility; initialize the state matrix and Q value consisting of depth, displacement / settlement and excavation sequence; S42, through ε - The greedy strategy selects the current excavation action. For excavation actions that are previously included in the offline training sample set, the corresponding displacement and settlement data are directly used. When encountering excavation actions outside the offline training sample set, the displacement and settlement data are obtained by learning the scheme simulated by learning the offline training sample set. S43. Execute reward calculation and strategy update based on the displacement and settlement data: S431. Calculate the current reward value based on the safety thresholds of displacement and settlement and the construction efficiency; S432, iteratively updating the Q value through the Bellman equation; S433, performing penalty determination based on the presence or absence of current excavation data in the historical database and the displacement / settlement exceeding the limit; S44, loop through S42-S43 until the training termination condition is met; S45. After the training is terminated, a convergence test is performed. When the average reward threshold reaches the set threshold and the Q value tends to be stable, the model is determined to have converged and the strategy test is started. S5. Result test and output: The trained Q-table is imported into the verification environment for strategy testing. A pure greedy strategy is used to select the optimal action in the Q-table and output the optimal excavation sequence that meets safety and efficiency requirements.
2. The method for determining the excavation sequence of deep and shallow foundation pits of a transfer channel using an intelligent algorithm according to claim 1, characterized in that: In step S42, ε - The greedy strategy selects the current excavation action, and its execution strategy is: Where ε is the exploration probability, 0≤ε≤1, Q(s,a) is the value function of action a in state s, argmaxa' Q(s,a') is the current optimal action, ε - The greedy strategy is to select the action with the highest current Q value with a probability of 1-ε, maximize the immediate reward, and select uniformly randomly among all actions with a probability of ε to discover potentially better actions.
3. The method for determining the excavation sequence of deep and shallow foundation pits of a transfer channel using an intelligent algorithm according to claim 1, characterized in that: In step S431, the current reward value is calculated based on the safety thresholds of displacement and settlement and the construction efficiency. The design mechanism of the current reward value is as follows: The expression of the reward function is: R=w p R progress +w s R safety +w c R completion +w q R sequence (2) Where w p is the progress weight coefficient, R progress For progress rewards, w s is the safety weight coefficient, R safety For safety punishment, w c To complete the reward weight coefficient, R completion To complete the reward, w q is the sequence penalty weight coefficient, R sequence It is a sequence penalty; The progress bonus is calculated based on the ratio of the total number of deep and shallow foundation pit excavations to the total number of excavations. The calculation formula is: Where N deep is the number of times the deep foundation pit has been excavated, N shallow is the number of times the shallow foundation pit has been excavated, N total is the total number of excavations; The safety penalty is calculated based on the difference between the displacement and the displacement safety threshold, and the difference between the settlement and the settlement safety threshold, and the formula is: Where, d is the maximum displacement of the ground-connected wall, d max is the displacement safety threshold of the ground-connected wall, s is the maximum surface settlement, s max is the safety threshold of surface settlement; When the excavation times of both deep and shallow foundation pits reach the design upper limit, and the displacement and settlement do not exceed the safety threshold, a one-time high completion reward will be granted. The formula is: When the same foundation pit is repeatedly excavated, the sequence penalty is calculated based on the ratio of the square of the length of consecutive identical actions to the total number of excavations. The formula is: Where c i is the length of the i-th segment of continuous identical action, and k is the number of segments of continuous identical action in the current excavation sequence.
4. The method for determining the excavation sequence of deep and shallow foundation pits of a transfer channel using an intelligent algorithm according to claim 1, characterized in that: In step S432, the Q value is iteratively updated using the Bellman equation, which is as follows: Q(s,a)←Q(s,a)+α[R+γmax a' Q(s',a')-Q(s,a)] (7) Where s is the current state, a is the current action, Q(s,a) is the current Q value, which represents the historical experience value of executing action a in state s, and the initial value is 0; max a' Q(s',a') is the maximum expected value of the next state; R is the immediate reward, which is calculated by the reward function described in S431, α is the learning rate, γ is the discount factor, and R+γmax a' Q(s′,a′)-Q(s,a)] is the temporal difference error, which reflects the difference between the actual experience and the expected value and is used to correct the Q table.
5. The method for determining the excavation sequence of deep and shallow foundation pits of a transfer passage using an intelligent algorithm according to claim 1, characterized in that: In step S433, penalty determination is performed based on the set dynamic penalty mechanism according to the displacement exceeding the limit, settlement exceeding the limit or data missing. The dynamic penalty mechanism is designed as follows: After each excavation action is executed, the Q-learning algorithm model generates a unique code based on the current excavation sequence and queries the historical database to determine the presence or absence of the unique code. The formula is: When the unique code does not exist in the historical database, it is determined that the unique code is missing and isDataValid = 0. At this time, the learning rate is dynamically adjusted and the learning rate is proportionally decayed to avoid noisy updates: α adjusted =0.2α (9) Where α is the learning rate set by S31 above; At this time, multi-dimensional penalty calculation is performed, and the penalty value consists of two parts: security penalty and data missing penalty; Safety penalty is calculated based on the displacement and settlement exceeding the limit: the corresponding formula is Safety penalty = -50[ / (·)(Δ d >d max )]+[ / (·)(Δ s >s max )] (10) Where / (·) is an indicator function, which takes 1 when the limit is exceeded and 0 otherwise; Data missing penalty: if the current state has no historical data support, that is, when the unique code is missing, an additional penalty is imposed: the corresponding formula is In addition, when the unique code exists in the historical database, the maximum displacement and settlement values of the response are obtained from the historical database. When the unique code does not exist in the historical database, while dynamically adjusting the learning rate, the existing data is learned and a displacement / settlement value is generated, and a contribution penalty is applied to the displacement / settlement in proportion. The formula is: Combining multi-dimensional penalties and contribution penalties, the complete penalty calculation process is as follows:
6. The method for determining the excavation sequence of deep and shallow foundation pits of a transfer passage using an intelligent algorithm according to claim 1, characterized in that: The training termination condition in step S44 is: The conditions for successful termination are as follows: The failure termination conditions are as follows: If the successful termination condition is met, the result is recorded and the next step of convergence determination is entered. If the failed termination condition is met, the calculation is stopped and returned to S41 to reinitialize the calculation.
7. The method for determining the excavation sequence of deep and shallow foundation pits of a transfer passage using an intelligent algorithm according to claim 1, characterized in that: The conditions for convergence determination in step S45 are as follows: Average reward threshold: N is the number of training rounds in the statistical window, R episode,i is the total reward of the i-th training round, τ is the preset threshold; Q value stability: σ(Q current -Q previous )<0.1 (17) σ represents the standard deviation, which calculates the Q value fluctuation of the last N updates; If the convergence judgment fails, the process returns to the S41 round initialization to continue the Q-value update calculation. If the convergence judgment succeeds, it means that the Q-learning algorithm model has learned a stable strategy, the Q-value changes slowly, and the average reward is close to the theoretical optimal value, and the calculation ends and enters the strategy test.
Citation Information
Patent Citations
Rapid intelligent prediction method for deformation of deep foundation pit in soft soil area
CN119720746A