Intelligent traffic control method based on bridge structure state and traffic state
Through the intelligent traffic control method combined with LSTM and DQN, bridge health status trend prediction and traffic strategy optimization are achieved, the dynamic balance of bridge monitoring and traffic control is solved, and intelligent decision-making support for bridge protection and traffic efficiency is improved.
Patent Information
- Application Number
- CN202510605272.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-04
AI Technical Summary
Existing bridge monitoring technology is difficult to predict future state changes. Traffic control measures lack real-time data and dynamic optimization of bridge structure status, making it difficult to take into account both bridge protection and traffic efficiency.
Long-term memory network (LSTM) is used to predict bridge health status trends, and reward functions are constructed in combination with deep Q network (DQN), and traffic control strategies are dynamically optimized to achieve intelligent balance between bridge protection and traffic efficiency.
Accurately predict future trends of bridge health scores, dynamically optimize traffic control strategies, reduce the need for manual intervention, improve system adaptability and operation efficiency, and provide efficient and scientific decision-making support.
Smart Images

Figure CN120260288A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bridge structure monitoring and traffic management, and in particular to an intelligent traffic control method based on bridge structure status and traffic status. Background Art
[0002] Bridges are an important part of traffic infrastructure, and their health status is directly related to traffic safety and passing efficiency. With the increase of the service life of bridges and the continuous increase of traffic loads, the problems of bridge structure damage and fatigue accumulation gradually appear, posing higher requirements for the safety and durability of bridges. In order to ensure the long-term use safety of bridges, the current bridge monitoring technology mainly relies on the health monitoring system to evaluate the bridge status by collecting the dynamic response data of the bridge structure.
[0003] However, the existing technologies have the following limitations in bridge health monitoring and traffic control: First, traditional monitoring methods usually can only perform static evaluation on the current status of bridges, and it is difficult to predict the future status change trend, and it is impossible to provide forward-looking decision-making support for bridge maintenance and traffic management. Second, most of the existing traffic control measures rely on preset rules or empirical judgments, lacking the dynamic optimization ability based on real-time monitoring data and bridge structure status, resulting in difficulty in balancing bridge protection and traffic efficiency. In addition, in complex traffic scenarios, it is difficult to quantify the relationship among bridge health score, traffic efficiency and control cost, and the existing technologies lack intelligent multi-objective optimization strategies.
[0004] With the development of artificial intelligence and big data technologies, it becomes possible to introduce deep learning and reinforcement learning technologies into bridge monitoring and traffic control. By combining the long short-term memory network (LSTM) to model and predict the time series data of the bridge health status, accurate prediction of the bridge status change trend can be achieved; by dynamically optimizing the traffic control strategy through the deep Q network (DQN), traffic efficiency can be maximized while ensuring the safety of the bridge structure. However, there has not been seen a complete technical solution for bridge health status prediction and traffic strategy adaptive optimization yet.
[0005] Therefore, there is an urgent need for an intelligent control method that combines real-time evaluation, trend prediction and adaptive learning optimization of bridge status to achieve a dynamic balance between bridge protection and traffic efficiency, while reducing the control cost, and providing efficient and scientific decision-making support for bridge maintenance units and traffic management departments. Summary of the Invention
[0006] To overcome the above-mentioned defects in the prior art, the present invention provides an intelligent traffic control method based on bridge structure state and traffic state. By introducing long short-term memory network (LSTM) and deep Q-network (DQN) optimization methods, combined with real-time monitoring data for state evaluation and future trend prediction, the traffic control strategy is dynamically optimized to achieve an intelligent balance between bridge protection and traffic efficiency.
[0007] To achieve the above object, the present invention adopts the following technical solutions, including:
[0008] An intelligent traffic control method based on bridge structure state and traffic state, comprising the following steps:
[0009] S1, Collect bridge structure monitoring data and traffic state data in real time;
[0010] S2, Evaluate the bridge structure monitoring data and calculate the current health score of the bridge in real time;
[0011] S3, Use the long short-term memory network to predict the future health score of the bridge;
[0012] S4, Define the state space of reinforcement learning according to the current health score and future health score of the bridge and the traffic state, define the action space of reinforcement learning based on the traffic control strategy, and use the deep Q-network to obtain the optimal traffic control strategy in combination with the reward function.
[0013] Preferably, in step S1, the bridge structure monitoring data includes acceleration, strain, lateral displacement and deflection; the traffic state data includes average vehicle speed, average load, heavy vehicle ratio and traffic flow density.
[0014] Preferably, in step S2, the current health score of the bridge is calculated by weighted calculation from multiple monitoring indicators according to the normal distribution model to quantify the current health state of the bridge. The calculation formula is as follows:
[0015]
[0016] Where, S c Is the current health score of the bridge, and the value range is [0,1]. The higher the value of S c , The healthier the bridge state; W i Is the weight of the i-th monitoring indicator in the bridge structure monitoring data, satisfying There are n monitoring indicators in the bridge structure monitoring data in total; S i Is the score of the i-th monitoring indicator, and the value range is [0,1]. The calculation method is as follows:
[0017]
[0018] Where, Vi is the actual value of the i-th monitoring index; V i max is the upper limit value allowed for the i-th monitoring index; V i min is the lower limit value allowed for the i-th monitoring index; V i ref is the reference value of the i-th monitoring index; σ i is the standard deviation of the i-th monitoring index.
[0019] Preferably, in step S3, the long short-term memory network is trained using the time series data of the bridge health score, and the trained long short-term memory network is used to predict the future health score of the bridge.
[0020] Preferably, in step S4, an optimal traffic control strategy is generated based on the deep Q network, which specifically includes the following steps:
[0021] S41. Define the state space. The state S is composed of the bridge structure state and the traffic state, including:
[0022] The current health score S of the bridge c ;
[0023] In the next k time steps, the predicted sequence S of the future health score of the bridge p (t + 1), S p (t + 2),..., S p (t + k);
[0024] Traffic state data, including traffic flow density q, average vehicle speed v c , proportion r of heavy-duty vehicles l , average load w c ;
[0025] S42. Define the action space A, which is composed of traffic control strategy combinations, including speed limit strategy, weight limit strategy, and diversion strategy;
[0026] S43. Construct a reward function. The reward R is used to quantify the advantages and disadvantages of different traffic control strategies. Considering the bridge protection, traffic efficiency, and control cost of the traffic control strategy, the expression of the reward function is:
[0027] R = ΔS bridge + λ1·T eff - λ2·C ctrl
[0028] where λ1 and λ2 are weight factors for adjusting the relative importance of traffic efficiency and control cost;
[0029] ΔS bridge represents the comprehensive health score of the bridge, ΔSbridge The higher the value, the better the health state of the bridge. ΔS bridge The goal is to maximize:
[0030]
[0031] T eff represents the traffic efficiency under the traffic control strategy. T eff The goal is to maximize it, and the calculation method is:
[0032]
[0033] where Q eff is the throughput per unit time, that is, the number of equivalent vehicles passing through per unit time; r l is the heavy load ratio, considering the structural burden brought by heavy-duty vehicles; δ is the penalty coefficient; q is the traffic flow density; T avg is the average passing time of vehicles; L is the total length of the bridge; v c is the average vehicle speed;
[0034] C ctrl represents the control cost of the traffic control strategy. C ctrl The goal is to minimize it, and the calculation method is as follows:
[0035] C ctrl = β1·(v normal - v c ) + β2·(w normal - w c ) + β3r d
[0036] where v normal is the designed vehicle speed; w normal is the expected average load of normal passing vehicles; r d is the diversion ratio; β1, β2, β3 are loss conversion coefficients;
[0037] S44. Obtain the optimal traffic control strategy based on the deep Q-network, as follows:
[0038] Input the current state S into the deep Q-network, and output the Q values corresponding to each action A; the Q value represents the cumulative reward that can be obtained by selecting action A in state S; the goal of the deep Q-network is to find the action A that maximizes the long-term cumulative reward through continuous updating;
[0039] where, the ε-greedy strategy is used for action selection, randomly select an unattempted action A with a probability of ε; select the action A with the largest current Q value with a probability of 1 - ε;
[0040] Execute the selected action A and implement the corresponding traffic control strategy for action A during the actual operation of the bridge. At the same time, record the current state S, the reward R obtained after executing the selected action A, and the new state S′ obtained after executing the selected action A, forming a quadruple (S, A, R, S′), and store this quadruple (S, A, R, S′) as a training sample of the deep Q-network in the experience replay pool.
[0041] Preferably, the deep Q-network continuously updates the network parameters through the gradient descent optimization method, as follows:
[0042] Sample training samples, randomly extract several quadruples (S, A, R, S′) from the experience replay pool to form a training mini-batch;
[0043] For each training sample, calculate the target Q value:
[0044]
[0045] where S′ is the new state after executing the current action A; A′ is the possible next action; γ is the discount factor, measuring the importance of future rewards; θ - is the target network parameter, and its value comes from the periodic copy of the current network parameter θ;
[0046] Use the mean squared error MSE as the loss function, calculate the average loss value of all training samples as the training error of the current deep Q-network to measure the difference between the predicted value and the target Q value of the current deep Q-network. The calculation method is as follows:
[0047]
[0048] where represents the expectation for multiple training samples; θ is the current network parameter of the deep Q-network; N is the number of batch samples used for training, and each sample corresponds to a state S i action A i and the target Q value y i ;
[0049] Every time a set training cycle passes, synchronously copy the current network parameter θ to the target network parameter θ - ;
[0050] The trained deep Q-network is used to output the optimal action A for any state S * :
[0051]
[0052] Preferably, if the future health score of the bridge indicates a downward trend in its health status, update the traffic control strategy. Using the method of step S4, based on the current health score and future health score of the bridge as well as the traffic status, call the deep Q-network to generate the optimal traffic control strategy; otherwise, maintain the current traffic control strategy.
[0053] Preferably, it further includes the following steps:
[0054] S5. Optimize the traffic control strategy in real time according to the current health score and future health score of the bridge and the traffic status.
[0055] S6. Dynamically feedback the changes in the bridge score and traffic efficiency after the implementation of the traffic control strategy, and update the traffic control strategy based on this feedback to achieve closed-loop optimization.
[0056] The present invention also provides an electronic device, which includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the intelligent traffic control method based on the bridge structure state and traffic state.
[0057] The present invention also provides a computer program product, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the intelligent traffic control method based on the bridge structure state and traffic state.
[0058] The advantages of the present invention are as follows:
[0059] (1) Compared with the existing technologies, the beneficial effects of the present invention are mainly reflected in: First, it proposes a method for predicting the trend of bridge state based on the long short-term memory network (LSTM), which can accurately predict the future change trend of the bridge health score and identify potential risks in advance; Second, combined with the deep Q-network (DQN), it dynamically optimizes the traffic control strategy by constructing a reward function, realizes the intelligent balance between bridge protection and traffic efficiency, reduces the need for manual intervention, and improves the self-adaptability and operation efficiency of the system.
[0060] (2) The method proposed by the present invention can intelligently optimize the traffic control strategy, thus more accurately responding to the changes in the bridge structure state and traffic flow fluctuations, and providing efficient and scientific decision-making assistance for bridge maintenance units and traffic management departments through intelligent prediction and real-time optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is a flowchart of an intelligent traffic control method based on the bridge structure state and traffic state of the present invention.
[0062] Figure 2It is the flowchart of the LSTM prediction of the bridge health trend of the present invention.
[0063] Figure 3 It is the framework diagram of the deep Q-network strategy optimization of the present invention.
[0064] Figure 4 It is the curve of the bridge health score and traffic efficiency change in the embodiment of the present invention. Detailed implementation manners
[0065] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0066] From Figures 1 - 3 As shown, an intelligent traffic control method based on the bridge structure state and traffic state includes the following steps:
[0067] S1. Collect bridge structure monitoring and traffic state data in real time to ensure the accuracy and integrity of the data.
[0068] S2. Based on the health score model, conduct real-time evaluation on the bridge structure monitoring data, calculate the current health score of the bridge, and judge the current state level of the bridge.
[0069] S3. Use the long short-term memory network (LSTM) to train the time series data of the bridge health score and predict the future health score of the bridge.
[0070] S4. Define the state space S based on the bridge structure state and traffic state, define the action space A based on the traffic control strategy, and use the deep Q-network (DQN) in combination with the reward function to generate the optimal traffic control strategy.
[0071] S5. Optimize the generated traffic control strategy in real time, including dynamically adjusting speed limits, weight limits, and diversion measures.
[0072] S6. Conduct dynamic feedback on the bridge score change and traffic efficiency change after the implementation of the traffic control strategy, update the action, that is, the traffic control strategy, and achieve closed-loop optimization.
[0073] The specific process of step S1 is as follows:
[0074] S11. Collect bridge structure monitoring data in real time, including acceleration, strain, lateral displacement, deflection, etc.
[0075] S12. Collect traffic status data in real time, including average vehicle speed, average load, traffic flow density (number of vehicles per unit time), and heavy vehicle ratio.
[0076] In step S2, the assessment of the bridge health status is specifically as follows:
[0077] S21. Build a bridge health scoring model, and calculate the current health score of the bridge by weighted calculation from multiple monitoring indicators according to the normal distribution model to quantify the current health status of the bridge. The calculation formula is as follows:
[0078]
[0079] Among them, S c is the current health score of the bridge, and its value range is [0, 1]. The higher the value of S c , the healthier the bridge status; W i is the weight of the i-th monitoring indicator in the bridge structure monitoring data, satisfying There are n monitoring indicators in the bridge structure monitoring data in total; S i is the score of the i-th monitoring indicator, and its value range is [0, 1]. The calculation method is as follows:
[0080]
[0081] Among them, V i is the actual value of the i-th monitoring indicator; V i max is the allowable upper limit value of the i-th monitoring indicator; V i min is the allowable lower limit value of the i-th monitoring indicator; V i ref is the reference value (i.e., the most ideal state) of the i-th monitoring indicator; σ i is the standard deviation of the i-th monitoring indicator, which is used to control the index fluctuation tolerance and can be set to corresponding to the 99.7% normal distribution range; when V i falls within the allowable interval, that is, [V i min , V i max , the closer it is to the reference value V i ref , the higher the score S i ; when V i exceeds the allowable interval, the score S i = 0.
[0082] S22. According to the bridge health scoring result, divide the current state level of the bridge: S c≥0.75, it is determined to be in a "healthy" state and no traffic control measures are required; 0.5 ≤ S c <0.75, it is determined to be in a "warning" state, and it is recommended to implement low-intensity control measures such as speed limit or weight limit; S c <0.5, it is determined to be in a "risk" state, and it is recommended to take comprehensive control measures such as speed limit, weight limit, and traffic diversion.
[0083] S23, dynamically update the bridge health score to ensure that the score can reflect the bridge health status in real time: dynamically adjust the weight W of the monitoring indicators according to the time variation characteristics of the monitoring indicators, expert experience, or historical impact analysis i ; the bridge health score model recalculates S according to a fixed time window c , and uses the latest score for traffic strategy optimization.
[0084] In step S3, the trend prediction of the bridge health status specifically includes the following steps:
[0085] S31, extract the historical bridge health scores within a fixed time window to form a time series data set:
[0086] D = {S c (t - n + 1), S c (t - n + 2),..., S c (t)}
[0087] Among them, t is the current time, and the sequence length is the fixed time window size n.
[0088] Divide the time series data set into a training set and a test set for model training and verification.
[0089] S32, model the bridge health score time series data through an LSTM network to capture the short-term and long-term dependencies of the bridge health score, and predict the bridge health scores for multiple future time steps:
[0090] {S p (t + 1), S p (t + 2), …, S p (t + k)}
[0091] Among them, S p (t + k) represents the predicted value of the bridge health score at the kth future time step.
[0092] S33, train and optimize the trend prediction model, and use the mean squared error (MSE) as the loss function to measure the prediction error:
[0093]
[0094] Among them, For the model to predict the health score, For the true health score, and N is the number of data samples.
[0095] Update the model parameters by the gradient descent method to minimize the loss function L.
[0096] S34. Use the trained LSTM model to input the latest time series data of the bridge health score, and predict the bridge health scores at multiple future time steps in real time. According to the predicted score results, further divide the health levels, which are used for risk interpretation, information display, and early warning prompts of the bridge management unit, and are not used as direct constraints for the action selection of the deep Q network.
[0097] In step S4, generating the optimal traffic control strategy based on DQN specifically includes the following steps:
[0098] S41. Define the state space. The state S is composed of the bridge structure state and the traffic state, including:
[0099] The current health score S of the bridge c ;
[0100] The predicted sequence S of the future health scores of the bridge within the next k time steps p (t + 1), S p (t + 2),..., S p (t + k); generated by the LSTM model, representing the change trend of the bridge health state within the next k time steps;
[0101] Traffic state data, including traffic flow density q, average vehicle speed v c , proportion r of heavy-duty vehicles l and average load w c ; In this embodiment, the traffic state data collected over a period of time is averaged or weighted to obtain the current traffic state data.
[0102] S42. Define the action space A. The action refers to the traffic control strategy, which is composed of different combinations of traffic control strategies, including speed limit strategies (such as speed limits of 70%, 50%, 30%), weight limit strategies (such as prohibiting vehicles over 30 tons or 40 tons from passing), and diversion strategies (no diversion, partial diversion, full diversion).
[0103] S43. Construct the reward function. The reward R is used to quantify the advantages and disadvantages of different traffic control strategies. Considering bridge protection, traffic efficiency, and control costs comprehensively, the expression of the reward function is:
[0104] R = -ΔS bridge +λ1·T eff -λ2·C ctrl
[0105] Among them, λ1 and λ2 are weight factors for adjusting the relative importance of traffic efficiency and control cost;
[0106] ΔS bridge represents the comprehensive health score of the bridge, reflecting the change trend of the bridge's health status at present and in a period of time in the future. The higher the value of ΔS bridge , the better the corresponding bridge health state. The goal of ΔS bridge is to be maximized, and the calculation method is:
[0107]
[0108] T eff represents the traffic efficiency under the traffic control strategy (such as average vehicle speed or the number of vehicles passing through per unit time). The goal of T eff is to be maximized, and the calculation method is:
[0109]
[0110] Among them, Q eff is the number of vehicles passing through per unit time, that is, the number of equivalent vehicles passing through per unit time; r l is the heavy load ratio, considering the structural burden brought by heavy-duty vehicles; δ is the penalty coefficient; q is the traffic flow density; T avg is the average passing time of vehicles; L is the total length of the bridge; v c is the average vehicle speed;
[0111] C ctrl represents the control cost of the traffic control strategy (such as economic losses caused by speed limits and traffic diversions). The goal of C ctrl is to be minimized, and the calculation method is:
[0112] C ctrl = β1·(v normal - v c ) + β2·(w normal - w c ) + β3r d
[0113] Among them, v normal is the designed vehicle speed; w normal is the expected average load of normal passing vehicles; r d is the diversion ratio; β1, β2, and β3 are loss conversion coefficients, which can be estimated by experience or adjusted by simulation.
[0114] S44. Input the state S into the Deep Q-Network (DQN) to output the Q-values for each action A. Use the ε-greedy policy to select an action: with a probability of ε, randomly select an action to explore new possibilities; with a probability of 1 - ε, select the action with the maximum current Q-value; execute the selected action A and feedback to calculate the reward R. Specifically, it is as follows:
[0115] Construct the DQN structure. Input the state S into the DQN to output the Q-values corresponding to each action A. The Q-value represents the reward that can be obtained by selecting action A in state S. The goal of the DQN is to find the strategy that maximizes the long-term cumulative reward by continuously updating Q(S,A).
[0116] Use the ε-greedy policy for action selection. With a probability of ε, randomly select action A to explore unattempted actions; with a probability of 1 - ε, select the action A with the maximum current Q-value.
[0117] Execute the selected action A and implement the traffic control strategy corresponding to action A during the actual operation of the bridge.
[0118] Record the current state S, the reward R obtained after executing the selected action A, and the new state S′ obtained after executing the selected action A to form a quadruple (S,A,R,S′), and store this quadruple (S,A,R,S′) as a training sample of the deep Q-network in the experience replay pool.
[0119] S45. The deep Q-network continuously updates the network parameters through the gradient descent optimization method, making the generated strategy more suitable for the current state of the bridge and the traffic target. Specifically, it is as follows:
[0120] (1) Sample training samples. Randomly draw several quadruples (S,A,R,S′) from the experience replay pool to form a training mini-batch; avoid the high correlation between consecutive states and improve the learning stability.
[0121] (2) For each training sample, calculate the target Q-value:
[0122]
[0123] where S′ is the new state after executing the current action A; A′ is the possible next action; γ is the discount factor, measuring the importance of future rewards; θ - are the target network parameters, whose values are from the periodic copy of the current network parameters θ, used to keep the target value stable and avoid the policy oscillation caused by the simultaneous update of the predicted value and the target value.
[0124] (3) The mean squared error (MSE) is used as the loss function to calculate the average loss value of all training samples, which serves as the training error of the current deep Q-network to measure the difference between the predicted value and the target Q-value of the current deep Q-network. The calculation method is as follows:
[0125]
[0126] Among them, represents the expectation (average) over multiple training samples; θ is the current network parameter of the deep Q-network, which is responsible for participating in policy training and updating in real time; N is the number of batch samples used for training, and each sample corresponds to a state S i , action A i and target Q-value y i .
[0127] (4) Every certain number of training cycles (such as every 100 steps), the parameters θ of the main network are synchronously copied to the target network θ - to improve the stability of Q-value estimation.
[0128] (5) As the training progresses, the deep Q-network will gradually converge. The trained deep Q-network is used to output the optimal action A for any state S * :
[0129]
[0130] This optimal action A * corresponds to the optimal traffic control strategy under the current state of the bridge, including the speed limit level, weight limit threshold, and diversion ratio.
[0131] Example 1: Real-time assessment and trend prediction of bridge health status
[0132] A variety of sensor devices are installed on an in-service bridge to collect bridge structure monitoring data and traffic status data in real time. The specific process is as follows:
[0133] 1. Data collection
[0134] Bridge structure monitoring data: acceleration a b = 110mm / s 2 , strain e b = 85με, lateral displacement x h = 3.2mm, deflection d = 4.5mm.
[0135] Traffic status data: average vehicle speed v c = 62km / h, traffic flow density q = 850 vehicles / hour, average load w c = 22 tons, heavy vehicle ratio r l = 15%.
[0136] 2. Bridge Health Score Calculation
[0137] According to the bridge structure health score calculation model, using the above monitoring indicators and their allowable ranges and reference values, the individual scores of each indicator are calculated respectively through the normal distribution scoring method. According to the preset weights (acceleration 35%, strain 25%, lateral displacement 20%, deflection 20%), the individual scores are weighted and calculated to obtain the current bridge health score. After calculation, the current bridge health score is 0.36. According to the bridge health score grading standard, it is determined that the current bridge is in a "risk" state.
[0138] 3. Bridge State Trend Prediction
[0139] To further judge the future health change trend of the bridge, the long short-term memory network (LSTM) is used to analyze the time series data of the bridge health score. The model uses the health score sequence of the past 1 hour as the training sample. The mean squared error (MSE) is used as the loss function during model training, and the parameters are optimized by the gradient descent method.
[0140] Through the trained model, predict the change of the bridge health score in the next 3 hours. The predicted score S p (t + 1) = 0.34, the predicted score S p (t + 2) = 0.33, the predicted score S p (t + 3) = 0.32.
[0141] Example 2. Optimization and Implementation of Traffic Control Strategy
[0142] Based on the bridge health score and trend prediction results of Example 1, use the deep Q-network (DQN) to generate traffic control strategies and implement them through specific measures to reduce the risk of fatigue damage to the bridge structure. The specific process is as follows:
[0143] 1. State Space Construction
[0144] Combining the current bridge score S c = 0.32 and the future trend prediction results {S p (t + 1) = 0.34, S p (t + 2) = 0.33, S p (t + 3) = 0.32}, as well as the traffic state data (average vehicle speed v c = 62 km / h, traffic flow density q = 850 vehicles / hour, average load w c = 22 tons, heavy truck proportion r l = 0.15), define the state space. The state S is modeled as a multi-dimensional vector:
[0145] S = [Sc ,S p (t + 1),S p (t + 2),S p (t + 3),v c ,q,w c ,r l
[0146] 2. Action Space Design
[0147] According to the bridge state and traffic flow characteristics, the action space includes multiple traffic control measures and their combinations:
[0148] Speed limit strategy: speed limit of 70%, speed limit of 50%, speed limit of 30%, no speed limit;
[0149] Weight limit strategy: prohibit vehicles over 40 tons, prohibit vehicles over 30 tons, prohibit vehicles over 20 tons, no weight limit;
[0150] Diversion strategy: diversion of 70%, diversion of 50%, diversion of 30%, no diversion;
[0151] Each action A in the action space i represents a specific combination, corresponding to a traffic control strategy plan. For example:
[0152] Action A1 = [speed limit of 50%, prohibit vehicles over 30 tons, diversion of 30%].
[0153] 3. Reward Function Construction
[0154] To achieve the comprehensive optimization of bridge health, traffic efficiency, and control cost, the following reward function is constructed:
[0155]
[0156] Among them, traffic efficiency T eff , according to the formula:
[0157]
[0158] Substitute L = 100m, δ = 1.5, q = 850, v c = 62, r l = 0.15, after calculation, the traffic efficiency is approximately 409.
[0159] Control cost C ctrl , according to the formula:
[0160] C ctrl = β1·(v normal - v c ) + β2·(w normal - w c ) + β3rd
[0161] Substitute the data v normal = 70 km / h, w normal = 30 t, v c = 35 km / h, w c = 20 t, r d = 30%, take β1 = 1.0 (1 unit cost is counted for every 1 km / h decrease in vehicle speed), β2 = 1.5 (1.5 unit costs are counted for every 1 ton decrease in weight limit), β3 = 100 (100 unit costs are counted for every 1% diversion). After calculation, the control cost is about 80.
[0162] Let the reward function weights λ1 = 0.6 and λ2 = 0.4.
[0163] 4. Policy Generation and Implementation
[0164] According to the current state and the above reward function, the deep Q-network calculates the Q-values for the action combinations and selects the optimal policy through the ε-greedy policy.
[0165] The optimal traffic control policy output in this round is:
[0166] A * = [Speed limit 50%, weight limit 20 tons, diversion 30%].
[0167] The system implements this policy through the traffic management platform via electronic induction screens, speed limit signs, weight limit notices, and road guidance systems to control vehicle passing behaviors.
[0168] Example 3. Feedback Optimization and Dynamic Adjustment
[0169] Based on Example 2, the traffic control policy is dynamically adjusted through the feedback optimization module to further improve the balance between bridge protection and traffic efficiency and achieve closed-loop adaptive optimization. The specific process is as follows:
[0170] 1. Data Feedback Collection and Analysis
[0171] The bridge score 1 hour after control rises from S c = 0.36 to S c = 0.40. Trend prediction shows that the score will remain stable in the next 3 hours, S p (t + 1) = 0.40, S p (t + 2) = 0.40, S p (t + 3) = 0.40; the traffic volume rises to q = 800 vehicles per hour, and the average vehicle speed increases to v c = 55 km / h. This indicates that while ensuring the structural safety of the bridge, the traffic operation efficiency has been improved.
[0172] 2. Policy Optimization and Update
[0173] After analyzing the feedback data, adjust the reward function weights to λ1 = 0.7 and λ2 = 0.3.
[0174] The updated optimal policy is: the speed limit is adjusted to 70%; the weight limit of 20 tons is maintained; the diversion ratio is reduced from 30% to 20%.
[0175] 3. Monitoring of Execution Effect
[0176] After 1 hour, the system feedback shows that the bridge score stabilizes at S c = 0.42, the predicted trend S p (t + 1) = 0.43, S p (t + 2) = 0.44, S p (t + 3) = 0.45, the vehicle speed rises back to v c = 60 km / h, the traffic flow recovers to q = 820 vehicles per hour, and the traffic pressure is effectively relieved.
[0177] Figure 4 This is the curve of the bridge health score and traffic efficiency change in the embodiment of the present invention, which reflects the dynamic change process of the bridge structure state and traffic operation state within 24 hours after the implementation of the traffic control strategy optimization. As shown in the figure, in the initial stage (0 hour), the bridge structure is subjected to greater stress and the traffic operation efficiency is limited. After implementing the traffic control strategy, the bridge health score rapidly increases within the first 4 hours, the structural risk is alleviated, and at the same time the traffic efficiency increases, indicating that the traffic operation condition is gradually improving; during the period of 5 - 10 hours, affected by the traffic backflow effect, the increase in traffic flow leads to an enhanced dynamic response of the bridge structure, the bridge health score shows a slight decline, and the traffic efficiency also fluctuates and decreases accordingly. At this time, the system dynamically adjusts the reward function weights according to the feedback data, updates the traffic control strategy, the traffic efficiency gradually recovers after adjustment, and the bridge health score tends to be stable after multiple rounds of optimization. Figure 4 It shows that the present invention can achieve a dynamic balance between the bridge structure safety and traffic efficiency through adaptive learning and optimization by combining the deep Q - network with the bridge health score, traffic efficiency, and traffic control cost.
[0178] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. An intelligent traffic control method based on bridge structure status and traffic status, characterized in that, It includes the following steps: S1. Collect bridge structure monitoring data and traffic status data in real time; S2. Evaluate the bridge structure monitoring data and calculate the current health score of the bridge in real time; S3. Use a long short-term memory network to predict the future health score of the bridge; S4. Define the state space of reinforcement learning according to the current health score and future health score of the bridge and the traffic status, define the action space of reinforcement learning based on the traffic control strategy, and use a deep Q-network in combination with the reward function to obtain the optimal traffic control strategy.
2. The intelligent traffic control method based on the bridge structure state and traffic state according to claim 1, wherein In step S1, the bridge structure monitoring data includes acceleration, strain, lateral displacement, and deflection; the traffic status data includes average vehicle speed, average load, heavy vehicle ratio, and traffic flow density.
3. An intelligent traffic control method based on bridge structure status and traffic status according to claim 1, characterized in that, In step S2, the current health score of the bridge is calculated by weighted calculation from multiple monitoring indicators according to the normal distribution model to quantify the current health status of the bridge. The calculation formula is as follows: Among them, S c is the current health score of the bridge, with a value range of [0, 1]. The higher the value of S c , the healthier the bridge state; W i is the weight of the i-th monitoring index in the bridge structure monitoring data, satisfying There are a total of n monitoring indexes in the bridge structure monitoring data; S i is the score of the i-th monitoring index, with a value range of [0, 1]. The calculation method is as follows: Among them, V i is the actual value of the i-th monitoring index; V i max is the upper allowable limit value of the i-th monitoring index; V i min is the lower allowable limit value of the i-th monitoring index; V i ref is the reference value of the i-th monitoring index; σ i is the standard deviation of the i-th monitoring index.
4. An intelligent traffic control method based on bridge structure status and traffic status according to claim 1, characterized in that, In step S3, the long short-term memory network is trained using the time series data of the bridge health score, and the trained long short-term memory network is used to predict the future health score of the bridge.
5. An intelligent traffic control method based on bridge structure status and traffic status according to claim 1, characterized in that, In step S4, the optimal traffic control strategy is generated based on the deep Q-network, which specifically includes the following steps: S41. Define the state space. The state S is composed of the bridge structure state and the traffic status, including: The current health score S of the bridge c ; The predicted sequence S of the future health scores of the bridge within the next k time steps p (t + 1), S p (t + 2),..., S p (t + k); Traffic state data, including traffic flow density q, average vehicle speed v c , proportion r of heavy-duty vehicles l and average load w c ; S42. Define the action space A, which is composed of a combination of traffic control strategies, including speed limit strategy, weight limit strategy, and diversion strategy; S43. Construct the reward function. The reward R is used to quantify the advantages and disadvantages of different traffic control strategies. Considering the bridge protection, traffic efficiency, and control cost of the traffic control strategy comprehensively, the expression of the reward function is: R = ΔS bridge + λ1·T eff - λ2·C ctrl where λ1 and λ2 are weight factors for adjusting the relative importance of traffic efficiency and control cost; ΔS bridge represents the comprehensive health score of the bridge. The higher the value of ΔS bridge , the better the corresponding health state of the bridge. The goal of ΔS bridge is to maximize: T eff represents the traffic efficiency under the traffic control strategy, and T eff is targeted to be maximized, and the calculation method is as follows: Among them, Q eff is the throughput per unit time, that is, the equivalent number of vehicles passing through per unit time; r l is the heavy load ratio, considering the structural burden brought by heavy-duty vehicles; δ is the penalty coefficient; q is the traffic flow density; T avg is the average passing time of vehicles; L is the total length of the bridge; v c is the average vehicle speed; C ctrl represents the control cost of the traffic control strategy, and C ctrl is targeted to be minimized and calculated as follows: C ctrl = β1·(v normal - v c ) + β2·(w normal - w c ) + β3r d Among them, v normal is the designed vehicle speed; w normal is the expected average load of normally passing vehicles; r d is the diversion ratio; β1, β2, and β3 are loss conversion coefficients; S44. Obtain the optimal traffic control strategy based on the deep Q-network, specifically as follows: Input the current state S into the deep Q-network, and output the Q values corresponding to each action A; the Q value represents the cumulative reward that can be obtained by selecting action A in state S; the goal of the deep Q-network is to find the action A that maximizes the long-term cumulative reward through continuous update; where the ε-greedy strategy is used for action selection, and an unattempted action A is randomly selected with a probability of ε; the action A with the largest current Q value is selected with a probability of 1 - ε; Execute the selected action A, and implement the traffic control strategy corresponding to action A during the actual operation of the bridge; at the same time, record the current state S, the reward R obtained after the execution of the selected action A, and the new state S' obtained after the execution of the selected action A, form a quadruple (S, A, R, S'), and store this quadruple (S, A, R, S') as a training sample of the deep Q-network in the experience replay pool.
6. The intelligent traffic control method based on the bridge structure state and traffic state according to claim 1, wherein, The deep Q-network continuously updates the network parameters through the gradient descent optimization method, specifically as follows: Sample training samples, and randomly extract several quadruples (S, A, R, S') from the experience replay pool to form a training mini-batch; For each training sample, calculate the target Q value: where S′ is the new state after the execution of the current action A; A′ is the possible next action; γ is the discount factor, measuring the importance of future rewards; θ - is the target network parameter, whose value comes from the periodic replication of the current network parameter θ; The mean squared error (MSE) is used as the loss function to calculate the average loss value of all training samples, which serves as the training error of the current deep Q-network to measure the difference between the predicted value and the target Q-value of the current deep Q-network. The calculation method is as follows: Among them, represents the expectation for multiple training samples; θ is the current network parameter of the deep Q-network; N is the number of batch samples used for training, and each sample corresponds to a state S i , action A i and the target Q value y i ; Every time a set training cycle passes, the current network parameter θ is synchronously copied to the target network parameter θ - ; The deep Q-network after training is used to output the optimal action A for any state S * : 。 7. An intelligent traffic control method based on bridge structure status and traffic status according to claim 1, characterized in that, If the future health score of the bridge shows a downward trend in its health status, the traffic control strategy is updated. Using the method of step S4, based on the current health score, future health score, and traffic status of the bridge, the deep Q-network is called to generate the optimal traffic control strategy. Otherwise, the current traffic control strategy is maintained.
8. An intelligent traffic control method based on bridge structure state and traffic state according to claim 1, characterized in that It further includes the following steps: S5, Optimize the traffic control strategy in real time according to the current health score, future health score, and traffic status of the bridge. S6, Dynamically feedback the changes in the bridge score and traffic efficiency after the implementation of the traffic control strategy, and update the traffic control strategy based on this feedback to achieve closed-loop optimization.
9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the intelligent traffic control method according to any one of claims 1 to 8, which is based on the bridge structure state and traffic state.
10. A computer program product, characterized in that, It includes a computer program / instructions. When the computer program / instructions are executed by the processor, it implements the intelligent traffic control method according to any one of claims 1 to 8, which is based on the bridge structure state and traffic state.
Citation Information
Cited By
Bridge load intelligent management method and system based on artificial intelligence, and medium
CN120975576A