A signal global induction real-time control method and system
By collecting traffic data in real time and combining it with fuzzy evaluation and reinforced game theory models to calculate intersection credit values, the scope of cooperation and signal duration allocation are dynamically adjusted, which solves the problem of insufficient global coordination in existing technologies and improves the operational efficiency and resource utilization efficiency of the road network.
Patent Information
- Application Number
- CN202511484638.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-10-17
AI Technical Summary
Existing technologies, when dealing with road network environments with dynamically changing traffic flow, do not fully consider the traffic flow correlation between intersections across the entire network. This results in a lack of global coordination in signal control, an imbalance in the allocation of green light resources, and causes localized congestion or low traffic efficiency.
Traffic status data is collected in real time by radar detectors, status values are generated using fuzzy comprehensive evaluation method, intersection credit values are calculated by combining multi-stage reinforced game model, and the cooperative range is dynamically adjusted by cooperative suppression distance field mechanism. Signal duration is allocated by combining three-level arbitration strategy, and credit values are dynamically updated to achieve global sensing real-time control.
It has achieved coordination and flexibility in traffic signal control at all intersections across the network, optimized the overall traffic stability and green light resource utilization efficiency of the road network, avoided imbalance in signal resource allocation, and improved the operational efficiency of the road network.
Smart Images

Figure CN120977130B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic digital data processing, and in particular to a real-time control method and system for global signal sensing. Background Technology
[0002] In scenarios where digital data processing technology is applied to traffic signal control, preset timing schemes or adaptive adjustment techniques based on local traffic data are often used to address the signal scheduling needs of multiple intersections within a road network. These technologies collect traffic state data from individual or local intersections, combine this data with fixed algorithms to calculate and adjust signal light durations, thereby achieving basic traffic coordination between intersections. Their technical characteristics include reliance on pre-set time period parameters or data feedback from a single intersection, with data processing focusing primarily on local traffic flow analysis. Signal control commands are generated through simple logical judgments, making them suitable for scenarios with relatively stable traffic flow and capable of ensuring basic traffic order within the road network to a certain extent.
[0003] However, existing technologies have limitations in dealing with dynamically changing road network environments. This problem stems from the fact that the technology design does not fully consider the traffic flow correlation between intersections across the entire network. Data processing is limited to traffic information at a single intersection or within a fixed area, failing to capture the dynamic changes and mutual influences of traffic conditions across the entire network in real time. This results in a lack of global coordination in signal control strategy adjustments. Consequently, some intersections may not receive signal durations adapted to the overall traffic flow, leading to an imbalance in green light resource allocation, causing localized congestion or low traffic efficiency, and affecting the overall smoothness of the road network. Solving this problem can effectively improve the coordination and flexibility of traffic signal control across the entire network, optimizing the overall traffic stability and green light resource utilization efficiency of the road network. Summary of the Invention
[0004] This invention provides a real-time control method and system for global signal sensing, which solves the technical problems mentioned above.
[0005] The technical solution of this invention to solve the above-mentioned technical problems is as follows: A real-time control method for global signal sensing, comprising the following steps:
[0006] Step 1: Collect traffic status data in real time using radar detectors, and generate status values from the traffic status data using a fuzzy comprehensive evaluation method;
[0007] Specifically, the process of generating state values using the fuzzy comprehensive evaluation method includes:
[0008] Based on the traffic status data collected by the radar detectors, the status values are obtained through fuzzy comprehensive evaluation after processing.
[0009] The traffic status data includes traffic flow, queue length, phase saturation, and intersection level. After time alignment and anomaly removal, it is normalized to form a standardized dataset. Then, a four-level equal-width triangular membership matrix is constructed based on the standardized dataset to obtain the membership matrix.
[0010] The comprehensive evaluation vector is calculated based on the membership matrix and the equal weight vector. Finally, the comprehensive evaluation vector is defuzzified by performing an inner product with the preset level score vector to generate a scalar state value.
[0011] Step 2: Using state values and historical compromise rates, calculate the credit value for each intersection based on a multi-stage reinforcement game model;
[0012] Specifically, the multi-stage reinforcement game model is as follows:
[0013] The cooperative utility is calculated by weighting the state value and the historical compromise rate; the basic benefit is calculated based on the state value archived in the previous control cycle and the state value in the current cycle, and the potential function is constructed by combining the state value and the historical compromise rate; when switching cycles, a reward shaping term is formed based on the potential function of the next cycle and the current cycle, and the reward shaping term is added to the basic benefit to obtain the immediate benefit.
[0014] A strategy value table is established and updated based on immediate benefits. The item with the highest value in the current period is selected as the long-term strategy value. Based on the long-term strategy value, a greedy rule with an exploration rate is used to determine the behavior selection. Finally, the credit value of each intersection is obtained by weighted synthesis based on the cooperation utility and the long-term strategy value.
[0015] Step 3: Combining the credit value and physical distance of each intersection, a cooperative inhibition distance field mechanism is used to dynamically adjust the cooperation range;
[0016] Specifically, the dynamic adjustment of the cooperative range using the cooperative suppression distance field mechanism includes:
[0017] Step 31: Combine the credit value and physical distance of each intersection to calculate the cooperation inhibition distance between intersections;
[0018] Step 32: Set the cooperation threshold and dynamically adjust the effective cooperation range of different intersections based on the calculated cooperation inhibition distance and cooperation threshold.
[0019] Step 4: Determine the intersections participating in the decision-making process based on the cooperation scope; use the credit values of the participating intersections to obtain the arbitration result through a three-level arbitration strategy, and allocate signal duration based on the arbitration result;
[0020] Specifically, the arbitration outcome obtained through the three-tier arbitration strategy includes:
[0021] Step 41: Based on the cooperative inhibition mechanism, select the intersections participating in the decision-making process and define them as candidate intersections;
[0022] Step 42: Calculate the green wave weight by combining the credit scores and phase urgency of the candidate intersections involved in the decision-making process;
[0023] Step 43: Obtain the dominant difference ratio, maximum proportion, and relative dispersion based on the green wave weights, and combine them with the three-level arbitration strategy to obtain the arbitration result;
[0024] Step 44: Based on the arbitration result, output the final signal duration allocation and notify each intersection to execute it.
[0025] Specifically, the types of arbitration resulting in the arbitration award include: full release arbitration, proportional allocation arbitration, and equal allocation arbitration.
[0026] The full release arbitration is as follows: when the dominant difference ratio is greater than or equal to the upper threshold of the dominant difference ratio and the maximum proportion is greater than or equal to the upper threshold of the maximum proportion, it is determined to be significantly leading and a full release strategy is adopted.
[0027] The proportional allocation arbitration is as follows: when the full release condition is not met, and the relative dispersion is not lower than the relative dispersion threshold, or the dominant difference ratio is between the lower threshold and the upper threshold of the dominant difference ratio, it is determined that the difference is large and a proportional allocation strategy is adopted.
[0028] The equal-division arbitration is as follows: when the dominant difference ratio is lower than the lower threshold of the dominant difference ratio and the relative dispersion is lower than the relative dispersion threshold, it is determined that the difference is very small, and the equal-division strategy is adopted.
[0029] Specifically, the full release strategy, the proportional allocation strategy, and the equal distribution strategy are as follows:
[0030] Full-allowance strategy: Set the green light duration of the intersection with the highest weight to the product of the full-allowance ratio and the total signal cycle, and set the green light duration of the other candidate intersections to 0.
[0031] Proportional allocation strategy: The green light duration of each candidate intersection is equal to the ratio of the green wave weight of that candidate intersection to the sum of the weights, multiplied by the total signal cycle.
[0032] Equal distribution strategy: Allocate equally according to the number of candidate intersections, and the green light duration of each candidate intersection is equal to the ratio of the total signal cycle to the number of candidate intersections;
[0033] Step 5: Dynamically update the credit score of each intersection based on the arbitration result;
[0034] Step 6: Calculate the standard deviation of the credit values of all intersections in the network and input it into the credit entropy stabilization mechanism to obtain the stability optimization strategy; count the number of times each intersection yields to other intersections in the past period, calculate the historical compromise rate and output it to Step 2.
[0035] Specifically, the stability optimization strategy includes the following steps:
[0036] Step 61: Calculate the standard deviation of the credit scores for all intersections across the network;
[0037] Step 62: Input the standard deviation of the credit score into the credit entropy stabilization mechanism, generate an optimization strategy based on the credit entropy value, and trigger system stability adjustment;
[0038] Specifically, the credit entropy stabilization mechanism is as follows:
[0039] Based on the standard deviation of credit scores, a baseline entropy value is set by combining the statistical mean of the standard deviation of credit scores during stable traffic periods across the entire network. The deviation between the standard deviation of credit scores and the baseline entropy value is calculated, and the deviation is substituted into the entropy value transformation function to obtain the credit entropy value. The stability threshold range is determined by back-calculation based on traffic failure cases across the entire network, and a stability optimization strategy is generated by combining the credit entropy value with threshold judgment rules. If the credit entropy value is less than or equal to the lower limit threshold, a credit equilibrium strategy is generated; if the credit entropy value is greater than or equal to the upper limit threshold, a credit convergence strategy is generated; if the credit entropy value is within the threshold range, a credit maintenance strategy is generated to trigger the system to dynamically adjust the credit score update rules for each intersection.
[0040] Step 63: Count the number of times the intersection yielded to other intersections over a period of time, calculate the historical compromise rate, and feed it back to Step 2;
[0041] A real-time signal global sensing control system specifically includes the following modules:
[0042] Traffic status data acquisition and evaluation module: Real-time traffic status data is acquired through radar detectors, and the traffic status data is used to generate status values using a fuzzy comprehensive evaluation method;
[0043] Intersection Credit Score Preliminary Calculation Module: Using state values and historical compromise rates, the credit score of each intersection is calculated based on a multi-stage reinforcement game model.
[0044] Dynamic adjustment module for cooperation range: Combining the credit value and physical distance of each intersection, the cooperation range is dynamically adjusted using a cooperation inhibition distance field mechanism;
[0045] Decision intersection determination and signal duration allocation module: determines the intersections participating in the decision based on the cooperation scope; uses the credit value of the participating intersections to obtain the arbitration result through a three-level arbitration strategy, and allocates signal duration based on the arbitration result;
[0046] Intersection Credit Value Dynamic Update Module: Dynamically updates the credit value of each intersection based on the arbitration results;
[0047] Credit Entropy Stabilization and Historical Compromise Rate Statistics Module: Calculates the standard deviation of the credit value of all intersections in the network and inputs it into the credit entropy stabilization mechanism to obtain the stability optimization strategy; counts the number of times each intersection yields to other intersections in the past period, calculates the historical compromise rate and outputs it to the intersection credit value initial calculation module.
[0048] Compared with the prior art, the beneficial effects of the present invention are:
[0049] This invention collects real-time road network operation status data through detectors, generates status values by combining fuzzy comprehensive evaluation methods, calculates node credit values based on a multi-stage reinforced game model, and dynamically adjusts the cooperation range by adopting a cooperative suppression distance field mechanism. This effectively solves the problem of lack of global coordination in signal control caused by insufficient consideration of the correlation of the operation status of all network nodes and data processing being limited to local information in existing technologies.
[0050] This method integrates multi-dimensional information such as real-time node operating status, historical compromise rate, and physical distance to construct a complete control logic from data collection and evaluation to credit value calculation, cooperation range adjustment, and signal duration allocation. It overcomes the limitations of traditional technologies that rely on preset parameters or local data and struggle to adapt to dynamic changes in node operating status within the road network. Based on a three-level arbitration strategy and a credit entropy stabilization mechanism, this invention achieves precise selection of participating decision-making nodes and reasonable allocation of signal duration. While ensuring coordinated control of all nodes in the network, it also ensures system stability. This not only avoids imbalances in signal resource allocation within the road network and reduces localized traffic inefficiencies, but also significantly improves the overall operating efficiency and resource utilization efficiency of the road network, aligning with the actual application needs of road network signal control. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 This is a flowchart of a real-time global signal sensing control method according to the present invention;
[0053] Figure 2 This is a functional block diagram of a global signal sensing real-time control system according to the present invention. Detailed Implementation
[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0055] Example 1:
[0056] Please see Figure 1 As shown, this embodiment provides a cross-configuration method for a global signal sensing real-time control method module, including:
[0057] Step 1: Collect traffic status data in real time using a radar detector, and generate status values from the traffic status data using a fuzzy comprehensive evaluation method.
[0058] Specifically, the fuzzy comprehensive evaluation method acquires a traffic state dataset within a unified sampling period, including four indicators: traffic flow, queue length, phase saturation, and intersection grade. Traffic flow is the number of vehicles passing the stop line per unit time; queue length is the actual distance from the tail of the vehicle queue to the stop line at any given moment; phase saturation is the ratio of the actual number of vehicles allowed to pass during a green light period to the theoretical capacity of the lane; and intersection grade is determined by road network design data based on road function and traffic organization. After time alignment and outlier removal, these four indicators are normalized to the [0, 1] interval, with each data point being a standardized value x, forming a standardized dataset to ensure that indicators with different dimensions have a unified measurement.
[0059] Based on the standardized dataset, a four-level equal-width triangular membership matrix is constructed. Specifically, each standardized value x is mapped to four levels: {smooth, normal, congested, and severely congested}. The anchor points are {0, 1 / 3, 2 / 3, 1}, and the bandwidth is 1 / 3. Each standardized value x is matched with one of the four anchor points. The membership degree is calculated using the formula: membership degree = max(0, 1-3×|x-anchor|), where max(0, 1-3×|x-anchor|) is the maximum value between 0 and 1-3×|x-anchor|.
[0060] Each item is calculated and summarized into a membership matrix, which is a 4×4 two-dimensional array. The rows are arranged in order according to the four indicator sets (traffic flow, queue length, phase saturation, intersection level), and the columns are arranged in order according to the level set (smooth, normal, congested, severe congested). The matrix elements take values in the range [0, 1], representing the membership degree of the corresponding indicator and level.
[0061] A comprehensive evaluation vector is obtained by combining the membership degree matrix with an equal-weight vector. Specifically, the equal-weight vector is set so that the weights of the four indicators are equal at 0.25. The weighted sum of the four levels is calculated by multiplying the weight vector by the membership degree matrix, and the four-dimensional comprehensive evaluation vector is output.
[0062] For example, the standardized dataset = {Traffic flow standardized value = 0.75, queue length standardized value ≈ 0.636, phase saturation standardized value ≈ 0.8125, intersection grade standardized value = 1.00}.
[0063] A membership matrix is obtained by mapping four-level equal-width triangular membership degrees to a standardized dataset. Specifically, the membership degrees of each indicator on {smooth, normal, congested, severe congestion} are as follows:
[0064] Traffic flow (0.75): {0, 0, 0.75, 0.25};
[0065] Queue length (0.636): {0, 0.09, 0.907, 0};
[0066] Phase saturation (0.8125): {0, 0, 0.562, 0.437};
[0067] Intersection level (1.00): {0, 0, 0, 1}.
[0068] The resulting matrix is [0, 0, 0.75, 0.25; 0, 0.09, 0.907, 0; 0, 0, 0.562, 0.437; 0, 0, 0, 1].
[0069] The state value is obtained by defuzzifying the comprehensive evaluation vector. Specifically, the grade rating vector is set to {low, lower, higher, highest} with an increasing scale of four monotonically increasing points within [0, 1]. For example, [0, 0.25, 0.5, 1] is used. The single scalar state value is calculated by the inner product of the grade rating vector and the comprehensive evaluation vector.
[0070] Step 2: Using state values and historical compromise rates, calculate the credit value for each intersection based on a multi-stage reinforcement game model;
[0071] Specifically, the multi-stage reinforcement game model includes:
[0072] The cooperation utility is calculated by weighting the state value and historical compromise rate, as follows: Cooperation Utility = Coefficient a × State Value + Coefficient b × Historical Compromise Rate, where coefficients a and b are constants used to balance the impact of immediate congestion on long-term cooperation. Coefficients a and b are determined during the deployment phase using historical data and the least squares method. Specifically, a linear regression is performed using the baseline return within the historical window as the dependent variable and the state value and historical compromise rate as independent variables. The non-negative parts of the regression coefficients are taken and normalized to a sum of 1, then solidified into coefficients a and b.
[0073] The base benefit is calculated using the state values archived from the previous control cycle and the state values of the current cycle. Base benefit = current state value - previous cycle state value, with negative values counted as zero. A potential function is constructed based on cooperative utility: potential function = coefficient c × state value + coefficient d × historical compromise rate. Coefficients c and d are separately calibrated during the deployment phase by searching for the value with the smallest oscillation amplitude within a finite candidate grid and fixing it as a universal constant for the entire network.
[0074] During cycle switching, the potential function recalculated from the next cycle and the potential function of the current cycle form a reward shaping term;
[0075] The reward shaping term = discount factor × potential function of the next period - potential function of the current period, where the discount factor is a constant used to control the weight of future returns in the current estimate. The discount factor is determined by setting the discount factor based on the expected number of convergence periods M, discount factor = 1 - 1 / M, and the optimal traffic efficiency index is selected by verifying historical data on a small-scale grid. The base return is added to the reward shaping term to obtain the immediate return, immediate return = base return + reward shaping term.
[0076] Based on immediate benefits, a strategy value table is established and updated through multiple stages. Specifically, the strategy value table is a numerical mapping of a fixed set of cooperation intensity levels, including weak cooperation, medium cooperation, and strong cooperation. For the currently selected cooperation level, the strategy value is updated according to the formula: New Strategy Value = (1 - Learning Rate) × Old Strategy Value + Learning Rate × (Maximum value among Immediate Benefit + Discount Factor × Old Strategy Value). Unselected levels are maintained only by a decay term, which is equal to 1 - Learning Rate. After the update, the strategy value table is output. The learning rate is a constant used to control the coverage strength of new observations on existing estimates. The learning rate is selected during the deployment phase through rolling validation, choosing the value that minimizes the fluctuation of state values.
[0077] The maximum entry in the strategy value table for the current period is used as the long-term strategy value, and this long-term strategy value is output. To improve exploration capabilities, a greedy rule with an exploration rate is adopted during the execution phase for behavioral selection: with the exploration rate as a constant, a non-maximum entry is randomly selected with a low probability to collect new information; otherwise, the maximum entry is selected. This selection only affects the formation of immediate benefits and reward shaping items in the next period and does not change the long-term strategy value already output in the current period. The exploration rate is initially determined and its lower limit is set during the deployment phase using small-grid validation of historical data, and decays periodically during the runtime using linear annealing to ensure sufficient exploration in the early stages and stable convergence in the later stages.
[0078] A credit score is obtained by weighting the cooperation utility and long-term strategy value. Specifically, the credit score = trade-off coefficient × cooperation utility + (1 - trade-off coefficient) × long-term strategy value, where the trade-off coefficient is a constant ranging from 0 to 1. During the deployment phase, the trade-off coefficient is selected to minimize the penalty for mismatch between average queue length and phase urgency. The credit score for each intersection is then output.
[0079] Step 3: Combining the credit score and physical distance at each intersection, a cooperative suppression distance field mechanism is used to dynamically adjust the cooperation range. Specifically, the cooperative suppression distance field mechanism includes the following steps:
[0080] Step 31: Combine the credit value and physical distance of each intersection to calculate the cooperation inhibition distance between intersections.
[0081] Specifically, the cooperation inhibition distance between each pair of intersections is calculated based on the credit score of each intersection and the physical distance between them. The formula for calculating the cooperation inhibition distance is:
[0082] Cooperative inhibition distance = ;
[0083] Wherein, physical distance represents the actual distance between the two intersections, credit value i and credit value j represent the credit values of intersection i and intersection j respectively, |credit value i - credit value j| represents the difference between the credit values of the two intersections, and 0.001 is used to avoid division by zero errors in the calculation and to ensure the smoothness of the calculation results.
[0084] This calculation yields the cooperation inhibition distance between each pair of intersections. If the credit values of two intersections differ significantly, the cooperation inhibition distance will increase, reflecting a lower cooperation potential between the two intersections, and the system will be more inclined to ignore their cooperation requests. Conversely, if the credit values of two intersections are similar, the cooperation inhibition distance will be smaller, indicating a higher probability of cooperation between them, and the system will prioritize their cooperation.
[0085] Step 32: Set the cooperation threshold and dynamically adjust the effective cooperation range of different intersections based on the calculated cooperation inhibition distance and the cooperation threshold;
[0086] Alternatively, the collaboration threshold can be set as follows:
[0087] Collaboration threshold = α × maximum physical distance + β × |credit score difference|;
[0088] The maximum physical distance is the maximum physical distance between intersections in the system. This parameter ensures that cooperation between distant intersections is appropriately ignored. Credit score difference refers to the absolute difference between the credit scores of two intersections, adjusted by a weighted coefficient β to ensure that intersections with large credit score differences are less likely to cooperate. α and β are weighting coefficients used to adjust the degree of influence of physical distance and credit score difference. Increasing α and β increases the cooperation threshold.
[0089] Optionally, the cooperation threshold can be dynamically adjusted based on factors such as traffic flow, intersection load, and real-time traffic conditions. When traffic flow in the area is high, the cooperation threshold is lowered so that more intersections can participate in cooperation, thereby speeding up the response of signal dispatching. When intersection load is high, the cooperation threshold is increased to prevent low-credit intersections from occupying too many resources.
[0090] When the calculated cooperation suppression distance exceeds the set cooperation threshold, the system will ignore intersections that do not meet the conditions, ensuring that there is sufficient cooperation potential between cooperating intersections. Specifically, intersections will only be included in the effective cooperation range if the cooperation suppression distance between them is less than the set cooperation threshold. If the conditions are not met, these intersections will be excluded and cannot participate in subsequent decision-making and signal duration allocation.
[0091] In this way, the system can dynamically adjust the cooperative relationship between intersections, making traffic signal scheduling more intelligent and flexible, ensuring that intersections with high credit and moderate traffic load participate in decision-making first, and improving overall traffic efficiency.
[0092] Step 4: Determine the intersections participating in the decision-making process based on the cooperation scope; use the credit values of the participating intersections to obtain an arbitration result through a three-level arbitration strategy, and allocate signal duration based on the arbitration result. Specifically, Step 4 includes the following steps:
[0093] Step 41: Based on the cooperative inhibition mechanism, select intersections that meet the conditions as candidate intersections to participate in the decision-making process;
[0094] The cooperation inhibition distance between any two intersections is compared with the cooperation threshold, and the credit value of the corresponding intersection is verified against the credit threshold. The intersections participating in the decision-making are then output for subsequent arbitration and timing.
[0095] Specifically, the cooperation inhibition mechanism filters candidate intersections in the order of "distance first, then credit": For each pair of intersections, denoted as intersection i and intersection j, when the cooperation inhibition distance (i,j) is less than or equal to the cooperation threshold, the credit values of these two intersections are continuously verified to be greater than or equal to the credit threshold; only when both "cooperation inhibition distance (i,j) ≤ cooperation threshold" and "credit value i ≥ credit threshold, credit value j ≥ credit threshold" are intersection i and intersection j marked as eligible to participate; all intersections marked as eligible to participate are aggregated to obtain the intersections participating in the decision-making process. Here, credit value i and credit value j represent the credit values of intersections i and j, respectively; the credit threshold is determined by the median of the credit value set of all intersections in the current period, i.e., the value at the middle position when sorting the credit values of the entire network by size, and is automatically updated with each control period.
[0096] In summary, the criteria for selecting candidate intersections to participate in the decision-making process are as follows:
[0097] The intersections involved in the decision-making process are defined as follows: {Intersection i | There exists intersection j such that (cooperative inhibition distance (i,j) ≤ cooperative threshold) and (credit value i ≥ credit threshold) and (credit value j ≥ credit threshold)}.
[0098] When the summary result is empty, to ensure that the decision can be executed, the intersection with the highest credit value is selected from all intersections in the network in descending order of credit value as the intersection to participate in the decision; when the summary result is not empty, the intersection to participate in the decision is directly output and passed to subsequent steps.
[0099] Step 42: Calculate the green wave weight for each intersection by combining the credit value and phase urgency of the participating intersections;
[0100] Specifically, the calculation of green wave weights is based on two factors: the intersection's weighted credit value and the phase urgency. The weighted credit value reflects the intersection's cooperation potential and traffic flow, while the phase urgency indicates the urgency at which the intersection needs a green light signal.
[0101] The weighted credit score is assessed based on the intersection's historical cooperation record and traffic flow. Specifically, the weighted credit score can be calculated as follows:
[0102] Weighted credit score = In this formula, credit value i is the credit value of intersection i, N is the total number of intersections involved in the decision-making process, and i and k are the index values of the number of intersections. This formula standardizes the credit values of all intersections, making the credit weight of each intersection proportional to its credit value.
[0103] Phase urgency refers to the degree of urgency at an intersection to require a green light due to traffic congestion or other factors. Phase urgency is calculated as the ratio of the number of vehicles queuing at intersection i to the maximum queue length at that intersection under worst-case conditions.
[0104] The final green wave weight combines the weighted credit score and the phase urgency score. Specifically, the green wave weight can be obtained using the following formula:
[0105] Green wave weight = β1 × weighted credit value + β2 × phase urgency;
[0106] β1 and β2 are weighting coefficients used to balance the influence of credit score and phase urgency on the green wave weight. Depending on the specific application scenario, these two coefficients can be determined through experimentation or system optimization. For example, if the system needs to prioritize traffic flow, a higher β1 can be set; conversely, a higher β2 can be set to address traffic congestion.
[0107] This method allows the system to dynamically assess the priority of each intersection, rationally allocate green light durations, and ensure smooth traffic flow. The green wave weight at each intersection serves as a crucial input in the execution of the three-level arbitration strategy, ultimately determining the allocation of signal durations.
[0108] Step 43: Based on the green wave weights, obtain the dominant difference ratio, maximum proportion and relative dispersion used to determine the type of arbitration, and obtain the arbitration result based on the three-level arbitration strategy of full release arbitration, proportional allocation arbitration and equal allocation arbitration;
[0109] Specifically, using the output green wave weight set, the total weight, maximum weight, second-largest weight, average weight, and standard deviation of the weights are first calculated. Based on these, the dominant difference ratio, maximum proportion, and relative dispersion used to determine the arbitration type are obtained. The total weight is equal to the sum of the green wave weights at each intersection; the maximum weight is the maximum value in the green wave weight set; the second-largest weight is the maximum value in the green wave weight set excluding the maximum value; the average weight is equal to the total weight divided by the number of intersections; and the standard deviation of the weights is the square root of the mean square deviation of the green wave weights at each intersection from the average weight. Therefore, the dominant difference ratio is defined as: (maximum weight - second-largest weight) ÷ total weight; the maximum proportion is defined as: maximum weight ÷ total weight; and the relative dispersion is defined as: standard deviation of the weights ÷ average weight. The total green light duration that can be allocated in this cycle is represented by the total signal cycle.
[0110] Specifically, the three-tier arbitration strategy is implemented according to the following threshold rules:
[0111] (a) Full approval: When the dominant difference ratio is greater than or equal to the upper threshold of the dominant difference ratio and the maximum proportion is greater than or equal to the upper threshold of the maximum proportion, it is judged as significantly leading and the full approval strategy is adopted;
[0112] (b) Proportional allocation: When the full release condition is not met, if the relative dispersion is greater than or equal to the relative dispersion threshold, or the dominant difference ratio is greater than or equal to the upper threshold of the dominant difference ratio and less than the lower threshold of the dominant difference ratio, it is determined that the difference is large and a proportional allocation strategy is adopted.
[0113] (c) Equal Distribution: When the dominant difference ratio is less than the lower threshold of the dominant difference ratio and the relative dispersion is less than the relative dispersion threshold, the difference is determined to be very small, and an equal distribution strategy is adopted. For example, the upper threshold of the dominant difference ratio can be set to 0.35, the lower threshold of the dominant difference ratio can be set to 0.20, the upper threshold of the maximum proportion can be set to 0.60, and the relative dispersion threshold can be set to 0.18.
[0114] Specifically, based on the arbitration type, the following strategies are generated: full release strategy, proportional allocation strategy, and equal distribution strategy:
[0115] Full-allowance strategy: The green light duration of the intersection with the highest weight is set to the product of the full-allowance ratio and the total signal cycle, while the green light duration of other candidate intersections is set to zero. The full-allowance ratio is a configurable variable representing the percentage of time granted to the dominant intersection under full-allowance conditions. For example, the full-allowance ratio can be set to 0.70.
[0116] Proportional allocation strategy: The green light duration of each candidate intersection is determined by the ratio of the green wave weight of that intersection to the sum of the weights, multiplied by the total signal cycle. The weight ratio is the ratio of the green wave weight of that intersection to the sum of the weights, used to characterize the relative importance of a single intersection in the candidate set.
[0117] Equal distribution strategy: The number of candidate intersections is allocated equally, and the green light duration of each candidate intersection is equal to the ratio of the total signal cycle to the number of candidate intersections; where the ratio of the total signal cycle to the number of candidate intersections is used to replace the division operation to express the equal distribution duration.
[0118] The green light durations for each intersection calculated from the above three scenarios are compiled into an intersection green light duration allocation table, which serves as the arbitration result of this step.
[0119] Step 44: Based on the arbitration result, output the final signal duration allocation and notify each intersection to execute it.
[0120] Specifically, a candidate duration table is generated by combining the intersection green light duration allocation table with the arbitration type and the total duration of a cycle, and then the candidate duration table is output. The total duration of a cycle refers to the total time available for each phase's operation and phase transition within a single control cycle.
[0121] A constraint check was performed using the candidate duration table combined with the phase transition time and the minimum pedestrian duration threshold to obtain a restricted duration table. The phase transition time is a fixed time configuration for the yellow and all-red lights at each intersection; the minimum pedestrian duration threshold is calculated as "crossroad width ÷ design walking speed + clearing time", where clearing time refers to the buffer time required for pedestrians who have entered the crossroads to complete crossing the street after the signal ends. Specifically, for each intersection in the candidate duration table: if the candidate green light duration is lower than the minimum pedestrian duration threshold, a lower limit is applied to raise it to the minimum pedestrian duration threshold; then an upper limit is set, defining the upper limit of the green light duration as "total duration of one cycle × maximum green time ratio - phase transition time of the intersection". The maximum green time ratio refers to the maximum proportion of a single phase's green light duration allowed to occupy in the total duration of a cycle. This ratio is set by the traffic control center based on historical traffic flow statistics and simulation optimization results. The lower limit ensures that the green light duration at each intersection is not lower than the minimum pedestrian time threshold, thus guaranteeing pedestrian safety when crossing the street. The upper limit ensures that the green light duration at each intersection does not exceed the upper limit jointly defined by the total duration of a cycle and the maximum green time ratio, thus preventing a single intersection from occupying too much green light time and affecting the overall cycle balance.
[0122] When a candidate green light duration exceeds the upper bound of the green light duration, it is taken as the upper bound of the green light duration. When the arbitration type is full clearance and the green wave weight corresponding to the intersection is the highest in the set, only the lower limit is checked and the upper limit is not executed. After completing the constraints, the restricted duration table is output as the input for the total cycle conservation correction.
[0123] The total allocatable green light duration is calculated and output using the restricted duration table and phase transition time. Specifically, the total allocatable green light duration = total duration of one cycle - sum of phase transition times for the entire network. The remaining time difference is then obtained by subtracting the sum of green light durations for all intersections in the restricted duration table from the total allocatable green light duration, and this remaining time difference is output. A conservation correction is then performed using the restricted duration table, the remaining time difference, and phase urgency, to obtain the conserved duration table.
[0124] Specifically, when the remaining time difference is positive, the increase ΔTi at each intersection is calculated using the formula:
[0125] ΔTi = (Phase urgency i ÷ ΣPhase urgency j) × Remaining time difference;
[0126] Distribute the time proportionally among intersections that have not reached the upper limit until the remaining time difference is zero or can no longer be increased;
[0127] When the remaining time difference is negative, the reduction ΔTi at each intersection is calculated using the formula:
[0128] ΔTi = (Amount that can be reduced i ÷ ΣAmount that can be reduced j) × |Remaining time difference|;
[0129] The amount that can be reduced, i = current green light duration - minimum pedestrian duration threshold (non-negative part), is reduced proportionally at each intersection until the total amount is conserved or can no longer be reduced.
[0130] Among them, phase urgency refers to the intensity of demand for additional green light time for the current traffic phase at each intersection, which is usually calculated based on real-time queue length, vehicle saturation, or delay estimation results.
[0131] The symbol i represents the target intersection number currently being calculated, the symbol j represents the set index of all intersections in the entire network, and Σ phase urgency j represents the summation of the phase urgency of all intersections.
[0132] The conserved duration table is first discretized to eliminate rounding errors, resulting in a discrete conserved duration table. Specifically, discretization involves rounding the duration of each intersection to the second level. The total difference generated by rounding is then used to add or subtract one second to each intersection according to phase urgency, from highest to lowest, until the total duration matches the total allocable green light duration. Next, the discrete conserved duration table is combined with the previous cycle duration allocation table and the change limit ratio to perform cross-cycle jitter suppression, resulting in a variable constraint duration table. The change limit ratio is a system parameter used to limit the maximum relative change in green light duration between two adjacent cycles for a single intersection. This ratio is pre-configured by the system based on road grade, historical traffic fluctuation characteristics, and control stability requirements. Specifically, the previous cycle duration allocation table is the execution record of the previous control cycle. For each intersection, the green light duration in the discrete conserved duration table is limited to the interval [previous cycle duration - change limit ratio × previous cycle duration, previous cycle duration + change limit ratio × previous cycle duration]. The difference caused by the restrictions is then proportionally distributed or reduced among the intersections that have not yet reached the restrictions, based on the phase urgency, until the total amount is conserved. The final output is a table of change constraint durations, which serves as input for safety and consistency verification.
[0133] Safety and consistency checks are performed using the variable constraint duration table to obtain the check results. Specifically, the safety check includes: verifying that there is no time overlap or conflict between the green light duration and phase transition time at each intersection within a total cycle duration; and verifying that the minimum pedestrian duration threshold is met for all pedestrians. The consistency check includes: when the arbitration type is equal distribution, verifying that the duration difference at each intersection does not exceed the equilibrium tolerance. The equilibrium tolerance is a system parameter used to limit the maximum allowable difference in the equal distribution scenario. The equilibrium tolerance is set by traffic managers based on road grade and regional fairness requirements. After the checks pass, the final signal duration allocation table is generated by combining the variable constraint duration table and the check results, and the final signal duration allocation table is output.
[0134] An execution notification is generated and output by combining the final signal duration allocation table with the arbitration type and phase transition time. Specifically, the execution notification includes at least: the effective period number, the planned effective time, the total duration of one period, the arbitration type, the green light duration at each intersection, the phase transition time at each intersection, a parameter summary, and a checksum. The effective period number is the period sequence number of the field controller, the planned effective time is calculated by the central system based on the current clock and period boundaries, and the checksum is generated using a 16-bit cyclic redundancy checksum and is used by the field controller to verify the integrity of the notification.
[0135] Step 5: Dynamically update the credit value of each intersection based on the arbitration result.
[0136] The arbitration result obtained in step four is combined with the credit value of each intersection generated in step two, and an updated credit value is calculated through a credit value update mechanism. The arbitration result refers to the three types of strategies (full release strategy, proportional allocation strategy, and equal allocation strategy) and the corresponding intersection green light duration allocation table generated in step four through a three-level arbitration strategy, reflecting the priority of intersections in signal duration allocation. The credit value of each intersection is a scalar calculated in step two based on state values and historical compromise rates, used to quantify the priority of intersection collaborative control.
[0137] Specifically, the credit score update mechanism is as follows: First, the direction of credit score adjustment is determined based on the arbitration result. If the intersection receives a longer green light duration than the previous cycle in the arbitration result, it is judged as a positive adjustment; if it receives a shorter green light duration, it is judged as a negative adjustment. Then, the credit score adjustment range is calculated. The adjustment range = original credit score × adjustment coefficient. The adjustment coefficient is obtained by calculating the ratio of the difference between the green light duration of the current cycle and the previous cycle to the green light duration of the previous cycle, and then weighting it by combining the benchmark coefficients corresponding to different strategies.
[0138] The baseline coefficient is a fundamental parameter reflecting the degree of impact of different traffic management strategies on credit score adjustments. Its setting method needs to be combined with the actual traffic management objectives.
[0139] Efficiency-first strategy: For road sections with high traffic volume, the baseline coefficient can be set to a higher value (such as 0.8~1.2) to make changes in green light duration more sensitive to credit value adjustments, thereby incentivizing intersections to quickly optimize traffic efficiency.
[0140] Balanced traffic management strategy: Applicable to areas with uneven road network traffic distribution, with a baseline coefficient set at a moderate value (e.g., 0.5~0.7), ensuring the passage of main roads while taking into account the needs of branch roads;
[0141] Emergency response strategy: In the event of a sudden traffic incident, the baseline coefficient of the relevant intersection can be temporarily increased to 1.5 or above, and the credit value weight of the emergency lane can be forcibly increased.
[0142] The result is then calculated according to the formula "Updated credit score = Original credit score ± Adjustment range". Positive adjustments are calculated by addition, and negative adjustments are calculated by subtraction.
[0143] Step 6: Calculate the standard deviation of the credit values of all intersections in the network and input it into the credit entropy stabilization mechanism to obtain the stability optimization strategy; count the number of times an intersection yields to other intersections in the past period, calculate the historical compromise rate and output it to Step 2.
[0144] Specifically, step six includes the following steps:
[0145] Step 61: Calculate the standard deviation of the credit scores of all intersections in the network.
[0146] The new credit score table is used to compile the intersection credit score set. The average credit score and the standard deviation of the network credit score are calculated in sequence, and the weighted standard deviation is calculated as needed. Finally, the standard deviation index is generated and output.
[0147] Specifically, the credit scores of each intersection participating in the decision-making process are extracted from the new credit score table for the current period, resulting in a set of intersection credit scores. The average credit score is then calculated using this set of intersection credit scores and the number of intersections, using the following formula:
[0148] Average credit score = (sum of credit scores at all intersections) ÷ number of intersections.
[0149] The number of intersections refers to the number of intersections that participate in the decision-making process and generate credit value updates in this cycle; the credit value of each intersection is the value in the interval [0, 100] after memory smoothing and boundary mapping in step six.
[0150] Specifically, the standard deviation of the entire network's credit score is calculated from the average credit score and the intersection credit score set. This standard deviation measures the relative dispersion of the credit scores. The formula is as follows:
[0151] Standard deviation = ;
[0152] Among them, (credit score of each intersection - average credit score)² is used to measure the degree of deviation of a single intersection from the network average; the larger the standard deviation value, the weaker the network coordination and stability.
[0153] Optionally, to highlight the impact of key intersections on overall volatility, a weighted standard deviation can also be calculated. Specifically, a weighting factor is first set for each intersection, and then the dispersion is measured in a weighted manner, using the following formula:
[0154] Weighted standard deviation = ;
[0155] The weighting factor refers to the non-negative weight reflecting the importance of the intersection, which can be composed of traffic flow weight × credit fluctuation weight. The traffic flow weight can be obtained by normalizing "the number of vehicles passing through the intersection in this period ÷ the total number of vehicles passing through the network". The credit fluctuation weight can be given as the normalized value of the fluctuation range of the intersection's credit value in the entire network over the past few periods (such as 3 periods). If there is no historical window or it is not necessary to amplify the influence of key intersections, the weighting factor can be set to 1, so that the weighted standard deviation degenerates into the unweighted standard deviation.
[0156] Step 62: Input the standard deviation of the credit value into the credit entropy stabilization mechanism, generate an optimization strategy based on the credit entropy value, and trigger system stability adjustment.
[0157] Using the standard deviation of credit scores as input, the credit entropy value is calculated through a credit entropy stabilization mechanism. This mechanism quantifies the degree of disorder in the distribution of credit scores across the entire network intersections, achieving stability analysis by combining the standard deviation of credit scores with network topology characteristics. The standard deviation of credit scores is an index of the dispersion of credit scores across the entire network intersections calculated in step 61.
[0158] Specifically, the credit entropy stabilization mechanism is as follows: First, a baseline entropy value is set, which is determined based on the statistical mean of the standard deviation of credit values during the stable traffic periods across the entire network over the past three months. The stable traffic periods refer to three time periods each day: 6:00-7:00 AM, 12:00-1:00 PM, and 10:00-11:00 PM. Dates affected by holidays and severe weather need to be excluded during the statistical analysis. Then, the deviation between the standard deviation of the credit value and the baseline entropy value is calculated: Deviation = |Standard deviation of credit value - Baseline entropy value| / Baseline entropy value. Subsequently, the deviation is substituted into the entropy value conversion function to obtain the credit entropy value, which is: Credit entropy value = Deviation × Entropy value scaling factor + Baseline entropy value.
[0159] The entropy scaling factor is used to adjust the weight of the deviation on the credit entropy value. The value is dynamically calibrated based on the number of nodes in the traffic network. The basic coefficient value is set according to the number of nodes in different intervals, and the final value is determined by linear interpolation based on the proportion of the number of nodes in the interval.
[0160] A stability optimization strategy is derived by combining credit entropy values with a stability threshold range and using threshold determination rules. The stability threshold range includes a lower entropy threshold and an upper entropy threshold. The lower threshold represents the critical value where the credit value distribution is too concentrated, and the upper threshold represents the critical value where the credit value distribution is too dispersed. Both are determined by back-calculating from traffic failure cases across the entire network over the past year. Specifically, cases where signal control failures were caused by abnormal credit value distribution are selected, the credit entropy value at the time of the case is extracted, and the 10th percentile of the case's entropy value is used as the lower threshold, and the 90th percentile as the upper threshold.
[0161] Specifically, the threshold determination rule is as follows:
[0162] If the credit entropy value is less than or equal to the lower limit threshold, a credit equilibrium strategy is generated: For intersections where the overall credit value is lower than the average credit value across the entire network, the credit value is adjusted according to the formula: "Adjusted credit value = Original credit value + (Average credit value across the entire network - Original credit value) × Equilibrium adjustment coefficient". The equilibrium adjustment coefficient controls the extent to which the credit value converges towards the average, and its value is determined based on the difference between the lower limit threshold and the current credit entropy value.
[0163] If the credit entropy value is greater than or equal to the upper limit threshold, a credit convergence strategy is generated: For intersections where the overall credit value is 20% or more above the average credit value of the entire network, the credit value is adjusted according to the formula: "Adjusted credit value = Original credit value - (Original credit value - Average credit value of the entire network × 1.2) × Convergence adjustment coefficient". The convergence adjustment coefficient is used to control the rate of decrease in the credit value of high-credit-value intersections, and its value is determined based on the difference between the current credit entropy value and the upper limit threshold.
[0164] If the credit entropy value is within the range of the upper and lower thresholds of the entropy value, a credit maintenance strategy is generated: keep the original credit value of each intersection unchanged, and only record the current credit value distribution state for stability analysis in subsequent periods.
[0165] The stability optimization strategy is used to trigger the system to dynamically adjust the credit value update rules for each intersection, thereby achieving stable control of the credit value distribution across the entire network.
[0166] Step 63: Count the number of times the intersection yielded to other intersections over a period of time, calculate the historical compromise rate, and feed it back to Step 2.
[0167] First, extract the time-series records from the statistics window in the arbitration log to generate a yielding number table and a decision-making number table. Then, calculate the historical compromise rate table from the two tables. Finally, write the historical compromise rate table into the historical compromise rate cache and feed it back to step two for use.
[0168] Specifically, the data in the statistics window is filtered from the arbitration log to generate a yield count table. The statistics window is a scrolling interval of "the most recent several signal cycles" (set by system parameters). The arbitration log includes fields such as "intersection identifier, cycle number, green light duration before allocation, green light duration after allocation, yield flag, and participation flag." The records for the same intersection in the statistics window are summarized based on the "yield flag" to obtain the yield count for each intersection; similarly, the records for the same intersection in the statistics window are summarized based on the "participation flag" to obtain the participation decision count for each intersection.
[0169] Specifically, using the number of times yielding and the number of times participating in decision-making as inputs, the historical compromise rate of each intersection is calculated to obtain a historical compromise rate table; where, historical compromise rate = number of times yielding ÷ number of times participating in decision-making. When the number of times participating in decision-making is 0, the historical compromise rate is defined as 0 to avoid uncertain results caused by dividing by zero.
[0170] Specifically, to reduce the impact of short-term fluctuations on subsequent credit calculations, the historical compromise rate table and the previous period's historical compromise rate cache are updated in place using exponential smoothing to obtain a new historical compromise rate cache. Exponential smoothing uses a single smoothing coefficient, which is set by system parameters. The larger the value, the more sensitive it is to the response to the current period's statistical results.
[0171] Finally, the updated historical compromise rate cache is fed back to step two as the historical compromise rate, and used to participate in the subsequent credit value calculation together with the status value received in step two.
[0172] Example 2:
[0173] Please see Figure 2 As shown, a real-time signal global sensing control system includes the following modules:
[0174] Traffic status data acquisition and evaluation module: Real-time traffic status data is acquired through radar detectors, and the traffic status data is used to generate status values using a fuzzy comprehensive evaluation method;
[0175] Intersection Credit Score Preliminary Calculation Module: Using state values and historical compromise rates, the credit score of each intersection is calculated based on a multi-stage reinforcement game model.
[0176] Dynamic adjustment module for cooperation range: Combining the credit value and physical distance of each intersection, the cooperation range is dynamically adjusted using a cooperation inhibition distance field mechanism;
[0177] Decision intersection determination and signal duration allocation module: determines the intersections participating in the decision based on the cooperation scope; uses the credit value of the participating intersections to obtain the arbitration result through a three-level arbitration strategy, and allocates signal duration based on the arbitration result;
[0178] Intersection Credit Value Dynamic Update Module: Dynamically updates the credit value of each intersection based on the arbitration results;
[0179] Credit Entropy Stabilization and Historical Compromise Rate Statistics Module: Calculates the standard deviation of the credit value of all intersections in the network and inputs it into the credit entropy stabilization mechanism to obtain the stability optimization strategy; counts the number of times each intersection yields to other intersections in the past period, calculates the historical compromise rate and outputs it to the intersection credit value initial calculation module.
Claims
1. A real-time control method for global signal sensing, characterized in that, include: Step 1: Collect traffic status data in real time using radar detectors, and generate status values from the traffic status data using a fuzzy comprehensive evaluation method; Step 2: Using state values and historical compromise rates, calculate the credit value for each intersection based on a multi-stage reinforcement game model; Step 3: Combining the credit value and physical distance of each intersection, a cooperative inhibition distance field mechanism is used to dynamically adjust the cooperation range; Step 4: Determine the intersections involved in the decision-making process based on the scope of cooperation; The credit scores of the intersections involved in the decision-making process are used to obtain arbitration results through a three-level arbitration strategy, and signal duration is allocated based on the arbitration results. Step 5: Dynamically update the credit score of each intersection based on the arbitration result; Step 6: Calculate the standard deviation of the credit values at all intersections in the network and input it into the credit entropy stabilization mechanism to obtain the stability optimization strategy; Count the number of times each intersection yields to other intersections over a period of time, calculate the historical compromise rate, and output it to step two; The process of generating state values using the fuzzy comprehensive evaluation method is as follows: Based on the traffic state data collected by the radar detector, the state values are obtained through the fuzzy comprehensive evaluation method after standardization. The traffic state data includes traffic flow, queue length, phase saturation, and intersection level. After time alignment and anomaly removal, the data is normalized to form a standardized dataset. Then, a four-level equal-width triangular membership degree is constructed based on the standardized dataset, and the membership degree matrix is obtained by combining the triangular membership degrees. The comprehensive evaluation vector is calculated based on the membership matrix and the equal weight vector. Finally, the state value is generated by defuzzifying the comprehensive evaluation vector and the grade score vector by performing the inner product.
2. The real-time control method for global signal sensing according to claim 1, characterized in that, The multi-stage reinforcement game model is as follows: The cooperative utility is calculated by weighting the state value and the historical compromise rate; the basic benefit is calculated based on the state value archived in the previous control cycle and the state value in the current cycle, and the potential function is constructed by combining the state value and the historical compromise rate; when switching cycles, a reward shaping term is formed based on the potential function of the next cycle and the current cycle, and the reward shaping term is added to the basic benefit to obtain the immediate benefit. Establish and update the strategy value table based on immediate returns, and select the largest item in the current period from the strategy value table as the long-term strategy value. Based on long-term strategic value, a greedy rule with an exploration rate is used to determine behavioral choices; finally, a weighted synthesis is performed based on cooperative utility and long-term strategic value to obtain the credit value of each intersection.
3. The real-time control method for global signal sensing according to claim 2, characterized in that, The method of dynamically adjusting the cooperative range using a cooperative suppression distance field mechanism includes: Step 31: Combine the credit value and physical distance of each intersection to calculate the cooperation inhibition distance between intersections; Step 32: Set the cooperation threshold and dynamically adjust the effective cooperation range of different intersections based on the cooperation inhibition distance and the cooperation threshold.
4. The real-time control method for global signal sensing according to claim 3, characterized in that, The arbitration outcome obtained through the three-tier arbitration strategy includes: Step 41: Based on the cooperative inhibition mechanism, select the intersections participating in the decision-making process and define them as candidate intersections; Step 42: Calculate the green wave weight by combining the credit scores and phase urgency of the candidate intersections involved in the decision-making process; Step 43: Obtain the dominant difference ratio, maximum proportion, and relative dispersion based on the green wave weights, and combine them with the three-level arbitration strategy to obtain the arbitration result; Step 44: Based on the arbitration result, output the final signal duration allocation and notify each intersection to execute it.
5. The real-time control method for global signal sensing according to claim 4, characterized in that, The types of arbitration outcomes include: full release arbitration, proportional allocation arbitration, and equal division arbitration. The full release arbitration is as follows: when the dominant difference ratio is greater than or equal to the upper threshold of the dominant difference ratio and the maximum proportion is greater than or equal to the upper threshold of the maximum proportion, it is determined to be significantly leading and a full release strategy is adopted. The proportional allocation arbitration is as follows: when the full release condition is not met, if the relative dispersion is greater than or equal to the relative dispersion threshold, or the dominant difference ratio is greater than or equal to the upper threshold of the dominant difference ratio and less than the lower threshold of the dominant difference ratio, it is determined that the difference is large and a proportional allocation strategy is adopted. The equal-division arbitration is as follows: when the dominant difference ratio is less than the lower threshold of the dominant difference ratio and the relative dispersion is less than the relative dispersion threshold, it is determined that the difference is very small, and the equal-division strategy is adopted.
6. The real-time control method for global signal sensing according to claim 5, characterized in that, The full release strategy, proportional allocation strategy, and equal distribution strategy are as follows: Full-allowance strategy: Set the green light duration of the intersection with the highest weight to the product of the full-allowance ratio and the total signal cycle, and set the green light duration of the other candidate intersections to 0. Proportional allocation strategy: The green light duration of each candidate intersection is equal to the ratio of the green wave weight of that candidate intersection to the sum of the weights, multiplied by the total signal cycle. Equal distribution strategy: Allocate equally according to the number of candidate intersections, and the green light duration of each candidate intersection is equal to the ratio of the total signal cycle to the number of candidate intersections.
7. The real-time control method for global signal sensing according to claim 6, characterized in that, The obtained stability optimization strategy includes the following steps: Step 61: Calculate the standard deviation of the credit scores for all intersections across the network; Step 62: Input the standard deviation of the credit score into the credit entropy stabilization mechanism, generate an optimization strategy based on the credit entropy value, and trigger system stability adjustment; Step 63: Count the number of times the intersection yielded to other intersections over a period of time, calculate the historical compromise rate, and feed it back to Step 2.
8. The real-time control method for global signal sensing according to claim 7, characterized in that, The credit entropy stabilization mechanism is as follows: Based on the standard deviation of the credit score, and combined with the statistical mean of the standard deviation of the credit score during stable traffic periods across the entire network, a benchmark entropy value is set; the deviation between the standard deviation of the credit score and the benchmark entropy value is calculated, and the deviation is substituted into the entropy value transformation function to obtain the credit entropy value; the stability threshold range is determined by back-calculation based on traffic failure cases across the entire network, and a stability optimization strategy is generated by combining the credit entropy value with threshold determination rules; If the credit entropy value is less than or equal to the lower limit threshold, a credit equilibrium strategy is generated; if the credit entropy value is greater than or equal to the upper limit threshold, a credit convergence strategy is generated; if the credit entropy value is within the threshold range, a credit maintenance strategy is generated, thereby triggering the system to dynamically adjust the credit value update rules for each intersection.
9. A real-time signal global sensing control system, used to implement the real-time signal global sensing control method according to any one of claims 1-8, characterized in that, include: Traffic status data acquisition and evaluation module: Real-time traffic status data is acquired through radar detectors, and the traffic status data is used to generate status values using a fuzzy comprehensive evaluation method; Intersection Credit Score Preliminary Calculation Module: Using state values and historical compromise rates, the credit score of each intersection is calculated based on a multi-stage reinforcement game model. Dynamic adjustment module for cooperation range: Combining the credit value and physical distance of each intersection, the cooperation range is dynamically adjusted using a cooperation inhibition distance field mechanism; Decision intersection determination and signal duration allocation module: determines the intersections participating in the decision-making process based on the aforementioned cooperation range; The credit scores of the intersections involved in the decision-making process are used to obtain arbitration results through a three-level arbitration strategy, and signal duration is allocated based on the arbitration results. Intersection Credit Value Dynamic Update Module: Dynamically updates the credit value of each intersection based on the arbitration results; Credit entropy stabilization and historical compromise rate statistics module: Calculates the standard deviation of credit values at all network intersections and inputs it into the credit entropy stabilization mechanism to obtain a stability optimization strategy; The system counts the number of times each intersection yields to other intersections over a period of time, calculates the historical compromise rate, and outputs it to the intersection credit value initial calculation module. Based on the cooperation scope, the system determines which intersections will participate in the decision-making process.
Citation Information
Patent Citations
Coordinated control oriented trunk line crossing correlation analysis and division method
CN105825690A
Road traffic signal control system and method
CN119649621A