Urban traffic jam analysis and prediction method based on gradient elevator

By using a gradient booster machine model to analyze and predict urban traffic congestion, this approach addresses the issues of insufficient data integration and prediction in traditional methods, thereby achieving efficient traffic management decision support.

CN121661832APending Publication Date: 2026-03-13LIAONING UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Traditional urban traffic congestion management methods lack dynamic optimization capabilities, have insufficient data integration, and are difficult to predict complex congestion problems. Existing prediction systems lack machine learning support, resulting in low traffic management efficiency.

Method used

The Gradient Boosting Machine (GBM) model is used to obtain historical traffic data from appropriate data sources, perform data preprocessing and outlier detection, construct feature vectors, and use the GBM model to predict traffic congestion and generate congestion warning information to support traffic management decisions.

Benefits of technology

It enables accurate prediction of urban traffic congestion, reduces the difficulty and cost of data labeling, improves prediction accuracy, and supports the dynamic adjustment and governance strategy formulation of traffic management departments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661832A_ABST
    Figure CN121661832A_ABST
Patent Text Reader

Abstract

An urban traffic congestion analysis and prediction method based on a gradient elevator comprises the following steps: 1) determining a target area for congestion prediction, and determining key input variables in combination with urban multi-dimensional elements; 2) selecting a plurality of areas with high urban congestion rate, capturing original traffic and meteorological data in a set time window, normalizing the original data, and storing the original data as analysis data of model training; 3) detecting data missing values and abnormal values, and reserving reasonable extreme values; 4) dividing a data set according to a time sequence, performing iterative training on the model by adopting a gradient elevator, and introducing a family algorithm to perform transverse verification; and 5) using the generalization ability of the index and time sequence cross validation evaluation model, and combining the thermodynamic diagram and the time period curve to carry out visualization and effect test on the prediction result to form traffic jam peak early warning and dispersion suggestions. The invention provides a traffic jam prediction preprocessing scheme, the prediction threshold is reduced, and the prediction accuracy and interpretability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of urban traffic management, and specifically designs an urban traffic congestion analysis and prediction method based on a gradient booster, which is applicable to traffic flow prediction and traffic data analysis of intelligent transportation systems (ITS). Background Technology

[0002] Traffic congestion prediction utilizes historical traffic data and machine learning algorithms to forecast traffic flow, speed, and congestion levels on road networks over a future period. Against the backdrop of increasing vehicle ownership and accelerated urban modernization, traditional urban congestion management methods reveal significant limitations. Generally, traffic light cycles at intersections are fixed and their timing rigid. For urban arterial roads where peak-hour traffic volumes significantly exceed off-peak hours, these cycles cannot keep pace with the drastic fluctuations in traffic flow. Currently, manual traffic management relies heavily on on-site traffic control by traffic police, resulting in long average processing times for emergencies. Furthermore, traditional traffic management systems suffer from insufficient data integration and fragmented multi-source data. For example, traffic flow, weather information, and traffic incident data are scattered across different management departments, making it difficult to predict complex congestion problems. Moreover, existing prediction systems are mostly rule-based and lack the dynamic optimization capabilities of machine learning. Summary of the Invention

[0003] To overcome the shortcomings of existing technologies, this invention provides a method for urban traffic congestion analysis and prediction based on gradient booster machines. By acquiring historical traffic data from appropriate data sources and performing good data preprocessing, the GBM model is used to predict future traffic congestion, providing forward-looking decision support for traffic management departments and effectively alleviating urban traffic problems.

[0004] The technical solution of this invention is: a method for urban traffic congestion analysis and prediction based on a gradient booster, comprising the following steps: Step 1) Select the traffic situation interface, data dimension and sampling period according to the traffic congestion characteristics of the target city through the map open platform to complete the data collection and form complete traffic data.

[0005] In step 1), multiple high-congestion areas are selected, including at least one of the following types: roads around hospitals, school areas, public transportation hubs, urban main roads, and freight corridors in industrial areas, as the target city's traffic congestion prediction areas. Data is collected for each area with an appropriate regional situation within a preset radius.

[0006] In step 1), the traffic congestion features include: constructing a morning and evening peak period labeling field based on hourly features, labeling the time periods of 7:00-9:00 and 17:00-19:00 as peak periods, and labeling the remaining time periods as off-peak periods.

[0007] Step 2) Perform missing value detection, outlier detection, and field format conversion on the traffic data to obtain a structured traffic sample dataset.

[0008] In step 2), outlier detection employs an outlier detection method based on interquartile range: The samples were grouped according to road signs and practice windows. The first quartile (Q1), third quartile (Q3), and interquartile range (IQR) were calculated as follows: IQR = Q3 - Q1. Congestion rates below Q1 - 1.5 were considered. IQR may be higher than Q3 + 1.5 IQR samples are marked as abnormal samples.

[0009] Step 3) Select feature vectors from the structured sample dataset, including time, coordinates, city, temperature, humidity, wind force level, traffic condition description, smooth flow rate, congestion rate, severe congestion rate, road condition level, and road condition description, and divide the sample data into training set and test set according to time order.

[0010] Step 4) Using the congestion rate in the training set as the regression target, initialize the model prediction value as the mean of the target variable. In each iteration: calculate the residual between the current model prediction value and the actual value; train a regression decision tree as the weak learner for the current iteration using the residual as the learning target; accumulate the output of the current weak learner to the model prediction value according to the preset learning rate, and update the model; iterate until the preset number of trees is reached or the loss function converges, and obtain the trained gradient booster traffic congestion prediction model.

[0011] In step 4), the gradient booster traffic congestion prediction model uses the mean squared error as the loss function and the negative gradient of the residual as the fitting target of the weak learner in each round.

[0012] Step 5) Test the gradient booster traffic congestion prediction model using a test set; The specific method in step 5) is as follows: input the feature vector corresponding to the test set into the gradient booster traffic congestion prediction model, output the congestion situation of the target road segment or target time, and generate a congestion spatiotemporal distribution map or congestion warning information based on the prediction results; if the gradient booster traffic congestion prediction model shows the expected accuracy on the test set and the loss function reaches the set convergence criterion, it means that the model training is complete, and the final gradient booster traffic congestion prediction model is output; if the accuracy does not reach the expected level or the loss function does not converge, return to step 4) to continue training until the accuracy meets the requirements or the loss function converges.

[0013] When evaluating the performance of the gradient booster traffic congestion prediction model on the test set, the root mean square error (RMSE) and the coefficient of determination (R²) are calculated, and the evaluation results are compared with prediction models based on XGBoost and LightGBM.

[0014] Step 6) Input the target city, latitude and longitude, and time to be detected into the final gradient booster traffic congestion prediction model obtained in Step 5) for detection, and obtain congestion warning information or congestion spatiotemporal distribution map generated based on the congestion prediction results.

[0015] The congestion warning information generated based on the congestion prediction results in step 6) is used to: guide traffic management departments to dynamically adjust signal timing, optimize the release phase and duration at key intersections, and formulate management strategies such as diversion and detours, increasing public transportation services to alleviate congestion, and adding BRT urban rapid public transportation lines during peak hours.

[0016] The beneficial effects of this invention are as follows: The advantage of this invention lies in its accurate prediction of urban traffic congestion based on the Gradient Boosting Machine (GBM) model. Furthermore, it defines key predictive variables and event elements for different types of traffic congestion problems (such as peak hours and specific road sections), ensuring these elements are more aligned with the actual operation of the traffic system rather than lumping all types of congestion together. This invention significantly reduces the annotation difficulty in subsequent model training, lowering data annotation costs and minimizing its impact on prediction accuracy. Through a series of steps, and utilizing the obtained high-quality traffic congestion test dataset, an effective congestion prediction and governance decision support system is formed. Attached Figure Description

[0017] Figure 1 This is a framework diagram for a method of analyzing and predicting urban traffic congestion based on gradient boosters.

[0018] Figure 2 This is a map showing the traffic congestion situation in Shenyang on January 1, 2025, based on the collected data and processed by the data processing flow of this invention.

[0019] Figure 3 This is a map showing the traffic congestion situation in Shenyang during the 2025 Spring Festival, based on the collected data and processed by the data processing flow of this invention.

[0020] Figure 4 This is a map showing the traffic congestion situation in Shenyang on a typical weekday in 2025, based on the collected data and processed by the data processing flow of this invention.

[0021] Figure 5This is a large-scale overall map of traffic congestion in Shenyang City, drawn based on the collected data and processed by the data processing flow of this invention.

[0022] Figure 6 Traffic congestion heat maps for 20 specific areas in Shenyang City are generated based on the collected data and processed by the data processing flow of this invention.

[0023] Figure 7 For based on Figure 4 A comparison chart showing the actual congestion situation and the model's predicted congestion situation.

[0024] Figure 8 This is a time comparison chart showing how the training time of Gradient Boosting Machine (GBM), Lightweight Gradient Boosting Machine (LightGBM), and Extreme Gradient Boosting Machine (XGBoost) changes with the number of training epochs on the same dataset.

[0025] Figure 9 Overall architecture diagram of a method for analyzing and predicting urban traffic congestion based on gradient boosters.

[0026] Specific implementation method

[0027] This invention provides a method for analyzing and predicting urban traffic congestion based on Gradient Boosting Machine (GBM), the framework of which is as follows: Figure 1 As shown.

[0028] Example 1: The present invention will be further described in detail below with reference to the accompanying drawings and relevant specific operations and examples. The example city is Shenyang, Liaoning Province.

[0029] (1) Select the target prediction area and choose a data sampling scheme according to the characteristics of the city.

[0030] For the urban traffic congestion analysis and prediction model, the main data sources are from Shenyang's traffic management department, traffic monitoring cameras, real-time traffic detection systems, and Shenyang's public transportation operation data. As a large city in China, Shenyang's traffic congestion problems are typically concentrated in certain specific areas, such as Qingnian Avenue, Chongshan Road, Jianshe Road, Baogong Street, Nanjing Street, Fangcheng Ring Road, and other city center, commercial districts, and major transportation hubs. This invention employs a multiple sampling method, with 20 data points and a circular area radius of 10 meters.

[0031] (2) Data acquisition is completed through the open map platform API.

[0032] Considering the economic feasibility requirements of practical deployment, and comprehensively comparing the advantages and disadvantages of several commonly used open geographic traffic data platforms in domestic research, this invention selects the Amap (Gaode Maps) open platform as the original traffic data acquisition interface. Based on the urban traffic characteristics of Shenyang, among the various API interfaces provided by Amap, the circular area traffic situation query, which better meets practical needs, was chosen. The circular area focuses on a specific region, optimizing computing resources while maintaining high data accuracy and preserving regional relevance to the greatest extent. It can quickly identify congestion hotspots based on real-time traffic data. The specific implementation utilizes Python combined with HTTP request methods to send requests to the Amap API, obtaining real-time and historical traffic data for the target road by specifying the city code, road name, and time parameters. The city code for Shenyang is "210100," and the API is periodically called to obtain all-day traffic data.

[0033] (3) Storage and structuring of raw traffic data.

[0034] While acquiring data, all data is converted into a unified structured format, and timestamps and road names are standardized to ensure data consistency and availability.

[0035] This invention employs a time-series partitioning method, dividing the dataset into a training set (80%) and a test set (20%) according to time sequence. The resulting training set is used to train a supervised model, while the test set is used to evaluate the model's final performance and verify whether the model is overfitting or underfitting.

[0036] (4) Detection and processing of missing values.

[0037] The data collection method used in this invention utilizes the Gaode Map Open Platform, a recognized professional platform for acquiring various types of geographic data and an HTTP interface approved by Gaode Maps, a leading developer in China. The selected area for this example is the main roads of Shenyang City, and the collected data is expected to have no missing values. However, for the integrity and adaptability of the overall workflow, missing value detection and processing are still performed on the dataset. Using Python, the dataset processed in the above steps is imported, and the `df.isnull().sum()` method is called to quickly calculate the total number of missing values ​​in the entire DataFrame and print the details of the number of missing values ​​in each column.

[0038] (5) Detection and handling of outliers.

[0039] In this invention, outlier handling must consider data characteristics, business scenarios, and specific model requirements. Traffic data typically exhibits a skewed multimodal distribution or contains numerous extreme values ​​and periodic data; therefore, the core detection method used is the interquartile range (IQR). IQR makes no assumptions about the data distribution, is more robust to occasional extreme values, and is directly correlated with business indicators. For the prediction model GBM selected in this invention, GBM has the advantage of autonomously learning sparse features, thus it can handle locally occurring outliers.

[0040] (6) Data standardization and feature engineering.

[0041] Before formally establishing the prediction model, the time field in the dataset needs to be parsed. This invention converts the time field into a year:month:day:hour:minute time format so that the model can better capture the temporal patterns of the data. To make certain percentage fields more understandable to the model, this invention processes them as follows: remove the percentage sign and convert them into floating values; for example, if the original field is 80%, the value is converted to 0.80 after processing. For the city field, this invention uses LabelEncoder to encode the city name into a numeric category. The coordinate field is split into longitude and latitude columns, which are added to the model as independent features. For feature engineering, this invention uses correlation analysis and recursive feature elimination methods to screen out key variables affecting congestion from dimensions such as historical congestion index, current vehicle speed, time, and weather conditions. Feature extraction and construction are based on the original data to further construct efficient features, such as the average vehicle speed change rate, short-term traffic flow fluctuation index, and congestion patterns during morning and evening peak hours calculated based on historical data. These derived features can enhance the model's ability to capture changes in traffic data. Since the numerical ranges of different features may vary greatly, this invention will further scale the dataset. For example, for vehicle speed (km / h) and congestion index (0-10), this invention uses the Z-score normalization method to scale all features to the same scale to improve the convergence speed and stability of the model.

[0042] (7) Visual analysis of urban temporal patterns.

[0043] From a subjective perspective of urban traffic congestion, the time-domain analysis of traffic congestion patterns should focus on holidays and morning / evening rush hours. Therefore, this invention selects 20 data collection points in Shenyang City on January 1, 2025, as an example. Without considering location factors, the Matplotlib library is used to plot the traffic congestion rate for that day. The horizontal axis represents 24-hour time, in units of one hour; the vertical axis represents the congestion rate (%). The traffic congestion situation in Shenyang City on January 1, 2025, is plotted as follows: Figure 2 As shown.

[0044] Generally, traffic congestion on holidays and regular workdays fluctuates on an hourly basis. For example, congestion on weekday mornings typically begins around 7:00 AM, but this time usually lags behind on regular holidays. By comparing traffic congestion trends on specific dates with those on regular dates, it's possible to investigate whether there are significant temporal differences in urban traffic congestion. The following chart illustrates traffic congestion in Shenyang during the 2025 Spring Festival and on a regular workday. Figure 3-4 As shown.

[0045] The images above allow us to explore the temporal characteristics of traffic congestion in Shenyang City based on original traffic congestion data. This invention defines a congestion rate greater than 5% as slow traffic and greater than 12% as severe congestion. Observing the images, we can intuitively draw the following conclusions: Regardless of the date, traffic congestion in Shenyang City generally begins around 7:30 AM each day. Congestion gradually increases on each studied road segment, typically exceeding the severe congestion threshold around 8:00 AM. From this point onward, the congestion rate continues to rise, approaching 16%, and the road conditions rapidly change from "slow traffic" to "severe congestion." Around 9:00 AM, the congestion situation drops sharply, returning to normal levels. Generally, we refer to this period of congestion as the "morning rush hour." Similarly, an evening rush hour occurs from approximately 5:30 PM to 7:00 PM each day, with congestion indices and patterns similar to the morning rush hour.

[0046] In other words, subsequent prediction models based on the GBM model do not require significant special adjustments for specific dates. The fluctuations in Shenyang's traffic congestion caused by special dates do not significantly contradict the traffic congestion patterns on ordinary weekdays. Therefore, the subsequent prediction models do not need to be adjusted from the perspective of holidays within an acceptable range of accuracy.

[0047] (8) Visual analysis of urban regional patterns.

[0048] Furthermore, a traffic congestion heatmap of Shenyang can be created using existing data. The resulting overall large-scale traffic congestion heatmap of Shenyang based on the dataset is shown below. Figure 5 As shown.

[0049] Shenyang, as an important city in Northeast China, is divided into major areas such as Heping District, Shenhe District, Huanggu District, Tiexi District, and Dadong District. The research area is focused on the first ring road area and main roads such as Qingnian Street, Beiling Street, Chongshan Road, and Jianshe Road, which effectively cover the city's administrative area. From the map, we can draw the following conclusions about the regional patterns of traffic congestion in Shenyang: congestion is mainly concentrated on the city's first ring road and the urban center radiation zone centered on the intersection of Qingnian Street and Zhongshan Road; the congestion situation is radial, spreading from the specific intersection to the relevant main roads, resulting in slow traffic. However, except for the morning peak discussed in (6), the traffic congestion situation in Shenyang is positive and improving; more specifically, adjusting the heat map ratio, for example Figure 6 As shown.

[0050] The map provides a more specific and clearer picture of the data collection areas. Congestion centers can be categorized as follows: ring roads centered around hospitals, school areas, key public transportation hubs, main urban roads such as Qingnian Avenue and Jianshe Road, and freight routes in tourist attractions and industrial areas. Congestion radiates outwards from these areas, but generally does not extend beyond the traffic intersections within those areas. For example, small-scale, short-term congestion may occur in transportation hub areas during peak train arrival times, while tidal congestion may occur on main roads, especially along the Qingnian Avenue Golden Corridor.

[0051] In summary, the 20 regional center points selected in this study have significant practical value for studying traffic congestion in Shenyang. When the model makes actual predictions, real-time traffic flow within a 30-meter radius of the key monitoring area can be collected as model data. At the same time, roads with recent construction can also be given special attention.

[0052] (9) Construction and training of GBM model.

[0053] Urban traffic congestion prediction is a typical regression problem. In the initialization of the GBM prediction model regression problem, since the mean is the best constant prediction value that minimizes the squared error, the mean of the target variable—the congestion rate—can be chosen as the initial prediction value; that is, the model predicts the mean of the target value for all samples. Then, the relevant hyperparameters of the first tree (parameters determined before model training) can be defined: `n_estimators` is a parameter describing the number of trees, i.e., weak learners, which can be adjusted according to the actual prediction performance of the model; `max_path` is the maximum depth of the decision tree, and its value determines the complexity of the weak learners. Here, 3 is chosen as the default value for the maximum tree depth. After determining the baseline of the prediction model, the error between the current model's prediction value and the true target value, i.e., the residual value, can be calculated. The residual value can be used to measure the prediction performance of the established prediction model. After completing the above two steps, the training of the first tree should naturally begin. The GBM model trains itself by progressively correcting the prediction error, and the first tree takes on the responsibility of predicting the current residual through feature splitting, thereby optimizing the model. This process is similar to the training process of traditional decision trees, i.e., GBDT. The weak learner of the GBM model is itself a theoretical definition of a decision tree, but the goal of the tree is not to fit the target variable, but to fit the residuals. For a typical regression problem such as urban traffic congestion prediction, the training goal of the first tree is to minimize the loss function; then, the optimal split point is found to divide the dataset. In this invention, the splitting rule of each tree is based on the features of the training set to minimize the mean squared error (MSE).

[0054] After training the first tree, the GBM prediction model needs to adjust the predicted values ​​using the learning rate. The learning rate is a regularization hyperparameter of the GBM model. It slows down the model's learning speed by reducing the influence of each tree on the overall prediction and encourages the model to generalize more. It controls the weight of each tree, i.e., the contribution of each learner to the model: a smaller learning rate will make smaller adjustments to the weights of each tree, thereby improving the model's robustness, but higher robustness will lead to an increase in the number of iterations; conversely, a larger learning rate may lead to overfitting. After comprehensive consideration, this invention sets the learning rate to 0.1. The training of the GBM model is based on an "additive model," that is, multiplying the new predicted value by the learning rate (i.e., the step size) and adding it to the baseline predicted value to achieve the purpose of "gradient boosting" in the function space, thereby gradually improving the model's predictive ability.

[0055] The training process of the GBM model includes the following core steps: (1) Initialize the predicted values: The model sets the predicted values ​​of all data points to the mean of the target variable (congestion rate).

[0056]

[0057] in, This is the initial predicted value. It is the sample size. This is the actual value.

[0058] (2) Calculate the residual: For each iteration, the residual, that is, the difference between the actual value and the current predicted value, must be calculated.

[0059]

[0060] in, It is the first The residual in the round of iteration, It is the sample size. It is the predicted value from the previous iteration.

[0061] (3) Training a weak learner: Using the residuals as the target variable, train a new decision tree. This tree attempts to minimize the squared error of the current residuals.

[0062]

[0063] in, It is the first Residual prediction of wheels, It is the first Tree pairs of features The prediction.

[0064] (4) Update the predicted value: based on the learning rate Update the forecast values.

[0065]

[0066] in, It is the learning rate.

[0067] (4) Iterative training.

[0068] At this point, the model initialization, the establishment and training of the first tree are complete, and subsequent iterations will be carried out based on the above.

[0069] (10) Evaluation of urban congestion prediction results.

[0070] The comprehensive comparison of the indicators selected in this invention is shown in Table 1 below.

[0071] Table 1: Evaluation Indicators and Comparison of Congestion Prediction Models

[0072] In this evaluation, 80% of the data was used for training and 20% for testing. During the training process of the GBM-based traffic congestion prediction model of this invention, both the training and testing errors gradually decreased, indicating that the model was gradually approaching the best fit and possessed good predictive ability. Specifically, with the increase in the number of iterations, the root mean square error (RMSE) of both the training and testing sets decreased significantly. From 0.03537 for training and 0.03536 for testing in the first round, the RMSE decreased to 0.00024428 for training and 0.00024300 for testing in the 100th round, with a significant reduction in error magnitude, indicating that the model's fitting effect was continuously optimized. After 100 rounds of iterative optimization, the model's final performance indicators showed the expected accuracy and stability. First, RMSE is a standard indicator for measuring the prediction error of a regression model. The smaller the value, the better the prediction effect. The root mean square error (RMSE) of the model constructed above on the test set is 0.000243, which means that the prediction error of the model is almost negligible within the allowable range. In addition, R2 is 0.99996, which is close to the ideal state, indicating that the model can explain 99.996% of the data variability and shows excellent fitting effect on the dataset. This result verifies the good fit of the model to the data, indicating that the correlation between the prediction result and the actual value is very strong. Finally, the mean absolute error (MAE) of the model is 0.000148, which means that the average difference between each predicted value and the actual value is very small, further confirming the high accuracy of the model. Based on the image drawing of the traffic congestion situation of Shenyang City on a typical weekday (7), the model is used to predict the traffic congestion situation of a typical weekday in the future. The time unit is taken as five minutes, and a prediction line graph is drawn. Placing the two broken lines on the same coordinate system clearly shows that the GBM prediction model constructed above can predict future traffic congestion relatively well within a certain allowable error range. Figure 7 As shown.

[0073] (11) Comparative verification of predictive effectiveness.

[0074] In the field of ensemble learning research in machine learning, Gradient Boosting Machine (GBM) and its derivative algorithms XGBoost and LightGBM, which use gradient boosting decision trees as base learners, have become mainstream. This paper collectively refers to them as the GBM algorithm family. The evolution of these three models reflects the complete development path of this algorithm family from theoretical breakthroughs to engineering optimization. Utilizing the commonalities of these three models being in the same algorithm family, the accuracy of the prediction capability of the GBM prediction model established in this paper is verified horizontally on the same dataset. The prediction capabilities of the three models are compared under the premise of controlling relevant hyperparameters, as shown in Table 2.

[0075] Table 2: Comparison of Predictive Capabilities of the GBM Algorithm Family

[0076] For urban traffic congestion prediction, the most crucial metric for the model should be time. Given predictable congestion rates, achieving the fastest possible prediction delivery is key to the industrial deployment of the model. On existing datasets, this paper compares the performance of the traditional GBM model with other GBM algorithms in terms of training time, given the same number of iterations. Specific performance details are as follows: Figure 8 As shown.

[0077] (12) Suggestions for signal timing and diversion optimization.

[0078] Based on the visualization of raw data and model predictions, the regularity of traffic congestion in Shenyang City allows the model of this invention to be used to optimize traffic signal control. The model can be applied to dynamic signal control systems to automatically adjust the switching frequency and duration of traffic lights based on real-time traffic flow; or to optimize the departure frequency of bus routes on relevant roads based on congestion prediction results, and to set up specific BRT (Bus Rapid Transit) urban public transportation lines; furthermore, it can collaborate with ride-sharing platforms to distribute carpooling coupons to users during peak congestion periods to reduce the proportion of single-motorized vehicle trips.

[0079] (13) Congestion warning and information release.

[0080] The GBM model can be used to assist in real-time traffic decision-making. By integrating the GBM model into Shenyang's ITS (Internet System for Traffic Management), and connecting with various dynamic traffic data, short-term traffic congestion heat maps for specific areas can be generated. Real-time congestion predictions can be published through traffic management department guidance screens or navigation apps, guiding drivers on relevant road sections to rationally plan their routes and achieving dynamic traffic diversion on municipal roads from a macro-control perspective. Furthermore, by leveraging the GBM model's free feature importance ranking, the main factors causing traffic congestion can be quantified, and reasonable priority management strategies can be formulated based on the weight ratio of these factors.

Claims

1. A method for urban traffic congestion analysis and prediction based on gradient booster, characterized in that, Includes the following steps: Step 1) Collect traffic data by selecting the traffic situation interface, data dimension and sampling period according to the traffic congestion characteristics of the target city through the map open platform, and form complete traffic data; Step 2) Perform missing value detection, outlier detection, and field format conversion on the traffic data to obtain a structured traffic sample dataset; Step 3) Select feature vectors from the structured sample dataset, including time, coordinates, city, temperature, humidity, wind force level, traffic condition description, smooth flow rate, congestion rate, severe congestion rate, road condition level, and road condition description, and divide the sample data into training set and test set according to time order. Step 4) Use the congestion rate in the training set as the regression target, initialize the model prediction value as the mean of the target variable, and in each iteration: calculate the residual between the current model prediction value and the actual value; The regression decision tree is trained using the residual as the learning objective, serving as the weak learner for the current round. The output of the current weak learner is accumulated into the model's predicted value according to the preset learning rate, and the model is updated; the iteration continues until the preset number of trees is reached or the loss function converges, and the trained gradient booster traffic congestion prediction model is obtained. Step 5) Test the gradient booster traffic congestion prediction model using a test set; Step 6) Input the target city, latitude and longitude, and time to be detected into the final gradient booster traffic congestion prediction model obtained in Step 5) for detection, and obtain congestion warning information or congestion spatiotemporal distribution map generated based on the congestion prediction results.

2. The urban traffic congestion analysis and prediction method based on gradient booster as described in claim 1, characterized in that, In step 1), multiple high-congestion areas are selected, including at least one of the following types: roads around hospitals, school areas, public transportation hubs, urban main roads, and freight corridors in industrial areas, as the target city's traffic congestion prediction areas. Data is collected for each area with an appropriate regional situation and a preset radius.

3. The urban traffic congestion analysis and prediction method based on gradient booster as described in claim 1, characterized in that, In step 2), outlier detection employs an outlier detection method based on interquartile range: The samples were grouped according to road signs and practice windows. The first quartile (Q1), third quartile (Q3), and interquartile range (IQR) were calculated as follows: IQR = Q3 - Q1. Congestion rates below Q1 - 1.5 were considered. IQR may be higher than Q3 + 1.5 IQR samples are marked as abnormal samples.

4. The urban traffic congestion analysis and prediction method based on gradient booster as described in claim 1, characterized in that, In step 1), the traffic congestion features include: constructing a morning and evening peak marking field based on hourly features, marking the time periods of 7:00-9:00 and 17:00-19:00 as peak periods, and marking the remaining time periods as off-peak periods.

5. The urban traffic congestion analysis and prediction method based on gradient booster as described in claim 1, characterized in that, In step 4), the gradient booster traffic congestion prediction model uses the mean squared error as the loss function and the negative gradient of the residual as the fitting target of each round of weak learner.

6. The urban traffic congestion analysis and prediction method based on gradient booster as described in claim 1, characterized in that, The specific method in step 5) is as follows: input the feature vector corresponding to the test set into the gradient booster traffic congestion prediction model, output the congestion situation of the target road segment or target time, and generate a congestion spatiotemporal distribution map or congestion warning information based on the prediction results; if the gradient booster traffic congestion prediction model shows the expected accuracy on the test set and the loss function reaches the set convergence criterion, it means that the model training is complete, and the final gradient booster traffic congestion prediction model is output; if the accuracy does not reach the expected level or the loss function does not converge, return to step 4) to continue training until the accuracy meets the requirements or the loss function converges.

7. The urban traffic congestion analysis and prediction method based on gradient booster as described in claim 6, characterized in that, When evaluating the performance of the gradient booster traffic congestion prediction model using the test set, the root mean square error (RMSE) and the coefficient of determination (R²) are calculated, and the evaluation results are compared with prediction models based on XGBoost and LightGBM.

8. The method for urban traffic congestion analysis and prediction based on gradient booster as described in claim 1, characterized in that, The congestion warning information generated based on the congestion prediction results in step 6) is used to: guide traffic management departments to dynamically adjust signal timing, optimize the release phase and duration at key intersections, and formulate management strategies such as diversion and detours, increasing public transportation services to alleviate congestion, and adding BRT urban rapid public transportation lines during peak hours.