Method for predicting traffic conflict risk stability of expressway interleaving area
By using causal regularization logistic regression model and binning algorithm in the fast road interleaving area, the causal correlation between traffic conflict risk and traffic flow characteristics is established, and the problem of sample selection deviation and insufficient model robustness in the existing technology is solved, real-time stable prediction and safety control of traffic conflict risks in the interleaving area is achieved.
Patent Information
- Application Number
- CN202510171798.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has problems of sample selection bias and insufficient model robustness in traffic conflict risk prediction in the rapid road intertwined areas, making it difficult to effectively identify key traffic flow characteristics that affect traffic conflict risks.
The causal regularization logistic regression model is adopted to convert traffic flow data into binary variables through the boxing algorithm, establish a causal relationship between traffic conflict risks in the interleaved area and traffic flow characteristics, identify key traffic flow characteristics and achieve real-time stable prediction.
The impact of sample selection bias on model prediction results is reduced, the robustness of traffic conflict prediction models is improved, and a theoretical basis is provided for identifying dangerous traffic flow characteristics and implementing targeted traffic safety control measures.
Smart Images

Figure CN120014880A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of expressway traffic safety, and in particular, relates to a method for predicting stable traffic conflict risks in an expressway interweaving zone. Background Art
[0002] Weaving areas are an important part of expressways. In weaving areas, the distance between the entrance and exit ramps is less than 800 meters, and vehicles entering and exiting the ramps are close to each other, resulting in frequent acceleration, deceleration and lane changes, complex and unstable traffic operation, and has always been a safety bottleneck that cannot be ignored on expressways. Therefore, traffic safety analysis of weaving areas helps to identify dangerous traffic flow characteristics and provide the necessary theoretical basis for timely taking control measures (such as ramp control, variable speed, lane change guidance, etc.).
[0003] In recent years, most of the existing studies have built traffic accident frequency prediction models in weaving areas based on historical accident data, analyzing the impact of factors such as traffic flow and road geometry design in weaving areas on traffic accident frequency. However, the historical accident data collection cycle is long and the quality requirements are high. With the rapid development of traffic safety theory, some scholars have begun to use indirect evaluation indicators of traffic safety, such as traffic conflicts, to establish traffic conflict risk prediction models and evaluate the real-time safety of traffic operation. However, most of the existing studies focus on urban intersections and basic sections of expressways, and the correlation between traffic flow operation characteristics in weaving areas and traffic conflict risks is insufficient.
[0004] At the same time, most of the existing traffic conflict risk models are based on observational data. The modeling process usually assumes that the test data set has the same distribution characteristics as the training data set. However, since the sampling process cannot be fully controlled, selection bias is very likely to occur, resulting in distribution bias between the test set and the training set, which in turn affects the accuracy of the model estimation parameters and reduces the robustness of the model results. Some studies have established heterogeneous models by assigning different regression parameter values to different samples to improve the accuracy of the prediction results, but the regression parameters of such models usually require strict distribution assumptions, and their prediction performance is greatly affected by the complexity of the model and the randomness of the samples. Summary of the invention
[0005] In view of the shortcomings in the prior art, the present invention provides a method for stable prediction of traffic conflict risks in the weaving area of expressways. The method takes the weaving area of expressways as the research object, and based on the stability of causal relationships, utilizes a causal regularized logistic regression model to establish a causal relationship between the traffic conflict risk in the weaving area and the traffic flow characteristics, identifies key traffic flow characteristics that affect the traffic conflict risk in the weaving area, and realizes real-time and stable prediction of the traffic conflict risk in the weaving area, thereby providing a certain theoretical basis for real-time improvement of traffic safety operation in the weaving area of expressways.
[0006] The present invention achieves the above technical objectives through the following technical means.
[0007] A stable prediction method for traffic conflict risk in weaving areas of expressways:
[0008] Based on the traffic conflict samples and non-traffic conflict samples in the weaving area, the corresponding traffic flow data are obtained respectively; the traffic flow data of traffic conflict and non-traffic conflict samples are converted into binary variables using the binning algorithm; using the traffic flow binary variable as the explanatory variable, a causal regularized logistic regression model is constructed to predict whether a traffic conflict will occur and identify the key traffic flow characteristics that affect the risk of traffic conflict.
[0009] Furthermore, the traffic conflict samples are obtained by calculating the traffic conflict index time-to-collision of the front and rear vehicles in the same lane, and the non-traffic conflict samples are randomly obtained from the vehicle trajectories without traffic conflicts in the weaving area.
[0010] Furthermore, the ratio of the traffic conflict samples to the non-traffic conflict samples is 1:k.
[0011] Furthermore, the traffic flow data of the traffic conflict sample is obtained from the trajectory data of all vehicles within a set time T before the traffic conflict occurs and within a range of radius r with the center of gravity of the rear vehicle as the center of the circle.
[0012] Furthermore, the traffic flow data of the non-traffic conflict sample is obtained from all vehicle trajectory data within a set time T when no traffic conflict occurs and within a radius r with the center of gravity of the rear vehicle of the non-traffic conflict sample as the center of the circle.
[0013] Furthermore, the traffic flow data includes the traffic flow through each lane of each section, the average vehicle speed, the speed variation coefficient, the speed difference of adjacent lanes, the section speed difference, the average headway between vehicles, the headway variation coefficient, and the headway difference between vehicles in adjacent lanes.
[0014] Furthermore, the section includes an upstream section and a downstream section, the upstream section is r meters upstream of the vehicle after the traffic conflict sample or the non-traffic conflict sample, and the downstream section is r meters downstream of the vehicle after the traffic conflict sample or the non-traffic conflict sample.
[0015] Furthermore, the binning algorithm includes classification and regression tree binning, chi-square binning and Kolmogorov-Smirnov binning, and the information value index value of the binary variable is calculated according to the information value index, and the binning algorithm corresponding to the maximum information value index value is selected to determine the optimal binning result.
[0016] Furthermore, the specific structure of the causal regularized logistic regression model is:
[0017]
[0018] In the formula: represents the weighted logistic regression loss function; i = 1, 2, ..., n represents the sample number, n is the total number of samples; W i represents the weight corresponding to the i-th sample; X i represents the set of explanatory variables corresponding to sample i; Y i represents the dependent variable corresponding to sample i, Y i =1 means traffic conflict occurs, Y i =0 means no traffic conflict occurs; β represents the model regression coefficient, the dimension of the model regression coefficient is m×1, m is the number of explanatory variables; t=1,2,…,m represents the explanatory variable number; X t represents the tth explanatory variable; X -t Is X t The set of all variables remaining after the values in the explanatory variable set X are set to zero, X -t T Represents X -t The transpose of ; W represents the global sample weight, the dimension of the global sample weight is n×m, W T represents the transpose of the global sample weight; ⊙ represents the Hadamard product; λ 1~ λ5 represents the model trade-off parameter; and the model trade-off parameter λ 1~ The initial values of λ5, model regression coefficient β, and global sample weight W are randomly generated, and 0<λ1≤20, 0<λ2≤1, 0<λ3≤1, 0<λ4≤1, 0<λ5≤20.
[0019] Furthermore, the key traffic flow characteristics that affect the risk of traffic conflict are determined by a model regression coefficient β in a causal regularized logistic regression model.
[0020] Compared with the prior art, the present invention has the following beneficial effects:
[0021] (1) A causal regularized logistic regression model was selected to establish a traffic conflict risk prediction model for the interweaving area of expressways, revealing the causal relationship between the traffic conflict risk and traffic flow characteristics in the interweaving area;
[0022] (2) By utilizing the stability of causal relationships in different environments, the impact of sample selection bias on model prediction results is reduced, thus improving the robustness of the traffic conflict prediction model;
[0023] (3) The traffic conflict risk prediction model for the interweaving area of the expressway obtained by the present invention can provide a certain theoretical basis for identifying the characteristics of dangerous traffic flows in the interweaving area and implementing targeted traffic safety control measures. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a flowchart of the method for stable prediction of traffic conflict risk in the interweaving area of the expressway according to the present invention;
[0025] FIG2( a ) is a schematic diagram of type A weaving in the weaving area of the expressway according to the present invention;
[0026] FIG2( b ) is a schematic diagram of the B-type weaving in the weaving area of the expressway according to the present invention;
[0027] FIG. 3( a ) is a diagram of the present invention using the American A-type interlaced area dataset (Dataset a_a ) as the training set, and the Chinese A-type interlaced area dataset (Dataset a_c ) is the predicted ROC curve corresponding to the test set;
[0028] FIG. 3( b ) is a diagram of the present invention using the American A-type interlaced area dataset (Dataset a_a ) as the training set, and the Chinese B-type interlaced area dataset (Dataset b_c ) is the predicted ROC curve corresponding to the test set;
[0029] FIG. 3( c ) is a diagram of the present invention using the Chinese A-type interlaced area dataset (Dataset a_c ) as the training set, and the Chinese B-type interlaced area dataset (Dataset b_c ) is the predicted ROC curve corresponding to the test set;
[0030] FIG. 3( d ) is a diagram of the present invention using the Chinese A-type interlaced area dataset (Dataset a_c ) as the training set, and the U.S. Type A interweaving area dataset as the test set (Dataset a_a )The corresponding predicted ROC curve diagram;
[0031] FIG. 3( e ) is a diagram of the present invention using the Chinese B-type interlaced area dataset (Dataset b_c ) as the training set, and the Chinese A-type interlaced area dataset (Dataset a_c ) is the predicted ROC curve corresponding to the test set;
[0032] FIG. 3( f ) is a diagram of the present invention using the Chinese B-type interlaced area dataset (Dataset b_c ) as the training set, using the American Type A interweaving area dataset (Dataset a_a ) is the predicted ROC curve corresponding to the test set;
[0033] Figure 4 It is the predicted MAPE distribution diagram of the three models described in the present invention. DETAILED DESCRIPTION
[0034] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments, but the protection scope of the present invention is not limited thereto.
[0035] like Figure 1 As shown, the present invention provides a method for predicting the risk of traffic conflicts in a weaving area of a fast road, the method comprising the following steps:
[0036] Step 1: Manually collect road geometry design information of the weaving area through real scenes and satellite maps, and use a drone to vertically shoot traffic flow operation video directly above the weaving area. The road geometry design information includes the weaving area type, weaving area length, and the number of lanes in the weaving area. The flight altitude of the drone when shooting the video is between 150 and 200 meters.
[0037] Step 2: For the weaving area video shot in step 1, the area between the entrance and exit ramps of the weaving area is taken as the study area, and the Tracker software is used to extract the driving trajectory data of each vehicle in the study area. The driving trajectory data of each vehicle includes the vehicle number, the speed, acceleration and position of the vehicle in the time step s, and the time step s = 0.3 seconds.
[0038] Step 3: For the vehicle trajectory data extracted in step 2, calculate the traffic conflict index time-to-collision (TTC) of the front and rear vehicles in the same lane. The calculation method is as follows:
[0039]
[0040] In formula (1): s f-l Represents the headway between the front and rear vehicles (unit: meter); L l Represents the length of the preceding vehicle (unit: meter); v f and v l Represent the speed of the rear vehicle and the front vehicle respectively (unit: m / s). The samples with TTC less than m seconds are marked as traffic conflict samples, and the time, location and front and rear vehicle numbers of each conflict sample are recorded, and m = 3 seconds.
[0041] Step 4. For each traffic conflict sample in step 3, within a set time T before the traffic conflict occurs, obtain all vehicle trajectory data within a radius r with the center of gravity of the rear vehicle as the center of the circle; based on the vehicle trajectory data, collect traffic flow data with the section r meters upstream of the rear vehicle after the traffic conflict (set as the upstream section) and the section r meters downstream of the rear vehicle (set as the downstream section); wherein, the set time T is 5-10 minutes before the traffic conflict occurs, and the radius r = 100 meters.
[0042] Step 5. For each traffic conflict sample in step 3, randomly obtain vehicle trajectories without traffic conflicts in the study area according to the ratio of traffic conflict samples to non-traffic conflict samples of 1:k to obtain non-traffic conflict samples; and based on the non-traffic conflict samples, within a set time T when no traffic conflict occurs, obtain all vehicle trajectory data within a radius r with the center of gravity of the rear vehicle of the non-traffic conflict sample as the center of the circle; based on the vehicle trajectory data, collect traffic flow data with the section r meters upstream of the rear vehicle of the non-traffic conflict (set as the upstream section) and the section r meters downstream of the rear vehicle (set as the downstream section); wherein, k=10, the set time T is 5-10 minutes before the traffic conflict occurs, and the radius r=100 meters.
[0043] Step 6: The traffic flow data of each traffic conflict sample in step 4 and the traffic flow data of each non-traffic conflict sample in step 5 are converted into binary variables using three binning algorithms, and the information value (IV) value of the binary variable is calculated according to the information value (IV), and the binning algorithm corresponding to the maximum IV value is selected to determine the optimal binning result. Among them, the three binning algorithms are classification and regression tree (CART) binning, chi-square binning, and Kolmogorov-Smirnov (KS) binning.
[0044] Step 7: Using the traffic flow binary variables of traffic conflict and non-traffic conflict obtained in step 6 as explanatory variables and whether a traffic conflict occurs as the dependent variable, a causal regularized logistic regression model is constructed to predict the risk of traffic conflict. The specific structure of the model is:
[0045]
[0046] In formula (2): represents the weighted logistic regression loss function; i = 1, 2, ..., n represents the sample number, n is the total number of samples; W i represents the weight corresponding to the i-th sample; X i represents the set of explanatory variables corresponding to sample i; Y i represents the dependent variable corresponding to sample i, Y i =1 means traffic conflict occurs, Y i =0 means no traffic conflict occurs; β represents the model regression coefficient, the dimension of the model regression coefficient is m×1, m is the number of explanatory variables; t=1,2,…,m represents the explanatory variable number; X t represents the tth explanatory variable; X -t Is X t The set of all variables remaining after the values in the explanatory variable set X are set to zero, X -t T Represents X -t The transpose of ; W represents the global sample weight, the dimension of the global sample weight is n×m, WT represents the transpose of the global sample weight; ⊙ represents the Hadamard product; λ 1~ λ5 represents the model trade-off parameter; and the model trade-off parameter λ 1~ The initial values of λ5, model regression coefficient β, and global sample weight W are randomly generated, and 0<λ1≤20, 0<λ2≤1, 0<λ3≤1, 0<λ4≤1, 0<λ5≤20.
[0047] Select an iterative algorithm to solve the above model. First, given the model trade-off parameter λ 1~ λ5, the initial values of the global sample weight moment W and the model regression coefficient β, fix the global sample weight W, and use the following formula to update the model regression coefficient β:
[0048]
[0049] When updating the model regression coefficient β using formula (3), formula (3) can be regarded as a standard least squares problem regularized with l1 and l2 norms, and solved by the proximal gradient algorithm with proximal operators. On this basis, the model regression coefficient β is fixed, and the global sample weight W is updated. Since the global sample weight W ≥ 0, W = w⊙w, where w is the square root matrix of W elements, and its dimension is n×m, W i =w i ⊙w i , w i It is W i The element square root matrix of , and its dimension is n × 1, the global sample weight W is updated using the following formula:
[0050]
[0051] The key to solving formula (4) is to calculate the partial derivative of F(w) with respect to w, that is, the above formula can be changed to:
[0052]
[0053] In formula (5), matrix 1 T =(1,1,…,1) m , m is the number of explanatory variables, and the partial derivatives of functions G(w) and G(w) with respect to w can be expressed as:
[0054]
[0055] Select gradient descent to solve formula (5), and after updating the model regression coefficient β and the global sample weight W, calculate the difference between two consecutive iterations of the overall loss function of the model. When the difference is less than the threshold δ, terminate the entire training process.
[0056] Step 8. Based on the causal regularized logistic regression model obtained in step 7, the explanatory variables corresponding to the model regression coefficient β that is significantly not equal to 0 within the p% confidence interval are identified as significantly correlated with the traffic conflict risk, and the model regression coefficient β is positive, indicating that it is causally positively correlated with the traffic conflict risk, otherwise it is causally negatively correlated, and finally the key traffic flow characteristics that affect the traffic conflict risk are identified; where p% = 95%.
[0057] Step 9: Divide the traffic conflict and non-traffic conflict samples of different types of weaving areas into training sets and test sets; use the training set to train the causal regularized logistic regression model, conventional logistic regression model and extreme gradient boosting algorithm (XGBoost) in step 8, and use the test set to compare and analyze the predictive stability and robustness of the three models, and select the receiver operating characteristic curve (ROC curve) and the root mean square error (RMSE) between the actual value and the predicted value as evaluation indicators. The larger the area under the ROC curve and the smaller the RMSE, the stronger the corresponding model stability and robustness. Among them, RMSE is calculated according to the following formula:
[0058]
[0059] Where i = 1, 2, ..., n represents the sample number, n is the total number of samples; y i Represents the true value of sample i; Represents the predicted value of sample i.
[0060] According to the above steps, the performance of the present invention in predicting traffic conflict risks was tested using real data from three weaving areas on expressways in Nanjing, China and California, USA. The three weaving areas are defined as type A weaving and type B weaving according to the number of lanes required for merging and diverging vehicles to change lanes. The schematic diagrams of the two weaving types are shown in Figure 1. Figure 2(a) , 2(b) The dataset corresponding to the China Type A interlaced area is named Dataset a_c , the dataset corresponding to the China Type B interlaced area is named Dataset b_c , the dataset corresponding to the Type A interlaced area in the United States is named Dataset a_a According to steps 1-5, traffic flow data of 561 traffic conflicts and 5610 non-traffic conflict samples in the study area of the three weaving areas were obtained. The traffic flow data of each sample includes the average vehicle speed, speed variation coefficient, average headway, headway variation coefficient, speed difference between the target lane and the adjacent lane, headway time difference between the target lane and the adjacent lane, speed difference of the target lane in the upstream and downstream sections, headway time difference, and traffic flow in the upstream and downstream sections.
[0061] The traffic flow data of each sample needs to be processed into the binary variable format required for the input of the causal regularized logistic regression model. According to the three binning algorithms in step 6, the optimal binary result of each traffic flow variable is determined according to the maximum IV value. During the model training process, the sample data of one of the three weaving areas is used as the training set, and the sample data of the remaining two weaving areas is used as the test set. For example, the Chinese Type A weaving area dataset Dataset a_c When used as a training set, the dataset corresponding to China's Type B interlaced area and the United States' Type A interlaced area b_c 、Dataset a_a Finally, taking whether a traffic conflict occurs as the dependent variable and the traffic flow binary variable as the explanatory variable, the causal regularized logistic regression model in step 7 is used to predict the risk of traffic conflict in the weaving area. The five trade-off parameters λ corresponding to the causal regularized logistic regression model under each training set are 1~ The optimal value of λ5 is shown in Table 1. The key traffic flow characteristics that affect the risk of traffic conflict in China's Type A and Type B weaving areas are obtained according to step 8, as shown in Table 2.
[0062] Table 1 Optimal values of trade-off parameters in the causal regularized logistic regression model
[0063]
[0064] Table 2 Key variables affecting traffic conflict risk in two types of intersections in China
[0065]
[0066] Note: A model regression coefficient with a sign of "+" indicates that the variable is positively correlated with the risk of traffic conflict, and a model regression coefficient with a sign of "-" indicates that the variable is negatively correlated with the risk of traffic conflict.
[0067] According to the model results, the risk of traffic conflict in China's Type A weaving area is significantly correlated with the traffic flow characteristics of the upstream section. For example, when the traffic volume of the upstream section is between 290-304 vehicles / 5 minutes, the risk of traffic conflict increases. When the average speed of the target lane of the upstream section is in an unstable traffic flow state between 12.4-14.9 m / s, the risk of traffic conflict is also high. In addition, a small speed difference between the target lane and the adjacent lane of the upstream section will also increase the risk of traffic conflict. Due to the difference in traffic operation between Type B and Type A weaving areas, the factors affecting the risk of traffic conflict in China's Type B weaving area are significantly heterogeneous with those in Type A weaving areas. For example, when the average speed of the target lane of the upstream section is in an unstable traffic flow state between 14.0-15.8 m / s, the risk of traffic conflict in the Type B weaving area is reduced. The greater the speed difference between the target lane and the adjacent lane of the upstream section, the higher the risk of traffic conflict. In addition, when the traffic volume of the upstream section is greater than 316 vehicles / 5 minutes, the risk of traffic conflict decreases, indicating that compared with the Type A weaving area, the Type B weaving area is more suitable for high traffic demand scenarios. Therefore, the method of the present invention can effectively identify key traffic flow characteristics that affect the risk of traffic conflict in weaving areas and has practical application value.
[0068] Finally, in order to verify the stability of the traffic conflict risk prediction results of the method of the present invention, according to step 9, a conventional logistic regression model and an XGBoost model were established based on the same training set. The prediction accuracy of the three models in different test sets was compared and analyzed using the ROC curve. The results are shown in Figure 2. Figure 3(a)-3(f) As shown. According to the area under the ROC curve (AUC), the causal regularized logistic regression model of the present invention has the largest AUC for the prediction results under most combinations of training sets and test sets. At the same time, Figure 4 The root mean square error distribution (RMSE) between the actual value and the predicted value of the dependent variable of the three models on different test sets is displayed. The results also show that the causal regularized logistic regression model of the present invention has the smallest MAPE under most combinations of training sets and test sets, indicating that the causal regularized logistic regression model proposed in the present invention can more stably predict the risk of traffic conflicts in interweaving areas and has certain practical application value.
[0069] The embodiments are preferred implementations of the present invention, but the present invention is not limited to the above-mentioned implementations. Any obvious improvements, substitutions or modifications that can be made by those skilled in the art without departing from the essential content of the present invention belong to the protection scope of the present invention.
Claims
1. A stable prediction method for traffic conflict risk in an expressway weaving zone, characterized by: Based on the traffic conflict samples and non-traffic conflict samples in the weaving area, the corresponding traffic flow data are obtained respectively; the traffic flow data of the traffic conflict samples and the traffic flow data of the non-traffic conflict samples are converted into binary variables using a binning algorithm; Taking the binary variable of traffic flow as the explanatory variable, a causal regularized logistic regression model is constructed to predict whether traffic conflict will occur and to identify the key traffic flow characteristics that affect the risk of traffic conflict.
2. The method for predicting the risk of traffic conflicts in the interweaving area of a fast road according to claim 1 is characterized in that: The traffic conflict samples are obtained by calculating the traffic conflict index time-to-collision of the front and rear vehicles in the same lane, and the non-traffic conflict samples are randomly obtained from the vehicle trajectories without traffic conflicts in the weaving area.
3. The method for predicting stable traffic conflict risk in an expressway weaving zone according to claim 1, characterized in that: The ratio of the traffic conflict samples to the non-traffic conflict samples is 1:k.
4. The method for predicting stable traffic conflict risk in an expressway weaving zone according to claim 1, characterized in that: The traffic flow data of the traffic conflict sample is obtained from the trajectory data of all vehicles within a set time T before the traffic conflict occurs and within a radius r with the center of gravity of the rear vehicle as the center of the circle.
5. The method for predicting stable traffic conflict risk in an expressway weaving zone according to claim 1, characterized in that: The traffic flow data of the non-traffic conflict sample is obtained from all vehicle trajectory data within a set time T when no traffic conflict occurs and within a radius r with the center of gravity of the rear vehicle of the non-traffic conflict sample as the center of the circle.
6. The method for predicting stable traffic conflict risk in an expressway weaving zone according to claim 1, characterized in that: The traffic flow data includes the traffic flow through each lane of each section, the average vehicle speed, the speed variation coefficient, the speed difference of adjacent lanes, the section speed difference, the average headway between vehicles, the headway variation coefficient, and the headway difference between vehicles in adjacent lanes.
7. The method for predicting stable traffic conflict risk in an expressway weaving zone according to claim 6, characterized in that: The section includes an upstream section and a downstream section. The upstream section is r meters upstream of the vehicle after the traffic conflict sample or the non-traffic conflict sample, and the downstream section is r meters downstream of the vehicle after the traffic conflict sample or the non-traffic conflict sample.
8. The method for predicting stable traffic conflict risk in an expressway weaving zone according to claim 1, characterized in that: The binning algorithms include classification and regression tree binning, chi-square binning and Kolmogorov-Smirnov binning, and the information value index value of the binary variable is calculated according to the information value index, and the binning algorithm corresponding to the maximum information value index value is selected to determine the optimal binning result.
9. The method for predicting stable traffic conflict risk in an expressway weaving zone according to claim 1, characterized in that: The specific structure of the causal regularized logistic regression model is: In the formula: represents the weighted logistic regression loss function; i = 1, 2, ..., n represents the sample number, n is the total number of samples; W i represents the weight corresponding to the i-th sample; X i represents the set of explanatory variables corresponding to sample i; Y i represents the dependent variable corresponding to sample i, Y i =1 means traffic conflict occurs, Y i =0 means no traffic conflict occurs; β represents the model regression coefficient, the dimension of the model regression coefficient is m×1, m is the number of explanatory variables; t=1,2,…,m represents the explanatory variable number; X t represents the tth explanatory variable; X -t Is X t The set of all variables remaining after the values in the explanatory variable set X are set to zero, X -t T Represents X -t The transpose of ; W represents the global sample weight, the dimension of the global sample weight is n×m, W T represents the transpose of the global sample weight; ⊙ represents the Hadamard product; λ 1~ λ5 represents the model trade-off parameter; and the model trade-off parameter λ 1~ The initial values of λ5, model regression coefficient β, and global sample weight W are randomly generated, and 0<λ1≤20, 0<λ2≤1, 0<λ3≤1, 0<λ4≤1, 0<λ5≤20.
10. The method for predicting stable traffic conflict risk in an expressway weaving zone according to claim 9, characterized in that: The key traffic flow characteristics that affect the risk of traffic conflict are determined by the model regression coefficient β in the causal regularized logistic regression model.