Traffic simulation parameter calibration method based on GMM clustering and Bayesian optimization

Different driving styles are identified through GMM clustering and the traffic simulation parameter calibration is combined with Bayesian optimization algorithm, which solves the problems of instability and low computational efficiency of existing methods when dealing with heteroscedastic noise and driving behavior differences, achieving a more efficient and reliable calibration effect.

CN120046459APending Publication Date: 2025-05-27SOUTHEAST UNIV +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411912765.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing traffic simulation parameter calibration methods have problems of instability and low computational efficiency when dealing with heteroscedastic noise and driving behavior differences.

Method used

Gaussian hybrid model (GMM) clustering is used to identify different driving styles, and parameter calibration is performed through Bayesian optimization algorithm, taking into account the influence of heteroscedastic noise, and improving calibration robustness and computational efficiency.

Benefits of technology

The heterogeneity of driving behavior is realized more accurately, the robustness and calculation efficiency of parameter calibration are improved, and the stability and reliability of calibration results are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046459A_ABST
    Figure CN120046459A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic simulation parameter calibration method based on GMM clustering and Bayesian optimization, and the method comprises the following steps: inputting and preprocessing data, including data collection, driving style clustering analysis and simulation environment construction; calculating simulation repetition times of the simulation parameter sample points; according to a simulation result of the simulation parameter sample point, estimating expectation and variance distribution of an objective function; and determining a next simulation parameter sampling point based on a Bayesian optimization method. The traffic simulation parameter calibration method considering the heterovariance noise is established, the accuracy of parameter calibration is improved by identifying different driving styles, the mean square percentage error is reduced from 20.2% to 3.1%, and an effective tool is provided for the accuracy of highway traffic simulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a traffic simulation parameter calibration method based on GMM clustering and Bayesian optimization, and belongs to the field of traffic simulation and optimization algorithms. Background Art

[0002] With the advancement of urbanization and the rapid growth of traffic demand, accurate traffic simulation has become an important tool for traffic planning and management. Microscopic traffic simulation has been widely used in traffic system analysis and evaluation because it can carefully depict traffic flow characteristics. In order to ensure the accuracy of the simulation results, it is necessary to calibrate the parameters of the simulation model so that the simulation output matches the actual observation data. In essence, this is a simulation-based optimization problem, and the goal is to make the simulation output match the actual measurement value as much as possible by adjusting the simulation parameters.

[0003] Existing parameter calibration methods mainly include the following categories: First, methods based on heuristic algorithms, such as genetic algorithms and simulated annealing, which can find the optimal solution under nonlinear conditions, but the computational cost grows exponentially with the number of parameters; second, methods based on gradients, such as the stochastic perturbation simultaneous approximation algorithm (SPSA), which have high computational efficiency but are prone to fall into local optimality; third, methods based on proxy models, such as Kriging interpolation, which improve optimization efficiency by constructing approximate models.

[0004] However, these methods still have the following problems in practical applications:

[0005] First, it is usually assumed that the simulation randomness is constant under any conditions, ignoring the differences in randomness in different scenarios. In fact, the randomness of traffic simulation will change with the change of traffic status, and in some cases show obvious heteroscedasticity characteristics. Ignoring this heteroscedasticity will lead to unstable calibration results.

[0006] Second, there is a lack of consideration for differences in driving behavior, and a unified parameter setting is often used. In actual traffic flow, different drivers show obvious behavioral differences, including differences in expected speed, following distance and other characteristics. It is difficult to accurately describe this behavioral heterogeneity using a unified parameter.

[0007] Third, the computational efficiency of existing calibration methods is low, especially when dealing with large-scale simulation scenarios. Traditional methods often require a large number of simulation evaluations, resulting in long computational time and difficulty meeting the needs of actual applications. At the same time, for cases with heteroscedastic noise, it is necessary to increase the number of repeated simulations to obtain reliable estimates, further increasing the computational burden.

[0008] In addition, existing methods rarely consider the variance when calibrating simulations, which makes the calibrated simulations less robust in practical applications. A good calibration method should not only consider the expected value of the simulation output, but also pay attention to its variance performance to ensure the stability and reliability of the calibration results. Summary of the invention

[0009] To solve the above problems, the present invention proposes a traffic simulation parameter calibration method based on GMM clustering and Bayesian optimization. The method of the present invention first uses a Gaussian mixture model to perform cluster analysis on driving behavior and identify different driving styles; then considers the influence of heteroscedastic noise and expresses the parameter calibration problem as a simulation-based robust optimization problem; finally, it is solved by a Bayesian optimization algorithm to improve the computational efficiency while ensuring the calibration accuracy.

[0010] In order to solve the above technical problems, the present invention provides a traffic simulation parameter calibration method based on GMM clustering and Bayesian optimization, comprising the following steps:

[0011] Step S1: input and preprocess data, including data collection, driving style cluster analysis and simulation environment construction;

[0012] Step S2: Calculate the number of simulation repetitions of the simulation parameter sample points;

[0013] Step S3: Estimate the expected value and variance distribution of the objective function according to the simulation results of the simulation parameter sample points;

[0014] Step S4: Determine the next simulation parameter sampling point based on the Bayesian optimization method, repeat the above steps for the sampling point, and iterate to find the optimal simulation parameter.

[0015] In step S1, data is input and preprocessed. First, highway ETC data, gantry data, and 100-meter average speed data are collected. ETC data must include information such as license plate number, entry and exit time, and driving direction; gantry data contains traffic volume and average speed data; and 100-meter mileage data contains spatial average speed data. The data is screened to eliminate outliers and unreasonable data, and a suitable time period is selected as the research sample.

[0016] Then, a Gaussian mixture model is used to cluster the vehicle driving characteristics. The main steps of GMM clustering include:

[0017] 1) Calculate the characteristic indicators of each vehicle, including average speed, speed standard deviation, etc.;

[0018] 2) Establish a Gaussian mixture model:

[0019]

[0020] Where x is the eigenvector, π k is the weight of the kth Gaussian distribution, μ k and∑ k are the mean vector and covariance matrix respectively;

[0021] 3) Use the expectation maximization algorithm to iteratively optimize model parameters;

[0022] 4) Driving styles are divided into aggressive, moderate and conservative types.

[0023] Finally, establish the simulation environment:

[0024] 1) Use SUMO to build a simulated road network in the study area;

[0025] 2) Set vehicle parameters such as expected speed according to different driving styles;

[0026] 3) Determine the initial value range of the parameter to be calibrated.

[0027] In step S2, the number of repetitions of the simulation parameter sample points is calculated. This step formulates the parameter calibration problem as an optimization problem considering heteroscedastic noise:

[0028] minE[f(θ)]

[0029] Constrained by:

[0030] Var[f(θ)]≤σc 2

[0031] θ l ≤θ≤θ m

[0032] Where θ represents the simulation parameter, f(θ) represents the difference between the simulation output and the actual measurement value, E represents the expectation, Var represents the variance, σc 2 represents the threshold of simulation variance, θ l and θ m They represent the lower and upper bounds of the parameters respectively. f(θ) is a non-convex, non-differentiable function with heteroscedastic noise.

[0033] The specific steps include:

[0034] 1) Perform minit initial simulations on the simulation parameter θ;

[0035] 2) Based on the current number of simulations, estimate the distribution of the expected E[f(θ)] and variance Var[f(θ)] corresponding to the simulation parameter θ;

[0036] 3) Check whether the condition Var[f(θ)]≤σc is satisfied 2 , if satisfied, continue, otherwise complete sampling;

[0037] 4) Check optimization conditions:

[0038] Calculate the probability that the current sampling point is better than the best sampling point:

[0039] p(E[f(θ)] <f * )≥1-αcE

[0040] Calculate the probability that the current sampling point is worse than the best sampling point:

[0041] p(E[f(θ)]>f * )≥1-αcE

[0042] where f * represents the optimal result obtained in the previous sampling, and αcE is the predefined significance level; if none of the above conditions are met, proceed to the next step. If any of the above conditions is met, the simulation parameter is no longer simulated, and the expected E[f(θ)] and variance Var[f(θ)] corresponding to the simulation parameter are obtained, and the sampling of the sampling point is completed;

[0043] 5) If none of the optimization conditions are met, perform madd additional simulations and check repeatedly.

[0044] In step S3, the expectation and variance distribution of the objective function are estimated based on the simulation results of the simulation parameter sample points. This step specifically includes:

[0045] 1) Use Bayesian inference to estimate the expectation E[f(θ)]:

[0046] p(E[f(θ)]|f(1),...,f(m))∝p(E[f(θ)])p(f(1),...,f(m)|E[f(θ)])

[0047] where f(1),...,f(m) denotes the evaluation of f(θ) after m simulation repetitions;

[0048] 2) Without losing generality, assume f(θ)~N(E[f(θ)],Var[f(θ)]), and obtain the posterior distribution of the expected E[f(θ)]:

[0049] p(E[f(θ)]|f(1),...,f(m),Var[f(θ)])=N(μm,σm 2 )

[0050] in, represents the mean of f;

[0051] 3) Estimate the variance Var[f(θ)]:

[0052]

[0053] In step S4, the next simulation parameter sampling point is determined based on the Bayesian optimization method. This step mainly includes constructing a Gaussian process proxy model considering heteroscedastic noise and designing a sampling strategy.

[0054] First, construct a Gaussian process agent model:

[0055] Define (θi, yi, ri) as the sampling point of Bayesian optimization, where θi is the simulation parameter, yi, ri are the estimated values ​​of E[f(θ)] and Var[f(θ)] respectively:

[0056] Assuming there are n sampling points, let D = {(θ1, y1, r1), ..., (θn, yn, rn)};

[0057] Establishing agent model based on Gaussian process;

[0058] Create a multivariate Gaussian distribution:

[0059] Yn~N(μn,Kn+R)

[0060] Where Yn=(y1,...,yn)T, μn is the expectation of the multivariate Gaussian distribution, and the initial μn is set to 0;

[0061] Construct the covariance matrix Kn:

[0062] Kn=[Cov(y1,y1)…Cov(y1,yn)

[0063] ………

[0064] Cov(yn,y1)...Cov(yn,yn)]

[0065] Define the kernel function:

[0066]

[0067] Where l is the hyperparameter of the kernel function;

[0068] Construct the diagonal matrix R:

[0069] The elements on the main diagonal of R are (r1 / m1,...,rn / mn) T .

[0070] Define another Gaussian process:

[0071] log(Var[f(θ)])~GP(·,·)

[0072] For a given θ′, the posterior distribution of log(r′) is:

[0073] log(r′)|θ~N(K*TKn-1Rn,K*-K*TKn-1K*)

[0074] Where Rn=(log(r1),...,log(rn)) T

[0075] Finally, two sampling strategies are designed:

[0076] Strategy 1: Take the expected improvement acquisition function as the objective function and use the proxy model of log(Var[f(θ)]) as the variance constraint;

[0077]

[0078] bound by;

[0079] τ(θ)≤log(σc)

[0080] θ l ≤θ≤θ m

[0081] where τ(θ) is the expectation of the log(r′)|θ distribution.

[0082] Strategy 2:

[0083] 1) Probability of introducing variance constraints:

[0084]

[0085] where γ(θ) is the variance of the log(r′)|θ distribution;

[0086] Let y* denote the minimum objective function value so far. The basic idea of ​​expected improvement is to sample a new point so that the expected value of y*-y(θ) reaches the maximum value:

[0087] maxE[y * -y(θ)]

[0088] 2) Modify the acquisition function:

[0089] maxpc(θ)E[y*-y(θ)]

[0090] Constrained by:

[0091] θ l ≤θ≤θ m

[0092] Solve the modified acquisition function:

[0093]

[0094] Constrained by:

[0095] θ l ≤θ≤θ m

[0096] where φ and are the cumulative distribution function and probability density function of the Gaussian distribution, μ(θ) is the expectation of y(θ)|θ, and σ(θ) is the variance of y(θ)|θ.

[0097] When the sampling point is determined in each iteration, the next sampling point θnew is selected according to strategy 1 or strategy 2; the simulation of θnew is repeated for an adaptive number of times; the objective function value is calculated and the proxy model is updated;

[0098] The preset maximum number of iterations is reached; the improvement in multiple consecutive iterations is less than the threshold; the predetermined accuracy requirements are met.

[0099] The specific steps of repeating the simulation for the adaptive number of times of θnew are:

[0100] Step 11.1 Input parameters:

[0101] Predefined significance level αcE, initial number of simulations minit, additional number of simulations madd;

[0102] Step 11.2 Perform the initial simulation:

[0103] When the input is the simulation parameter θ, minit initial simulations are performed;

[0104] Step 11.3 Estimate the distribution and check the conditions:

[0105] Estimate the distribution of the expected E[f(θ)] and variance Var[f(θ)] corresponding to the simulation parameter θ, and check whether the condition Var[f(θ)]<σc is satisfied 2 ;

[0106] Step 11.4 Check the quality judgment conditions:

[0107] Check whether p(E[f(θ)] <f * )≥1-αcE or p(e[f(θ)]>f * )≥1-αcE

[0108] where f * represents the best result obtained previously;

[0109] Step 11.5 Perform additional simulations:

[0110] If none of the conditions are met, perform madd additional simulations for the sampling point and return to step 11.3.

[0111] According to another aspect of the present invention, the method of the present invention is applied to the following scenarios:

[0112] Calibration of highway traffic flow simulation parameters;

[0113] Modeling vehicle behavior considering different driving styles;

[0114] Simulation systems that need to consider the effects of heteroscedastic noise;

[0115] Large-scale simulation optimization problems with limited computing resources.

[0116] According to another aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of the traffic simulation parameter calibration method based on GMM clustering and Bayesian optimization of the present invention are implemented.

[0117] According to another aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the traffic simulation parameter calibration method based on GMM clustering and Bayesian optimization of the present invention are implemented.

[0118] Compared with the prior art, the above method of the present invention has the following beneficial effects:

[0119] 1. The heteroscedastic noise in traffic simulation is taken into account to improve the robustness of parameter calibration;

[0120] 2. GMM clustering is introduced to identify different driving styles, which more accurately describes the heterogeneity of driving behavior;

[0121] 3. The Bayesian optimization algorithm is used for parameter calibration, which improves the computational efficiency while ensuring accuracy; BRIEF DESCRIPTION OF THE DRAWINGS

[0122] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0123] Figure 1 A schematic diagram of a method framework of a preferred embodiment of the present invention;

[0124] Figure 2 A driving style comparison diagram based on GMM clustering according to a preferred embodiment of the present invention; (a) clustering result of driving behavior of small cars; (b) clustering result of driving behavior of small cars;

[0125] Figure 3 A comparison diagram of the calibration effect of the parameters after calibration and the default parameters of a preferred embodiment of the present invention;

[0126] Figure 4 The present invention is a method flow chart of a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0127] To make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the embodiments of the present invention will be described clearly and completely in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them.

[0128] Unless otherwise defined, technical or scientific terms used herein shall have the same meaning as commonly understood by one of ordinary skill in the art to which the invention belongs.

[0129] Example 1: Figure 1-4 As shown, a traffic simulation parameter calibration method based on GMM clustering and Bayesian optimization includes three main steps: simulation environment construction and data processing, driving style clustering and parameter calibration.

[0130] Step S1. Data input and preprocessing;

[0131] In this paper, the section from K188 to K215 of a certain highway is taken as the research area. First, three types of data are collected and sorted:

[0132] ETC data: Contains OD information such as license plate number, entry and exit time, and driving direction. Each record contains the complete trip information of the vehicle in the study area and can be used to calculate the average driving speed of the vehicle.

[0133] Gantry data: record the average speed of vehicles passing through, counted every 10 minutes.

[0134] Hundred-meter mileage data: records the spatial average speed of each 100-meter section of road, provided by the online map platform.

[0135] By analyzing the spatiotemporal speed distribution map, the non-congested period data on September 6, 2023 were selected for GMM cluster analysis.

[0136] Data screening criteria include:

[0137] 1. Select the vehicle data between the two farthest gantries to ensure sufficient driving distance;

[0138] 2. Eliminate local congestion data to ensure free flow status;

[0139] 3. According to the characteristics of the vehicle models, the vehicles are divided into two categories: small vehicles (sedans, vans) and large vehicles (passenger cars, trucks). GMM clustering is used for each type of vehicle:

[0140] 1. Calculate characteristic indicators: average speed, speed standard deviation, etc.

[0141] 2. Clustering using GMM algorithm:

[0142]

[0143] Where x is the eigenvector, πk is the weight of the kth Gaussian distribution, μk and ∑k are the mean vector and covariance matrix respectively;

[0144] 3. Use the expectation maximization algorithm to iteratively optimize model parameters

[0145] 4. Driving styles are classified into aggressive, moderate and conservative

[0146] Finally, SUMO was used to build a simulation model of the research section: the reference speed for small cars was 120 km / h, and the reference speed for large cars was 100 km / h.

[0147] The SpeedFactor parameter adopts a truncated normal distribution:

[0148] SpeedFactor=normc(mean,deviation,lowerCutOff,upperCutOff)

[0149] Step S2. Calculate the number of repetitions of the simulation parameter sample points:

[0150] This step first performs minit initial simulations. Then, based on the current number of simulations, the distribution of the expected E[f(θ)] and variance Var[f(θ)] corresponding to the simulation parameter θ is estimated, and the condition Var[f(θ)]<σc is checked to see whether it satisfies 2 .

[0151] Where θ represents the simulation parameter, f(θ) represents the difference between the simulation output and the actual measurement value, E represents the expectation, Var represents the variance, σc 2 represents the threshold of simulation variance, θ l and θ m They represent the lower and upper bounds of the parameters respectively. f(θ) is a non-convex, non-differentiable function with heteroscedastic noise.

[0152] If the condition is met, continue to check the optimization conditions:

[0153] Calculate the probability that the current sampling point is better than the best sampling point:

[0154] p(E[f(θ)] <f * )≥1-acE

[0155] Calculate the probability that the current sampling point is worse than the best sampling point:

[0156] p(E[f(θ)]>f* )≥1-acE

[0157] where f * represents the best result obtained in the previous sampling, and αcE is the predefined significance level.

[0158] If none of the optimization conditions are met, perform madd additional simulations and check repeatedly.

[0159] Step S3 estimates the expected and variance distribution of the objective function:

[0160] Use Bayesian inference to estimate the expectation E[f(θ)]:

[0161] p(E[f(θ)]|f(1),...,f(m))∝p(E[f(θ)])p(f(1),...,f(m)|E[f(θ)])

[0162] Assuming f(θ)~N(E[f(θ)],Var[f(θ)]), we can get the posterior distribution of the expected E[f(θ)]:

[0163] p(E[f(θ)]|f(1),...,f(m),Var[f(θ)])=N(μm,σm 2 )

[0164] At the same time, calculate the variance Var[f(θ)]:

[0165]

[0166] Step S4 determines the next sampling point based on Bayesian optimization:

[0167] The next simulation parameter sampling point is determined based on the Bayesian optimization method. This step mainly includes constructing a Gaussian process surrogate model considering heteroscedastic noise and designing a sampling strategy.

[0168] First, construct a Gaussian process agent model:

[0169] Define (θi, yi, ri) as the sampling point of Bayesian optimization, where θi is the simulation parameter, yi, ri are the estimated values ​​of E[f(θ)] and Var[f(θ)] respectively:

[0170] Assuming there are n sampling points, let D = {(θ1, y1, r1), ..., (θn, yn, rn)};

[0171] Establishing agent model based on Gaussian process;

[0172] Create a multivariate Gaussian distribution:

[0173] Yn~N(μn,Kn+R)

[0174] Finally, two sampling strategies are designed:

[0175] Strategy 1: Take the expected improvement acquisition function as the objective function and use the proxy model of log(Var[f(θ)]) as the variance constraint;

[0176]

[0177] bound by;

[0178] τ(θ)≤log(σc)

[0179] θ l ≤θ≤θ m

[0180] where τ(θ) is the expectation of the log(r′)|θ distribution.

[0181] Strategy 2:

[0182] 1) Probability of introducing variance constraints:

[0183]

[0184] where γ(θ) is the variance of the log(r′)|θ distribution;

[0185] Let y* denote the minimum objective function value so far. The basic idea of ​​expected improvement is to sample a new point so that the expected value of y*-y(θ) reaches the maximum value:

[0186] maxE[y * -y(θ)]

[0187] 2) Modify the acquisition function:

[0188] maxpc(θ)E[y*-y(θ)]

[0189] Constrained by:

[0190] θ l ≤θ≤θ m

[0191] Solve the modified acquisition function:

[0192]

[0193] Constrained by:

[0194] θ l ≤θ≤θ m

[0195] where φ and are the cumulative distribution function and probability density function of the Gaussian distribution, μ(θ) is the expectation of y(θ)|θ, and σ(θ) is the variance of y(θ)|θ.

[0196] Experimental verification shows that (such as Figure 2 and Figure 3 As shown in Figure 2, both strategies can reduce the estimated value of the objective function by about 35% at a low computational cost. In terms of solution accuracy, strategy 2 is better than strategy 1. Compared with the MAPE of 20.2% using default parameters, this method reduces the error to 3.1%. A MAPE of 5.01% was achieved in the model robustness test, verifying the stability and scalability of the method.

[0197] In summary, the present invention combines Gaussian mixture model clustering and Bayesian optimization algorithm, facing the highway traffic environment, first uses ETC data, gantry data and 100-meter mileage data to identify different driving style characteristics, and then considers the influence of heteroscedastic noise, and performs parameter calibration through Bayesian optimization, thereby achieving high-precision calibration of traffic simulation parameters.

[0198] Embodiment 2:

[0199] The computer-readable storage medium of this embodiment stores a computer program thereon, and when the program is executed by a processor, the steps in the traffic simulation parameter calibration method based on GMM clustering and Bayesian optimization of Embodiment 1 are implemented.

[0200] The computer-readable storage medium of this embodiment may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal; the computer-readable storage medium of this embodiment may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card, a secure digital card, a flash memory card, etc. equipped on the terminal; further, the computer-readable storage medium may also include both an internal storage unit of the terminal and an external storage device.

[0201] The computer-readable storage medium of this embodiment is used to store computer programs and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.

[0202] Embodiment 3:

[0203] The computer device of this embodiment includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the traffic simulation parameter calibration method based on GMM clustering and Bayesian optimization of Example 1 are implemented.

[0204] In this embodiment, the processor may be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, readily available programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0205] Those skilled in the art will appreciate that the disclosed content of the embodiments may be provided as methods, systems, or computer program products. Therefore, the present solution may adopt the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Moreover, the present solution may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program codes.

[0206] The present solution is described with reference to the flowchart and / or block diagram of the method and computer program product according to the embodiment of the present solution. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions; these computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0207] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0208] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0209] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0210] The examples described in the present invention are merely descriptions of the preferred implementation modes of the present invention, and are not intended to limit the concept and scope of the present invention. Without departing from the design concept of the present invention, various modifications and improvements made to the technical solutions of the present invention by engineers and technicians in this field should all fall within the protection scope of the present invention.

Claims

1. A traffic simulation parameter calibration method based on GMM clustering and Bayesian optimization, characterized in that: The steps include: Input and preprocess data, including data collection, driving style cluster analysis, and simulation environment construction; Calculate the number of simulation repetitions for the simulation parameter sample points; Estimate the expected and variance distribution of the objective function based on the simulation results of the simulation parameter sample points; The next simulation parameter sampling point is determined based on the Bayesian optimization method, and the above steps are repeated for the simulation parameter sampling points, iterating to find the optimal simulation parameters.

2. The method according to claim 1, characterized in that The specific steps for inputting and preprocessing data are as follows: Step S1.1 Data collection and screening: Select highway sections within the study area and collect ETC data; Obtain the traffic flow and average speed data of the gantry on the corresponding road section; Collect average speed data for 100-meter sections in the study area; Eliminate abnormal data and select data samples of time periods suitable for parameter calibration; Step S1.2 Driving style recognition: Calculate the characteristic index of each vehicle; Cluster analysis using Gaussian mixture model: Where x is the eigenvector, π k is the weight of the kth Gaussian distribution, μ k and∑ k are the mean vector and covariance matrix respectively; Iteratively optimize model parameters using the expectation-maximization algorithm; According to the clustering results, driving styles are divided into aggressive, conservative and moderate types; Step S1.3 Simulation environment construction: Use SUMO to build a simulated road network in the study area; Set vehicle parameters; Determine the initial range of the parameter to be calibrated.

3. The method according to claim 1, characterized in that The specific steps for calculating the number of simulation repetitions of simulation parameter sample points are: Among them, minit represents the initial simulation times, madd represents the additional simulation times, and these parameters are used to control the simulation calculation times z; The robust calibration problem of traffic simulation considering heteroscedasticity is solved and mathematically expressed as a simulation-based robust optimization problem: minE[f(θ)] Constrained by: Var[f(θ)]≤σc 2 i l ≤θ≤θ m Where θ represents the simulation parameter, f(θ) represents the difference between the simulation output and the actual measurement value, E represents the expectation, Var represents the variance, σc 2 represents the threshold of simulation variance, θ l and θ m They represent the lower and upper bounds of the parameters respectively. f(θ) is a non-convex, non-differentiable function with heteroscedastic noise. Step S2.1 performs initial simulation on the simulation parameter θ: Set the initial simulation times minit; Repeat the simulation for minit times for the simulation parameter θ of the sampling point; Record the output result f(i) of each simulation, i=1,2,...,minit; Step S2.2 estimates the expectation and variance: Based on the current number of simulations, estimate the distribution of the expected E[f(θ)] and variance Var[f(θ)] corresponding to the simulation parameter θ; Check whether the constraint Var[f(θ)]≤σc is satisfied 2 ; If satisfied, proceed to the next step; if not satisfied, complete the sampling of the sampling point; Step S2.3 Check the optimization conditions: Calculate the probability that the current sampling point is better than the best sampling point: p(E[f(θ)] <f * )≥1-αcE Calculate the probability that the current sampling point is worse than the best sampling point: p(E[f(θ)]>f * )≥1-αcE where f * represents the best result obtained in the previous sampling, and acE is the predefined significance level; if none of the above conditions are met, proceed to the next step. If any of the above conditions is met, the simulation parameter is no longer simulated, and the expected E[f(θ)] and variance Var[f(θ)] corresponding to the simulation parameter are obtained, and the sampling of the sampling point is completed; Step S2.4: perform madd simulation repetitions for the sampling points with simulation parameter θ; then execute step S2.3 to check whether the conditions are met and iterate the loop.

4. The method according to claim 1, characterized in that According to the simulation results of the simulation parameter sample points, the specific steps for estimating the expected and variance distributions of the objective function are: Step S3.1: When calculating the expectation E[f(θ)], Bayesian inference is used to estimate E[f(θ)]: p(E[f(θ)]|f(1),...,f(m))∝p(E[f(θ)])p(f(1),...,f(m)|E[f(θ)]) where f(1), ..., f(m) denotes the evaluation of f(θ) after m simulation repetitions; Step S3.2: Assuming f(θ)~N(E[f(θ)], Var[f(θ)]), the posterior distribution of the expected E[f(θ)] is: p(E[f(θ)]|f(1),...,f(m),Var[f(θ)])=N(μm,σm 2 ) in, represents the mean of f; Step S3.3: The formula for estimating the variance Var[f(θ)] is:

5. The method according to claim 1, characterized in that The specific steps of determining the next simulation parameter sampling point based on the Bayesian optimization method, repeating the above steps for the simulation parameter sampling points, and iterating to find the optimal simulation parameters are as follows: Step S4.1 constructs a Gaussian process surrogate model considering heteroscedastic noise: Define (θi, yi, ri) as the sampling point of Bayesian optimization, where θi is the simulation parameter, yi, ri are the estimated values ​​of E[f(θ)] and Var[f(θ)] respectively: Assuming there are n sampling points, let D = {(θ1, y1, r1), ..., (θn, yn, rn)); Establishing agent model based on Gaussian process; Step S4.2 Design sampling strategy: Strategy 1: Take the expected improvement acquisition function as the objective function and use the proxy model of log(Var[f(θ)]) as the variance constraint; Strategy 2: Introduce the probability of variance constraint pc(θ) and modify the objective function of the acquisition function; Step S4.3 selects new sampling points based on the sampling strategy and updates the model.

6. The method according to claim 5, characterized in that Strategy 1 involves solving the following optimization problem; bound by; τ(θ)≤log(σc) i l ≤θ≤θ m where τ(θ) is the expectation of the log(r′)|θ distribution.

7. The method according to claim 5, characterized in that Strategy 2 includes the following steps: The probability pc(θ) of introducing variance constraints to perform point stratification is: where γ(θ) is the variance of the log(r′)|θ distribution; Let y* denote the minimum objective function value so far. The basic idea of ​​expected improvement is to sample a new point so that the expected value of y*-y(θ) reaches the maximum value: maxE[y * -y(θ)] The modified acquisition function is: maxpc(θ)E[y*-y(θ)] Constrained by: i l ≤θ≤θ m Solve for the modified acquisition function: Constrained by: i l ≤θ≤θ m where Φ and are the cumulative distribution function and probability density function of the Gaussian distribution, μ(θ) is the expectation of y(θ)|θ, and σ(θ) is the variance of y(θ)|θ.

8. The method according to claim 5, characterized in that The construction of the Gaussian process surrogate model includes: Step 4.1.1 Establish multivariate Gaussian distribution: Yn~N(μn,Kn+R) Where Yn=(y1,…,yn)T, μn is the expectation of the multivariate Gaussian distribution, and the initial μn is set to 0; Step 4.1.2 Construct the covariance matrix Kn: Cov(yn,y1)...Cov(yn,yn)] Step 4.1.3 Define the kernel function: Where l is the hyperparameter of the kernel function; Step 4.1.4 Construct the diagonal matrix R: The elements on the main diagonal of R are (r1 / m1, ..., rn / mn) T .

9. The method according to claim 8, characterized in that The objective function distribution for unsampled points (θ′, y′) is: [Fn;y′]~N([μn;μ * ],[Kn+R,K * ;K * T,K ** ]) in, K * =(Cov(y1,y′),...,Cov(yn,y′)) T K ** =Cov(y′,y′) For a given θ′, the posterior distribution of y′ is: y′|θ′~N(K * T(Kn+R)-1(Fn-μn),K ** -K * T(Kn+R)-1K * ) 10. The method according to claim 9, characterized in that Introduce the approximation of Var[f(θ)]: Define another Gaussian process: log(Var[f(θ)])~GP(·,·) For a given θ′, the posterior distribution of log(r′) is: log(r′)|θ~N(K * TKn-1Rn,K ** -K * TKn-1K * ) Where Rn=(log(r1),...,log(rn)) T .

11. The method according to claim 1, characterized in that The final implementation process includes: Each time the sampling point is determined iteratively: Select the next sampling point θnew according to strategy 1 or strategy 2; repeat the simulation of θnew for an adaptive number of times; calculate the objective function value and update the agent model; Stopping criteria: The preset maximum number of iterations is reached; the improvement in multiple consecutive iterations is less than the threshold; the predetermined accuracy requirements are met.

12. The method according to claim 11, characterized in that The specific steps of repeating the simulation for the adaptive number of times of θnew are: Step 11.1 Input parameters: Predefined significance level αcE, initial number of simulations minit, additional number of simulations madd; Step 11.2 Perform the initial simulation: When the input is the simulation parameter θ, minit initial simulations are performed; Step 11.3 Estimate the distribution and check the conditions: Estimate the distribution of the expected E[f(θ)] and variance Var[f(θ)] corresponding to the simulation parameter θ, and check whether the condition Var[f(θ)]<σc is satisfied 2 ; Step 11.4 Check the quality judgment conditions: Check whether p(E[f(θ)] <f * )≥1-αcE or p(E[f(θ)]>f * )≥1-αcE where f * represents the best result obtained previously; Step 11.5 Perform additional simulations: If none of the conditions are met, perform madd additional simulations for the sampling point and return to step 11.

3.

13. The method according to claim 1, characterized in that The method is applied to the following scenarios: Calibration of highway traffic flow simulation parameters; Modeling vehicle behavior considering different driving styles; Simulation systems that need to consider the effects of heteroscedastic noise; Large-scale simulation optimization problems with limited computing resources.

Citation Information

Cited By

  • Parameter optimization method, system and equipment for time-dependent diffusion magnetic resonance imaging

    CN120779309A

  • Method, system and apparatus for parameter optimization of time-dependent diffusion magnetic resonance imaging

    CN120779309B