Method for identifying abnormal variable-speed vehicle based on ETC (Electronic Toll Collection) data

By constructing a speed difference correction function optimized by the quantitative model of multi-factor vehicle speed difference and genetic algorithm, combined with the GMM Gaussian hybrid model, the identification accuracy of abnormal speed vehicles in ETC data is solved, and the accuracy of traffic management and the operation efficiency of intelligent traffic systems are improved.

CN120279719APending Publication Date: 2025-07-08CHONGQING UNIV +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510339619.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing method of identifying abnormally variable speed vehicles based on ETC data does not fully consider factors such as traffic flow and vehicle proportion, resulting in poor identification accuracy, and lack of unified standards for each detection indicator on the time and space scale, which affects the accuracy of traffic management and strategy formulation.

Method used

By constructing a quantitative model for the impact of vehicle speed difference that takes into account multiple factors, combining genetic algorithms and GMM Gaussian hybrid model, comprehensively considering factors such as traffic flow and vehicle proportions, optimizing the speed difference correction function, and using ETC data to match time and space and identify abnormally variable speed vehicles.

Benefits of technology

It realizes accurate identification of abnormally variable speed vehicles, improves the accuracy of traffic management and the effectiveness of strategy formulation, and improves the operation efficiency and safety of intelligent transportation systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279719A_ABST
    Figure CN120279719A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent traffic, and discloses a method for identifying an abnormal variable-speed vehicle based on ETC data, and the method comprises the steps: S1, carrying out the space-time matching of the ETC data of a highway, and obtaining the space-time matching data of adjacent road sections; s2, based on the space-time matching data, calculating indexes according to periods; s3, constructing a speed difference correction function considering the indexes; s4, on the basis of the assumption that the driving behaviors have consistency, continuously adjusting parameter values of a correction function by using a genetic algorithm and taking the maximum correlation of correction speed differences of adjacent road sections as a target; s5, calculating the change value of the correction speed difference of the adjacent road sections of each vehicle, and calculating the speed difference value between the vehicle in the current period and the average vehicle speed in the period; and S6, constructing a GMM (Gaussian Mixture Model) to obtain a judgment standard for dividing an abnormal variable-speed vehicle and a normal driving vehicle. According to the invention, the abnormal variable-speed vehicle can be accurately identified from the complex ETC data, and reliable support is provided for traffic management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent transportation, and particularly relates to a method for identifying abnormal speed-changing vehicles based on ETC data. The present invention identifies vehicles with abnormal driving behaviors through ETC devices deployed on highways. Background Art

[0002] In the context of the rapid development of intelligent transportation, it is crucial to efficiently and accurately identify abnormal speed-changing vehicles for ensuring road safety and optimizing traffic management. ETC (Electronic Toll Collection) as an advanced toll collection technology is widely used in scenarios such as highways, generating a large amount of ETC data. This data covers rich information such as the time and speed of vehicles passing through ETC gantries, providing strong support for traffic state analysis and vehicle behavior research.

[0003] Although the ETC data resources are rich, the current methods for identifying abnormal vehicles based on these data have obvious defects. On the one hand, the existing identification indicators are relatively single, usually only focusing on vehicle speed changes, failing to fully explore the potential value of ETC data and ignoring many factors that have a significant impact on vehicle driving states, such as traffic flow and the proportion of large vehicles. Traffic flow directly determines the degree of road congestion, which in turn affects the normal driving speed and driving behaviors of vehicles; due to their own characteristics such as size and load, large vehicles will cause different degrees of interference to the overall operation of the traffic flow at different proportions, and will also affect the driving states of other vehicles. However, the existing technologies do not comprehensively consider these factors, resulting in poor accuracy of the identification results and being unable to fully and truly reflect the vehicle driving conditions.

[0004] On the other hand, the detection indicators lack a unified standard in terms of time and space scales. The detection data of different sections and different time periods are difficult to effectively integrate and correlate, and it is easy to deviate when analyzing vehicle driving states, resulting in the inability to accurately identify abnormal speed-changing vehicles. This problem not only affects the accurate judgment of traffic conditions by traffic management departments, but also makes it difficult to formulate targeted and effective traffic management strategies, restricting the efficient operation and development of intelligent transportation systems. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a method for identifying abnormal speed-changing vehicles based on ETC data. The present invention aims to solve the problems that the existing methods do not fully consider influencing factors and the detection indicators lack a unified standard in terms of time and space scales, resulting in poor accuracy in identifying abnormal speed-changing vehicles.

[0006] The method of the present invention deeply explores the potential value of ETC data, takes into account various key factors affecting the driving state of vehicles such as traffic flow and the proportion of large vehicles, and by constructing a quantitative model for the impact of vehicle speed difference considering multiple factors, unifies the spatio-temporal scale of section detection indicators, and realizes the accurate characterization of the vehicle driving environment.

[0007] The present invention provides a method for identifying abnormally variable-speed vehicles based on ETC data, including the following steps:

[0008] S1. Perform spatio-temporal matching on the ETC data of the highway to obtain the spatio-temporal matching data of adjacent sections;

[0009] S2. Based on the spatio-temporal matching data of each section obtained in step S1, calculate indicators in cycles;

[0010] The indicators include the traffic flow saturation of each section, the proportion of large vehicles, the adjacent vehicle speed dispersion, the average driving speed in the cycle, and the maximum traffic flow value of each section;

[0011] S3. Considering the linear and non-linear relationships between factors, construct a speed difference correction function considering the traffic flow saturation of the section, the proportion of large vehicles, the adjacent vehicle speed dispersion, and the average speed in the cycle;

[0012] S4. Based on the assumption that driving behaviors are consistent, use the genetic algorithm to continuously adjust the parameter values of the correction function with the goal of maximizing the correlation of the corrected speed difference between adjacent sections, so as to determine the appropriate parameters;

[0013] S5. Calculate the change value of the corrected speed difference between adjacent sections of each vehicle, and calculate the speed difference value between the vehicle in this cycle and the average vehicle speed in the cycle;

[0014] S6. Construct a GMM Gaussian mixture model with the change value of the corrected speed difference between adjacent sections and the difference between the vehicle in this cycle and the average vehicle speed in the cycle as characteristic parameters, and obtain the judgment criteria for dividing abnormally variable-speed vehicles from normally driving vehicles.

[0015] Further, in step S1, the spatio-temporal matching data of adjacent sections includes: traffic flow, speed, and the proportion of large vehicles.

[0016] Further, step S2 includes the following sub-steps:

[0017] S2.1 Calculate the average driving speed in the cycle

[0018]

[0019] In the formula, v i is the average driving speed of the i-th vehicle on the corresponding section; n is the number of vehicles passing through the end gantry of the corresponding section in the cycle; dab is the length of the corresponding road section; is the time when the i-th vehicle passes through the ETC gantry at the starting point of the corresponding road section; is the time when the i-th vehicle passes through the ETC gantry at the ending point of the corresponding road section;

[0020] S2.2 Calculate the road section flow saturation q;

[0021]

[0022] In the formula, Q is the number of vehicles passing through the gantry at the starting point of the corresponding road section within a 5-minute cycle; Q max is the maximum value of the number of vehicles passing through the gantry at the starting point of the corresponding road section within a 5-minute cycle in historical data, that is, the maximum flow value of the corresponding road section;

[0023] S2.3 Calculate the proportion p of large vehicles;

[0024]

[0025] In the formula, Q T is the number of large vehicles passing through the gantry at the starting point of the corresponding road section within a 5-minute cycle;

[0026] S2.4 Calculate the dispersion σ of the adjacent vehicle speed difference;

[0027]

[0028] In the formula, v i is the average driving speed of the i-th vehicle on the corresponding road section; n is the number of vehicles passing through the gantry at the ending point of the corresponding road section within a 5-minute cycle.

[0029] Furthermore, in the step S3, the speed difference correction function constructed considering the road section flow saturation, the proportion of large vehicles, the adjacent vehicle speed dispersion and the cycle average speed is:

[0030]

[0031] In the formula, β1, β2, β3, β4 are undetermined influence coefficients; Δv i is the difference between the average driving speed v i of the target vehicle within the cycle and the cycle average driving speed .

[0032] Furthermore, the step S4 includes the following sub-steps:

[0033] S4.1 Initialize the population;

[0034] Randomly generate a set of initial solutions as the population, and each individual consists of the undetermined influence coefficients β1, β2, β3, β4 in the speed difference correction function;

[0035] S4.2 Fitness function;

[0036] The Pearson correlation coefficient r is used as the fitness function to measure the correlation of the corrected speed differences between adjacent road segments;

[0037] Assume that the sequences of corrected speed differences for two adjacent road segments are C l ={C 1,l , C 2,l ,......, C n,l} and C l-1 ={C 1,l-1 , C 2,l-1 ,......, C n,l-1}, where n is the number of samples and l is the road segment number. Then the calculation formula for the Pearson correlation coefficient r is:

[0038]

[0039] In the formula, and are the average values of C l and C l-1 respectively; the value range of r is (-1, 1). The closer r is to 1, the higher the degree of correlation between C l and C l-1 ;

[0040] S4.3 Selection operation - Roulette wheel selection method;

[0041] I. Calculate the proportion of the fitness of each individual in the total fitness of the population, that is, the probability that the individual is selected for reproduction. The calculation formula for the probability P j that the individual is selected is:

[0042]

[0043] In the formula, r j is the fitness of individual j; is the total fitness of the population; N is the number of individuals in the population;

[0044] II. Select individuals from the population according to the probability through the roulette wheel method and enter the next-generation population;

[0045] S4.4 Crossover operation;

[0046] Randomly select two individuals from the selected individuals as parents, and exchange some genes of the two parents with a certain crossover probability to generate new offspring individuals;

[0047] The parts of the genes exchanged by the two parents are the partial values of β1, β2, β3, and β4;

[0048] S4.5 Mutation operation;

[0049] Assume the mutation probability P m ;

[0050] The mutation probability P m is used to adjust the mutation degree according to the characteristics of the corresponding road section when performing genetic algorithm operations on the relevant parameters of a certain road section;

[0051] The mutation probability P m is used to determine whether to mutate and the mutation amplitude when performing mutation operations on the undetermined influence coefficients β1, β2, β3, and β4 of the speed difference correction function;

[0052] S4.6 Judge the termination condition;

[0053] Take fitness convergence as the termination condition, and set the convergence accuracy as ∈;

[0054] That is, when the fitness difference corresponding to the road section is less than ∈ in two adjacent iterations, it is considered that the algorithm converges and the iteration stops;

[0055] S4.7 Output the calibration parameters;

[0056] The output calibration parameters are the results optimized with the goal of maximizing the correlation of the corrected speed difference between adjacent road sections after multiple rounds of genetic algorithm iteration.

[0057] Furthermore, the step S5 includes the following sub-steps:

[0058] S5.1 Calculate the change amount ΔC of the corrected speed difference between adjacent road sections of each vehicle i ;

[0059] For vehicle i passing through the downstream gantry of the target detection road section l within cycle t, assume its corrected speed differences in two adjacent road sections l and l-1 are C l and C l-1 , then the change amount ΔC of the corrected speed difference between adjacent road sections i is:

[0060] ΔC i = C l - C l-1

[0061] S5.2 Calculate the speed difference Δv between the vehicle in this cycle and the average vehicle speed in the cycle i ;

[0062]

[0063] In the formula, v i is the speed of vehicle i in this cycle; is the average speed of all vehicles in the cycle.

[0064] Further, the step S6 includes the following sub-steps:

[0065] S6.1 Construct a feature vector;

[0066] Take the change amount ΔC of the corrected speed difference of adjacent road segments i and the difference Δv between the vehicle speed in this cycle and the average cycle speed i as features, and construct a two-dimensional feature vector x for each vehicle i = [ΔC i , Δv i ;

[0067] The feature vectors of all vehicles form a data set X = {x1, x2, ……, x n}), where n is the total number of vehicles;

[0068] S6.2 Initialize the GMM Gaussian mixture model;

[0069] For each Gaussian distribution in the GMM model, its probability density function is:

[0070]

[0071] In the formula, μ k represents the mean of the k-th Gaussian distribution; represents the variance of the k-th Gaussian distribution;

[0072] S6.3 EM algorithm;

[0073] Use the EM algorithm to estimate the parameters to be solved in the GMM model

[0074] S6.4 Calculate the maximum membership probability;

[0075] For the feature vector x corresponding to each vehicle i , calculate the probability P(x i ∈G k ) that it belongs to each Gaussian distribution, and then find the maximum value from them, that is, the maximum membership probability, denoted as:

[0076]

[0077] where k = 1, 2 ……, K; K is the number of Gaussian distributions, and K clustering centers; G k represents the k-th Gaussian distribution in the Gaussian mixture model;

[0078] S6.5 Identify abnormal vehicles;

[0079] According to the results of the Gaussian mixture model, for each data point xi Calculate the probability that it belongs to each Gaussian distribution;

[0080] Set a probability threshold τ to identify abnormal vehicles;

[0081] For vehicle i, if then it is determined that vehicle i is an abnormal variable-speed vehicle.

[0082] Beneficial effects:

[0083] The present invention proposes a method for identifying abnormal variable-speed vehicles based on ETC data. Through the close cooperation of the speed difference correction function, genetic algorithm, and GMM Gaussian distribution fitting technology, it can accurately identify abnormal variable-speed vehicles.

[0084] First of all, the speed difference correction function plays a key role. It comprehensively considers various important factors affecting the driving state of vehicles, such as the degree of traffic flow dispersion and vehicle types. By constructing such a function, the impact of vehicle speed differences under different traffic flow states can be accurately quantified. For example, in sections with dense traffic flow and low dispersion, and in cases where the proportion of large vehicles is relatively high, the meaning represented by the vehicle speed difference is different from the normal situation. The speed difference correction function can quantify these differences to make the calculation of the speed difference more in line with the actual traffic conditions.

[0085] Next, the genetic algorithm is used to optimize the parameters of the speed difference correction function. Based on the reasonable assumption that vehicle driving behaviors are consistent, the genetic algorithm aims to maximize the correlation of the corrected speed difference between adjacent sections and continuously adjusts the parameter values of the correction function. In this process, by steps such as initializing the population, selection operation, crossover operation, and mutation operation, it simulates the natural evolution process to find the optimal parameter combination. The parameters obtained in this way can fully consider the consistency of vehicle driving behaviors, making the speed difference correction function more accurately reflect the speed change characteristics of vehicles in different traffic environments.

[0086] Finally, in combination with the GMM Gaussian distribution fitting technology, a scientific classification of the vehicle driving state is achieved. After being processed by the speed difference correction function and optimized by the genetic algorithm, calculate the difference between the change value of the corrected speed difference between adjacent sections of each vehicle and the vehicle speed and the average vehicle speed within the cycle, and use these differences as key feature data. The GMM Gaussian distribution fitting technology fits the distribution of these feature data through the combination of multiple Gaussian distributions. The data characteristics of normally driving vehicles will be concentrated in certain Gaussian distributions, while the data characteristics of abnormal variable-speed vehicles will deviate from these main distributions. By setting an appropriate probability threshold and comparing the probability that the data points belong to each Gaussian distribution with the threshold value, it is possible to judge from the probability perspective whether the vehicle is an abnormal variable-speed vehicle. For example, when the maximum probability that a vehicle data point belongs to all Gaussian distributions is lower than the set threshold, it can be determined that the vehicle is an abnormal variable-speed vehicle.

[0087] Through the collaborative operation of this series of technologies, the present invention can accurately identify abnormal speed-changing vehicles from complex ETC data, providing reliable support for traffic management.

[0088] Other advantages, objectives, and features of the present invention will, to some extent, be described in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Brief Description of the Drawings

[0089] Figure 1 is a flowchart of a method for identifying abnormal speed-changing vehicles based on ETC data according to the present invention;

[0090] Figure 2 is a display diagram of ETC data. Detailed Description of the Embodiments

[0091] To make the technical solutions, advantages, and objectives of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings of the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of this application.

[0092] As Figure 1 shown, the present invention provides a method for identifying abnormal speed-changing vehicles based on ETC data, including the following steps:

[0093] S1. Perform spatio-temporal matching on highway ETC data to obtain spatio-temporal matching data for adjacent sections;

[0094] Highway ETC data is as Figure 2 shown;

[0095] The spatio-temporal matching data for adjacent sections includes: traffic flow, speed, and large vehicle proportion.

[0096] S2. Based on the spatio-temporal matching data for each section obtained in step S1, calculate indicators by cycle;

[0097] The indicators include the traffic flow saturation degree, large vehicle proportion, adjacent vehicle speed dispersion degree, cycle average driving speed, and maximum traffic flow value for each section;

[0098] S2.1 Calculate the cycle average driving speed

[0099]

[0100] Wherein, v i is the average driving speed of the i-th vehicle on the corresponding road section; n is the number of vehicles passing through the end gantry of the corresponding road section within a cycle; d ab is the length of the corresponding road section; is the time when the i-th vehicle passes through the starting ETC gantry of the corresponding road section; is the time when the i-th vehicle passes through the end ETC gantry of the corresponding road section;

[0101] S2.2 Calculate the road section flow saturation q;

[0102]

[0103] Wherein, Q is the number of vehicles passing through the starting gantry of the corresponding road section within a 5-minute cycle; Q max is the maximum value of the number of vehicles passing through the starting gantry of the corresponding road section within a 5-minute cycle in historical data, that is, the maximum flow value of the corresponding road section;

[0104] S2.3 Calculate the large vehicle proportion p;

[0105]

[0106] Wherein, Q T is the number of large vehicles passing through the starting gantry of the corresponding road section within a 5-minute cycle;

[0107] S2.4 Calculate the dispersion σ of the adjacent vehicle speed difference;

[0108]

[0109] Wherein, v i is the average driving speed of the i-th vehicle on the corresponding road section; n is the number of vehicles passing through the end gantry of the corresponding road section within a 5-minute cycle;

[0110] Within the cycle, according to the passing order of vehicles through the end gantry, its speed set is V = {v1, v2,..., v n}.

[0111] S3. Comprehensively consider the linear and non-linear relationships between factors, and construct a speed difference correction function considering road section flow saturation, large vehicle proportion, adjacent vehicle speed dispersion and cycle average speed;

[0112] The constructed speed difference correction function considering road section flow saturation, large vehicle proportion, adjacent vehicle speed dispersion and cycle average speed is:

[0113]

[0114] Wherein, β1, β2, β3, and β4 are undetermined influence coefficients; Δv i is the average driving speed v of the target vehicle within a cycle i and the average driving speed of the cycle difference.

[0115] S4. Based on the assumption that driving behaviors are consistent, use the genetic algorithm to continuously adjust the parameter values of the correction function with the goal of maximizing the correlation of the corrected speed differences between adjacent road segments, thereby determining the appropriate parameters;

[0116] S4.1 Initialize the population;

[0117] Randomly generate a set of initial solutions as the population. Each individual consists of the undetermined influence coefficients β1, β2, β3, and β4 in the speed difference correction function, representing a possible parameter combination. The population size is usually determined according to the complexity of the problem and computing resources. For example, set the population size to N, and the recommended value is 50 - 200.

[0118] S4.2 Fitness function;

[0119] The goal of the genetic algorithm is to maximize the correlation of the corrected speed differences between adjacent road segments. The Pearson correlation coefficient can be used as the fitness function to measure this correlation.

[0120] Assume that the sequences of corrected speed differences for two adjacent road segments are C l ={C 1,l , C 2,l ,......, C n,l} and C l-1 ={C 1,l-1 , C 2,l-1 ,......, C n,l-1}, where n is the number of samples and l is the road segment number. Then the calculation formula for the Pearson correlation coefficient r is:

[0121]

[0122] In the formula, and are the averages of C l and C l-1 respectively; the value range of r is (-1, 1). The closer r is to 1, the higher the degree of correlation between C l and C l-1 .

[0123] If r = -1, it indicates that C l and C l-1 have a completely negative linear correlation; if r = 1, it indicates that C l and C l-1 have a completely positive linear correlation; if r = 0, Cl and C l-1 do not have a linear correlation.

[0124] In the scenario of this patent, we hope that r is as close to 1 as possible, that is, the correlation of the corrected speed differences between adjacent road segments is the largest. For example, when the r value calculated for a certain two groups of adjacent road segment corrected speed difference sequences is 0.9, it indicates that they have a strong positive correlation, and the corresponding parameter combination is more in line with the optimization goal in the genetic algorithm.

[0125] S4.3 Selection operation - Roulette wheel selection method;

[0126] I. Calculate the proportion of the fitness of each individual in the total fitness of the population. This proportion determines the probability that the individual is selected for reproduction. The probability P j is calculated by the formula:

[0127]

[0128] where r j is the fitness of individual j; is the total fitness of the population; N is the number of individuals in the population;

[0129] II. Through the roulette wheel method, select individuals from the population according to the probability. Individuals with higher fitness have a greater chance of being selected and enter the next generation of the population;

[0130] S4.4 Crossover operation;

[0131] Randomly select two individuals from the selected individuals as the parents, and exchange some of their genes (i.e., some values of β1, β2, β3, β4) with a certain crossover probability (such as P c ) to generate new offspring individuals. For example, for two parent individuals P1 = (β1, β2, β3, β4) and P2 = (β1, β2, β3, β4), perform a crossover operation at a certain crossover point (such as the kth gene position) to generate offspring individuals.

[0132] C1 = (β 1,1 , ……, β k,1 , β k+1,2 , ……, β 4,2 )

[0133] C2 = (β 1,2 , ……, β k,2 , β k+1,1 , ……, β 4,1 )

[0134] S4.5 Mutation operation;

[0135] The mutation operation usually involves a mutation probability and a mutation method. Assume the mutation probability is Pm , where m represents mutation, specifically the mutation probability here, which needs to be distinguished from the probabilities before and after. For example, when performing genetic algorithm operations on the relevant parameters of a certain road section, the mutation probability P m can adjust the degree of mutation according to the characteristics of this road section. In the actual code implementation, when mutating the undetermined influence coefficients β1, β2, β3, β4 of the speed difference correction function, it can be based on P m to determine whether to mutate and the mutation amplitude. Usually, P m can take the value of 0.01, but considering the complexity of traffic flow conditions, it is recommended to take the value of 0.2.

[0136] S4.6 Judge the termination condition;

[0137] Taking fitness convergence as the termination condition, let the convergence accuracy be ∈. That is, when the fitness difference corresponding to the road section is less than ∈ in two adjacent iterations, it is considered that the algorithm converges and the iteration stops. Considering the complex situation of traffic flow, the recommended value of the accuracy is 10 -6 .

[0138] S4.7 Output the calibration parameters;

[0139] When the genetic algorithm meets the set termination condition, it enters the stage of outputting the calibration parameters. In this stage, the algorithm will output the key parameters required to determine the speed difference correction function, that is, the undetermined influence coefficients β1, β2, β3, β4. These parameters are the results optimized through multiple rounds of genetic algorithm iterations with the goal of maximizing the correlation of the corrected speed differences between adjacent road sections.

[0140] In practical applications, the output calibration parameters will be directly used in the calculation of the speed difference correction function. For example, in the scenario of highway ETC data processing, the traffic conditions of different road sections are different. These parameters can accurately correct the speed difference according to the traffic flow saturation q, the proportion p of large vehicles, the dispersion σ of adjacent vehicle speeds, and the average speed of each road section. For example, in a busy road section with a high traffic flow saturation and a large proportion of large vehicles, the parameters obtained through the genetic algorithm will play a role in the speed difference correction function, reasonably adjusting the corrected speed difference, and then more accurately identifying the abnormal speed-changing vehicles on this road section.

[0141] At the same time, the output calibration parameters also have the characteristics of being storable and reusable. The traffic management department or relevant researchers can store these parameters in the database. When it is necessary to process the same road section again later, there is no need to perform complex genetic algorithm calculations again. They can directly call the stored parameters, which greatly improves the efficiency of identifying abnormal speed-changing vehicles. Moreover, these parameters also provide important data support for further studying traffic flow characteristics and optimizing traffic management strategies, helping to improve the operation efficiency and safety of the entire intelligent transportation system.

[0142] S5. Calculate the change value of the corrected speed difference between adjacent road segments for each vehicle, and calculate the speed difference between the vehicle in this cycle and the average speed of the cycle;

[0143] S5.1 Calculate the change amount ΔC of the corrected speed difference between adjacent road segments i ;

[0144] For vehicle i passing through the downstream gantry of the target detection road segment l within cycle t, assume the corrected speed differences between it in two adjacent road segments l (target detection road segment) and l - 1 (previous road segment) are C l and C l-1 , respectively. Then the change amount ΔC of the corrected speed difference between adjacent road segments i is:

[0145] ΔC i = C l - C l-1

[0146] S5.2 Calculate the speed difference Δv between the vehicle in this cycle and the average speed of the cycle i ;

[0147]

[0148] In the formula, v i is the speed of vehicle i in this cycle; is the average speed of all vehicles within the cycle.

[0149] S6. Construct a GMM Gaussian mixture model with the change value of the corrected speed difference between adjacent road segments and the difference between the vehicle in this cycle and the average speed within the cycle as characteristic parameters, and obtain the judgment criterion for dividing abnormal speed-changing vehicles and normal driving vehicles;

[0150] S6.1 Construct a feature vector;

[0151] Take the change amount ΔC of the corrected speed difference between adjacent road segments i and the speed difference Δv between the vehicle in this cycle and the average cycle speed i as features, and construct a two-dimensional feature vector x i = [ΔC i , Δv i for each vehicle;

[0152] The feature vectors of all vehicles form a data set X = {x1, x2, ……, x n}, where n is the total number of vehicles;

[0153] S6.2 Initialize the GMM Gaussian mixture model;

[0154] The GMM model (Gaussian Mixture Model) is a mixture model composed of multiple Gaussian distributions. Each Gaussian distribution represents a subgroup or cluster in the data. The GMM model describes the distribution of the data by estimating the mean, variance, and weight of each Gaussian distribution.

[0155] For a Gaussian model formed by a one-dimensional random variable X, its probability density function is:

[0156]

[0157] In the formula, μ represents the expectation of the data set; σ 2 represents the variance.

[0158] For a two-dimensional random variable X with d = 2, the probability density function of the Gaussian mixture distribution composed of K Gaussians is:

[0159]

[0160] In the formula, is the parameter vector to be estimated; α k represents the coefficient of the k-th distribution and satisfies μ k represents the mean of the k-th Gaussian distribution; represents the variance of the k-th Gaussian distribution.

[0161] For each Gaussian distribution in the GMM model, its probability density function is:

[0162]

[0163] In the formula, μ k represents the mean of the k-th Gaussian distribution; represents the variance of the k-th Gaussian distribution;

[0164] S6.3 EM algorithm;

[0165] To estimate the parameters to be solved in the GMM model usually the EM algorithm (Expectation Maximization) is used. The EM algorithm is an iterative optimization algorithm for estimating the parameters of a probability model with latent variables. Due to its simple iterative operation and the ability to flexibly consider latent variables, the EM algorithm is widely used in machine learning algorithms.

[0166] The core idea of the EM algorithm is to continuously alternate between the E step (Expectation) and the M step (Maximization) until the parameters converge. In the E step, according to the current parameter estimation, the posterior probability of the latent variable is calculated. In the M step, according to the posterior probability of the latent variable obtained in the E step, the new model parameters are obtained by maximizing the likelihood function.

[0167] The specific steps of the algorithm are as follows:

[0168] Step 1: Initialize the GMM model parameters and determine the number k of Gaussian distributions;

[0169] Step 2: In the E step, calculate the probability of the sample x i from the k-th component, that is, the posterior probability of the latent variable: Introduce the latent variable Z = {z1, z2, z3,..., z n}}, and the posterior probability y of the latent variable ik is:

[0170]

[0171] Step 3: In the M step, solve for the new parameters by maximizing the likelihood function based on the posterior probability of the latent variable. And here M is a proprietary term of the algorithm and needs to be distinguished from the m representing mutation in front;

[0172] The expression of the log-likelihood function of the GMM model is:

[0173]

[0174] Take the derivative of μ in the formula k and set it to zero, then μ k can be obtained as:

[0175]

[0176] Take the derivative of in the formula and set it to zero, then can be obtained as:

[0177]

[0178] Since α k ∈(0, 1) and satisfies In addition to maximizing L(X), introduce the Lagrange multiplier λ, then there is:

[0179]

[0180] Take the derivative of α in the formula k and set it to zero, then α k can be obtained as:

[0181]

[0182] Step 4: Update the parameters and repeat steps 2 and 3 until the likelihood function converges. Finally, obtain the parameters of the GMM model.

[0183] In step 4 of the EM algorithm, after multiple rounds of iteration until the likelihood function converges, we obtain the exact parameters of the Gaussian mixture model (GMM), including the mean μ of each Gaussian distribution k , the covariance matrix ∑ k , and the mixing coefficient α k (k = 1, 2 ……, K, where K is the number of Gaussian distributions and the number of K cluster centers). These parameters characterize the distribution of vehicle feature data in the normal driving state.

[0184] At this time, for each two-dimensional feature vector X i in the dataset, which is composed of the change in the corrected speed difference between adjacent road segments of the vehicle and the difference between the vehicle's speed in this cycle and the average speed of the cycle, we can calculate the probability density of it belonging to each Gaussian distribution based on the determined GMM model.

[0185] According to the probability density function formula of the Gaussian distribution:

[0186]

[0187] In the formula, d = 2 (because X i is a two-dimensional feature vector); |∑ k | is the determinant of the covariance matrix ∑ k ; is its inverse matrix.

[0188] Combined with the mixing coefficient α k , we can further obtain the probability that the data point x i belongs to the k-th Gaussian distribution, that is:

[0189]

[0190] S6.4 Calculate the maximum membership probability;

[0191] For the feature vector x i corresponding to each vehicle, calculate the probability P(x i ∈ G k ) that it belongs to each Gaussian distribution, and then find the maximum value among them, that is, the maximum membership probability, denoted as:

[0192]

[0193] where k = 1, 2 ……, K; K is the number of Gaussian distributions and the number of K cluster centers; G k represents the k-th Gaussian distribution in the Gaussian mixture model;

[0194] This maximum value represents the degree of closeness between the vehicle data characteristics and the most matching distribution among all Gaussian distributions. If there are K clustering centers, then for each vehicle, the value with the highest probability is found among the K distributions.

[0195] S6.5 Identification of abnormal vehicles;

[0196] According to the results of the Gaussian mixture model, for each data point x i Calculate the probability that it belongs to each Gaussian distribution. Generally, the data points of abnormal vehicles deviate from the main Gaussian distribution. Therefore, an abnormal vehicle can be identified by setting a probability threshold τ. For example, if it is found that more than 95% of the vehicles are normal vehicles and the probability of belonging to a certain Gaussian distribution is greater than 0.6, then initially τ can be set to 0.6.

[0197] For vehicle i, if Then this vehicle is considered an abnormal variable-speed vehicle. This is because for a vehicle traveling normally, its data characteristics should have a relatively high matching probability with a certain Gaussian distribution in the GMM model; while the data characteristics of an abnormal variable-speed vehicle have a relatively low matching probability with all normal distributions, which means that its driving pattern is significantly different from the normal situation.

[0198] It is hereby declared that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A method for identifying abnormal speed-changing vehicles based on ETC data, characterized in that, It includes the following steps: S1. Perform spatio-temporal matching on highway ETC data to obtain spatio-temporal matching data of adjacent sections; S2. Calculate indicators periodically based on the spatio-temporal matching data of each section obtained in step S1; The indicators include the traffic saturation degree of each section, the proportion of large vehicles, the dispersion degree of adjacent vehicle speeds, the average driving speed of the period, and the maximum traffic volume value of each section; S3. Considering the linear and non-linear relationships between factors comprehensively, construct a speed difference correction function considering the traffic saturation degree of the section, the proportion of large vehicles, the dispersion degree of adjacent vehicle speeds and the average speed of the period; S4. Based on the assumption that driving behaviors are consistent, use the genetic algorithm to continuously adjust the parameter values of the correction function with the goal of maximizing the correlation of the corrected speed difference between adjacent sections, so as to determine the appropriate parameters; S5. Calculate the change value of the corrected speed difference between adjacent sections of each vehicle, and calculate the speed difference between the vehicle in this period and the average vehicle speed of the period; S6. Construct a GMM Gaussian mixture model with the change value of the corrected speed difference between adjacent sections and the difference between the vehicle in this period and the average vehicle speed in the period as characteristic parameters, and obtain the judgment criteria for dividing abnormal speed-changing vehicles and normal driving vehicles.

2. The method for identifying abnormal speed-changing vehicles based on ETC data according to claim 1, wherein: In step S1, the spatio-temporal matching data of adjacent sections includes: traffic volume, speed, and proportion of large vehicles.

3. The method for identifying abnormally variable-speed vehicles based on ETC data according to claim 2, characterized in that: Step S2 includes the following sub-steps: S2.1 Calculate the average driving speed over a period Where, v i is the average driving speed of the i-th vehicle on the corresponding road section; n is the number of vehicles passing through the end gantry of the corresponding road section within a cycle; d ab is the length of the corresponding road section; is the time when the i-th vehicle passes through the starting ETC gantry of the corresponding road section; is the time when the i-th vehicle passes through the end ETC gantry of the corresponding road section; S2.2 Calculate the traffic saturation degree q of the section; Where Q is the number of vehicles passing through the gantry at the starting point of the corresponding road section within a 5-minute period; Q max is the maximum value of the number of vehicles passing through the gantry at the starting point of the corresponding road section within a 5-minute period in the historical data, that is, the maximum flow value of the corresponding road section; S2.3 Calculate the proportion p of large vehicles; Where Q T is the number of large vehicles passing through the gantry at the starting point of the corresponding road section within a 5-minute period; S2.4 Calculate the dispersion degree σ of the adjacent vehicle speed difference; where v i is the average driving speed of the i-th vehicle on the corresponding section; n is the number of vehicles passing through the end gantry of the corresponding section within a 5-minute cycle.

4. The method for identifying abnormal speed-changing vehicles based on ETC data according to claim 3, wherein: In step S3, the constructed speed difference correction function considering the traffic saturation degree of the section, the proportion of large vehicles, the dispersion degree of adjacent vehicle speeds and the average speed of the period is: where β1, β2, β3, β4 are undetermined influence coefficients; Δv i is the average driving speed v of the target vehicle within the cycle i and the average driving speed of the cycle difference.

5. A method for identifying abnormally variable-speed vehicles based on ETC data according to claim 4, characterized in that: Step S4 includes the following sub-steps: S4.1 Initialize the population; Randomly generate a group of initial solutions as the population, and each individual consists of the undetermined influence coefficients β1, β2, β3, β4 in the speed difference correction function; S4.2 Fitness function; Use the Pearson correlation coefficient r as the fitness function to measure the correlation of the corrected speed difference between adjacent sections; Assume that the sequence of corrected speed differences between two adjacent road segments is C l ={C 1,l , C 2,l ,......, C n,l} and C l-1 ={C 1,l-1 , C 2,l-1 ,......, C n,l-1}, where n is the number of samples and l is the road segment number. Then the calculation formula for the Pearson correlation coefficient r is: In the formula, and are the average values of C l and C l-1 respectively; the value range of r is (-1, 1). The closer r is to 1, the higher the correlation degree between C l and C l-1 is; S4.3 Selection operation - Roulette wheel selection method; I. Calculate the proportion of the fitness of each individual in the total fitness of the population, that is, the probability that the individual is selected for reproduction, and the probability \(P\) that the individual is selected j The calculation formula is as follows: where r j is the fitness of individual j; is the total fitness of the population; N is the number of individuals in the population; II. Select individuals from the population according to the probability by means of roulette wheel and enter the next generation population; S4.4 Crossover operation; Randomly select two individuals from the selected individuals as the parents, and exchange part of the genes of the two parents with a certain crossover probability to generate new offspring individuals; The part of the genes exchanged by the two parents is the partial values of β1, β2, β3, β4; S4.5 Mutation operation; Assume mutation probability P m ; The mutation probability P m is used to adjust the degree of mutation according to the characteristics of the corresponding road section when performing genetic algorithm operations on the parameters related to a certain road section; The mutation probability P m is used to determine whether to mutate and the amplitude of mutation when mutating the undetermined influence coefficients β1, β2, β3, and β4 of the speed difference correction function; S4.6 Judge the termination condition; Take fitness convergence as the termination condition, and set the convergence accuracy as ∈; That is, when the fitness difference corresponding to the section is less than ∈ in two adjacent iterations, it is considered that the algorithm converges and the iteration stops; S4.7 Output the calibrated parameters; The output calibrated parameters are the results optimized with the goal of maximizing the correlation of the corrected speed difference between adjacent sections after multiple rounds of genetic algorithm iteration.

6. The method for identifying abnormal speed-changing vehicles based on ETC data according to claim 5, characterized in that: Step S5 includes the following sub-steps: S5.1 Calculate the change value ΔC of the corrected speed difference between adjacent road segments for each vehicle i ; For the downstream gantry of vehicle i passing through the target detection section l within cycle t, let the corrected speed differences of vehicle i at two adjacent sections l and l-1 be C l and C l-1 , respectively. Then the change ΔC i in the corrected speed difference between adjacent sections is: ΔC i = C l - C l-1 S5.2 Calculate the speed difference Δv between the vehicle in this cycle and the average speed of the cycle i ; where v i is the speed of vehicle i in this cycle; is the average speed of all vehicles in the cycle.

7. The method for identifying abnormally variable-speed vehicles based on ETC data according to claim 6, wherein: Step S6 includes the following sub-steps: S6.1 Construct the feature vector; The change amount ΔC of the corrected speed difference between adjacent road segments i and the difference Δv between the vehicle speed in this cycle and the average cycle speed i are used as features to construct a two-dimensional feature vector x for each vehicle i = [ΔC i , Δv i ; The feature vectors of all vehicles form a data set X = {x1, x2, ……, x n}, where n is the total number of vehicles; S6.3 EM algorithm; For each Gaussian distribution in the GMM model, its probability density function is: where μ k represents the mean of the k-th Gaussian distribution; represents the variance of the k-th Gaussian distribution; S6.3 EM algorithm; Use the EM algorithm to estimate the parameters to be solved in the GMM model S6.4 Calculate the maximum membership probability; For the feature vector x corresponding to each vehicle i , calculate the probability P(x i ∈ G k ) that it belongs to each Gaussian distribution, and then find the maximum value among them, that is, the maximum membership probability, denoted as: where k = 1, 2..., K; K is the number of Gaussian distributions and the number of K cluster centers; G k represents the k-th Gaussian distribution in the Gaussian mixture model; S6.5 Identify abnormal vehicles; Based on the results of the Gaussian mixture model, for each data point x i calculate the probability that it belongs to each Gaussian distribution; Set a probability threshold τ to identify abnormal vehicles; For vehicle i, if then it is determined that vehicle i is an abnormal shifting vehicle.