Expressway traffic state division method oriented to unbalanced data
By enhancing and optimizing the cluster centers of highway traffic data, and combining particle swarm optimization and genetic algorithms, the limitations of fixed threshold division and data imbalance problems were solved, achieving more accurate and stable traffic state division.
Patent Information
- Application Number
- CN202511163051.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-07
AI Technical Summary
Existing methods for classifying highway traffic conditions rely on fixed thresholds, making it difficult to adapt to traffic conditions in different regions and road types. Furthermore, clustering models lack robustness and accuracy when faced with imbalanced data.
By enhancing traffic data, the initial cluster centers are optimized using particle swarm optimization and genetic algorithms. Combined with fuzzy C-means clustering, a fitness function is constructed by introducing traffic state jump rate and upstream-downstream consistency to perform dynamic partitioning.
It improves the accuracy and stability of traffic state classification, adapts to traffic operation states in different regions and road types, and enhances the global search capability and reliability of the clustering model in complex environments.
Smart Images

Figure CN120913399A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a highway traffic state division method for unbalanced data, and belongs to the field of highway traffic flow analysis. BACKGROUND
[0002] Traffic state refers to the road traffic operation in a certain time and space range, and is used to reflect the traffic efficiency and safety, and is usually classified and described by states such as smooth and congestion. The accurate division and discrimination of traffic state are of great significance for travelers to reasonably plan their travel and for traffic management departments to implement corresponding control measures.
[0003] Although there are many traffic state division standards, the division method is generally a static classification based on fixed thresholds. This method has certain limitations in practical application. On the one hand, the threshold division method is too absolute and cannot accurately reflect the traffic operation in the critical state. On the other hand, traffic operation state is affected by many factors such as road structure and weather state, and the traffic flow characteristics of different regions are significantly different, so it is difficult for fixed threshold division to accurately reflect the actual traffic state.
[0004] In addition, when dividing the traffic state, the distribution characteristics of the input sample greatly determine the effect of the clustering model. However, in the actual operation process of the highway, the frequency of occurrence of various traffic states is often significantly unbalanced, and the data proportion of the running state is usually large. This imbalance in the number of samples of different categories will obviously affect the robustness of the clustering model and the rationality of the clustering result. SUMMARY
[0005] In order to solve the problem of how to adapt to the traffic operation state division under different regions and different road types, the present application provides a highway traffic state division method for unbalanced data.
[0006] The highway traffic state division method for unbalanced data provided by the present application comprises the following steps:
[0007] The collected traffic data is analyzed, and the minority class traffic data is enhanced to obtain an enhanced data set;
[0008] The traffic state is dynamically divided: the enhanced data set is taken as the input of the particle swarm optimization algorithm for global search to generate a group of candidate cluster centers, the candidate cluster centers are taken as the initial cluster centers of the fuzzy C-means clustering method, in the iteration process, the candidate cluster centers are updated by using the genetic algorithm, the fitness function value of the genetic algorithm is calculated according to the traffic state jump rate of the current candidate cluster center, the consistency of the upstream and downstream traffic states and the objective function of the fuzzy C-means clustering method, and the final cluster centers are obtained by iteration.
[0009] The final cluster centers are de-buzzed, each data point is assigned to the class with the maximum membership, and the class division result is output.
[0010] As preferred, the fitness function value is:
[0011]
[0012] wherein, is the normalized fuzzy C-means clustering method objective function value, is the normalized traffic state jump rate, is the normalized upstream and downstream traffic state consistency, , and are the weights of , and .
[0013] As preferred,
[0014]
[0015]
[0016]
[0017] wherein, , , , , , , represent the number of samples required to determine the weights of , and , represents the normalized objective function value of the i-th sample, represents the normalized objective function mean value, represents the normalized traffic state jump rate of the i-th sample, represents the normalized traffic state jump rate mean value, represents the normalized traffic state jump rate of the i-th sample, represents the normalized traffic state jump rate mean value.
[0018] As preferred, the method for analyzing the collected traffic data and enhancing the minority class traffic data comprises:
[0019] The collected traffic data is used as the original data, the original data label is set, and the minority class traffic data is determined;
[0020] By setting the target number of each minority class data in the original data through proportion or custom setting, setting the number of k-neighbors, using the SMOTE method to generate synthetic samples;
[0021] Transition samples are generated between adjacent classes in the original data through linear interpolation, and corresponding labels are set;
[0022] The original data, the synthetic samples generated using the SMOTE method, and the transition samples are synthesized to obtain an enhanced data set.
[0023] Preferably, the process of generating synthetic samples using the SMOTE method comprises:
[0024]
[0025] In the formula, x j is a synthetic sample generated by the SMOTE method, x i is a minority class data sample, x nn is a randomly selected sample in the k-neighbors, is a random weight factor between 0 and 1.
[0026] Preferably, the method of generating transition samples through linear interpolation comprises:
[0027]
[0028] In the formula, x t is a generated transition sample, x a and x b are samples in two adjacent classes, respectively, is a random weight factor between 0 and 1.
[0029] Preferably, the method of dynamically dividing traffic states comprises:
[0030] Step 1: Taking the enhanced data set as the input of the particle swarm optimization algorithm, performing global search, and generating a set of candidate cluster centers, taking the candidate cluster centers as the initial cluster centers of the fuzzy C-means clustering method;
[0031] Step 2: According to the current cluster center and the enhanced data set, the fuzzy C-means clustering method is used to obtain the cluster center; in the initial iteration, the current cluster center is the initial cluster center;
[0032] Step 3: Calculate the objective function corresponding to the cluster center obtained by the fuzzy C-means clustering method, calculate the current fitness function value according to the objective function, if the current fitness function value meets the convergence compared with the fitness function value obtained in the previous iteration, output the cluster center, and go to step 5, otherwise, go to step 4;
[0033] Step 4, performing selection, crossover, mutation operations of the genetic algorithm, updating particle velocity, position, and mapping to the current cluster center, and going to Step 2;
[0034] Step 5, performing de-fuzzification processing on the final clustering result, assigning each data point in the enhanced data set to the class with the maximum membership degree, and outputting the class division result.
[0035] Preferably, Step 2 comprises:
[0036] Step 21, randomly generating an n x c membership matrix and performing normalization processing, c being the number of initial cluster centers;
[0037] Step 22, updating the membership in the membership matrix according to the current cluster center;
[0038] Step 23, updating the current cluster center according to the membership;
[0039] Step 24, if the current cluster center meets the convergence compared with the cluster center obtained in the previous iteration, outputting the obtained cluster center, otherwise, going to Step 22.
[0040] Preferably, the objective function of the fuzzy C-means clustering method is:
[0041]
[0042]
[0043] In the formula, n is the total number of samples, c is the number of clusters, x i is the i-th sample, v j is the j-th cluster center, is the membership of the sample x i to the cluster center v j , the membership is between [0, 1], m is the fuzzy weighting index, k represents the feature dimension of the sample, l = 1, 2,..., k, is the Euclidean distance of the i-th data to the j-th cluster center, is the l-th feature dimension of the i-th sample, is the l-th feature dimension of the j-th cluster center.
[0044] The beneficial effects of this invention are as follows: This invention improves upon the method of classifying traffic states based on fixed thresholds by dynamically classifying them according to the distribution characteristics of actual traffic data. This allows for adaptation to traffic state classification in different regions and road types, providing a methodological reference for improving the perception and management of highway traffic states. This invention introduces an oversampling data augmentation method, effectively expanding the number of minority class samples and solving the problem of uneven distribution of states among different categories in actual traffic data. This optimizes the distribution of clustering input data and significantly improves the accuracy and stability of traffic state classification. This invention improves clustering through a global optimization mechanism, optimizing the selection method of initial cluster centers, enhancing the model's global search capability, effectively overcoming the shortcomings of traditional FCM clustering algorithms that are sensitive to initial cluster centers and prone to getting trapped in local optima, and enhancing the reliability of clustering results in complex traffic environments. Attached Figure Description
[0045] Figure 1 This is a flowchart illustrating the highway traffic state classification method of this application;
[0046] Figure 2 This is a scatter plot of the original traffic flow data.
[0047] Figure 3 A scatter plot showing the distribution of data after augmentation;
[0048] Figure 4 The result of traffic condition classification;
[0049] Figure 5 This shows the actual traffic conditions at a testing site during a certain holiday. Detailed Implementation
[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0051] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.
[0052] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but this is not intended to limit the scope of the invention.
[0053] The highway traffic state classification method for imbalanced data in this embodiment includes:
[0054] Step 1, analyze the collected traffic data, and enhance the minority class traffic data to obtain an enhanced data set; to address the adverse effects of data imbalance on clustering results, step 1 enhances the minority class samples to improve the balance and continuity of sample distribution.
[0055] Step 2, dynamically divide the traffic state: use the enhanced data set as the input of the particle swarm optimization algorithm for global search, generate a set of candidate cluster centers, use the candidate cluster centers as the initial cluster centers of the fuzzy C-means clustering method, update the candidate cluster centers using the genetic algorithm in the iteration process, calculate the fitness function value of the genetic algorithm according to the traffic state jump rate of the current candidate cluster center, the consistency of the upstream and downstream traffic states, and the objective function of the fuzzy C-means clustering method, and iterate to obtain the final cluster centers. De-fuzzification is performed on the final cluster centers, each data point is assigned to the class with the maximum membership degree, and the class division result is output.
[0056] Fuzzy C-means clustering (FCM) is a clustering analysis method based on distance measurement, which aims to divide a set of data into multiple categories, so that the data points in the same category have high similarity, and the data points between different categories have large differences.
[0057] Although the fuzzy C-means clustering algorithm has good performance in handling fuzzy boundary problems, it still has certain limitations. The selection of initial cluster centers greatly determines the clustering results, and lacks global search ability, which is easy to fall into local optimal solution. The fuzzy C-means clustering algorithm is sensitive to noise and outliers, which can easily lead to the shift of cluster centers. It needs to be improved by optimization algorithm to improve the selection of initial cluster centers and improve the global search ability of clustering. The present embodiment improves the selection of initial cluster centers by using particle swarm optimization algorithm, and improves the global search ability of clustering by improving the fitness function.
[0058] The fitness function usually uses the objective function of FCM, but this method only considers the distance between data points and cluster centers. To improve the clustering effect and better conform to the actual traffic flow running state, the fitness function is improved by introducing traffic state jump rate (TSR) and upstream and downstream consistency (UDC) to construct the fitness function and evaluate the clustering effect of the current cluster center. The smaller the fitness, the higher the membership degree of the data points to their category, the better the clustering compactness, and the more consistent with the actual situation of traffic flow. Specifically, the present embodiment also gives the objective function of FCM, traffic state jump rate (TSR) and upstream and downstream consistency (UDC) a normalized threshold, and the fitness function is as follows:
[0059]
[0060] Wherein, J represents the objective function of FCM, TSR represents the traffic state jump rate, reflecting the stability of the time dimension traffic state, the smaller the better, UDC represents the consistency of upstream and downstream traffic state, reflecting the unity of the spatial dimension traffic state, the larger the better, the three indexes are the maximum and minimum normalized values. The enhanced sample does not have real time and detection point information, so when calculating the TSR and UDC indexes, only the original data is used, and when calculating the objective function J, all samples are used.
[0061] The calculation formula is as follows:
[0062] The traffic state jump rate of the pth traffic flow detection point:
[0063]
[0064] Wherein represents the traffic state label of the ith time point and the pth detection point, N p represents the total number of valid time points of the pth detection point, I(·) is an indicator function, which is 1 when the bracket condition is met, otherwise it is 0, and it is considered as a jump when the traffic state change amplitude of adjacent time is different by 2 labels or more.
[0065] The total jump rate is obtained by weighted average of the traffic state jump rate of each traffic flow detection point, and the calculation formula is as follows:
[0066]
[0067] Wherein P is the number of traffic flow detection points.
[0068] The consistency of upstream and downstream traffic state:
[0069]
[0070]
[0071] Wherein Q represents the logarithm of upstream and downstream detection points, represents the traffic state label of the jth pair of upstream detection points, represents the traffic state label of the jth pair of downstream detection points, sim is the repetition rate function, and T represents the total number of valid time points.
[0072] , , The weights of each index are respectively, which are determined according to the variance ratio method, because the greater the weight variance, the greater the fluctuation, the higher the discrimination, and the greater the influence on the fitness function, therefore a higher weight should be assigned, and the calculation formula is as follows:
[0073]
[0074]
[0075]
[0076] wherein, , , , , , , respectively represent the number of samples required to determine the weights of , and , represents the normalized objective function value of the i-th sample, represents the normalized objective function mean value, represents the normalized traffic state jump rate of the i-th sample, represents the normalized traffic state jump rate mean value, represents the normalized traffic state jump rate of the i-th sample, represents the normalized traffic state jump rate mean value.
[0077] The embodiment improves the method of dividing the traffic state depending on the fixed threshold, dynamically divides according to the distribution characteristics of the actual traffic data, and can adapt to the traffic operation state division under different regions and different road types, providing a method reference for the traffic state perception and management level improvement of the expressway. To solve the adverse effect of the above-mentioned unbalanced data on the clustering effect, the embodiment introduces the synthetic minority over-sampling method to over-sample the minority class samples, and combines the transition sample interpolation strategy to improve the balance and continuity of the sample distribution. Specifically, the method for enhancing the minority class traffic data by analyzing the collected traffic data includes:
[0078] The collected traffic data is used as the original data, the original data is labeled according to the distribution of the original data, the original data label is set, and the minority class traffic data is determined;
[0079] The target number of each minority class data in the original data is set by proportion or self-definition, the number of k neighbors is set: one of the k nearest neighbors in the feature space is randomly selected, and a synthetic sample is generated using the SMOTE method;
[0080] Transition samples are generated by linear interpolation between adjacent categories in the original data, and the corresponding labels are set;
[0081] The original data, the synthetic samples generated using the SMOTE method, and the transition samples are combined into an enhanced data set.
[0082] In this embodiment, new synthetic samples are generated between the sample and its neighbors by using the SMOTE method, which can effectively increase the number of minority class samples and expand the distribution range of minority class samples, while avoiding the overfitting problem caused by simple replication. The process of generating synthetic samples using the SMOTE method is as follows:
[0083]
[0084] where x j is the synthetic sample generated by the SMOTE method, x i is the minority class data sample, x nn is a randomly selected sample in the k-neighbors, is a random weight factor between 0 and 1.
[0085] To enhance the continuity of the class boundary region, a transition sample synthesis mechanism is introduced to generate new samples by interpolating between adjacent classes, increase the data distribution of the class boundary region, and improve the data enhancement effect. The basic idea is to select a pair of samples from two adjacent classes, interpolate their feature vectors according to a certain weight ratio, generate transition samples between the two classes, supplement the sparse or blank areas between classes, and the interpolation calculation formula is as follows.
[0086]
[0087] where x t is the generated transition sample, x a and x b are the samples in the two adjacent classes, is a random weight factor between 0 and 1.
[0088] The original data, SMOTE new data and transition samples are synthesized and summarized to construct the enhanced complete data set, and the enhanced data is analyzed to verify the balance of the samples.
[0089] The oversampling data enhancement method introduced in this embodiment effectively expands the number of minority class samples, solves the problem of uneven distribution of various states in actual traffic data, optimizes the distribution of clustering input data, and significantly improves the accuracy and stability of traffic state division.
[0090] The fuzzy C-means clustering algorithm selected in this embodiment dynamically divides the traffic state. Fuzzy C-means (FCM) is a clustering analysis method based on distance measurement, which aims to divide a set of data into multiple categories, so that the data points in the same category have high similarity, and the data points between different categories have large differences.
[0091] Although the FCM algorithm has good performance in dealing with fuzzy boundary problems, it still has certain limitations. The selection of the initial clustering center of the algorithm largely determines the clustering result, lacks global search ability, and is prone to fall into a local optimal solution. The FCM algorithm is sensitive to noise and outliers, which can easily lead to the deviation of the clustering center. It needs to be improved by optimization algorithm to improve the selection of the initial clustering center and enhance the global search ability of clustering.
[0092] To improve the accuracy and robustness of the FCM algorithm in traffic state division, an improved FCM clustering model based on GAPSO is constructed in this embodiment. The model combines the advantages of particle swarm optimization algorithm and genetic algorithm, combines the global search ability of genetic algorithm and the rapid convergence of particle swarm algorithm, repeatedly optimizes the clustering center in the whole iteration process, alleviates the defects of FCM algorithm being sensitive to initial value and prone to local optimum, and improves the global optimality and stability of the clustering result. In this framework, FCM is used as the clustering main body, the enhanced data is used as the input, the particle swarm optimization algorithm is used for global search in the solution space, a set of reasonable distribution and strong convergence candidate clustering centers are generated, and a more optimal initial input is provided for FCM. A limited number of FCM iterations are performed on the selected initial position to refine the clustering center. The traffic state jump rate and the consistency of upstream and downstream traffic states are introduced to construct an improved fitness function that integrates the temporal and spatial characteristics of traffic, which is used to evaluate the clustering effect of the current particle. To improve search diversity, genetic algorithm is introduced to disturb the particle swarm, new solutions are generated through selection, crossover and mutation operations, and the particle swarm position and speed are updated to improve the particle swarm in a better direction. After multiple iterations, the globally optimal clustering center and label division result are output. Specifically, in this embodiment, the method for dynamically dividing traffic states includes:
[0093] Step 1: Determine the key parameters of the algorithm, including the number of clusters C, the maximum number of iterations, the fuzzy weight coefficient m, the crossover probability P c and mutation probability P m of genetic algorithm, etc. Normalize the original traffic volume and average speed data to eliminate the influence of dimension difference between variables, and use the min-max normalization method to process the data. The positions and velocities of a number of particles are randomly generated by the particle swarm algorithm, and the position of each particle is used as a set of FCM initial clustering centers.
[0094] Step 2: According to the current clustering center and the enhanced data set, the fuzzy C-means clustering method is used to obtain the clustering center; in the initial iteration, the current clustering center is the initial clustering center;
[0095] Step 3, calculate the objective function corresponding to the clustering center obtained by the fuzzy C-means clustering method, calculate the current fitness function value according to the objective function, if the current fitness function value compared with the fitness function value obtained by the previous iteration meets the convergence, output the clustering center, turn to step 5, otherwise, turn to step 4, the convergence judgment formula is:
[0096]
[0097] F (t) Table the fitness function value at the tth iteration, Take 10 -5 .
[0098] Step 4, execute the selection, crossover and mutation operations of genetic algorithm, update the particle speed and position, generate new particle swarm, increase the diversity of particle swarm, help to jump out of local optimum, and map to the current clustering center, turn to step 2;
[0099] Step 5, de-fuzzification processing is performed on the final clustering result, each data point in the enhanced data set is assigned to the class with the maximum membership degree, and the class division result is output.
[0100] Fuzzy C-means clustering (Fuzzy c-means, FCM) is a clustering analysis method based on distance measurement, which aims to divide a group of data into multiple categories, so that the data points in the same category have high similarity, and the data points between different categories have large difference.
[0101] The objective of FCM algorithm is to minimize the weighted squared error objective function, so that the clustering result is optimal, and the definition of the objective function is as follows:
[0102]
[0103]
[0104] In the formula, n is the total number of samples, c is the number of clustering, x i is the ith sample, v j is the jth clustering center, The membership degree of sample x i to clustering center v j is between [0, 1], the membership degree is between [0, 1], m is the fuzzy weighted index, k represents the feature dimension of the sample, l = 1, 2,..., k, The Euclidean distance of the ith data to the jth clustering center is represented by d The lth feature dimension of the ith sample is represented by x The lth feature dimension of the jth clustering center is represented by v
[0105] Step 21, randomly generate an n x c membership matrix and normalize it, c is the number of initial cluster centers;
[0106] Step 22, update the membership in the membership matrix according to the current cluster center:
[0107]
[0108] Step 23, update the current cluster center according to the membership:
[0109]
[0110] Step 24, if the current cluster center meets the convergence compared with the cluster center obtained in the last iteration, output the obtained cluster center, otherwise, go to step 22. In this embodiment, for each particle corresponding cluster center, several rounds of FCM iteration are performed, the center point and membership matrix are updated, and when the maximum iteration number or convergence condition is met, the output is output. The convergence judgment formula is:
[0111]
[0112] Wherein , represents all cluster centers in the tth round, Take 10 -5 .
[0113] Output the final membership matrix U and cluster center set V, divide each sample point into the class with the maximum membership, realize de-fuzzification, and obtain the class label corresponding to the sample point.
[0114] This embodiment improves clustering through global optimization mechanism, optimizes the selection method of initial cluster center, improves the global search ability of the model, effectively overcomes the defects that the traditional FCM clustering algorithm is sensitive to the initial cluster center and easy to fall into local optimum, and enhances the reliability of the clustering result in complex traffic environment.
[0115] Embodiment:
[0116] 1. Data source and distribution
[0117] The data comes from four continuous traffic flow detection facilities on a four-lane highway, and the total number of effective data is 35596 under five-minute granularity, each data contains traffic volume and average speed in the corresponding period. Among them, the data with average speed above 80km / h accounts for more than 90%, and the number of samples in low speed state is small, and the data distribution is obviously uneven. The data distribution is shown in Figure 2 .
[0118] 2. Over-sampling setting and result
[0119] SMOTE method is used to enhance the traffic data. According to the empirical distribution of traffic flow and average speed and the statistical results of the original data, the preliminary classification is carried out.
[0120] Class 0: speed > 80, class 1: 60 < speed ≤ 80 and traffic volume > 150, class 2: speed ≤ 60 and traffic volume > 200, class 3: speed ≤ 60 and 100 < traffic volume ≤ 200, class 4: speed ≤ 60 and 0 ≤ traffic volume ≤ 100, class 5: unclassified data, and class 6: transition generated samples. In the algorithm, the number of k neighbors is set to 3, which ensures that the newly generated samples are distributed in the local feature space, avoiding the blurring of the class boundary caused by cross-cluster interpolation, thereby improving the representativeness and rationality of the generated samples. Python is used to realize oversampling processing to enhance the balance of classes.
[0121] After sampling processing, a total of 21846 new data are added, and the total amount of data is 57442. The data distribution is shown in Figure 3 As shown in the figure, the low-speed data is significantly increased, effectively improving the imbalance of the distribution of traffic flow and average speed in each class of samples.
[0122] 3. Traffic state division
[0123] (1) Parameter setting
[0124] The related parameters of the GAPSO-FCM model are set as follows:
[0125] In fuzzy C-means clustering, the number of clustering categories is set to 5, the fuzzy coefficient m = 2.0, the maximum number of iterations of each FCM sub-iteration is 5 steps, and the final optimal center point is further iterated for a maximum of 20 steps. The convergence threshold is set to 1 × 10 -5 .
[0126] In the GAPSO optimization part, the population size is set to 30 particles, and the maximum number of iterations is 50 rounds. In the PSO module parameters, the inertia weight ω = 0.72, and the individual cognitive factor and group social factor are c1 = c2 = 1.49. The GA module uses classical genetic operations, the crossover probability is set to 0.7, the mutation probability is 0.1, and the roulette wheel selection mechanism is used to update the position of the new generation of particles. In each iteration, the particle position is refined by FCM and used for fitness evaluation, and the objective function is the clustering performance function in FCM.
[0127] (2) Data normalization
[0128] The minimum and maximum normalization method is used to map the original data to the interval [0, 1], and the formula is as follows:
[0129]
[0130] wherein x is the original data, x min , x max is the minimum and maximum value of the data, is the normalized value.
[0131] (3) Traffic state division result
[0132] The traffic state is divided into 5 categories in the embodiment, corresponding to the smooth, stable, slow, crowded and jammed states respectively. The clustering center, sample quantity and state characteristics of each category are shown in Table 1, and the state division visualization result is shown in Figure 4 .
[0133] Table 1 Traffic state division
[0134]
[0135] (4) Traffic state
[0136] The traffic state data of a detection point in a holiday is selected for drawing, as shown in Figure 5 .
[0137] The traffic state does not jump in the time and space dimensions, and is gradually changed from the smooth, stable to the slow, crowded and jammed states, which conforms to the basic law of traffic flow operation.
[0138] Although the present application is described herein with reference to particular embodiments, it is to be understood that these examples are merely illustrative of the principles and applications of the present application. It is therefore to be understood that numerous modifications can be made to the illustrative embodiments and that other arrangements can be devised without departing from the spirit and scope of the present application as defined by the appended claims. It is to be understood that the features of the dependent claims can be combined with those of the parent application in the alternative. It is also to be understood that features described with respect to one embodiment can be used in other embodiments.
Claims
1. A method for classifying highway traffic states for imbalanced data, characterized in that, The method comprises the following steps: analyzing the collected traffic data and performing enhancement processing on the minority traffic data to obtain an enhanced data set; dynamically dividing the traffic state: taking the enhanced data set as the input of the particle swarm optimization algorithm, performing global search, generating a set of candidate cluster centers, taking the candidate cluster centers as the initial cluster centers of the fuzzy C-means clustering method, in the iteration process, updating the candidate cluster centers by using the genetic algorithm, calculating the fitness function value of the genetic algorithm according to the traffic state jump rate of the current candidate cluster center, the consistency of the upstream and downstream traffic states and the objective function of the fuzzy C-means clustering method, and performing iteration to obtain the final cluster center; de-fuzzification processing is performed on the final cluster center, each data point is assigned to the class with the maximum membership degree, and the class division result is output.
2. The method of classifying highway traffic states for imbalanced data according to claim 1, wherein, The fitness function value is: wherein is the normalized fuzzy C-means clustering method objective function value, is the normalized traffic state jump rate, is the normalized consistency of upstream and downstream traffic states, , and are weights of , and , respectively.
3. The highway traffic state division method for unbalanced data according to claim 2, characterized in that, wherein, , , , , , , respectively represent the number of samples required to determine the weights of , and , represents the normalized objective function value of the i-th sample, represents the normalized objective function mean value, represents the normalized traffic state jump rate of the i-th sample, represents the normalized traffic state jump rate mean value, represents the normalized traffic state jump rate of the i-th sample, represents the normalized traffic state jump rate mean value.
4. The method of claim 1, wherein, the method for analyzing the collected traffic data and performing enhancement processing on the minority traffic data comprises: collecting the traffic data as original data, setting the original data labels, and determining the minority traffic data; setting the target number of each minority data in the original data by proportion or self-definition, setting the number of k neighbors, and generating synthetic samples by using the SMOTE method; transition samples are generated by linear interpolation between adjacent classes in the original data, and corresponding labels are set; the original data, the synthetic samples generated by using the SMOTE method and the transition samples are combined to form an enhanced data set.
5. The method of classifying highway traffic conditions for imbalanced data according to claim 4, wherein, The process of generating synthetic samples by using the SMOTE method comprises: where x j is a synthetic sample generated by the SMOTE method, x i is a minority class data sample, x nn is a randomly selected sample from the k-neighbors, is a random weight factor between 0 and 1.
6. The method of classifying highway traffic conditions for imbalanced data according to claim 4, wherein, The method for generating transition samples by linear interpolation comprises: where x t is the generated transition sample, x a and x b are the samples in the two adjacent classes, respectively, is a random weight factor between 0 and 1.
7. The method of claim 3, wherein the method is characterized by, The method for dynamically dividing the traffic state comprises: Step 1: taking the enhanced data set as the input of the particle swarm optimization algorithm, performing global search, generating a set of candidate cluster centers, and taking the candidate cluster centers as the initial cluster centers of the fuzzy C-means clustering method; Step 2: obtaining the cluster center by using the fuzzy C-means clustering method according to the current cluster center and the enhanced data set; in the initial iteration, the current cluster center is the initial cluster center; Step 3: calculating the objective function corresponding to the cluster center obtained by the fuzzy C-means clustering method, calculating the current fitness function value according to the objective function, if the current fitness function value meets the convergence compared with the fitness function value obtained in the previous iteration, outputting the cluster center and turning to step 5, otherwise, turning to step 4; Step 4: performing selection, crossover and mutation operations of the genetic algorithm, updating the particle speed and position, and mapping to the current cluster center, and turning to step 2; Step 5: de-fuzzification processing is performed on the final cluster result, each data point in the enhanced data set is assigned to the class with the maximum membership degree, and the class division result is output.
8. The method of classifying highway traffic conditions for imbalanced data according to claim 7, wherein, Step 2 comprises: Step 21: randomly generating an nxc membership matrix and performing normalization processing, and c is the number of initial cluster centers; Step 22: updating the membership in the membership matrix according to the current cluster center; Step 23: updating the current cluster center according to the membership. Step 24, if the current clustering center meets the convergence compared with the clustering center obtained in the previous iteration, output the obtained clustering center, otherwise, go to step 22. 9.The method of classifying highway traffic states for unbalanced data according to claim 7, wherein, The objective function of the fuzzy C-means clustering method is: where n is the total number of samples, c is the number of clusters, x i is the ith sample, v j is the jth cluster center, is the membership of the sample x i to the cluster center v j , membership is between [0, 1], m is the fuzzy weighting exponent, k represents the feature dimension of the sample, l = 1, 2,..., k, is the Euclidean distance from the ith data to the jth cluster center, is the lth feature dimension of the ith sample, is the lth feature dimension of the jth cluster center.
10. A device for classifying highway traffic states for imbalanced data, comprising a storage device, a processor, and a computer program stored in the storage device and executable on the processor, wherein, The processor executes the computer program to realize the steps of the expressway traffic state division method for unbalanced data as claimed in any one of claims 1 to 9.