Traffic prediction method based on multilayer K-nearest neighbor, storage medium and computing device

High correlation state vectors are constructed through multi-layer K nearest neighbor algorithm and screened the European-style distance and amplitude change trend nearest neighbor data. Combined with the support vector regression algorithm, the problem of insufficient spatial and temporal feature extraction in urban road networks is solved, and the accuracy of short-term traffic flow prediction is improved.

CN120355044AActive Publication Date: 2025-07-22CHANGAN UNIV

Patent Information

Application Number
CN202510846564.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively extract the spatial and temporal characteristics of urban road networks, resulting in insufficient prediction accuracy of short-term traffic flows.

Method used

A multi-layer K nearest neighbor algorithm is used to construct a high correlation state vector, and the nearest neighbor data is filtered through the Euclidean distance and amplitude change trend, and the support vector regression algorithm is used for prediction.

Benefits of technology

It improves the prediction accuracy of urban traffic data, can effectively deal with the strong nonlinearity and randomness of the data, and improves the accuracy of short-term traffic flow prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355044A_ABST
    Figure CN120355044A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent traffic, and discloses a traffic prediction method based on multi-layer K-nearest neighbor, a storage medium and computing equipment, and the method comprises the steps: S1, collecting traffic flow data; s2, calculating a correlation coefficient of time sequence data between sampling points, selecting # imgabs0 # neighbors by adopting a K nearest neighbor algorithm, and constructing current and historical state vectors; s3, selecting # imgabs 1 # neighbors by adopting a K nearest neighbor algorithm according to the Euclidean distance; s4, calculating a first-order difference value of the traffic flow data of the sampling points, and constructing a current state vector and an amplitude change trend vector of Euclidean distance neighbor data; adopting a K nearest neighbor algorithm to select # imgabs2 neighbors; and S5, predicting the traffic flow by using a support vector regression algorithm. According to the method, the high-correlation state vector is constructed, and dual screening of Euclidean distance and amplitude variation trend is performed on the state vector, so that the characteristics of strong nonlinearity and randomness of urban traffic data can be effectively handled, and the accuracy of a prediction algorithm is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent transportation, and particularly relates to a traffic prediction method based on multi-layer K-nearest neighbor, a storage medium, and a computing device. Background Art

[0002] With the rapid growth of urban population and motor vehicles, road traffic problems such as traffic jams and traffic accidents have become increasingly prominent, and have become a bottleneck restricting urban development. The intelligent transportation system (ITS) is a new traffic technology proposed to improve the traffic environment. Short-term traffic prediction, as an important part of the intelligent transportation system, is an indispensable means in travel planning and traffic control, especially important for urban roads with high congestion, higher spatio-temporal complexity. Reliable and accurate prediction results can improve the operation efficiency of the entire traffic system and reduce the probability of traffic accidents. Generally, it is considered that short-term traffic prediction has a prediction time span of no more than 15 minutes. Therefore, accurately and quickly predicting traffic is the key to alleviating traffic congestion and reducing traffic accidents.

[0003] Currently, the algorithms for traffic prediction can generally be divided into three categories: The first category is the algorithms based on traditional statistical models, mainly including the historical mean method, the Kalman filter method, and time series, etc. The structures of these algorithms are simple, and it is difficult to accurately model for a traffic system with complex nonlinearity and randomness, especially for short-term traffic prediction problems. The second category is the algorithms based on deep learning models, mainly including long short-term memory networks, convolutional neural networks, graph neural networks, etc. These algorithms can accurately model complex traffic scenarios and have high prediction accuracy, but they have high requirements for the quantity and quality of data, and there are also problems such as difficult parameter tuning and slow algorithm convergence speed. Especially for different scenarios, it is often necessary to re-tune parameters and train, increasing the complexity of prediction. The third category is the algorithms based on traditional machine learning models, mainly including the K-nearest neighbor method, the support vector machine method, and the Bayesian network, etc. These algorithms can handle the nonlinearity and randomness in traffic prediction and have a relatively fast calculation speed.

[0004] As a relatively common traffic flow prediction algorithm, K-nearest neighbor non-parametric regression has achieved good prediction results in traffic flow prediction on sections such as urban expressways and highways. However, for the urban road network with a more complex spatio-temporal structure and stronger scene time-variability, the traditional K-nearest neighbor method can only match the state vectors through the Euclidean distance, and it is difficult to effectively extract the spatio-temporal features of the road network. Therefore, how to quickly and effectively extract the spatio-temporal features of the road network to improve the short-term traffic flow prediction accuracy is an urgent problem to be solved. Summary of the Invention

[0005] The purpose of the present invention is to provide a traffic prediction method, a storage medium, and a computing device based on multi-layer K-nearest neighbor, so as to effectively solve the problem that it is difficult to effectively extract spatio-temporal features of road networks in the above-mentioned prior art. Specifically, first, use the multi-layer K-nearest neighbor algorithm to perform spatio-temporal correlation analysis on the traffic data of road segments in the collected road network, and construct a strongly correlated state vector; secondly, select the nearest neighbor data of the Euclidean distance through the Euclidean distance between the current state vector and the historical state vector; on this basis, construct an amplitude change trend state vector through the first-order difference of the nearest neighbor data of the Euclidean distance, and finally select the nearest neighbor data with the most similar Euclidean distance and amplitude change trend. Finally, according to the finally selected nearest neighbor data, obtain the true value of the next moment of the nearest neighbor data and the data to be predicted respectively, construct input and output data, and combine the support vector regression algorithm for prediction to improve the short-term prediction accuracy of urban traffic flow.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a traffic prediction method based on multi-layer K-nearest neighbor, including the following steps: S1. Set 1 sampling point on each road segment in the target area, and collect traffic flow data of each sampling point in real time; S2. According to the traffic flow data of all sampling points obtained in S1, characterize the time series data of different sampling points under different delay periods; calculate the correlation coefficient of the time series data between sampling points under different delay periods, and select nearest neighbors using the K-nearest neighbor algorithm according to the correlation coefficient; construct the current state vector and the historical state vector according to the selected nearest neighbors; S3. Calculate the Euclidean distance between the current state vector and the historical state vector, and select nearest neighbor data of the Euclidean distance using the K-nearest neighbor algorithm according to the Euclidean distance; S4. Calculate the first-order difference value of the traffic flow data of any sampling point b at time , and normalize the positive and negative parts of the difference value respectively; according to the normalization result of the first-order difference value, calculate the corresponding difference value when the cumulative distribution function of the difference value takes values at steps of 10% in the range of 10% to 90%; define the amplitude change trend according to the obtained corresponding difference value; calculate the first-order difference value of the current state vector and the first-order difference value of the nearest neighbor data of the Euclidean distance obtained in S3, and construct the amplitude change trend vectors of the current state vector and the nearest neighbor data of the Euclidean distance obtained in S3; according to the current state vector and the nearest neighbor data of the Euclidean distance obtained in S3, The Euclidean distance of the amplitude change trend vectors of the data with the closest Euclidean distance is used, and finally nearest neighbors are selected using the K-nearest neighbor algorithm; S5. According to the nearest neighbor data finally selected in S4, obtain the first element of the historical state vector of the next moment of this nearest neighbor data in the historical feature database sample; Sort the selected elements in ascending order as the input of the support vector regression algorithm, and use the support vector regression algorithm to predict the sampling point to be predicted a at time for traffic flow prediction.

[0007] In a second aspect, the present invention provides a computer-readable storage medium storing a program, including: The program includes instructions that, when executed by a computing device, cause the computing device to execute the traffic prediction method based on multi-layer K-nearest neighbor of the present invention.

[0008] In a third aspect, the present invention provides a computing device, including: A processor, a memory, and a program, wherein the program is stored in the memory and configured to be executed by the processor, and the program includes instructions for executing the traffic prediction method based on multi-layer K-nearest neighbor of the present invention.

[0009] Compared with the prior art, the present invention has at least the following beneficial effects: (1) The traffic prediction method based on multi-layer K-nearest neighbor of the present invention. By constructing highly correlated state vectors, extracting Euclidean distance and amplitude change trend nearest neighbor data through the multi-layer K-nearest neighbor algorithm, and combining the support vector regression algorithm at the same time, short-term prediction of traffic flow is realized. Compared with the prior art, this method can construct highly correlated state vectors, perform double screening of Euclidean distance and amplitude change trend on the state vectors, effectively cope with the strong non-linearity and randomness of urban traffic data, and improve the accuracy of the prediction algorithm.

[0010] (2) The present invention uses the K-nearest neighbor method to quantify the correlation size between sampling points under different delay periods based on Pearson correlation, so as to construct highly correlated state vectors. Compared with the traditional traffic prediction method based on the K-nearest neighbor algorithm, this method can improve the speed and accuracy of state vector construction.

[0011] (3) The designed nearest neighbor matching mechanism of the present invention uses the Euclidean distance between the current state vector and the historical state vector to screen the Euclidean distance nearest neighbor data. Compared with the prior art, this method can make full use of the spatio-temporal characteristics of the road network and improve the accuracy of nearest neighbor data screening.

[0012] (4)Based on the selected Euclidean distance nearest neighbor data, calculate the first-order difference of the Euclidean distance nearest neighbor data, construct an amplitude change trend state vector, and then use the K-nearest neighbor method to select the nearest neighbor data with the most similar Euclidean distance and amplitude change trend for prediction. Compared with the prior art, this method can perform double screening on the Euclidean distance and amplitude change trend of the state vector, improve the quality of the nearest neighbor data of the K-nearest neighbor method, and improve the accuracy of prediction.

[0013] (5)According to the finally selected nearest neighbor data, combine the support vector regression algorithm to realize short-term traffic flow prediction. This method can effectively address the characteristics of strong nonlinearity and randomness of urban traffic data and improve the accuracy of the prediction algorithm.

[0014] In summary, the traffic prediction method based on multi-layer K-nearest neighbor disclosed in the present invention can effectively address the characteristics of strong nonlinearity and randomness of urban traffic data by constructing highly correlated state vectors and performing double screening on the Euclidean distance and amplitude change trend of the state vectors, thereby improving the accuracy of the prediction algorithm. Brief Description of the Drawings

[0015] Figure 1 It is a schematic diagram of the sampling position; Figure 2 It is a schematic diagram of the actual road network structure of the experimental data of the present invention; Figure 3 It is for the present invention at different nearest neighbor values The prediction error result graph of sampling point 1 in the training set; among them, the left half and the right half are the MAE and MSE results respectively.

[0016] Figure 4 It is for the present invention at different nearest neighbor values and The prediction error result graph of sampling point 1 in the training set. Among them, the left half and the right half are the MAE and MSE results respectively.

[0017] The content of the present invention will be further explained below in conjunction with the drawings and specific embodiments. Specific Embodiments

[0018] The traffic prediction method based on multi-layer K-nearest neighbor of the present invention first quantifies the correlation between sampling points under different delay cycles in the road network by using the K-nearest neighbor algorithm based on the traffic sampling information of the target road network, and constructs a strongly correlated state vector; secondly, uses the K-nearest neighbor algorithm again to select the Euclidean distance nearest neighbor data through the Euclidean distance between the current state vector and the historical state vector; on this basis, by calculating the first-order difference of the Euclidean distance nearest neighbor data, an amplitude change trend state vector is designed, and the K-nearest neighbor algorithm is used to finally select the nearest neighbor data, whose distances and amplitude change trends are the most similar; finally, according to the finally selected nearest neighbor data, the short-term prediction of traffic data is realized by combining the support vector regression algorithm.

[0019] The traffic prediction method based on multi-layer K-nearest neighbor given by the present invention includes the following steps: S1. Set 1 sampling point on each road section in the target area, and collect the traffic flow data of each sampling point in real time. The traffic flow data of all sampling points in the target area is expressed as:

[0020]

[0021] In the formula: —The traffic flow data of the sampling points in the target area; —The traffic flow data of the sampling point i ; m —The number of sampling points in the target area; j —The sampling time; —The traffic flow data of the sampling point i at j time; —The number of sampling times.

[0022] See Figure 1 , the red circles represent road intersections, the black line segments represent road sections, the arrows represent the road driving directions, there is one sampling point on each road section, and there are a total of 13 sampling points.

[0023] S2. According to the traffic flow data of all sampling points obtained in S1, characterize the time series data of different sampling points under different delay cycles; calculate the correlation coefficients of the time series data between sampling points under different delay cycles, and select nearest neighbors by using the K-nearest neighbor algorithm; construct the current state vector and the historical state vector according to the selected nearest neighbors.

[0024] S2 specifically includes the following sub-steps: S201. According to the traffic flow data of the sampling points obtained in S1, represent the time series data of the sampling point to be predicted a in the delay period as:

[0025] In the formula: a — the sampling point to be predicted, ; t — the current moment; τ — the delay period; — the length of the time series. Since the traffic flow data has strong periodicity, the value does not need to be too large. In the present invention, the historical data of the previous week is selected for subsequent calculation of the correlation size; — the traffic flow data of the sampling point to be predicted a at moment; — the time series data of the sampling point to be predicted a in the delay period ; S202. To quantify the correlation size between sampling points under different delay periods, calculate and Pearson correlation coefficient , and the calculation formula is as follows:

[0026] In the formula: b — the sampling point, which can refer to any sampling point in the road network and can be the same as the sampling point to be predicted a ; — the time series data and covariance; — the time series data of the sampling point to be predicted a in the delay period when it is 0 (i.e., the original time series data); — the time series data of the sampling point b in the delay period ; , — 、 Standard deviation of — Correlation coefficient with , representing the degree of correlation between the two. In particular, when , it indicates Correlation coefficient between ; S203. With as the current state vector, as the historical state vector, the absolute value of the Pearson correlation coefficient is the distance between the current state vector and the historical state vector ( The larger the value, the stronger the correlation). When ( is the maximum delay period) and (sampling point b needs to traverse all sampling points in the road network (including the sampling point a to be predicted)), the K-nearest neighbor algorithm is used to select nearest neighbors from the historical state vectors in descending order , , , , , . According to the nearest neighbors, a set of sampling points under different delay periods is obtained:

[0027] In the formula: —Set of time series data of sampling points under different delay periods, and the elements in this set are highly correlated; —Among all historical state vectors, the time series data of sampling point under delay period has the largest correlation with as the i th largest, where ; S204. The state vector for prediction is composed of elements in the set , which can effectively reflect the similarity and change trend of the current state in historical data. Specifically, the current state vector and the historical state vector are expressed as:

[0028]

[0029]

[0030] wherein: — the capacity ratio of the historical feature database, ; — sampling point at the traffic flow data at the moment; — the number of samples in the historical feature database. In the present invention, the first data are selected from S obtained in S1 as the samples of the historical feature database, where is the number of sampling moments.

[0031] Particularly, since the selection in S203 is from large to small, and the sampling point a in the delay period is 0 for the time series data and its own ( ) correlation coefficient , the correlation is the strongest. Therefore, the first element in the current state vector is the variable to be predicted , and the current state vector and the historical state vector are respectively expressed as follows:

[0032]

[0033] Reasonably constructing the state vector is the key to accurate prediction of the K-nearest neighbor algorithm, and S2 provides a necessary basis for the algorithm prediction.

[0034] S3. Calculate the Euclidean distance between the current state vector and the historical state vector , and select Euclidean distance nearest neighbor data according to the Euclidean distance by using the K-nearest neighbor algorithm.

[0035] Specifically, calculate the Euclidean distance between the current state vector and the historical state vector , and arrange all the obtained Euclidean distances from small to large, and select the historical state vectors with the smallest Euclidean distance from the current state vector as the Euclidean distance nearest neighbor data. Calculate the th in thes The corresponding time of the data of the nearest neighbor with Euclidean distance j is , then The data of the nearest neighbor with Euclidean distance is expressed as: .

[0036] Among them, refers to the index of the corresponding time of the s -th data of the nearest neighbor with Euclidean distance among the data of the nearest neighbor with Euclidean distance

[0037] S4. Calculate the first-order difference value of the traffic flow data of the sampling point b at the time , and normalize the positive and negative parts of the difference value respectively; according to the normalization result of the first-order difference value, calculate the corresponding difference value when the cumulative distribution function of the difference value takes values at 10% - 90% in steps of 10%; define the amplitude change trend of according to the obtained corresponding difference value; calculate the first-order difference value of the current state vector and the first-order difference value of the data of the nearest neighbor with Euclidean distance obtained in S3, and construct the amplitude change trend vector of the current state vector and the data of the nearest neighbor with Euclidean distance obtained in S3; according to the Euclidean distance of the amplitude change trend vector of the current state vector and the data of the nearest neighbor with Euclidean distance obtained in S3, finally select nearest neighbors by using the K-nearest neighbor algorithm .

[0038] S4 specifically includes the following sub-steps: S401. When , , calculate the first-order difference value of the traffic flow data of the sampling point b at the time , which is expressed as , and obtain all the difference values; preferably, all the difference values can be normalized to obtain the normalized difference values, specifically, normalize the positive and negative parts of the difference values to between 0 and 1 and between - 1 and 0 respectively; S402. According to all the difference values obtained in S401, calculate the corresponding difference values of the cumulative distribution function of all the difference values when the cumulative distribution function takes 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% , , , , , , , , , , as shown in Table 1, where the cumulative distribution function , x represents the difference value, i ∈ 1~9; Table 1 reflects the cumulative distribution of the amplitude change trend of the state variable.

[0039] Table 1 Cumulative Distribution of Amplitude Change Trend of State Variable

[0040] S403. The corresponding difference value of the cumulative distribution function obtained in step S402 defines the sampling point b at j time traffic flow data amplitude change trend as: ; wherein, —amplitude change trend of the traffic flow data at the sampling point b at j time traffic flow data ; S404. Calculate the first-order difference value of the current state vector and the first-order difference value of the Euclidean distance nearest neighbor data obtained in S3, respectively expressed as: and ; S405. According to the definition in S403 and the first-order difference value obtained in S404, construct the amplitude change trend vectors of the current state vector and the Euclidean distance nearest neighbor data obtained in S3, respectively expressed as:

[0041]

[0042] In the formula: —amplitude change trend vector of the current state vector ; —amplitude change trend vector of the Euclidean distance nearest neighbor data obtained in S3, ; S406. Calculate the amplitude change trend vector of the current state vector obtained in step S405 and amplitude change trend vectors of the Euclidean distance between , and arrange them in ascending order. Finally, select the smallest ( ) as the final nearest neighbor data, and their distance and amplitude change trends are the most similar to the target sampling point. Count the corresponding time g of the j th Euclidean distance nearest neighbor data among the nearest neighbor data as ; then the finally selected ( ) nearest neighbor data are expressed as:

[0043] S5. According to the nearest neighbor data finally selected in S4, obtain the first element of the historical state vector of the next moment of these final nearest neighbor data in the historical feature database sample; after sorting the selected elements in ascending order, use the support vector regression algorithm to predict the traffic flow at the sampling point a at moment.

[0044] S5 specifically includes the following sub-steps: S501. According to the final nearest neighbor data , , obtain the historical state vector of the next moment of these final nearest neighbor data in the historical feature database sample. According to , obtain the first element of the historical state vector of the next moment of the final nearest neighbor data, and form them into a variable

[0045] wherein, — the time corresponding to the historical state vector of the final nearest neighbor data; S502. After sorting the elements in a in ascending order, use the support vector regression algorithm as the input, and then use the support vector regression algorithm to predict the traffic flow at the sampling point at

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0047] The data used in this experiment includes the traffic flow data of 13 road segments (sampling points) connected by four intersections in the local road network of Furong District, Changsha City for 20 days. The actual road network distribution is as Figure 2 shown. The data collection interval is 5 minutes, and 288 pieces of data are collected daily, recording the traffic flow data of 13 sampling points for a total of 20 days from September 17, 2013 to October 6, 2013. Generally, 15 minutes is used as the time interval for short-term prediction. Therefore, in the present invention, the 3 pieces of data collected within every 15 minutes are accumulated to form one piece of data, thereby obtaining the sampling period data of 15 minutes. There are 96 pieces of daily data for each sampling point, and a total of 1920 pieces of data for 20 days.

[0048] In the present invention, the first 1440 pieces of data (1 - 1440) in the first 15 days are used to construct the historical feature database; the first 280 pieces of data (1441 - 1720) in the data of the last 5 days are used as the training set to determine , and the values of and train the support vector regression algorithm; the last 200 pieces (1721 - 1920) are used as the test set to evaluate the prediction results of the algorithm. In addition, the present invention selects the RBF radial basis kernel function as the kernel function of the support vector regression algorithm, and the penalty factor and the kernel parameter are respectively set to 0.5, . The present invention selects sampling point 1 as the sampling point to be predicted, and selects three metric functions, namely the mean absolute error (MAE), the mean absolute percentage error (MAPE), and the mean square error (MSE), as the evaluation indicators of the model performance. The formulas of the three metric functions are respectively defined as:

[0049]

[0050]

[0051] Among them, represents the true value; represents the predicted value; n represents the number of experimental data. The smaller the values of MAE, MSE, and MAPE, the better the prediction effect of the method.

[0052] To determine the value of this invention constructs corresponding state vectors for different values of Figure 3 and selects the nearest neighbor data in the training set using the K-nearest neighbor algorithm, calculates the mean of the nearest neighbor data as the prediction result, and uses it to compare the influence of different values of on the prediction accuracy. Please refer to which shows the prediction error results when the nearest neighbor values are 5, 10, 15, 20, 25, and 30 respectively (since the change trends of the three errors are basically the same, only the performances of MAE and MSE are analyzed). From the results, it can be seen that the prediction error Table 2 Cumulative Distribution of the Amplitude Change Trend of State Variables

[0053] Therefore, the amplitude change trend can be expressed as follows:

[0054] For and the values of Figure 4 please refer to to know the MAE and MSE results of the prediction error of sampling point 1 on the training set when . Among them, the red line represents the error of the multi-layer K-nearest neighbor algorithm in the optimal case when takes different values from 1 to 50; the blue line represents the degraded multi-layer K-nearest neighbor, that is, the error of the multi-layer K-nearest neighbor algorithm that loses the amplitude change trend matching when . When is very small, the amplitude change trend of the Euclidean distance nearest neighbor data selected is also closest to the current state vector, so the two curves basically coincide; when the amplitude change trend matching comes into play, and at Among the Euclidean distance nearest neighbor data, further select nearest neighbor data that are most similar to the current state vector change trend for prediction. Therefore, the subsequent multi-layer K-nearest neighbor errors are all smaller than the degraded multi-layer K-nearest neighbor errors; When is too large, too many irrelevant interference data are selected. Therefore, as increases, both the multi-layer K-nearest neighbor algorithm error and the degraded multi-layer K-nearest neighbor algorithm error show an increasing trend. Starting from MAE and MSE errors increase as increases, and the MAE and MSE errors of the degraded multi-layer K-nearest neighbor and multi-layer K-nearest neighbor algorithms reach the minimum at ( ), respectively. Therefore, determine the degraded multi-layer K-nearest neighbor , multi-layer K-nearest neighbor ( ) for prediction.

[0055] Finally, in the model of the present invention , and take the values of 11, 15, and 13, respectively. Furthermore, verify the prediction performance of the present invention on the test set, and the prediction error results are shown in Table 3.

[0056] Table 3 Prediction Errors of Different Prediction Algorithms at Sampling Point 1

[0057] The above analysis shows that the present invention uses the multi-layer K-nearest neighbor algorithm to extract the spatio-temporal characteristics of the road network, constructs a strongly correlated state vector, extracts the nearest neighbor data with the most similar Euclidean distance and amplitude change trend, and then combines the support vector regression algorithm for prediction, which can generate more accurate prediction results. The results show that the MAE = 6.0396, MSE = 61.4873, and MAPE = 0.0884 of the prediction algorithm proposed by the present invention are the best among the three indicators compared with the prediction algorithm based on the traditional K-nearest neighbor.

[0058] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0059] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or multiple blocks.

[0060] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or multiple blocks.

[0061] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or multiple blocks.

[0062] The above content is only to illustrate the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the claims of the present invention.

Claims

1. A traffic prediction method based on multi-layer K-nearest neighbors, characterized in that It includes the following steps: S1. Set 1 sampling point at each road section in the target area and collect the traffic flow data of each sampling point in real time; S2. Characterize the time - series data of different sampling points under different delay periods based on the traffic - flow data of all sampling points obtained in S1; calculate the correlation coefficients of the time - series data between sampling points under different delay periods, and select nearest neighbors according to the correlation coefficients using the K - nearest - neighbor algorithm; construct the current state vector and the historical state vector based on the selected nearest neighbors; S3. Calculate the Euclidean distance between the current state vector and the historical state vector, and select nearest neighbor data with Euclidean distance using the K-nearest neighbor algorithm; S4. Calculate any sampling point b At the moment Traffic flow data Calculate the first-order difference value, and normalize the positive and negative parts of the difference value respectively; According to the normalization result of the first-order difference value, calculate the corresponding difference values when the cumulative distribution function of the difference value takes values at intervals of 10% in the range of 10% to 90%; Define according to the obtained corresponding differential value of the amplitude change trend; Calculate the first-order difference value of the current state vector and the first-order difference values of the Euclidean distance nearest neighbor data, and construct the amplitude change trend vectors of the current state vector and the Euclidean distance nearest neighbor data; according to the Euclidean distances of the amplitude change trend vectors of the current state vector and the Euclidean distance nearest neighbor data, finally select nearest neighbors using the K-nearest neighbor algorithm; S5. Based on the final selection in S4, the nearest neighbor data, obtain the first element of the historical state vector of the next moment of these final nearest neighbor data from the historical feature database samples; sort the selected elements in ascending order and use them as the input of the support vector regression algorithm, and use the support vector regression algorithm to predict the traffic flow at the sampling point a at time.

2. The traffic prediction method based on multi-layer K-nearest neighbors according to claim 1, wherein In S1, the traffic flow data of all sampling points in the target area is expressed as: In the formula: — Traffic flow data of sampling points within the target area; — Traffic flow data of sampling points i ; m — The number of sampling points within the target area; j — Sampling moment; — Sampling point i At j Traffic flow data at the moment; — The number of sampling instants.

3. The traffic prediction method based on multi-layer K-nearest neighbor according to claim 2, wherein S2 specifically includes the following sub-steps: S201. Represent the time series data of the sampling point to be predicted at the delay period based on the traffic flow data of the sampling point obtained in S1 as follows: a at the delay period as: In the formula: a — Sampling point to be predicted, ; t — Current moment; τ — Delay period; — the length of the time series; — Sampling point to be predicted a Traffic flow data at time — Sampling point to be predicted a Time series data under the delay period ; S202. To quantify the correlation magnitude between sampling points under different delay periods, calculate and Pearson correlation coefficient : In the formula: b — Sampling point, which can refer to any sampling point in the road network and can be the same as the sampling point to be predicted a Same; — Time series data and covariance of; — Sampling point to be predicted a In the delay period Time series data with a delay of 0; — Sampling point b Time series data during the delay period ; , — , standard deviation of; — and correlation coefficient; when it indicates and correlation coefficient; S203. Using as the current state vector, as the historical state vector, and the absolute value of the Pearson correlation coefficient as the distance between the current state vector and the historical state vector. When and , is the maximum delay period. Using the K-nearest neighbor algorithm, select nearest neighbors from the historical state vectors in descending order , , , , , . According to the nearest neighbors, obtain the sets of sampling points at different delay periods: In the formula: — A set of time series data of sampling points under different delay periods; — Among all historical state vectors, the sampling points The time series data at the delay period has the maximum correlation with for the i th largest, where ; S204. Represent the current state vector and the historical state vector as follows: In the formula: — Proportion of the historical feature database capacity, ; — Sampling point At the traffic flow data at that moment; — The number of samples in the historical feature database; Current state vector The first element in Is the variable to be predicted When the current state vector And the historical state vector Are respectively expressed as follows: 。 4. The traffic prediction method based on multi-layer K-nearest neighbor according to claim 3, characterized in that, S3 specifically includes the following operations: calculating the current state vector and the historical state vector to calculate the Euclidean distance between them, arranging all the obtained Euclidean distances in ascending order, and selecting the historical state vector with the smallest Euclidean distance from the current state vector as the Euclidean distance nearest neighbor data. ​ 5. The traffic prediction method based on multi-layer K-nearest neighbors according to claim 4, wherein , and take values of 11, 15, and 13 respectively.

6. A computer-readable storage medium storing a program, characterized in that, It includes: The program includes instructions, which when executed by a computing device, cause the computing device to execute the traffic prediction method based on multi-layer K-nearest neighbor as described in any one of claims 1 to 5.

7. A computing device, characterized in that, It includes: A processor, a memory and a program, wherein the program is stored in the memory and configured to be executed by the processor, and the program includes instructions for executing the traffic prediction method based on multi-layer K-nearest neighbor as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Urban road vehicle running speed forecasting method based on road network characteristics

    CN104464304A

  • Highway pavement rainfall distribution estimation method, storage medium and computing equipment

    CN112036630A

Cited By

  • Rainfall phase state identification system based on dynamic threshold value

    CN120892753A