A road traffic flow prediction method in a distributed system
By using Hadoop and Spark to process large-scale historical vehicle data in distributed systems and establishing machine learning regression models, the problem of long running time of traditional stand-alone prediction methods is solved, and efficient and accurate road traffic prediction is achieved.
Patent Information
- Application Number
- CN202111298708.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-04
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-11-04
AI Technical Summary
When traditional single-machine prediction methods process large-scale historical vehicle data, the running time is long, making it difficult to achieve real-time road traffic prediction.
Using a distributed system, a distributed environment is built through Hadoop and Spark, and using MapReduce to process data in parallel, establish a machine learning regression model, and predict road traffic.
Under large-scale data conditions, parallel processing by distributed system significantly shortens the calculation time and improves the accuracy and efficiency of the prediction model.
Smart Images

Figure CN114117892B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of traffic technology, and in particular to a method for predicting road traffic flow in a distributed system. Background Art
[0002] With the continuous development of society, the number of urban motor vehicles continues to increase, and the problem of urban road congestion is serious, which has become an important problem hindering the development of urban travel. Accurate short-term traffic flow prediction provides data support for managers to make timely adjustments to alleviate the traffic pressure in the city. The traditional single-machine prediction method takes a long time to run when the data volume is large. Summary of the invention
[0003] Purpose of the invention: In order to solve the technical problems existing in the background technology, the present invention proposes a road traffic flow prediction method based on a distributed system, the purpose of which is to establish a prediction model to predict the road traffic flow of the day by means of a distributed system parallel processing method under the condition of huge historical data.
[0004] The present invention comprises the following steps:
[0005] Step 1: Build the Hadoop distributed environment, which includes the HDFS distributed file system and the Spark distributed environment of the server;
[0006] Step 2: Obtain all historical vehicle data for a road and predict the real-time vehicle data for the day;
[0007] Step 3: Based on the historical vehicle data in step 2, divide it by hour and use Hadoop to calculate the traffic volume for each hour of the day, which is used as the traffic volume within an hour;
[0008] Step 4: Calculate the vector distance between all historical traffic flows and the predicted real-time traffic flow for the day;
[0009] Step 5: For the data obtained in step 4, select different K values and calculate the deviation and variance of the distance between the first to the Kth vector. When the deviation and variance are minimized, the optimal K value is obtained.
[0010] Step 6: Based on the optimal K value, select the traffic flow data of the remaining days to train the machine learning regression model
[0011] Step 7: According to the optimal K value, select K pieces of historical traffic flow data and input them into each regression model to obtain more than two sets of prediction values for the next hour and the corresponding root mean square error values;
[0012] Step 8: Based on the multiple sets of prediction values and corresponding root mean square error values obtained in step 7, use K nearest neighbor pattern matching to predict the road traffic flow in the next hour of the day.
[0013] In step 2, the historical vehicle data and the real-time vehicle data of the predicted day, the data structure includes the recording time and the vehicle identification number.
[0014] In step 2, the historical vehicle data and the real-time vehicle data for the predicted day are saved to the HDFS distributed file system.
[0015] In step 3, data is read from the HDFS distributed file system, and the read vehicle record data is processed in distributed parallel to obtain the daily and hourly vehicle flow and store it in the HDFS distributed file system. The specific steps include the following:
[0016] Step 3-1: Use <month-day-time, vehicle number> as the input key-value pair;
[0017] Step 3-2: Count the number of vehicles in all data at month-day-hour, and output the key-value pair as <month-day-hour, number of vehicles>;
[0018] Step 3-3: Classify the key-value pairs in step 3-2 according to the time period of each day, count the situation of each day, and output as <month-day, (hour, number of vehicles)>;
[0019] Step 3-4: Integrate the data of the same day and output it as <month-day, {(hour, number of vehicles), (hour, number of vehicles)…}>.
[0020] In step 4, Hadoop's MapReduce process is used for calculation. MapReduce includes the Map phase and the Reduce phase, which specifically include the following steps:
[0021] Step 4-1, read data s for n days n ={x n1 ,x n2 ,…,x n24} and the predicted data for the day q = {y 1 ,y 2 ,…,y k}, where x ni represents the number of people at time i on day n, y i Represents the number of people at time i on that day.
[0022] Step 4-2: Distributed calculation of vector distance The key-value pair <i,s i >As input to the Map stage, <L i ,i> as the output key-value pair of the Map stage;
[0023] Step 4-3, in the Reduce phase, <L i,i> sort in descending order and swap the parameter positions to <i,L i >Output key-value pairs and temporarily store the results in the HDFS distributed file system.
[0024] In step 5, the K value is the coefficient in the K nearest neighbor pattern matching, and the first to Kth distances are selected, that is, M i ,=1,2,…,, and calculate their deviation and variance in a distributed manner. The specific calculation steps are as follows:
[0025] Step 5-1, read the result of step 4-3, distributed computing {M i |i=1,2,…,K}, specifically, first generate key-value pairs <i,(M 1 ,M 2 ,…M i )>, as the input of the Map stage,<i,(B,V)> As the output of the Map stage
[0026] Step 5-2, when the K value varies between 1 and n, there is an optimal K value that minimizes the deviation and variance, and the K value at this time is selected.
[0027] In step 6, the regression models include a linear regression model, a decision tree regression model, a random forest regression model, and a gradient boosting tree regression model on a distributed system. These models are provided by the Machine Learning Library (MLLib) on Spark. The steps of training each regression model include:
[0028] Step 6-1, take the data of the previous k hours as the feature value, and the data of the k+1th hour as the target value;
[0029] Step 6-2, normalize the data of the previous k hours;
[0030] Step 6-3, train each regression model using K-fold cross validation method;
[0031] Step 6-4, calculate the root mean square error value RMSE of each regression model i .
[0032] Step 7 includes: inputting the data of the forecast day into each regression model to obtain more than two sets of forecast values χ in ,i=1,2,3,4,n=1,2,…k,χ in It indicates the predicted value of the nth data corresponding to the i-th regression model. When i is 1, 2, 3, and 4, it corresponds to the linear regression model, decision tree regression model, random forest regression model, and gradient boosting tree regression model, respectively.
[0033] Step 8 includes:
[0034] Step 8-1, using the formula The traffic flow forecast at time k+1 in the historical K days is obtained by calculating the regression model weighted by the root mean square error value, h n Indicates the traffic flow forecast at k+1 on the day corresponding to the nth data;
[0035] Step 8-2, considering the correlation between historical traffic flow and real-time data, the Tanimoto coefficient is introduced to improve the vector distance formula of KNN. The improved vector distance is λ n = n * n , where the Tanimoto coefficient
[0036] Step 8-3, using the formula Calculate the traffic flow prediction result at k+1 on the same day, y is the traffic flow prediction result at k+1 on the same day.
[0037] The present invention has the following advantages and beneficial effects: processing huge data in parallel on a distributed system, distributing the data to different nodes, and making them run on different nodes at the same time, thereby speeding up the calculation speed. In K-nearest neighbor matching, multiple machine learning regression models are integrated, all historical data are effectively used to predict the road traffic flow in the next hour, and the model accuracy is relatively high. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0039] Figure 1 Shown is a diagram of the steps for implementing the method of the present invention.
[0040] Figure 2 Shown is a distributed system structure diagram of the present invention.
[0041] Figure 3 Shown is a specific implementation flow chart of the present invention. DETAILED DESCRIPTION
[0042] Example
[0043] like Figure 1 , Figure 2 , Figure 3 As shown, this embodiment provides a road traffic flow prediction method based on a distributed system, Hadoop and its common component cluster installation, specifically including:
[0044] 1. Use VMware to virtualize 3 servers and build a Hadoop distributed environment.
[0045] 2. Based on the Hadoop distributed environment, build a Spark distributed environment with 3 servers.
[0046] Get the historical vehicle data of a road within a year and the real-time vehicle data of the predicted day in the format of (month-day-hour, vehicle number). The historical vehicle data on October 9 is shown in Table 1:
[0047] Table 1
[0048] time Vehicle number 10-9-11 c5h4431g 10-9-11 B38k96s8 …… …… 10-9-12 Df3154f3 10-9-12 L63m1f9t …… ……
[0049] The real-time vehicle data as of 12:00 on October 21 is shown in Table 2:
[0050] Table 2
[0051] time Vehicle number 10-21-11 P88o9r6t 10-21-11 T1r6ddf3 …… …… 10-21-12 8A9jr6x2 10-21-12 5j6sw9er
[0052] Use the hadoop fs-put command to upload historical vehicle data and predict the real-time vehicle data of the day and save them to the HDFS distributed file system.
[0053] Build the Eclipse development environment under Linux, configure the Hadoop plug-in, and import the org.apache.hadoop.fs package to support opening files, reading and writing files, deleting files, etc.
[0054] Use FileSystem.open(Path f) to open the file from the HDFS distributed file system to read the data, and perform distributed parallel processing on the massive vehicle data read to obtain the traffic flow of each day and hour in the historical data, and calculate the vector distance between all historical traffic flows and the predicted real-time traffic flow of the day. The specific processing is as follows:
[0055] 1. Use <month-day-time, vehicle number> as the input key-value pair.
[0056] 2. In the reduce function, count the number of vehicles in "month-day-time" in all data, and the output key-value pair is <month-day-time, number of vehicles>.
[0057] 3. In the map function, the key-value pairs in the previous step are classified according to the time period of each day, and the situation of each day is counted, and the output is <month-day, (hour, number of vehicles)>.
[0058] Fourth, in the reduce function, the key-value pairs in the previous step are processed, and the data of the same day are integrated together, and the output is <month-day, {(hour, number of vehicles), (hour, number of vehicles)...}>. For example, the data of October 9 is shown in Table 3:
[0059] Table 3
[0060] Hour Number of vehicles 5 362 6 305 7 448 8 627 9 536 …… ……
[0061] 5. Use the key-value pair from the previous step as the input of the Map stage and calculate the vector distance between it and the key-value pair data of the day Take <month-day, vector distance> as the output of the Map stage.
[0062] 6. In the Reduce stage, the above results are sorted in descending order and temporarily stored in the HDFS distributed file system. For example, in Table 4 below, the closest vector distance to the real-time vehicle data as of 12:00 on October 21 is May 26, which is 53; followed by May 12, with a vector distance of 55, and July 13, with a vector distance of 58.
[0063] Table 4
[0064] date Vector distance 5-26 53 5-12 55 7-13 58 …… ……
[0065] In the spark-shell command line environment, use the sc.textFile function to read the <month-day, vector distance> file.
[0066] When the K value varies between 1 and n, the deviation and variance of the first K data are calculated to find the optimal K value that minimizes the deviation and variance. Here, K is taken as 3.
[0067] Use the Machine Learning Library (MLLib) on Spark to build linear regression, decision tree regression, random forest regression, and gradient boosted tree regression models. The training and prediction steps are as follows:
[0068] Step 1: Use spark.createDataFrame to read the card swipe count file for each day and hour in the historical data, and use the data of the first k hours of each day as the feature value and the data of the kth hour as the target value.
[0069] Step 2: Use Normalizer in MLLib to normalize the data.
[0070] Step 3: Perform K-fold cross validation using CrossValidator in MLLib to train each model.
[0071] Step 4: Use the RegressionEvaluator in MLLib to calculate the root mean square error RMSE of each model i .
[0072] Input K historical data into each model, that is, use the model.transform function in MLLib to obtain multiple sets of prediction values χ in ,i=1,2,3,4,n=1,2,…k,i represents various models, and n represents the corresponding K data. Taking the first historical data, that is, the data on May 26, as an example, the prediction results at 13:00 are shown in Table 5 below:
[0073] Table 5
[0074] Model Predicted value RMSE Linear Regression 1324.28 8.9 Decision Tree Regression 1485.63 12.7 Random Forest Regression 1347.71 5.6 Gradient Boosted Trees 1262.29 6.3
[0075] Using the formula Calculate the traffic flow prediction at time k+1 in the historical K days obtained by weighting multiple machine learning regression models according to the root mean square error value, h n Indicates the traffic flow forecast at k+1 on the day corresponding to the nth data. Taking the first historical data, that is, the data on May 26, as an example, h 1 =1337.62.
[0076] Considering the correlation between historical traffic flow and real-time data, the Tanimoto coefficient is introduced to improve the vector distance formula of KNN. The improved vector distance is λ n =T n *M n , where the Tanimoto coefficient
[0077] Using the formula Calculate the traffic flow prediction result at k+1 on the same day, y is the traffic flow prediction result at k+1 on the same day. For example, from the data in Table 6 below, we can get y=1410, that is, the method predicts that the traffic flow at 13:00 on October 21 is 1410.
[0078] Table 6
[0079] date <![CDATA[Vector distance M n > Tanimoto coefficient <![CDATA[Corrected vector distance λ n > <![CDATA[Predicted value h of vehicle flow n > 5-26 53 0.94 49.82 1337.26 5-12 55 0.89 48.95 1467.61 7-13 58 0.83 48.14 1426.90
[0080] The present invention provides a method for predicting road traffic flow in a distributed system. There are many methods and ways to implement the technical solution. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented using existing technologies.
Claims
1. A method for predicting road traffic flow in a distributed system. It is characterized in that The steps include: Step 1: Build the Hadoop distributed environment, which includes the HDFS distributed file system and the Spark distributed environment of the server; Step 2: Obtain all historical vehicle data for a specific road and predict the real-time vehicle data for the day; Step 3: Based on the historical vehicle data, divide it by hour and use Hadoop to calculate the traffic volume for each hour of the day, which is used as the traffic volume within an hour; Step 4: Distributed calculation of the vector distances between all historical traffic flows and the predicted real-time traffic flows for the day, and sorting them in descending order; Step 5: For the data obtained in step 4, select different K values and calculate the deviation and variance of the distance between the first to the Kth vector. When the deviation and variance are minimized, the optimal K value is obtained. Step 6: According to the optimal K value, select the traffic flow data of the remaining days to train each regression model; Step 7: According to the optimal K value, select K pieces of historical traffic flow data and input them into each regression model to obtain more than two sets of prediction values for the next hour and the corresponding root mean square error values; Step 8: Based on the predicted value and RMS error value obtained in step 7, use K nearest neighbor pattern matching to predict the traffic flow in the next hour of the day.
2. The method according to claim 1, It is characterized in that In step 2, the historical traffic flow data and the real-time traffic flow data predicted for the day, the data structure includes recording time and vehicle identification number.
3. The method according to claim 2, It is characterized in that In step 2, the historical traffic flow data and the predicted real-time traffic flow data for the day are saved to the HDFS distributed file system.
4. The method according to claim 3, It is characterized in that In step 3, data is read from the HDFS distributed file system, and the read vehicle record data is processed in distributed parallel to obtain the daily and hourly vehicle flow and store it in the HDFS distributed file system. The specific steps include the following: Step 3-1: Use <month-day-time, vehicle number> as the input key-value pair; Step 3-2: Count the number of vehicles in all data at month-day-hour, and output the key-value pair as <month-day-hour, number of vehicles>; Step 3-3: Classify the key-value pairs in step 3-2 according to the time period of each day, count the situation of each day, and output as <month-day, (hour, number of vehicles)>; Step 3-4: Integrate the data of the same day and output it as <month-day, {(hour, number of vehicles), (hour, number of vehicles)…}>.
5. The method according to claim 4, It is characterized in that In step 4, Hadoop's MapReduce process is used for calculation. MapReduce includes the Map phase and the Reduce phase, which specifically include the following steps: Step 4-1, read n days of data s n ={x n1 , x n2 , ..., x n24 } and the predicted data for the day q = {y 1 ,y 2 , ..., y k }, where x ni represents the number of people at time i on day n, y i Represents the number of people at time i on that day; Step 4-2: Distributed calculation of vector distance The key-value pair <i,s i >As input to the Map stage, <L i , i> as the output key-value pair of the Map stage; Step 4-3, in the Reduce phase, <L i , i> sort in descending order and swap the parameter positions to <i,L i >Output key-value pairs and temporarily store the results in the HDFS distributed file system.
6. The method according to claim 5, It is characterized in that Step 5 includes: Step 5-1, read the result of step 4-3, distributed computing {M i |i=1,2,…,K}, the deviation B and variance V are generated first, and the key-value pairs are generated <i,(M 1 , M 2 , …M i )>, as the input of the Map stage,<i,(B,V)> As the output of the Map stage; Step 5-2, when the K value varies between 1 and n, there is an optimal K value that makes the deviation B and variance V the lowest, and the K value at this time is selected.
7. The method according to claim 6, It is characterized in that In step 6, the regression models include a linear regression model, a decision tree regression model, a random forest regression model, and a gradient boosting tree regression model on a distributed system, and the steps of training each regression model include: Step 6-1, take the data of the previous k hours as the feature value, and the data of the k+1th hour as the target value; Step 6-2, normalize the data of the previous k hours; Step 6-3, train each regression model using K-fold cross validation method; Step 6-4, calculate the root mean square error value RMSE of each regression model i .
8. The method according to claim 7, It is characterized in that Step 7 includes: inputting the data of the forecast day into each regression model to obtain more than two sets of forecast values χ in ,i=1,2,3,4,n=1,2,...k,χ in It indicates the predicted value of the nth data corresponding to the i-th regression model. When i is 1, 2, 3, and 4, it corresponds to the linear regression model, decision tree regression model, random forest regression model, and gradient boosting tree regression model, respectively.
9. The method according to claim 8, It is characterized in that Step 8 includes: Step 8-1, using the formula The traffic flow forecast at time k+1 in the historical K days is obtained by calculating the regression model weighted by the root mean square error value, h n Indicates the traffic flow forecast at k+1 on the day corresponding to the nth data; Step 8-2, considering the correlation between historical traffic flow and real-time data, the Tanimoto coefficient is introduced to modify the vector distance of KNN. The improved vector distance is λ n =T n *M n , where the Tanimoto coefficient Step 8-3, using the formula Calculate the traffic flow prediction result at k+1 on the same day, y is the traffic flow prediction result at k+1 on the same day.
Citation Information
Patent Citations
K neighbor data prediction method based on MapReduce
CN104573331A
Real-time bus passenger flow prediction method based on neighbor regression
CN108415885A