A traffic data analysis method and system
By constructing a dynamic spatio-temporal map and formulating differentiated line planning, the problem of existing systems neglecting regional differences is solved, and accurate prediction of public transportation needs and balanced services are achieved.
Patent Information
- Application Number
- CN202411168738.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-08-23
AI Technical Summary
Existing traffic data analysis methods and systems often ignore the demand differences and service gaps between different regions, resulting in insufficient optimization of the allocation of public transportation resources and unbalanced development of public transportation services, which cannot meet users' public transportation travel needs.
By collecting multi-source data, a comprehensive urban transportation data system is formed, a dynamic spatio-temporal map is constructed to analyze traffic data in different time and space dimensions, predict public transportation travel needs in different regions, and differentiated route plans are formulated based on travel needs and public transportation service levels, and finally recommend public transportation routes to users based on user travel data.
It realizes accurate prediction of public transportation needs in different regions, improves the balance of public transportation services, meets the public transportation travel needs of regional users, and optimizes the allocation of public transportation resources.
Smart Images

Figure CN119181238B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and particularly to a traffic data analysis method and system. Background Art
[0002] In recent years, with the acceleration of the urbanization process and the continuous increase in traffic demand, the optimization of urban traffic management and public transportation systems has become a key research area. Traffic data analysis technology has played an important role in this process. Traditional urban traffic management systems mainly rely on a single data source, such as traffic flow monitors and GPS data, to monitor and manage traffic conditions in real time. However, the limitation of this method lies in the singularity of its data source, which often makes it difficult to comprehensively reflect the complexity and dynamic changes of urban traffic.
[0003] To solve this problem, in recent years, researchers have proposed a method of multi-source data fusion, integrating data from different sources to obtain more comprehensive traffic information. This method can better capture the diversity and complexity of traffic flow and improve the accuracy of traffic prediction. However, although multi-source data fusion technology has improved the comprehensiveness and accuracy of traffic data to a certain extent, existing systems often ignore the demand differences and service gaps between different regions, resulting in sub-optimal allocation of public transportation resources, uneven development of public transportation services in different regions, and inability to well meet the needs of users' public transportation trips. Summary of the Invention
[0004] In view of the problems existing in the above-mentioned existing traffic data analysis methods and systems, the present invention is proposed.
[0005] Therefore, the problem to be solved by the present invention is that existing systems often ignore the demand differences and service gaps between different regions, resulting in sub-optimal allocation of public transportation resources, uneven development of public transportation services in different regions, and inability to well meet the needs of users' public transportation trips.
[0006] To solve the above technical problems, the present invention provides the following technical solution: A traffic data analysis method, which includes collecting multi-source data to form a comprehensive urban traffic data system, constructing a dynamic spatio-temporal graph to analyze traffic data in different time and space dimensions to obtain the public transportation travel demands of different regions; evaluating the public transportation service levels of different regions, and formulating differentiated route plans according to the travel demands and public transportation service levels of different regions; predicting the vehicle traffic density of different regions and recommending public transportation routes for users according to user travel data.
[0007] As a preferred embodiment of the traffic data analysis method of the present invention, the collection of multi-source data to form a comprehensive urban traffic data system means determining the main data sources including traffic monitoring data, GPS data, and social media platform data, collecting data from each data source in real time and preprocessing the data, converting the time and date of all data sources into a unified format and unifying the geographical location information of the data, storing the preprocessed data into a central database to form a comprehensive urban traffic data system, and dividing the city into different regions according to geographical areas.
[0008] As a preferred embodiment of the traffic data analysis method of the present invention, the construction of a dynamic spatio-temporal graph to analyze traffic data in different time and space dimensions to obtain the public transportation travel demand in different regions means defining a dynamic graph structure, extracting population and traffic feature data from the urban traffic data system, selecting traffic intersections and public transportation stations in the city as nodes, and using a GIS system and an urban traffic layout map to locate the geographical location of each node;
[0009] Determining the connection edges between nodes based on the actual traffic routes and public transportation connection lines, integrating the population and traffic feature data into the nodes, strengthening the attributes of the edges using traffic and travel time data, and dynamically updating the weights of the edges according to the real-time traffic flow data;
[0010] Updating the status of the nodes and edges in the graph according to the real-time data obtained from the urban traffic data system;
[0011] Constructing a spatio-temporal tensor network based on the defined nodes and edges, including a spatial tensor graph and a temporal tensor graph:
[0012] The construction of the spatial tensor graph means measuring the actual distance between nodes by collecting GIS data, and determining the spatial relationship between nodes based on the geographical distance and traffic flow connection between nodes;
[0013] Calculating the connection weights between nodes based on the traffic flow and commuting data between nodes, converting the spatial relationship and weights into a multi-dimensional tensor, with each dimension representing a node, and the connection between nodes being represented by the weights of the tensor;
[0014] The construction of the temporal tensor graph means collecting the time series data of each node, including the daily traffic flow and the impact of special events, creating independent layers for each time point, with each layer representing the traffic state within a time period, and defining the connection between layers through time dependence description;
[0015] Optimize the spatio-temporal tensor network using the projected entangled pair state algorithm, convert the data in the spatio-temporal tensor network into a multi-dimensional array form, select the tensor format according to the dimension and structure of the data, set the core parameters of the projected entangled pair state algorithm, including the virtual dimension and the number of optimization iterations, and use simulation data to test different parameter settings to select the best parameter combination;
[0016] Input the multi-dimensional data into the projected entangled pair state algorithm for tensor decomposition processing, simplify the complex multi-dimensional data into an easy-to-manage and analyze format, and use the gradient descent algorithm to regularly optimize the parameters of the tensor network according to the latest spatio-temporal data;
[0017] Design a multi-layer graph convolutional network structure, where each layer is responsible for processing the features of the nodes and integrating the information of neighboring nodes, and configure the network layers through residual connections;
[0018] Use the graph convolution formula to update the feature representation of each node:
[0019]
[0020] In the formula, is the feature representation of node v in the (l + 1)-th layer, is the feature vector of node v in the l-th layer, N(v) is the set of neighbor nodes of node v, W (l) is the learnable weight matrix in the l-th layer, σ is the non-linear activation function ReLU, c vu is the number of neighbors of node v, u is the neighbor node of node v, is the feature representation of node u in the l-th layer;
[0021] For each node, aggregate the features of itself and its neighbors to obtain a new node representation, use historical traffic data for model training and define the cross-entropy loss function, calculate the gradient of the loss function with respect to each weight through the backpropagation algorithm, apply gradient descent to update the weights, and during the training process, calculate the model loss in real time. If the model loss does not decrease in three consecutive epochs, stop the training to obtain the graph neural network model;
[0022] Deploy the graph neural network model in the central database to receive traffic data in real time and predict the travel demand of public transportation in each region.
[0023] As a preferred solution of the traffic data analysis method described in the present invention, wherein: the evaluation of the public transportation service level in different regions refers to obtaining the vehicle load rate, the on-time rate of trips, and the passenger flow of the region from the urban traffic data system, and obtaining the line coverage rate and station accessibility of the public transportation service through GIS and
[0024] Group the collected data according to time periods;
[0025] Analyze the vehicle load rate data to identify problem areas and time periods with overloading and low vehicle utilization;
[0026] Analyze the on-time performance rate data of shifts to evaluate the reliability of public transportation services during peak traffic periods;
[0027] Analyze the passenger flow data to evaluate the satisfaction of service demands for each major route and station;
[0028] Analyze the route coverage rate to evaluate the coverage of public transportation services in different regions;
[0029] Analyze the platform accessibility data to analyze the average walking distance of residents to access the nearest station;
[0030] Integrate the obtained analysis results to identify areas with insufficient public transportation service levels.
[0031] As a preferred solution of the traffic data analysis method described in the present invention, wherein: the formulation of a differentiated route plan according to the travel demands and public transportation service levels in different regions means using the K-means clustering algorithm to classify different regions of the city into high-demand areas and low-demand areas according to traffic demands, and preparing a demand characteristic report for each clustering region, including high-demand time periods, passenger preferences, and service deficiency points;
[0032] Based on the travel demands of public transportation in each region and the public transportation service levels in each region, identify areas with high public transportation travel demands but insufficient public transportation service levels, use GIS tools and data visualization techniques to map the specific locations where the public transportation travel demands do not match the public transportation service levels, and take this region as the primary route optimization target;
[0033] For areas with insufficient public transportation service levels, formulate strategies to increase public transportation resources, including increasing shifts, extending operating hours, adding new routes, and adjusting existing routes;
[0034] For areas with excessive public transportation service levels, reduce the allocation of public transportation resources in this region, optimize existing routes, merge uneconomical public transportation routes, and reallocate resources to regions with higher public transportation travel demands.
[0035] As a preferred solution of the traffic data analysis method described in the present invention, wherein: the prediction of vehicle traffic density in different regions means using a convolutional neural network to construct a traffic density prediction model, including an input layer, a convolutional layer, a pooling layer, and an output layer;
[0036] Use historical traffic flow data as the training set to input into the traffic density prediction model for iterative training. Define the loss function and the Adam optimizer to iteratively optimize the model parameters. When the loss of the demand prediction model no longer significantly decreases during continuous iteration, stop the iteration and output the updated model parameters to update the traffic density prediction model;
[0037] Input the real-time traffic flow data into the traffic density prediction model to obtain the future vehicle traffic density in each region.
[0038] As a preferred solution of the traffic data analysis method of the present invention, wherein: the step of recommending public transportation routes for users according to user travel data refers to collecting the user's historical travel data including common routes, travel time and frequency, and preference settings, preprocessing the collected data, and using the K-means clustering algorithm to classify users into morning peak commuters and weekend travelers according to user data. Use the Apriori algorithm to identify the user's frequent travel combinations, including common starting stations and destination stations, integrate the future vehicle traffic density data, and generate the feature vectors of each planned public transportation route, including route length, vehicle traffic density, and historical user selection frequency;
[0039] Use the random forest model, input the historical user travel data as training data into the random forest model for training, use the cross-validation method to optimize the number of trees and the depth of the trees in the random forest for iteration until the model performance no longer significantly improves, then stop the iteration to obtain the trained random forest model. Input the generated feature vectors into the random forest model to obtain the scores of each planned public transportation route, sort the public transportation routes in descending order of scores, and select the public transportation route with the highest score as the optimal public transportation route and push it to the user in real time.
[0040] Another object of the present invention is to provide a traffic data analysis system, which includes,
[0041] A data collection module for collecting traffic data from multiple data sources, preprocessing it, and storing it in the central database to form a comprehensive urban traffic data system;
[0042] A data analysis module for constructing a dynamic graph structure and a spatio-temporal tensor network based on the real-time traffic data in the urban traffic data system, and predicting the travel demand of public transportation in each region through a graph neural network;
[0043] A horizontal evaluation module for analyzing the traffic data in each region in real time and evaluating the public transportation service level in each region;
[0044] A route planning module for planning the public transportation routes in each region based on the travel demand and service level of public transportation in each region;
[0045] A personalized travel module for recommending public transportation routes to users in real time based on the planned public transportation routes and according to the users' travel data and the vehicle traffic density of each public transportation route.
[0046] A computer device, comprising: a memory and a processor; the memory stores a computer program, and when the processor executes the computer program, the steps of the traffic data analysis method are implemented.
[0047] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the traffic data analysis method are implemented.
[0048] The beneficial effects of the present invention are as follows: By collecting and integrating multi-source data, the present invention constructs a dynamic spatio-temporal graph to analyze traffic data in different time and space dimensions, thereby accurately predicting the public transportation demand in different regions. It can not only capture the complex characteristics of traffic flow more comprehensively, but also formulate differentiated route plans based on the evaluation results of the public transportation service level and recommend the optimal public transportation routes for users based on the predicted traffic density, improving the balance of public transportation services and meeting the public transportation travel needs of regional users. Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 It is a schematic flow chart of the traffic data analysis method.
[0051] Figure 2 It is an implementation schematic diagram of the evaluation of the public transportation service level in different regions.
[0052] Figure 3 It is a schematic structural diagram of the traffic data analysis system. Detailed Embodiments
[0053] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed embodiments of the present invention in conjunction with the drawings of the specification.
[0054] Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0055] Second, the "one embodiment" or "embodiment" referred to herein means a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not all refer to the same embodiment, nor is it an embodiment that is separate or selectively exclusive of other embodiments.
[0056] Embodiment 1
[0057] Referring to Figure 1 and Figure 2 , which is the first embodiment of the present invention. This embodiment provides a traffic data analysis method. The traffic data analysis method includes:
[0058] S1. Collect multi-source data to form a comprehensive urban traffic data system, construct a dynamic spatio-temporal graph, analyze traffic data in different time and space dimensions, and obtain the public transportation travel demands in different regions;
[0059] Specifically, collecting multi-source data to form a comprehensive urban traffic data system means determining the main data sources including traffic monitoring data, GPS data, and social media platform data, collecting data from each data source in real time and preprocessing the data, converting the time and date of all data sources into a unified format and unifying the geographical location information of the data, storing the preprocessed data into a central database to form a comprehensive urban traffic data system, and dividing the city into different regions according to geographical areas.
[0060] By integrating data from traffic monitoring, GPS, and social media, a comprehensive urban traffic data system is constructed. This integration of multi-source data can provide richer and multi-dimensional traffic information, make up for the deficiencies of a single data source, improve the accuracy of traffic condition analysis, store the preprocessed data in a central database for centralized management and maintenance, provide efficient data storage and retrieval capabilities. This centralized management can improve the security and availability of data, support large-scale data analysis and processing, provide fast data retrieval and query capabilities, support real-time data analysis and applications, and divide the city into different regions according to geographical locations, which helps to manage and analyze traffic data more precisely. This regional division can help identify the traffic characteristics and demands of different regions and support differentiated traffic management strategies.
[0061] Furthermore, constructing a dynamic spatio-temporal graph to analyze traffic data in different time and space dimensions and obtain the public transportation travel demands in different regions means defining a dynamic graph structure, extracting population and traffic characteristic data from the urban traffic data system, selecting traffic intersections and public transportation stations in the city as nodes, and using the GIS system and the urban traffic layout map to locate the geographical location of each node;
[0062] Determine the connecting edges between nodes based on the actual traffic routes and public transportation connection lines, integrate population and traffic characteristic data into the nodes, strengthen the attributes of the edges using traffic and travel time data, and dynamically update the weights of the edges according to real-time traffic flow data;
[0063] Update the states of the nodes and edges in the graph according to the real-time data obtained from the urban traffic data system;
[0064] Construct a spatio-temporal tensor network based on the defined nodes and edges, including a spatial tensor graph and a temporal tensor graph:
[0065] The construction of the spatial tensor graph refers to measuring the actual distance between nodes by collecting GIS data, and determining the spatial relationship between nodes based on the geographical distance and traffic flow connection between nodes;
[0066] Calculate the connection weights between nodes based on the traffic flow and commuting data between nodes, convert the spatial relationship and weights into a multi-dimensional tensor, where each dimension represents a node, and the connection between nodes is represented by the weight of the tensor;
[0067] The construction of the temporal tensor graph refers to collecting the time series data of each node, including the daily traffic flow and the impact of special events (holidays, large-scale events), creating independent layers for each time point, where each layer represents the traffic state within a time period, and defining the connection between layers through time dependence description (how the traffic flow in the previous time period affects the next time period);
[0068] Use the projected entangled pair state algorithm to optimize the spatio-temporal tensor network, convert the data in the spatio-temporal tensor network into a multi-dimensional array form, select the tensor format according to the dimension and structure of the data, set the core parameters of the projected entangled pair state algorithm, including the virtual dimension and the number of optimization iterations, and use simulation data to test different parameter settings to select the best parameter combination;
[0069] Input the multi-dimensional data into the projected entangled pair state algorithm for tensor decomposition processing, simplify the complex multi-dimensional data into an easy-to-manage and analyze format, and use the gradient descent algorithm to regularly optimize the parameters of the tensor network according to the latest spatio-temporal data;
[0070] Design a multi-layer graph convolutional network structure, where each layer is responsible for processing the features of the nodes and integrating the information of neighboring nodes, and configure the network layers through residual connections;
[0071] Use the graph convolution formula to update the feature representation of each node:
[0072]
[0073] In the formula, is the feature representation of node v in the l+1 layer, is the feature vector of node v in the l-th layer, N(v) is the set of neighbor nodes of node v, and W (l) is the learnable weight matrix of the l-th layer, σ is the non-linear activation function ReLU, and c vu is the number of neighbors of node v, u is a neighbor node of node v, is the feature representation of node u in the l-th layer;
[0074] For each node, aggregate the features of itself and its neighbors to obtain a new node representation, use historical traffic data for model training and define a cross-entropy loss function, calculate the gradient of the loss function with respect to each weight through the backpropagation algorithm, apply gradient descent to update the weights, during the training process, calculate the model loss in real time, if the model loss does not decrease in three consecutive epochs, then stop training to obtain the graph neural network model;
[0075] Deploy the graph neural network model in the central database to receive traffic data in real time and predict the travel demand of public transportation in each region.
[0076] The dynamic graph structure refers to a data structure in which the attributes of nodes and edges change dynamically over time. The dynamic graph structure is used in this embodiment to describe the real-time changes of the urban traffic network, reflecting the dynamic updates of traffic flow and node characteristics. The spatio-temporal tensor network is a multi-dimensional data structure that combines the time and space dimensions. In this embodiment, the spatio-temporal tensor network is used to describe the complex spatio-temporal relationships of traffic data, including a spatial tensor graph and a time tensor graph. The projected entangled pair state algorithm is an algorithm for processing and optimizing tensor networks;
[0077] By tensor decomposition, complex multi-dimensional data is converted into a more manageable and analyzable form. The dynamic spatio-temporal graph structure can update the states of nodes and edges according to real-time traffic flow data, reflecting the latest traffic conditions. This dynamic update mechanism ensures the real-time nature and accuracy of the traffic model. By integrating population and traffic characteristic data, as well as traffic routes and public transportation connection lines, the dynamic spatio-temporal graph provides a comprehensive perspective for traffic data analysis, enabling better capture of the complexity of the urban traffic network. The spatio-temporal tensor network provides a multi-dimensional perspective on traffic data by integrating data from spatial and temporal dimensions. This integration enables the model to capture the temporal and spatial dependencies of traffic flow, improving the accuracy of traffic prediction. Using the projected entangled pair state algorithm to optimize the spatio-temporal tensor network simplifies complex multi-dimensional data into an easily manageable and analyzable format, improving the efficiency and accuracy of data processing. Through multi-layer convolutional operations and residual connections, the GCN can effectively integrate the features of nodes and their neighbors, enhancing the accuracy of node feature representation. Using the gradient descent algorithm and cross-entropy loss function to optimize model parameters in real-time ensures that the model continuously improves its performance during training. The GCN model can process and predict traffic data in real-time, providing the latest traffic demand prediction results to support traffic management and optimization decisions. By continuously optimizing and updating model parameters, the GCN model can provide highly accurate traffic predictions, improving the operating efficiency of the public transportation system.
[0078] S2. Evaluate the public transportation service levels in different regions, and formulate differentiated route plans according to the travel demands and public transportation service levels in different regions;
[0079] Specifically, evaluating the public transportation service levels in different regions means obtaining the vehicle load rate, schedule punctuality rate, and passenger flow volume of the region from the urban traffic data system, and obtaining the line coverage rate and station accessibility of public transportation services through GIS and
[0080] Group the collected data according to time periods;
[0081] Analyze the vehicle load rate data to identify problem areas and time periods with overloading and low vehicle utilization;
[0082] Analyze the schedule punctuality rate data to evaluate the reliability of public transportation services during peak traffic periods;
[0083] Analyze the passenger flow volume data to evaluate the satisfaction of service demands for major lines and stations;
[0084] Analyze the line coverage rate to evaluate the coverage of public transportation services in different regions;
[0085] Analyze the platform accessibility data to analyze the average walking distance of residents to the nearest stations;
[0086] Integrate the obtained analysis results to identify areas with insufficient public transportation service levels.
[0087] By obtaining vehicle load factor, schedule punctuality rate, and passenger flow data from the urban transportation data system, and combining with the line coverage rate and station accessibility data obtained from the GIS system, the comprehensiveness and accuracy of the data are ensured. This comprehensive data collection method can provide a reliable basis for subsequent analysis. By analyzing the vehicle load factor data, problem areas and time periods with overloading and low vehicle utilization can be identified. Additional transportation capacity is needed in overloaded areas, while resource allocation needs to be optimized in areas with low utilization. After identifying the problem areas, measures can be taken to improve service quality, reduce passenger congestion, and enhance the travel experience of passengers. By analyzing the schedule punctuality rate data, the reliability of public transportation services during peak traffic hours can be evaluated, and problems such as schedule delays and unpunctuality can be identified. For schedules with low punctuality rates, measures can be taken to improve punctuality to ensure the timeliness and reliability of public transportation services. By analyzing the passenger flow data, the satisfaction of service demand for each major line and station can be evaluated, and areas with excessive or insufficient demand can be identified. Based on the analysis results of passenger flow, the lines and schedule frequencies can be adjusted to ensure that public transportation services can meet the actual needs of passengers. Analyze the line coverage rate data to evaluate the coverage of public transportation services in different areas and identify areas with insufficient services. By analyzing the station accessibility data, the average walking distance of residents to access the nearest station can be evaluated, and areas with poor accessibility can be identified. Integrate all analysis results to identify areas with insufficient public transportation service levels, which can help decision-makers formulate improvement measures. By identifying and improving areas with insufficient services, the service level and efficiency of the entire public transportation system can be enhanced.
[0088] Furthermore, formulating a differentiated line plan according to the travel demands and public transportation service levels in different regions means using the K-means clustering algorithm to classify different regions of the city into high-demand areas and low-demand areas based on traffic demands, and preparing a demand characteristic report for each clustering region, including high-demand time periods, passenger preferences, and points of insufficient service;
[0089] Based on the travel demands of public transportation in each region and the public transportation service levels in each region, identify areas with high public transportation travel demands but insufficient public transportation service levels. Use GIS tools and data visualization techniques to map the specific locations where public transportation travel demands do not match public transportation service levels, and take this area as the primary line optimization target;
[0090] For areas with insufficient public transportation service levels, formulate strategies to increase public transportation resources, including increasing schedules, extending operating hours, adding new lines, and adjusting existing lines;
[0091] For areas with an oversupply of public transportation service levels, reduce the allocation of public transportation resources in these areas, optimize existing routes, merge uneconomical public transportation routes, and reallocate resources to areas with higher public transportation travel demands.
[0092] Through the K-means clustering algorithm, different urban areas can be accurately divided into high-demand areas and low-demand areas according to traffic demands. This can better understand the traffic demand characteristics of each area, thereby formulating more effective public transportation planning strategies. The K-means algorithm can process a large amount of data and perform automated analysis, reducing the errors and workload of manual classification and improving work efficiency. By analyzing the public transportation travel demands and service levels of each area, identify areas with high public transportation travel demands but insufficient service levels. These problem areas can be the key targets for optimizing public transportation resources to ensure that limited resources are used where they are most needed. Use GIS tools and data visualization techniques to map the specific locations where public transportation travel demands do not match service levels, providing intuitive and visual decision-making support. For areas with insufficient service levels, increasing the number of trips, extending operating hours, adding new routes, and adjusting existing routes can significantly improve the public transportation service levels in these areas and meet the travel demands of passengers. By increasing resource investment and optimizing route design, improve the coverage and convenience of public transportation services, reduce passengers' waiting times and travel inconveniences. For areas with an oversupply of public transportation service levels, reducing resource allocation can save operating costs, reallocate the saved resources to areas with higher demands, and improve resource utilization efficiency. By optimizing existing routes and merging uneconomical routes, reduce duplicate and inefficient services and improve the operating efficiency of the overall route network.
[0093] S3. Predict the vehicle traffic density in different areas and recommend public transportation routes for users based on user travel data;
[0094] Specifically, predicting the vehicle traffic density in different areas means using a convolutional neural network to construct a traffic density prediction model, including an input layer, a convolutional layer, a pooling layer, and an output layer;
[0095] Use historical traffic flow data as the training set and input it into the traffic density prediction model for iterative training. Define a loss function and an Adam optimizer to iteratively optimize the model parameters. When the loss of the demand prediction model no longer decreases significantly during consecutive iterations, stop the iteration and output the model parameters to update the traffic density prediction model;
[0096] Input the real-time traffic flow data into the traffic density prediction model to obtain the future vehicle traffic density in each area.
[0097] Through the convolutional layer of the Convolutional Neural Network (CNN), the spatio-temporal features in traffic flow data can be effectively extracted, improving the prediction accuracy of the model. The pooling layer reduces the computational complexity of the model through dimensionality reduction, while preventing overfitting problems and improving the generalization ability of the model. By using historical traffic flow data as the training set, the CNN model can learn the historical patterns and rules of traffic flow, improving the accuracy of prediction. By defining the loss function and using the Adam optimizer to iteratively optimize the model parameters, it is ensured that the model can continuously converge to the optimal solution during training, improving the prediction performance. Inputting real-time traffic flow data into the trained CNN model can obtain the predicted results of future vehicle traffic density in each region in real time, providing timely traffic information.
[0098] Furthermore, recommending public transportation routes for users based on user travel data means collecting users' historical travel data including common routes, travel times and frequencies, and preference settings, preprocessing the collected data, and using the K-means clustering algorithm to classify users into morning rush-hour commuters and weekend travelers according to user data. The Apriori algorithm is used to identify users' frequent travel combinations, including common starting and destination stations, integrating future vehicle traffic density data, and generating feature vectors for each planned public transportation route, including route length, vehicle traffic density, and historical user selection frequency.
[0099] Using the random forest model, the historical user travel data is used as training data to input into the random forest model for training. The cross-validation method is used to optimize the number of trees and the depth of the trees in the random forest for iteration until the model performance no longer improves significantly, then stop the iteration to obtain the trained random forest model. Input the generated feature vectors into the random forest model to get the scores of each planned public transportation route. Sort the public transportation routes in descending order of scores, select the public transportation route with the highest score as the optimal public transportation route, and push it to the user in real time.
[0100] The Apriori algorithm is a classic algorithm for frequent itemset mining and association rule learning. In this embodiment, the Apriori algorithm is used to identify users' frequent travel combinations, such as common starting stations and destination stations. By collecting users' historical travel data, including common routes, travel times, frequencies, and preference settings, it is possible to comprehensively understand users' travel habits and needs, providing a basis for subsequent personalized recommendations. Using the K-means clustering algorithm to classify users into morning rush-hour commuters and weekend travelers based on user data can identify the travel patterns of different types of users, enhancing the pertinence and effectiveness of recommendations. Using the Apriori algorithm to identify users' frequent travel combinations can understand the common starting stations and destination stations of users, providing an important reference for route recommendations. Integrating future vehicle traffic density data and generating feature vectors for each planned public transportation route, including route length, vehicle traffic density, and historical user selection frequency, ensures that the recommendation model considers multiple factors. Using historical user travel data as training data to input into a random forest model for training, and optimizing the number and depth of trees through cross-validation, ensures the stable and excellent performance of the model on different datasets. Inputting real-time traffic flow data into a traffic density prediction model to obtain the future vehicle traffic density of each region ensures that the recommendation results can reflect the latest traffic conditions. Inputting the generated feature vectors into the trained random forest model to obtain the scores of each planned public transportation route, sorting the public transportation routes in descending order of scores, and selecting the public transportation route with the highest score as the optimal public transportation route and pushing it to users in real time.
[0101] Embodiment 2
[0102] Referring to Figure 3 , this is the second embodiment of the present invention. This embodiment is different from the previous one and provides a traffic data analysis system, which includes
[0103] A data collection module for collecting traffic data from multiple data sources, preprocessing it, and storing it in a central database to form a comprehensive urban traffic data system;
[0104] A data analysis module for constructing a dynamic graph structure and a spatio-temporal tensor network based on real-time traffic data in the urban traffic data system and predicting the travel demand of public transportation in each region through a graph neural network;
[0105] A horizontal evaluation module for analyzing the traffic data of each region in real time and evaluating the public transportation service level of each region;
[0106] A route planning module for planning the public transportation routes in each region based on the travel demand and service level of public transportation in each region;
[0107] The personalized travel module is used to recommend public transportation routes to users in real time based on the planned public transportation routes and according to the users' travel data and the vehicle traffic density of each public transportation route.
[0108] If the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes.
[0109] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0110] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts (electronic devices) having one or more wirings, portable computer disk cartridges (magnetic devices), random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), optical fiber devices, and portable compact disc read-only memories (CDROMs). Additionally, a computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.
[0111] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
Claims
1. A traffic data analysis method, characterized in that: include, Collect multi-source data to form a comprehensive urban traffic data system, build a dynamic spatiotemporal graph to analyze traffic data in different time and space dimensions to obtain public transportation travel needs in different regions; Evaluate the public transportation service levels in different regions and formulate differentiated route plans based on the travel needs and public transportation service levels in different regions; Predict vehicle traffic density in different areas and recommend public transportation routes for users based on user travel data; the collection of multi-source data to form a comprehensive urban traffic data system refers to determining the main data sources including traffic monitoring data, GPS data and social media platform data, collecting data from various data sources in real time and preprocessing the data, converting the time and date of all data sources into a unified format and unifying the geographical location information of the data, storing the preprocessed data in a central database to form a comprehensive urban traffic data system and dividing the city into different areas according to geographical regions; the construction of a dynamic spatiotemporal graph to analyze traffic data in different time and space dimensions to obtain public transportation travel needs in different areas refers to defining a dynamic graph structure, extracting population and traffic characteristic data from the urban traffic data system, selecting the city's traffic intersections and public transportation stations as nodes, and locating the geographical location of each node using a GIS system and an urban traffic layout map; Determine the connecting edges between nodes based on actual traffic routes and public transportation links, integrate population and traffic characteristic data into nodes, use traffic and travel time data to enhance edge attributes, and dynamically update edge weights based on real-time traffic flow data; Update the states of nodes and edges in the graph based on real-time data obtained from the urban traffic data system; Construct a spatiotemporal tensor network based on the defined nodes and edges, including a spatial tensor graph and a temporal tensor graph: The construction of the spatial tensor graph refers to measuring the actual distance between nodes by collecting GIS data, and determining the spatial relationship between nodes based on the geographical distance and traffic flow connection between nodes; The connection weights between nodes are calculated based on the traffic flow and commuting data between nodes, and the spatial relationship and weight are converted into a multi-dimensional tensor. Each dimension represents a node, and the connection between nodes is represented by the weight of the tensor. The construction of the time tensor graph refers to collecting the time series data of each node, including the daily traffic flow and the impact of special events, creating an independent layer for each time point, each layer represents the traffic status in a time period, and defining the connection between layers through time dependency description; Use the projected entanglement pair algorithm to optimize the space-time tensor network, convert the data in the space-time tensor network into a multidimensional array form, select the tensor format according to the dimension and structure of the data, set the core parameters of the projected entanglement pair algorithm, including the virtual dimension and the number of optimization iterations, and use simulation data to test different parameter settings to select the best parameter combination; Input multidimensional data into the projected entanglement pair algorithm for tensor decomposition processing, simplify the complex multidimensional data into a format that is easy to manage and analyze, and use the gradient descent algorithm to regularly optimize the parameters of the tensor network based on the latest spatiotemporal data; Design a multi-layer graph convolutional network structure, where each layer is responsible for processing node features and integrating information about neighboring nodes, and configures the network layer through residual connections; Update the feature representation of each node using the graph convolution formula: In the formula, is the feature representation of node v at layer l+1, is the feature vector of node v in the lth layer, N(v) is the set of neighbor nodes of node v, and W (l) is the learnable weight matrix of the lth layer, σ is the nonlinear activation function ReLU, c vu is the number of neighbors of node v, u is the neighbor node of node v, is the feature representation of node u at layer l; For each node, aggregate its own and neighbor's features to obtain a new node representation, use historical traffic data to train the model and define the cross entropy loss function, calculate the gradient of the loss function with respect to each weight through the back propagation algorithm, apply gradient descent to update the weight, and calculate the model loss in real time during the training process. If the model loss does not decrease in three consecutive cycles, stop training to obtain the graph neural network model. The graph neural network model is deployed in the central database to receive traffic data in real time and predict the travel demand of public transportation in various regions.
2. The traffic data analysis method according to claim 1, characterized in that: The evaluation of the public transportation service level in different regions refers to obtaining the vehicle load rate, on-time rate and passenger flow of the region from the urban transportation data system, obtaining the line coverage and station accessibility of the public transportation service through GIS, and grouping the collected data according to time periods; Analyze vehicle load factor data to identify problem areas and time periods of overload and low vehicle utilization; Analyze on-time performance data to assess the reliability of public transportation services during peak traffic periods; Analyze passenger flow data to assess whether service demand is being met on major routes and stations; Analyze route coverage to assess the coverage of public transportation services in different areas; Analyze station accessibility data to analyze the average walking distance residents need to visit the nearest station; The obtained analysis results are integrated to identify areas with insufficient public transportation services.
3. The traffic data analysis method according to claim 2, characterized in that: The said formulating differentiated route planning according to the travel demand and public transportation service level of different regions refers to using the K-means clustering algorithm to classify different areas of the city into high-demand areas and low-demand areas according to traffic demand, and compiling a demand characteristic report for each cluster area, including high-demand time periods, passenger preferences and service shortage points; Based on the public transportation travel demand and public transportation service level of each region, identify the areas where public transportation travel demand is high but public transportation service level is insufficient. Use GIS tools and data visualization technology to map the specific locations where public transportation travel demand and public transportation service level do not match, and take this area as the primary route optimization target; For areas with insufficient public transportation services, formulate strategies to increase public transportation resources, including increasing frequency, extending operating hours, adding new routes, and adjusting existing routes; For areas with excess public transportation services, reduce the allocation of public transportation resources in the area, optimize existing routes and merge uneconomical public transportation routes, and reallocate resources to areas with higher public transportation travel demand.
4. The traffic data analysis method according to claim 3, characterized in that: The prediction of vehicle traffic density in different areas refers to building a traffic density prediction model using a convolutional neural network, including an input layer, a convolution layer, a pooling layer, and an output layer; Use historical traffic flow data as a training set to input the traffic density prediction model for iterative training, define the loss function and Adam optimizer to iteratively optimize the model parameters, and stop iterating when the loss of the demand prediction model no longer decreases significantly during continuous iterations, output the model parameters and update the traffic density prediction model; The real-time traffic flow data is input into the traffic density prediction model to obtain the future vehicle traffic density of each area.
5. The traffic data analysis method according to claim 4, characterized in that: The recommending of public transportation routes for users based on user travel data refers to collecting historical travel data of users including commonly used routes, travel time and frequency, and preference settings, preprocessing the collected data and using a K-means clustering algorithm to classify users into morning rush hour commuters and weekend travelers based on user data, using an Apriori algorithm to identify frequent travel combinations of users, including commonly used starting and destination stations, integrating future vehicle traffic density data and generating a feature vector for each planned public transportation route, including route length, vehicle traffic density, and historical user selection frequency; Using the random forest model, historical user travel data is used as training data to input into the random forest model for training. The cross-validation method is used to optimize the number of trees and the depth of the random forest and iterate until the model performance is no longer significantly improved. Then the iteration is stopped to obtain a trained random forest model. The generated feature vector is input into the random forest model to obtain the score of each planned public transportation route. The public transportation routes are sorted in descending order by score, and the public transportation route with the highest score is selected as the optimal public transportation route and pushed to the user in real time.
6. A traffic data analysis system based on the traffic data analysis method according to any one of claims 1 to 5, characterized in that: include, Data collection module, used to collect traffic data from multiple data sources and store it in the central database after pre-processing to form a comprehensive urban traffic data system; The data analysis module is used to build a dynamic graph structure and a spatiotemporal tensor network based on real-time traffic data in the urban traffic data system and predict the travel demand of public transportation in each region through a graph neural network; Level assessment module, used to analyze the traffic data of each area in real time to assess the public transportation service level of each area; Route planning module, used to plan public transportation routes in various regions based on the travel demand and service level of public transportation in various regions; The personalized travel module is used to recommend public transportation routes to users in real time based on the planned public transportation routes and according to the user's travel data and the vehicle traffic density of each public transportation route.
7. A computer device comprising: Memory and processor; The memory stores a computer program, wherein the processor implements the steps of the traffic data analysis method according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the traffic data analysis method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Public transport station layout optimization method
CN117196197A
Intelligent bus scheduling method and system
CN117671992A