Highway-oriented traffic big data traceability analysis method and device in early stage and storage medium

By using multi-source data analysis and latent category models, the problem of insufficient early-stage traffic data collection for highways was solved, resulting in optimized highway construction plans and improving the accuracy and real-time performance of traffic planning.

CN120636165BActive Publication Date: 2025-11-18HUNAN COMM RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511104255.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-18
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

Existing technologies struggle to obtain comprehensive and accurate traffic data in the early stages of highway traffic data collection and analysis, failing to reflect changes in traffic conditions in real time, resulting in insufficient precision in traffic flow prediction and management optimization decisions.

Method used

By acquiring multi-source traffic data, including highway gantry data, toll station entrance and exit data, mobile phone signaling data, traffic survey data, and geographic information system data, we conduct source analysis of vehicle travel routes, extract travel behavior characteristics, and use latent category models for analysis to generate highway corridor optimization schemes and interchange setting optimization schemes.

Benefits of technology

It provides comprehensive data support, identifies traffic flow hotspots, reveals user travel patterns, optimizes highway construction plans, meets traffic demands and complies with national land spatial planning, and improves the scientific nature and feasibility of planning decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636165B_ABST
    Figure CN120636165B_ABST
Patent Text Reader

Abstract

The application discloses a highway early-stage-oriented traffic big data traceability analysis method and device and a storage medium, relates to the technical field of intelligent transportation systems, and comprises the following steps: acquiring multi-source traffic data including highway gantry data, toll station entrance and exit data, mobile phone signaling data, traffic survey data, geographic information system data and facility point data; performing vehicle driving path traceability analysis according to the highway gantry data and the toll station entrance and exit data, and obtaining path frequency distribution results; obtaining travel behavior characteristics according to the mobile phone signaling data, the traffic survey data, the geographic information system data and the facility point data; analyzing the travel behavior characteristics by means of a latent category model to obtain traffic behavior characteristics; and superimposing the path frequency distribution results, the traffic behavior characteristics and territorial space limiting elements to generate a highway corridor belt optimization scheme and an interworking setting optimization scheme. The application can quantitatively analyze traffic behavior characteristics to optimize highway early-stage planning decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent transportation systems technology, and in particular to methods, devices and storage media for tracing and analyzing traffic big data in the early stages of highway construction. Background Technology

[0002] Currently, traffic data collection and analysis in highway pre-planning work mainly rely on traditional methods such as manual surveys, traffic flow observation stations, and toll station data. While these methods can provide traffic flow and vehicle movement information to some extent, they have significant limitations. For example, manual surveys and sampling analysis struggle to obtain comprehensive and accurate traffic data, leading to insufficient precision in traffic flow prediction and traffic behavior analysis. Furthermore, these methods are ill-suited for handling large-scale traffic data, cannot reflect real-time changes in traffic conditions, and are insufficient to support dynamic traffic management and optimization decisions. Therefore, how to quantitatively analyze traffic behavior characteristics to optimize highway pre-planning decisions has become an urgent problem to be solved. Summary of the Invention

[0003] The purpose of this application is to provide a method, device and storage medium for tracing and analyzing traffic big data in the early stages of highway construction, aiming to solve the technical problem of how to quantitatively analyze traffic behavior characteristics to optimize early-stage highway planning decisions.

[0004] To achieve the above objectives, this application proposes a method for tracing and analyzing traffic big data in the early stages of highway construction, the method comprising:

[0005] Acquire multi-source traffic data, including highway gantry data, toll station entrance and exit data, mobile phone signaling data, traffic survey data, geographic information system data, and facility point data;

[0006] Based on the highway gantry data and the toll station entrance and exit data, a vehicle travel path tracing analysis was performed to obtain the path frequency distribution results;

[0007] Based on the mobile phone signaling data, the traffic survey data, the geographic information system data, and the facility location data, travel behavior characteristics are obtained;

[0008] The travel behavior characteristics are analyzed using a latent category model to obtain traffic behavior characteristics;

[0009] By overlaying the path frequency distribution results, the traffic behavior characteristics, and the land space constraints, an optimization scheme for highway corridors and an optimization scheme for interchange settings are generated.

[0010] Furthermore, to achieve the above objectives, this application also proposes a traffic big data tracing and analysis device for early-stage highway construction, the device comprising:

[0011] The data acquisition module is used to acquire multi-source traffic data, including highway gantry data, mobile phone signaling data, traffic survey data, geographic information system data, and facility point data.

[0012] The path tracing module is used to perform vehicle travel path tracing analysis on the highway gantry data to obtain path frequency distribution results.

[0013] The feature extraction module is used to obtain travel behavior features based on the mobile phone signaling data, the traffic survey data, the geographic information system data, and the facility point data;

[0014] The feature analysis module is used to analyze the travel behavior features through a latent category model to obtain traffic behavior features;

[0015] The planning module is used to overlay the path frequency distribution results, the traffic behavior characteristics, and land space constraints to generate highway corridor optimization schemes and interchange setting optimization schemes.

[0016] Furthermore, to achieve the above objectives, this application also proposes a traffic big data source tracing and analysis device for the early stages of highway construction. The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the traffic big data source tracing and analysis method for the early stages of highway construction as described above.

[0017] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the above-described method for tracing and analyzing traffic big data in the early stages of highway construction.

[0018] One or more technical solutions proposed in this application have at least the following technical effects:

[0019] First, the traffic source tracing and analysis system acquires multi-source traffic data, including highway gantry data, toll station entrance and exit data, mobile phone signaling data, traffic survey data, geographic information system data, and facility point data. This provides a comprehensive data foundation for subsequent analysis. Next, the system uses gantry and toll station data to conduct vehicle travel path tracing analysis, generating path frequency distribution results that visually demonstrate the usage frequency of different paths, helping to identify traffic flow hotspots. Then, the system combines mobile phone signaling data, traffic survey data, geographic information system data, and facility point data to extract travel behavior characteristics, revealing the travel patterns and preferences of different users and providing more detailed user behavior data for traffic planning. Through latent category modeling, the system further derives traffic behavior characteristics, classifies complex travel behaviors, identifies groups with similar travel patterns, and provides a basis for formulating targeted traffic policies. Finally, the system overlays path frequency distribution results, traffic behavior characteristics, and land space constraints to generate optimized highway corridor and interchange settings. This step comprehensively considers traffic demand, land space constraints, and traffic behavior characteristics to ensure that highway construction meets traffic demand, complies with land space planning, and is technically and economically feasible. It also enables quantitative analysis of traffic behavior characteristics to optimize early-stage highway planning decisions. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0021] Figure 1 This is a flowchart illustrating an embodiment of the traffic big data source tracing and analysis method for highway pre-construction in this application.

[0022] Figure 2 A schematic diagram of some highways and gantries in Hunan Province provided as an example of the traffic big data source tracing and analysis method for highway pre-construction in this application;

[0023] Figure 3 This is a schematic diagram of traffic flow distribution on a highway section, provided as an example of the traffic big data source tracing and analysis method for highway pre-construction in this application.

[0024] Figure 4 This is a schematic diagram showing the distribution of truck stopping points in a certain urban agglomeration, provided as an example of the traffic big data source tracing and analysis method for the early stage of highway construction in this application.

[0025] Figure 5This is a flowchart illustrating Embodiment 2 of the method for tracing and analyzing traffic big data in the early stages of highway construction, as provided in this application.

[0026] Figure 6 This is a schematic diagram of the K-means clustering effect provided in Embodiment 2 of the method for tracing and analyzing traffic big data in the early stages of highway construction, as presented in this application.

[0027] Figure 7 This is a schematic diagram of the module structure of a traffic big data tracing and analysis device for the early stages of highway construction, as described in this application embodiment.

[0028] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0029] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application. To better understand the technical solutions of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0030] It should be noted that the executing entity of this application embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or traffic source tracing analysis system capable of realizing the above functions. The following uses a traffic source tracing analysis system as an example to describe this embodiment and the following embodiments.

[0031] Based on this, the embodiments of this application provide a method for tracing and analyzing traffic big data in the early stages of highway construction, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the traffic big data source tracing and analysis method for the early stages of highway construction according to this application.

[0032] In this embodiment, the method for tracing and analyzing traffic big data in the early stages of highway construction includes steps S10 to S50:

[0033] Step S10: Obtain multi-source traffic data, including highway gantry data, toll station entrance and exit data, mobile phone signaling data, traffic survey data, geographic information system data, and facility point data.

[0034] It should be noted that highway gantry data refers to the data collected by the ETC gantry system installed along the highway, including specific information about vehicles passing through the gantry, such as vehicle license plate number, direction of travel, passage time, toll mileage, vehicle type code, etc. This data can reflect the specific travel path and time of the vehicle on the highway.

[0035] Table 1. Examples of partial highway gantry data

[0036]

[0037] Please refer to Figure 2 , Figure 2 This diagram illustrates a portion of Hunan Province's expressways and gantries, providing an example of the traffic big data source tracing analysis method for highway pre-planning in this application. Blue lines represent expressway routes, while brown diamond-shaped markers along the routes indicate the locations of ETC gantries. These gantries are key facilities on expressways used for electronic toll collection and traffic monitoring, recording the specific time and location information of vehicle passage. The expressway network in the diagram shows the main traffic arteries in the region, connecting several important cities and areas, providing fundamental support for transportation in Hunan Province. Through the data collected by these gantries, traffic planners can conduct vehicle travel path source tracing analysis to understand traffic flow distribution, thereby optimizing highway pre-planning and traffic management strategies, and improving road use efficiency and safety.

[0038] Toll station entry and exit data refers to the data recorded when vehicles enter or exit highway toll stations, including vehicle license plate number, entry and exit time, toll station number, and toll amount. This data allows us to understand the vehicle's origin and destination, as well as the travel time and mileage on the highway.

[0039] Table 2 Examples of data for some toll station entrances and exits

[0040]

[0041] Mobile signaling data refers to the signal data exchanged between base stations and mobile terminals in a mobile communication network. It includes the mobile phone's location information, call duration, and SMS sending time. Through this data, users' activity trajectories and travel patterns can be obtained, thereby inferring traffic flow and travel demand.

[0042] Traffic survey data refers to information such as traffic flow, vehicle speed, and vehicle type collected manually or automatically. This data is usually obtained through facilities such as traffic flow observation stations and manual survey points, and can reflect the traffic conditions of a specific road section or area.

[0043] Geographic Information System (GIS) data refers to data related to geographic spatial location, including information such as topography, landforms, hydrology, meteorology, land use, and ecological protection red lines. It is usually presented in the form of maps and can provide detailed information about geographic space.

[0044] Facility point data refers to the specific location and attribute information related to transportation facilities, such as toll stations, service areas, interchanges, bridges, and tunnels. It includes information such as the geographical location, type, scale, and function of the facilities, and is used to analyze the layout and function of transportation facilities and to assess the impact of transportation facilities on traffic flow.

[0045] Understandably, the traffic source tracing and analysis system first connects to the highway management department's data interface to obtain real-time vehicle passage information recorded by the highway gantry system, ensuring the timeliness and accuracy of the data. Second, the system extracts vehicle entry and exit data from the toll station management system to supplement and verify the gantry data, constructing a complete vehicle travel trajectory. Then, the system obtains mobile phone signaling data through a data sharing protocol established with mobile communication operators. This data reflects users' activity trajectories and travel patterns, providing a more comprehensive perspective for traffic flow analysis. Next, the system obtains traffic survey data from traffic survey departments. This data, collected manually or automatically, is used to verify and supplement other data sources, improving the reliability of the analysis. Afterward, the system obtains geographic information system (GIS) data from geographic information departments for analyzing the selection and evaluation of highway corridors, ensuring that highway construction is coordinated with land use planning. Finally, the system collects facility point data to assess the impact of traffic facilities on traffic flow and optimize the layout of traffic facilities. By integrating these multi-source data, the traffic source tracing and analysis system can provide comprehensive and accurate data support for preliminary highway work.

[0046] Step S20: Based on the highway gantry data and the toll station entrance / exit data, perform vehicle travel path tracing analysis to obtain path frequency distribution results.

[0047] It should be noted that the path frequency distribution result refers to the statistical statistics of the frequency with which vehicles choose different travel routes on the highway network within a specific time period. Specifically, it is determined by analyzing highway gantry data and toll station entrance and exit data to identify the specific path taken by each vehicle from the entrance toll station to the exit toll station, and to count the number of times each path is selected.

[0048] As an example, the step of performing vehicle travel path tracing analysis based on the highway gantry data and the toll station entrance / exit data to obtain the path frequency distribution result includes: generating an initial vehicle trajectory point sequence based on the gantry location information and timestamp in the highway gantry data; extracting the entrance toll station location, entrance time, exit toll station location, and exit time from the toll station entrance / exit data; merging the entrance toll station location and entrance time as the starting point and the exit toll station location and exit time as the ending point with the initial vehicle trajectory point sequence in chronological order to obtain a complete vehicle path sequence; mapping each location point in the complete vehicle path sequence to the corresponding road segment according to a preset road segment topology network to generate a road segment traffic sequence; performing frequency statistics on the road segment traffic sequences with the same entrance toll station location and the same exit toll station location to obtain frequency statistics results; and converting the frequency statistics results into probability distribution data to obtain the path frequency distribution result.

[0049] Gantry location information refers to the specific geographical location of each gantry in a highway gantry system. It typically includes information such as the gantry number, longitude, and latitude, and is used to determine the specific location of a vehicle when it passes through the gantry.

[0050] A timestamp is a record of the specific time a vehicle passes through a gantry or toll station, usually recorded in the format of year, month, day, hour, minute, and second.

[0051] Initial vehicle trajectory point sequence This refers to a sequence of initial trajectory points generated based on gantry location information and timestamps, showing the vehicle's initial path along the highway. These points are arranged chronologically, reflecting the vehicle's initial travel path at a specific time and location. (Entrance toll station location) This refers to the specific geographical location of the toll station a vehicle passes through when entering a highway, typically including the toll station number, longitude, latitude, and other information. Entry time. This refers to the specific time a vehicle passes through the entrance toll station when entering the highway, usually recorded in the format of year, month, day, hour, minute, and second. Exit toll station location. This refers to the specific geographical location of the toll station a vehicle passes through when leaving the highway, usually including the toll station number, longitude, latitude, and other information.

[0052] Export time This refers to the specific time a vehicle passes through an exit toll station when leaving the highway, usually recorded in the format of year, month, day, hour, minute, and second. The starting point is the beginning of the vehicle's journey, determined by the location of the entrance toll station and the entry time, used to identify the vehicle's specific location and time of entry onto the highway. The ending point is the end of the vehicle's journey, determined by the location of the exit toll station and the exit time, used to identify the vehicle's specific location and time of departure from the highway.

[0053] Complete vehicle route sequence This refers to a continuous path from the entrance to the exit of a highway, formed by combining the vehicle's starting point upon entering the highway, the initial sequence of vehicle trajectory points recorded by various gantries during its journey, and the vehicle's exit point, arranged chronologically. This path records in detail every position and corresponding time of the vehicle on the highway, accurately reflecting the vehicle's actual travel trajectory.

[0054] Predefined road segment topology network refers to the predefined topology of a highway network, including the connection relationships and location information of each road segment. It is used to map the driving path of a vehicle to a specific road segment and is the basis for path frequency distribution analysis.

[0055] Road segment traffic sequence This refers to the sequence of road segments that a vehicle passes through on a highway. It is generated by mapping each location point in the complete vehicle path sequence to a specific road segment in a preset road segment topology network. It reflects the actual driving path of the vehicle on the highway and is key data for path frequency distribution analysis.

[0056] Frequency statistics refer to the number of times each different route was selected within a specific time period for all vehicle routes with the same entrance and exit toll stations. This result is presented numerically, showing the usage frequency of each route, i.e., how many vehicles chose that route. For example, if 100 vehicles depart from toll station A and arrive at toll station B, and 60 vehicles chose route 1, 30 vehicles chose route 2, and 10 vehicles chose route 3, then the frequency of route 1 is 60, the frequency of route 2 is 30, and the frequency of route 3 is 10.

[0057] Probability distribution data refers to converting frequency statistics into probability form, representing the probability of each path being selected.

[0058] First, the traffic source tracing analysis system analyzes highway gantry data to extract the gantry number j, gantry location information (x, y), and timestamp t from each record. Based on the chronological order of the timestamps, it concatenates the records of the same vehicle passing through different gantries to form the initial trajectory point sequence of the vehicle.

[0059]

[0060] This step is to initially outline the vehicle's trajectory on the highway. Next, the system filters the tollbooth entry and exit data to identify the vehicle's entry tollbooth. Entry time Exit toll station Exit toll station time These points form the starting and ending points of the path, used to determine the vehicle's complete journey. Then, the starting points, the initial trajectory point sequence, and the ending points are merged in chronological order to form the complete vehicle path sequence.

[0061]

[0062] The purpose of this step is to construct the complete driving path of the vehicle from the entrance to the exit, in order to more accurately analyze the actual driving situation of the vehicle. Next, the system uses a pre-defined road segment topology network... By using nearest neighbor matching, each location point in the path sequence is matched with a road segment in the network to obtain the vehicle's road segment passage sequence. This allows for the refinement of vehicle travel paths down to specific road segments, providing a more detailed basis for path frequency analysis. Subsequently, for all vehicles with the same entrance-exit pairs (…),… , The traffic sequences of road segments are aggregated, and it is assumed that... For all paths in the set, count the unique paths for each path. Frequency of selection :

[0063]

[0064] This step quantifies the usage of different paths, providing data support for subsequent probability analysis. Finally, the system converts the frequency statistics into probability distribution data, calculating the proportion of each path's usage frequency to the total path usage frequency to obtain the path frequency distribution results.

[0065]

[0066] This leads to the inlet-outlet pair ( , The set of path probability distributions under:

[0067]

[0068] This set intuitively reflects the probability of vehicles traveling on different paths, providing a scientific basis for early-stage highway planning and decision-making, helping decision-makers understand the distribution patterns of traffic flow, and optimize highway network design.

[0069] Step S30: Based on the mobile phone signaling data, the traffic survey data, the geographic information system data, and the facility point data, obtain travel behavior characteristics.

[0070] It should be noted that travel behavior characteristics refer to the various characteristics and patterns exhibited by individuals or groups during transportation. These characteristics reflect information such as travelers' preferences, travel purposes, travel time and spatial distribution.

[0071] As an example, the step of obtaining travel behavior characteristics based on the mobile phone signaling data, the traffic survey data, the geographic information system data, and the facility point data includes: associating trajectory points in the mobile phone signaling data with the travel destination field in the traffic survey data to generate an initial trajectory chain; correcting the positioning deviation of the initial trajectory chain based on the road network topology layer in the geographic information system data to obtain a reference trajectory chain; associating the functional area codes in the facility point data with the reference trajectory chain to generate fused trajectory chain data; performing spatiotemporal slicing processing on the fused trajectory chain data to obtain travel time period units; calculating the origin-destination distribution entropy value and path fluctuation entropy value of the travel time period unit; and marking travel behavior characteristics according to the range of the origin-destination distribution entropy value and the path fluctuation entropy value.

[0072] Track points refer to user location information recorded in mobile phone signaling data, which typically includes longitude, latitude, and timestamp, reflecting the user's specific location at different times.

[0073] The travel purpose field refers to the specific purpose of a user's trip recorded in the traffic survey data, such as going to work, school, shopping, leisure, etc. These fields provide users' motivation and background information for their trips, which helps to understand traffic demand and behavioral characteristics under different travel purposes.

[0074] Initial trajectory chain It refers to the preliminary trajectory sequence generated by associating trajectory points in mobile phone signaling data with the travel purpose field in traffic survey data. It reflects the user's travel path and purpose in different time periods and is the basis for analyzing user travel behavior.

[0075] The road network topology layer refers to the layer in geographic information system data that describes the road network structure, including information such as road connections, directions, and lengths. This data is used to correct the positioning deviation of trajectory points and ensure that the trajectory chain is consistent with the actual road network.

[0076] Reference trajectory chain It refers to the trajectory chain after correction by the road network topology layer, which more accurately reflects the user's driving path on the actual road network.

[0077] Functional area coding refers to the codes used in facility point data to identify different functional areas, such as commercial areas, residential areas, and industrial areas. These codes provide attribute information about the user's destination and departure points, which helps to analyze the user's travel patterns between different functional areas.

[0078] Fusion trajectory chain data This refers to the data generated by combining the reference trajectory chain with the functional area code. It not only includes the user's travel path, but also the functional area information traversed by the path.

[0079] Travel time unit This refers to travel data within a specific time period obtained by processing fused trajectory chain data into time slices. These units are used to analyze users' travel behavior patterns in different time periods.

[0080] Origin and destination distribution entropy values It refers to an indicator that quantifies the randomness of the distribution of origin and destination points for user travel. It assesses the diversity and uncertainty of user travel by calculating the probability distribution entropy value of the origin and destination point distribution.

[0081] Path fluctuation entropy It refers to an indicator that quantifies the degree of fluctuation in user travel paths. It assesses the stability and variability of user travel paths by calculating the probability distribution entropy value of path fluctuations.

[0082] First, the traffic source tracing and analysis system uses data matching algorithms (such as the BF algorithm) to identify trajectory points in mobile phone signaling data. With the travel purpose field in traffic survey data Perform correlation to construct an initial trajectory chain. ,in This step, which identifies the travel purpose corresponding to the point, aims to imbue the trajectory data with semantic information about the travel purpose, facilitating subsequent analysis of travel behavior patterns. Secondly, the system utilizes the road network topology layer from the geographic information system data. The initial trajectory chain is corrected by using spatial analysis methods, such as nearest neighbor matching or geofencing, to accurately map the trajectory points onto the actual road network, correcting the positioning deviation of the trajectory points and obtaining a more accurate reference trajectory chain.

[0083]

[0084] in, This indicates the number of the specific road onto which the projection occurs.

[0085] This step ensures consistency between the trajectory data and the actual road network, improving the accuracy of the trajectory data.

[0086] Then, the system encodes the functional areas in the facility point data. By associating with the reference trajectory chain and using nearest neighbor matching technology, each trajectory point is matched with the functional areas it passes through to generate fused trajectory chain data:

[0087]

[0088] This represents the vehicle's movement behavior between specific functional areas. Next, the system processes the fused trajectory chain data at preset time intervals. Spatiotemporal slicing is performed to obtain a set of travel time period units:

[0089]

[0090] in, User The travel behavior of the j-th travel time segment.

[0091] Then, the system calculates each travel time segment unit. The origin and destination distribution entropy value and path fluctuation entropy The calculation method is as follows:

[0092] ,

[0093] in, Let be the probability that the starting point or ending point belongs to the i-th functional area. Let be the probability that the j-th path is selected for a given OD pair.

[0094] Finally, the system according to and The numerical range maps travel behavior to a set of tags. This yields the final behavioral feature vector (travel behavior features):

[0095]

[0096] This step aims to classify the quantified travel behavior characteristics in order to better understand and compare the travel habits and needs of different users, and to provide a scientific basis for traffic planning and management.

[0097] Step S40: Analyze the travel behavior characteristics using a latent category model to obtain traffic behavior characteristics.

[0098] It should be noted that Latent Class Model (LCM) is a statistical analysis method used to identify and classify individuals or groups with similar characteristics. In traffic big data analysis, LCM divides travelers into different categories or groups by analyzing travel behavior characteristics. These categories reflect different travel patterns or behavioral characteristics, such as commuters.

[0099] Traffic behavior characteristics refer to the various features and patterns exhibited by individuals or groups during traffic travel. These characteristics reflect information such as travelers' preferences, travel purposes, and the time and spatial distribution of their trips. These characteristics include, but are not limited to, the distribution of origin and destination points, route selection, the regularity of travel time, and travel purpose. By analyzing these characteristics, different travel modes, such as commuting, leisure, and shopping, can be identified, and the temporal and spatial distribution patterns of these modes can be further understood.

[0100] Understandably, the traffic source tracing and analysis system first organizes the collected travel behavior characteristic data, and then analyzes each individual traveler... behavioral characteristics Represented in eigenvector form:

[0101]

[0102] Where d represents the dimension of travel behavior characteristics, such as origin-destination distribution, route selection, travel purpose, travel time, functional area type, etc.

[0103] Secondly, with Using C as the feature variables, we construct a latent category model to classify travelers. Then the mixed distribution of the overall behavioral characteristics can be expressed as:

[0104]

[0105] Then, the Expectation-Maximization (EM) algorithm is used to estimate the model parameters, obtaining the prior probabilities and conditional distribution parameters for each category, and calculating the parameters for each traveler. Posterior probabilities of each category:

[0106]

[0107] in, It refers to a potential travel category.

[0108] And based on the principle of maximum a posteriori probability, the travelers Divide into its most likely potential categories:

[0109]

[0110] Next, based on the analysis results of the latent category model, the system extracts the traffic behavior features of each category and obtains specific and actionable feature information from the model's output.

[0111] Step S50: Overlay the path frequency distribution results, the traffic behavior characteristics, and the land space constraints to generate the highway corridor optimization scheme and the interchange setting optimization scheme.

[0112] It should be noted that the restrictive elements of land space refer to various elements that need to be strictly controlled in land space planning. The distribution and protection requirements of these elements restrict the selection of highway corridors and the setting of interchanges, including basic farmland, ecological red lines, nature reserves, and water source protection areas.

[0113] The highway corridor optimization scheme refers to the layout and route selection plan for highway corridors proposed based on comprehensive consideration of traffic demand, land space constraints, and traffic behavior characteristics. This scheme aims to improve the scientific and rational nature of highway construction, ensuring that highway construction meets traffic demands, complies with land space planning requirements, and minimizes its impact on the ecological environment and socio-economic development.

[0114] Interchange optimization schemes refer to layout and design plans for highway interchanges that comprehensively consider traffic demand, traffic behavior characteristics, and land space constraints. These schemes aim to improve the efficiency of the highway network, optimize traffic flow organization and diversion, and reduce the impact on the ecological environment and socio-economic development. For example, by setting up interchanges near truck stopping points, truck traffic efficiency can be improved, traffic congestion reduced, and regional economic development promoted.

[0115] As an example, the steps of superimposing the path frequency distribution results, traffic behavior characteristics, and land space constraints to generate highway corridor optimization schemes and interchange setting optimization schemes include: converting the path frequency distribution results into a road network traffic heat map, and superimposing the spatial distribution of traffic demand in the traffic behavior characteristics onto the road network traffic heat map to obtain a traffic demand distribution map; constructing a spatial constraint map based on land space constraints, and weighting and superimposing the traffic demand distribution map with the spatial constraint map to generate a comprehensive suitability layer; performing minimum cost path analysis based on the comprehensive suitability layer to generate a preliminary highway corridor scheme; performing isochronous circle calculations based on truck GPS trajectory data and preset time thresholds to determine interchange setting hotspot areas; conducting land space coordination assessments, traffic function matching degree assessments, and engineering implementation feasibility assessments on the preliminary highway corridor scheme and the interchange setting hotspot areas, and generating highway corridor optimization schemes and interchange setting optimization schemes based on the assessment results.

[0116] A road network traffic heatmap is a visualization tool used to show the distribution of traffic flow on different road segments in a highway network. By converting the path frequency distribution results into a color-coded map, the map intuitively shows which road segments have high traffic flow and which have low traffic flow.

[0117] Please refer to Figure 3 , Figure 3 This diagram illustrates the traffic flow distribution of highway sections, as provided in Embodiment 1 of the traffic big data source tracing and analysis method for highway pre-construction, as presented in this application. Different colored lines represent different traffic flow levels. In the diagram, each highway is divided into eight levels based on its traffic volume, with different colored lines representing the sections with the lowest to highest traffic volumes. These lines visually reflect the usage of major traffic arteries. For example, sections with the highest traffic volume may be located in areas with frequent economic activity or connecting important cities, while sections with lower traffic volume may be located in areas with relatively low traffic demand. This detailed traffic flow distribution information is crucial for traffic planning and management, helping decision-makers identify traffic bottlenecks, optimize road network structure, and formulate effective traffic management strategies to improve the operational efficiency and service quality of the entire transportation system.

[0118] Spatial distribution of traffic demand refers to the traffic demand of travelers in different geographical locations, reflecting the frequency and purpose of travel in different areas, and providing detailed geographical distribution information for traffic planning.

[0119] A traffic demand distribution map is a map generated by overlaying a road network traffic heat map with the spatial distribution of traffic demand. It integrates the geographical distribution of traffic flow and travel demand, providing a more comprehensive view of traffic demand and helping planners identify areas with high traffic demand and high traffic volume.

[0120] Spatial constraint maps are maps constructed based on national spatial constraints. They show the areas that highway construction needs to avoid, clarifying the limiting conditions for highway construction and reducing damage to the ecological environment.

[0121] The comprehensive suitability layer is a map generated by weighted overlay of a traffic demand distribution map and a spatial constraint map. It takes into account both traffic demand and land space constraints, and through weight allocation, generates a map reflecting the suitability of different areas. This helps planners identify areas that meet both traffic demand and land space planning, providing a scientific basis for the optimization of highway corridors.

[0122] The preliminary highway corridor plan is a highway alignment scheme generated based on minimum cost path analysis using a comprehensive suitability layer. It determines the optimal path from the starting point to the end point by analyzing the comprehensive suitability layer, taking into account traffic demand and land space constraints.

[0123] Truck GPS trajectory data refers to the driving trajectory data of trucks recorded by GPS devices, including information such as timestamps and latitude and longitude, reflecting the actual driving situation of the trucks.

[0124] The preset time threshold refers to the time limit set when analyzing truck stop points, used to distinguish between short stops and actual loading and unloading activities. In this embodiment, the preset time threshold is 30 minutes.

[0125] Interchange hotspot areas refer to areas with a high density of truck stops, determined by isochronous circle calculations. These are typically loading and unloading points or logistics nodes for trucks and are priority areas for setting up interchanges.

[0126] Please refer to Figure 4 , Figure 4 This diagram illustrates the distribution of truck stop points in a regional urban agglomeration, as provided in Example 1 of the traffic big data source tracing and analysis method for highway pre-planning in this application. The diamond-shaped markers indicate the locations where trucks stop. These stop points are mainly concentrated along major traffic arteries, logistics parks, industrial zones, and commercial centers within the urban agglomeration, reflecting frequent loading and unloading activities by trucks in these areas. The dense distribution of stop points in the diagram reveals logistics demand hotspots within the urban agglomeration, providing important reference information for optimizing traffic facility layout, improving logistics efficiency, and alleviating traffic congestion. This data can be used to further analyze truck travel routes and assess traffic flow distribution, thereby providing a scientific basis for highway pre-planning and traffic management, and enhancing the adaptability and service capacity of the transportation system.

[0127] The assessment of land spatial coordination refers to evaluating whether the highway corridor plan and interchange plan comply with the requirements of the land spatial plan, and ensuring that highway construction will not have an adverse impact on the ecological environment and economy.

[0128] Traffic function matching assessment refers to evaluating whether the highway corridor plan and interchange setting plan meet traffic demand, and ensuring that highway construction can effectively meet traffic flow and travel needs.

[0129] Project feasibility assessment refers to evaluating the feasibility of highway corridor and interchange design schemes in actual construction. By analyzing factors such as terrain conditions, construction difficulty, and cost, it ensures that highway construction is technically and economically feasible.

[0130] The assessment results refer to the comprehensive results of the land space coordination assessment, the traffic function matching assessment, and the project implementation feasibility assessment. These results provide a scientific basis to help decision-makers select the optimal highway construction plan, ensuring that highway construction not only conforms to the land space plan and meets traffic needs, but is also technically and economically feasible.

[0131] First, the path frequency distribution results are processed using geographic information system software. Mapped to the highway network A traffic heatmap of the road network is constructed by weighting the traffic flow. Different color values ​​are assigned based on frequency, with darker colors indicating higher traffic volume, thus visually displaying the traffic flow distribution across different road segments. Secondly, the spatial distribution of travel demand from travel group C is overlaid onto this heatmap. The geographic coordinates and demand intensity of traffic demand data are extracted. Using the overlay analysis function of GIS, these demand points are spatially matched with heat maps, and the demand intensity is overlaid onto the corresponding locations in the form of transparency or color depth to obtain a traffic demand distribution map that comprehensively reflects traffic flow and demand. This provides a more comprehensive basis for highway planning. Next, a land space constraint cost layer is constructed based on land space constraint elements. Each spatial unit Cost value It is positively correlated with its corresponding restriction level. And the traffic demand distribution map... With space-constrained cost layers Weighted overlay is performed to generate a comprehensive suitability layer. :

[0132]

[0133] in, and For the weight parameters, satisfying .

[0134] In the appropriate layer The Dijkstra algorithm is used to find the minimum cost path from the starting point s to the ending point t. This process generates a preliminary highway corridor plan, aiming to find the optimal route while meeting traffic demand and spatial constraints. Simultaneously, isochronous circles are calculated based on truck GPS trajectory data and preset time thresholds. Centered on the truck's stopping point, calculate the range that the truck can reach within a preset time threshold, and statistically analyze the trajectory density of each stopping area. Select those with a density higher than the threshold The areas are designated as hotspots for interchange construction. Finally, the preliminary highway corridor plan and interchange hotspots are evaluated from three dimensions: spatial coordination, traffic function matching, and engineering feasibility. This involves checking whether the plan conforms to the national spatial planning, meets traffic demands, and is technically and economically feasible. Based on the evaluation results, the plan is adjusted and optimized to generate the final optimized highway corridor plan and interchange construction plan. This step ensures the plan's scientific validity, rationality, and feasibility.

[0135] This embodiment provides a method for tracing and analyzing traffic big data in the early stages of highway construction. First, the traffic tracing and analysis system acquires multi-source traffic data, including highway gantry data, toll station entrance and exit data, mobile phone signaling data, traffic survey data, geographic information system (GIS) data, and facility point data, providing a comprehensive data foundation for subsequent analysis. Next, the system uses gantry and toll station data to perform vehicle travel path tracing analysis, generating path frequency distribution results that visually display the usage frequency of different paths, helping to identify traffic flow hotspots. Then, the system combines mobile phone signaling data, traffic survey data, GIS data, and facility point data to extract travel behavior characteristics, revealing the travel patterns and preferences of different users, providing more detailed user behavior data for traffic planning. Through latent category modeling, the system further derives traffic behavior characteristics, classifies complex travel behaviors, identifies groups with similar travel patterns, and provides a basis for formulating targeted traffic policies. Finally, the system overlays path frequency distribution results, traffic behavior characteristics, and land space constraints to generate optimized highway corridor and interchange settings. This step comprehensively considers traffic demand, land space constraints, and traffic behavior characteristics to ensure that highway construction meets traffic demand, complies with land space planning, and is technically and economically feasible. It also enables quantitative analysis of traffic behavior characteristics to optimize early-stage highway planning decisions.

[0136] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 5 , Figure 5 This is a flowchart illustrating the second embodiment of the traffic big data source tracing and analysis method for the early stages of highway construction according to this application. Step S40 of the traffic big data source tracing and analysis method for the early stages of highway construction includes steps S41 to S43:

[0137] Step S41: Based on the travel behavior characteristics, perform travel group clustering analysis using a latent category model to obtain group classification results.

[0138] It should be noted that the group classification results refer to the classification of travelers into different categories after analyzing their travel behavior characteristics using a latent category model. These categories reflect groups with similar travel patterns and characteristics, such as commuters, leisure travelers, business travelers, drivers, taxi users, and truck drivers. Each category has its unique travel behavior characteristics, such as travel time, travel purpose, and route selection.

[0139] As an example, the steps of performing cluster analysis on the travel group based on the travel behavior characteristics to obtain the group classification results through a latent category model include: extracting travel mode features, travel time features, and personal attribute features from the travel behavior characteristics; inputting the travel mode features, travel time features, and personal attribute features into a latent category model for latent variable analysis to obtain a feature latent variable representation; performing cluster assignment on the feature latent variable representation to obtain an initial group classification result; performing silhouette coefficient analysis on the initial group classification result to determine the optimal number of classifications; and using the optimal number of classifications as a clustering quantity parameter to perform latent variable analysis and cluster assignment again to obtain the group classification result.

[0140] Travel mode characteristics refer to the specific mode of transportation chosen by travelers, such as walking, cycling, public transportation, cars, and taxis. In this embodiment, travel mode characteristics are obtained by analyzing multi-source data, including mobile phone signaling data and traffic survey data.

[0141] Travel time characteristics refer to the distribution of travelers' travel times throughout the day, including the start and end times of their trips. In this embodiment, travel time characteristics are obtained by analyzing mobile phone signaling data and traffic survey data, reflecting the travel habits of different travelers at different times of the day.

[0142] Personal attribute characteristics refer to attribute information related to an individual traveler, such as age, gender, occupation, and income level. In this embodiment, personal attribute characteristics are obtained by analyzing traffic survey data and mobile phone signaling data, which may contain user registration information or attributes inferred through data analysis.

[0143] Feature latent variable representation refers to the mathematical representation of latent variables obtained after analyzing travel behavior characteristics through latent category models. These latent variables reflect the underlying structure and patterns behind travel behavior characteristics and are a simplification and abstraction of complex data by the model.

[0144] The initial group classification result refers to the preliminary classification result obtained by clustering and assigning based on the latent variable representation of features after the latent category model analysis. It divides travelers into different categories, and each category represents a travel behavior pattern.

[0145] The optimal number of clusters refers to the best number of clusters determined through silhouette coefficient analysis. In this embodiment, the silhouette coefficient is used to evaluate the classification quality under different numbers of clusters. By calculating the silhouette coefficient of each cluster, the number of clusters with the highest silhouette coefficient is selected as the optimal number of clusters.

[0146] The cluster number parameter refers to the number of clusters set in cluster analysis. It is determined based on the optimal number of clusters and is used to re-analyze latent variables and assign clusters.

[0147] First, the traffic source tracing and analysis system uses data filtering and field extraction techniques to separate data fields related to travel mode, travel time, and personal attributes from travel behavior characteristics, constructing a feature dataset for modeling. Second, the system inputs the feature dataset into a latent category model for latent variable analysis, employing the Expectation-Maximization (EM) algorithm to estimate the probability that each individual traveler belongs to different latent categories, thereby identifying potential travel group structures. The posterior probability calculation formula for the latent category model is as follows:

[0148]

[0149] in, z represents the probability that individual i belongs to category c; z represents a certain implicit group type. C represents the feature vector of individual i; C refers to the total number of categories.

[0150] Then, the system uses the K-means clustering algorithm to classify travelers into groups. By minimizing the Euclidean distance between samples and cluster centers, each traveler is assigned to the nearest group center, resulting in the initial group classification. The calculation formula is as follows:

[0151]

[0152] In the formula, It is the feature vector of the i-th sample point. is the vector of the center points of the j-th cluster, and d is the feature dimension.

[0153] Next, the system performs silhouette coefficient analysis on the initial group classification results to evaluate the clustering effect and determine the optimal group size. The silhouette coefficient calculation formula is as follows:

[0154]

[0155] in, It refers to the silhouette coefficient of individual i, which is used to measure the quality of clustering results; This refers to the average distance of individual i to the nearest non-self-category center; It refers to the average distance from individual i to its own category center.

[0156] Finally, the system determines the optimal number of classes with the goal of maximizing the silhouette coefficient. Using these parameters as the final input parameters for latent class models and cluster analysis, the EM algorithm and K-means algorithm are applied again to obtain the final group classification results, providing a scientific basis for traffic planning and management.

[0157] Please refer to Figure 6 , Figure 6 This diagram illustrates the K-means clustering effect provided in Embodiment 2 of the method for tracing and analyzing traffic big data in the early stages of highway construction, as presented in this application. The diagram visualizes the results of clustering travel behavior characteristics using the K-means clustering algorithm. Each cluster is represented by a marker of a different shape, from 1 to 9, corresponding to different groups of travelers. Points within each group are connected by line segments, representing individuals with similar travel characteristics, such as travel time, cost, and service level. The distribution of clusters reveals the aggregation patterns of travelers in the feature space, helping to identify groups with similar travel behaviors. For example, some clusters may represent commuters during peak hours, while others may represent truck drivers active at night. This clustering analysis helps transportation planners understand the travel needs and behavioral patterns of different groups, providing a scientific basis for optimizing traffic services and facility layout.

[0158] As an example, the step of clustering and assigning the latent variable representations of features to obtain an initial population classification result includes: randomly selecting a preset number of sample points from the latent variable representations of features as initial cluster centers; calculating the Euclidean distance from each latent variable representation of features to all the initial cluster centers; assigning each latent variable representation of features to the cluster represented by the nearest initial cluster center according to the Euclidean distance; calculating the mean of the latent variable representations of features in each cluster, and updating the cluster center position according to the mean; and outputting the initial population classification result when the coordinate difference between the updated cluster center position and the unupdated cluster center position is less than a preset convergence threshold.

[0159] The preset number refers to the number of cluster centers that are set in advance during cluster analysis. It is set based on preliminary analysis or experience and is used to initialize the clustering process. In this embodiment, it is 9.

[0160] Initial cluster centers refer to the few sample points randomly selected at the beginning of the clustering algorithm as the starting points for clustering.

[0161] A cluster is a set of data points grouped together based on their similarity in cluster analysis. It is formed by calculating the Euclidean distance from the latent variable representation of each feature to the initial cluster center and assigning the data points to the nearest cluster center. Each cluster represents a group of travel behavior patterns with similar characteristics.

[0162] The mean is the average value calculated for all data points in each cluster during cluster analysis. It is used to update the position of the cluster centers and gradually optimize the clustering results.

[0163] Table 3. Examples of average values ​​for each of the nine K-means clustering clustering methods.

[0164]

[0165] The coordinate difference refers to the distance between the updated cluster center position and the original cluster center position in cluster analysis, and is used to determine whether the clustering algorithm has converged.

[0166] A preset convergence threshold is a small value set in cluster analysis to determine whether the locations of cluster centers have converged. In this embodiment, the preset convergence threshold is 0.01.

[0167] First, the traffic source tracing analysis system randomly selects a predetermined number of sample points from the latent variable representations of features as initial cluster centers. Specifically, a random sampling algorithm is used to select a specific number of sample points from the dataset; these points serve as the initial reference points for subsequent cluster assignments. Second, the system calculates the Euclidean distance from each latent variable representation of a feature to all initial cluster centers. The Euclidean distance formula quantifies the similarity between each data point and a cluster center; shorter distances indicate higher similarity. Then, based on the calculated Euclidean distances, each latent variable representation of a feature is assigned to the cluster represented by the nearest initial cluster center. This process compares the distances of each data point to each cluster center, classifying data points to the nearest cluster center, thus initially forming different cluster groups. Next, the system calculates the mean of the latent variable representations of features within each cluster, i.e., averaging each dimension of all data points in each cluster to obtain new cluster center positions. This process updates the cluster centers, making them closer to the center of the data points within the cluster, thus improving the accuracy of clustering. Finally, when the difference between the coordinates of the updated cluster center and the original cluster center is less than the preset convergence threshold, the system considers the clustering process to have converged and outputs the initial population classification result. This step uses a very small threshold to determine whether the cluster centers are stable, ensuring the stability and reliability of the clustering results.

[0168] Step S42: Construct a hybrid Logit model based on the group classification results, and calculate the mode of transport selection probability based on the hybrid Logit model.

[0169] It should be noted that the hybrid Logit model is a statistical model used to analyze individual choice behavior. It extends the traditional Logit model by allowing for heterogeneity of parameters among individuals. In this embodiment, the hybrid Logit model is used to analyze travelers' choice behavior among different modes of transportation, taking into account that different travelers may be influenced by various factors when choosing a mode of transportation, such as travel time, travel cost, and personal preferences.

[0170] The probability of choosing a mode of transportation refers to the likelihood that a traveler will choose a particular mode of transportation in a given travel situation. It reflects the tendency of different travelers to choose different modes of transportation after considering factors such as travel time, cost, and personal preferences. These probabilities help transportation planners understand the attractiveness of different modes of transportation, predict changes in traffic demand, and thus optimize the layout of transportation facilities and formulate more effective transportation policies.

[0171] As an example, the steps of constructing a hybrid Logit model based on the group classification results and calculating the transportation mode selection probability based on the hybrid Logit model include: extracting a travel feature set for each group from the group classification results, the travel feature set including travel time features, travel cost features, and service level features; constructing an observable utility component based on the travel time features, travel cost features, and service level features; constructing a utility function of the hybrid Logit model based on a preset random error component and the observable utility component to obtain an initial hybrid Logit model; estimating the parameters of the initial hybrid Logit model using the maximum likelihood estimation method to obtain a target hybrid Logit model; and calculating and summing the selection probabilities of each group under different transportation modes based on the target hybrid Logit model to obtain a transportation mode selection probability distribution.

[0172] A travel feature set refers to a series of travel-related features extracted from group classification results. These features reflect various factors considered by different groups when traveling, including travel time characteristics, travel cost characteristics, and service level characteristics. Travel time characteristics involve the time travelers spend on different modes of transportation, such as commuting time and travel time; travel cost characteristics include the costs travelers need to pay for using different modes of transportation, such as fares and fuel costs; and service level characteristics reflect the service quality of transportation modes, such as the frequency of public transportation and the degree of road congestion. These features describe the attractiveness of different modes of transportation and are key factors influencing travelers' choice of transportation mode.

[0173] The observable utility component refers to the utility part that can be constructed based on the travel feature set in a hybrid Logit model. It reflects the benefits or costs that travelers can clearly perceive and quantify when choosing a mode of transportation.

[0174] The pre-defined random error component refers to the error term introduced in a mixed Logit model to account for the randomness and unobservable factors in individual choice behavior, such as personal preferences and ad hoc decisions. The random error term is typically assumed to follow a certain distribution, such as a normal or log-normal distribution, to account for the randomness and unobservable factors in individual choice behavior.

[0175] The utility function, in a mixed Logit model, is a mathematical expression representing the total utility of a traveler choosing a particular mode of transportation. It reflects the combined influence of various factors that travelers consider when choosing a mode of transportation. It is the core part of the mixed Logit model and is used to calculate the probability of a traveler choosing different modes of transportation.

[0176] The initial hybrid Logit model refers to a hybrid Logit model constructed based on the travel feature set and preset random error components before parameter estimation, without having undergone parameter estimation and optimization.

[0177] The objective mixture Logit model refers to the final model obtained after estimating and optimizing the parameters of the initial mixture Logit model using the maximum likelihood estimation method.

[0178] First, the traffic source tracing analysis system extracts the travel feature set for each group from the group classification results:

[0179]

[0180] Among them, travel feature set This includes travel time, travel costs, and service level characteristics.

[0181] The specific approach involves selecting data fields related to travel time, travel costs, and service levels for each group. Travel time features may include average commute time and travel duration, while travel cost features cover transportation expenses and fuel costs. Service level features involve the frequency of public transportation services and the degree of road congestion. These feature sets provide the necessary data foundation for subsequent model construction.

[0182] Secondly, based on the extracted travel time, travel cost, and service level features, observable utility components are constructed. By assigning weight coefficients to each feature, these features are linearly combined into a utility expression. Then, the utility function of the hybrid Logit model is constructed according to the preset random error component and the observable utility component. Specifically, the observable utility component is added to the random error term to obtain the complete utility function, as shown in the following formula:

[0183]

[0184] In the formula, This represents the utility of individual i for travel mode j; These are independent variables, representing characteristics such as travel time, travel cost, and service level; These are regression coefficients, representing the strength of the influence of each feature; It is a random error term.

[0185] This embodiment sets random parameters, assuming that the regression coefficients of the independent variables in the model vary randomly among different travelers, in order to capture the heterogeneity among individuals. Specifically:

[0186]

[0187] in, It is the mean of the parameters, representing the average preference of individuals for travel mode j; It is the standard deviation of the parameter, representing the heterogeneity of individual preferences.

[0188] Next, the parameters of the initial mixed Logit model are estimated using the maximum likelihood estimation method. Specifically, a likelihood function is constructed, representing the probability of observed data occurring given the model parameters. An optimization algorithm, such as the Newton-Raphson method or a quasi-Newton method, is used to maximize the likelihood function, thereby obtaining the optimal model parameters. The purpose of parameter estimation is to enable the model to better fit the observed data and improve the model's predictive accuracy. Finally, the selection probabilities of each group under different modes of transportation are calculated based on the target mixed Logit model and summarized, representing the selection probabilities in the mixed Logit model. The expression is:

[0189]

[0190] in, It represents the probability that individual i chooses mode j. This represents the utility of individual i for travel mode j.

[0191] This embodiment assumes that the observed selection data is That is, if individual i chooses travel mode j, then Otherwise, it is 0. Then the likelihood function... This can be represented as the joint probability of all individuals and choices:

[0192]

[0193] in, It is the set of model parameters to be estimated; It represents the probability that individual i chooses mode j. It is the observed selection indicator variable.

[0194] Since the model contains random parameters, it needs to be converted into a log-likelihood function, calculated as follows:

[0195]

[0196] This embodiment uses the Newton-Raphson method to calculate the parameters in the mixed logit model and sets the initial parameter vector of the model. The initial values ​​of the parameter vector are obtained through preliminary regression estimation.

[0197] Calculate the log-likelihood function Relative to parameter vector The first derivative of is calculated using the following formula:

[0198]

[0199] in, Let be the gradient of the log-likelihood function.

[0200] Calculate the log-likelihood function Relative to parameter vector The Hessian matrix is ​​the second derivative, calculated using the following formula:

[0201]

[0202] This embodiment is based on the current parameters. ,gradient and Hessian matrix The Newton-Raphson method is used to update the model parameters, and the calculation formula is as follows:

[0203]

[0204] in, It is the inverse of the Hessian matrix. This represents the gradient under the current parameters.

[0205] The changes after parameter updates are checked, and if the parameter change is less than a preset threshold... If the model converges, the iteration stops. The convergence criteria are as follows:

[0206]

[0207] In the formula, The norm of a vector This is the set convergence accuracy threshold.

[0208] Through multiple iterative optimizations, the log-likelihood function is converged, yielding the final coefficient values ​​for each feature variable in relation to the travel mode classification. If the mean and variance of any random parameter are not significant at the 0.1 significance level, they are set as fixed parameters, and the estimation process is repeated until all parameters are significant.

[0209] Finally, using the estimated model parameters and the utility function, the probability of each group choosing each mode of transportation is calculated. Summarizing these probabilities yields a probability distribution of mode of transportation choice, providing a scientific basis for transportation planning and management, helping decision-makers understand the travel preferences of different groups, optimize the layout of transportation facilities, and formulate more effective transportation policies.

[0210] Step S43: Map the traffic mode selection probability to geospatial grid cells to generate a spatial distribution of traffic demand, and use the spatial distribution of traffic demand as a traffic behavior feature.

[0211] It should be noted that geospatial grid cells refer to the regular or irregular grid units that divide geographic space in a geographic information system. Each grid cell has a clearly defined geographical location and boundaries, and is used to map the probability of mode of transport selection to a specific geographic location, so as to intuitively display the travel demand of different areas on a map. These grid cells can be square, hexagonal, or other shapes, depending on the required accuracy of the analysis and the resolution of the data.

[0212] Spatial distribution of traffic demand refers to the distribution of traffic demand in different areas of a geographic space. It reflects the intensity and distribution characteristics of traffic demand in different areas and is usually displayed in the form of a map, which can intuitively show which areas have high traffic demand and which areas have low traffic demand.

[0213] Understandably, firstly, the traffic source tracing analysis system assigns the calculated traffic mode selection probability to the corresponding grid cell based on the coordinates and extent of the geospatial grid cells in the geographic information system. This process is achieved through spatial interpolation or spatial allocation algorithms, ensuring that the probability value within each grid cell accurately reflects the traffic mode selection tendency of the area. Secondly, the system generates a spatial distribution map of traffic demand by summarizing and analyzing these probability values ​​assigned to the grid cells. This map visually displays the intensity of traffic demand in different areas using color coding or numerical annotation, helping to identify traffic demand hotspots and low-demand areas. Finally, the system integrates this spatial distribution map of traffic demand as a traffic behavior feature into the decision-making process of traffic planning and management, providing detailed geospatial data for optimizing the layout of traffic facilities and formulating traffic policies.

[0214] This embodiment first performs cluster analysis of travel groups based on travel behavior characteristics using a latent category model to obtain group classification results. This step can identify similarities among travelers in terms of mode of transport selection, travel time, and personal attributes, dividing travelers into different groups and providing more detailed user group information for transportation planning. Next, a hybrid Logit model is constructed based on the group classification results, and the probability of mode of transport selection is calculated. The hybrid Logit model considers the heterogeneity and randomness of traveler choice behavior, and can more accurately predict the probability of travelers choosing a certain mode of transport in different situations, providing a scientific basis for the layout and optimization of transportation facilities. Finally, the probability of mode of transport selection is mapped to geospatial grid cells to generate a spatial distribution of traffic demand, and this spatial distribution of traffic demand is used as a traffic behavior feature. This process, through geographic information system technology, combines probability with geospatial location, intuitively displaying the intensity of traffic demand in different areas, helping transportation planners identify traffic demand hotspots, and providing detailed geospatial evidence for the layout and optimization of transportation facilities. The entire process provides comprehensive and scientific support for transportation planning and management.

[0215] This application also provides a traffic big data tracing and analysis device for early-stage highway construction; please refer to [reference needed]. Figure 7 The aforementioned traffic big data tracing and analysis device for early-stage highway construction includes:

[0216] Data acquisition module 10 is used to acquire multi-source traffic data, including highway gantry data, mobile phone signaling data, traffic survey data, geographic information system data, and facility point data;

[0217] The path tracing module 20 is used to perform vehicle travel path tracing analysis on the highway gantry data to obtain path frequency distribution results.

[0218] The feature extraction module 30 is used to obtain travel behavior features based on the mobile phone signaling data, the traffic survey data, the geographic information system data, and the facility point data;

[0219] The feature analysis module 40 is used to analyze the travel behavior features through a latent category model to obtain traffic behavior features;

[0220] Planning module 50 is used to overlay the path frequency distribution results, the traffic behavior characteristics, and land space constraints to generate highway corridor optimization schemes and interchange setting optimization schemes.

[0221] This application provides a traffic big data source tracing and analysis device for the early stages of highway construction. The device includes: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the traffic big data source tracing and analysis method for the early stages of highway construction described in Embodiment 1 above.

[0222] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the traffic big data source tracing and analysis method for the early stages of highway construction as described in the above embodiments.

[0223] The traffic big data source tracing and analysis device, equipment, and computer-readable storage medium for highway pre-planning provided in this application, employing the traffic big data source tracing and analysis method for highway pre-planning in the above embodiments, can solve the technical problem of how to quantitatively analyze traffic behavior characteristics to optimize highway pre-planning decisions. Compared with the prior art, the beneficial effects of the traffic big data source tracing and analysis device, equipment, and computer-readable storage medium for highway pre-planning provided in this application are the same as the beneficial effects of the traffic big data source tracing and analysis method for highway pre-planning provided in the above embodiments, and other technical features are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0224] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for tracing and analyzing traffic big data in the early stages of highway construction, characterized in that, The method includes: Acquire multi-source traffic data, including highway gantry data, toll station entrance and exit data, mobile phone signaling data, traffic survey data, geographic information system data, and facility point data. The traffic survey data refers to traffic flow, vehicle speed, and vehicle type information. Based on the highway gantry data and the toll station entrance and exit data, a vehicle travel path tracing analysis was performed to obtain the path frequency distribution results; Based on the mobile phone signaling data, the traffic survey data, the geographic information system data, and the facility location data, travel behavior characteristics are obtained; The travel behavior characteristics are analyzed using a latent category model to obtain traffic behavior characteristics; By superimposing the path frequency distribution results, the traffic behavior characteristics, and the land space constraints, an optimization scheme for highway corridors and an optimization scheme for interchange settings are generated. The step of analyzing the travel behavior characteristics using a latent category model to obtain traffic behavior characteristics includes: Based on the aforementioned travel behavior characteristics, a latent category model is used to perform travel group clustering analysis to obtain group classification results; A hybrid Logit model is constructed based on the group classification results, and the probability of mode of transportation selection is calculated based on the hybrid Logit model. The probability of choosing a mode of transportation is mapped to a geospatial grid cell to generate a spatial distribution of traffic demand, and the spatial distribution of traffic demand is used as a traffic behavior feature. The steps of superimposing the path frequency distribution results, traffic behavior characteristics, and land space constraints to generate highway corridor optimization schemes and interchange setting optimization schemes include: The path frequency distribution results are converted into a road network traffic heat map, and the spatial distribution of traffic demand in the traffic behavior characteristics is superimposed on the road network traffic heat map to obtain a traffic demand distribution map. A spatial constraint map is constructed based on the land space constraints, and the traffic demand distribution map is weighted and overlaid with the spatial constraint map to generate a comprehensive suitability layer; Based on the comprehensive suitability layer, a minimum cost path analysis is performed to generate a preliminary highway corridor plan. Based on the truck GPS trajectory data and preset time thresholds, isochronous circles are calculated to determine hotspot areas for interchange settings. The preliminary highway corridor plan and the hot spots for interchange placement are assessed for land space coordination, traffic function matching, and project implementation feasibility. Based on the assessment results, optimized highway corridor plans and optimized interchange placement plans are generated. The traffic function matching assessment refers to assessing whether the preliminary highway corridor plan and the hot spots for interchange placement meet traffic demand.

2. The method as described in claim 1, characterized in that, The steps of performing travel group clustering analysis based on the travel behavior characteristics and obtaining group classification results through a latent category model include: Extract travel mode features, travel time features, and personal attribute features from the travel behavior features; The travel mode characteristics, travel time characteristics, and personal attribute characteristics are input into a latent category model for latent variable analysis to obtain a latent variable representation of the features. Clustering and assignment are performed on the latent variable representations of the features to obtain the initial group classification results; The initial group classification results are analyzed using silhouette coefficients to determine the optimal number of classifications. Using the optimal number of classifications as the clustering quantity parameter, latent variable analysis and cluster assignment were performed again to obtain the population classification results.

3. The method as described in claim 2, characterized in that, The step of clustering and assigning the latent variable representations of the features to obtain the initial group classification results includes: A predetermined number of sample points are randomly selected from the latent variable representation of the features as initial cluster centers; Calculate the Euclidean distance from each of the latent feature variables to all the initial cluster centers; Based on the Euclidean distance, each of the latent feature variables is assigned to the cluster represented by the nearest initial cluster center; Calculate the mean of the latent variable representation of the features in each cluster, and update the cluster center position based on the mean; When the difference between the updated cluster center location and the original cluster center location is less than a preset convergence threshold, the initial population classification result is output.

4. The method as described in claim 1, characterized in that, The steps of constructing a hybrid Logit model based on the group classification results and calculating the probability of mode selection based on the hybrid Logit model include: The travel feature set of each group is extracted from the group classification results. The travel feature set includes travel time features, travel cost features, and service level features. An observable utility component is constructed based on the travel time characteristics, travel cost characteristics, and service level characteristics. The utility function of the hybrid Logit model is constructed based on the preset random error component and the observable utility component, thus obtaining the initial hybrid Logit model; The parameters of the initial mixed Logit model are estimated using the maximum likelihood estimation method to obtain the target mixed Logit model; The probability of each group choosing different modes of transportation is calculated based on the target hybrid Logit model and then summarized to obtain the probability distribution of mode of transportation selection.

5. The method as described in claim 1, characterized in that, The step of obtaining travel behavior characteristics based on the mobile phone signaling data, the traffic survey data, the geographic information system data, and the facility location data includes: The trajectory points in the mobile phone signaling data are associated with the travel purpose field in the traffic survey data to generate an initial trajectory chain; Based on the road network topology layer in the geographic information system data, the positioning deviation of the initial trajectory chain is corrected to obtain a reference trajectory chain; The functional area codes in the facility point data are associated with the reference trajectory chain to generate fused trajectory chain data; The fused trajectory chain data is spatiotemporally sliced ​​to obtain travel time period units; Calculate the origin-destination distribution entropy and path fluctuation entropy of the travel time period unit; Travel behavior characteristics are marked based on the range of the origin-destination distribution entropy value and the path fluctuation entropy value.

6. The method as described in claim 1, characterized in that, The step of performing vehicle travel path tracing analysis based on the highway gantry data and the toll station entrance / exit data to obtain the path frequency distribution results includes: Based on the gantry location information and timestamp in the highway gantry data, an initial vehicle trajectory point sequence is generated; Extract the entrance toll station location, entrance time, exit toll station location, and exit time from the toll station entrance and exit data; Using the entrance toll station location and entrance time as the starting point, and the exit toll station location and exit time as the ending point, these are merged with the initial vehicle trajectory point sequence in chronological order to obtain a complete vehicle path sequence. Based on the preset road segment topology network, each location point in the complete vehicle path sequence is mapped to the corresponding road segment to generate a road segment passage sequence; Frequency statistics are performed on the traffic sequences of the road segments that have the same entrance toll station location and the same exit toll station location to obtain frequency statistics results; The frequency statistics results are converted into probability distribution data to obtain the path frequency distribution results.

7. A traffic big data tracing and analysis device for early-stage highway construction, characterized in that, The device includes: The data acquisition module is used to acquire multi-source traffic data, including highway gantry data, mobile phone signaling data, traffic survey data, geographic information system data, and facility point data. The traffic survey data refers to traffic flow, vehicle speed, and vehicle type information. The path tracing module is used to perform vehicle travel path tracing analysis on the highway gantry data to obtain path frequency distribution results. The feature extraction module is used to obtain travel behavior features based on the mobile phone signaling data, the traffic survey data, the geographic information system data, and the facility point data; The feature analysis module is used to analyze the travel behavior features using a latent category model to obtain traffic behavior features. The step of analyzing the travel behavior features using a latent category model to obtain traffic behavior features includes: performing travel group clustering analysis based on the travel behavior features using a latent category model to obtain group classification results; constructing a hybrid Logit model based on the group classification results, and calculating the mode selection probability based on the hybrid Logit model; mapping the mode selection probability to geospatial grid cells to generate a spatial distribution of traffic demand, and using the spatial distribution of traffic demand as traffic behavior features. The planning module is used to overlay the path frequency distribution results, traffic behavior characteristics, and land space constraints to generate optimized highway corridor schemes and optimized interchange design schemes. The steps of overlaying the path frequency distribution results, traffic behavior characteristics, and land space constraints to generate optimized highway corridor schemes and optimized interchange design schemes include: converting the path frequency distribution results into a road network traffic flow heatmap, and overlaying the spatial distribution of traffic demand from the traffic behavior characteristics onto the road network traffic flow heatmap to obtain a traffic demand distribution map; constructing a spatial constraint map based on land space constraints, and then combining the traffic demand distribution map with the spatial constraint map. A weighted overlay is performed to generate a comprehensive suitability layer; based on the comprehensive suitability layer, a minimum cost path analysis is conducted to generate a preliminary highway corridor plan; based on truck GPS trajectory data and preset time thresholds, isochronous circles are calculated to determine hotspot areas for interchange placement; the preliminary highway corridor plan and the hotspot areas for interchange placement are assessed for land space coordination, traffic function matching, and engineering feasibility, and optimized highway corridor and interchange placement plans are generated based on the assessment results. The traffic function matching assessment refers to evaluating whether the preliminary highway corridor plan and the hotspot areas for interchange placement meet traffic demands.

8. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the method for tracing and analyzing traffic big data in the early stages of highway construction as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Freight logistics feature analysis method based on highway networking toll collection data

    CN118536884A

  • Expressway traffic distribution method and system based on neural network

    CN119169826A