Individual activity-trip chain simulation generation method and system based on multi-source data
By integrating mobile phone signaling data and activity log survey data, a dynamic transfer probability matrix is established and the activity chain sequence is generated using the Monte Carlo method, which solves the problems of data source heterogeneity and limited sample size in the prior art, and the accurate simulation generation of individual activity-travel chains is achieved.
Patent Information
- Application Number
- CN202510257609.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-05
AI Technical Summary
It is difficult for the existing technology to effectively integrate mobile phone signaling data and activity log survey data to realize the coordinated construction of individual activity-travel chains, resulting in limited activity type identification, residence time threshold constraints, travel mode identification and track accuracy.
By integrating the mobile phone signaling data and activity log survey data of the time period set in the research area, a dynamic transfer probability matrix was established, and the activity chain sequence was generated using the Monte Carlo method, and typical activity chain patterns were identified through hierarchical clustering and K-means algorithms, and finally a travel chain was constructed and a spatiotemporal trajectory was generated.
The accurate simulation generation of individual activity-travel chains is achieved, the activity type identification, travel mode identification and trajectory accuracy are improved, and the data source heterogeneity and limited sample size in traditional methods are overcome.
Smart Images

Figure CN120218765A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of urban traffic planning and population behavior analysis, and particularly relates to a method and system for simulating and generating individual activity-travel chains based on multi-source data. Background Art
[0002] The research on human activity-travel behavior is of great significance to traffic planning and urban management. In the Activity-based Travel Theory, travel demand is regarded as the spatio-temporal conversion need generated when people participate in different types of activities, which is different from the view of the traditional Four-step Model that simplifies travel as the displacement between the origin and the destination. Under this theoretical framework, an Activity Chain describes the sequential combination of a series of activities carried out by an individual within a specific observation period (such as one day), including information on the time dimension such as activity type, start time, and duration. The corresponding Travel Chain is the manifestation of the Activity Chain in the spatial dimension, recording decision-making elements such as spatial location selection, transportation mode, and route selection during the process of an individual completing the activity sequence.
[0003] Currently, collecting information on residents' activity-travel chains mainly relies on the activity log survey method. Respondents record detailed information on their daily activities and travels through paper questionnaires or online logs, including the type, time, location of activities, and the transportation mode used, etc. This method can obtain relatively complete individual behavior data, but there are also some limitations: the survey sample size is limited, the implementation cost is high, and the data update cycle is long. In addition, due to the complexity of population activity-travel behavior, it is difficult for traditional survey methods to achieve large-scale and long-term continuous observations.
[0004] With the development of mobile Internet technology, mobile signaling data is the communication log recorded between user terminals and base stations in the mobile communication network. Due to its advantages such as wide coverage, high spatio-temporal accuracy, and low acquisition cost, it has been widely used in population mobility research. However, mobile signaling data only contains information on the location changes of user devices and cannot directly identify the activity categories and travel purposes at the stop points. This data characteristic leads to inherent limitations in reconstructing activity-travel chains using mobile signaling data: characteristics such as limited activity type identification, stop duration threshold constraints, transportation mode discrimination, and trajectory accuracy deficiencies result in significant uncertainties in the key element identification link of activity-travel chains.
[0005] By analyzing the characteristics of these two types of data, it can be found that they are complementary: the activity log survey data contains rich behavioral decision-making information and can clearly reflect the activity type and travel purpose; the mobile phone signaling data has advantages in terms of sample size, spatio-temporal continuity, and update timeliness, and can better reflect the movement patterns at the group level. The behavioral patterns in the activity logs can assist in explaining the semantic content of the mobile phone signaling data, while the large-scale spatio-temporal characteristics of the mobile phone signaling data help to verify and expand the behavioral patterns observed in the activity logs.
[0006] In the field of activity-travel chain reconstruction, the existing technical solutions have the following main limitations:
[0007] (1) Although the mobile phone signaling data has advantages in terms of sample size and spatio-temporal continuity, its data characteristics make it impossible to directly obtain key behavioral information such as activity types and travel choices.
[0008] (2) The activity log survey data can provide rich details of behavioral decisions, but due to the limited survey scale and the difficulty of data acquisition, it is difficult to meet the need for continuous observation of a large-scale group.
[0009] Based on the above analysis, how to effectively integrate these two heterogeneous data sources and achieve the collaborative construction of activity chains and travel chains has become a key scientific issue in current research. The core challenges include: how to establish the association mechanism between the two types of data, how to extract and transfer the behavioral patterns in the activity logs, and how to verify the fusion results and construct an evaluation system, etc. Solving these problems not only has important theoretical value, but also has important practical significance for improving the refined level of urban governance and optimizing the supply of public services. Summary of the Invention
[0010] The present invention is made to solve the above problems, and aims to provide an individual activity-travel chain simulation and generation method and system based on multi-source data.
[0011] The present invention provides a method for simulating and generating individual activity-trip chains based on multi-source data, which has the following characteristics: integrating mobile signaling data and activity log survey data in a set time period of the research area, including the following steps: S10, obtaining population types, activity types, and the corresponding basic transition probability matrix, activity rhythm characteristics, activity duration distribution characteristics, time-varying activity participation intensity characteristics, and traffic travel mode characteristics based on the activity log survey data, and obtaining the time start and end intervals, spatio-temporal distribution characteristics, population type and structure proportion, and the in-out space characteristics of the research area based on the mobile signaling data; S20, dynamically modifying the basic transition probability matrix using the time-varying activity participation intensity characteristics to obtain a dynamic transition probability matrix; S30, controlling the overall time range with the time start and end intervals, controlling the overall sample structure with the population type and structure proportion, determining the initial activity state with the time-varying activity participation intensity characteristics, using the activity duration distribution characteristics as the verification constraint condition, extracting the conditional transition probability row vectors from the dynamic transition probability matrix, and then using the Monte Carlo method to generate an activity chain sequence; S40, combining hierarchical clustering and the K-means algorithm to perform pattern recognition and classification on the activity chain sequence to obtain several typical activity chain patterns; S50, based on the spatial positions of the key activity nodes in the typical activity chain patterns, constructing a spatio-temporal distribution probability matrix using the spatio-temporal distribution characteristics and then combining with POI / AOI data to construct a spatial selection probability model, and selecting the spatial positions of all activities of all people in the activity chain to obtain complete activity chain data including spatial position selection; S60, constructing a trip chain according to the spatial selection probability model, and combining with the traffic travel mode characteristics, generating a spatio-temporal trajectory through a path planning service, constructing a trip chain trajectory data set, and realizing time series optimization through an adaptive distribution parameter adjustment and a hierarchical sampling strategy.
[0012] In the method for simulating and generating individual activity-trip chains based on multi-source data provided by the present invention, it may also have the following characteristics: wherein, in step S10, the population types include working population, tourist population, business population, commercial consumption population, and local resident population, and the activity types include life services, cultural and sports leisure, employment work, tourism sightseeing, residential rest, shopping consumption, and dining activities.
[0013] In the method for simulating and generating individual activity-trip chains based on multi-source data provided by the present invention, it may also have the following characteristics: wherein, in step S10, after statistically counting the activity type transfer frequencies in a certain time period using the overlapping sliding window method, the model parameters of the model based on the mixed-state first-order Markov chain model are iteratively optimized through the EM algorithm until convergence to obtain the basic transition probability matrix.
[0014] In the individual activity-travel chain simulation generation method based on multi-source data provided by the present invention, there may also be such a feature: wherein the traffic travel mode characteristics include several travel mode types and their selection probabilities, and the travel mode selection probability is represented by constructing a probability distribution model: Where P m (d) is the probability of choosing travel mode m at distance d, α m is the initial choice tendency of the mode, β m is the distance sensitivity parameter, γ m is the basic selection probability, and d is the travel distance.
[0015] The individual activity-trip chain simulation generation method based on multi-source data provided by the present invention may also have the following features: wherein, the constraint condition is constructed:
[0016]
[0017] For different population types, the probability distribution model parameters are fitted to the survey data using the least squares method, and normalization is used to ensure that the sum of the selection probabilities of different travel modes at any distance is 1.
[0018] The individual activity-travel chain simulation generation method based on multi-source data provided by the present invention may also have the following features: wherein step S20 includes the following sub-steps: S21, correcting the basic transition probability matrix according to the following formula: p k,t =P base I k (t), where P base represents the basic transition probability matrix, I k (t) is the time-varying activity participation intensity feature, k represents the type of people, t represents the time, and p k,t represents the modified transition probability matrix of type k population at time t; S22, for p k,t =P base I k The zero probability term in (t) is set with a lower limit of ε = 10 -6 , and normalize the result to get the dynamic transfer probability matrix.
[0019] In the individual activity-travel chain simulation generation method based on multi-source data provided by the present invention, it can also have the following characteristics: wherein, step S30 includes the following sub-steps: S31, based on the type of crowd and the structural proportion, probability sampling is performed to determine the type of crowd to which the individual to be generated belongs, based on the time start and end interval of the type of crowd, the time boundary of the activity chain is determined, based on the time-varying activity participation intensity characteristics of the selected crowd type at the initial moment, probability sampling is performed to determine the starting activity state; S32, from the dynamic transition probability matrix: Extract the conditional transition probability row vector: P k (t, i) = [p k (t, i, 1), p k (t, i, 2), …, p k (t, i, n)], where p k (t, i, j) represents the probability that the population type k transfers from state i to state j at time t. The conditional transition probability row vector P k (t, i) represents the conditional probability distribution of transferring from the current state i to all possible next states; S33. Based on the conditional transition probability row vector P k (t, i), use the Monte Carlo method to select the activity state at the next moment, and repeat this process until the preset end time is reached to obtain a complete individual activity chain sequence; S34. Construct a two-dimensional verification framework, and evaluate the goodness of fit between the individual activity chain sequence and the activity duration distribution characteristics through the one-sample KS test to ensure the quality requirements and reasonable randomness of the samples.
[0020] In the method for simulating and generating an individual activity-trip chain based on multi-source data provided by the present invention, it may also have the following characteristics: Among them, step S40 includes the following sub-steps: S41. Construct a comprehensive distance metric for the activity chain sequence: d ij = w1·E(S i , S j ) + w2·T(S i , S j ) + w3·L(S i , S j ), where S i represents the i-th activity chain sequence, S j represents the j-th activity chain sequence, d ij represents the comprehensive difference degree between two activity chains (S i , S j ), E(S i , S j ) is the edit distance of the activity type sequence obtained by calculating the minimum number of edit operations required for sequence matching, T(S i , S j ) is the time overlap degree calculated using the Jaccard coefficient, L(S i , S j ) is the duration difference based on the standardized Euclidean distance, and w1, w2, and w3 are weight coefficients; S42. Randomly select several activity chain sequences as samples for different population types, construct a distance matrix, and then use the Ward method for hierarchical clustering. After calculating the silhouette coefficients for different numbers of clusters K, select the K value corresponding to the maximum silhouette coefficient of different population types as the optimal number of clusters. The silhouette coefficient Where a is the average distance between a sample and samples in the same cluster, and b is the average distance between a sample and samples in the nearest neighbor cluster; S43, based on the optimal number of clusters, perform clustering analysis on all active chain sequences using the K-means algorithm to obtain the final K typical active chain patterns.
[0021] In the method for simulating and generating individual activity-trip chains based on multi-source data provided by the present invention, it may further have the following feature: wherein, in step S50, the selection probability of key activity nodes is determined by the product of the activity spatio-temporal distribution probability and the basic attraction of the venue: Where, P t is the activity spatio-temporal distribution probability, R is the venue score, N is the number of reviews, Pr is the average price, and the selection probability of other activities in the typical activity chain pattern is determined by the activity spatio-temporal distribution probability, the basic attraction of the venue, and the spatial distance from adjacent activities: Where, D is the Euclidean straight-line distance between adjacent activities.
[0022] The present invention also provides an individual activity-trip chain simulation and generation system based on multi-source data, which has the following features: it uses the individual activity-trip chain simulation and generation method based on multi-source data of any one of the foregoing, including: a data preprocessing and feature extraction unit, which is used to obtain population types and activity types, as well as the corresponding basic transition probability matrix, activity rhythm characteristics, activity duration distribution characteristics, time-varying activity participation intensity characteristics, and traffic travel mode characteristics based on activity log survey data, obtain the time start and end intervals, spatio-temporal distribution characteristics, population type and structure ratio, and research area access space characteristics based on mobile phone signaling data, and use the time-varying activity participation intensity characteristics to dynamically correct the basic transition probability matrix to obtain a dynamic transition probability matrix; an activity chain generation model, connected to the data preprocessing and feature extraction unit, which is used to control the overall time range with the time start and end intervals, control the overall sample structure with the population type and structure ratio, determine the initial activity state with the time-varying activity participation intensity characteristics, use the activity duration distribution characteristics as the verification constraint condition, extract the conditional transition probability row vector from the dynamic transition probability matrix, and then use the Monte Carlo method to generate an activity chain sequence; a clustering and verification analysis unit, connected to the activity chain generation model, which is used to combine the hierarchical clustering and K-means algorithms to perform pattern recognition and classification on the activity chain sequence to obtain several typical activity chain patterns; a spatial location selection unit, connected to the clustering and verification analysis unit and the data preprocessing and feature extraction unit, based on the spatial locations of the key activity nodes in the typical activity chain patterns, constructs a spatio-temporal distribution probability matrix using the spatio-temporal distribution characteristics and then combines with POI / AOI data to construct a spatial selection probability model, and then selects the spatial locations of all activities of all people in the activity chain to obtain complete activity chain data including spatial location selection; and a trip chain generation unit, connected to the spatial location selection unit and the data preprocessing and feature extraction unit, constructs a trip chain according to the spatial selection probability model, combines with the traffic travel mode characteristics, generates a spatio-temporal trajectory through a path planning service, constructs a trip chain trajectory data set, and realizes time series optimization through an adaptive distribution parameter adjustment and a hierarchical sampling strategy.
[0023] Functions and effects of the invention
[0024] The present invention provides an individual activity-trip chain simulation and generation method and system based on multi-source data. By integrating the advantages of two types of data (mobile phone signaling data + activity log survey data), the activity-trip chain of residents is reconstructed in order to more accurately describe the spatio-temporal behavior patterns of the population and provide a new analysis perspective and technical solution for urban planning and management.
[0025] The present invention realizes the complementary advantages of two types of data by establishing an association mechanism between mobile phone signaling data and activity log data: on the one hand, it uses the behavior patterns in the activity log to identify the activity types and travel purposes in the mobile phone signaling data, and on the other hand, it verifies and expands the behavior rules of the activity log with the large-scale trajectory features of the mobile phone signaling data.
[0026] At the technical implementation level, the present invention solves the following key problems:
[0027] (1) Spatiotemporal alignment of heterogeneous data sources: Construct a spatiotemporal alignment framework to achieve the matching of mobile phone signaling data and activity log data in the time and space dimensions.
[0028] (2) Activity-travel chain generation model: Establish a model based on the fused data to simulate the all-day activities and travel behaviors of individuals on typical working days.
[0029] (3) Data processing and calculation methods: Develop efficient data processing and model calculation methods to ensure the feasibility of reconstructing the behaviors of a large population.
[0030] (4) Model evaluation and verification system: Construct an evaluation system from multiple dimensions such as spatiotemporal distribution characteristics and behavior rule consistency to verify the reliability of the reconstruction results. Description of the Drawings
[0031] Figure 1 is a flowchart of a method for simulating and generating an individual's activity-travel chain based on multi-source data according to an embodiment of the present invention;
[0032] Figure 2 is the stationary time proportion of seven activity types after long-term evolution in step S123 of an embodiment of the present invention;
[0033] Figure 3 is the activity rhythm characteristic of the dining activity in step S131 of an embodiment of the present invention;
[0034] Figure 4 is the activity duration distribution characteristic of the working population in step S143 of an embodiment of the present invention;
[0035] Figure 5 is the time-varying activity participation intensity characteristic of the working population in step S152 of an embodiment of the present invention;
[0036] Figure 6 is a schematic diagram of the probability distribution of traffic mode selection at different travel distances in step S162 of an embodiment of the present invention;
[0037] Figure 7 is a schematic diagram of the overall two-dimensional joint probability distribution of the time start and end intervals in step S171 of an embodiment of the present invention. Detailed implementation manners
[0038] In order to make the technical means, creative features, achieved purposes and functions of the present invention easy to understand, the following embodiments will specifically elaborate on a method and system for simulating and generating individual activity-trip chains based on multi-source data of the present invention in conjunction with the accompanying drawings.
[0039] <Embodiment>
[0040] This embodiment uses mobile signaling data and activity log survey data within a set time period in the research area.
[0041] (1) Mobile signaling data: By recording the information interaction between mobile phone users among different base stations, the spatio-temporal positions of users can be obtained relatively accurately, and then the dynamic flow trajectories of the crowd in the urban space can be restored. This embodiment uses the anonymous mobile signaling data of Shanghai in November 2023 provided by China Unicom's Smart Footprint Platform. This data set contains 8 fields, generating approximately 27.299 million records per day on average, covering 15 million users. At the spatial scale, using the geohash7 grid coding provided by the data platform, the Shanghai urban area is divided into regular grid cells with a side length of about 150 meters, which can support the analysis of crowd dynamics at a relatively high spatial granularity. The information fields of the mobile signaling data are shown in Table 1 below:
[0042] Table 1 (Schematic diagram of information fields of mobile signaling data)
[0043]
[0044] (2) Activity log survey data: Taking the Lujiazui area in Shanghai as an example, this embodiment conducts activity log questionnaire surveys in a combination of online and offline methods, and a total of 389 valid questionnaires are recovered. The survey content covers individual socioeconomic attributes (including age, gender, occupation, family structure, income level, education level, etc.) and complete activity-trip chain information on typical working days. In the activity classification system, residents' daily activities are classified into 11 main categories, including home activities, work, study, shopping, dining, leisure and entertainment, culture and sports, business, etc. For each activity, the respondents are required to record in detail its time attributes (start time, end time, duration, activity planning), spatial attributes (specific location of the activity place), travel characteristics (transportation mode, travel time consumption, approximate path selection), and social attributes (composition of the accompanying crowd). The schematic diagram of individual attributes of the activity log survey data and the schematic diagram of the activity-trip chain of the activity log survey data are shown in Table 2 and Table 3 below respectively:
[0045] Table 2 (Schematic diagram of individual attributes of activity log survey data)
[0046] ID Age Gender Occupation Family Structure Monthly Income (yuan) Educational Attainment 1 28 2 3 A1 20000 Master
[0047] Table 3 (Activity Log Survey Data Activity - Travel Chain Schematic)
[0048]
[0049] Figure 1 is a flowchart of a method for simulating and generating an individual activity - travel chain based on multi - source data according to an embodiment of the present invention. As Figure 1 shown, this embodiment provides a method for simulating and generating an individual activity - travel chain based on multi - source data, which integrates mobile signaling data and activity log survey data for a set time period in the research area, and includes the following steps S10 - S60:
[0050] S10, data pre - processing and feature extraction, including the following sub - steps S11 - S17:
[0051] S11, obtaining population types and activity types based on activity log survey data:
[0052] (1) Standardize the coding of activity types, and divide residents' daily behaviors into seven categories: life services, cultural and sports leisure, employment, tourism, living and resting, shopping and consumption, and dining activities.
[0053] The classification description of activity types is shown in Table 4 below:
[0054] Table 4 (Classification Description of Activity Types)
[0055] Activity Type Mainly Included Activities Life Services (A1) Daily service activities such as haircuts, beauty treatments, laundry, repairs, medical care, finance, and postal services Cultural and Recreational (A2) Leisure activities such as sports and fitness, movie watching, cultural activities, and park recreation Employment and Work (A3) Formal work and learning activities such as office employment, education and training, and business activities Tourism and Sightseeing (A4) Sightseeing and playing activities such as scenic spot visits, cultural experiences, and city explorations Residence and Rest (A5) Daily living and sleeping rest activities at the place of residence or temporary accommodation Shopping and Consumption (A6) Shopping activities at places such as shopping malls, supermarkets, and specialty stores Food and Beverage Activities (A7) Dining and gathering activities at places such as restaurants, fast food restaurants, and cafes
[0056] (2) Divide the population in the research area (Lujiazui area) into five categories: working population, tourism population, business population, commercial consumption population, and local resident population.
[0057] The classification description of population types is shown in Table 5 below:
[0058] Table 5 (Classification Description of Population Types)
[0059] Population Type Definition and Explanation Working Population People who work in fixed office locations within the Lujiazui area Tourist Population Visitors who come mainly for sightseeing Business Population People who come for temporary work such as business meetings, negotiations, and training Commercial Consumption Population Visiting population mainly for consumption activities such as shopping, dining, and leisure entertainment Local Resident Population Resident population with a fixed residence within the Lujiazui area but without a fixed workplace
[0060] Among them, in this embodiment: except for the above five main population types, other types of people (such as handling emergencies, short - term transfers, etc.) account for a very low proportion in actual investigations, and their visiting behaviors are highly sudden and accidental. Therefore, they are not included in the analysis scope of this embodiment.
[0061] S12, constructing a basic transition probability matrix, including the following sub - steps S121 - S123:
[0062] S121. Using the overlapping sliding window method, for the activity features at time t, analyze all 30-minute observation windows containing t: [t - 30 min, t], [t - 25 min, t + 5 min],..., [t - 15 min, t + 15 min],..., [t, t + 30 min]. The windows slide in steps of 5 minutes to form multi-view observation samples with an overlapping rate of 83.3%. Count the transfer frequencies between the seven activity types within each observation window, and take the arithmetic mean of the statistical results of all windows containing time t to establish the transfer count matrix at this moment:
[0063] N = {n ij}
[0064] where n ij represents the average number of transfers from activity i to activity j;
[0065] S122. Based on the mixed-state first-order Markov chain model, set the state space S = {s1, s2,..., s7} corresponding to the seven types of activities, and the transition probability matrix P = {p ij} satisfies ∑p ij = 1. Introduce the EM algorithm for parameter estimation: In the E step, calculate the posterior distribution of the hidden state sequence Z, and in the M step, maximize the expectation to update the parameter θ new , and iterate until convergence to obtain the optimal basic transition probability matrix P base . The basic transition probability matrix P base is shown in Table 6 below. Each row represents the probability of transferring from a certain activity to other activities:
[0066] Table 6 (Schematic of the basic transition probability matrix P base )
[0067]
[0068] S123. Solve the characteristic equation of the obtained basic transition probability matrix P base , and verify the steady-state distribution characteristics of the Markov chain by calculating its largest eigenvalue λ max and the corresponding eigenvector π.
[0069] To verify the effectiveness of the basic transition probability matrix P base , solve its characteristic equation, calculate the largest eigenvalue λ max and its corresponding normalized eigenvector π. The numerical analysis results show that the eigenvalue spectrum and its corresponding eigenvectors satisfy the convergence properties of the Markov chain. Among them, π, as the steady-state probability distribution vector, depicts the proportion of the stationary time of the seven activity types after long-term evolution, which is highly consistent with the original sample distribution, verifying the feasibility and accuracy of the construction of the basic transition matrix, as Figure 2 shown.
[0070] S13. Extract activity rhythm features, including the following sub-steps S131 - S132:
[0071] S131. Based on the 30-minute time unit, count the number of participants in various activities, and calculate the activity participation probability at time point t:
[0072]
[0073] Among them, P i (t) is the participation probability of activity i at time point t, n i (t) is the number of people participating in activity i at time point t, and N(t) is the total number of people at time point t.
[0074] S131. Connect the calculated discrete probability points to construct an initial probability density curve. Apply the cubic spline interpolation method to the discrete probability sequence to construct the final activity rhythm probability density curve as the activity rhythm feature, and the interpolation coefficient is determined by the least squares method to ensure node continuity.
[0075] Specifically, in this step, the activity rhythm features taking the dining activity as an example are as Figure 3 shown.
[0076] S14. Extract activity duration distribution features, including the following sub-steps S141 - S143:
[0077] S141. Based on the start and end time differences of each activity in each sample, obtain the duration of a single activity.
[0078] S142. For each sample, construct three-layer features: 1. Overall duration feature (accumulate the total duration of all activities within the sample); 2. Duration feature by type (count the total duration of different activities according to activity type); 3. Proportion feature by type (calculate the proportion of the duration of each type of activity).
[0079] S143. Aggregate the feature data of the five population type samples respectively, and perform probability density fitting using the Gaussian kernel function. Determine the optimal bandwidth parameter through K-fold cross-validation, and finally obtain the three feature distributions of the overall duration distribution, type duration distribution, and type proportion distribution of various populations in the study area as the activity duration distribution features.
[0080] Specifically, in this step, the activity duration distribution features taking the working population as an example are as Figure 4 shown.
[0081] S15. Based on the activity rhythm features of the seven activity types obtained in step S13 and the activity duration distribution features of the five population types obtained in step S14, construct time-varying activity participation intensity features, including the following sub-steps S151 - S152:
[0082] S151. To achieve the organic combination of activity rhythm characteristics and activity duration distribution characteristics, first, the activity duration distribution characteristics of seven activities for five types of population are optimized and extracted. The optimization goal is to minimize the mean square error between the fixed proportion parameter and the distribution curve:
[0083]
[0084] Subject to the following constraints:
[0085]
[0086] where α ki is the fixed proportion parameter of activity i, is its duration proportion distribution density function. Solve this constrained optimization problem by the Lagrange multiplier method to obtain the benchmark participation intensity vector α ki =(α k1 ,…,α k7 ).
[0087] S152. Combine the benchmark participation intensity with the rhythm probability density curve to construct a time-varying participation intensity function as the time-varying activity participation intensity feature:
[0088]
[0089] where I k,i (t) represents the intensity of type k population participating in activity i at time t, and p i (t) is the rhythm probability density of activity i. The denominator ensures that the sum of all activity intensities at any time is 1, achieving normalization across time and activity types.
[0090] Specifically, in this step, the time-varying activity participation intensity feature taking the working population as an example is as Figure 5 shown.
[0091] S16. Extract the characteristics of transportation modes, including the following steps S161 - S162:
[0092] S161. Based on the detailed transportation information recorded in the activity log survey data, integrate the transportation modes into four categories: driving, bus, cycling, and walking. The integration process follows the criteria of similarity in vehicle characteristics and proximity in usage scenarios, and merges the transportation modes with similar physical characteristics and service functions. The classification descriptions of the four transportation modes are as shown in Table 7 below:
[0093] Table 7 (Classification descriptions of the four transportation modes)
[0094]
[0095] S162. On the basis of the classification in step S161, construct a probability distribution model for the five types of people to choose various transportation modes at different travel distances. This model adopts an exponential decay form:
[0096]
[0097] In the formula, P m (d) is the probability of choosing mode m at distance d, α m is the initial selection tendency of the mode, β m is the distance sensitivity parameter, γ m is the basic selection probability, d is the travel distance,
[0098]
[0099] are the constraint conditions.
[0100] For the five types of people, the parameters of the probability distribution model in this step are fitted to the survey data by the least squares method, and normalization processing is adopted to ensure that the sum of the selection probabilities of the four modes at any distance is 1. As Figure 6 shown, this model obtains the probability distributions of different people choosing transportation modes at various travel distances.
[0101] S17. Extract the time start and end intervals, spatio-temporal distribution characteristics, population type and structure ratio, and the entry and exit space characteristics of the research area based on mobile phone signaling data, including the following sub-steps S171 to S174:
[0102] S171. Extract the time start and end intervals:
[0103] Based on the residence activity sequence of mobile phone signaling data, construct the activity time characteristics of the population in the research area. For each person with a residence activity in the research area, count the start time of the first residence activity and the end time of the last residence activity in the research area on the same day to form a time characteristic key-value pair (t start , t end ). Through the kernel density estimation method, construct a two-dimensional joint probability distribution of the time characteristics for the five different types of people in the research area as the time start and end intervals. The schematic diagram of the overall two-dimensional joint probability distribution of the time start and end intervals is as Figure 7 shown.
[0104] S172. Extract the spatio-temporal distribution characteristics:
[0105] The research area is divided into grids using geohash7 encoding, and the spatio-temporal dimensions of population activities are characterized in combination with 30-minute time units. For each time unit, the number of people engaged in three types of activities, namely staying at home, working, and others, in each grid is counted, and the proportion distribution of the total number of activities in that time unit is calculated. Finally, a feature matrix reflecting the spatial selection probability of various activities in different time periods is constructed as the spatio-temporal distribution feature, which is used to characterize the spatial preferences of the population when engaged in different activities at different times, as shown in Table 8:
[0106] Table 8 (Statistical table of the distribution and proportion of the number of people engaged in activities by time period based on grids)
[0107]
[0108] S173. Extract the proportion of population types and structures:
[0109] Based on the user residence behavior characteristics provided by the mobile signaling data platform, the population in the research area is first divided into three categories: working population (with work activities in the area), resident population (only with residence activities), and other activity population (neither work nor residence activities) according to the determination rules of work and residence activities. After data cleaning and outlier screening, the working population accounts for about 35.9%, the resident population accounts for 21.9%, and the other activity population accounts for 42.2%. Further combined with the activity log survey data, the other activity population is subdivided into tourist population mainly for sightseeing, business population mainly for business negotiation, and commercial consumption population mainly for shopping and dining according to the activity purpose, and finally the proportion of the five types of populations is formed, as shown in Table 9:
[0110] Table 9 (Proportion of the population in Lujiazui area)
[0111] Population Type Proportion Working Population 34.0% Local Residents in Lujiazui 20.8% Tourist Population 20.4% Business Population 15.5% Visiting Commercial Consumption Population 9.3%
[0112] S174. Extract the spatial characteristics of entering and leaving the research area:
[0113] Analyze the spatio-temporal distribution characteristics of users entering and leaving the research area. For each user trajectory, first identify the moment when it crosses the boundary of the research area, and determine the arrival time of entering the research area and the departure time of leaving the research area. At the same time, extract the location information (grid code) and activity type attributes of the last activity point before entering the area and the first activity point after leaving the area. The grid information before arrival and after departure may be the actual location code or "0", where "0" indicates that the user is only active within the research area during that time period.
[0114] Furthermore, the distribution probabilities of five groups of people in different combinations of entry and exit spaces are counted. A space combination consists of four dimensions: the location before arrival - activity type and the location after departure - activity type. Through attributes such as the grid before arrival, the activity type before arrival, the arrival time, the grid after departure, the activity type after departure, and the departure time, the sample quantity of each combination and its probability among this group of people are completely recorded, as shown in Table 10:
[0115] Table 10 (Example of Probabilities of Regional Entry and Exit Space Combinations)
[0116]
[0117] For the need of modeling time - varying activity selection, there are limitations in using a fixed basic transition probability matrix. First, there are significant differences in the activity preferences and transfer characteristics of people in different time periods; second, the activity participation patterns of various groups of people change dynamically over time. If the transition probabilities are directly calculated by time period based on survey data, the reliability of the calculation results will be insufficient due to sparse samples. To solve this problem, the basic transition probability matrix is dynamically corrected through the following steps S20.
[0118] S20. Construct a dynamic transition probability matrix, including the following sub - steps S21 - S22:
[0119] S21. Correct the basic transition probability matrix according to the following formula:
[0120] p k,t =P base ·I k (t)
[0121] In the formula, p k,t represents the corrected transition probability matrix, P base represents the basic transition probability matrix, k represents the type of people, t represents the time, p k,t represents the corrected transition probability matrix of people of type k at time t, and I k (t) is the time - varying activity participation intensity characteristic (the probability vector of people of type k participating in various activities at time t).
[0122] This operation realizes the dynamic adjustment of the transition probability by multiplying the basic transition probability by the real - time activity participation probability.
[0123] S22. Considering the numerical calculation stability, set a lower limit ε = 10 -6 for the zero - probability terms in the product result. Subsequently, to maintain the validity of the transition probability, normalize the processed result row - by - row so that the sum of probabilities in each row is 1, and finally obtain the dynamic transition probability matrix of various groups of people in different time periods, as shown in Table 11 below.
[0124] Table 11 (Dynamic Matrix of Activity Transfer Probabilities for Lujiazui Residents at 8:00)
[0125]
[0126] S30. Generate an activity chain sequence, including the following sub-steps S31 to S34:
[0127] S31. Conduct probability sampling based on the population type and its structural proportion to determine the population type to which the individual to be generated belongs; the time start and end intervals of this type of population determine the time boundary of the activity chain; conduct probability sampling based on the time-varying activity participation intensity characteristics of the selected population type at the initial moment to determine the starting activity state.
[0128] S32. The sequence generation uses a discrete-time simulation method with a 30-minute time step, and realizes the dynamic evolution of the activity state through iterative update. At each time step t, according to the current population type k and time point t, extract the conditional transfer probability row vector from the dynamic transfer probability matrix: P k (t,i) = [p k (t,i,1), p k (t,i,2), …, p k (t,i,n)].
[0129] Among them, p k (t,i,j) represents the probability that the population type k transfers from state i to state j at time t, and the conditional transfer probability row vector P k (t,i) represents the conditional probability distribution of transferring from the current state i to all possible next states.
[0130] S33. Based on the conditional transfer probability row vector P k (t,i), use the Monte Carlo method to select the activity state at the next moment, and repeat this process until the preset end time is reached to obtain a complete individual activity chain sequence.
[0131] In the above steps S31 to S33, the generation of the activity chain uses a probability sequence generation method based on the Markov chain. This method assumes that the activity state of an individual at any moment is only related to the current state and time, and characterizes the state transfer characteristics through the dynamic transfer probability matrix.
[0132] S34. Result verification and optimization:
[0133] Construct a two-dimensional verification framework, and evaluate the fitting degree between the individual activity chain sequence and the classified duration features and classified proportion features extracted by S14 through the one-sample KS test. Calculate the D statistic and the corresponding p-value for each dimension, and take the average value of the p-values of various activities as the goodness-of-fit index F1 for the classified duration feature dimension and the goodness-of-fit index F2 for the classified proportion feature respectively. The evaluation indicators of the two dimensions are linearly combined with a weight of 0.5:0.5 and normalized to obtain the comprehensive confidence level C I 。
[0134] To ensure the generation quality and maintain appropriate random diversity, a strategy combining hard threshold constraint and probability screening is adopted: when C I <0.2, directly filter the samples; when C I ≥0.2, randomly screen with a probability p f litered =(1 - C I ). This dual screening mechanism not only ensures the basic quality requirements of the samples but also maintains reasonable randomness within the high-confidence range.
[0135] An example of the screening results of individual activity chains based on the two-dimensional verification framework in this step is shown in Table 12 below:
[0136] Table 12 (Example of Screening Results of Individual Activity Chains Based on the Two-Dimensional Verification Framework)
[0137] Individual ID Population Type Credibility Serial Number Time Period Activity Type Duration (h) 7440 Tourist Population 0.46 1 10:00-11:00 A4 (Tourism and Sightseeing) 1.0 7440 Tourist Population 0.46 2 11:00-12:00 A7 (Food and Beverage Activities) 1.0 7440 Tourist Population 0.46 3 12:00-16:30 A4 (Tourism and Sightseeing) 4.5
[0138] In the above steps S31 - S34, an activity chain generation model based on dynamic transition probability is provided. This model takes the dynamically adjusted activity transition probability matrix as the core, combines the two-dimensional joint distribution characteristics of activity start and end times, the time-varying activity participation intensity characteristics, and the population composition ratio to realize the simulation generation from macroscopic statistical laws to individual microscopic activity chains. The model adopts a probability-based sequence generation method and ensures the reliability of the generation results through multi-level verification mechanisms such as the total activity duration distribution, the classified activity duration distribution, and the activity type proportion distribution.
[0139] The model includes two types of conditions: generation constraint conditions and verification constraint conditions.
[0140] The generation constraint conditions include four aspects:
[0141] (1) Mainly driven by the dynamic transition probability matrix obtained in step S20, which depicts the activity transition probability characteristics of different populations in different time periods.
[0142] (2) Constraining the overall time range of the activity chain with the two-dimensional joint probability distribution of time start and end in the time start and end interval extracted in step S171.
[0143] (3) Control the overall sample structure with the population type and its structural proportion extracted in step S173.
[0144] (3) Determine the initial activity state with the time-varying activity participation intensity feature constructed in step S15.
[0145] The verification constraints are based on the three-layer distribution features extracted in step S142: 1. The overall duration feature; 2. The duration feature by type; 3. The proportion feature by type. They are used to evaluate the reliability of the generated results. These probability constraints jointly constitute the complete boundary conditions for the generation and verification of individual activity chains.
[0146] S40. Obtain the typical activity chain patterns, including the following sub-steps S41 to S43:
[0147] S41. Construct the comprehensive distance metric for the activity chain sequence:
[0148] d ij = w1·E(S i , S j ) + w2·T(S i , S j ) + w3·L(S i , S j )
[0149] In the formula, S i represents the i-th activity chain sequence, S j represents the j-th activity chain sequence, and d ij represents the comprehensive difference degree between two activity chains (S i , S j ). E(S i , S j ) is the edit distance of the activity type sequence obtained by calculating the minimum number of edit operations required for sequence matching. T(S i , S j ) is the time overlap degree calculated using the Jaccard coefficient. L(S i , S j ) is the duration difference based on the standardized Euclidean distance. w1, w2, and w3 are weight coefficients. Specifically, in this embodiment, w1 = 0.4, w2 = w3 = 0.3.
[0150] S42. Randomly select 200 activity chain sequences for each of the 5 population types as samples. After constructing the distance matrix, apply the Ward method for hierarchical clustering. Calculate the silhouette coefficient for different numbers of clusters K, and then select the K value corresponding to the maximum silhouette coefficient of the 5 population types as the optimal number of clusters. Silhouette coefficient:
[0151]
[0152] Wherein, a is the average distance between the sample and the samples in the same cluster, and b is the average distance between the sample and the samples in the nearest neighbor cluster.
[0153] S43. Based on the optimal number of clusters K determined for the five types of people, the K-means algorithm is used to perform clustering analysis on all 10,000 activity chain sequences. In the specific implementation process, first, randomly select K activity chains from the dataset as the initial clustering centers. In the iterative optimization stage, calculate the comprehensive distance from each activity chain to each clustering center, and divide it into the cluster with the minimum distance. After each round of iteration, update the clustering centers by calculating the average features of all activity chains in each cluster, including the mode of the activity type sequence, the average time overlap degree, and the average duration. When the change in the cluster division results of two consecutive iterations is less than the preset threshold, it is considered that the algorithm converges, and finally K typical activity chain patterns are obtained.
[0154] Specifically, in this embodiment, after obtaining the final K typical activity chain patterns, clustering result verification is also performed:
[0155] Compare the typical activity chain patterns obtained by clustering with the activity log survey data obtained in step S10, and combine expert evaluation and field investigation to verify its rationality. The verification results show that the typical activity chain patterns obtained by clustering have a high consistency with the actual samples, conform to the field observations, and can better reflect the daily activity patterns of residents. The clustering results of the typical activity chains of the working population are shown in Table 13 below:
[0156] Table 13 (Schematic diagram of the clustering results of the typical activity chains of the working population)
[0157]
[0158] In the above steps S41 to S43, for the generated 10,000 activity chain sequence data, a clustering analysis framework based on the edit distance is constructed, and a two-level cross-validation mechanism is designed. By calculating the comprehensive distance in three dimensions of activity type conversion, time overlap degree, and duration, the activity chains are clustered and grouped to identify typical behavior patterns. At the same time, combined with the actual sample comparison and expert evaluation methods, a complete evaluation index system is established to ensure the reliability and practical value of the clustering results.
[0159] S50. Construct complete activity chain data including spatial location selection, including the following sub-steps S51 to S53:
[0160] S51. Determine the spatial locations of the most restrictive key activities in the activity chain. The spatial choices of these key activities will guide and restrict the spatial distributions of other activities. The activity chains of five typical population types in the study area all have their specific key activity nodes, as shown in Table 14 below. Based on the spatial locations of these key nodes, combined with the principle of minimum spatial resistance and facility attractiveness, construct a spatial selection probability distribution model for other activities.
[0161] Table 14 (Characteristics of Key Activity Nodes of Five Population Types)
[0162]
[0163] S52. Using the spatio-temporal distribution characteristic data obtained in Step S172, take the geohash7 encoding (spatial accuracy of about 153 meters) as the spatial unit and 30 minutes as the time unit to construct a spatio-temporal distribution probability matrix for work activities, residential activities and other activities in the study area.
[0164] Combined with the POI point data (including attribute information such as ratings, review numbers, prices, etc.) and AOI polygon geographic element data in the study area, construct a spatial selection probability model.
[0165] For the key activity nodes in the activity chain, their selection probabilities are determined by the product of the activity spatio-temporal distribution probability and the basic attractiveness of the venue:
[0166]
[0167] where P t is the activity spatio-temporal distribution probability, R is the venue rating, N is the number of reviews, and Pr is the average price.
[0168] For other activities in the activity chain, the spatial distance to adjacent activities needs to be additionally considered:
[0169]
[0170] where D is the Euclidean straight-line distance to adjacent activities. After determining the selection probability of the spatial grid based on the time and type of the activity, use the Huff gravity model to select specific POI points within the specific grid.
[0171] S53. Select the spatial locations of all activities of all populations in the activity chain. For each activity, according to its time, type, and the characteristics of the previous activity location, etc., use the spatial selection probability model in Step S52 to calculate the selection probabilities of each candidate location. After normalizing the probabilities, use the Monte Carlo method for random sampling to determine the specific spatial location of the activity, and finally obtain the complete activity chain data including the spatial location selection.
[0172] The schematic results of the selection of the spatial position of the activity chain in this step are shown in Table 15:
[0173] Table 15 (Schematic of the selection results of the spatial position of the activity chain)
[0174]
[0175] S60. Construct a travel chain trajectory dataset, including the following sub-steps S61 to S64:
[0176] S61. Travel chain initialization:
[0177] Construct a complete travel chain based on the activity chain data with spatial positions. For each activity chain, first obtain the start time of its first activity and the end time of its last activity as the arrival time and departure time. According to the population type to which the user belongs, based on the spatial characteristics of the access and egress spaces of the research area extracted in step S174 and the data in Table 10, randomly match a set of combinations of pre-arrival grid - activity type and post-departure grid - activity type through normalized probabilities.
[0178] When the pre-arrival grid identifier is 0, it means that the user's initial activity is within the research area. At this time, set the position of the first activity of the activity chain as the starting point of the travel chain; otherwise, use the pre-arrival grid as the starting point of the travel chain. Similarly, when the post-departure grid identifier is 0, set the position of the last activity of the activity chain as the end point of the travel chain; otherwise, use the post-departure grid as the end point of the travel chain. For the grid spatial unit, determine a set of longitude and latitude coordinates as spatial position information according to the activity type from the corresponding POI points according to the probability distribution. After determining the start and end points through the above process, construct the OD pairs and their departure times between adjacent activity points in the travel chain.
[0179] S62. Mode of transportation selection:
[0180] Calculate the straight-line distance of each OD pair in the travel chain based on the transportation mode characteristics extracted in step S16, and combine the population type to which the user belongs to determine the probability of mode of transportation selection for each segment of the travel. Use the Monte Carlo sampling method to determine the specific mode of transportation adopted for each segment of the travel from the probability distribution of each mode of transportation.
[0181] S63. Route planning and trajectory generation:
[0182] Based on the route planning service provided by the AutoNavi Map Open Platform, the generation of routes for various transportation modes is realized. When making API calls, core parameters such as the starting and ending longitude and latitude coordinates and route planning strategies need to be configured, and route queries for travel modes such as walking, cycling, driving, and public transportation are completed through different interfaces. The API response data contains basic information such as the sequence of route points, travel distance, and travel time, as well as data fields specific to the transportation mode: driving includes real-time traffic conditions and tolls, while public transportation includes transfer plans and fare information. The returned data is parsed and standardized to generate a unified format data set containing spatio-temporal trajectories, travel costs, and traffic conditions. The data for each travel segment is integrated in chronological order to construct a complete travel chain trajectory data set.
[0183] An example of the travel chain trajectory data set in this step is shown in Table 16 below:
[0184] Table 16 (Example of travel chain trajectory data)
[0185]
[0186] S64, correction:
[0187] This step provides a method for correcting time offsets based on activity planning. Through adaptive distribution parameter adjustment and hierarchical sampling strategies, the temporal optimization of the activity chain is achieved. This step establishes a quantitative mapping relationship between the planning level and time constraints, and introduces an asymmetric probability distribution model to describe the characteristics of time offsets, thus ensuring the rationality and effectiveness of the correction results.
[0188] In terms of time constraint modeling, the activity planning level k (1 - 5) is mapped to the maximum acceptable time offset T(k) = 35 - 6k minutes. This mapping relationship enables activities with the strongest planning (k = 5) to allow only ±5 minutes of offset, while activities with the weakest planning (k = 1) can accept ±29 minutes of offset, effectively characterizing the differences in time elasticity of different types of activities. Based on this, the Beta(α,β) distribution is used to describe the characteristics of time offsets, where α = k and β = 6 - k, enabling the distribution shape to be dynamically adjusted with the planning level. High-planning activities exhibit obvious skewness characteristics, while low-planning activities tend to approach a symmetric distribution.
[0189] The time correction process is implemented using a hierarchical acceptance-rejection sampling strategy. First, the value range [-T(k), T(k)] of the Beta distribution corresponding to each activity is normalized to the interval [0, 1], and then time offsets are generated for non-head and non-tail activity pairs in the travel chain. For adjacent activity pairs (Ai, Ai+1) with overlapping time periods, the system needs to reserve a time window for travel activities through time offset. Specifically, the system generates time offsets for both activities on both sides simultaneously, creating a necessary time interval through two-way expansion. When the initial sampling fails to meet the travel time requirements, the system compresses the distribution ranges of the two activities through parameter transformation α′ = 1.2α, β′ = 1.2β, and continues joint sampling. If a satisfactory result is not obtained after reaching the maximum number of attempts (20 times), the maximum acceptable time offset T(k) is expanded by 1.5 times and the above process is repeated. An example of correction through this step is shown in Table 17 below:
[0190] Table 17 (Example of Time Correction for Activity-Travel Chain on a Sample Weekday)
[0191]
[0192] This embodiment also provides an individual activity-travel chain simulation generation system based on multi-source data, which uses the individual activity-travel chain simulation generation method in this embodiment, including a data preprocessing and feature extraction unit, an activity chain generation model, a clustering and verification analysis unit, a spatial location selection unit, and a travel chain generation unit.
[0193] The data preprocessing and feature extraction unit is used to obtain the population type, activity type, and the corresponding basic transition probability matrix, activity rhythm characteristics, activity duration distribution characteristics, time-varying activity participation intensity characteristics, and traffic travel mode characteristics based on the activity log survey data according to the methods in steps S10 to S20. Based on the mobile phone signaling data, it obtains the time start and end intervals, spatio-temporal distribution characteristics, population type and structure ratio, and the characteristics of the access space in the research area, and dynamically corrects the basic transition probability matrix using the time-varying activity participation intensity characteristics to obtain the dynamic transition probability matrix.
[0194] The activity chain generation model is connected to the data preprocessing and feature extraction unit, and is used to control the overall time range with the time start and end intervals, control the overall sample structure with the population type and structure ratio, determine the initial activity state with the time-varying activity participation intensity characteristics, and use the activity duration distribution characteristics as the verification constraint condition. After extracting the conditional transition probability row vector from the dynamic transition probability matrix, it uses the Monte Carlo method to generate the activity chain sequence.
[0195] The clustering and verification analysis department is connected to the activity chain generation model and is used to perform pattern recognition and classification on the activity chain sequence according to the method in step S40 by combining hierarchical clustering and the K-means algorithm to obtain several typical activity chain patterns.
[0196] The spatial location selection department is connected to the clustering and verification analysis department and the data preprocessing and feature extraction department. It is used to construct a spatio-temporal distribution probability matrix based on the spatial locations of key activity nodes in the typical activity chain patterns according to the method in step S50, and then combine POI / AOI data to construct a spatial selection probability model. After that, it selects the spatial locations of all activities of all people in the activity chain to obtain the complete activity chain data including spatial location selection.
[0197] The travel chain generation department is connected to the spatial location selection department and the data preprocessing and feature extraction department. It is used to construct a travel chain according to the spatial selection probability model according to the method in step S60, and combine the characteristics of traffic travel modes. Through the path planning service, it generates a spatio-temporal trajectory, constructs a travel chain trajectory data set, and realizes temporal optimization through an adaptive distribution parameter adjustment and a hierarchical sampling strategy.
[0198] Functions and effects of the embodiment
[0199] This embodiment has achieved a key breakthrough in data fusion: Existing research mostly relies on a single data source, each with limitations. Although mobile phone signaling data has complete spatio-temporal trajectory information, it is difficult to identify specific activity types and travel purposes. Activity log surveys contain detailed behavior information, but the sample size is limited and there may be memory biases. This embodiment realizes the effective fusion of the two types of data by constructing a multi-dimensional feature system. While maintaining the spatio-temporal representativeness of mobile phone signaling data, it incorporates the behavior pattern features of activity logs, thus more accurately describing the activity patterns of urban populations.
[0200] This embodiment overcomes the technical difficulties of multi-source data fusion: Aiming at the differences in spatio-temporal accuracy, sampling frequency, and information dimensions of different data sources, this embodiment establishes a systematic data fusion framework. It combines microscopic behavior patterns with macroscopic spatio-temporal distributions through a dynamic transition probability matrix and constructs a multi-level verification mechanism to ensure the reliability of the fusion results, providing a new idea for the integration of multi-source heterogeneous data in urban research.
[0201] This embodiment proposes a more practical modeling scheme: Compared with traditional methods that rely on simple rule matching or random sampling, the model framework based on dynamic transition probability provided by this embodiment can better reflect the spatio-temporal characteristics and behavior constraints of crowd activities by introducing time-varying activity participation intensities and key node guidance strategies, improving the accuracy of simulation results.
[0202] This embodiment has practical application values in many aspects: (1) In urban planning, this embodiment can analyze the spatio-temporal activity patterns of different types of people, provide a basis for optimizing the layout of public service facilities and improving the land use function structure, and help identify spatial function mismatches and insufficient facility supply during the urban renewal process; (2) In traffic management, through the dynamic analysis of population activity characteristics, traffic demand changes can be predicted, providing a reference for optimizing bus line networks, planning parking facilities, and evaluating the effectiveness of traffic policies; (3) In the construction of smart cities, this embodiment realizes the dynamic update of activity-travel characteristics, can provide data support for urban management platforms, and helps optimize resource allocation and emergency response mechanisms; (4) In terms of urban resilience, this method can simulate the changes in population activities under special circumstances, evaluate the carrying capacity of the urban system, and provide a reference for formulating emergency plans.
[0203] Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for simulating and generating individual activity-travel chains based on multi-source data, characterized in that: The mobile phone signaling data and activity log survey data for a set period of time in the study area are integrated, including the following steps: S10, based on the activity log survey data, obtaining the population type and activity type and the basic transition probability matrix corresponding to the two, activity rhythm characteristics, activity duration distribution characteristics, time-varying activity participation intensity characteristics and transportation mode characteristics; based on the mobile phone signaling data, obtaining the time start and end interval, spatiotemporal distribution characteristics, population type and structure proportion, and study area entry and exit space characteristics; S20, dynamically modifying the basic transition probability matrix using the time-varying activity participation intensity feature to obtain a dynamic transition probability matrix; S30, controlling the overall time range with the time start and end intervals, controlling the overall sample structure with the population type and structure proportion, determining the initial activity state with the time-varying activity participation intensity characteristics, using the activity duration distribution characteristics as verification constraints, extracting the conditional transition probability row vector from the dynamic transition probability matrix, and generating an activity chain sequence using the Monte Carlo method; S40, combining hierarchical clustering and K-means algorithm, performing pattern recognition and classification on the activity chain sequence to obtain several typical activity chain patterns; S50, based on the spatial positions of the key activity nodes in the typical activity chain pattern, after constructing a spatiotemporal distribution probability matrix using the spatiotemporal distribution characteristics and combining the POI / AOI data to construct a spatial selection probability model, all activities of all people in the activity chain are selected in spatial positions to obtain complete activity chain data including spatial position selection; S60, constructing a travel chain according to the spatial selection probability model, and generating a spatiotemporal trajectory through a path planning service in combination with the characteristics of the transportation mode, constructing a travel chain trajectory data set, and realizing timing optimization through adaptive distribution parameter adjustment and hierarchical sampling strategy.
2. The method for simulating and generating individual activity-travel chains based on multi-source data according to claim 1, characterized in that: in, In step S10, the types of people include working people, tourists, business people, commercial consumers and local residents. The types of activities include life services, cultural and sports leisure, employment, tourism, residential leisure, shopping and consumption, and catering activities.
3. The method for simulating and generating individual activity-travel chains based on multi-source data according to claim 1, characterized in that: in, In step S10, after the overlapping sliding window method is used to count the activity type transfer frequency in a certain time period, the model parameters based on the mixed state first-order Markov chain model are iteratively optimized by the EM algorithm until convergence to obtain the basic transfer probability matrix.
4. The method for simulating and generating individual activity-trip chains based on multi-source data according to claim 1, characterized in that: in, In step S10, the traffic mode characteristics include several types of travel modes and their selection probabilities. The probability of choosing a travel mode is represented by constructing a probability distribution model: Where P m (d) is the probability of choosing travel mode m at distance d, α m is the initial choice tendency of the mode, β m is the distance sensitivity parameter, γ m is the basic selection probability, and d is the travel distance.
5. The method for simulating and generating individual activity-trip chains based on multi-source data according to claim 4, characterized in that: in, Build constraints: For different types of people, the probability distribution model parameters are fitted to the survey data by the least squares method, and normalization is used to ensure that the sum of the selection probabilities of different travel modes at any distance is 1.
6. The method for simulating and generating individual activity-trip chains based on multi-source data according to claim 1, characterized in that: in, Step S20 includes the following sub-steps: S21, modifying the basic transition probability matrix according to the following formula: p k,t =P base ·I k (t) Where P base represents the basic transition probability matrix, I k (t) is the time-varying activity participation intensity feature, k represents the type of people, t represents the time, and p k,t Represents the modified transition probability matrix of type k population at time t; S22, for p k,t =P base I k The zero probability term in (t) is set with a lower limit of ε = 10 -6 , and normalize the result to obtain the dynamic transfer probability matrix.
7. The method for simulating and generating individual activity-trip chains based on multi-source data according to claim 1, characterized in that: in, Step S30 includes the following sub-steps: S31, performing probability sampling based on the population type and structure proportion to determine the population type to which the individual to be generated belongs, determining the time boundary of the activity chain based on the time start and end interval of the population of this type, and performing probability sampling based on the time-varying activity participation intensity characteristics of the selected population type at the initial moment to determine the starting activity state; S32, from the dynamic transition probability matrix: Extract the conditional transition probability row vector from: P k (t,i)=[p k (t,i,1),p k (t,i,2),…,p k (t,i,n)], Among them, p k (t,i,j) represents the probability of population type k transferring from state i to state j at time t, and the conditional transfer probability row vector P k (t,i) represents the conditional probability distribution of transitioning from the current state i to all possible next states; S33, based on the conditional transition probability row vector P k (t,i), the Monte Carlo method is used to select the activity state at the next moment, and the process is repeated until the preset end time is reached to obtain a complete individual activity chain sequence; S34, construct a two-dimensional verification framework, and evaluate the fit between the individual activity chain sequence and the activity duration distribution characteristics through a single-sample KS test to ensure the quality requirements and reasonable randomness of the sample.
8. The method for simulating and generating individual activity-trip chains based on multi-source data according to claim 1, characterized in that: in, Step S40 includes the following sub-steps: S41, constructing a comprehensive distance metric of the activity chain sequence: d ij =w1·E(S i ,S j )+w2·T(S i ,S j )+w3·L(S i ,S j ) In the formula, S i represents the sequence of the i-th activity chain, S j represents the jth activity chain sequence, d ij Represents two active chains (S i ,S j ), E(S i ,S j ) is the edit distance of the activity type sequence obtained by calculating the minimum number of edit operations required for sequence matching, T(S i ,S j ) is the time overlap calculated using the Jaccard coefficient, L(S i ,S j ) is the duration difference based on the standardized Euclidean distance, w1, w2, w3 are weight coefficients; S42, randomly select several activity chain sequences as samples for different population types, construct a distance matrix and apply Ward method to perform hierarchical clustering, calculate the silhouette coefficient under different cluster numbers K, and then select the K value corresponding to the maximum value of the silhouette coefficient of different population types as the optimal cluster number. Where a is the average distance between the sample and the samples in the same cluster, and b is the average distance between the sample and the samples in the nearest neighbor cluster; S43, based on the optimal number of clusters, using a K-means algorithm to perform cluster analysis on all activity chain sequences to obtain final K typical activity chain patterns.
9. The method for simulating and generating individual activity-trip chains based on multi-source data according to claim 1, characterized in that: in, In step S50, the selection probability of the key activity node is determined by the product of the activity spatiotemporal distribution probability and the basic attractiveness of the venue: Among them, P t is the probability of activity spatiotemporal distribution, R is the venue rating, N is the number of reviews, Pr is the average price, The selection probability of other activities in the typical activity chain pattern is determined by the spatiotemporal distribution probability of the activity, the basic attractiveness of the venue, and the spatial distance to adjacent activities: Where D is the Euclidean straight-line distance between adjacent activities.
10. An individual activity-travel chain simulation generation system based on multi-source data, characterized in that: The method for simulating and generating an individual activity-travel chain based on multi-source data as described in any one of claims 1 to 9 comprises: A data preprocessing and feature extraction unit is used to obtain the population type and activity type and the basic transition probability matrix corresponding to the two, activity rhythm characteristics, activity duration distribution characteristics, time-varying activity participation intensity characteristics and transportation mode characteristics based on the activity log survey data, and to obtain the time start and end intervals, spatiotemporal distribution characteristics, population type and structure proportions, and study area entry and exit space characteristics based on the mobile phone signaling data, and use the time-varying activity participation intensity characteristics to dynamically correct the basic transition probability matrix to obtain a dynamic transition probability matrix; an activity chain generation model, connected to the data preprocessing and feature extraction unit, for controlling the overall time range with the time start and end intervals, controlling the overall sample structure with the population type and structure proportion, determining the initial activity state with the time-varying activity participation intensity characteristics, using the activity duration distribution characteristics as verification constraints, extracting the conditional transition probability row vector from the dynamic transition probability matrix, and then using the Monte Carlo method to generate an activity chain sequence; A clustering and verification analysis unit is connected to the activity chain generation model and is used to combine hierarchical clustering and K-means algorithm to perform pattern recognition and classification on the activity chain sequence to obtain several typical activity chain patterns; A spatial position selection unit is connected to the clustering and verification analysis unit and the data preprocessing and feature extraction unit, and based on the spatial positions of the key activity nodes in the typical activity chain pattern, after constructing a spatiotemporal distribution probability matrix using the spatiotemporal distribution features and combining the POI / AOI data to construct a spatial selection probability model, selects the spatial positions of all activities of all people in the activity chain to obtain complete activity chain data including spatial position selection; and The travel chain generation unit is connected to the spatial position selection unit and the data preprocessing and feature extraction unit, constructs a travel chain according to the spatial selection probability model, and generates a spatiotemporal trajectory through a path planning service in combination with the characteristics of the transportation mode, constructs a travel chain trajectory data set, and realizes timing optimization through adaptive distribution parameter adjustment and hierarchical sampling strategy.
Citation Information
Patent Citations
Group activity data collection method and system based on multisource space-time trajectory data
CN106211071A
Multimode traffic distribution model construction method based on mobile phone signaling data
CN112133090A
Active chain reconstruction method and system combined with multi-source spatio-temporal data
CN114595300A
Visual analysis method for exploring dynamic division of urban functional areas based on semantic fusion model
CN116227791A
Multi-agent-based space-time activity analysis and prediction method for urban empty-nest elderly
CN119514872A