A Method and System for Simulating Individual Activities and Travel Chains Based on Multi-Source Data
By integrating mobile signaling data and activity log data, and using the Monte Carlo method and clustering algorithm to generate individual activity-travel chains, the data fusion problem in existing technologies is solved, and accurate simulation of individual activity-travel chains is achieved, providing a more refined analysis tool for urban planning.
Patent Information
- Application Number
- CN202510257609.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-03-05
AI Technical Summary
Existing technologies struggle to effectively integrate mobile signaling data and activity log survey data to achieve large-scale, long-term reconstruction of individual activity-travel chains, resulting in limited activity type identification and insufficient travel mode discrimination, which fails to meet the needs of urban planning and management.
By establishing a correlation mechanism between mobile signaling data and activity log data, integrating multi-source data, generating activity chain sequences using the Monte Carlo method, and combining hierarchical clustering and K-means algorithm for pattern recognition, a spatiotemporal distribution probability model is constructed to generate complete activity-travel chain data.
It achieves accurate simulation of individual activity-travel chains, provides a more refined perspective for urban planning and management analysis, and enhances the ability to describe the spatiotemporal behavior patterns of data.
Smart Images

Figure CN120218765B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of urban traffic planning and crowd behavior analysis, specifically involving a method and system for simulating and generating individual activity-travel chains based on multi-source data. Background Technology
[0002] Research on human activities and travel behavior is of great significance to transportation planning and urban management. In Activity-based Travel Theory, travel demand is considered as the spatiotemporal transformation needs arising from people participating in different types of activities. This differs from the traditional four-step model, which simplifies travel as a mere displacement between origin and destination. Within this theoretical framework, an activity chain describes the temporal sequence of activities performed by an individual within a specific observation period (e.g., a day), including temporal information such as activity type, start time, and duration. The corresponding travel chain is the spatial manifestation of the activity chain, recording the decision-making elements such as spatial location selection, mode of transportation, and route selection during the completion of the activity sequence.
[0003] Currently, collecting information on residents' activity-travel chains primarily relies on activity log surveys. Respondents record detailed information about their daily activities and travel through paper questionnaires or online logs, including the type, time, location, and mode of transportation used. This method can obtain relatively complete individual behavioral data, but it also has some limitations: limited sample size, high implementation costs, and long data update cycles. Furthermore, due to the complexity of population activity-travel behavior, traditional survey methods struggle to achieve large-scale, long-term continuous observation.
[0004] With the development of mobile internet technology, mobile signaling data, which is the communication log recorded between user terminals and base stations in mobile communication networks, has been widely used in population mobility research due to its advantages such as wide coverage, high spatiotemporal accuracy, and low acquisition cost. However, mobile signaling data only contains information on changes in the location of user devices and cannot directly identify the activity category and travel purpose at the point of stay. This data characteristic leads to inherent limitations in the reconstruction of activity-travel chains: limited activity type identification, threshold constraints on stay duration, insufficient travel mode discrimination, and insufficient trajectory accuracy, resulting in significant uncertainties in the identification of key elements of the activity-travel chain.
[0005] Analyzing the characteristics of these two types of data reveals their complementarity: activity log survey data contains rich information on behavioral decisions, clearly reflecting activity types and travel purposes; mobile phone signaling data, on the other hand, has advantages in sample size, spatiotemporal continuity, and timely updates, and can better reflect movement patterns at the group level. Behavioral patterns in activity logs can help interpret the semantic content of mobile phone signaling data, while the large-scale spatiotemporal characteristics of mobile phone signaling data help verify and expand the behavioral patterns observed in activity logs.
[0006] In the field of activity-travel chain restructuring, existing technical solutions have the following main limitations:
[0007] (1) Although mobile signaling data has advantages in terms of sample size and spatiotemporal continuity, its data characteristics make it impossible to directly obtain key behavioral information such as activity type and travel choice.
[0008] (2) Activity log survey data can provide rich details of behavioral decision-making, but due to the limited survey scale and the difficulty in data acquisition, it is difficult to meet the needs of continuous observation of large-scale groups.
[0009] Based on the above analysis, how to effectively integrate these two types of heterogeneous data sources to achieve the collaborative construction of activity chains and travel chains has become a key scientific issue in current research. The core challenges include: how to establish a correlation mechanism between the two types of data, how to extract and transfer behavioral patterns from activity logs, and how to verify the fusion results and construct an evaluation system. Solving these problems not only has significant theoretical value but also important practical implications for improving the precision of urban governance and optimizing the supply of public services. Summary of the Invention
[0010] This invention is made to solve the above problems, and aims to provide a method and system for simulating and generating individual activity-travel chains based on multi-source data.
[0011] This invention provides a method for simulating and generating individual activity-travel chains based on multi-source data. It integrates mobile phone signaling data and activity log survey data for a specified time period within a research area, and includes the following steps: S10, obtaining population types and activity types, as well as their corresponding basic transition probability matrices, activity rhythm characteristics, activity duration distribution characteristics, time-varying activity participation intensity characteristics, and transportation mode characteristics based on activity log survey data; obtaining time start and end intervals, spatiotemporal distribution characteristics, population type and structural proportions, and spatial characteristics of entry and exit from the research area based on mobile phone signaling data; S20, dynamically correcting the basic transition probability matrix using time-varying activity participation intensity characteristics to obtain a dynamic transition probability matrix; S30, controlling the overall time range with the time start and end intervals, controlling the overall sample structure with the population type and structural proportions, determining the initial activity state with the time-varying activity participation intensity characteristics, and... Using the dynamic duration distribution characteristics as validation constraints, conditional transition probability row vectors are extracted from the dynamic transition probability matrix, and then the Monte Carlo method is used to generate activity chain sequences. In S40, hierarchical clustering and the K-means algorithm are combined to perform pattern recognition and classification on the activity chain sequences to obtain several typical activity chain patterns. In S50, based on the spatial locations of key activity nodes in the typical activity chain patterns, a spatiotemporal distribution probability matrix is constructed using spatiotemporal distribution characteristics. This matrix is then combined with POI / AOI data to construct a spatial selection probability model. Finally, spatial location selection is performed on all activities of all people in the activity chain to obtain complete activity chain data including spatial location selection. In S60, travel chains are constructed based on the spatial selection probability model. Combined with transportation mode characteristics, spatiotemporal trajectories are generated through path planning services. A travel chain trajectory dataset is constructed, and temporal optimization is achieved through adaptive distribution parameter adjustment and hierarchical sampling strategies.
[0012] The individual activity-travel chain simulation generation method based on multi-source data provided by the present invention may also have the following features: in step S10, the population types include working people, tourists, business people, commercial consumers and local residents, and the activity types include life services, cultural and sports leisure, employment, tourism, residential rest, shopping and consumption and catering activities.
[0013] The individual activity-travel chain simulation generation method based on multi-source data provided by the present invention may also have the following features: In step S10, after using the overlapping sliding window method to count the activity type transfer frequency in a certain time period, the model parameters based on the mixed-state first-order Markov chain model are iteratively optimized by the EM algorithm until the basic transfer probability matrix is obtained.
[0014] The individual activity-travel chain simulation generation method based on multi-source data provided in this invention may also have the following feature: the transportation mode features include several travel mode types and their selection probabilities, and the travel mode selection probabilities are represented by constructing a probability distribution model. In the formula, P m (d) represents the probability of choosing mode of transportation m at a distance d, α m As the initial choice tendency, β m γ is the distance sensitivity parameter. m Let d be the probability of choosing the basic option, and d be the travel distance.
[0015] The individual activity-travel chain simulation generation method based on multi-source data provided by this invention may also have the following feature: wherein, the constraint conditions are constructed as follows:
[0016]
[0017] For different population types, the probability distribution model parameters are fitted to the survey data using the least squares method, and normalization is used to ensure that the sum of the probabilities of choosing different travel modes at any distance is 1.
[0018] The individual activity-travel chain simulation generation method based on multi-source data provided by this invention may also have the following feature: step S20 includes the following sub-step: S21, the basic transition probability matrix is corrected according to the following formula: p k,t =P base ·I k (t), where P base Let I represent the basic transition probability matrix. k (t) represents the time-varying activity participation intensity characteristic, k represents the population type, t represents the time, and p k,t S22 represents the modified transition probability matrix of group type k at time t; for p k,t =P base ·I k The zero probability term in (t) is set to a lower limit ε = 10. -6 The results are then normalized to obtain the dynamic transition probability matrix.
[0019] The individual activity-travel chain simulation generation method based on multi-source data provided by this invention may also have the following features: Step S30 includes the following sub-steps: S31, determining the population type of the individual to be generated by probability sampling based on population type and structural proportion; determining the time boundary of the activity chain based on the start and end time interval of the population type; and determining the initial activity state by probability sampling based on the time-varying activity participation intensity characteristics of the selected population type at the initial moment; S32, from the dynamic transition probability matrix: Extract the conditional transition probability row vector: P k (t,i)=[p k (t,i,1),p k (t,i,2),…,p k [(t,i,n)], where p k (t,i,j) represents the probability that population type k transitions from state i to state j at time t, and the conditional transition probability row vector P k (t,i) represents the conditional probability distribution of transitioning from the current state i to all possible next states; S33, based on the conditional transition probability row vector P k (t,i), the Monte Carlo method is used to select the activity state at the next moment, and the process is repeated until the preset end time is reached to obtain a complete individual activity chain sequence; S34, a two-dimensional verification framework is constructed, and the goodness of fit between the individual activity chain sequence and the activity duration distribution characteristics is evaluated by single sample KS test to ensure the quality requirements and reasonable randomness of the samples.
[0020] The individual activity-travel chain simulation generation method based on multi-source data provided by this invention may also have the following feature: step S40 includes the following sub-steps: S41, constructing a comprehensive distance metric for the activity chain sequence: d ij =w1·E(S) i ,S j )+w2·T(S i ,S j )+w3·L(S i ,S j In the formula, S i Let S represent the sequence of the i-th activity chain. j Let d represent the sequence of the j-th activity chain. ij Indicates two activity chains (S) i ,S j The overall difference between E(S) and E(S) i ,S j T(S) is the edit distance of an activity-type sequence obtained by calculating the minimum number of edit operations required for sequence matching. i ,S j L(S) represents the time overlap calculated using Jaccard coefficients. i ,S j S42 represents the duration difference based on standardized Euclidean distance, where w1, w2, and w3 are weighting coefficients; For each different population type, several activity chain sequences are randomly selected as samples. After constructing a distance matrix, Ward's method is applied for hierarchical clustering. The silhouette coefficients under different numbers of clusters K are calculated, and the K value corresponding to the maximum silhouette coefficient for each population type is selected as the optimal number of clusters. Where a is the average distance between a sample and samples in the same cluster, and b is the average distance between a sample and samples in the nearest neighbor cluster; S43, based on the optimal number of clusters, the K-means algorithm is used to perform cluster analysis on all activity chain sequences to obtain the final K typical activity chain patterns.
[0021] The individual activity-travel chain simulation generation method based on multi-source data provided by this invention may also have the following feature: in step S50, the selection probability of key activity nodes is determined by the product of the activity spatiotemporal distribution probability and the basic attractiveness of the location. Among them, P t Here, R represents the activity's spatiotemporal distribution probability, N represents the venue rating, N represents the number of reviews, and Pr represents the average price. The selection probability of other activities in a typical activity chain pattern is determined by the activity's spatiotemporal distribution probability, the venue's basic attractiveness, and the spatial distance to adjacent activities. Where D is the Euclidean linear distance between adjacent activities.
[0022] This invention also provides a multi-source data-based individual activity-travel chain simulation generation system, characterized by using any of the aforementioned multi-source data-based individual activity-travel chain simulation generation methods, including: a data preprocessing and feature extraction unit, used to obtain population type and activity type, as well as their corresponding basic transition probability matrix, activity rhythm features, activity duration distribution features, time-varying activity participation intensity features, and transportation mode features based on activity log survey data; to obtain time start and end intervals, spatiotemporal distribution features, population type and structural proportions, and study area entry and exit spatial features based on mobile phone signaling data; and to dynamically correct the basic transition probability matrix using time-varying activity participation intensity features to obtain a dynamic transition probability matrix; and an activity chain generation model, connected to the data preprocessing and feature extraction unit, used to control the overall time range by the time start and end intervals, control the overall sample structure by the population type and structural proportions, determine the initial activity state by the time-varying activity participation intensity features, and use the activity duration distribution features as verification constraints to obtain the dynamic transition probability matrix. After extracting the conditional transition probability row vectors from the matrix, the Monte Carlo method is used to generate activity chain sequences. The clustering and validation analysis unit, connected to the activity chain generation model, combines hierarchical clustering and the K-means algorithm to perform pattern recognition and classification on the activity chain sequences, obtaining several typical activity chain patterns. The spatial location selection unit, connected to the clustering and validation analysis unit and the data preprocessing and feature extraction unit, constructs a spatiotemporal distribution probability matrix based on the spatial locations of key activity nodes in typical activity chain patterns, and then combines this matrix with POI / AOI data to construct a spatial selection probability model. This model then selects the spatial locations of all activities for all people in the activity chain, resulting in complete activity chain data including spatial location selection. Finally, the travel chain generation unit, connected to the spatial location selection unit and the data preprocessing and feature extraction unit, constructs travel chains based on the spatial selection probability model. Combining this with transportation mode characteristics, it generates spatiotemporal trajectories through path planning services, constructs a travel chain trajectory dataset, and optimizes the time series through adaptive distribution parameter adjustment and hierarchical sampling strategies.
[0023] The role and effect of invention
[0024] This invention provides a method and system for simulating and generating individual activity-travel chains based on multi-source data. By integrating the advantages of two types of data (mobile phone signaling data + activity log survey data), it reconstructs residents' activity-travel chains, aiming to more accurately describe the spatiotemporal behavioral patterns of the population and provide new analytical perspectives and technical solutions for urban planning and management.
[0025] This invention establishes a correlation mechanism between mobile signaling data and activity log data to achieve complementary advantages between the two types of data: on the one hand, it uses behavioral patterns in activity logs to identify activity types and travel purposes in mobile signaling data; on the other hand, it leverages the large-scale trajectory features of mobile signaling data to verify and expand the behavioral patterns in activity logs.
[0026] At the technical implementation level, this invention solves the following key problems:
[0027] (1) Spatiotemporal alignment of heterogeneous data sources: Construct a spatiotemporal alignment framework to achieve matching of mobile signaling data and activity log data in time and space dimensions.
[0028] (2) Activity-Travel Chain Generation Model: Based on fused data, a model is established to simulate the activities and travel behavior of an individual on a typical workday.
[0029] (3) Data processing and calculation methods: Develop efficient data processing and model calculation methods to ensure the feasibility of large-scale population behavior reconstruction.
[0030] (4) Model evaluation and verification system: An evaluation system is constructed from multiple dimensions such as spatiotemporal distribution characteristics and consistency of behavioral patterns to verify the reliability of the reconstruction results. Attached Figure Description
[0031] Figure 1 This is a flowchart of an embodiment of the present invention for a method of simulating and generating individual activity-travel chains based on multi-source data;
[0032] Figure 2 This refers to the percentage of the seven activity types in the stationary time after long-term evolution in step S123 of an embodiment of the present invention.
[0033] Figure 3 This refers to the activity rhythm characteristics of the catering activity in step S131 of an embodiment of the present invention;
[0034] Figure 4 This refers to the activity duration distribution characteristics of the working population in step S143 of an embodiment of the present invention;
[0035] Figure 5 This refers to the time-varying activity participation intensity characteristics of the working population in step S152 of an embodiment of the present invention;
[0036] Figure 6 This is a schematic diagram of the probability distribution of transportation mode selection under different travel distances in step S162 of an embodiment of the present invention;
[0037] Figure 7 This is a schematic diagram of the two-dimensional joint probability distribution of the start and end time intervals in step S171 of an embodiment of the present invention. Detailed Implementation
[0038] To make the technical means, creative features, objectives and effects of the present invention easy to understand, the following embodiments, in conjunction with the accompanying drawings, specifically illustrate a method and system for simulating and generating individual activities and travel chains based on multi-source data.
[0039] <Example>
[0040] This embodiment uses mobile phone signaling data and activity log survey data within a specified time period in the study area.
[0041] (1) Mobile signaling data: By recording the information interaction between mobile phone users and different base stations, the spatiotemporal location of users can be obtained relatively accurately, thereby reconstructing the dynamic flow trajectory of people in urban space. This embodiment uses anonymous mobile signaling data of Shanghai in November 2023 provided by China Unicom's Smart Footprint Platform. This dataset contains 8 fields, generates an average of about 27.299 million records per day, and covers 15 million users. In terms of spatial scale, the geohash7 grid encoding provided by the data platform is used to divide the Shanghai area into regular grid units with a side length of about 150 meters, which can support high spatial granularity of population dynamic analysis. The information fields of the mobile signaling data are shown in Table 1 below:
[0042] Table 1 (Illustration of Information Fields in Mobile Signalling Data)
[0043]
[0044] (2) Activity Log Survey Data: This embodiment takes the Lujiazui area of Shanghai as an example, and conducts an activity log questionnaire survey using a combination of online and offline methods, collecting 389 valid questionnaires. The survey content covers individual socioeconomic attributes (including age, gender, occupation, family structure, income level, education level, etc.) and complete activity-travel chain information on a typical workday. In terms of activity classification system, residents' daily activities are summarized into 11 main categories, including home activities, work, study, shopping, dining, leisure and entertainment, cultural and sports activities, and business activities. For each activity, the survey requires respondents to record its time attributes (start time, end time, duration, activity planning), spatial attributes (specific location of the activity venue), travel characteristics (mode of transportation, travel time, approximate route selection), and social attributes (composition of the people traveling with the activity). The individual attributes of the activity log survey data and the activity-travel chain of the activity log survey data are shown in Tables 2 and 3 below, respectively:
[0045] Table 2 (Illustration of Individual Attributes in Activity Log Survey Data)
[0046] ID age gender Profession Family Structure Monthly income (yuan) Education level 1 28 2 3 A1 20000 master
[0047] Table 3 (Activity Log Survey Data - Activity-Travel Chain Diagram)
[0048]
[0049] Figure 1 This is a flowchart illustrating an embodiment of the present invention of a method for simulating and generating individual activity-travel chains based on multi-source data. Figure 1 As shown, this embodiment provides a method for simulating and generating individual activity-travel chains based on multi-source data, which integrates mobile phone signaling data and activity log survey data of a set time period in the study area, including the following steps S10 to S60:
[0050] S10, Data preprocessing and feature extraction, including the following sub-steps S11 to S17:
[0051] S11, Obtain the population type and activity type based on activity log survey data:
[0052] (1) Standardize the coding of activity types and divide residents’ daily behaviors into seven categories: life services, cultural and sports leisure, employment and work, tourism and sightseeing, residential rest, shopping and consumption, and catering activities.
[0053] The activity types are categorized as shown in Table 4 below:
[0054] Table 4 (Categorization of Activity Types)
[0055] Activity type Mainly includes activities Lifestyle Services (A1) Hairdressing, beauty, laundry, repair, medical, financial, postal and other daily service activities Arts, Sports and Leisure (A2) Leisure activities such as sports and fitness, movie viewing and entertainment, cultural activities, and park recreation Employment (A3) Formal work and learning activities such as office employment, education and training, and business activities. Tourism and sightseeing (A4) Sightseeing and recreational activities such as sightseeing tours, cultural experiences, and city exploration. Residential and recreational areas (A5) Daily living, sleeping, and other rest activities in a place of residence or temporary accommodation Shopping expenses (A6) Shopping activities in shopping malls, supermarkets, specialty stores, and other similar venues Food and beverage activities (A7) Dining and social activities in restaurants, fast food restaurants, cafes, and other similar venues.
[0056] (2) The population in the study area (Lujiazui area) is divided into five categories: working population, tourist population, business population, commercial consumer population, and local resident population.
[0057] The classification of population types is shown in Table 5 below:
[0058] Table 5 (Explanation of Population Type Classification)
[0059] People type Definitions working people People who have a fixed office space in the Lujiazui area Tourists Tourists who visit primarily for sightseeing purposes Business professionals People visiting for temporary work purposes such as business meetings, negotiations, and training. Commercial consumers Visitors whose primary purpose is consumption activities such as shopping, dining, and leisure and entertainment Local residents Residents with fixed residences in the Lujiazui area, but without fixed workplaces.
[0060] In this embodiment, apart from the five main types of people mentioned above, other types of people (such as those handling emergencies or those on short-term transit) account for a very small percentage in the actual survey, and their visits are highly spontaneous and accidental. Therefore, they are not included in the analysis scope of this embodiment.
[0061] S12, Construct the basic transition probability matrix, including the following sub-steps S121~S123:
[0062] S121, using the overlapping sliding window method, for the activity characteristics at time t, analyzes all 30-minute observation windows containing t: [t-30min,t], [t-25min,t+5min], ..., [t-15min,t+15min], ..., [t,t+30min]. The windows slide in 5-minute increments, forming a multi-view observation sample with an overlap rate of 83.3%. Within each observation window, the transition frequencies between the seven activity types are counted, and the arithmetic mean of the statistical results for all windows containing time t is taken to establish the transition count matrix for that time.
[0063] N = {n ij}
[0064] In the formula, n ij This represents the average number of transitions from activity i to activity j;
[0065] S122, based on a mixed-state first-order Markov chain model, sets the state space S = {s1, s2, ..., s7} to correspond to seven types of activities, and the transition probability matrix P = {p ij} satisfies ∑p ij =1. The EM algorithm is introduced for parameter estimation: the E-step calculates the posterior distribution of the hidden state sequence Z, and the M-step maximizes the expected update parameter θ. new The optimal fundamental transition probability matrix P is obtained by iterating until convergence. base The basic transition probability matrix P base The following table 6 illustrates the probability of moving from one activity to another:
[0066] Table 6 (Basic Transition Probability Matrix P) base (Illustration)
[0067]
[0068] S123, regarding the obtained basic transition probability matrix P base Solve the characteristic equation by calculating its largest eigenvalue λ. max The corresponding eigenvector π is used to verify the steady-state distribution characteristics of the Markov chain.
[0069] To verify the basic transition probability matrix P base To determine the validity of the eigenvalue, solve its characteristic equation and calculate the largest eigenvalue λ. max And its corresponding normalized eigenvector π. Numerical analysis results show that the eigenvalue spectrum and its corresponding eigenvector satisfy the convergence property of a Markov chain, where π, as the steady-state probability distribution vector, characterizes the proportion of the seven activity types in the stationary time after long-term evolution, which is highly consistent with the original distribution of the sample, verifying the feasibility and accuracy of the construction of the basic transition matrix, such as... Figure 2 As shown.
[0070] S13, Extract activity rhythm features, including the following sub-steps S131~S132:
[0071] S131, Based on a 30-minute time unit, count the number of participants in various activities and calculate the probability of participation in the activity at time point t:
[0072]
[0073] Among them, P i (t) represents the probability of participation in activity i at time point t, n i N(t) represents the number of participants in activity i at time t, and N(t) represents the total number of participants at time t.
[0074] S131: Connect the calculated discrete probability points to construct the initial probability density curve. Apply cubic spline interpolation to the discrete probability sequence to construct the final activity rhythm probability density curve as the activity rhythm feature. The interpolation coefficients are determined by the least squares method to ensure node continuity.
[0075] Specifically, in this step, taking catering activities as an example, the rhythmic characteristics of the activities are as follows: Figure 3 As shown.
[0076] S14, Extract the activity duration distribution features, including the following sub-steps S141~S143:
[0077] S141, the duration of a single activity is obtained based on the difference between the start and end times of each activity in each sample.
[0078] S142, construct three layers of features for each sample: 1. Overall duration feature (cumulate the duration of all activities within the sample); 2. Type duration feature (statistically calculate the total duration of different activities according to activity type); 3. Type proportion feature (calculate the duration proportion of each type of activity).
[0079] S143. The feature data of the five types of population samples are summarized respectively, and the probability density is fitted by Gaussian kernel function. The optimal bandwidth parameter is determined by K-fold cross-validation. Finally, the three feature distributions of the overall duration distribution, type duration distribution and type proportion distribution of each type of population in the study area are used as the activity duration distribution features.
[0080] Specifically, in this step, the activity duration distribution characteristics of the working population are as follows: Figure 4 As shown.
[0081] S15, Based on the activity rhythm characteristics of the seven activity types obtained in step S13 and the activity duration distribution characteristics of the five population types obtained in step S14, construct time-varying activity participation intensity characteristics, including the following sub-steps S151 to S152:
[0082] S151, to achieve an organic combination of activity rhythm characteristics and activity duration distribution characteristics, the activity duration distribution characteristics of seven activities across five population groups are first optimized and extracted. The optimization objective is set as minimizing the root mean square error between the fixed percentage parameter and the distribution curve.
[0083]
[0084] Subject to the following constraints:
[0085]
[0086] Where, α ki For activity i, the fixed percentage parameter. Let its duration proportion distribution density function be defined. The constrained optimization problem is solved using the Lagrange multiplier method to obtain the baseline participation intensity vector α for each type of activity. ki =(α k1 ,…,α k7 ).
[0087] S152 combines the baseline participation intensity with the rhythm probability density curve to construct a time-varying participation intensity function as a characteristic of time-varying activity participation intensity:
[0088]
[0089] Among them, I k,i (t) represents the intensity of participation of group k at time t in activity i, p i (t) represents the rhythm probability density of activity i. The denominator ensures that the sum of the intensities of all activities is 1 at any given time, achieving normalization across time and activity type.
[0090] Specifically, in this step, the time-varying activity participation intensity characteristics of the working population are as follows: Figure 5 As shown.
[0091] S16, Extract transportation mode features, including the following steps S161~S162:
[0092] S161, based on detailed travel information recorded in the activity log survey data, integrates transportation modes into four categories: driving, public transportation, cycling, and walking. The integration process follows the principles of similarity in vehicle characteristics and similarity in usage scenarios, merging travel modes with similar physical characteristics and service functions. The classification of the four transportation modes is explained in Table 7 below:
[0093] Table 7 (Classification of the Four Modes of Transportation)
[0094]
[0095] S162, Based on the classification in step S161, construct a probability distribution model for the five groups of people choosing various modes of transportation for different travel distances. This model adopts an exponential decay form:
[0096]
[0097] In the formula, P m (d) represents the probability of choosing method m at a distance d, α m As the initial choice tendency, β m γ is the distance sensitivity parameter. m Let d be the probability of basic choice, and d be the travel distance.
[0098]
[0099] These are constraints.
[0100] For the five population groups, the probability distribution model parameters in this step were fitted to the survey data using the least squares method, and normalization was applied to ensure that the sum of the probabilities of choosing the four methods at any distance was 1. For example... Figure 6 As shown, the model yields the probability distribution of different groups of people choosing different modes of transportation for various travel distances.
[0101] S17, based on the extraction of mobile phone signaling data, including the start and end time intervals, spatiotemporal distribution characteristics, population type and structural proportions, and spatial characteristics of entry and exit in the study area, includes the following sub-steps S171 to S174:
[0102] S171, Extraction time start and end interval:
[0103] Based on the activity sequence of mobile phone signaling data, the activity time characteristics of the population in the study area are constructed. For each person with activity in the study area, the start time of their first activity and the end time of their last activity in the study area on that day are calculated to form time feature key-value pairs (t). start , t end Using kernel density estimation, two-dimensional joint probability distributions of temporal characteristics were constructed for five different population groups within the study area, serving as the start and end time intervals. A schematic diagram of the overall two-dimensional joint probability distribution of the start and end time intervals is shown below. Figure 7 As shown.
[0104] S172, Extracting spatiotemporal distribution features:
[0105] Geohash7 encoding was used to divide the study area into grids, and 30-minute time units were combined to characterize the spatiotemporal dimensions of population activities. For each time unit, the number of people engaged in three types of activities—staying at home, working, and other—within each grid was counted, and the proportion of each activity to the total number of people active in that time unit was calculated. Finally, a feature matrix reflecting the spatial selection probability of various activities in different time periods was constructed as a spatiotemporal distribution feature to characterize the spatial preferences of the population when engaging in different activities at different times, as shown in Table 8:
[0106] Table 8 (Statistical table of the distribution and proportion of participants in different time periods based on grid)
[0107]
[0108] S173, Extracting population type and structural proportion:
[0109] Based on user dwell time behavior characteristics provided by the mobile signaling data platform, the population in the study area was first divided into three categories according to the rules for determining work and residence activities: working population (those with work activities in the area), residence population (those with only residence activities), and other activity population (those with neither work nor residence activities). After data cleaning and outlier removal, the working population accounted for approximately 35.9%, the residence population for 21.9%, and the other activity population for 42.2%. Further combining activity log survey data, the other activity population was subdivided according to activity purpose into tourists primarily engaged in sightseeing, business people primarily engaged in business negotiations, and commercial consumers primarily engaged in shopping and dining. The final composition of these five groups is shown in Table 9.
[0110] Table 9 (Population Composition Ratio in Lujiazui Area)
[0111] People type percentage working people 34.0% Lujiazui local residents 20.8% Tourists 20.4% Business professionals 15.5% Visiting commercial consumers 9.3%
[0112] S174, Extracting spatial features of entry and exit in the study area:
[0113] Analyze the spatiotemporal distribution characteristics of users entering and leaving the study area. For each user trajectory, first identify the moment it crosses the boundary of the study area, determining its arrival time upon entering the study area and its departure time upon leaving. Simultaneously, extract the location information (grid encoding) and activity type attribute of the last activity point before entering the area and the first activity point after leaving the area. The grid information before arrival and after departure may be the actual location encoding or "0", where "0" indicates that the user was only active within the study area during that time period.
[0114] Furthermore, the probability distribution of the five population groups under different spatial combinations was statistically analyzed. A spatial combination consists of four dimensions: location-activity type before arrival and location-activity type after departure. Using attributes such as grid points before arrival, activity type before arrival, arrival time, grid points after departure, activity type after departure, and departure time, the sample size and probability of each combination within that population group were fully recorded, as shown in Table 10.
[0115] Table 10 (Example of spatial combination probability for regional entry and exit)
[0116]
[0117] To address the need for modeling time-varying activity choices, using a fixed basic transition probability matrix has limitations. First, the activity preferences and transition characteristics of different groups vary significantly across different time periods. Second, the activity participation patterns of various groups change dynamically over time. Directly calculating transition probabilities based on survey data across different time periods results in unreliable calculations due to sample sparsity. To solve this problem, the basic transition probability matrix is dynamically corrected using the following step S20.
[0118] S20, Construct the dynamic transition probability matrix, including the following sub-steps S21 to S22:
[0119] S21, the basic transition probability matrix is corrected according to the following formula:
[0120] p k,t =P base ·I k (t)
[0121] In the formula, p k,t Let P represent the modified transition probability matrix. base Let p represent the basic transition probability matrix, k represent the population type, t represent time, and p represent the time interval. k,t I represents the modified transition probability matrix of group k at time t. k (t) represents the time-varying activity participation intensity feature (the probability vector of type k people participating in various activities at time t).
[0122] This operation achieves dynamic adjustment of the transition probability by multiplying the base transition probability by the real-time activity participation probability.
[0123] S22, Considering the stability of numerical computation, a lower limit ε = 10 is set for the zero-probability term in the product result. -6 Subsequently, to maintain the validity of the transition probabilities, the processed results were normalized row by row so that the sum of the probabilities in each row was 1, and finally the dynamic transition probability matrix of various groups of people at different time periods was obtained, as shown in Table 11 below.
[0124] Table 11 (Dynamic Matrix of Activity Transition Probability for Lujiazui Residents at 8:00 AM)
[0125]
[0126] S30, Generate the active chain sequence, including the following sub-steps S31 to S34:
[0127] S31. Based on the population type and structural proportion, probability sampling is performed to determine the population type to which the individual to be generated belongs; the start and end time interval of this population type determines the time boundary of the activity chain; based on the time-varying activity participation intensity characteristics of the selected population type at the initial time, probability sampling is performed to determine the initial activity state.
[0128] S32, sequence generation employs a discrete-time simulation method with a 30-minute step size, achieving dynamic evolution of the activity state through iterative updates. At each time step t, based on the current population type k and time point t, the dynamic transition probability matrix is: Extract the conditional transition probability row vector: P k (t,i)=[p k (t,i,1),p k (t,i,2),…,p k (t,i,n)].
[0129] Where, p k (t,i,j) represents the probability that population type k transitions from state i to state j at time t, and the conditional transition probability row vector P k (t,i) represents the conditional probability distribution of transitions from the current state i to all possible next states.
[0130] S33, based on the conditional transition probability row vector P k (t,i) uses the Monte Carlo method to select the activity state at the next moment, and repeats the process until the preset end time is reached to obtain a complete individual activity chain sequence.
[0131] In steps S31 to S33 above, the activity chain is generated using a probabilistic sequence generation method based on Markov chains. This method assumes that an individual's activity state at any given time depends only on the current state and time, and characterizes the state transition features through a dynamic transition probability matrix.
[0132] S34, Result Validation and Optimization:
[0133] A two-dimensional validation framework was constructed, and the goodness of fit between the individual activity chain sequence and the categorical duration and categorical proportion features extracted by S14 were evaluated using a one-sample KS test. For each dimension, the D statistic and corresponding p-value were calculated, and the average p-values for each activity category were used as the goodness-of-fit index F1 for the categorical duration feature dimension and F2 for the categorical proportion feature dimension, respectively. The evaluation indices of the two dimensions were linearly combined with a weight of 0.5:0.5 and normalized to obtain the overall confidence score C. I .
[0134] To ensure generation quality and maintain appropriate random diversity, a strategy combining hard threshold constraints and probabilistic screening is adopted: when C I When C < 0.2, samples are filtered directly; when C I When ≥0.2, with probability p f litered =(1-C I Random selection is performed. This dual selection mechanism ensures both the basic quality requirements of the samples and maintains reasonable randomness within a high confidence range.
[0135] Table 12 below shows an example of the individual activity chain screening results based on the two-dimensional verification framework in this step:
[0136] Table 12 (Example of Individual Activity Chain Screening Results Based on Two-Dimensional Verification Framework)
[0137] Individual ID People type Credibility Serial Number Time period Activity type Duration (h) 7440 Tourists 0.46 1 10:00-11:00 A4 (Tourism and Sightseeing) 1.0 7440 Tourists 0.46 2 11:00-12:00 A7 (Food and Beverage Activities) 1.0 7440 Tourists 0.46 3 12:00-16:30 A4 (Tourism and Sightseeing) 4.5
[0138] Steps S31 to S34 above provide an activity chain generation model based on dynamic transition probability. This model uses a dynamically adjusted activity transition probability matrix as its core, combining the two-dimensional joint distribution characteristics of activity start and end times, time-varying activity participation intensity characteristics, and population composition ratios to simulate and generate activity chains from macroscopic statistical patterns to individual microscopic levels. The model employs a probability-based sequence generation method and ensures the reliability of the generated results through multi-level verification mechanisms, including the distribution of total activity duration, the distribution of activity duration by type, and the distribution of activity type proportions.
[0139] The model contains two types of conditions: generation constraints and verification constraints.
[0140] Generating constraints involves four aspects:
[0141] (1) The dynamic transition probability matrix obtained in step S20 is the main driver, which describes the activity transition probability characteristics of different groups of people in different time periods.
[0142] (2) Use the two-dimensional joint probability distribution of the start and end times in the time start and end interval extracted in step S171 to constrain the overall time range of the activity chain.
[0143] (3) The overall sample structure is controlled by the population type and structural proportion extracted in step S173.
[0144] (3) Determine the initial activity state using the time-varying activity participation intensity characteristics constructed in step S15.
[0145] The validation constraints are based on the three-layer distribution features extracted in step S142: 1. Overall duration features; 2. Type-specific duration features; 3. Type-specific proportion features. These are used to evaluate the reliability of the generated results. These probabilistic constraints together constitute the complete boundary conditions for the generation and validation of individual activity chains.
[0146] S40, Obtain the typical activity chain pattern, including the following sub-steps S41 to S43:
[0147] S41, Construct a comprehensive distance metric for the activity chain sequence:
[0148] d ij =w1·E(S) i ,S j )+w2·T(S i ,S j )+w3·L(S i ,S j )
[0149] In the formula, S i Let S represent the sequence of the i-th activity chain. j Let d represent the sequence of the j-th activity chain. ij Indicates two activity chains (S) i ,S j The overall difference between E(S) and E(S) i ,S j T(S) is the edit distance of an activity-type sequence obtained by calculating the minimum number of edit operations required for sequence matching. i ,S j L(S) represents the time overlap calculated using Jaccard coefficients. i ,S j The duration difference is based on standardized Euclidean distance, and w1, w2, and w3 are weighting coefficients. Specifically, in this embodiment, w1 = 0.4, w2 = w3 = 0.3.
[0150] S42. For each of the five population types, 200 activity chain sequences are randomly selected as samples. After constructing a distance matrix, Ward's method is applied for hierarchical clustering. The silhouette coefficients under different numbers of clusters K are calculated, and the K value corresponding to the maximum silhouette coefficient for each of the five population types is selected as the optimal number of clusters. Silhouette coefficient:
[0151]
[0152] In the formula, a is the average distance between the sample and samples in the same cluster, and b is the average distance between the sample and samples in the nearest neighbor cluster.
[0153] S43, based on the optimal number of clusters K determined from five population groups, uses the K-means algorithm to perform cluster analysis on all 10,000 activity chain sequences. In the specific implementation, K activity chains are first randomly selected from the dataset as initial cluster centers. During the iterative optimization phase, the comprehensive distance from each activity chain to each cluster center is calculated, and the chain is assigned to the cluster with the smallest distance. After each iteration, the cluster centers are updated by calculating the average characteristics of all activity chains in each cluster, including the mode pattern of the activity type sequence, the average time overlap, and the average duration. When the change in cluster partitioning results between two consecutive iterations is less than a preset threshold, the algorithm is considered to have converged, yielding the final K typical activity chain patterns.
[0154] Specifically, in this embodiment, after obtaining the final K typical activity chain patterns, clustering results are verified:
[0155] The typical activity chain patterns obtained from clustering were compared with the activity log survey data obtained in step S10, and their rationality was verified by expert evaluation and field investigation. The verification results show that the typical activity chain patterns obtained from clustering have a high degree of consistency with the actual samples, conform to field observation, and can well reflect the daily activity patterns of residents. The clustering results of typical activity chains of the working population are shown in Table 13 below:
[0156] Table 13 (Symptoms of clustering results for typical activity chains of the working population)
[0157]
[0158] In steps S41-S43 above, a clustering analysis framework based on edit distance was constructed for the generated 10,000 activity chain sequence data, and a two-level cross-validation mechanism was designed. By calculating the comprehensive distance across three dimensions—activity type transformation, time overlap, and duration—the activity chains were clustered to identify typical behavioral patterns. Simultaneously, a complete evaluation index system was established, combining actual sample comparison and expert evaluation methods, to ensure the reliability and practical value of the clustering results.
[0159] S50, Construct complete activity chain data including spatial location selection, including the following sub-steps S51 to S53:
[0160] S51 identifies the spatial locations of the most constraining key activities in the activity chain, whose spatial choices will guide and constrain the spatial distribution of other activities. The activity chains of the five typical population groups within the study area each have their specific key activity nodes, as shown in Table 14 below. Based on the spatial locations of these key nodes, and combining the principle of minimum spatial resistance and facility attractiveness, a probability distribution model for the spatial choices of other activities is constructed.
[0161] Table 14 (Key Activity Node Characteristics of Five Population Types)
[0162]
[0163] S52. Using the spatiotemporal distribution feature data obtained in step S172, a spatiotemporal distribution probability matrix of work activities, residential activities and other activities within the study area is constructed, with geohash7 encoding (spatial accuracy of about 153 meters) as the spatial unit and 30 minutes as the time unit.
[0164] By combining POI location data (including attribute information such as ratings, number of reviews, and prices) with AOI areal geographic feature data within the study area, a spatial selection probability model is constructed.
[0165] For key activity nodes in the activity chain, their selection probability is determined by the product of the activity's spatiotemporal distribution probability and the site's fundamental attractiveness:
[0166]
[0167] Among them, P t Let R be the event's spatiotemporal distribution probability, R be the venue rating, N be the number of reviews, and Pr be the average price.
[0168] For other activities in the activity chain, the spatial distance to adjacent activities needs to be considered in addition:
[0169]
[0170] Where D is the Euclidean straight-line distance between adjacent activities. After determining the selection probability of the spatial grid based on the time and type of the activity, the Huff gravity model is used to select specific POI locations within the specific grid.
[0171] S53. Select the spatial location for all activities of all people in the activity chain. For each activity, calculate the selection probability of each candidate location using the spatial selection probability model from step S52, based on its time, type, and the location of preceding activities. After normalizing the probabilities, use the Monte Carlo method for random sampling to determine the specific spatial location of the activity, ultimately obtaining complete activity chain data including spatial location selection.
[0172] The results of the activity chain spatial location selection in this step are shown in Table 15:
[0173] Table 15 (Illustration of Activity Chain Spatial Location Selection Results)
[0174]
[0175] S60, Construct the travel chain trajectory dataset, including the following sub-steps S61 to S64:
[0176] S61, Trip Chain Initialization:
[0177] A complete travel chain is constructed based on activity chain data with spatial location. For each activity chain, the start time of the first activity and the end time of the last activity are first obtained as the arrival and departure times. Based on the user's population type, and using the spatial features of the study area extracted in step S174 and the data in Table 10, a set of pre-arrival grid-activity type and post-departure grid-activity type combinations are randomly matched using normalized probabilities.
[0178] When the arrival grid identifier is 0, it indicates that the user's initial activity is within the study area, and the first activity position in the activity chain is set as the starting point of the travel chain; otherwise, the arrival grid is used as the starting point of the travel chain. Similarly, when the departure grid identifier is 0, the last activity position in the activity chain is set as the ending point of the travel chain; otherwise, the departure grid is used as the ending point of the travel chain. For each grid spatial unit, a set of latitude and longitude coordinates is determined from the corresponding POI points according to the activity type and a probability distribution as spatial location information. After determining the start and end points through the above process, OD pairs and their departure times between adjacent activity points in the travel chain are constructed.
[0179] S62, Transportation Options:
[0180] Based on the travel mode features extracted in step S16, the straight-line distance between each OD pair in the travel chain is calculated, and the probability of choosing a travel mode for each segment is determined by combining the user's population type. Using Monte Carlo sampling, the specific travel mode used for each segment is determined from the probability distribution of each travel mode.
[0181] S63, Path Planning and Trajectory Generation:
[0182] Based on the route planning service provided by the Gaode Map Open Platform, this system generates routes for various modes of transportation. API calls require configuration of core parameters such as origin and destination latitude and longitude coordinates and route planning strategies. Different interfaces enable route queries for walking, cycling, driving, and public transportation. API response data includes basic information such as waypoint sequences, travel distance, and time, as well as mode-specific data fields: driving includes real-time traffic conditions and toll fees, while public transportation includes transfer options and fare information. The returned data is parsed and standardized to generate a unified format dataset containing spatiotemporal trajectories, travel costs, and traffic information. Data from each travel segment is integrated chronologically to construct a complete travel chain trajectory dataset.
[0183] Example of the travel chain trajectory dataset in this step is shown in Table 16 below:
[0184] Table 16 (Example of travel chain trajectory data)
[0185]
[0186] S64, Correction:
[0187] This step provides a time offset correction method based on activity planning. Through adaptive distribution parameter adjustment and a hierarchical sampling strategy, it optimizes the timing of activity chains. This step establishes a quantitative mapping relationship between planning level and time constraints, and introduces an asymmetric probability distribution model to describe time offset characteristics, thereby ensuring the rationality and effectiveness of the correction results.
[0188] In terms of time constraint modeling, the activity planning level k(1-5) is mapped to the maximum acceptable time offset T(k) = 35-6k minutes. This mapping relationship allows activities with the strongest planning (k=5) to have an offset of only ±5 minutes, while activities with the weakest planning (k=1) can accept an offset of ±29 minutes, effectively characterizing the time flexibility differences between different types of activities. Based on this, a Beta(α,β) distribution is used to describe the time offset characteristics, where α=k and β=6-k, allowing the distribution shape to dynamically adjust with the planning level. Highly planned activities exhibit obvious skewed characteristics, while lowly planned activities tend to have a symmetrical distribution.
[0189] The time correction process employs a stratified accept-reject sampling strategy. First, the Beta distribution values [-T(k), T(k)] corresponding to each activity are standardized to the [0,1] interval. Then, time offsets are generated for non-first and last activity pairs in the travel chain. For adjacent and overlapping activity pairs (Ai, Ai+1), the system needs to reserve a time window for the travel activity through time offsets. Specifically, the system generates time offsets for both activities simultaneously, creating the necessary time interval through bidirectional expansion. When the initial sampling fails to meet the travel time requirements, the system compresses the distribution range of the two activities by changing the parameters α′ = 1.2α and β′ = 1.2β, and continues joint sampling. If a satisfactory result is still not obtained after reaching the maximum number of attempts (20 times), the maximum acceptable time offset T(k) is increased by 1.5 times and the above process is repeated. Examples of corrections achieved through this step are shown in Table 17 below.
[0190] Table 17 (Example of weekday activities - travel chain time correction for a sample)
[0191]
[0192] This embodiment also provides an individual activity-travel chain simulation generation system based on multi-source data, which uses the individual activity-travel chain simulation generation method based on multi-source data in this embodiment, including a data preprocessing and feature extraction unit, an activity chain generation model, a clustering and verification analysis unit, a spatial location selection unit, and a travel chain generation unit.
[0193] The data preprocessing and feature extraction unit is used to obtain population type and activity type, as well as the corresponding basic transition probability matrix, activity rhythm characteristics, activity duration distribution characteristics, time-varying activity participation intensity characteristics, and transportation mode characteristics based on the activity log survey data according to the methods in steps S10 to S20. It also obtains the start and end time interval, spatiotemporal distribution characteristics, population type and structural proportion, and entry and exit spatial characteristics of the study area based on mobile phone signaling data. The basic transition probability matrix is dynamically corrected using the time-varying activity participation intensity characteristics to obtain the dynamic transition probability matrix.
[0194] The activity chain generation model connects the data preprocessing and feature extraction units. It is used to control the overall time range by the start and end time interval, control the overall sample structure by the population type and structural proportion, determine the initial activity state by the time-varying activity participation intensity characteristics, and use the activity duration distribution characteristics as verification constraints. After extracting the conditional transition probability row vector from the dynamic transition probability matrix, the Monte Carlo method is used to generate the activity chain sequence.
[0195] The clustering and validation analysis unit is connected to the activity chain generation model and is used to perform pattern recognition and classification on the activity chain sequence according to the method in step S40, combining hierarchical clustering and the K-means algorithm, to obtain several typical activity chain patterns.
[0196] The spatial location selection unit connects the clustering and validation analysis unit and the data preprocessing and feature extraction unit. It is used to select the spatial location of all activities of all people in the activity chain according to the method in step S50, based on the spatial location of key activity nodes in the typical activity chain pattern, construct a spatiotemporal distribution probability matrix using spatiotemporal distribution features, and then construct a spatial selection probability model by combining POI / AOI data. This allows for the selection of the spatial location of all activities of all people in the activity chain, resulting in complete activity chain data containing spatial location selection.
[0197] The trip chain generation unit is connected to the spatial location selection unit and the data preprocessing and feature extraction unit. It is used to construct trip chains according to the spatial selection probability model according to the method in step S60, and generate spatiotemporal trajectories through path planning services in combination with the characteristics of transportation modes. It constructs a trip chain trajectory dataset and achieves temporal optimization through adaptive distribution parameter adjustment and hierarchical sampling strategy.
[0198] The role and effect of the embodiments
[0199] This embodiment achieves a key breakthrough in data fusion: existing studies mostly rely on single data sources, each with its limitations: while mobile phone signaling data possesses complete spatiotemporal trajectory information, it is difficult to identify specific activity types and travel purposes; activity log surveys contain detailed behavioral information, but the sample size is limited and may be subject to memory bias. This embodiment achieves effective fusion of the two types of data by constructing a multi-dimensional feature system. While maintaining the spatiotemporal representativeness of mobile phone signaling data, it incorporates behavioral pattern features from activity logs, thereby more accurately describing the activity patterns of urban populations.
[0200] This embodiment overcomes the technical challenges of multi-source data fusion: Addressing the differences in spatiotemporal precision, sampling frequency, and information dimensions among different data sources, this embodiment establishes a systematic data fusion framework. By combining micro-level behavioral patterns with macro-level spatiotemporal distributions through a dynamic transition probability matrix, and constructing a multi-level verification mechanism to ensure the reliability of the fusion results, it provides a new approach for integrating multi-source heterogeneous data in urban research.
[0201] This embodiment proposes a more realistic modeling scheme: compared with traditional methods that rely on simple rule matching or random sampling, the model framework based on dynamic transition probability provided in this embodiment, by introducing time-varying activity participation intensity and key node guidance strategies, can better reflect the spatiotemporal characteristics and behavioral constraints of crowd activities, thereby improving the accuracy of simulation results.
[0202] This embodiment has multiple practical application values: (1) In terms of urban planning, this embodiment can analyze the spatiotemporal activity patterns of different types of people, providing a basis for optimizing the layout of public service facilities and improving the land use function structure, and helping to identify spatial function mismatch and insufficient facility supply in the process of urban renewal; (2) In terms of traffic management, through dynamic analysis of the characteristics of people's activities, it can predict changes in traffic demand, providing a reference for optimizing the public transport network, planning parking facilities, and evaluating the effectiveness of traffic policies; (3) In terms of smart city construction, this embodiment realizes the dynamic updating of activity-travel characteristics, which can provide data support for the urban management platform and help optimize resource allocation and emergency response mechanisms; (4) In terms of urban resilience, this method can simulate changes in people's activities under special circumstances, assess the carrying capacity of the urban system, and provide a reference for the formulation of emergency plans.
[0203] Those skilled in the art should understand that this invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to this invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A method for generating individual activity-travel chain simulation based on multi-source data, characterized in that, Fusing mobile phone signaling data and activity log survey data of a study area setting time period, including the following steps: S10, obtaining population type and activity type, and corresponding basic transition probability matrix, activity rhythm characteristics, activity duration distribution characteristics, time-varying activity participation intensity characteristics and traffic travel mode characteristics based on the activity log survey data, obtaining time start and end interval, space-time distribution characteristics, population type and structure proportion, and research area entry and exit space characteristics based on the mobile phone signaling data, Wherein, the activity type transition frequency of a time period is counted by using overlapping sliding window method, and the model parameters based on mixed state first-order Markov chain model are iteratively optimized by EM algorithm until convergence to obtain the basic transition probability matrix, The traffic travel mode characteristics include several travel mode types and their selection probabilities, and the travel mode selection probability is represented by constructing a probability distribution model: , wherein is the probability of choosing mode m for distance d, is the mode initial choice propensity, is the distance sensitivity parameter, is the base choice probability, d is the trip distance, Wherein, the constraint condition is constructed: , For different population types, the parameters of the probability distribution model are fitted to the survey data by least squares method, and normalization processing is adopted to ensure that the sum of selection probabilities of different travel modes at any distance is 1; S20, dynamically correcting the basic transition probability matrix using the time-varying activity participation intensity characteristics to obtain a dynamic transition probability matrix, including the following sub-steps: S21, correcting the basic transition probability matrix as follows: , wherein, denotes the base transition probability matrix, is the time-varying activity engagement intensity feature, k denotes the population type, t denotes the time instant, denotes the modified transition probability matrix for the population of type k at time instant t, S22, to a lower limit ε = 10 is set for the zero probability term in -6 and the results are normalized to obtain the dynamic transition probability matrix; S30, controlling the overall time range by the time start and end interval, controlling the overall sample structure by the population type and structure proportion, determining the initial activity state by the time-varying activity participation intensity characteristics, extracting the conditional transition probability row vector from the dynamic transition probability matrix, and using the Monte Carlo method to generate activity chain sequences, including the following sub-steps: S31, determining the population type to which the individual to be generated belongs by probability sampling based on the population type and structure proportion, determining the time boundary of the activity chain based on the time start and end interval of the selected population type, and determining the initial activity state by probability sampling based on the time-varying activity participation intensity characteristics of the selected population type at the initial time, S32, from the dynamic transition probability matrix extracts the conditional transition probability row vector: , wherein, denotes the probability of the crowd type k transitioning from state i to state j at time t, the conditional transition probability row vector denotes the conditional probability distribution of transitioning from the current state i to all possible next states, S33, based on the conditional transition probability row vector The activity state of next time is selected by Monte Carlo method, and the process is repeated until the preset end time is reached, and a complete individual activity chain sequence is obtained. S34, constructing a two-dimensional verification framework, and evaluating the fitting degree of the individual activity chain sequence and the activity duration distribution characteristics by single sample KS test to ensure the quality requirements and reasonable randomness of the sample; S40, combining hierarchical clustering and K-means algorithm to perform pattern recognition and classification on the activity chain sequences to obtain several typical activity chain patterns, including the following sub-steps: S41, constructing a comprehensive distance metric of the activity chain sequence: , wherein, denotes the i-th activity chain sequence, denotes the j-th activity chain sequence, denotes the integrated difference between two activity chains , is the edit distance of the activity type sequences obtained by calculating the minimum number of editing operations required for sequence matching, is the time overlap degree calculated using the Jaccard coefficient, is the duration difference based on the normalized Euclidean distance, , , is a weight coefficient, S42, respectively randomly extract several activity chain sequences as samples for different population types, construct a distance matrix, and then apply the Ward method for hierarchical clustering. The optimal clustering number is selected as the K value corresponding to the maximum profile coefficient of different population types after calculating the profile coefficient under different clustering numbers K. The profile coefficient wherein is the average distance between the sample and the samples in the same cluster, is the average distance between the sample and the nearest neighbor cluster sample, S43, based on the optimal cluster number, using K-means algorithm to perform clustering analysis on all activity chain sequences to obtain the final K typical activity chain patterns; S50, based on the space position of the key activity node in the typical activity chain mode, constructing a space selection probability model by using the space-time distribution feature and combining POI / AOI data, and selecting the space position of all activities of all people in the activity chain to obtain complete activity chain data containing space position selection, Wherein, the selection probability of the key activity node is determined by the product of the activity spatio-temporal distribution probability and the place basic attractiveness: Wherein, is the activity spatio-temporal distribution probability, is the place score, is the number of comments, is the average price, The selection probability of other activities in the typical activity chain mode is determined by the activity space-time distribution probability, the place-based attraction, and the spatial distance from adjacent activities: where D is the Euclidean straight-line distance of adjacent activities. S60, constructing a travel chain according to the space selection probability model, and generating a space-time trajectory through a path planning service by combining the traffic mode feature, constructing a travel chain trajectory data set, and realizing time sequence optimization through an adaptive distribution parameter adjustment and a hierarchical sampling strategy.
2. The individual activity-travel chain simulation generation method based on multi-source data according to claim 1, characterized in that: in step S10, the crowd types include working crowds, tourist crowds, business crowds, commercial consumer crowds, and local residence crowds, and the activity types include life service, culture and leisure, employment, tourism, residence and rest, shopping and consumption, and catering activities. wherein The individual activity-travel chain simulation generation method based on multi-source data according to claim 1 or 2, comprising: a data preprocessing and feature extraction part for obtaining crowd types and activity types, and corresponding basic transition probability matrices, activity rhythm features, activity duration distribution features, time-varying activity participation intensity features, and traffic mode features based on the activity log survey data, obtaining time start and end intervals, space-time distribution features, crowd types and structure proportions, and research area entry and exit space features based on the mobile phone signaling data, and using the time-varying activity participation intensity features to dynamically correct the basic transition probability matrices to obtain dynamic transition probability matrices; 3. A multi-source data based individual activity-travel chain simulation generation system, characterized in that, an activity chain generation model connected with the data preprocessing and feature extraction part, for controlling the overall time range by using the time start and end intervals, controlling the overall sample structure by using the crowd types and structure proportions, determining the initial activity state by using the time-varying activity participation intensity features, extracting conditional transition probability row vectors from the dynamic transition probability matrices by using the activity duration distribution features as verification constraint conditions, and generating activity chain sequences by using a Monte Carlo method; a clustering and verification analysis part connected with the activity chain generation model, for combining hierarchical clustering and K-means algorithm to perform mode recognition and classification on the activity chain sequences to obtain several typical activity chain modes; a space position selection part connected with the clustering and verification analysis part and the data preprocessing and feature extraction part, for constructing a space-time distribution probability matrix by using the space-time distribution features based on the space position of the key activity node in the typical activity chain mode, constructing a space selection probability model by combining POI / AOI data, and selecting the space position of all activities of all people in the activity chain to obtain complete activity chain data containing space position selection; and The trip chain generation unit is connected with the space position selection unit and the data preprocessing and feature extraction unit, constructs a trip chain according to the space selection probability model, combines the traffic mode features, generates a space-time trajectory through a path planning service, constructs a trip chain trajectory data set, and realizes time sequence optimization through an adaptive distribution parameter adjustment and a hierarchical sampling strategy.
Citation Information
Patent Citations
Group activity data collection method and system based on multisource space-time trajectory data
CN106211071A
Multimode traffic distribution model construction method based on mobile phone signaling data
CN112133090A