A public transport passenger travel behavior space-time semantic similarity measurement method

By transforming public transportation passenger travel behavior attributes into word vectors and utilizing an improved word-shift distance metric, the problem of insufficient spatiotemporal semantic relevance in existing technologies is solved, achieving a more accurate measure of travel behavior similarity and supporting public transportation demand modeling and policy evaluation.

CN115718799BActive Publication Date: 2025-10-21BEIJING UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211492149.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-25
Publication Date
2025-10-21
Estimated Expiration
2042-11-25

AI Technical Summary

Technical Problem

Existing methods for measuring the similarity of public transport passenger travel behavior cannot effectively reflect the spatiotemporal semantic relevance of travel behavior, resulting in reduced accuracy of similarity measurement.

Method used

Using natural language processing techniques, the Word2vec model is used to transform travel attributes into word vectors. An improved word-shift distance is used to measure the spatiotemporal semantic similarity between travel sequences, constructing multidimensional travel sequences and capturing the spatiotemporal semantic correlation between different granularities.

Benefits of technology

It improves the accuracy of passenger travel behavior similarity measurement, enabling a better depiction of the changing patterns and complexity of passengers' daily travel behavior, and providing support for public transportation demand modeling and policy evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115718799B_ABST
    Figure CN115718799B_ABST
Patent Text Reader

Abstract

The application discloses a kind of public transport passenger travel behavior space-time semantic similarity measurement method, the implementation steps of this method include: (1) based on the card data of passenger extraction departure place, departure time, travel mode, activity type and destination 5 kinds of travel attributes, further construct public transport passenger individual travel sequence;(2) the travel attribute in passenger travel sequence is expressed as discrete variable;(3) the travel behavior attribute and travel sequence are analogized as word and sentence respectively, and the travel attribute is converted into word vector by applying Word2vec model, so as to realize the space-time semantic embedding representation of travel sequence;(4) the space-time semantic similarity of passenger multi-day travel behavior is measured by using improved word shift distance.The application solves the defect that traditional travel behavior similarity measurement model cannot consider its space-time semantic correlation, can provide support for passenger market segmentation, individual travel demand modeling and public transport policy making etc..
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for measuring the spatiotemporal semantic similarity of public transportation passenger travel behaviors, and belongs to the field of public transportation data mining applications. Background Art

[0002] The travel behavior of individual public transportation passengers is highly cyclical and predictable, but also subject to randomness, influenced by daily activity demands and other external factors. With the increasing proportion of residents engaging in multi-tasking travel and changes in work styles, diverse activity demands may complicate individual travel decision-making processes. Studying the similarities in passengers' daily travel behaviors and revealing the degree of variation and regularity in individual repetitive travel behaviors over medium and long periods of time can provide support for understanding travelers' refined travel needs.

[0003] In the current study of individual passenger travel behavior based on multi-source public transportation smart card data, the Sequence Alignment Model (SAM) is usually used to measure the similarity of individual multi-day travel behavior. For example, Liu S et al. in "Exploring travel pattern variability of public transport users through smart card data: role of gender and age" (IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 5, pp. 4247-4256) and Lin Pengfei et al. in "Study on daily similarity of individual activity chains of public transportation passengers" (Transportation Systems Engineering and Information, 2020, 20(6): 178-183, 204) each constructed a multidimensional sequence to characterize travel behavior to reflect the rich information and interdependence in card swiping data. The travel sequence is usually composed of attributes such as departure time, travel purpose and travel mode, and companions; on this basis, the Levenshtein distance and PrefixSpan algorithm are used to measure the similarity of passenger travel behavior. However, methods like the Levenshtein distance can only reflect the structural and categorical similarity of travel behavior sequences, but cannot reflect the spatiotemporal correlation between trip attributes, resulting in reduced accuracy in measuring the similarity between two trip sequences. Therefore, considering the spatiotemporal semantic correlation of travel behavior attributes will help better characterize the similarity of passengers' daily travel behaviors.

[0004] In the field of natural language processing, in order to measure the similarity between two sentences, word embedding technology is usually used to represent each word in the sentence as a word vector containing semantics, and the similarity between the two sentences can be measured using a distance function. Currently, natural language processing technology has been applied to the field of transportation. By converting the longitude and latitude coordinates of continuous trajectories into word vectors and embedding spatial semantic information, the spatial similarity of the two trajectories is measured. Considering that travel sequences have similar structures and characteristics as text data, each travel sequence reflects the mutual relationship and spatiotemporal constraints between activities and trips, and adjacent activities, and this relationship is similar to the structural characteristics of natural language, that is, words with the same context have similar semantics. Therefore, the present invention uses natural language processing technology to measure the spatiotemporal semantic similarity of passengers' daily travel behavior. Summary of the Invention

[0005] This paper aims to provide a method for measuring the spatiotemporal semantic similarity of public transit passengers' travel behavior, useful for analyzing long-term travel behavior patterns and understanding the decision-making mechanisms and complexity of individual travel behaviors under different spatiotemporal conditions. Based on passengers' public transit card swipe data, this method extracts travel attribute information from multiple dimensions to construct multidimensional travel sequences for individual public transit passengers. The Word2vec model is used to capture the spatiotemporal semantic correlations between travel attributes at different granularities. An improved word shift distance is used to measure the spatiotemporal semantic similarity between travel sequences, characterizing the daily changes in passengers' travel behavior.

[0006] The technical solution of the present invention is a method for measuring the spatiotemporal semantic similarity of public transportation passenger travel behavior, including the following technical solutions:

[0007] Step 1: Construct individual travel sequences of public transport passengers.

[0008] Step 1.1 Construction of individual passenger travel chain

[0009] Based on multi-source data such as passengers' smart card transaction data of buses, subways, and public bicycles, spatial vector data of station routes, and vehicle operation data, multi-source data fusion is used to construct the passenger's individual travel chain. The individual travel chain should include the passenger's card number, smart card type, travel mode, travel start and end time, the station name and longitude and latitude of the starting and ending points, travel distance, and other information.

[0010] Step 1.2 Active Extraction

[0011] Passenger trip chain data is sorted by departure time, and the starting and ending stations of each trip are extracted to form the passenger's activity site set. Each passenger's activity site set is clustered using the DBSCAN algorithm, which clusters spatially adjacent stations near the activity site. The DBSCAN algorithm uses the Haversine distance for distance calculation, with the neighborhood radius r and minimum sample set to 500 meters and 1, respectively.

[0012] Step 1.3 Identify the location of residence

[0013] Considering that most passengers' travel behavior is symmetrical, that is, the destination of a passenger's last trip of the day is the same as the departure point of their first trip of the day; the departure point of the first trip of the day is the same as the destination of the last trip of the previous day, and both are located near the passenger's residence. Therefore, the present invention uses the starting and ending points of the passenger's first and last trips of the day to identify the passenger's residence. The specific steps are as follows:

[0014] S1. Select the travel chain data of a passenger and sort them in ascending order by departure time.

[0015] S2. If the number of travel chains for the passenger on that day is greater than or equal to 2, the first and last travel chains are considered the first and last trips of the day, respectively. If the number of travel chains is 1, the travel chains with a departure time before 12:00 are defined as the first trip of the day, and the travel chains with a departure time after 12:00 are defined as the last trip of the day.

[0016] S3. Extract the departure point of all first trips and the destination of the last trip of the passenger during the study period, and define the most frequent travel location as the passenger's residence.

[0017] S4. Repeat the above steps until all travelers are traversed, and then end the algorithm.

[0018] Step 1.4 Activity Type Inference

[0019] First, based on the passenger's current trip chain t, the starting point of the adjacent trip chain t+1, and the end point of the adjacent trip chain t-1, the passenger's active state is identified and the start and end times of the activity are calculated. The specific steps are as follows:

[0020] S1. Extract the trip chain data of a traveler and sort them in ascending order by departure time. If trip chain t is the first trip in the period, or the interval between trip chain t and trip chain t-1 is greater than 1 day, the traveler is considered to be active at the starting point of trip chain t before the departure time of trip chain t.

[0021] S2. When trip chain t and trip chain t-1 occur on the same day, or on the second day after trip chain t-1, and the end point of trip chain t-1 is the same as the starting point of trip chain t, the passenger is considered to be in an active state; if they are different, the passenger is considered to have used non-public transportation to travel during this period.

[0022] S3. When trip chain t and trip chain t+1 occur on the same day, or on the day before trip chain t-1, proceed as in S2.

[0023] S4. When the interval between trip chain t and trip chain t+1 is greater than 1 day, or trip chain t is the passenger's last trip in the period, the passenger is considered to be active at the end of trip t from the end of trip chain t to the end of the day.

[0024] S5. Repeat the above steps until the travel chains of all travelers are traversed, and then the algorithm ends.

[0025] Then, the activity type of each passenger's trip is inferred based on the passenger's smart card type, the frequency of visiting the activity location, and the start and end time of the activity. The inference steps are as follows:

[0026] S1. The activity place with the highest frequency of visits outside of the place of residence is defined as the first activity place, and the remaining activity places other than the "place of residence" and the "first activity place" are defined as "other activity places."

[0027] S2. If the travel destination is located at the first activity location, and the activity start and end time is between 5:00-23:00, it is defined as "work", "study" and "living out" for ordinary cards, student cards and senior citizen cards respectively.

[0028] S3. If the destination of the trip is the passenger’s “place of residence”, the activity type is defined as “home”.

[0029] S4. If the destination of the trip is "other activity places", the activity type is defined as "other".

[0030] The inference rules for summarizing activity types are shown in Table 1.

[0031] Table 1 Inference rules for passenger activity types

[0032]

[0033] Based on the above steps, the five types of travel attributes of each passenger's travel chain are extracted, namely, the starting point, travel mode, departure time, activity type, and destination. All the travel chains of the passenger in one day are spliced ​​together in the order of departure time to obtain the passenger's travel sequence for one day, namely Sequence p,d ={trip k(startPoint,travelMode,departureTime,activityType,endPoint),|k=1,2,···,N}, where Sequence p,d represents the travel sequence of the pth traveler on the dth day, trip k It represents the kth trip of the traveler on that day. startPoint, travelMode, departureTime, activityType and endPoint represent the five types of travel attributes: starting point, travel mode, departure time, activity type and destination respectively.

[0034] Step 2: Discretization representation of the trip sequence.

[0035] For the five types of travel attributes in the travel sequence, namely, travel origin, travel mode, departure time, activity type and destination, all are represented by discretized variables. The specific steps are as follows:

[0036] S1. Use 6-digit Geohash to represent the latitude and longitude of the departure and destination points.

[0037] S2. For trip chains without transfers, use the strings "bus," "subway," and "bike" to represent bus, subway, and public rental bicycle, respectively. For trip chains with transfers, use "to" to connect the two modes. For example, bus to subway, subway to bus, and bus to bus are represented as "bustosubway," "subwaytobus," and "bustobus," respectively.

[0038] S3. Divide a day into 24 time periods at hourly granularity, and represent them using a string consisting of the string "hour" and the time period label. For example, if a passenger departs between 6:00 and 6:59, it is represented as "hour06".

[0039] S4. Represent the five activity types of “going home”, “going to work”, “going to school”, “going out” and “other” as five character strings “home”, “work”, “study”, “main” and “other” respectively.

[0040] Step 3: Embed spatiotemporal semantic information based on the Word2vec model.

[0041] To make trip sequences better reflect the interrelationships and spatiotemporal constraints between activities and trips, as well as between adjacent activities, we analogize trip behavior attributes and trip sequences to words and sentences, respectively. The collection of all passenger trip sequences constitutes a document. All trip attributes in a document constitute a vocabulary, with a vocabulary length of V. Each trip attribute is represented using a one-hot encoding of length N.

[0042] The Skip-gram framework in the Word2vec model is used to train travel attributes into word vectors to capture the spatiotemporal semantic correlation between different granularities of each attribute. The Skip-gram framework is a neural network structure consisting of an input layer, a hidden layer, and an output layer. For the travel attribute indexed as i in the vocabulary, v is used to represent the travel attribute. i and u i Represents the vector when it is the center word and context word. For a given travel attribute w c , generate any upper and lower adjacent travel attributes w o The conditional probability of can be calculated by the softmax function of the vector dot product:

[0043]

[0044] u o is the travel attribute w o As the vector of the context word, v c is the travel attribute v c As the center word vector; for a given travel document of length L, the likelihood function of the model is the probability of generating adjacent upper and lower travel attributes given any travel attribute as the center word:

[0045]

[0046] Where w (l) is the travel attribute with index l, and m is the context window.

[0047] The parameters of the model are the center word vector and context word vector of all travel attributes in the vocabulary. The parameters of the model are trained by maximizing the likelihood function, that is, minimizing the loss function Loss:

[0048]

[0049] The final trained center word vector is the word vector representation of the trip attribute. All trip attributes in the trip sequence are represented by word vectors, and the word vector representation of the trip sequence is obtained.

[0050] Step 4: Calculate the spatiotemporal semantic similarity of the line sequence based on the improved word shift distance.

[0051] Assume that passenger p’s travel activity sequence for any two days is p,d ={w i |i=1,2,···N} and Sequence p,b ={w j |j=1,2,···M}, the improved word shift distance is used to calculate the minimum shift distance required to transform all travel attributes in one travel sequence into another travel sequence.

[0052] The word shift distance uses the Euclidean distance to measure the difference in word vectors. Since the Euclidean distance is an unlimited quantity, it is not convenient for intuitive perception of similarity. Therefore, the present invention uses the cosine distance to represent the travel attribute w i The semantics of the trip is transformed into travel attributes w j Word travel cost, i.e.

[0053] Where, v i and v j Represent the travel attributes w i and w j word vectors.

[0054] In order to obtain the global minimum moving distance of the two sequences, it is converted into a linear programming problem:

[0055]

[0056] Where, γ i,j is the travel attribute w i To travel attribute w j The transfer amount, travel attribute w i All removals γ i,j Should be equal to its own modulus || v i ||, travel attribute w j The amount of displacement should be equal to its own modulus || v j ||, at the same time, in order to highlight the importance of different travel attributes, the word vector is normalized, that is, and

[0057] Convert the calculated minimum word shift distance between the two sequences into sequence similarity The calculation is shown as follows:

[0058]

[0059] The value range of is [0,1], The closer it is to 1, the more similar the passenger's travel behavior is.

[0060] The beneficial effects of the present invention are mainly manifested in:

[0061] Based on multi-source data, including public transportation smart card data, this method uses multidimensional travel sequences to accurately characterize passengers' daily travel behavior, reflecting the interdependencies and spatiotemporal constraints between trips and activities, and between adjacent activities. Based on natural language processing techniques, this method embeds spatiotemporal semantic correlations into trip attributes and uses an improved word shift distance to measure the similarity of passenger travel behaviors. This overcomes the limitation of traditional travel behavior similarity measurement models that cannot account for spatiotemporal semantic correlations, and can provide support for public transportation travel demand modeling, market segmentation, and policy evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 A flow chart of the method of the present invention;

[0063] Figure 2 Visualization results of word vectors based on t-SNE; DETAILED DESCRIPTION

[0064] The present invention is further described below with reference to the accompanying drawings and examples. The method for measuring the spatiotemporal semantic similarity of public transportation passenger travel behavior comprises the following steps:

[0065] Step 1: Construct individual travel sequences of public transport passengers.

[0066] Step 1.1 Construction of individual passenger travel chain

[0067] Based on multimodal public transportation travel data, including passenger smart card data and station and line attribute data, and referring to the "Public Transportation Travel Feature Extraction Method Based on Multimodal Bus Data Matching" disclosed in Chinese invention patent application number CN201510068077.7, multiple travel stages with the same travel purpose were integrated in the order of departure time according to transfer time thresholds and transfer walking distance thresholds. Individual travel chain data for Beijing from April to May 2018 were obtained. The data sample is shown in Table 2.

[0068] Table 2 Sample passenger individual trip chain data

[0069]

[0070]

[0071] Step 1.2: Active Extraction

[0072] Step 1.3: Identify the location of residence

[0073] Step 1.3: Activity Type Inference

[0074] Based on the above steps, we extract five travel attributes for each passenger's trip chain: origin, travel mode, departure time, activity type, and destination. All of the passenger's trip chains for the day are concatenated in chronological order of departure time to obtain the passenger's daily travel sequence. For example, the travel sequence for passenger ***50000603*** on April 2, 2018, is {("Dabailou", "Public Transportation", "2018-04-02 09:33:01", "Going to Work", "Heyi Farm"), ("Heyi Farm", "Public Transportation", "2018-04-02 18:26:01", "Home", "Dabailou")}.

[0075] Step 2: Discretization representation of the trip sequence.

[0076] For the five types of travel attributes in the travel sequence, namely, travel origin, travel mode, departure time, activity type and destination, all are represented by discretized variables. Finally, the passenger’s travel sequence for one day is obtained. An example of the discretized representation of travel sequence data is shown in Table 3.

[0077] Table 3 Examples of discretized representation of passenger travel sequences

[0078]

[0079] Step 3: Embed spatiotemporal semantic information based on the Word2vec model

[0080] The travel behavior attributes and travel sequences are compared to words and sentences respectively, and the collection of all travel sequences of passengers constitutes a document. All travel attributes in the document constitute a vocabulary, which consists of 3051 words. Each travel attribute is represented by a one-hot encoding with a length of 100 dimensions. The present invention uses the Word2Vec of the genism package to train word vectors, and the context window is set to 3. The center word vector finally obtained by training is the word vector representation of the travel attribute. The t-SNE dimensionality reduction technology is used to reduce the word vector from 100 dimensions to 2 dimensions for visualization, as shown in the attached figure. Figure 2 As shown. Figure 2 It can be seen that words with the same semantic attributes are identified and grouped into several clusters; at the same time, in each cluster, the distance between words with similar semantics is smaller. Taking departure time as an example, Figure 2 The distance between the two adjacent time periods of 7 o'clock and 8 o'clock is small, while the distance between the two time periods of 7 o'clock and 17 o'clock is relatively large, that is, the distance between time periods will gradually increase with the increase of time interval.

[0081] Step 4: Calculate the spatiotemporal behavior similarity of the line sequence based on the improved word shift distance.

[0082] The improved word shift distance is used to calculate the similarity of each passenger's travel sequences on any two days. Taking the three passengers with card numbers "***50000603***," "***52384710***," and "***50001026***" as an example, the spatiotemporal semantic similarity of each passenger's travel sequences on any two days within a week is shown in Table 4.

[0083] Table 4. Spatial-temporal semantic similarity of passengers’ travel sequences on any two days

[0084]

[0085]

[0086] Assuming three scenarios of changes in departure time, travel mode, and activity type, the proposed model is compared with the SAM model. The comparison results are shown in Table 5. As shown in Table 5, the proposed model can well capture the spatiotemporal semantic correlation between travel attributes.

[0087] Table 5 Model comparison results

[0088]

[0089] Trip sequences 1, 2, and 3 have different departure times. The similarities between sequence 1 and sequence 2, and sequence 1 and sequence 3 are calculated using the SAM model and the method proposed in this paper, respectively. The results of the SAM model are both 0.8. However, for the method based on word shift distance, as the departure time changes from 7 o'clock to 9 o'clock, the similarity decreases from 0.940 to 0.810, indicating that the method proposed in this paper can better capture time correlation.

[0090] Sequences 1 and 4 have different travel modes, but the subway station in sequence 1 (Beijing University of Technology West Gate Station) and the bus station in sequence 4 (Beijing University of Technology Station) are spatially adjacent and are both located near the passengers' departure points. The SAM model treats the two stations as independent stations and ignores the spatial proximity relationship between the stations. The method proposed in this paper can better capture the spatial correlation between stations by introducing Geohash encoding of the stations and training them as word vectors.

[0091] The similarity measurement method based on word shift distance also reflects the relative importance of travel behavior attributes. For example, sequence 1 and sequence 4, and sequence 1 and sequence 5 both have two different attributes, and the similarity calculated by the SAM model is 0.6. That is, the contribution of travel mode and activity type to similarity is equivalent. However, the calculation results of the present invention show that the change of activity type has a greater impact on the similarity of travel behavior than the change of travel mode.

[0092] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A method for measuring the spatiotemporal semantic similarity of public transportation passenger travel behavior, characterized by: The following steps are involved: Step 1: Construct individual travel sequences of public transport passengers; Step 2: Discrete representation of travel sequence; Step 3: Embed spatiotemporal semantic information based on the Word2vec model; Step 4: Measure the spatiotemporal semantic similarity of the row sequence based on the improved word shift distance; The discretization representation of the travel sequence described in step 2 is to represent the five travel attributes of the travel sequence, namely, the travel origin, travel mode, departure time, activity type, and destination, using discretized variables. The specific steps are as follows: S1. Use 6-digit Geohash to represent the latitude and longitude of the departure and destination points. S2. For trip chains without transfers, use the strings "bus," "subway," and "bike" to represent bus, subway, and public rental bicycles, respectively. For trip chains with transfers, use "to" to connect the two modes. S3. Divide a day into 24 time periods at hourly granularity, using a string consisting of the string "hour" and the time period label. S4. Represent the five activity types "home", "work", "school", "outing", and "other" as five strings: "home", "work", "study", "main", and "other" respectively; In step 3, spatiotemporal semantic information is embedded based on the Word2vec model. The specific steps are: The travel behavior attributes and travel sequences are compared to words and sentences respectively. The set of all passenger travel sequences constitutes a document. All travel attributes in the document are represented by a one-hot encoding of length V. Each travel attribute is represented by a one-hot encoding of length N. The Skip-gram framework in the Word2vec model is used to train travel attributes into word vectors to capture the spatiotemporal semantic correlation between different granularities of each attribute; the Skip-gram framework is a neural network structure consisting of an input layer, a hidden layer, and an output layer; for the travel attribute indexed as i in the vocabulary, v is used to represent the travel attribute. i and u i Represents the vector when it is the center word and context word; for a given travel attribute w c , generate any upper and lower adjacent travel attributes w o The conditional probability of is calculated by the softmax function of the vector dot product: Where u o is the travel attribute w o As the vector of the context word, v c is the travel attribute v c The vector when it is the central word; For a given travel document of length L, the likelihood function of the model is the probability of generating adjacent upper and lower travel attributes given any travel attribute as the central word: Where w (l) is the travel attribute with index l, and m is the context window; The model parameters are the center word vector and context word vector of all travel attributes in the vocabulary; the model objective is to minimize the loss function Loss, which is the following log-likelihood function: The central word vector obtained through training is the word vector representation of the travel attribute. All travel attributes in the travel sequence are represented by word vectors, and the word vector representation of the travel sequence is obtained.

2. The method for measuring spatiotemporal semantic similarity of public transportation passenger travel behavior according to claim 1 is characterized in that: The steps of constructing the individual travel sequence described in step 1 specifically include: Step 1: Construct an individual passenger travel chain based on multi-source data fusion. The individual travel chain includes the passenger's card number, smart card type, travel mode, travel start and end time, the name and longitude and latitude of the starting and ending stations, and travel distance. Step 2: Sort the passenger's trip chain data by departure time, extract the starting and ending stations of each trip, form the passenger's activity station set, and cluster the passenger activity station set using the DBSCAN algorithm; Step 3: Use the starting and ending points of the passenger's first and last trips of the day to identify the passenger's residence. The specific steps are as follows: S1. Select the trip chain data of a passenger and sort them in ascending order by departure time; S2. If the number of trip chains for the passenger on that day is greater than or equal to 2, the first and last trip chains are considered the first and last trips of the day, respectively. If the number of trip chains is 1, the trip chains with a departure time before 12:00 are defined as the first trip of the day, and the trip chains with a departure time after 12:00 are defined as the last trip of the day. S3. Extract the departure point of all first trips and the destination of the last trip of the passenger during the study period, and define the most frequent travel location as the passenger's residence; S4. Repeat steps S1-S3 until all travelers are traversed; Step 4: Based on the passenger's current trip chain t, the starting point of the adjacent trip chain t+1, and the end point of the adjacent trip chain t-1, identify whether the passenger is in an active state, calculate the start and end times of the activity, and implement activity type inference. The specific steps are as follows: S1. Extract the trip chain data of a traveler and sort them in ascending order by departure time. If trip chain t is the first trip in the period, or the interval between trip chain t and trip chain t-1 is greater than 1 day, the traveler is considered to be active at the starting point of trip chain t before the departure time of trip chain t. S2. If trip chain t and trip chain t-1 occur on the same day, or on the second day after trip chain t-1, and the end point of trip chain t-1 is the same as the starting point of trip chain t, the passenger is considered to be in an active state; if they are different, the passenger is considered to have traveled by non-public transportation; S3. If trip chain t and trip chain t+1 occur on the same day, or on the day before trip chain t-1, proceed as in S2. S4. If the interval between trip chain t and trip chain t+1 is greater than one day, or trip chain t is the passenger's last trip in the period, the passenger is considered active at the destination of trip t from the end of trip chain t to the end of the day. S5. Repeat the above steps until all travel chains of all travelers are traversed; Then, the activity type of each passenger's trip is inferred based on the passenger's smart card type, the frequency of visiting the activity location, and the start and end time of the activity. The inference steps are as follows: S1. The most frequently visited activity location outside of the place of residence is defined as the primary activity location, and the remaining activity locations other than "place of residence" and "primary activity location" are defined as "other activity locations"; S2. If the travel destination is at the first activity location and the activity starts and ends between 5:00 AM and 11:00 PM, then the categories are defined as "work," "study," and "living out" for Standard, Student, and Senior Citizen Cardholders, respectively. S3. If the destination of the trip is the passenger's "place of residence", the activity type is defined as "home"; S4. If the destination of the trip is "other activity location", the activity type is defined as "other"; All the travel chains of the passenger in one day are spliced ​​together in the order of departure time to obtain the passenger's travel sequence for one day, that is, Sequence p,d ={trip k (startPoint,travelMode,departureTime,activityType,endPoint),|k=1,2,···,N}, where Sequence p,d represents the travel sequence of the pth traveler on the dth day, trip k It represents the kth trip of the traveler on that day. startPoint, travelMode, departureTime, activityType and endPoint represent the five types of travel attributes: starting point, travel mode, departure time, activity type and destination respectively.

3. The method for measuring spatiotemporal semantic similarity of public transportation passenger travel behavior according to claim 1 is characterized in that: In step 4, the spatiotemporal semantic similarity of the line sequence is calculated based on the improved word shift distance. The specific steps are: Assume that passenger p’s travel activity sequence for any two days is p,d ={w i |i=1,2,···N} and Sequence p,b ={w j |j=1,2,···M}, the improved word shift distance is used to calculate the minimum shift distance required to transform all travel attributes in one travel sequence into another travel sequence; Use cosine distance to represent the travel attribute w i The semantics of the trip is transformed into travel attributes w j Word travelcost, that is Where, v i and v j Represent the travel attributes w i and w j word vectors; In order to obtain the global minimum moving distance of the two sequences, it is converted into a linear programming problem: Where, γ i,j is the travel attribute w i To travel attribute w j The transfer amount, travel attribute w i All removals γ i,j Equal to its own modulus || v i ||, travel attribute w j The amount of displacement is equal to its own modulus || v j ||, and normalize the word vector, that is, and Convert the calculated minimum word shift distance between the two sequences into sequence similarity The calculation formula is as follows: The value range of is [0,1], The closer it is to 1, the more similar the passenger's travel behavior is.

Citation Information

Patent Citations

  • A Public Transportation Travel Feature Extraction Method Based on Multi-modal Bus Data Matching

    CN104766473B

  • Method and system for inferring travel purpose of individual passengers in urban rail transit

    CN113642625A

  • Natural language outputs for path prescriber model simulation for nodes in a time-series network

    US20210365643A1