Method for predicting passenger travel preferences based on division of passenger relationship integrated network community
By building a passenger relationship integration network and dividing communities, key passengers are found to predict other passengers' travel preferences, and the existing technology is solved to predict passengers' travel preferences in a comprehensive transportation environment, achieving more accurate and efficient prediction results.
Patent Information
- Application Number
- CN202211191549.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-09-28
AI Technical Summary
In the comprehensive transportation environment, it is difficult for the prior art to effectively use information from multiple transportation modes to predict passenger travel mode preferences, and the identified passenger relationship cannot be verified, resulting in low accuracy of the prediction results.
By preprocessing passenger ticket booking data, a passenger relationship integration network is built, and network communities are divided to find key passengers to predict other passengers' travel preferences.
In a comprehensive transportation environment, passengers' preferences for various modes of transportation can be more accurately predicted, reduce data collection costs, eliminate interference from sparse relationships, and reduce calculation amount and intermediate data.
Smart Images

Figure CN115577877B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of passenger travel data processing, and particularly relates to a method for predicting passenger travel preferences based on the division of a passenger relationship integrated network community. Background Art
[0002] Deeply understanding passengers' preferences for transportation modes is crucial for railways and airlines to improve their service levels and product quality. For decision-makers (Policy Makers) in civil aviation and railway-related departments and enterprises, understanding passengers' travel mode preferences is conducive to market segmentation, cultivating passenger loyalty, and improving the quality of decision-making and services.
[0003] The prior art "A method for predicting the travel of civil aviation passengers" calculates the correlation degree between passengers through the route co-occurrence degree and attribute co-occurrence degree of passengers. Among them, the route co-occurrence degree refers to the number of route co-occurrences between passengers, and the attribute co-occurrence degree refers to whether the age, gender, average discount, and average mileage of passengers are the same. After that, this method conducts topic modeling on passengers, routes, and airlines, discovers the potential topics and their distribution among them, and combines the three to obtain passengers' preferences for airlines and routes. After the above two steps are completed, this method constructs a passenger correlation graph for subsequent prediction. This prior art can play a certain role in a single transportation field, but it cannot make better use of the information provided by the comprehensive transportation environment with multiple transportation modes coexisting. At the same time, the passenger relationships identified by the prior art cannot be verified, and thus the impact of passenger relationships on predicting passengers' travel mode preferences cannot be fully utilized for prediction. Moreover, due to the interference caused by sparse passenger relationships not being excluded, the accuracy of the prediction results may be further reduced.
[0004] The prior art "A method and device for predicting passengers' travel modes" first obtains data on passengers' age, gender, income, travel frequency, travel purpose, and their travel mode choices through SP and RP surveys. After obtaining the data, this method conducts a correlation analysis on the above features and the feature of travel mode choice, and uses the features with statistically significant correlations as the main features. After obtaining the main features, this method conducts a clustering analysis on the data to obtain passenger classifications, sets the passenger classifications as a feature, and conducts a correlation analysis with other passenger features such as age and gender together with passengers' travel mode choices to obtain the main influencing features affecting passengers' travel modes. Finally, this method constructs a prediction model through a Logit model combined with the above main influencing features. After collecting the travel characteristics of passengers to be predicted, it predicts the travel modes of the passengers to be predicted. This prior art can predict passengers' travel modes to a certain extent in intercity transportation, but when it comes to cross-city transportation, due to the high cost of data collection and the neglect of the impact of passenger relationships, it is difficult to obtain a more comprehensive prediction result.
[0005] The prior art "A method for constructing an integrated passenger relationship network based on comprehensive transportation big data" mainly constructs an integrated passenger relationship network across transportation modes based on comprehensive transportation big data, enabling government departments and transportation and tourism enterprises to better understand the characteristic connotations and changing laws of relationships among passengers from both micro and macro levels, enhancing the understanding of passenger travel behaviors, and providing theoretical and methodological support driven by big data for their management and service decisions. However, it has a complex network construction method. This method requires constructing the passenger relationship network of a single transportation mode first and then integrating each network, resulting in a large amount of calculation and generating a large number of intermediate data. Summary of the Invention
[0006] In view of the above deficiencies in the prior art, a method for predicting passenger travel preferences based on the division of an integrated passenger relationship network community provided by the present invention solves the following problems:
[0007] 1. Most of the prior art is limited to a specific transportation mode field, such as civil aviation, etc., and has little effect on predicting the preferences of passengers for multiple transportation modes in a comprehensive transportation environment.
[0008] 2. The prior art infers passenger associations through route co-occurrence and attribute co-occurrence. However, such a method cannot fully describe the real passenger relationships, making it difficult to analyze passenger travel mode preferences through passenger relationships.
[0009] 3. The data collection cost is high and the impact brought by passenger relationships is ignored, making it difficult to obtain a more comprehensive prediction result.
[0010] 4. The prediction of passenger travel mode preferences is interfered by sparse passenger relationships, resulting in low accuracy of the prediction results.
[0011] 5. The prior art needs to construct the passenger relationship network of a single transportation mode first and then integrate each network, resulting in problems of large calculation amount and generating a large number of intermediate data.
[0012] To achieve the above invention purpose, the technical solution adopted by the present invention is: A method for predicting passenger travel preferences based on the division of an integrated passenger relationship network community, including the following steps:
[0013] S1. Preprocess and store the passenger ticket booking data to obtain a passenger personal information table and a passenger ticket booking data table;
[0014] S2. Construct an integrated passenger relationship network according to the passenger personal information table and the passenger ticket booking data table;
[0015] S3. Divide the integrated passenger relationship network to obtain an integrated passenger relationship network community;
[0016] S4. Find the key passengers in each passenger relationship integrated network community;
[0017] S5. Predict the travel preferences of other passengers within each passenger relationship integrated network community based on the key passengers.
[0018] Furthermore, the step S1 includes the following sub-steps:
[0019] S11. Eliminate the abnormal data from the passenger ticket booking data;
[0020] S12. Fill in the missing values in the passenger ticket booking data after eliminating the abnormal data to obtain the filled passenger ticket booking data;
[0021] S13. Encrypt the passenger privacy data in the filled passenger ticket booking data;
[0022] S14. Encode the encrypted filled passenger ticket booking data to obtain the passenger ticket booking data with a unified format;
[0023] S15. Store the passenger ticket booking data with a unified format using a hash table to obtain the passenger personal information table and the passenger ticket booking data table.
[0024] The beneficial effects of the above further solution are as follows: First, eliminate the abnormal data to exclude the influence of abnormal data, and fill in the missing values to avoid the impact on the follow-up.
[0025] Furthermore, the step S2 includes the following sub-steps:
[0026] S21. Traverse the passenger personal information table, and use the passenger IDs corresponding to the passengers with travel records in different transportation modes as the nodes of the passenger relationship integrated network to construct the network node set V0;
[0027] S22. Traverse the passenger ticket booking data table, extract the passengers involved in more than two passengers in the same order in the passenger ticket booking data table to form passenger relationship hyperedges, and use each passenger in each passenger relationship hyperedge as a node;
[0028] S23. When the intersection of the network node set V0 and the passenger relationship hyperedges is not empty, add the passenger relationship hyperedges to the passenger relationship hyperedge set;
[0029] S24. Take the union of the passenger relationship hyperedge set and the network node set V0 to obtain the passenger relationship integrated network node set;
[0030] S25. Obtain the passenger relationship integrated network H=(V, E) according to the passenger relationship integrated network node set, where H is the passenger relationship integrated network, E is the passenger relationship hyperedge set, and V is the passenger relationship integrated network node set.
[0031] Further, the step S3 includes the following sub-steps:
[0032] S31. Calculate the similarity between each pair of nodes according to the distance between the nodes in the passenger relationship integration network;
[0033] S32. Divide the similar nodes into the same community according to the similarity between each pair of nodes to obtain an initial community;
[0034] S33. Calculate the information entropy of the initial community;
[0035] S34. According to the information entropy of the initial community, add or remove a passenger relationship hyperedge from the initial community, and calculate the change value of the information entropy when adding or removing the passenger relationship hyperedge;
[0036] S35. Start adding or removing from the passenger relationship hyperedge with the largest change in information entropy to the initial community;
[0037] S36. Determine whether the community information entropy decreases after adding or removing the passenger relationship hyperedge to or from the initial community. If so, add or remove the nodes within the passenger relationship hyperedge to or from the initial community until the community information entropy of the initial community reaches the lowest value, and obtain a passenger relationship integration network community with the division completed. If not, use the initial community as the passenger relationship integration network community;
[0038] S37. Select the next initial community, and jump to step S35 until all initial communities are traversed to obtain multiple passenger relationship integration network communities.
[0039] The beneficial effect of the above further solution is that: The present invention divides the passenger relationship integration network community with the goal of minimizing the community information entropy. During the community division, the division process can be realized by continuously adding a hyperedge set to the community or removing a hyperedge set from the community. Each time a hyperedge is added to or removed from the community, it may cause a decrease in the community information entropy. Each time a hyperedge set is selected to be added to or removed from the community, start with the hyperedge set that can make the information entropy value decrease the fastest, so as to accelerate the division speed and achieve a better community division effect.
[0040] Further, the formula for calculating the similarity between nodes in the step S31 is:
[0041] W(v i ,v j )=(1 - g(v i ,v j ))·(1 + W1(v i ,v j )·a ij + W2(v i ,v j ))
[0042]
[0043]
[0044]
[0045] Among them, W(v i , v j ) is the similarity between the i-th node v i and the j-th node v j , g(v i , v j ) is the Jaccard distance between the node v i and the node v j , W1(v i , v j ) is the similarity gain of direct connection between the node v i and the node v j , W2(v i , v j ) is the similarity gain of indirect connection between the node v i and the node v j , a ij is the state of whether the i-th node v i and the j-th node v j are adjacent. Adjacent is 1, and non - adjacent is 0. f() is the sine function, which is used to calculate the influence degree of the direct connection between two nodes on the similarity. g(v j , u) is the Jaccard distance between the j-th node v j and the neighbor node u. g(v i , u) is the Jaccard distance between the i-th node v i and the neighbor node u. f(1 - g(v i , v j )) is the influence degree of the local similarity between the i-th node v i and the j-th node v j on the similarity of the direct connection between the two nodes. f(1 - g(v j , u)) is the influence degree of the local similarity between the j-th node v j and the neighbor node u on the similarity of the direct connection between the two nodes. f(1 - g(v i , u)) is the influence degree of the local similarity between the i-th node v i and the neighbor node u on the similarity of the direct connection between the two nodes. d(v i ) is the out - degree of the node v i . d(v j ) is the out - degree of the node v j . u is the node v i and the node v jCommon neighbor nodes when indirectly connected, N(v i ) is the set of passenger relationship hyperedges containing node v i , N(v j ) is the set of passenger relationship hyperedges containing node v j .
[0046] The beneficial effect of the above further solution is that the present invention takes into account both the similarity of direct connection and indirect connection, making the measurement of similarity more reasonable and better used for identifying the initial community.
[0047] Furthermore, the formula for calculating the information entropy of the initial community in step S33 is as follows:
[0048] s = ∑ v∈C s v
[0049] s v = -p(v)ln p(v) - (1 - p(v))ln(1 - p(v))
[0050]
[0051] where s is the information entropy of the initial community, s v is the information entropy of node v, C is the initial community, v is any node, C(v) is the initial community containing node v, e is the passenger relationship hyperedge, p(v) is the probability that the set of passenger relationship hyperedges containing node v is included in the initial community C(v), and E(v) is the set of sets of passenger relationship hyperedges containing node v.
[0052] The beneficial effect of the above further solution is that the information entropy calculation amount is low, there is no need to calculate the global entropy, the algorithm complexity is reduced, and the quality of the community division result can be measured.
[0053] Furthermore, step S4 includes the following sub-steps:
[0054] S41. Calculate the influence and activity of each node in each passenger relationship integrated network community;
[0055] S42. Calculate the node criticality value according to the influence and activity;
[0056] S43. According to the node criticality value, sort the nodes in descending order, and take the passengers corresponding to the nodes ranked in the front as the key passengers.
[0057] Furthermore, the formula for calculating the influence of nodes in step S41 is:
[0058]
[0059] Among them, I(ζ) is the influence of node ζ in the passenger relationship integrated network community, I0 is the basic influence of the node in the passenger relationship integrated network community, I(u) is the influence of the neighboring nodes of node ζ, h(ζ) is the excess degree of node ζ in the passenger relationship integrated network community, and h(u) is the excess degree of the neighboring nodes of node ζ;
[0060] The activity formula of the node in step S41 is:
[0061]
[0062] Among them, L(ζ) is the activity of node ζ in the passenger relationship integrated network community, cotravel(ζ) is the number of trips of the passenger corresponding to node ζ in the passenger relationship integrated network community and the passengers corresponding to other nodes in the passenger relationship integrated network community, and freq(ζ) is the total number of trips of the passenger corresponding to node ζ in the passenger relationship integrated network community;
[0063] The formula for calculating the node criticality value in step S42 is:
[0064] P(ζ)=I(ζ)·L(ζ)
[0065] Among them, P(ζ) is the critical value of node ζ in the passenger relationship integrated network community, I(ζ) is the influence of node ζ in the passenger relationship integrated network community, and L(ζ) is the activity of node ζ in the passenger relationship integrated network community.
[0066] The beneficial effect of the above further scheme is: the present invention classifies passengers with similar activities according to similarity, obtains multiple passenger relationship integrated network communities, and then finds passengers with high influence and activity in each passenger relationship integrated network community as key passengers, and uses the key passengers to measure the travel modes of other passengers.
[0067] Furthermore, the step S5 includes the following sub-steps:
[0068] S51. Obtaining the travel mode preference of key passengers according to the passenger personal information table;
[0069] S52. In each passenger relationship integrated network community, according to the travel mode preferences of key passengers, calculate the preference values of other passengers for different transportation modes;
[0070] S53. Taking the maximum value of the passenger's preference values for different modes of transportation as the most preferred mode of transportation for the corresponding passenger.
[0071] Furthermore, the formula for calculating the preference values of other passengers for different modes of transportation in step S52 is:
[0072]
[0073] Among them, U′ k is the preference value of other passengers for the k-th transportation mode, O is the set composed of all combinations of passenger relationship integrated network communities, c is any passenger relationship integrated network community belonging to the set O, χ is the passenger corresponding to any node belonging to the passenger relationship integrated network community c, is the status of whether the passenger χ is a key passenger in the passenger relationship integrated network community c, The passenger χ is a key passenger, The passenger χ is not a key passenger, t k (χ) is the number of times the passenger χ travels by the k-th transportation mode, and n is the number of types of transportation modes.
[0074] The beneficial effects of the above further solution are as follows: After obtaining the key passengers, the travel mode preferences of other passengers within the community can be predicted through the key passengers. The influence range of the passengers is set within the community, which can not only ensure that most of the passengers who may be affected by them are closely related to them or closely related to the passengers who are closely related to them, but also ensure that the interference of sparsity is excluded.
[0075] In summary, the beneficial effects of the present invention are as follows:
[0076] 1. The present invention predicts the travel mode preferences of passengers in the comprehensive transportation environment, so that the prediction of passengers' travel mode preferences can be carried out on the basis of more complete passenger travel data (corresponding to more complete and comprehensive passenger characteristics). In addition, since the data used in the present invention can be collected through the business information systems of various transportation modes, the cost of collecting data can be effectively reduced.
[0077] 2. The present invention identifies practical and reliable passenger relationships through the co-travel behaviors extracted from passenger data, describes the passenger relationships with the passenger relationship integrated network, and predicts the travel mode preferences of passengers based on the description of more complete and comprehensive passenger relationships.
[0078] 3. The present invention will design a method for dividing passenger communities so as to divide passengers into reasonable communities, and predict the travel mode preferences of passengers according to the passenger influence-activity, combined with the individual hierarchical characteristics of passengers and the characteristics of passenger relationships.
[0079] 4. The influence range of passengers is set within the community, which can not only ensure that most of the passengers who may be affected by them are closely related to them or closely related to the passengers who are closely related to them, but also ensure that the interference of sparsity is excluded.
[0080] 5. The present invention divides the passenger relationship integrated network community into passenger relationship integrated network communities, and classifies passengers with high similarity into one passenger relationship integrated network community. The key passengers in the passenger relationship integrated network community are used to characterize the transportation methods of other passengers. The calculation amount is small and a large amount of intermediate data will not be generated. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 A flowchart of a method for predicting passenger travel preferences based on passenger relationship integrated network community division;
[0082] Figure 2 An example diagram for predicting passengers' travel mode preferences;
[0083] Figure 3 This is the prediction result diagram of passengers’ travel mode preference. DETAILED DESCRIPTION
[0084] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0085] like Figure 1 As shown, a method for predicting passenger travel preferences based on passenger relationship integrated network community division includes the following steps:
[0086] S1. Preprocess and store passenger booking data to obtain a passenger personal information table and a passenger booking data table;
[0087] Passenger booking data refers to the large amount of data that can be automatically collected by the business information systems of transportation departments and enterprises such as civil aviation and railways during their operations. Passenger booking data mainly includes ticket purchase order number, passenger ID, age, gender, departure place, arrival place, travel time, distance, fare and other information. In passenger transport data, each passenger booking record usually contains an order number that identifies the booking record. If multiple records have the same order number, it means that the passengers in these records have made a joint booking, that is, a joint travel behavior. To protect data privacy, private data such as passenger IDs are encrypted.
[0088] The step S1 comprises the following sub-steps:
[0089] S11. Eliminate abnormal data from passenger booking data;
[0090] S12, filling the missing parts of the passenger booking data after removing the abnormal data with null values to obtain filled passenger booking data;
[0091] S13. Encrypt the passenger privacy data in the filled passenger booking data;
[0092] S14. Encode the encrypted filled passenger booking data to obtain passenger booking data with a unified format;
[0093] Under different transportation modes, attributes such as the date representation format and the encoding of the departure location may vary. By unifying the encoding, the format is unified.
[0094] S15. Use a hash table to store the passenger booking data with a unified format to obtain a passenger personal information table and a passenger booking data table.
[0095] In this embodiment, step S15 is specifically: using the passenger's ID as the key, and the various attributes of the passenger as the key value to store the passenger personal information table; using the order number as the key, and one or more passengers involved in the order and the travel-related attributes (distance, fare, etc.) of the corresponding order as the key value to construct the passenger booking data table.
[0096] The passenger personal information table includes information at the individual level (age, gender, address, etc.) and travel information (the first travel time, the last travel time, the total travel frequency, the total mileage, the total cost, etc. of the passenger under different transportation modes); the passenger booking data table includes the passenger's ID, train number / flight number (transportation tool identification number), departure location, destination, departure time, arrival time, transportation mode name or number, etc.
[0097] S2. Construct a passenger relationship integration network based on the passenger personal information table and the passenger booking data table;
[0098] The said step S2 includes the following sub-steps:
[0099] S21. Traverse the passenger personal information table, and use the passenger IDs corresponding to the passengers with travel records in different transportation modes as the nodes of the passenger relationship integration network to construct a network node set V0;
[0100] S22. Traverse the passenger booking data table, extract the passengers involved in more than two passengers in the same order in the passenger booking data table to form a passenger relationship hyperedge, and use each passenger in each passenger relationship hyperedge as a node;
[0101] S23. When the intersection of the network node set V0 and the passenger relationship hyperedge is not empty, add the passenger relationship hyperedge to the passenger relationship hyperedge set;
[0102] S24. Take the union of the passenger relationship hyperedge set and the network node set V0 to obtain a passenger relationship integration network node set;
[0103] S25. Integrate the network node set of passenger relationships to obtain a passenger relationship integration network, \(H=(V, E)\), where \(H\) is the passenger relationship integration network, \(E\) is the set of passenger relationship hyperedges, and \(V\) is the set of network nodes of the passenger relationship integration network.
[0104] In the subsequent content, the nodes referred to are the nodes in the passenger relationship hyperedges in step S22 and the nodes in step S21, that is, all nodes will participate in the processes described below.
[0105] S3. Divide the passenger relationship integration network to obtain passenger relationship integration network communities;
[0106] The step S3 includes the following sub-steps:
[0107] S31. Calculate the similarity between each pair of nodes according to the distance between the nodes in the passenger relationship integration network;
[0108] The formula for calculating the similarity between nodes in step S31 is:
[0109] W(v i , v j )=(1 - g(v i , v j ))·(1 + W1(v i , v j )·a ij + W2(v i , v j )
[0110]
[0111]
[0112]
[0113] where W(v i , v j ) is the similarity between the \(i\)-th node \(v i and the \(j\)-th node \(v j , g(v i , v j ) is the Jaccard distance between the node \(v i and the node \(v j , W1(v i , v j ) is the similarity gain of direct connection between the node \(v i and the node \(v j , W2(v i , v j ) is the similarity gain of indirect connection between the node \(v i and the node \(v j , and \(a\)ij is the state of whether the $i$-th node $v$ i and the $j$-th node $v$ j are adjacent. If adjacent, it is 1; if not, it is 0. $f()$ is the sine function, which is used to calculate the influence degree of the direct connection between two nodes on the similarity. $g(v$ j , $u)$ is the Jaccard distance between the $j$-th node $v$ j and its neighbor node $u$. $g(v$ i , $u)$ is the Jaccard distance between the $i$-th node $v$ i and its neighbor node $u$. $f(1 - g(v$ i , $v$ j )) is the influence degree of the local similarity between the $i$-th node $v$ i and the $j$-th node $v$ j on the similarity of the direct connection between the two nodes. $f(1 - g(v$ j , $u))$ is the influence degree of the local similarity between the $j$-th node $v$ j and its neighbor node $u$ on the similarity of the direct connection between the two nodes. $f(1 - g(v$ i , $u))$ is the influence degree of the local similarity between the $i$-th node $v$ i and its neighbor node $u$ on the similarity of the direct connection between the two nodes. $d(v$ i ) is the out-degree of node $v$ i . $d(v$ j ) is the out-degree of node $v$ j . $u$ is the common neighbor node when node $v$ i and node $v$ j are indirectly connected. $N(v$ i ) is the set of passenger relationship hyperedges containing node $v$ i . $N(v$ j ) is the set of passenger relationship hyperedges containing node $v$ j .
[0114] S32. According to the similarity between each node, divide the similar nodes into the same community to obtain the initial community;
[0115] Specifically, step S32 is as follows:
[0116] S321. Select any node in the passenger relationship integrated network as the initial community;
[0117] S322. Judge whether the similarity between the nodes in the initial community and their adjacent nodes is higher than the average similarity in the passenger relationship integrated network. If so, add the adjacent nodes to the initial community. If not, select the next node that has not been added to the initial community and jump to step S322 as the initial community until all nodes find their initial communities.
[0118] S33. Calculate the information entropy of the initial community;
[0119] In step S33, the formula for calculating the information entropy of the initial community is as follows:
[0120] s = ∑ v∈C s v
[0121] s v = -p(v) ln p(v) - (1 - p(v)) ln(1 - p(v))
[0122]
[0123] where s is the information entropy of the initial community, s v is the information entropy of node v, C is the initial community, v is any node, C(v) is the initial community containing node v, e is the passenger relationship hyperedge, p(v) is the probability that the set of passenger relationship hyperedges containing node v is included in the initial community C(v), and E(v) is the set of sets of passenger relationship hyperedges containing node v.
[0124] S34. According to the information entropy of the initial community, add or remove passenger relationship hyperedges from the initial community, and calculate the change value of the information entropy of adding or removing passenger relationship hyperedges;
[0125] Select the hyperedge set with the largest change in information entropy according to the change value of the information entropy of adding or removing the hyperedge set.
[0126] S35. Start adding or removing from the passenger relationship hyperedge with the largest change in information entropy to the initial community;
[0127] S36. Determine whether the community information entropy decreases after the passenger relationship hyperedge is added to or removed from the initial community. If so, add or remove the nodes in the passenger relationship hyperedge from the initial community until the community information entropy of the initial community reaches the lowest value, and obtain a passenger relationship integrated network community with the division completed. If not, use the initial community as the passenger relationship integrated network community;
[0128] S37. Select the next initial community and jump to step S35 until all initial communities are traversed to obtain multiple passenger relationship integrated network communities.
[0129] S4. Find the key passengers in each passenger relationship integrated network community;
[0130] Step S4 includes the following sub-steps:
[0131] S41. Calculate the influence and activity of each node in each passenger relationship integrated network community;
[0132] In step S41, the formula for calculating the influence of a node is:
[0133]
[0134] Among them, I(ζ) is the influence of node ζ in the passenger relationship integrated network community, I0 is the basic influence of the node in the passenger relationship integrated network community, I(u) is the influence of the neighboring nodes of node ζ, h(ζ) is the excess degree of node ζ in the passenger relationship integrated network community, and h(u) is the excess degree of the neighboring nodes of node ζ;
[0135] The activity formula of the node in step S41 is:
[0136]
[0137] Among them, L(ζ) is the activity of node ζ in the passenger relationship integrated network community, cotravel(ζ) is the number of trips of the passenger corresponding to node ζ in the passenger relationship integrated network community and the passengers corresponding to other nodes in the passenger relationship integrated network community, and freq(ζ) is the total number of trips of the passenger corresponding to node ζ in the passenger relationship integrated network community;
[0138] S42. Calculate the node criticality value based on influence and activity;
[0139] The formula for calculating the node criticality value in step S42 is:
[0140] P(ζ)=I(ζ)·L(ζ)
[0141] Among them, P(ζ) is the critical value of node ζ in the passenger relationship integrated network community, I(ζ) is the influence of node ζ in the passenger relationship integrated network community, and L(ζ) is the activity of node ζ in the passenger relationship integrated network community.
[0142] S43. Arrange the nodes in descending order according to the node criticality values, and take the passengers corresponding to the nodes in the front order as the key passengers.
[0143] Key travelers refer to those who can significantly influence the travel preferences of other travelers. This type of travelers should have two characteristics. On the one hand, they have strong influence and occupy a core position in the relationship network; on the other hand, they have the possibility of influencing others, that is, the possibility of establishing relationships with others. This can be calculated by the proportion of their choice of traveling together to their number of trips, so it can be called activity. After calculating the influence and activity of the travelers, combine the two in the form of a product, sort the nodes within each community, and take the top 25% of the travelers as key travelers (25% has been verified to be the proportion of key travelers that can make subsequent predictions most accurate).
[0144] S5. Predict the travel preferences of other passengers within each passenger relationship integrated network community based on key passengers.
[0145] The step S5 includes the following sub-steps:
[0146] S51. Obtain the travel mode preference of the key passenger according to the passenger personal information form;
[0147] S52. In each passenger relationship integrated network community, calculate the preference values of other passengers for different transportation modes according to the travel mode preference of the key passenger;
[0148] The formula for calculating the preference values of other passengers for different transportation modes in the step S52 is:
[0149]
[0150] where U′ k is the preference value of other passengers for the k-th transportation mode, O is the set composed of all passenger relationship integrated network community combinations, c is any passenger relationship integrated network community belonging to the set O, χ is the passenger corresponding to any node belonging to the passenger relationship integrated network community c, is the status of whether the passenger χ is a key passenger in the passenger relationship integrated network community c, the passenger χ is a key passenger, the passenger χ is not a key passenger, t k (χ) is the number of times the passenger χ travels by the k-th transportation mode, and n is the number of transportation modes.
[0151] S53. Take the maximum value among the preference values of the passenger for different transportation modes as the most preferred transportation mode of the corresponding passenger.
[0152] Figure 2 、 3 shows a process of predicting the travel mode of a passenger. Figure 2 There are three communities in total. The communities are represented by dashed boxes. The nodes within the same community are in the same dashed box. The sizes of the communities are 15, 7, and 9 respectively. The labels of the nodes represent the passenger influence ability values calculated by the passenger influence - activity model. Taking the community with 15 nodes as an example, the passenger influence ability values of the nodes with high influence and high activity are 0.068, 0.051, 0.031, and 0.029 respectively. The travel mode preferences of the other nodes in this community except these four nodes can be determined by the travel mode preferences of these four nodes. Figure 3 shows Figure 2Results of predicting passengers' travel mode preferences, where the labels of the nodes represent the predicted values of passengers' travel mode preferences (the nodes used for prediction are represented as true values). This predicted value is the average travel mode preference of high-influence and high-activity passenger nodes in the community.
[0153] The beneficial effects brought by the present invention include:
[0154] 1. It can be used to analyze passengers' preferences for different transportation modes in a comprehensive transportation environment. For the current competition and cooperation among various transportation modes, the prediction of passengers' travel mode preferences based on a single transportation mode is difficult to effectively reflect passengers' choices among multiple transportation modes, thus unable to obtain accurate prediction results. The present invention is carried out in a comprehensive transportation environment and can accurately predict passengers' preferences for multiple transportation modes.
[0155] 2. Effectively utilize multi-modal transportation data to improve the prediction accuracy. Traditional prediction methods can only combine the travel data of a single transportation mode, and most of the data collection costs are high and the scale is small. The present invention can, on the basis of passengers' ticket booking data, combine multi-modal transportation data to fully and comprehensively display passengers' travel characteristics and more accurately predict passengers' travel mode preferences.
[0156] 3. The description of passenger relationships is complete and comprehensive, and the interference caused by sparse passenger relationships is excluded. The present invention makes full use of passenger relationships to predict passengers' travel mode preferences. Passengers who purchase tickets through the same order are very likely to have an actual relationship, and such a relationship is more complete in a comprehensive transportation environment. Therefore, the present invention makes full use of passenger relationships in predicting passengers' travel mode preferences. Moreover, the present invention divides the passenger relationship network into closely connected communities through community division and conducts predictions within the communities, which can exclude the interference caused by sparse connections.
Claims
1. A passenger travel preference prediction method based on the division of a passenger relationship integrated network community, characterized in that The following steps are involved: S1. Preprocess and store passenger booking data to obtain a passenger personal information table and a passenger booking data table; S2. Based on the passenger personal information table and the passenger booking data table, a passenger relationship integration network is constructed, which is specifically as follows: S21, traverse the passenger personal information table, use the passenger IDs corresponding to passengers who have travel records in different modes of transportation as nodes of the passenger relationship integration network, and construct a network node set V0; S22, traversing the passenger booking data table, extracting the same order in the passenger booking data table involving more than two passengers, forming passenger relationship hyperedges, and taking each passenger in each passenger relationship hyperedge as a node; S23. When the intersection of the network node set V0 and the passenger relationship hyperedge is not empty, add the passenger relationship hyperedge to the passenger relationship hyperedge set; S24, performing a union of the passenger relationship hyperedge set and the network node set V0 to obtain a passenger relationship integrated network node set; S25. Obtain a passenger relationship integrated network according to the passenger relationship integrated network node set, H = (V, E), where H is the passenger relationship integrated network, E is the passenger relationship hyperedge set, and V is the passenger relationship integrated network node set; S3. Divide the passenger relationship integrated network to obtain passenger relationship integrated network communities, which are as follows: S31, calculating the similarity between nodes according to the distance between nodes of the passenger relationship integrated network; S32, according to the similarity between the nodes, similar nodes are divided into the same community to obtain an initial community; S33, calculating the information entropy of the initial community; S34, according to the information entropy of the initial community, a passenger relationship hyperedge is added or exited from the initial community, and a change value of the information entropy of the passenger relationship hyperedge is calculated; S35, join or exit the initial community starting from the passenger relationship hyperedge that has the greatest change in information entropy; S36, judging whether the community information entropy decreases after the passenger relationship hyperedge joins or exits the initial community, if so, adding or exiting the nodes in the passenger relationship hyperedge until the community information entropy of the initial community reaches a minimum value, thereby obtaining a passenger relationship integrated network community that has been divided, if not, taking the initial community as the passenger relationship integrated network community; S37, selecting the next initial community and jumping to step S35, until all initial communities are traversed to obtain multiple traveler relationship integrated network communities; S4. Find key travelers in each passenger relationship integration network community, specifically: S41. Calculate the influence and activity of each node in each passenger relationship integrated network community; S42. Calculate the node criticality value based on influence and activity; S43, arranging the nodes in descending order according to the node criticality values, and taking the passengers corresponding to the nodes arranged in front as the key passengers; S5. Based on key passengers, predict the travel preferences of other passengers in each passenger relationship integrated network community.
2. The passenger travel preference prediction method based on the division of the passenger relationship integrated network community according to claim 1, wherein The step S1 comprises the following sub-steps: S11. Eliminate abnormal data from passenger booking data; S12, filling the missing parts of the passenger booking data after removing the abnormal data with null values to obtain filled passenger booking data; S13. Encrypt the passenger privacy data in the filled passenger booking data; S14. Encode the encrypted filled passenger booking data to obtain passenger booking data with a unified format; S15. Store the passenger booking data with a unified format using a hash table to obtain a passenger personal information table and a passenger booking data table.
3. The passenger travel preference prediction method based on the division of the passenger relationship integrated network community according to claim 1, wherein The formula for calculating the similarity between computing nodes in step S31 is: W(v i ,v j ) = (1 - g(v i ,v j ))·(1 + W1(v i ,v j )·a ij + W2(v i ,v j )) Among them, W(v i , v j ) is the similarity between the i-th node v i and the j-th node v j . g(v i , v j ) is the Jaccard distance between node v i and node v j . W1(v i , v j ) is the similarity gain of direct connection between node v i and node v j . W2(v i , v j ) is the similarity gain of indirect connection between node v i and node v j . a ij is the adjacent state of the i-th node v i and the j-th node v j . If adjacent, it is 1; if not adjacent, it is 0. f() is the sine function, which is used to calculate the influence degree of the direct connection between two nodes on the similarity. g(v j , u) is the Jaccard distance between the j-th node v j and its neighbor node u. g(v i , u) is the Jaccard distance between the i-th node v i and its neighbor node u. f(1 - g(v i , v j )) is the influence degree of the local similarity between the i-th node v i and the j-th node v j on the similarity of the direct connection between the two nodes. f(1 - g(v j , u)) is the influence degree of the local similarity between the j-th node v j and its neighbor node u on the similarity of the direct connection between the two nodes. f(1 - g(v i , u)) is the influence degree of the local similarity between the i-th node v i and its neighbor node u on the similarity of the direct connection between the two nodes. d(v i ) is the out-degree of node v i . d(v j ) is the out-degree of node v j . u is the common neighbor node when node v i and node v j are indirectly connected. N(v i ) is the set of passenger relationship hyperedges containing node v i . N(v j ) is the set of passenger relationship hyperedges containing node v j .
4. The method for predicting passenger travel preferences based on the division of the passenger relationship integrated network community according to claim 3, wherein The formula for calculating the information entropy of the initial community in step S33 is: s = ∑ v∈C s v s v = -p(v)ln p(v) - (1 - p(v))ln(1 - p(v)) Among them, s is the information entropy of the initial community, and s v is the information entropy of node v, C is the initial community, v is any node, C(v) is the initial community containing node v, e is the passenger relationship hyperedge, p(v) is the probability that the set of passenger relationship hyperedges containing node v is included in the initial community C(v), and E(v) is the set of sets of passenger relationship hyperedge sets containing node v.
5. The passenger travel preference prediction method based on the division of the passenger relationship integrated network community according to claim 4, wherein The formula for calculating the influence of a node in step S41 is: in, Integrate nodes in the network community for traveler relations I0 is the basic influence of nodes in the passenger relationship integrated network community, and I(u) is the influence of nodes. The influence of neighbor nodes, Integrate nodes in the network community for traveler relations The degree of h(u) is the node The over-performance of neighbor nodes; The formula for the activity level of a node in step S41 is: Among them, is the activity of the node in the passenger relationship integrated network community ; is the number of times the passengers corresponding to the node in the passenger relationship integrated network community travel together with the passengers corresponding to other nodes in the passenger relationship integrated network community, and is the total number of trips of the passengers corresponding to the node in the passenger relationship integrated network community. The formula for calculating the key degree value of a node in step S42 is: Among them, is the key degree value of node in the passenger relationship integrated network community, is the influence of node in the passenger relationship integrated network community, is the activity of node in the passenger relationship integrated network community.
6. The passenger travel preference prediction method based on the division of the passenger relationship integrated network community according to claim 1, characterized in that, Step S5 includes the following sub-steps: S51. Obtain the travel mode preference of key passengers according to the passenger personal information table; S52. In each passenger relationship integration network community, calculate the preference values of other passengers for different transportation modes according to the travel mode preference of key passengers; S53. Take the maximum value among the preference values of passengers for different transportation modes as the most preferred transportation mode of the corresponding passenger.
7. The passenger travel preference prediction method based on the division of the passenger relationship integrated network community according to claim 6, characterized in that, The formula for calculating the preference values of other passengers for different transportation modes in step S52 is: Among them, U k ′ is the preference value of other passengers for the k-th transportation mode, Ο is the set composed of all passenger relationship integration network communities, c is any passenger relationship integration network community belonging to the set Ο, χ is the passenger corresponding to any node belonging to the passenger relationship integration network community c, is the status of whether the passenger χ is a key passenger in the passenger relationship integration network community c, The passenger χ is a key passenger, The passenger χ is not a key passenger, t k (χ) is the number of times the passenger χ travels by the k-th transportation mode, and n is the number of types of transportation modes.
Citation Information
Patent Citations
Overlapping community structure detection method and system of civil aviation passenger relationship network
CN111368213A
Fusion type passenger relation network construction method based on comprehensive traffic big data
CN113806450A