A method for predicting the up and down passenger flow of the subway and a processing terminal

By obtaining the historical passenger records and OD data of the subway network, combining clustering algorithms, distinguishing individual passengers and groups, calculating the probability and path of riding, the problem of inaccurate prediction of upstream and downstream and transfer stations in the existing technology is solved, and more accurate passenger flow prediction is achieved.

CN113902180BActive Publication Date: 2025-07-04PCI TECH GRP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111139593.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-27
Publication Date
2025-07-04
Estimated Expiration
2041-09-27

AI Technical Summary

Technical Problem

The existing subway passenger flow prediction technology uses the subway station as the granularity to predict, and it is impossible to accurately distinguish between passenger flow in the upstream and downstream directions and passenger flow in the transfer station, resulting in inaccurate prediction results.

Method used

By obtaining the historical passenger ride records of the target subway network, using OD data and clustering algorithms, distinguishing individual passengers and groups, calculating the passenger's ride probability and path, and combining factors such as date type and site type to predict the number of people up and down.

Benefits of technology

It improves the accuracy of subway passenger flow forecasts, especially at transfer stations, reduces the impact of accidental behavior, is more in line with passengers' actual ride behavior, and provides data to support subway management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113902180B_ABST
    Figure CN113902180B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting the up and down passenger flow of a subway and a processing terminal. The method includes statistically calculating the boarding probabilities of each boarding path of individual passengers according to OD data, first converting the group of passengers into individual passengers, and then calculating the boarding probabilities of each boarding path of the group of passengers, so as to determine the number of up and down passengers at the current station based on the boarding probabilities of each boarding path, thereby completing the passenger flow prediction. Compared with the conventional subway passenger flow prediction method, the present invention can predict the passenger flow intensity from each current station as the boarding station to any other station as the alighting station based on the up and down passenger flow prediction, is suitable for transfer stations, reduces accidental behaviors, and improves the accuracy of up and down passenger flow prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of subway passenger flow prediction, and particularly relates to a method for predicting the up and down passenger flows of a subway and a processing terminal. Background Art

[0002] The subway has increasingly become an important means of transportation for people to travel. While bringing huge economic benefits to society, the safety of the subway has also become an important part of daily management and operation. In order to avoid traffic jams, crowded platforms, and deploy subway staff such as security in advance, it is necessary to be able to predict the passenger flow of the subway in advance, especially to be able to make real-time predictions, so as to change the deployment at any time and meet the subway management needs of different passenger flows.

[0003] However, the existing subway passenger flow prediction technologies mainly predict the subway passenger flow with the subway station as the granularity, that is, to predict the inbound and outbound passenger flow intensity of a subway station. However, such passenger flow prediction has many deficiencies and defects due to the relatively rough granularity, including being unable to understand from which direction passengers enter and leave the station, and thus being unable to predict which direction has more train passenger flow and which direction has relatively less train passenger flow. For example, for the same line, the passenger flow intensities in the up and down directions during the same time period (such as the peak commuting hours) are different, so the existing subway passenger flow prediction hardly helps in the management of such up and down direction passenger flows. It also includes that for transfer stations, the inbound passenger flow intensity may be weak, but the station is often very congested because the transfer passenger flow intensity is too large. If the inbound and outbound passenger flows predicted based on the subway station as the granularity cannot predict such a situation. In addition, the existing subway passenger flow prediction often does not distinguish between individual passenger behaviors and group passenger behaviors, which will lead to a large deviation between the subway passenger flow prediction and the real passenger flow, and the authenticity of the prediction results remains to be improved. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, one of the purposes of the present invention is to provide a method for predicting the up and down passenger flows of a subway, which can solve the problems mentioned in the background art;

[0005] Another purpose of the present invention is to provide a processing terminal, which can solve the problems mentioned in the background art.

[0006] The technical solution for achieving one of the purposes of the present invention is: A method for predicting the up and down passenger flows of a subway, comprising the following steps:

[0007] Step 1: Obtain the historical passenger riding records of the target subway network within a unit time period. The historical passenger riding records include OD data. Form a set of individual passengers with unique passenger identity identifiers, and form a set of group passengers with the remaining passengers;

[0008] Step 2: Calculate the boarding probability of each passenger in the passenger individual set choosing any real boarding path according to the OD data. The real boarding path standard is the reliability of the passenger's real boarding path.

[0009] Step 3: Convert the passengers in the passenger group set into passenger individuals according to the OD data, so as to regard them as the passengers in the passenger individual set in Step 2, and then calculate the boarding probability of the passengers in the passenger group set choosing any real boarding path.

[0010] Step 4: Determine all possible paths from the current station as the boarding station to other stations as the alighting stations. Divide the passengers entering the current station into the current passenger individual set and the current passenger group set according to whether they have a unique identity identifier.

[0011] For the passengers in the current passenger individual set, calculate the number of passengers in the up and down directions at the boarding time according to the boarding probability obtained in Step 2. Take the average value of the number of passengers in the up direction and the average value of the number of passengers in the down direction at each boarding time as the first part of the up and down passenger numbers at the current station at the same current time.

[0012] For the passengers in the current passenger group set, calculate the second part of the up and down passenger numbers at each boarding time in real time according to the boarding probability obtained in Step 3.

[0013] Take the sum of the first part of the up passenger number and the second part of the up passenger number at the same boarding time as the predicted up passenger number at the current time at the current station, and take the sum of the first part of the down passenger number and the second part of the down passenger number at the same boarding time as the predicted down passenger number at the current time at the current station, thus completing the prediction of the up and down passenger numbers of the subway passenger flow.

[0014] Further, in Step 2, calculate all possible paths corresponding to each OD data of each passenger in the passenger individual set, and determine the real boarding path corresponding to each OD data from all possible paths. One OD data corresponds to several real boarding paths. The possible path refers to the boarding route composed of all available stations that can reach the alighting station from the boarding station.

[0015] Further, to determine the real boarding path from all possible paths, its specific implementation includes the following steps:

[0016] If the current possible path does not contain a transfer station, the current possible path is used as the real boarding path of the current OD data.

[0017] If the current possible path contains a transfer station, then use the current possible path that meets Condition 1 as a member of the first possible path set, and traverse all possible paths corresponding to the OD data to obtain the first possible path set.

[0018] Condition 1: The error between the theoretical travel time T of the current possible path and the actual travel time L corresponding to the OD data is within the preset threshold time t.

[0019] If there is only one possible path in the first set of possible paths, then this possible path is taken as the actual travel path for the current OD data. If there are two or more possible paths in the first set of possible paths, then continue to delete the possible paths that meet any of the situations in Condition 2 from the first set of possible paths to obtain the second set of possible paths.

[0020] Condition 2 is a condition formed based on the entry time and the surrounding locations covered by the possible path.

[0021] If there is only one possible path in the second set of possible paths, then this possible path is taken as the actual travel path corresponding to the current OD data; if there are zero possible paths in the second set of possible paths, then the current possible path corresponding to the highest priority in the corresponding situation in Condition 2 is taken as the actual travel path corresponding to the current OD data; if there are two or more possible paths in the second set of possible paths, determine whether there is a possible path that meets Condition 3. If there is, then the possible path that meets Condition 3 is taken as the actual travel path corresponding to the current OD data. If not, then randomly select a path from the second set of possible paths as the actual travel path corresponding to the current OD data.

[0022] Condition 3 is a condition formed based on the transfer stations of the possible path and the time at the transfer stations and the preset time.

[0023] Furthermore, the value of the preset threshold time t should be such that at least one possible path meets Condition 1.

[0024] Furthermore, Condition 3 specifically is: The current possible path includes transfer stations, and the transfer time of at least one transfer station is less than q times the theoretical travel time of other possible paths, where 0 < q < 1.

[0025] Furthermore, Condition 2 includes the following four situations:

[0026] a. The entry time is during the morning rush hour on a weekday, the exit time is during the evening rush hour on a weekday, and the transfer stations in the current possible path are stations covered by the office area or the commercial area.

[0027] b. The day of the entry time is a holiday or a weekend, and the current possible path includes intermediate stations within a preset range from tourist attractions or commercial centers. Intermediate stations refer to other stations in the current possible path except for the entry station and the exit station.

[0028] c. The day of the arrival time is within the n days before a holiday or the peak return period of a holiday, where n ≥ 1, and the current possible path contains an intermediate station covering any one of the transportation hubs such as an airport, a railway station, a passenger station, or a wharf;

[0029] d. The day of the arrival time is during a special event or a large-scale event, and the current possible path contains an intermediate station passing through the location where the special event or large-scale event is held.

[0030] Furthermore, the specific implementation of the riding probability for each passenger to select any real riding path includes the following steps:

[0031] Group all the OD data in the historical riding records, group the same OD data into the same group, and different OD data into different groups.

[0032] In the same group of OD data, take the ratio of the sum of all the same real riding paths of the current passenger to the sum of all the real riding paths under this group of OD data as the probability of the current real riding path, so as to calculate the riding probability of selecting any real riding path in a group of OD data.

[0033] Furthermore, in step 3, converting the passengers in the passenger group set into individual passengers according to the OD data, its specific implementation includes:

[0034] Group the passengers in the passenger group set according to the OD data, group the passengers with the same OD data into the same group, and different OD data into different groups. Construct several features for the passengers within each group of OD data according to the OD data.

[0035] Use the clustering algorithm to cluster each group of OD data according to the constructed features to obtain the clustering results. One clustering result corresponds to several classification results. One classification result contains one clustering center, and one classification result corresponds to several passengers.

[0036] Take the average value of the theoretical travel times of all the passengers under each clustering center as the travel time of the current clustering center, so as to calculate the travel time corresponding to each clustering center. Regard each clustering center as a virtual passenger.

[0037] Thus, the passengers in the passenger group set are converted into individual passengers.

[0038] Furthermore, in step 3, calculating the riding probability of the passengers in the passenger group set to select any real riding path, its specific implementation includes:

[0039] After obtaining the virtual passenger, the theoretical travel time of the virtual passenger is the travel time of the cluster center. All possible paths of each OD data of the virtual passenger and the corresponding theoretical travel time of each possible path are obtained. The virtual passengers are processed according to step 2, and the boarding probabilities and theoretical travel times of each virtual passenger for each real boarding path are calculated.

[0040] Calculate the boarding path formed by the virtual passenger from the current station to the target station and the probability sub_path_probability of each sub-path under this real boarding path according to formula ③:

[0041]

[0042] In the formula, Q represents the total number of groups, J represents the total number of sub-paths included in the real boarding path corresponding to each cluster center, M represents the total number of all real boarding paths corresponding to each cluster center, and prob represents the boarding probability of each real boarding path calculated according to step 2.

[0043] The technical solution for achieving the second object of the present invention is: A processing terminal, which includes:

[0044] A memory for storing program instructions;

[0045] A processor for running the program instructions to execute the steps of the subway up and down passenger flow prediction method.

[0046] The beneficial effects of the present invention are as follows: Compared with the conventional subway passenger flow prediction method, the present invention can predict the passenger flow intensity from each current station as the boarding station to any other station as the alighting station based on the up-and-down passenger flow prediction, including being able to adapt to transfer stations, and can play a role in providing data support for subsequent passenger flow intensity management and guidance. In addition, the present invention takes into account the individual behaviors and group behaviors of passengers. For individual passenger behaviors, they have certain regularity and stability. For group passenger behaviors, clustering algorithms can be used to cluster and predict the possible paths of these passengers without unique identity identifiers, thus avoiding the contingency caused by the inability to locate group passenger behaviors from interfering with the prediction results, reducing accidental behaviors, and further increasing the accuracy of up-and-down passenger flow prediction in subway stations. Finally, when determining the passenger's travel path, instead of simply assigning a single path to the passenger, it is considered that the passenger has multiple paths to choose from, and the probability of each path selection is different. Based on probability theory, the passenger's travel path is determined, which is more in line with the actual travel behavior of passengers and can also avoid reducing the accuracy of the model due to the problem of single path selection. At the same time, when determining the specific travel path of passengers, the priorities of multiple factors such as date type (N days before holidays, during holidays, peak return periods after holidays, peak morning and evening periods on weekdays), special events, and station types are fully considered, making the passenger's path closer to their actual travel path. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic flowchart of Embodiment 1;

[0048] Figure 2 It is a schematic diagram of a target subway network;

[0049] Figure 3 It is a schematic diagram of a processing terminal. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following further describes the specific embodiments of the present application in detail with reference to the accompanying drawings. It can be understood that the specific embodiments described herein are only used to explain the present application, rather than limiting the present application. Additionally, it should be noted that for the sake of description, only parts related to the present application are shown in the drawings rather than all of the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The processing can be terminated when its operations are completed, but there can also be additional steps not included in the drawings. The processing can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0051] Example 1

[0052] Reference Figure 1 , a method for predicting the up and down passenger flow of the subway, comprising the following steps:

[0053] Step 1: Obtain the historical passenger travel records of the target subway network within a unit time period. The historical passenger travel records include the boarding station, the alighting station, the boarding time, the alighting time, and the passenger identity identifier. The boarding station and the alighting station form an OD data. Calculate the real travel time of the passengers according to the boarding and alighting times, and divide the passengers into a passenger individual set and a passenger group set according to the passenger identity identifier. The passenger individual set refers to the set composed of passengers with unique identity identifiers, and the passenger group set refers to the set composed of passengers without unique identity identifiers. The passenger identity identifier determines whether the passenger has a unique identity identifier according to the way the passenger buys the subway ticket. For example, passengers who buy tickets by means of a registered one-card pass, WeChat payment, Alipay payment, UnionPay card payment, etc. have a unique corresponding identity identifier, so these passengers are regarded as individuals; while passengers who buy single-trip tickets in cash or other ways do not have a unique corresponding identity identifier, so these passengers are regarded as a group.

[0054] The unit time period can be selected according to the actual situation. For example, select the time period from 10:00 to 11:00 in the morning on August 1st. Within the unit time period, the same passenger may have more than one travel record, that is, there are records of taking multiple subways. Each travel record corresponds to an OD data, but one OD data can correspond to multiple travel records. The OD data refers to the boarding and alighting record composed of the boarding station and the alighting station and characterizing the travel direction. For example, if the boarding station is a1 and the alighting station is b1, then an OD data (a1, b1) is formed, and the travel direction is from station a1 to station b1, but the OD data (a1, b1) can have multiple travel records, including travel records with different times or different paths but the same boarding and alighting stations. It should be noted that for two same stations, due to different travel directions (upward and downward), they respectively belong to two OD data. For example, for these two stations a1 and b1, the OD data (a1, b1) means that the boarding station is a1 and the alighting station is b1, and the OD data (b1, a1) means that the boarding station is b1 and the alighting station is a1.

[0055] In addition, the up and down here refers to the relative concept in direction, not specifically defining a certain specific direction as up or down. Taking a current station as an example, if the direction leaving the current station is upward, then the direction returning to the current station is downward; conversely, if the direction leaving the current station is downward, then the direction returning to the current station is upward.

[0056] Line station information and train operation time information can be obtained from the target subway network. The line station information includes all stations and the distances between adjacent stations. The train operation time includes the train operation route map, the arrival time at each station, and the dwell time at each station.

[0057] Step 2: Calculate all possible paths corresponding to each OD data of each passenger in the passenger individual set and the theoretical travel time corresponding to each possible path.

[0058] Reference Figure 2 , a possible path refers to a riding route composed of all accessible stations that can reach the outbound station from the inbound station. Figure 2 is a schematic diagram of a target subway network. The dots in the figure represent a station. From Figure 2 the inbound station a1 to the outbound station b1 in, the possible paths at least include two riding routes [a1, M1, c1, b1] and [a1, M1, M2, b1]. Therefore, all possible paths of the OD data (a1, b1) are {[a1, M1, c1, b1], [a1, M1, M2, b1]}.

[0059] Among them, finding all possible paths of the OD data in the target subway network can be achieved by any one of algorithms such as Dijkstra algorithm, depth-first algorithm, breadth-first algorithm, etc. After calculating all possible paths, for a certain possible path, combined with the train operation time information, the theoretical travel time corresponding to the current possible path can be calculated. Traverse all possible paths, and then calculate the theoretical travel time of each possible path.

[0060] Specifically, for a path without transfer stations, the theoretical travel time path_consume_time can be calculated according to formula ①:

[0061]

[0062] In the formula, N represents the total number of stations of the current possible path. For example, the total number of stations N of the current possible path [a1, M1, c1, b1] is 4. running_time represents the running time between adjacent stations, which can be directly calculated from the train operation time information. dwell_time represents the dwell time of the train at each intermediate station, and i represents the i-th group of adjacent stations.

[0063] For a possible path with transfer stations, the theoretical travel time path_consume_time can be calculated according to formula ②:

[0064]

[0065] In the formula, M represents the total number of transfer stations on the current possible path. For example, for the current possible path [a1, M1, c1, b1], there are only two transfer stations M1 and c1, so M = 2. sub_path_consume_time represents the theoretical travel time of the sub-paths divided according to whether they are on the same line for the current possible path. The theoretical travel time of the sub-paths sub_path_consume_time is calculated according to formula ①. trans_time represents the theoretical transfer time of the passengers at the transfer stations. The theoretical transfer time can be preset artificially according to the actual transfer distance. The farther the transfer distance, the longer the theoretical transfer time; the closer the transfer distance, the shorter the theoretical transfer time. i represents sub-path i, and j represents transfer station j.

[0066] Similarly, taking the path [a1, M1, c1, b1] in Figure 2 as an example, since station a1 and M1 are on the same line, station M1 and station c1 are on the same line, and station c1 and station b1 are on the same line, therefore, the path [a1, M1, c1, b1] is divided into 3 sub-paths according to whether they are on the same line. The 3 sub-paths are {[a1, M1], [M1, c1], [c1, b1]}. Among them, both station M1 and station c1 are transfer stations, so M = 2.

[0067] It should be noted that the transfer stations here refer to whether there are transfer stations that transfer from one line to another on the current possible path. For example, for the path [a1, M1, c1, b1] in Figure 2 , since it transfers from station a1 to station c1 on another line, so here station M1 is a transfer station. Similarly, there is no transfer station in the path [a1, M1, M2]. Although both station M1 and station M2 are transfer stations in the actual sense (common stations that can transfer from one line to another), but for the current possible path [a1, M1, M2], since this possible path does not need to transfer from one line to another, therefore, neither station M1 nor station M2 is regarded as a transfer station on this path.

[0068] After calculating all possible paths corresponding to each OD data of each passenger and the theoretical travel time corresponding to each possible path, determine the real travel path corresponding to each OD data from all possible paths. One OD data corresponds to several real travel paths. One OD data may have more than one real travel path, that is, the real travel paths screened from all possible paths corresponding to the OD data are a set.

[0069] Specifically, for the current possible path, if the current possible path does not include a transfer station, then the current possible path is taken as the actual riding path of the current OD data; if the current possible path contains a transfer station, then the current possible path that meets Condition 1 is taken as a member of the first set of possible paths, and all possible paths corresponding to the OD data are traversed to obtain the first set of possible paths. Condition 1 is as follows:

[0070] Condition 1: The error between the theoretical travel time T of the current possible path and the actual travel time L corresponding to the OD data is within the preset threshold time t, that is, ||L - T|| ≤ t. The value of t can be adjusted according to actual needs. However, it should be noted that the value of t should be such that at least one possible path meets Condition 1. Otherwise, the value of t needs to be continuously increased. The actual travel time L corresponding to the OD data is also the difference between the departure time at the departure station and the arrival time at the arrival station as mentioned above.

[0071] If there is only 1 possible path in the first set of possible paths, then this possible path is taken as the actual riding path of the current OD data; if there are more than 2 possible paths in the first set of possible paths, then continue to delete the possible paths that meet any of the situations in Condition 2 from the first set of possible paths to obtain the second set of possible paths. Condition 2 is as follows:

[0072] a. The arrival time is during the morning rush hour on a weekday, the departure time is during the evening rush hour on a weekday, and the transfer station in the current possible path is a station covered by the office area or the commercial area. Among them, the morning and evening rush hours can be preset according to actual needs. Similarly, the office area and the commercial area are also defined according to needs. For example, set 7 am to 9 am as the morning rush hour, 5 pm to 7 pm as the evening rush hour, set the CBD area as the office area, and the stations within a certain distance from this CBD area are regarded as stations covered by the office area, and set the scope of a certain pedestrian street as the commercial area.

[0073] b. The day of the arrival time is a holiday or a weekend, and the current possible path contains intermediate stations that are very close to tourist attractions or commercial centers. Intermediate stations refer to other stations in the current possible path except the arrival station and the departure station. Whether it is very close to a tourist attraction or a commercial center can be judged by comparing the distance between the intermediate station and the tourist attraction or commercial center with the preset distance threshold. For example, if the distance between the intermediate station and the tourist attraction or commercial center ≤ the preset distance threshold, it can be regarded as very close, otherwise, it cannot be regarded as very close.

[0074] c. The day of the entry time is within the n days before a holiday or during the peak return period of a holiday, where n ≥ 1, and the current possible path contains intermediate stops that cover transportation hubs such as airports, railway stations, passenger stations, and docks. That is, if the distance of a certain intermediate stop to a transportation hub that can cover airports, railway stations, passenger stations, docks, etc. is very close, it can be regarded as covered. For example, in real life, most railway stations usually have subway stations inside, and this subway station is regarded as an intermediate stop that can cover the railway station.

[0075] d. The day of the entry time is during the holding of a special event or a large-scale event, and the current possible path contains intermediate stops that pass through the location where the special event or large-scale event is held. The special event or large-scale event can be defined according to the actual situation. For example, a national or international marathon race held in a city can be regarded as a large-scale event. Another example is that a meeting between two high-ranking officials held in a certain city can be regarded as a special event.

[0076] If there is only one possible path in the second set of possible paths, then this possible path is used as the actual travel path corresponding to the current OD data; if there are zero possible paths in the second set of possible paths, then sort them in ascending order of priorities a - d, and retain the current possible path corresponding to the highest priority in the second set of possible paths. That is, filter those that meet the situations of a - d according to the priority level (priority d > c > b > a), and retain the possible path that meets the situation with the highest priority in condition two in the second set of possible paths. For example, if a certain path in the first set of possible paths only meets d in condition two, then retain it in the second set of possible paths. Otherwise, continue to check if there is a path in the first set of possible paths that only meets c in condition two. If so, retain it in the second set of possible paths, and so on, to ensure that there is at least one or more possible paths in the second set of possible paths (if there is only one possible path, then this possible path is used as the actual travel path corresponding to the current OD data); if there are two or more possible paths in the second set of possible paths, determine if there is a possible path that meets condition three. If so, then use the possible path that meets condition three as the actual travel path corresponding to the current OD data. If not, randomly select a path from the second set of possible paths as the actual travel path corresponding to the current OD data. Among them, condition three is as follows:

[0077] Condition three: The current possible path contains transfer stations, and there is at least one transfer station whose transfer time is less than q times the theoretical travel time of other possible paths, where 0 < q < 1. In this embodiment, q is set to 0.2.

[0078] By following the above steps, the real travel path of each OD data of each passenger can be calculated, so as to calculate all the real travel paths corresponding to each OD data, and further calculate all the real travel paths corresponding to all the OD data of all passengers.

[0079] Group all the OD data in the historical travel records. The same OD data of each passenger is divided into the same group, that is, the set of OD data with the same entry station and exit station. Different OD data are divided into different groups. Since the OD data represents the travel direction from the entry station to the exit station, that is, it characterizes the up and down directions of the subway passenger flow. The OD data in the same group represents the same subway passenger flow direction, and the OD data in different groups represents different subway passenger flow directions. Thus, the OD data in the same group can reflect the passenger flow in the current direction (upward or downward). In the OD data of the same group, the OD data of each passenger includes multiple real travel paths. The real travel paths of the OD data recorded at different entry times may have the same real travel path, that is, there are overlapping real travel paths among the OD data recorded at different entry times in the same group of OD data. Therefore, the real travel path of the OD data in the same group is the sum of the real travel paths of all the OD data corresponding to the current passenger in this group.

[0080] For example, in a set of OD data (a1, b1), for the current passenger, it includes 5 boarding times, namely boarding time 1, boarding time 2, boarding time 3, boarding time 4, and boarding time 5. Boarding time 1 corresponds to 3 actual travel paths in this OD data (a1, b1), denoted as actual travel path a, actual travel path b, and actual travel path c; boarding time 2 corresponds to 4 actual travel paths in this OD data (a1, b1), denoted as actual travel path a, actual travel path d, actual travel path e, and actual travel path f; boarding time 3 corresponds to 4 actual travel paths in this OD data (a1, b1), denoted as actual travel path b, actual travel path c, actual travel path d, and actual travel path e; boarding time 4 corresponds to 5 actual travel paths in this OD data (a1, b1), denoted as actual travel path a, actual travel path b, actual travel path g, actual travel path h, and actual travel path j; boarding time 5 corresponds to 4 actual travel paths in this OD data (a1, b1), denoted as actual travel path a, actual travel path c, actual travel path d, and actual travel path f. Therefore, the sum of all actual travel paths in this set of OD data (a1, b1) is 3 + 4 + 4 + 5 + 4 = 20. The sum of actual travel path a is 4, the sum of actual travel path b is 3, the sum of actual travel path c is 3, the sum of actual travel path d is 3, the sum of actual travel path e is 2, the sum of actual travel path f is 2, the sum of actual travel path g is 1, the sum of actual travel path h is 1, and the sum of actual travel path j is 1.

[0081] In the same set of OD data, the ratio of the sum of the same actual travel paths of the current passenger to the sum of all actual travel paths under this set of OD data is used as the probability of the current actual travel path, so as to calculate the travel probability of selecting any actual travel path in a set of OD data. Therefore, among the above 5 boarding times, in the OD data (a1, b1), the travel probability of selecting actual travel path a is 4 / 20, the travel probability of selecting actual travel path b is 3 / 20, the travel probability of selecting actual travel path c is 3 / 20, the travel probability of selecting actual travel path d is 3 / 20, the travel probability of selecting actual travel path e is 2 / 20, the travel probability of selecting actual travel path f is 2 / 20, the travel probability of selecting actual travel path g is 1 / 20, the travel probability of selecting actual travel path h is 1 / 20, and the travel probability of selecting actual travel path j is 1 / 20. Since the travel probability of selecting any actual travel path has been calculated, by traversing all OD data of all passengers, the travel probability of selecting a certain actual travel path among all possible paths of all passengers is known. Therefore, the travel probability and theoretical travel time of each actual travel path of the passenger can be calculated.

[0082] Step 3: Group the passengers in the passenger group according to the OD data. Passengers with the same OD data are grouped into the same group, and passengers with different OD data are grouped into different groups. For the passengers within each group of OD data, construct the following five features:

[0083] Feature 1: Real travel time. According to the entry time and exit time of each passenger's ride, the real travel time corresponding to each ride path can be obtained.

[0084] Feature 2: First date type. Determine the first date type corresponding to the ride path according to whether the day of the entry time is a holiday, weekend, or weekday.

[0085] Feature 3: Second date type. Determine the second date type corresponding to the ride path according to whether the day of the entry time is any day from Monday to Sunday.

[0086] Feature 4: Riding time period. Determine the hourly period corresponding to each passenger's ride path according to the entry time. For example, if the entry time is 7:05, the riding time period is 7 o'clock. Another example, if the entry time is 8:56, the riding time period is 8 o'clock.

[0087] Feature 5: Card swiping type. Determine the card swiping type according to the passenger's ticket purchase method. For example, using a transportation card (such as the Yangcheng Tong in Guangzhou Metro), a QR code (such as the WeChat ride code), or a single - way ticket purchased with cash.

[0088] Use the clustering algorithm to cluster each group of OD data according to the above five features, that is, use the clustering algorithm to perform clustering processing on each group of OD data according to these five features. The passengers in a group of OD data get a clustering result, and a clustering result corresponds to several classification results. That is, all the passengers in the same group of OD data are divided into several classes. A classification result contains a clustering center, a classification result corresponds to several passengers, and the set of passengers corresponding to all classification results is all the passengers in the same group of OD data.

[0089] The clustering algorithm can use existing clustering algorithms. In this embodiment, the DBSCAN algorithm is used to cluster the passengers.

[0090] Take the average value of the theoretical travel times of all passengers under each clustering center as the travel time of the current clustering center, so as to calculate the travel time corresponding to each clustering center.

[0091] Regarding each cluster center as a passenger, it is denoted as a virtual passenger for this purpose. The theoretical travel time of the virtual passenger is the travel time of this cluster center. All possible paths of the OD data of the virtual passenger (i.e., the OD data on the purchased one-way ticket) and the corresponding theoretical travel time for each possible path are obtained. Therefore, the virtual passengers can be processed according to step 2 to calculate the boarding probabilities and theoretical travel times of each virtual passenger for each real boarding path.

[0092] After calculating the boarding probabilities and theoretical travel times of the real boarding paths corresponding to each cluster center, they can be merged to calculate the boarding paths formed from the current station to the target station and the probabilities of each sub-path under this real boarding path. The probability of the sub-path sub_path_probability is calculated according to formula ③:

[0093]

[0094] In the formula, Q represents the total number of groups, that is, the total number of clusters, J represents the total number of sub-paths included in the real boarding path corresponding to each cluster center, M represents the total number of all real boarding paths corresponding to each cluster center, and prob represents the boarding probabilities of each real boarding path calculated according to step 2.

[0095] Step 4: Using the current station as the boarding station and other stations in the target subway network as the alighting stations, all possible paths from the boarding station to the alighting stations are obtained. The passengers entering the current station are divided into a set of individual passengers and a set of group passengers according to whether they have a unique identity identifier.

[0096] According to the processing in step 2, count the up and down numbers of the set of individual passengers at each boarding time in all possible paths formed with the current station as the boarding station and other stations in the target subway network as the alighting stations. The average value of the up numbers and the average value of the down numbers at each boarding time are used as the first part of the up and down numbers of the current station at the same current time. For example, when counting the current station a, according to the processing in step 2, it is counted that the boarding times of the passengers in the set of individual passengers include 8 am, 10 am, and 5 pm, a total of 3 boarding times. These 3 boarding times are the corresponding time periods of 3 days, namely 1 day, 2 days, and 3 days. The average value of the total up and down numbers at each boarding time for these 3 days is used as the up and down numbers at the same current boarding time. Therefore, the average value of the up numbers at 8 am for 3 days is used as the up numbers at 8 am of the current time, the average value of the up numbers at 10 am for 3 days is used as the up numbers at 10 am of the current time, and the average value of the up numbers at 5 pm for 3 days is used as the up numbers at 5 pm of the current time.

[0097] According to the processing in step 3, passengers without a unique identity identifier at the current time are counted in real time, and the number of passengers in the up and down directions of the second part is counted based on the ticket purchase information of passengers without a unique identity identifier. Since the passengers without a unique identity identifier purchase single-trip tickets, and the single-trip tickets include the boarding station and the alighting station, the number of passengers in the up and down directions of this part of passengers can be counted in real time.

[0098] The sum of the number of passengers in the up direction of the first part and the number of passengers in the up direction of the second part is used as the predicted number of passengers in the up direction at the current time of the current station, and the sum of the number of passengers in the down direction of the first part and the number of passengers in the down direction of the second part is used as the predicted number of passengers in the down direction at the current time of the current station, thus completing the prediction of the number of passengers in the up and down directions of the subway passenger flow.

[0099] Compared with the conventional subway passenger flow prediction method, the present invention can predict the passenger flow intensity from each current station as the boarding station to any other station as the alighting station based on the up and down passenger flow prediction, including being able to adapt to transfer stations, and can play a role in supporting data for subsequent passenger flow intensity management and guidance. In addition, the present invention takes into account the individual behavior and group behavior of passengers. For the individual behavior of passengers, it has a certain regularity and stability, while for the group behavior of passengers, clustering can also be performed through a clustering algorithm to predict the possible paths of these passengers without a unique identity identifier, thus avoiding the contingency caused by the inability to locate the group behavior of passengers from interfering with the prediction results, reducing accidental behaviors, and further increasing the accuracy of the up and down passenger flow prediction of the subway station. Finally, when determining the passenger's travel path, it is not simply to assign a single path to the passenger, but it is considered that the passenger has multiple paths to choose from, and the probability of each path selection is different. Based on probability theory, the passenger's travel path is determined, which is more in line with the actual travel behavior of passengers and can also avoid reducing the accuracy of the model due to the problem of single path selection. At the same time, when determining the specific travel path of passengers, the priorities of multiple factors such as date type (N days before holidays, holidays, peak return periods after holidays, peak morning and evening periods on weekdays), special events, and station types are fully considered, making the passenger's path closer to their actual travel path.

[0100] Embodiment 2

[0101] Reference Figure 3 , this embodiment also provides a processing terminal, which includes:

[0102] A memory 101 for storing program instructions;

[0103] A processor 102 for running the program instructions to execute the steps of the subway up and down passenger flow prediction method.

[0104] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices create means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0105] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0106] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0107] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0108] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for predicting the up and down passenger flow of the subway, characterized in that, It includes the following steps: Step 1: Obtain the historical passenger travel records of the target subway network within a unit time period. The historical passenger travel records include OD data. Group the passengers with unique passenger identity identifiers into a passenger individual set, and group the remaining passengers into a passenger group set; Step 2: Calculate the travel probability of each passenger in the passenger individual set choosing any real travel path according to the OD data. The real travel path is characterized by the reliability of the passenger's real travel path; Step 3: Convert the passengers in the passenger group set into passenger individuals according to the OD data, and thus regard them as the passengers in the passenger individual set in Step 2, so as to calculate the travel probability of the passengers in the passenger group set choosing any real travel path; Step 4: Determine all possible paths from the current station as the inbound station to other stations as the outbound stations. Divide the passengers entering the current station into a current passenger individual set and a current passenger group set according to whether they have a unique identity identifier, Calculate the number of up and down passengers at the inbound time for the passengers in the current passenger individual set according to the travel probability obtained in Step 2. Take the average value of the number of up passengers and the average value of the number of down passengers at each inbound time as the first part of the up and down passenger numbers of the current station at the same current time, Calculate the second part of the up and down passenger numbers at each inbound time in real time for the passengers in the current passenger group set according to the travel probability obtained in Step 3, Take the sum of the first part of the up passenger number and the second part of the up passenger number at the same inbound time as the predicted up passenger number at the current time of the current station, and take the sum of the first part of the down passenger number and the second part of the down passenger number at the same inbound time as the predicted down passenger number at the current time of the current station, so as to complete the prediction of the up and down passenger numbers of the subway passenger flow, In Step 3, converting the passengers in the passenger group set into passenger individuals according to the OD data, its specific implementation includes: Group the passengers in the passenger group set according to the OD data. Those with the same OD data are grouped into the same group, and those with different OD data are grouped into different groups. Construct several features for the passengers within each group of OD data according to the OD data, Use the clustering algorithm to cluster each group of OD data according to the constructed features to obtain a clustering result. One clustering result corresponds to several classification results. One classification result contains one clustering center, and one classification result corresponds to several passengers, Take the average value of the theoretical travel times of all passengers under each clustering center as the travel time of the current clustering center, so as to calculate the travel time corresponding to each clustering center. Regard each clustering center as a virtual passenger, Thus, the passengers in the passenger group set are converted into passenger individuals.

2. The subway up-and-down passenger flow prediction method according to claim 1, characterized in that In Step 2, calculate all possible paths corresponding to each OD data of each passenger in the passenger individual set, and determine the real travel path corresponding to each OD data from all possible paths. One OD data corresponds to several real travel paths. The possible path refers to the travel route composed of all available stations that can reach the outbound station from the inbound station.

3. The subway up and down passenger flow prediction method according to claim 2, wherein In step 2, the real ride path is determined from all possible paths, and its specific implementation includes the following steps: If the current possible path does not contain a transfer station, the current possible path is used as the real ride path of the current OD data. If the current possible path contains a transfer station, the current possible path that meets Condition 1 is taken as a member of the first set of possible paths, and all possible paths corresponding to the OD data are traversed to obtain the first set of possible paths. Condition 1: The error between the theoretical travel time T of the current possible path and the real travel time L corresponding to the OD data is within the preset threshold time t. If there is only one possible path in the first set of possible paths, this possible path is used as the real ride path of the current OD data. If there are two or more possible paths in the first set of possible paths, then continue to delete the possible paths that meet any of the situations in Condition 2 from the first set of possible paths to obtain the second set of possible paths. Condition 2 is the condition formed based on the entry time and the surrounding locations covered by the possible path. If there is only one possible path in the second set of possible paths, this possible path is used as the real ride path corresponding to the current OD data; if there are zero possible paths in the second set of possible paths, then the current possible path corresponding to the highest priority in the corresponding situation of Condition 2 is used as the real ride path corresponding to the current OD data; if there are two or more possible paths in the second set of possible paths, it is determined whether there is a possible path that meets Condition 3. If so, the possible path that meets Condition 3 is used as the real ride path corresponding to the current OD data. If not, a path is randomly selected from the second set of possible paths as the real ride path corresponding to the current OD data. Condition 3 is the condition formed based on the transfer station of the possible path and the time at the transfer station and the preset time.

4. The subway up and down passenger flow prediction method according to claim 3, characterized in that In step 2, the value of the preset threshold time t should be such that at least one possible path meets Condition 1.

5. The subway up and down passenger flow prediction method according to claim 3, wherein In step 2, Condition 3 specifically is: the current possible path contains a transfer station, and the transfer time of at least one transfer station is less than q times the theoretical travel time of other possible paths, where 0 < q < 1.

6. The subway up and down passenger flow prediction method according to claim 3, wherein, In step 2, Condition 2 includes the following four situations: a. The entry time is during the morning rush hour on a weekday, the exit time is during the evening rush hour on a weekday, and the transfer station in the current possible path is a station covered by an office area or a commercial area. b. The day of the entry time is a holiday or a weekend, and the current possible path contains an intermediate station within a preset range from a tourist attraction or a commercial center. An intermediate station refers to other stations in the current possible path except the entry station and the exit station. c. The day of the entry time is n days before a holiday or during the peak return period of a holiday, where n ≥ 1, and the current possible path contains an intermediate station covering any one of the transportation hubs such as an airport, a railway station, a passenger station, or a wharf. d. The day of the entry time is during a special event or a large-scale event, and the current possible path contains an intermediate station passing through the location of the special event or the large-scale event.

7. The subway up and down passenger flow prediction method according to claim 1, wherein In step 2, the specific implementation of the boarding probability of each passenger choosing any real boarding path includes the following steps: Group all OD data in the historical boarding records. The same OD data is grouped into the same group, and different OD data is grouped into different groups. In the same group of OD data, the ratio of the sum of the current passenger's identical real boarding paths to the sum of all real boarding paths under this group of OD data is used as the probability of the current real boarding path, so as to calculate the boarding probability of choosing any real boarding path in a group of OD data.

8. The subway up-and-down passenger flow prediction method according to claim 1, characterized in that In step 3, calculating the boarding probability of passengers in the passenger group set choosing any real boarding path, its specific implementation includes: After obtaining the virtual passengers, the theoretical travel time of the virtual passengers is the travel time of the cluster center. All possible paths of each OD data of the virtual passengers and the corresponding theoretical travel time of each possible path are obtained. The virtual passengers are processed according to step 2, and the boarding probability and theoretical travel time of each virtual passenger choosing each real boarding path are calculated. Calculate the boarding path formed by the virtual passengers from the current station to the target station and the probability sub_path_probability of each sub-path under this real boarding path according to formula ③: ------③ In the formula, Q represents the total number of groups, J represents the total number of sub-paths included in the real boarding path corresponding to each cluster center, M represents the total number of all real boarding paths corresponding to each cluster center, and prob represents the boarding probability of each real boarding path calculated according to step 2.

9. A processing terminal, characterized in that, It includes: A memory for storing program instructions; A processor for running the program instructions to execute the steps of the subway up and down passenger flow prediction method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Distribution method for urban rail transit

    CN103761589A

  • Subway short-term passenger flow forecasting method based on LS-SVM and real-time big data

    CN109308543A