A method and system for dynamically estimating all-day passenger flow distribution of urban rail transit network
By constructing a knowledge base linking weather and passenger flow and a dual matching mechanism, the problem of the difficulty in considering the impact of meteorological conditions and emergencies in existing technologies has been solved, and dynamic and accurate estimation of the passenger flow distribution of urban rail transit throughout the day has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 湖南承希科技有限公司
- Filing Date
- 2025-12-09
- Publication Date
- 2026-06-09
AI Technical Summary
Existing methods for predicting urban rail transit passenger flow are difficult to accurately account for the impact of weather conditions and emergencies, resulting in a large deviation between the estimated results and the actual passenger flow. Furthermore, the lack of a systematic similarity measurement criterion affects the reliability of the estimated results.
By integrating card swipe records, meteorological information, and event log data, a knowledge base linking weather and passenger flow is constructed to quantify the impact of emergencies. A dual matching mechanism of periodic attributes and waveform morphology is used to filter historical reference cases, enabling dynamic estimation of passenger flow distribution throughout the day.
It improves the accuracy and reliability of passenger flow estimation, can dynamically adapt to the long-term evolution of passenger flow patterns in the network, adapt to sudden event disturbances in different areas, and achieve accurate estimation of passenger flow distribution throughout the day.
Smart Images

Figure CN121658842B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban rail transit technology, and in particular to a method and system for dynamically estimating the daily passenger flow distribution of an urban rail transit network. Background Technology
[0002] As the backbone of public transportation in large cities, urban rail transit bears increasing passenger load. Accurately predicting passenger flow patterns throughout the day is crucial for train schedule planning, platform passenger flow management, and staff allocation. However, rail transit passenger flow is influenced by a variety of factors, including changes in weather conditions, large-scale events in the surrounding area, and temporary equipment malfunctions. These factors intertwine, resulting in complex spatiotemporal fluctuations in passenger flow.
[0003] Existing passenger flow forecasting methods primarily rely on statistical analysis or time series modeling of historical passenger flow data, with limited consideration of the impact of weather conditions on passenger travel intentions and a lack of quantitative assessment mechanisms for the disruptive effects of unforeseen events. When encountering rainfall or station equipment malfunctions, traditional methods often struggle to adjust forecasting strategies in a timely manner, leading to significant discrepancies between estimated results and actual passenger flow. Furthermore, existing methods lack systematic similarity metrics when selecting historical reference cases, making it difficult to simultaneously consider both date-period attributes and passenger flow pattern characteristics, thus affecting the reliability of the estimation results. Summary of the Invention
[0004] This invention discloses a method and system for dynamically estimating the daily passenger flow distribution of an urban rail transit network. It aims to integrate multi-source data such as card swipe records, meteorological information, and event logs to construct a knowledge base linking weather and passenger flow, quantify the impact of emergencies, and filter historical reference cases based on a dual matching mechanism of periodic attributes and waveform morphology to achieve dynamic estimation of the daily passenger flow distribution of the network, providing data support for operation scheduling decisions.
[0005] The first aspect of this invention proposes a method for dynamically estimating the daily passenger flow distribution of an urban rail transit network, comprising the following steps:
[0006] Acquire passenger card swipe data, historical weather data, and special event record data from the AFC system, and perform time-series alignment on the passenger card swipe data and the historical weather data to generate a time-series coupled dataset;
[0007] The passenger card swipe data is used to perform time-period passenger volume statistics to generate a passenger volume distribution sequence. Passenger flow fluctuation characteristics are extracted from the passenger volume distribution sequence, and the passenger flow fluctuation characteristics are mapped to the historical weather data to form a weather-passenger flow knowledge graph.
[0008] The special event record data is parsed to generate an event coding sequence, and the event coding sequence is evaluated to generate an event impact factor.
[0009] Based on the weather-passenger flow knowledge graph, a similar case set is generated by performing similarity retrieval on the time-series coupled dataset. The similar case set is then filtered using the event influence factor to generate a filtered case set. The filtered case set is then subjected to time distance calculation and weekday attribute matching to generate periodic similarity. Based on the passenger flow fluctuation characteristics, the filtered case set is then subjected to morphological contour matching to generate waveform fit.
[0010] Based on the period similarity and the waveform fit, the selected case set is subjected to confidence modulation to generate a passenger volume estimate. The passenger volume estimate is compared with the actual passenger volume to generate an optimized weight parameter. Based on the optimized weight parameter, the estimated passenger flow distribution for the whole day is output to the operation management terminal.
[0011] A second aspect of this invention proposes a dynamic estimation system for the all-day passenger flow distribution of an urban rail transit network, comprising:
[0012] The data acquisition module is used to acquire passenger card swipe data, historical weather data, and special event record data from the AFC system, and to perform time-series alignment on the passenger card swipe data and the historical weather data to generate a time-series coupled dataset.
[0013] The feature extraction module is used to perform time-period passenger volume statistics on the passenger card swiping data to generate a passenger volume distribution sequence, extract passenger flow fluctuation features from the passenger volume distribution sequence, and map the passenger flow fluctuation features to the historical weather data to form a weather-passenger flow knowledge graph.
[0014] The event parsing module is used to perform event type parsing on the special event record data to generate an event coding sequence, and to evaluate the impact range of the event coding sequence to generate an event impact factor.
[0015] The case retrieval module is used to perform similarity retrieval on the time-series coupled dataset based on the weather-passenger flow knowledge graph to generate a set of similar cases, filter the set of similar cases through the event influence factor to generate a set of filtered cases, perform time distance calculation and weekday attribute matching on the set of filtered cases to generate periodic similarity, and perform morphological contour matching on the set of filtered cases based on the passenger flow fluctuation characteristics to generate waveform fit.
[0016] The estimation output module is used to generate a passenger volume estimate by performing confidence modulation on the selected case set based on the period similarity and the waveform fit, perform error analysis on the passenger volume estimate and the actual passenger volume to generate optimized weight parameters, and output the estimated passenger flow distribution for the whole day to the operation management terminal based on the optimized weight parameters.
[0017] The beneficial effects of this invention are reflected in the following points: First, the constructed weather-passenger flow knowledge graph structurally expresses the correlation between meteorological conditions and passenger flow tidal patterns. By quantifying the confidence of the correlation between different weather types and passenger flow characteristics through edge weights, passenger flow estimation can quickly locate similar historical passenger flow patterns along high-weight paths based on the weather conditions of the day, solving the problem of traditional methods simply using meteorological factors as input features while ignoring their intrinsic correlation with passenger flow fluctuations. Second, in terms of quantifying event impact, for the near-end station group, a transfer coefficient is used to amplify the passenger flow deviation value to reflect the network propagation effect of transfer hubs. For the far-end station group, a dual weighting of distance attenuation and time attenuation is used to characterize the spatiotemporal reduction law of the impact. Compared with the traditional method of treating the event impact as a uniform distribution, this method can more accurately assess the degree of differentiated disturbance of sudden events in different areas of the network. Finally, the confidence score of candidate cases is obtained by a dual matching mechanism of period similarity and waveform fit, which takes into account the periodicity of the week attribute and the morphological characteristics of the peak period. The weight parameters are dynamically adjusted by estimation error feedback, so that the passenger flow estimation model can continuously self-optimize, adapt to the long-term evolution of the network passenger flow pattern, and achieve dynamic and accurate estimation of the passenger flow distribution throughout the day.
[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0019] The accompanying drawings illustrate specific examples of the technical solutions described in this invention and, together with the detailed embodiments, form part of the specification, serving to explain the technical solutions, principles, and effects of this invention.
[0020] Unless otherwise specified or defined, the same reference numerals in different figures represent the same or similar technical features, and different reference numerals may be used to represent the same or similar technical features.
[0021] Figure 1 This is a flowchart illustrating the method for dynamically estimating the daily passenger flow distribution of an urban rail transit network according to the present invention.
[0022] Figure 2 This is a structural block diagram of a dynamic estimation system for the all-day passenger flow distribution of an urban rail transit network according to the present invention. Detailed Implementation
[0023] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0024] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0025] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0026] The technical solutions of the embodiments of this application will be described below.
[0027] like Figure 1 As shown, this embodiment of the invention provides a method for dynamically estimating the daily passenger flow distribution of an urban rail transit network, including the following steps S110-S150:
[0028] Step S110: Obtain passenger card swiping data, historical weather data, and special event record data from the AFC system, and perform time-series alignment on the passenger card swiping data and historical weather data to generate a time-series coupled dataset.
[0029] Specifically, passenger card swipe data, historical weather data, and special event records from the AFC system are acquired. Passenger card swipe data is retrieved from the urban rail transit AFC system database. This data includes entry and exit card swipe records, with each record containing four core fields: card number, station code, swipe time, and transaction type. The passenger card swipe data undergoes integrity verification, removing records with missing timestamps or abnormal station codes, retaining valid data for subsequent analysis. During peak hours, a single subway line can generate hundreds of thousands of passenger card swipe records daily, necessitating timestamp indexing to improve subsequent query efficiency. Historical weather data is obtained from the meteorological department's data interface. This data includes meteorological indicators such as temperature, humidity, rainfall, and wind speed, recorded hourly at each monitoring station. Missing values in the historical weather data are imputed by interpolating the mean of adjacent time periods to form a complete historical weather data sequence. During heavy rain or extreme heat, rainfall and temperature indicators in the historical weather data significantly impact passenger flow fluctuations, requiring the continuity of historical weather data collection. Special event record data is obtained from the operations management platform. This data includes the occurrence time, duration, and impact range of events such as equipment malfunctions, large-scale crowd control, emergencies, and large-scale events in the surrounding area. The special event record data is standardized in format, unifying event records from different sources into a standard format of time-event type-affected site. During holidays or after large concerts, large-scale crowd control events in the special event record data can significantly impact the passenger flow distribution of surrounding sites, necessitating the complete preservation of the spatiotemporal scope information of the special event record data.
[0030] In some embodiments, performing time-series alignment of the passenger card swipe data and the historical weather data to generate a time-series coupled dataset includes: extracting entry and exit timestamps from the passenger card swipe data to form a time-label sequence; associating the time-label sequence and the historical weather data by time period to form a time period-weather correspondence table; identifying weather abrupt change points from the time period-weather correspondence table to form a change marker set; and performing segmented alignment based on the change marker set to form a time-series coupled dataset.
[0031] A time-stamp sequence is formed by extracting entry and exit timestamps from passenger card swipe data. The swipe time field, accurate to the second, is read from each record of passenger card swipe data, recording the precise moment the passenger passed through the gate. The timestamp of the entry swipe record is marked as the entry time, and the timestamp of the exit swipe record is marked as the exit time, distinguishing the inflow and outflow directions of passenger flow. All timestamps in the passenger card swipe data are sorted chronologically to form a time-stamp sequence. The time-stamp sequence records the time nodes of all card swipe actions in the passenger card swipe data, reflecting the distribution characteristics of passenger flow over time. The time-stamp sequence is deduplicated, merging multiple records within the same second, retaining only the statistical value of card swipes for that second. During the morning peak period from 7:30 to 8:30, the time-stamp sequence typically exhibits a dense distribution, potentially containing dozens of card swipe records per second, while during the period before the end of operations in the early morning, the time-stamp sequence exhibits a sparse distribution.
[0032] A time-segment-weather mapping table is formed by associating time-stamped sequences with historical weather data. Based on the timestamps in the time-stamped sequences, the time interval to which each timestamp belongs is determined, with each time interval divided into 15-minute units, resulting in 96 time segments per day. Historical weather data is divided according to the same time intervals, and linear interpolation is used to convert hourly historical weather data into 15-minute data, ensuring consistent time granularity between historical weather data and time-stamped sequences. A mapping relationship is established between the time segments of the time-stamped sequences and weather indicators in historical weather data, forming the time-segment-weather mapping table. Each record in the time-segment-weather mapping table includes the start and end times of the time segment, the number of card swipes within that time segment, and the corresponding weather indicator values such as temperature and rainfall. The time-segment-weather mapping table is checked for completeness to ensure that each time segment has a corresponding weather data record. During sudden afternoon thunderstorms in summer, the rainfall indicator in the time-segment-weather mapping table shows a sudden increase from zero, while the number of card swipes in the corresponding time segment fluctuates significantly. The time-segment-weather mapping table provides a data foundation for analyzing the impact of weather on passenger flow.
[0033] A weather abrupt change point is identified from the time-time-weather correspondence table to form an abrupt change marker set. The magnitude of weather index changes between adjacent time periods in the time-time-weather correspondence table is analyzed, and the inter-time period differences for indicators such as temperature and rainfall are calculated. When the difference in weather indicators between adjacent time periods in the time-time-weather correspondence table exceeds a preset abrupt change threshold, the boundary of that time period is marked as a weather abrupt change point. The temperature abrupt change threshold is set to a change exceeding 3 degrees Celsius within 30 minutes, and the rainfall abrupt change threshold is set to a change from rainless to rainy conditions or a sudden change in rainfall intensity. All time periods in the time-time-weather correspondence table are traversed, and all identified weather abrupt change points are arranged chronologically to form the abrupt change marker set. The abrupt change marker set records the time points in the time-time-weather correspondence table where weather conditions change significantly; these points are usually accompanied by changes in passenger flow patterns. The abrupt change marker set is used to guide subsequent segmentation and alignment operations, using the abrupt change points in the abrupt change marker set as segmentation boundaries. During winter cold waves, the abrupt change marker set records the time points when temperatures plummet; the passenger flow characteristics before and after these points may differ significantly, requiring segmentation processing to improve the accuracy of passenger flow estimation.
[0034] A time-series coupled dataset is formed by segmenting and aligning data based on a mutation marker set. Using weather abrupt changes in the mutation marker set as segment boundaries, the time-series data is divided into multiple time intervals with relatively stable weather conditions, maintaining relatively consistent weather characteristics within each interval. Based on the segment boundaries of the mutation marker set, passenger card-swiping data is timestamped and aligned with weather data within each segment to ensure accurate correspondence between card-swiping counts and weather indicators at the same time. The data from each segment, divided according to the mutation marker set, is then concatenated chronologically to form a complete time-series coupled dataset. The time-series coupled dataset retains the segment boundary information from the mutation marker set and labels each segment with its corresponding weather type, facilitating subsequent analysis and identification of passenger flow patterns under different weather conditions. The time-series coupled dataset undergoes quality checks to ensure accurate correspondence between card-swiping data and weather data for each time period, with no time misalignment or data loss. Taking a weekday as an example, if a sudden rainfall occurs at 14:00 on that day, the mutation marker set will mark 14:00 as the mutation point. The time-series coupled dataset is divided into two segments, before and after the rainfall, so that the subsequent passenger flow estimation model can establish prediction rules for different weather conditions.
[0035] Step S120: Perform time-period passenger volume statistics on passenger card swipe data to generate a passenger volume distribution sequence, extract passenger flow fluctuation characteristics from the passenger volume distribution sequence, and map the passenger flow fluctuation characteristics to historical weather data to form a weather-passenger flow knowledge graph.
[0036] Specifically, passenger card swipe data is statistically analyzed by time period to generate a passenger volume distribution sequence. Passenger card swipe data is grouped and aggregated by time period, with a statistical granularity of 15 minutes, accumulating the number of swipe records within each time period. Entry and exit records in the passenger card swipe data are statistically analyzed separately; the number of entry swipes reflects passenger inflow, and the number of exit swipes reflects passenger outflow. The number of passengers entering and exiting the station for each time period is recorded separately, forming a passenger volume distribution sequence. The passenger volume distribution sequence is indexed by time period and includes three indicators: passenger volume entering the station, passenger volume exiting the station, and total passenger volume for that time period. The passenger volume distribution sequence is smoothed using a 3-point moving average to eliminate occasional abnormal fluctuations in passenger card swipe data. The smoothed passenger volume calculation formula is: V_s(t)=[V(t-1)+V(t)+V(t+1)] / 3, where V(t) is the original passenger volume at time t, and V_s(t) is the smoothed passenger volume. Smoothing makes the passenger volume distribution sequence better reflect the overall trend of passenger flow and eliminates random disturbances in individual time periods. During the morning rush hour on weekdays, the passenger volume entering stations between 7:30 and 8:30 typically reaches its peak for the day, while the passenger volume exiting stations shows peak characteristics at commercial area stations during this period. On weekends or holidays, the peak period of the passenger volume distribution sequence shifts to around 10:00, and the peak intensity is relatively lower than on weekdays.
[0037] In some embodiments, extracting passenger flow fluctuation characteristics from the passenger volume distribution sequence includes: performing tidal gradient analysis on the passenger volume distribution sequence to form a passenger flow change rate curve; identifying commuter tidal critical points from the passenger flow change rate curve to form a peak-valley marker sequence; performing time period positioning on the peak-valley marker sequence to form peak periods and valley periods; and forming passenger flow fluctuation characteristics based on the tidal amplitude intensity of the peak periods and valley periods.
[0038] Tidal gradient analysis is performed on the passenger volume distribution sequence to generate a passenger flow change rate curve. The difference in passenger volume between adjacent time periods in the passenger volume distribution sequence is calculated; this difference reflects the increase or decrease in passenger flow per unit time. Dividing the passenger volume difference between time periods in the passenger volume distribution sequence by the time interval yields the passenger flow change rate for each time period. The formula for the change rate is: R(t) = [V(t) - V(t-1)] / Δt, where V(t) is the passenger volume at time period t, and Δt is the time interval. A positive change rate R(t) indicates that passenger flow is in an upward phase, while a negative change rate indicates that passenger flow is in a downward phase. Arranging the change rates of all time periods in the passenger volume distribution sequence in chronological order forms the passenger flow change rate curve. The passenger flow change rate curve describes the dynamic changes in the passenger volume distribution sequence; the alternation of positive and negative values in the curve reflects the tidal fluctuation pattern of passenger flow. The passenger flow rate change curve was normalized using the formula: R_norm(t) = R(t) / R_max, where R_max is the maximum absolute value of the rate of change. The normalized rate of change ranges from -1 to 1, facilitating horizontal comparisons of passenger flow rate change curves across different dates and lines. On typical weekdays, the passenger flow rate change curve exhibits a sustained positive value and a rapid increase between 6:00 AM and 7:30 AM, indicating a rapid influx of passengers into the subway network. This characteristic reflects the formation process of the morning rush hour.
[0039] For example, the step of identifying commuter tidal critical points from the passenger flow change rate curve to form a peak-valley marker sequence includes: performing an operating period window search based on the passenger flow change rate curve to form a candidate extreme point set; classifying and aggregating the candidate extreme point set by peak and off-peak periods to form an extreme point cluster; selecting representative points from the extreme point cluster to form commuter tidal critical points; and generating a peak-valley marker sequence according to the time period distribution of the commuter tidal critical points.
[0040] A candidate extreme value set is formed by searching the passenger flow change rate curve during operating hours. A sliding search window is set on the passenger flow change rate curve, with a window width of 30 minutes, covering two adjacent statistical periods. The search window is moved step by step along the time axis of the passenger flow change rate curve, with a step size of one statistical period. Local maxima and local minima of the curve are identified at each window position. When the value of the passenger flow change rate curve at a certain moment is greater than the values at other moments within the window, that moment is marked as a local maximum, corresponding to the moment when the passenger flow increase rate reaches a local peak. When the value of the passenger flow change rate curve at a certain moment is less than the values at other moments within the window, that moment is marked as a local minimum, corresponding to the moment when the passenger flow decrease rate reaches a local peak or passenger flow growth tends to stagnate. All local extreme values found on the passenger flow change rate curve are summarized to form a candidate extreme value set. The candidate extreme value set contains all possible tidal turning points in the passenger flow change rate curve, but it may contain spurious extreme values caused by noise, requiring further screening. Extreme points with an absolute change rate less than 20% of the mean in the candidate extreme point cluster are marked. These low-significance extreme points are preferentially removed in subsequent screening. At transfer stations with frequent passenger flow fluctuations, the number of extreme points in the candidate extreme point cluster may reach more than 10, while at suburban stations with relatively stable passenger flow, the number of extreme points in the candidate extreme point cluster is usually between 4 and 6.
[0041] The candidate extreme point set is classified and aggregated into extreme point clusters based on peak and off-peak periods. Extreme points are categorized into different time periods according to their temporal location. Extreme points located between 6:30 and 9:30 are assigned to the morning peak category, those between 17:00 and 20:00 to the evening peak category, and those in other time periods to the off-peak category. Spatially, extreme points within the same category are spatially aggregated. Adjacent extreme points with time intervals less than 15 minutes are merged into the same cluster using a temporal proximity criterion. The rate of change of each extreme point is retained during merging for subsequent representative point selection. The combinations of extreme points formed after classification and aggregation are defined as extreme point clusters, each cluster corresponding to a characteristic time period within a tidal cycle. Extreme point clusters organize the scattered extreme points in the candidate extreme point set into structured categories, with extreme points within each cluster sharing similar temporal locations and tidal characteristics. The validity of extreme point clusters is verified. If a cluster contains only a single low-significance extreme point, the cluster is marked as a noise cluster and removed in subsequent processing. On a typical workday, extreme point clusters usually form several main categories, such as morning peak rising clusters, morning peak peak clusters, midday trough clusters, evening peak peak clusters, and nighttime trough clusters.
[0042] Representative points are selected from extreme point clusters to form commuter tidal critical points. Within each extreme point cluster, the most representative point is selected based on the significance of the extreme points as the representative of that cluster. The absolute value of the passenger flow change rate corresponding to each extreme point in the extreme point cluster is calculated, and the extreme point with the largest absolute value of the change rate is selected as the representative point of that extreme point cluster. The representative point has the highest absolute value of the passenger flow change rate among all extreme points in the cluster. If there are multiple extreme points with similar change rates in an extreme point cluster, and the difference in change rates is less than 5%, then the extreme point with the middle time position is selected as the representative point to more accurately reflect the center moment of the tidal turning point. The representative points selected from each extreme point cluster are summarized to form the commuter tidal critical point set. The commuter tidal critical points are the most representative turning points selected from the extreme point clusters, and their number is much smaller than the original candidate extreme points, usually 4 to 6. Each commuter tidal critical point is labeled with its corresponding extreme point cluster category and extreme value type, including peak and valley types. The passenger flow change rate corresponding to each critical point is also recorded as a significance indicator. A temporal sequence verification of the commuter tidal critical points is performed to ensure that the extreme value types of adjacent critical points alternate. If consecutive critical points of the same type appear, the aggregation results of the extreme point clusters are checked back. On weekdays with distinct morning and evening peak characteristics, commuter tidal critical points typically include key nodes such as the morning peak start point, morning peak point, evening peak start point, and evening peak point.
[0043] A peak-valley marker sequence is generated based on the time-segment distribution of commuting tide critical points. The commuting tide critical points are arranged chronologically to form a time-ordered sequence. Based on the extreme value type of each critical point, a peak or valley marker is assigned; local maxima are marked as peaks, and local minima as valleys. The timestamp, extreme value type, and time-segment category of each commuting tide critical point are encapsulated as marker records in the format (timestamp, peak / valley marker, time-segment category), arranged chronologically to form the peak-valley marker sequence. Each record in the peak-valley marker sequence also includes a significant value of the rate of change of that critical point, used to determine the expansion width when expanding subsequent time-segment boundaries. Logical verification is performed on the peak-valley marker sequence to ensure that peak and valley markers alternate. If consecutive markers of the same type appear, the identification results of the commuting tide critical points need to be checked back. When the time interval between adjacent markers in the peak-valley marker sequence is less than 30 minutes, it is checked for over-segmentation, and adjacent time periods changing in the same direction are merged if necessary. On weekends or holidays, the distribution of peaks and valleys in the peak-valley marker sequence may differ significantly from that on weekdays. The peak time may shift later and the number of peaks and valleys may decrease to 3 to 4. For example, the peak time of the morning rush hour on weekends may shift from 8:00 on weekdays to around 10:30.
[0044] The peak and valley marker sequences are used to determine peak and valley periods. Based on the timestamps of each critical point in the sequence, the time interval of each critical point is determined. Using the peak critical point as the center, several time periods are extended forward and backward to form the boundary of the peak period; the width of this extension is dynamically determined based on the duration of passenger flow increases and decreases. Similarly, using the valley critical point as the center, the boundary of the valley period is formed by extending forward and backward. The morning peak point in the sequence is located to determine the start and end times of the morning peak period; on a typical weekday, the morning peak period is usually from 7:15 to 8:45. The evening peak point is also located to determine the start and end times of the evening peak period; on a typical weekday, the evening peak period is usually from 17:30 to 19:00. Valley periods include midday valleys and nighttime valleys; midday valleys are usually from 10:00 to 11:30, and nighttime valleys are usually from 21:00 until the end of operations. The precise boundaries between peak and trough periods are dynamically adjusted based on the actual location of the critical points in the peak-trough marker sequence. The time period boundaries may differ for different routes and different dates.
[0045] Passenger flow fluctuation characteristics are formed based on the tidal amplitude intensity during peak and trough periods. The average passenger volume during peak periods is calculated, denoted as V_peak, reflecting the passenger flow intensity level during peak hours. The average passenger volume during trough periods is calculated, denoted as V_valley, reflecting the baseline passenger flow level during off-peak periods. The difference between the average passenger volume during peak and trough periods is defined as the tidal amplitude, calculated using the formula: A = V_peak - V_valley. A larger tidal amplitude indicates a more significant difference between peak and off-peak passenger flow. The ratio of the average passenger volume during peak periods to the average passenger volume during trough periods is defined as the tidal intensity ratio, calculated using the formula: S = V_peak / V_valley. The tidal intensity ratio reflects the relative amplitude of passenger flow fluctuations. Integrating indicators such as the tidal amplitude during the morning peak, the tidal amplitude during the evening peak, and the tidal intensity ratio during the morning and evening peak periods, passenger flow fluctuation characteristics are formed. Tidal amplitude is normalized using the formula: A_norm = A / A_max, where A is the current tidal amplitude and A_max is the maximum tidal amplitude for the same historical period. The normalized amplitude ranges from 0 to 1. Passenger flow fluctuation characteristics are represented as a vector, denoted as F = [A_norm_morning, A_norm_evening, S_morning, S_evening]. Each dimension of the vector is a dimensionless indicator, facilitating subsequent feature comparison and knowledge graph construction, and comprehensively describing the fluctuation pattern of passenger volume over time. At shuttle stations in large residential areas, the morning peak tidal amplitude is usually greater than the evening peak tidal amplitude, reflecting that these stations are dominated by commuter passengers who leave early and return late.
[0046] In some embodiments, mapping the passenger flow fluctuation features to the historical weather data to form a weather-passenger flow knowledge graph includes: materializing the passenger flow fluctuation features into nodes to generate a passenger flow feature node set; mapping the passenger flow feature node set to the historical weather data to form a weather-passenger flow association edge; quantifying the association confidence through the weather-passenger flow association edge to generate an edge weight set; and constructing a weather-passenger flow knowledge graph based on the passenger flow feature node set, the weather-passenger flow association edge, and the edge weight set.
[0047] The passenger flow fluctuation characteristics are entityized into a set of passenger flow feature nodes. Each dimension of the passenger flow fluctuation characteristics is transformed into entity nodes in a knowledge graph, with each node representing a type of passenger flow fluctuation characteristic. The tidal amplitude index in the passenger flow fluctuation characteristics is discretized and classified into three levels—high amplitude, medium amplitude, and low amplitude—based on the comparison between the amplitude value and the historical average. High amplitude corresponds to an amplitude value greater than 1.2 times the historical average, low amplitude corresponds to an amplitude value less than 0.8 times the historical average, and medium amplitude corresponds to the remaining cases. The tidal intensity ratio index in the passenger flow fluctuation characteristics is also discretized and classified into three levels: strong fluctuation, medium fluctuation, and weak fluctuation. Strong fluctuation corresponds to an intensity ratio greater than 3, and weak fluctuation corresponds to an intensity ratio less than 1.5. The time location of peak periods in the passenger flow fluctuation characteristics is classified, based on the standard peak time, into categories such as early peak shift, standard early peak, and late peak. A shift more than 30 minutes earlier is classified as early peak, and a shift more than 30 minutes later is classified as late peak. All entity nodes obtained from the transformation of passenger flow fluctuation characteristics are aggregated to form a passenger flow characteristic node set. The passenger flow characteristic node set contains three types of entities: amplitude level nodes, fluctuation intensity nodes, and peak period type nodes. Each node is named in the format of "category_level", such as "amplitude_high", "fluctuation_strong", "early peak_early shift", etc. On routes with a high proportion of commuter passenger flow, high amplitude nodes and strong fluctuation nodes appear more frequently in the passenger flow characteristic node set.
[0048] A weather-passenger flow association edge is formed by mapping the passenger flow feature node set with historical weather data. The dates corresponding to each node in the passenger flow feature node set are analyzed, and the weather conditions recorded in the historical weather data for those dates are found. Weather types in the historical weather data are converted into weather entity nodes, including categories such as sunny, cloudy, rainy, high-temperature, and low-temperature nodes. High temperature is defined as a daily maximum temperature exceeding 35 degrees Celsius, and low temperature is defined as a daily minimum temperature below 5 degrees Celsius. A correspondence is established between passenger flow nodes in the passenger flow feature node set and weather nodes in the historical weather data. When a passenger flow node and a weather node appear simultaneously on a certain date, an association edge is created between them. The pairs of nodes that appear simultaneously in the passenger flow feature node set and the historical weather data are counted, and the association edges between all node pairs are summarized to form a weather-passenger flow association edge set. The weather-passenger flow association edge describes the co-occurrence relationship between passenger flow characteristics in the passenger flow feature node set and weather conditions in the historical weather data. During the rainy season, the number of edges connecting rainy nodes and high-amplitude nodes in the weather-passenger flow association edge is usually larger, reflecting the impact of rainfall on passenger flow fluctuations.
[0049] Edge weight sets are generated by quantifying the confidence of weather-passenger flow association edges. The number of times each node pair connected by each edge in the weather-passenger flow association edge co-occurs in historical data; a higher co-occurrence frequency indicates a more stable association. A confidence index is calculated for each edge in the weather-passenger flow association edge, defined as the ratio of the number of node pair co-occurrences to the total number of weather node occurrences, calculated as: Conf(A→B)=Count(A∩B) / Count(A), where A is the weather node, B is the passenger flow feature node, and the confidence value ranges from 0 to 1. Edges with confidence values below a threshold of 0.1 are pruned, removing occasional weak association edges. The confidence value of each retained edge in the weather-passenger flow association edge is recorded to form an edge weight set. The edge weight set assigns a quantifiable strength value to each edge in the weather-passenger flow association edge; a higher weight indicates a stronger association between weather conditions and passenger flow characteristics. In the edge weight set, the edge weight connecting the rainy day node and the high amplitude node is usually relatively high, reaching 0.6 to 0.8, indicating that rainy days are strongly correlated with high fluctuations in passenger flow.
[0050] A weather-passenger flow knowledge graph is constructed based on a set of passenger flow feature nodes, weather-passenger flow association edges, and edge weights. All nodes in the passenger flow feature node set are designated as passenger flow feature layer nodes in the knowledge graph, and weather type nodes are designated as weather condition layer nodes. Weather-passenger flow association edges are used as cross-layer edges connecting nodes in the knowledge graph, with each edge representing the association between weather conditions and passenger flow features. Weights from the edge weight set are assigned to the attribute fields of the corresponding edges in the knowledge graph, giving each edge a quantifiable association strength. The weather-passenger flow knowledge graph is then constructed by integrating the passenger flow feature node set, weather-passenger flow association edges, and edge weights. The knowledge graph organizes the association between weather and passenger flow in a graph structure, represented as G=(V_flow∪V_weather,E,W), where V_flow is the passenger flow feature node set, V_weather is the weather node set, E is the association edge set, and W is the edge weight set. When estimating passenger flow, the weather-passenger flow knowledge graph can be searched for related passenger flow feature nodes based on the weather conditions of the day to obtain the typical passenger flow fluctuation pattern under the weather conditions. When searching, the associated path with the highest edge weight is selected first.
[0051] Step S130: Perform event type parsing on special event record data to generate event coding sequences, and evaluate the impact range of the event coding sequences to generate event impact factors.
[0052] Specifically, event type parsing is performed on special event record data to generate an event coding sequence. The event description field in the special event record data is read, and the event's type category is determined based on the description. Equipment failure events in the special event record data are further subdivided into subtypes such as turnstile failure, escalator failure, signal failure, and vehicle failure, with each subtype assigned a unique type code. Operational control events in the special event record data are further subdivided into subtypes such as large passenger flow restrictions, temporary shutdowns, and extended operating hours, with each subtype assigned a corresponding type code. External impact events in the special event record data are further subdivided into subtypes such as surrounding performances, sporting events, and severe weather warnings, with each subtype assigned a corresponding type code. The type code, occurrence time, duration, and affected sites of each event record in the special event record data are encapsulated into a coded record, arranged chronologically to form an event coding sequence. The event coding sequence is a structured representation of the special event record data, transforming unstructured event descriptions into a computable coded form. At sites near large sports stadiums, sports event codes appear frequently in event coding sequences, especially during the event's closing time when multiple related coding records appear in clusters.
[0053] In some embodiments, the step of assessing the impact range of the event encoding sequence to generate an event impact factor includes: identifying the event occurrence sites based on the event encoding sequence to form an event center point; performing spatial radiation analysis on the event center point to form an impact coverage area; statistically analyzing the affected sites within the impact coverage area to form an impact site set; and generating an event impact factor based on the passenger flow deviation and time decay characteristics of the impact site set.
[0054] The event center is identified by identifying the site where the event occurred based on the event coding sequence. The affected site field of each coding record in the event coding sequence is read; this field records the site code where the event directly occurred or was first affected. For equipment failure events in the event coding sequence, the affected site is the site where the faulty equipment is located, and the site code is retrieved from the equipment asset management database. For operation control events in the event coding sequence, the affected site is the site where flow restriction or shutdown measures are implemented, and the site code is extracted from the operation dispatch instructions. For external impact events in the event coding sequence, the affected site is the subway station closest to the event location, and the site code is calculated based on the distance between the event location coordinates and the station coordinates. The affected sites of each event in the event coding sequence are extracted and marked as the event center. The event center records the code, latitude and longitude coordinates, and line information of the site where the event occurred. When a signal failure occurs at a transfer hub station, that station is marked as the event center, and subsequent analysis radiates outwards from this point to surrounding stations. The event center is geographically labeled to obtain its latitude and longitude location information in the network space. In incidents occurring at transfer hubs, the impact of the incident center usually spreads outward along multiple lines, affecting a wider area than at ordinary stations.
[0055] Spatial radiation analysis is performed on the event's epicenter to form an impact coverage area. Using the event epicenter as the center, an initial radiation radius is set according to the event type: for equipment failure events, the radius is set to 3 station intervals; for large passenger flow control events, the radius is set to 5 station intervals; and for external activity events, the radius is set to 4 station intervals. The impact area is extended outwards from the event epicenter along the subway line, including line sections within the radiation radius. The radiation distance is measured by the number of stations rather than geographical distance, because passenger travel decisions are primarily influenced by station accessibility rather than physical distance. For event epicenters at transfer stations, radiation analysis is conducted along each intersecting line, forming multi-directional radiation paths. The number of radiation paths equals the number of transfer lines at the station where the event epicenter is located. All line sections and station areas within the event epicenter's radiation range are merged to form the impact coverage area. The impact coverage area describes the spatial extent that the event's impact may reach, exhibiting a spatial distribution characteristic that decreases outwards from the event epicenter. The boundary of the impact coverage area is dynamically adjusted according to the severity of the event; the impact coverage area for severe events can extend to 1.5 times the initial radiation radius. When a fault occurs at a station at the end of the line, the affected coverage area mainly extends in one direction, while when a fault occurs at a station in the middle of the line, the affected coverage area expands in both directions simultaneously.
[0056] A set of affected sites is formed by statistically analyzing the affected sites within the impact coverage area. All sites within the impact coverage area are traversed, and their codes and attribute information are extracted. The distance between each site and the event center point within the impact coverage area is calculated, and the topological distance of each site to the event center point is recorded. The topological distance is defined as the minimum number of sites between two sites. Sites within the impact coverage area are classified according to their passenger flow level: sites with an average daily passenger flow exceeding 50,000 are marked as high-passenger-flow sites, sites with an average daily passenger flow between 10,000 and 50,000 are marked as medium-passenger-flow sites, and sites with an average daily passenger flow below 10,000 are marked as low-passenger-flow sites. The codes, topological distances, and passenger flow level information of all sites within the impact coverage area are summarized to form the set of affected sites. The set of affected sites records all sites within the impact coverage area that may be affected by the event, and each site includes its spatial relationship attribute with the event center point. The set of affected sites is sorted according to their topological distance from the event center point, from closest to furthest. Equipment failures occurring during the morning rush hour typically have a more significant impact on stations with high passenger flow, resulting in more pronounced passenger congestion.
[0057] For example, the step of forming an event impact factor based on the degree of passenger flow deviation and time decay characteristics of the affected station set includes: dividing the affected station set into a near-end station group and a far-end station group according to spatial distance; differentiating between transfer stations and non-transfer stations in the near-end station group and accumulating passenger flow deviation values to form a core deviation; applying time decay weighting to the passenger flow deviation values of the far-end station group to form an edge deviation; and forming an event impact factor based on the fusion of the core deviation and the edge deviation.
[0058] The affected station set is divided into near-end station group and far-end station group based on spatial distance. A distance threshold of 3 stations is set between each station in the affected station set and the event center point to divide the stations into two groups. Stations with a topological distance less than or equal to the threshold are assigned to the near-end station group; these stations are closer to the event location and are directly and significantly affected. Stations with a topological distance greater than the threshold are assigned to the far-end station group; these stations are farther from the event location and are affected more indirectly. When a transfer hub station experiences a failure, stations within 3 stations of the hub are assigned to the near-end station group, radiating outwards around the event center point. Stations 4 stations or more away are assigned to the far-end station group, typically located at the end of the line. The near-end station group usually includes the event center point and 2 to 3 adjacent stations, ranging from 3 to 7 stations. The far-end station group includes the remaining stations in the affected station set. The number of stations and cumulative passenger flow were counted separately for the near-end station group and the far-end station group. The near-end station group usually has fewer stations than the far-end station group, but the impact of a single station is higher.
[0059] The core deviation is formed by differentially accumulating passenger flow deviation values between transfer stations and non-transfer stations in the near-end station group. Stations with transfer functions are identified within the near-end station group. Transfer stations are nodes where multiple lines intersect, and their passenger flow deviation has an amplifying effect on the network. The passenger flow deviation value of transfer stations in the near-end station group is amplified by multiplying it by a transfer coefficient. Since transfer stations connect multiple lines, their passenger flow deviation will be transmitted to other lines along the transfer corridor; therefore, the weight is amplified to 1.5 times that of non-transfer stations. The transfer coefficient is determined based on the number of lines intersecting at the station, calculated using the formula: β = 1 + 0.5 × (n - 1), where n is the number of lines intersecting at the station, and β is the transfer coefficient. According to this formula, the coefficient is 1.5 for two-line transfer stations, 2.0 for three-line transfer stations, and 2.5 for four-line transfer stations. The passenger flow deviation value for non-transfer stations in the near-end station group remains unchanged, i.e., the coefficient is 1.0, and the impact of non-transfer stations is limited to a single line. The core deviation value (D_core) is obtained by summing the differentiated passenger flow deviation values of all stations in the near-end station group. For example, if 3000 people are stranded at a transfer station, it is counted as 4500 people (multiplied by 1.5), and 2000 people are stranded at an adjacent ordinary station, it is still counted as 2000 people. After summing, the core deviation value is 6500 people, reflecting the amplification effect of transfer stations. When the event center is located at a two-line transfer station, the passenger flow deviation contribution of adjacent three-line transfer stations may account for more than 40% of the core deviation value. The core deviation value reflects the comprehensive intensity of the impact of the event on the stations in the near-end station group. The weighted contribution of transfer stations is higher than that of non-transfer stations. When a failure occurs at a large transfer hub, the value of the core deviation value is usually significantly higher than that of a failure at an ordinary station.
[0060] The passenger flow deviation values of the remote station group are weighted by time decay to form the edge deviation. The topological distance between each station in the remote station group and the event center point is calculated. The farther the station is, the more delayed and weaker the influence transmission. A distance decay coefficient and a time decay coefficient are applied to the passenger flow deviation values of each station in the remote station group. Both decay coefficients are calculated using an exponential decay function, with the formula: γ=exp(-k·d)×exp(-λ·t), where d is the topological distance between the station and the event center point (in terms of the number of stations), k is the distance decay constant (0.2), indicating that the influence intensity decreases by approximately 18% for each additional station, t is the time elapsed after the event (in hours), and λ is the time decay constant (0.15). The passenger flow deviation values of all stations in the remote station group are multiplied by their corresponding decay coefficients and then summed to obtain the edge deviation D_edge. For example, a station 5 stations away from the incident center and 1 hour after the incident actually had 1,000 more passengers, but after the combined effects of distance and time, only about 450 were counted, demonstrating the diminishing impact of remote stations. In the early stages of an incident, remote stations have not yet perceived the impact, and the marginal deviation is close to zero. As stranded passengers move to the nearest stations, the marginal deviation peaks 20 to 30 minutes after the incident. After the incident ends and passengers gradually disperse, the marginal deviation decays to normal levels within 1 hour. The marginal deviation reflects the degree to which remote station groups are indirectly affected by the incident, and its value is usually smaller than the core deviation.
[0061] The event impact factor is formed by the fusion of core deviation and edge deviation. The core and edge deviations are weighted and summed, with the core deviation weight set to 0.7 and the edge deviation weight set to 0.3. The higher weight of the core deviation is because near-end stations directly bear the impact of the event and experience severe passenger backlog, while the lower weight of the edge deviation is because the impact of far-end stations is delayed and can be mitigated through scheduling and traffic diversion. First, the base impact is calculated using the formula: I_base = α × [(0.7 × D_core + 0.3 × D_edge) / N] × T / T_ref, where D_core is the core deviation (in person-times), D_edge is the edge deviation (in person-times), N is the total number of stations in the affected station set, T is the event duration (in hours), T_ref is the reference duration (set to 1 hour), and α is the event type correction coefficient. The correction factor for equipment failure events is 1.0, for large passenger flow control events it is 0.8, and for external event events it is 0.6. These correction factors reflect the differences in the degree of passenger flow disruption caused by different types of events. An event impact factor I is obtained by interval mapping of the basic impact quantity. The mapping formula is: I = 10 × I_base / I_max, where I_max is the maximum historical basic impact quantity. The event impact factor is normalized to a standard range of 0 to 10 through interval mapping; a larger value indicates a more significant overall impact of the event on the network's passenger flow. The event impact factor of signal failures at transfer hubs during the morning rush hour is typically 7 to 8, while the event impact factor of gate failures at terminal stations during off-peak hours is typically only 2 to 3. When a signal failure occurs at a transfer hub during the morning rush hour, the passenger flow deviation at the transfer station in the core deviation is amplified by the coefficient. In addition, the passenger flow base is large during the morning rush hour, and the passenger flow continues to accumulate during the duration of the failure, so the operation scheduling needs to activate the emergency diversion plan. On the other hand, when a short-term gate failure occurs at a small station at the end of the off-peak hours, the core deviation and the marginal deviation are both small, and normal operation can be restored with only routine handling.
[0062] Step S140: Based on the weather-passenger flow knowledge graph, perform similarity retrieval on the time-series coupled dataset to generate a set of similar cases. Filter the set of similar cases by event influence factors to generate a set of filtered cases. Perform time distance calculation and weekday attribute matching on the set of filtered cases to generate periodic similarity. Perform morphological contour matching on the set of filtered cases based on passenger flow fluctuation characteristics to generate waveform fit.
[0063] In some embodiments, the step of generating a similar case set by performing similarity retrieval on the time-series coupled dataset based on the weather-passenger flow knowledge graph includes: pre-segmenting the time-series coupled dataset by weekday type according to the weather-passenger flow knowledge graph to form a domain-specific case pool; extracting query feature vectors from the domain-specific case pool to form a retrieval benchmark; performing distance calculation based on the retrieval benchmark to form a similarity list; and performing threshold filtering on the similarity list to form a similar case set.
[0064] Based on the weather-passenger flow knowledge graph, the time-series coupled dataset is pre-domained according to weekday type to form a domain case pool. Weather node types corresponding to the weather conditions of the day are extracted from the weather-passenger flow knowledge graph to determine the weather category of the day. According to the weekday attribute of the day, the time-series coupled dataset is divided into two basic domains: a weekday set containing historical data from Monday to Friday, and a weekend subset containing historical data from Saturday and Sunday. There are fundamental differences in passenger flow patterns between weekdays and weekends. Weekdays exhibit a clear morning and evening double-peak commuting characteristic, while weekends show a single-peak leisure travel characteristic in the afternoon. Mixing the two types of data for retrieval would lead to mutual interference of features. Within the basic domains, a secondary partitioning is performed based on the weather node types in the weather-passenger flow knowledge graph, assigning historical dates of different weather types such as sunny, cloudy, and rainy days to their corresponding subdomains. The data subset formed after the time-series coupled dataset is partitioned by both weekday type and weather type is defined as the domain case pool. The domain-specific case pool organizes historical data in the time-series coupled dataset according to periodic and weather attributes, narrowing the search scope and avoiding mismatches such as using weekend data to predict weekdays or using sunny data to predict rainy days. When estimating passenger flow on rainy weekdays, the domain-specific case pool only includes historical data for rainy weekdays, excluding weekend and sunny data. If the database were searched directly from the full dataset without domain-specific analysis, it might incorrectly match high-passenger-flow cases from sunny weekends, leading to a significant overestimation of passenger flow on rainy weekdays.
[0065] The retrieval benchmark is formed by extracting query feature vectors from the domain-specific case pool. Based on the day's weather conditions and passenger flow characteristics, a query vector for retrieval is constructed. Attribute values of the day's weather nodes, including meteorological indicators such as temperature range, rainfall level, and humidity range, are extracted from the weather-passenger flow association information corresponding to the domain-specific case pool. The day's meteorological indicators are combined with the associated passenger flow characteristic indicators in the domain-specific case pool to form a multi-dimensional query feature vector. The retrieval benchmark is a vectorized representation of the day's weather and passenger flow characteristics, essentially creating a "feature profile" for that day. Subsequent retrieval will search the domain-specific case pool for the historical date most similar to this profile. The vector dimensions of the retrieval benchmark include weather type encoding, normalized temperature value, rainfall level, and tidal amplitude level. The retrieval benchmark vector is normalized to unify the value range of each dimension to between 0 and 1, eliminating the influence of dimensional differences. The retrieval benchmark will serve as a reference standard for similarity calculation; each historical date in the domain-specific case pool will be measured against this retrieval benchmark in terms of distance. When estimating passenger flow during the humid and hot weather of the plum rain season, the humidity and rainfall dimensions in the retrieval benchmark are set to high values. During the retrieval, historical cases with dry and sunny weather in the domain case pool will be automatically excluded, and priority will be given to matching dates that are also in high humidity and rainy weather, because hot and humid weather will inhibit passengers' willingness to travel, and the passenger flow pattern is significantly different from that of sunny weather.
[0066] A similarity list is generated by performing distance calculations based on the retrieval benchmark. Weather and passenger flow characteristics for each historical date in the sub-domain case pool are also constructed as feature vectors, forming a set of candidate vectors to be matched. The Euclidean distance between the retrieval benchmark vector and each candidate vector in the sub-domain case pool is calculated. Euclidean distance measures the straight-line distance between two vectors in multidimensional space; the smaller the distance, the more similar the two dates are in terms of weather and passenger flow characteristics. The Euclidean distance is converted into a similarity score using the formula: S = 1 / (1+d), where d is the Euclidean distance between the retrieval benchmark and the candidate vectors, and S is the similarity score, ranging from 0 to 1, with smaller distances resulting in higher scores. After completing the distance calculations for all historical dates in the sub-domain case pool, they are sorted from highest to lowest similarity score to form a similarity list. This similarity list essentially ranks the historical dates in the sub-domain case pool according to their "similarity to today," with the highest-ranked dates being the closest to the current date in terms of weather conditions and passenger flow characteristics, and thus having the highest reference value. When searching for information on cold waves during winter, the top-ranked dates in the similarity list are mostly dates that have historically experienced similar cold waves and temperature drops. While dates with normal temperatures are also in the domain-specific case pool, they rank lower in the similarity list due to significant differences in temperature, and are therefore given lower priority for filtering.
[0067] A threshold-based filtering process is applied to the similarity list to form a set of similar cases. A similarity threshold of 0.6 is set to filter the similarity list, retaining only historical dates with similarity scores exceeding the threshold. Historical dates with similarity scores below the threshold are removed, as these dates have significantly different characteristics from the current day, and even if they are in the same sub-domain case pool, their customer flow patterns still have limited matching degree with the current day. If the number of dates exceeding the threshold in the similarity list is too small (less than 10), the threshold is appropriately lowered to 0.5 to broaden the candidate range and ensure sufficient historical cases for subsequent analysis. If the number of dates exceeding the threshold is too large (more than 50), the threshold is appropriately increased to 0.7 to select the most similar cases, avoiding the introduction of too many marginally similar cases that dilute prediction accuracy. The historical dates retained after threshold filtering in the similarity list are summarized to form a set of similar cases. The set of similar cases is a subset of historical cases in the similarity list that are closest to the characteristics of the current day, and its number is usually controlled between 15 and 30, ensuring both the diversity of reference cases and eliminating interference from insufficient similarity. When estimating passenger flow on extreme weather days such as typhoons, since similar weather conditions are rare in history, the number of cases in the similar case set may only be 5 to 8. In this case, it is necessary to appropriately relax the threshold or include adjacent weather types in the search scope to obtain sufficient reference cases.
[0068] A filtered case set is generated by filtering similar case sets using event impact factors. The event impact factor values for each historical date in the similar case set are read. The event impact factor reflects whether the date was significantly affected by special events such as equipment failure, large-scale passenger flow control, or surrounding activities. Historical dates in the similar case set whose event impact factor I exceeds a threshold of 5 are marked. Exceeding this threshold indicates that the passenger flow on that date was significantly disturbed, and its passenger flow data deviated from normal patterns. Abnormal dates with event impact factors exceeding the threshold are removed from the similar case set, retaining historical dates with relatively normal passenger flow patterns. If a historical date with a history of large-scale delays due to signal failure is retained in the similar case set, the passenger flow data on that date will exhibit abnormal peak shifts and backlog characteristics. Using it as a reference for normal dates would seriously mislead prediction results; therefore, it needs to be identified and removed using event impact factors. If there are special events on that day, such as large-scale events at nearby stadiums, then historical dates with similar event impact factor levels should be prioritized in the similar case set to match the passenger flow patterns under the influence of the event. The result of filtering similar cases using event impact factors is defined as the selected case set. The selected case set removes historical dates affected by abnormal events, retaining cases with more stable passenger flow reference value, typically containing 10 to 20 candidate dates. When estimating passenger flow during the National Day holiday, the selected case set retains historical cases with normal passenger flow during the same National Day holiday, while removing National Day dates with abnormal passenger flow due to equipment failure or unforeseen accidents, ensuring that the reference cases reflect the true passenger flow patterns during the holiday.
[0069] The selected case set is analyzed using time distance calculation and weekday attribute matching to generate periodic similarity. The calendar distance between each historical date in the selected case set and the current date is calculated; the calendar distance is the number of days between two dates. Time proximity is calculated using an exponential decay function, P_time=exp(-0.02×Δd), where Δd is the calendar distance. The closer the date, the higher the score, as recent passenger flow patterns better reflect current travel habits and network operation status. The weekday attribute is extracted from each historical date in the selected case set, including seven categories from Monday to Sunday. Significant differences exist in passenger flow across different weeks. Mondays often see a large number of passengers returning to their work cities with luggage, resulting in higher passenger flow intensity than midweek. On Fridays, some passengers leave the city earlier, and passenger flow concentrates at railway hubs during the evening rush hour. Historical dates in the selected case set with the same weekday attribute as the current date are assigned a weekday matching score P_week: 1.0 for a perfect match, 0.7 for weekdays with different weekdays, and 0.3 for no match between weekdays and weekends. The formula for calculating periodic similarity is S_period = 0.4 × P_time + 0.6 × P_week, with a value ranging from 0 to 1. A higher periodic similarity indicates a closer alignment between the historical date's periodic position and the current date. When estimating passenger flow on Fridays, periodic similarity assigns higher weight to historical Friday cases in the selected case set because Friday evening rush hour features a unique overlap of homecoming and outbound travel. Using Wednesday cases to predict Friday's evening rush hour would underestimate passenger flow pressure.
[0070] In some embodiments, the step of performing morphological contour matching on the selected case set based on the passenger flow fluctuation characteristics to generate waveform fit includes: extracting curves from the selected case set to form a case waveform sequence; aligning the case waveform sequence with the passenger flow fluctuation characteristics using peak anchor points to form an alignment offset; performing peak-period priority distance measurement on the alignment offset to generate a morphological difference value; and performing differential weighting on the morphological difference value according to the passenger flow fluctuation characteristics to generate waveform fit.
[0071] Curve extraction was performed on the selected case set to form a case waveform sequence. Passenger volume values for each time period of the day were extracted from the historical data of each date in the selected case set, resulting in 96 time periods with a 15-minute granularity. The passenger volume values for each of the 96 time periods from each historical date in the selected case set were arranged chronologically to form a one-dimensional time-series vector. The time-series vectors for each historical date in the selected case set were normalized, scaling the passenger volume values to the range of 0 to 1. Normalization eliminated the influence of differences in total passenger flow across different dates, allowing subsequent comparisons to focus on the temporal distribution pattern of passenger flow rather than absolute quantity. The normalized time-series vectors for all historical dates in the selected case set were summarized to form a case waveform sequence set. The case waveform sequence describes the temporal distribution pattern of passenger flow for each historical date in the selected case set. Each waveform reflects the tidal change profile of passenger flow on that date, including peak occurrence time, peak duration, and peak-to-valley difference, among other morphological characteristics. In weekday case waveform sequences, the waveform curves usually show a clear bimodal shape, with the morning peak occurring around 8 a.m. and the evening peak occurring around 6 p.m. In contrast, weekend case waveform sequences show a relatively flat single peak or no obvious peak.
[0072] The alignment offset is generated by aligning the case waveform sequence with the peak anchor points of passenger flow fluctuation characteristics. The peak times of the morning and evening peaks are extracted from the passenger flow fluctuation characteristics of the day and used as the alignment anchor points. The peak times of the morning and evening peaks are identified from the waveform curves of historical dates in the case waveform sequence to locate the anchor points for each case. Peak times may vary between dates; for example, the morning peak may occur 15 to 20 minutes earlier during the school opening season, while the evening peak may end half an hour later during large exhibitions. If this time offset is not corrected, it can lead to false morphological differences when comparing waveforms. The time deviation between the morning peak anchor points of historical dates in the case waveform sequence and the morning peak anchor points in the passenger flow fluctuation characteristics of the day is calculated. The deviation value can be positive or negative; a positive value indicates that the peak time of a historical case was later than the current day, and a negative value indicates that the peak time of a historical case was earlier than the current day. Similarly, the time deviation between the evening peak anchor points of historical dates in the case waveform sequence and the evening peak anchor points in the passenger flow fluctuation characteristics of the day is calculated to obtain the evening peak offset. The morning and evening peak offsets for each historical date in the case waveform sequence are recorded to form alignment offsets. Alignment offsets reflect the differences in peak-hour positioning between historical cases and the current day's passenger flow fluctuations in the case waveform sequence. By using alignment offsets, the case waveform sequence can be shifted along the time axis, aligning the peak times of each case with the current day before morphological comparison. On the first workday after students' winter break, due to the overlap of commuter and student passenger flow, the morning peak time is earlier than usual. The alignment offsets will show that the morning peak time for most ordinary weekday cases in the case waveform sequence is later than the current day, requiring the waveforms of these cases to be shifted forward before morphological comparison.
[0073] A peak-period priority distance metric is used to measure the alignment offset to generate a morphological difference value. Based on the magnitude of the alignment offset, the waveforms of each historical date in the case waveform sequence are shifted along the time axis to align the peak anchor point with the passenger flow fluctuation characteristics of the current day. The shifted case waveform sequence is then placed in the same time reference system as the current day. The shifted case waveform sequence is compared with the passenger volume time-series curve of the current day, and the difference in passenger volume for each time period is calculated. A higher weight is assigned to the difference during peak periods: 1.5 for peak periods, 1.0 for off-peak periods, and 0.5 for low-peak periods. This differentiated weighting is because the accuracy of passenger flow estimation during peak periods is more critical for train scheduling and passenger organization decisions; estimation errors during peak periods may lead to excessively high train occupancy rates or wasted capacity. The morphological difference value is obtained by summing the squared weighted differences between each historical date and the current day in the case waveform sequence. The morphological difference value reflects the overall deviation in shape between the historical cases and the current day's passenger flow curve in the case waveform sequence; the smaller the value, the more similar the waveform shape. After alignment offset correction, the morphological difference value eliminates the differences caused solely by the timing of peak hours, focusing instead on the similarity of the waveform profile itself. In a historical case, although the morning peak was 15 minutes later than the current day's peak, after anchor point alignment and translation, the peak height, rise slope, duration, and other morphological characteristics of the two are highly consistent. In this case, the morphological difference value will be very small, indicating that its waveform profile is actually very similar to that of the current day, making it a high-quality reference case.
[0074] Waveform fit is generated by differentially weighting morphological differences based on passenger flow fluctuation characteristics. Indicators such as tidal amplitude and tidal intensity ratio are extracted from the daily passenger flow fluctuation characteristics to guide the differential weighting strategy. Historical cases with smaller morphological differences are assigned higher base scores; the base score is inversely proportional to the morphological difference value, with smaller morphological differences resulting in higher base scores. Based on the tidal amplitude level in the passenger flow fluctuation characteristics, historical cases with similar amplitudes in the waveform sequence are assigned amplitude matching bonuses. Tidal amplitude reflects the degree of difference between passenger flow peaks and troughs within a day; larger amplitudes indicate higher passenger flow concentration during peak periods, while smaller amplitudes indicate a more even distribution of passenger flow throughout the day. An amplitude difference within 10% is awarded 0.1 points, and an amplitude difference within 20% is awarded 0.05 points. Based on the tidal intensity ratio level in the passenger flow fluctuation characteristics, historical cases with similar intensity ratios in the waveform sequence are assigned intensity matching bonuses; an intensity ratio difference within 0.5 is awarded 0.1 points. The waveform fit score is calculated by combining the base score obtained from the morphological difference value conversion with the amplitude matching score and the intensity matching score. The waveform fit score ranges from 0 to 1 and comprehensively measures the degree of similarity between each historical case in the waveform sequence and the current day in terms of waveform profile, amplitude intensity, peak and trough characteristics. If the current day is a commuter-dominated weekday with large amplitude, using a leisure-dominated weekend case with smaller amplitude for prediction will result in a significant underestimation of peak passenger flow due to amplitude mismatch, even if the morphological difference values are not large after alignment. Therefore, the waveform fit score will be lowered due to amplitude mismatch, preventing it from being selected as a primary reference case.
[0075] Step S150: Based on the period similarity and waveform fit, confidence modulation is applied to the selected case set to generate passenger volume estimates. Error analysis is performed between the passenger volume estimates and the actual passenger volume to generate optimized weight parameters. Based on the optimized weight parameters, the estimated passenger flow distribution for the whole day is output to the operation management terminal.
[0076] Specifically, passenger volume estimates are generated by confidence modulation of the selected case set based on periodic similarity and waveform fit. The periodic similarity and waveform fit of each historical date in the selected case set are combined to calculate the overall confidence score for each case. The overall confidence score is calculated using a weighted summation method, with the formula: C = W1 × S_period + W2 × S_wave, where S_period is the periodic similarity; S_wave is the waveform fit; W1 and W2 are weighting coefficients, initially set to 0.4 and 0.6 respectively, with W2 being higher because waveform shape has a more direct impact on passenger flow estimation; C is the overall confidence score, ranging from 0 to 1. The historical dates in the selected case set are sorted according to the overall confidence score, and the top 5 to 10 cases with the highest confidence are selected as reference cases. The number of reference cases is dynamically adjusted according to the size of the selected case set. The passenger volume time-series curves of the reference cases are weighted and averaged according to their confidence scores, with higher confidence scores carrying greater weight. The weighting coefficient is the sum of the confidence scores of each case and the sum of the confidence scores of all reference cases. The weighted average passenger volume time-series curves are then smoothed to eliminate local fluctuations caused by differences between cases, resulting in smoothed passenger volume estimates. These passenger volume estimates are based on the comprehensive prediction results of high-confidence cases selected from the case set, encompassing passenger volume estimates for 96 time periods throughout the day. Each time period's estimate includes three indicators: inbound passenger volume, outbound passenger volume, and total passenger volume. On dates where weather conditions and weekday attributes closely match historical cases, the accuracy of the passenger volume estimates is typically high, with the average relative error controlled within 8%.
[0077] Error analysis is performed between the estimated and actual passenger volume to generate optimized weighting parameters. After the end of the day's operations, actual passenger volume statistics are retrieved from the AFC database, and the actual passenger volume for each time period is compared with the estimated passenger volume. The absolute and relative errors between the estimated and actual passenger volume are calculated. The absolute error is the absolute value of the difference between the two, and the relative error is the ratio of the absolute error to the actual passenger volume; the relative error better reflects the accuracy of the estimate. The mean and standard deviation of the relative errors for each time period throughout the day are calculated. The mean reflects the systematic bias of the estimate, and the standard deviation reflects the stability of the estimate. The error distribution characteristics of the passenger volume estimate in different time periods are analyzed to identify time intervals with larger errors, with particular attention paid to the error levels during the morning peak (7:00-9:00) and evening peak (17:00-19:00), as the estimation accuracy during peak hours has a greater impact on train scheduling and passenger flow management decisions. Based on the error analysis results, the weighting coefficients in the comprehensive confidence score calculation are adjusted. If cases with high period similarity show better prediction results, the value of W1 is increased; if cases with high waveform fit show better prediction results, the value of W2 is increased. The weighting adjustment step size is set to 0.05. Simultaneously, based on the error differences between peak and off-peak periods, a time-period correction factor is generated, applying corrections to periods with larger errors in the next estimation. The adjusted weighting coefficients and time-period correction factors are summarized to form the optimized weighting parameters. These optimized weighting parameters are derived from feedback comparing passenger volume estimates with actual data, and are used to improve the accuracy of the next passenger flow estimation.
[0078] Based on optimized weight parameters, the estimated passenger flow distribution for the entire day is output to the operations management terminal. The optimized weight parameters are applied to the passenger flow estimation model, updating the weight configuration and time-period correction factor for the comprehensive confidence score calculation. The confidence score of each case in the selected case set is recalculated based on the updated weight configuration, generating optimized passenger volume estimates. These optimized passenger volume estimates are then organized into a passenger flow distribution table by time period. The table includes estimated inbound passenger volume, estimated outbound passenger volume, estimated total passenger volume, and estimated confidence intervals for each time period. The confidence intervals are calculated based on historical error distribution. The passenger flow distribution table is converted into a chart format, generating a time-series passenger flow curve. The curve uses time as the horizontal axis and passenger volume as the vertical axis, while also marking the peak positions and expected peak passenger volumes for morning and evening peak hours. Passenger flow heatmaps are generated for each station. The heatmaps are arranged with stations as rows and time periods as columns, using color intensity to represent passenger flow intensity, facilitating the identification of high-passenger-flow stations and high-passenger-flow periods by operations management personnel. Passenger flow distribution tables and visualization charts are encapsulated into standard data messages and pushed to operation and management terminals via data interfaces. Terminal types include dispatch center screens, station staff handheld devices, and management workstations. The dispatch center screen displays passenger flow distribution estimation results from a network-wide perspective, supporting operations dispatchers in adjusting train schedules and making decisions regarding the deployment of standby trains. Station staff handheld devices push local passenger flow estimation curves and peak-hour warnings, facilitating advance preparation for passenger flow management.
[0079] To implement the above-described method embodiment, a method for dynamically estimating the daily passenger flow distribution of an urban rail transit network is proposed to achieve the corresponding functionalities and technical effects. See also... Figure 2 , Figure 2 This diagram illustrates a structural block diagram of a dynamic estimation system 200 for all-day passenger flow distribution in an urban rail transit network, as provided in an embodiment of this application. For ease of explanation, only the parts relevant to this embodiment are shown. The dynamic estimation system 200 for all-day passenger flow distribution in an urban rail transit network, as provided in this embodiment of the application, includes:
[0080] Data acquisition module 201 is used to acquire passenger card swiping data, historical weather data and special event record data of AFC system, and perform time-series alignment on the passenger card swiping data and the historical weather data to generate a time-series coupled dataset;
[0081] Feature extraction module 202 is used to perform time-period passenger volume statistics on the passenger card swiping data to generate a passenger volume distribution sequence, extract passenger flow fluctuation features from the passenger volume distribution sequence, and map the passenger flow fluctuation features to the historical weather data to form a weather-passenger flow knowledge graph.
[0082] Event parsing module 203 is used to perform event type parsing on the special event record data to generate an event coding sequence, and to perform impact range assessment on the event coding sequence to generate an event impact factor;
[0083] The case retrieval module 204 is used to perform similarity retrieval on the time-series coupled dataset based on the weather-passenger flow knowledge graph to generate a set of similar cases, filter the set of similar cases through the event influence factor to generate a set of filtered cases, perform time distance calculation and weekday attribute matching on the set of filtered cases to generate periodic similarity, and perform morphological contour matching on the set of filtered cases based on the passenger flow fluctuation characteristics to generate waveform fit.
[0084] The estimation output module 205 is used to generate a passenger volume estimate by performing confidence modulation on the selected case set based on the period similarity and the waveform fit, perform error analysis on the passenger volume estimate and the actual passenger volume to generate optimized weight parameters, and output the estimated passenger flow distribution for the whole day to the operation management terminal based on the optimized weight parameters.
[0085] The aforementioned urban rail transit network all-day passenger flow dynamic estimation system 200 can implement the urban rail transit network all-day passenger flow dynamic estimation method of the above method embodiment. The options in the above method embodiment are also applicable to this embodiment, and will not be detailed here. The remaining contents of this application embodiment can refer to the contents of the above method embodiment, and will not be repeated in this embodiment.
[0086] The purpose of the above embodiments is to reproduce and derive the technical solution of the present invention by way of example, and to fully describe the technical solution, purpose and effect of the present invention. The purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosure of the present invention, and not to limit the scope of protection of the present invention.
[0087] The above embodiments are not an exhaustive list based on the present invention, and there may be many other embodiments not listed. Any substitutions and improvements made without departing from the concept of the present invention are within the protection scope of the present invention.
Claims
1. A method for dynamically estimating the daily passenger flow distribution of an urban rail transit network, characterized in that, include: Acquire passenger card swipe data, historical weather data, and special event record data from the AFC system, and perform time-series alignment on the passenger card swipe data and the historical weather data to generate a time-series coupled dataset; The passenger card swipe data is used to perform time-period passenger volume statistics to generate a passenger volume distribution sequence. Passenger flow fluctuation characteristics are extracted from the passenger volume distribution sequence, and the passenger flow fluctuation characteristics are mapped to the historical weather data to form a weather-passenger flow knowledge graph. The process involves parsing the special event record data to generate an event coding sequence, and then assessing the impact range of the event coding sequence to generate an event impact factor. This includes: identifying the event occurrence sites based on the event coding sequence to form an event center point; performing spatial radiation analysis on the event center point to form an impact coverage area; statistically analyzing the affected sites within the impact coverage area to form an affected site set; and generating an event impact factor based on the passenger flow deviation degree and time decay characteristics of the affected site set. Specifically, generating the event impact factor based on the passenger flow deviation degree and time decay characteristics of the affected site set includes: dividing the affected site set into a near-end site group and a far-end site group based on spatial distance; differentially accumulating passenger flow deviation values between transfer stations and non-transfer stations in the near-end site group to form a core deviation; applying time decay weighting to the passenger flow deviation values of the far-end site group to form a marginal deviation; and merging the core deviation and the marginal deviation to form the event impact factor. The process of generating a similar case set by performing similarity retrieval on the time-series coupled dataset based on the weather-passenger flow knowledge graph includes: pre-segmenting the time-series coupled dataset by weekday type according to the weather-passenger flow knowledge graph to form a domain-specific case pool; extracting query feature vectors from the domain-specific case pool to form a retrieval benchmark; performing distance calculation based on the retrieval benchmark to form a similarity list; performing threshold filtering on the similarity list to form a similar case set; filtering the similar case set using the event influence factor to generate a filtered case set; performing time distance calculation and weekday attribute matching on the filtered case set to generate periodic similarity; and performing morphological contour matching on the filtered case set based on the passenger flow fluctuation characteristics to generate waveform fit. Based on the period similarity and the waveform fit, the selected case set is subjected to confidence modulation to generate a passenger volume estimate. The passenger volume estimate is compared with the actual passenger volume to generate an optimized weight parameter. Based on the optimized weight parameter, the estimated passenger flow distribution for the whole day is output to the operation management terminal.
2. The method according to claim 1, characterized in that, The step of performing time-series alignment between the passenger card swipe data and the historical weather data to generate a time-series coupled dataset includes: Based on the passenger card swipe data, the entry and exit timestamps are extracted to form a time tag sequence; The time-stamp sequence and the historical weather data are correlated by time period to form a time period-weather correspondence table; A set of abrupt weather changes is formed by identifying weather change points from the time period-weather correspondence table. Based on the mutation marker set, segmentation and alignment are performed to form a temporally coupled dataset.
3. The method according to claim 1, characterized in that, Extracting passenger flow fluctuation features from the passenger volume distribution sequence includes: Tidal gradient analysis is performed on the passenger volume distribution sequence to generate a passenger flow change rate curve; The commuter tidal critical points are identified from the passenger flow change rate curve to form a peak-valley marker sequence; The peak and valley marker sequence is used to locate peak periods and valley periods; Passenger flow fluctuation characteristics are formed based on the tidal amplitude intensity during the peak and trough periods.
4. The method according to claim 1, characterized in that, The step of mapping the passenger flow fluctuation characteristics to the historical weather data to form a weather-passenger flow knowledge graph includes: The passenger flow fluctuation characteristics are materialized into nodes to generate a passenger flow characteristic node set; The weather-passenger flow association edge is formed by mapping the passenger flow feature node set with the historical weather data. The edge weight set is generated by performing association confidence quantification on the weather-passenger flow association edge; A weather-passenger flow knowledge graph is constructed based on the set of passenger flow feature nodes, the weather-passenger flow association edges, and the set of edge weights.
5. The method according to claim 1, characterized in that, The step of performing morphological contour matching on the selected case set based on the passenger flow fluctuation characteristics to generate waveform fit includes: Curve extraction is performed on the selected case set to form a case waveform sequence; The waveform sequence of the case is aligned with the passenger flow fluctuation characteristics to form an alignment offset by aligning the peak anchor point; The alignment offset is used to generate morphological difference values by prioritizing distance measurement during peak periods. Based on the passenger flow fluctuation characteristics, the morphological difference values are weighted differently to generate waveform matching degree.
6. The method according to claim 3, characterized in that, The step of identifying commuter tidal critical points from the passenger flow change rate curve to form a peak-valley marker sequence includes: Based on the passenger flow change rate curve, a candidate extreme value point set is formed by searching the operating period window. The candidate extreme point set is classified and aggregated into extreme point clusters by peak-to-peak classification; Representative points are selected from the cluster of extreme points to form the commuting tidal critical points; Peak-valley marker sequences are generated based on the time period distribution of the commuting tidal critical points.
7. A dynamic estimation system for the all-day passenger flow distribution of an urban rail transit network, characterized in that, include: The data acquisition module is used to acquire passenger card swipe data, historical weather data, and special event record data from the AFC system, and to perform time-series alignment on the passenger card swipe data and the historical weather data to generate a time-series coupled dataset. The feature extraction module is used to perform time-period passenger volume statistics on the passenger card swiping data to generate a passenger volume distribution sequence, extract passenger flow fluctuation features from the passenger volume distribution sequence, and map the passenger flow fluctuation features to the historical weather data to form a weather-passenger flow knowledge graph. The event analysis module is used to perform event type analysis on the special event record data to generate an event coding sequence, and to evaluate the impact range of the event coding sequence to generate an event impact factor. This includes: identifying the event occurrence sites based on the event coding sequence to form an event center point; performing spatial radiation analysis on the event center point to form an impact coverage area; statistically analyzing the affected sites within the impact coverage area to form an affected site set; and forming an event impact factor based on the passenger flow deviation degree and time decay characteristics of the affected site set. Specifically, forming the event impact factor based on the passenger flow deviation degree and time decay characteristics of the affected site set includes: dividing the affected site set into a near-end site group and a far-end site group according to spatial distance; differentiating and accumulating passenger flow deviation values between transfer stations and non-transfer stations in the near-end site group to form a core deviation; applying time decay weighting to the passenger flow deviation values of the far-end site group to form a marginal deviation; and forming an event impact factor by fusing the core deviation and the marginal deviation. The case retrieval module is used to perform similarity retrieval on the time-series coupled dataset based on the weather-passenger flow knowledge graph to generate a set of similar cases. This includes: pre-segmenting the time-series coupled dataset by weekday type according to the weather-passenger flow knowledge graph to form a domain-specific case pool; extracting query feature vectors from the domain-specific case pool to form a retrieval benchmark; performing distance calculation based on the retrieval benchmark to form a similarity list; performing threshold filtering on the similarity list to form a set of similar cases; filtering the set of similar cases using the event impact factor to generate a filtered case set; performing time distance calculation and weekday attribute matching on the filtered case set to generate periodic similarity; and performing morphological contour matching on the filtered case set based on the passenger flow fluctuation characteristics to generate waveform fit. The estimation output module is used to generate a passenger volume estimate by performing confidence modulation on the selected case set based on the period similarity and the waveform fit, perform error analysis on the passenger volume estimate and the actual passenger volume to generate optimized weight parameters, and output the estimated passenger flow distribution for the whole day to the operation management terminal based on the optimized weight parameters.
Citation Information
Patent Citations
DAS driving strategy optimization method driven by subway line real-time data
CN120363970A
Integrated dispatch decision-making system and method for hub supporting multi-modal transportation coordination, and device
WO2025232023A1