Ontology-based station-city collaborative data integration and planning prediction method

By constructing an urban rail transit station-city collaborative ontology and utilizing a multi-task spatiotemporal graph convolutional neural network, the problem of inconsistency among multi-source data was solved, the interpretability and executability of urban rail transit planning were realized, and the cost of strategy iteration was reduced.

CN121581579APending Publication Date: 2026-02-27BEIJING JIAOTONG UNIV

Patent Information

Application Number
CN202511919829.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

In existing urban rail transit planning, inconsistent data sources and spatial units make it difficult to reproduce planning assessments, difficult to explain the basis for intervention, and high costs for strategy iteration. Furthermore, the prediction and planning processes are disconnected.

Method used

We construct an urban rail transit station-city collaborative ontology, perform semantic annotation and semantic fusion, use a multi-task spatiotemporal graph convolutional neural network to jointly predict passenger flow and built environment indicators, and optimize feature contribution through Shapley additive interpretation value and genetic algorithm to generate planning indicator thresholds and intervention measure parameters.

Benefits of technology

It achieves unified and traceable integration of multi-source data at the scale of pedestrian service areas in stations, and simultaneously outputs interpretable prediction results of passenger flow and built environment indicators, providing actionable planning inputs and reducing the cost of strategy iteration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121581579A_ABST
    Figure CN121581579A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of urban rail transit station-city collaborative planning, in particular to an ontology-based station-city collaborative data integration and planning prediction method, which comprises the following steps of: obtaining rail transit station passenger flow data, resident travel behavior data and station periphery built environment index data; forming a space-time sample sequence according to the unified space-time granularity of the site walking service area; constructing an urban rail transit station-city cooperation ontology, and carrying out semantic annotation and semantic fusion on the space-time sample sequence to generate a feature sequence; inputting the feature sequence into a multi-task space-time diagram convolutional neural network prediction model to output a passenger flow prediction result and establish an environment index prediction result; and calculating a feature contribution degree based on a Shapley additive interpretation value, optimizing a background sample set by using a genetic algorithm to determine a key action element set, outputting a planning index threshold and an intervention measure parameter, and realizing an interpretable station-city collaborative prediction and planning decision closed loop.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of urban rail transit station-city collaborative planning, and in particular to a station-city collaborative data integration and planning prediction method based on ontology. BACKGROUND

[0002] With the increase in the density of urban rail transit networks, the space form around the station, public transport connection and the travel behavior of the crowd present strong coupling evolution characteristics. The planning department needs to simultaneously analyze the passenger flow trend and the change of the built environment index at the micro service area scale to support the collaborative decision of line organization, station update and space control. In existing practices, the multi-source data caliber is not unified, the spatial unit is inconsistent, and the prediction and planning links are disconnected, often forming a chain between "data collection-prediction analysis-planning implementation", which leads to difficulties in reproducing planning evaluation, difficulties in explaining intervention basis, and high cost of strategy iteration. SUMMARY

[0003] The present application provides a station-city collaborative data integration and planning prediction method based on ontology, which at least solves the problems of difficulty in unifying the semantics and spatio-temporal caliber of multi-source data, difficulty in jointly predicting passenger flow and spatial indicators and forming an interpretable planning decision input in the station-city collaborative scenario.

[0004] In one aspect, the present application provides a station-city collaborative data integration and planning prediction method based on ontology, comprising the following steps: Obtain rail transit station passenger flow data, resident travel behavior data and station surrounding built environment index data, and form a spatio-temporal sample sequence with station identification and time index according to the unified spatio-temporal granularity of the station walking service area; Construct a city rail transit station-city collaborative ontology, and perform semantic annotation and semantic fusion on the spatio-temporal sample sequence based on the city rail transit station-city collaborative ontology to generate a feature sequence; Input the feature sequence into a multi-task spatio-temporal graph convolutional neural network prediction model to output passenger flow prediction results and built environment index prediction results; Calculate the feature contribution degree based on the Shapley additive explanation value, optimize the selection of the background sample set used to calculate the Shapley additive explanation value based on a genetic algorithm, determine a key role element set based on the feature contribution degree, and generate planning index thresholds and intervention measure parameters based on the key role element set.

[0005] In one possible implementation, the station surrounding built environment index data includes point of interest density index, land use mixing degree index, land use intensity index, building density index, road network density index, road network accessibility index, residential population density index, employment density index, green coverage index and public transport connection index.

[0006] In a possible implementation, the forming of the space-time sample sequence with the station identifier and the time index in the station walking service area includes: performing spatial matching on the rail transit station passenger flow data, the resident travel behavior data, and the station surrounding built environment index data to associate to the station walking service area, and performing time index alignment on the spatially matched data to generate the space-time sample sequence.

[0007] In a possible implementation, the constructing of the urban rail transit station-city collaborative ontology includes: constructing a top-level ontology, a domain ontology, a task ontology, and an application ontology, defining a rail transit station concept, a station walking service area concept, a resident travel behavior concept, and a built environment index concept in the urban rail transit station-city collaborative ontology, and defining an interaction relationship constraint between the resident travel behavior concept and the built environment index concept and an association relationship constraint between the resident travel behavior concept and the rail transit station concept.

[0008] In a possible implementation, the semantic annotation of the space-time sample sequence based on the urban rail transit station-city collaborative ontology includes: establishing a corresponding rule of a data field and a concept and a data attribute in the urban rail transit station-city collaborative ontology, and converting the space-time sample sequence into an ontology instance according to the corresponding rule.

[0009] In a possible implementation, the semantic fusion of the space-time sample sequence based on the urban rail transit station-city collaborative ontology includes: performing instance merging of the same station identifier and the same time index on the ontology instance, and determining a value according to a preset conflict resolution rule for a conflicting data attribute; and generating a feature sequence according to a preset feature sorting rule of the rail transit station passenger flow data attribute, the resident travel behavior data attribute, and the station surrounding built environment index data attribute.

[0010] In a possible implementation, the multi-task space-time graph convolutional neural network prediction model constructs a graph node based on a rail transit station, and determines a graph connection relationship based on a geographical distance between rail transit stations and a station-to-station trip association relationship represented by resident travel behavior data.

[0011] In a possible implementation, the multi-task space-time graph convolutional neural network prediction model includes a spatial feature extraction network and a temporal feature extraction network, the spatial feature extraction network generates a dynamic adjacency relationship based on a self-attention mechanism and performs graph convolution calculation on a node feature corresponding to the graph connection relationship, the temporal feature extraction network performs time series modeling on the feature sequence based on a time dimension attention mechanism, and the consistency of the output of the spatial feature extraction network and the output of the temporal feature extraction network is constrained based on a space-time consistency loss function in a training process, and a multi-task output structure of the passenger flow prediction output and the built environment index prediction output is set.

[0012] In one possible implementation, calculating the feature contribution based on Shapley additive interpretation values ​​includes: calculating the contribution of each feature in the feature sequence to the prediction results for both passenger flow prediction results and built environment indicator prediction results, and determining the set of key contributing factors based on the feature contribution.

[0013] In one possible implementation, optimizing the selection of the background sample set used to calculate the Shapley additive interpretation value based on a genetic algorithm includes: performing selection, crossover, and mutation on the background sample set to determine the target background sample set; generating planning indicator thresholds and intervention parameters includes: transforming the set of key elements into planning indicators, and establishing a correspondence between elements, indicators, thresholds, and interventions based on the planning indicators to output planning indicator thresholds and intervention parameters.

[0014] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: By constructing a collaborative ontology between urban rail transit stations and cities and implementing semantic annotation and fusion of spatiotemporal sample sequences, a unified caliber and traceable integration of multi-source heterogeneous data at the scale of station pedestrian service areas were achieved. By using a multi-task spatiotemporal graph convolutional neural network to simultaneously output passenger flow prediction results and built environment indicator prediction results, joint prediction of station-city elements was realized. By calculating feature contribution using Shapley additive interpretation values ​​and optimizing the background sample set using a genetic algorithm, the stability and interpretability of key role element identification were achieved. By outputting planning indicator thresholds and intervention measure parameters through the correspondence between elements, indicators, thresholds, and intervention measures, a closed-loop transformation from prediction results to executable planning inputs was realized. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the execution flow of the method of the present invention; Figure 2 This is a schematic diagram of the AST-GCN model architecture in this invention. Detailed Implementation

[0016] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0018] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0019] Station-city collaborative data refers to a spatiotemporal data set centered around urban rail transit stations and their pedestrian service areas, capable of depicting the correlation chain of "travel behavior—passenger flow changes—spatial element evolution." This type of data typically includes information such as rail transit station passenger flow statistics, resident travel behavior records, and built environment indicators around the stations, using station identifiers and time indexes as unified anchor points to achieve joint expression across sources and scales. The value of station-city collaborative data lies not in the independent use of a single data source, but in placing behavioral and spatial changes within the same analytical framework through unified semantic definitions and spatiotemporal granularity. This allows passenger flow prediction and built environment indicator prediction to mutually constrain and interpret each other, providing traceable data support for setting planning indicator thresholds and outputting intervention parameters. Based on this, this invention proposes an ontology-based station-city collaborative data integration and planning prediction method to generate standardized feature sequences and interpretable outputs at the station pedestrian service area scale that can be used for prediction and planning decisions.

[0020] like Figure 1 As shown, an ontology-based method for station-city collaborative data integration and planning prediction includes the following steps: Acquire passenger flow data, resident travel behavior data, and built environment indicator data around the stations for rail transit stations, and form a spatiotemporal sample sequence with station identification and time index according to the unified spatiotemporal granularity of the pedestrian service area of ​​the station; In the embodiments of this specification, passenger flow data at rail transit stations is obtained by aggregating data from automatic fare collection systems, turnstile counting systems, or video passenger flow counting systems, and is recorded using station identifiers and collection times. Resident travel behavior data is obtained by parsing mobile communication signaling data, public transportation card swipe data, or travel order data, extracting behavioral elements such as travel origin, travel destination, and travel time. Built environment indicator data around stations is calculated from layers of interest, land use, road network, and green space in a geographic information system. For each rail transit station, a pedestrian service area is generated based on the shortest walking path of the road network, and resident travel behavior data and built environment indicator data around the station are spatially correlated to the pedestrian service area. Furthermore, the above three types of data are aggregated at a uniform time granularity, and missing values ​​are processed and outliers are removed, ultimately forming a spatiotemporal sample sequence with station identifiers and time indices as keys, and passenger flow characteristics, travel behavior characteristics, and built environment indicators as content.

[0021] The data on the built environment around the station include indicators such as point of interest density, land use mix, land use intensity, building density, road network density, road network accessibility, residential population density, employment density, green space coverage, and public transportation connection.

[0022] In this embodiment, the built environment index data around the station is used to characterize the spatial form and connection conditions within the pedestrian service area of ​​the rail transit station, and is correlated with rail transit station passenger flow data and resident travel behavior data at a unified spatiotemporal granularity. The basic built environment data comes from geographic information systems and urban statistical databases, including rail transit station entrances / exits, road networks, land use status, building data, points of interest (POIs) data, green space data, residential population and employment distribution data, and bus stop and route data. After performing coordinate system one and merging duplicate records on the spatial data, the boundary of the station's pedestrian service area is generated based on the shortest walking path of the road network, starting from the rail transit station entrances / exits. The station's pedestrian service area serves as the unified spatial unit for index calculation. The POI density index is the result of normalizing the number of POIs within the station's pedestrian service area according to the area of ​​the station's pedestrian service area. The land use mix index is obtained by calculating the area proportion of each land use type and characterizing the diversity of land use structure based on information entropy normalization. The land use intensity index is obtained by summarizing the total building area within the station's pedestrian service area and normalizing it according to the land area of ​​the station's pedestrian service area. The building density index is obtained by summarizing the building footprint area within the pedestrian service area of ​​the station and normalizing it according to the area of ​​the pedestrian service area. The road network density index is obtained by statistically analyzing the total length of the road centerlines within the pedestrian service area of ​​the station and normalizing it according to the area of ​​the pedestrian service area. The road network accessibility index is obtained by generating grid sampling points within the pedestrian service area of ​​the station, calculating the shortest walking time from the sampling points to the entrance / exit of the rail transit station, and using the statistical results of the shortest walking time to characterize the accessibility level. The residential population density index and the employment density index are the results of normalizing the number of residents and employees within the pedestrian service area of ​​the station according to the area of ​​the pedestrian service area, respectively. The green space coverage index is the proportion of green space area within the pedestrian service area of ​​the station. The public transportation connection index is obtained by statistically analyzing the number of bus stops, the number of bus routes covered, and the walking time from the entrance / exit of the rail transit station to the nearest bus stop within the pedestrian service area of ​​the station, and generating a connection level characterization value according to a preset synthesis rule. The built environment index data is updated according to a preset update cycle and mapped to the corresponding time index of the spatiotemporal sample sequence according to the time index alignment rule.

[0023] The process of forming a spatiotemporal sample sequence with station identifiers and time indexes according to the unified spatiotemporal granularity of the station pedestrian service area includes: performing spatial matching on the passenger flow data of rail transit stations, resident travel behavior data, and the built environment index data around the stations to associate them with the station pedestrian service area, and performing time index alignment on the spatially matched data to generate a spatiotemporal sample sequence.

[0024] In the embodiments of this specification, the passenger flow data, resident travel behavior data, and surrounding built environment indicator data of rail transit stations are derived from a multi-source heterogeneous data system, including but not limited to rail transit ticketing and card swiping data, mobile phone signaling data, point of interest data, land use status data, public transport card swiping data, and passenger flow survey data, to cover the key data attributes required for the "behavior-space" interaction at the micro-scale around the station. To achieve "forming a spatiotemporal sample sequence with station identification and time index according to the unified spatiotemporal granularity of the station's pedestrian service area," the boundary of the station's pedestrian service area is first established: starting from the station entrance / exit, the walkable range is calculated based on the road network, limited to the spatial range corresponding to a five-minute walking circle, and this range is used as the unified carrier unit for spatial matching.

[0025] During the spatial matching phase, the three types of data are associated with the pedestrian service areas of the stations respectively: For passenger flow data of rail transit stations, station identifiers are directly assigned based on the station number or the station number to which the turnstile belongs, and data from different entrances and exits of the same station are merged into a unified station identifier; For resident travel behavior data, the departure location, arrival location, and occurrence time of each record are extracted, the location is mapped to geographic coordinates, and a point-to-area inclusion judgment is performed. Records falling into a certain station's pedestrian service area are assigned the station identifier; When the same location falls into multiple station pedestrian service areas at the same time, the station with the shortest walking time in the road network is used as the belonging station to avoid duplicate counting; For the built environment indicator data around the station, spatial statistics are performed based on the boundary of the station's pedestrian service area to obtain indicator values ​​that correspond one-to-one with the station identifiers, and the indicator caliber is kept consistent throughout the entire process.

[0026] During the time index alignment phase, a unified spatiotemporal granularity is determined. After converting the time fields of each data source to a unified time zone and a unified time format, the time index is divided according to a preset time slice. For event-type data such as passenger flow and behavior data, aggregation is performed under the same station identifier according to the time index to form statistics corresponding to the time index. For relatively slow-changing data such as built environment indicators, they are mapped to the corresponding time index according to their update cycle, and the values ​​are kept continuous and consistent between adjacent time indices. When a station identifier lacks event-type data in a certain time index, it is recorded as a null value and processed by the subsequent semantic fusion phase according to the data source priority and consistency rules. After completing the above spatial matching and time index alignment, a spatiotemporal sample sequence with "station identifier - time index" as the primary key is generated, which serves as the data input basis for subsequent ontology instantiation, semantic annotation and semantic fusion, as well as for spatiotemporal graph convolutional prediction models.

[0027] Construct an urban rail transit station-city collaborative ontology, and perform semantic annotation and semantic fusion on spatiotemporal sample sequences based on the urban rail transit station-city collaborative ontology to generate feature sequences; In this embodiment, when constructing the urban rail transit station-city collaborative ontology, a conceptual hierarchy is pre-established, including rail transit stations, station pedestrian service areas, rail transit station passenger flow, resident travel behavior, station surrounding built environment indicators, and time attributes. Object relationships and data attributes are configured for each concept, forming a reusable ontology structure. Based on the ontology structure, a correspondence rule is established between data fields and concepts and data attributes. Station identifiers, time indices, and each data field from the spatiotemporal sample sequence are written into the corresponding ontology instance to complete semantic annotation. During semantic fusion, station identifiers and time indices are used as instance merging keys to perform entity alignment and attribute consistency checks on multi-source ontology instances. When conflicting values ​​exist for the same data attribute, the value is determined according to preset conflict resolution rules, and missing data attributes are filled in according to preset filling rules. Subsequently, based on preset feature sorting rules, rail transit station passenger flow features, resident travel behavior features, and station surrounding built environment indicator features are extracted from the fused ontology instance to generate a feature sequence for input to the subsequent multi-task spatiotemporal graph convolutional neural network prediction model.

[0028] The construction of the urban rail transit station-city collaborative ontology includes: constructing a top-level ontology, a domain ontology, a task ontology, and an application ontology. In the urban rail transit station-city collaborative ontology, the concepts of rail transit station, station pedestrian service area, resident travel behavior, and built environment indicators are defined, as well as the interaction constraints between the resident travel behavior concept and the built environment indicator concept, and the association constraints between the resident travel behavior concept and the rail transit station concept.

[0029] In this embodiment, the urban rail transit station-city collaborative ontology is used to unify rail transit station passenger flow data, resident travel behavior data, and built environment indicator data around the stations under the same semantic framework to support subsequent semantic annotation, semantic fusion, and feature sequence generation. When constructing the urban rail transit station-city collaborative ontology, a top-level ontology is first built to define common concepts such as "entities, events, indicators, and spatiotemporal attributes" and their inheritance relationships, and to stipulate naming rules and data type constraints for station identifiers and time indexes to ensure semantic consistency of different data sources in spatiotemporal fields. Secondly, a domain ontology is constructed to refine the concepts of rail transit station, station pedestrian service area, resident travel behavior, and built environment indicator. The station pedestrian service area concept is used to represent the spatial boundary at the micro-scale of the station and to aggregate and express behavior and spatial indicators within the same spatial unit. Then, a task ontology is constructed to define the interaction constraints between the concept of resident travel behavior and the concept of built environment indicators. These interaction constraints must include at least the following relationship types: "travel behavior occurs within the pedestrian service area of ​​a station," "travel behavior is affected by built environment indicators," and "built environment indicators are updated as travel behavior changes." Constraints are also placed on the direction, value range, and temporal consistency of these relationships. Simultaneously, the association constraints between the concept of resident travel behavior and the concept of rail transit stations are defined. These association constraints must include at least the following relationship types: "the station to which the travel behavior belongs," "the station passenger flow increment generated by the travel behavior," and "the passenger flow statistics caliber of the station corresponding to the travel behavior." Constraints are also placed on the consistency of station identifiers and the rules for merging multi-source records. Finally, an application ontology is constructed to define instance verification rules and data quality rules for planning and prediction scenarios, including station identifier uniqueness verification, time index continuity verification, indicator value legality verification, and priority rules for resolving conflicting values. The urban rail transit station-city collaborative ontology is stored using machine-readable ontology description files and provides query interfaces in the form of concepts, object relationships, and data attributes. This enables the ontology to map spatiotemporal sample sequence fields to ontology instances during the semantic annotation stage, and to complete cross-source alignment and consistency checks based on the interaction and association constraints during the semantic fusion stage, thereby ensuring the uniformity and reproducibility of subsequent feature sequence extraction.

[0030] Semantic annotation of spatiotemporal sample sequences based on the urban rail transit station-city collaborative ontology includes: establishing correspondence rules between data fields and concepts and data attributes in the urban rail transit station-city collaborative ontology, and converting spatiotemporal sample sequences into ontology instances according to the correspondence rules.

[0031] In this embodiment, semantic annotation is used to convert data fields in spatiotemporal sample sequences into ontology instances consistent with the urban rail transit station-city collaborative ontology, ensuring that multi-source data is reusable, verifiable, and traceable under the same semantic caliber. Specifically, a corresponding rule file of "data field - ontology concept - data attribute" is pre-established, providing at least the field source identifier, field name, corresponding ontology concept, corresponding data attribute, data type constraint, unit constraint, and value validity constraint for each data field; among them, the station identifier field corresponds to the identifier attribute of the rail transit station concept, the time index field corresponds to the time index attribute in the spatiotemporal attributes, the station pedestrian service area field corresponds to the spatial range attribute of the station pedestrian service area concept, the rail transit station passenger flow field corresponds to the passenger flow statistics attribute of the rail transit station passenger flow concept, the resident travel behavior field corresponds to the behavioral attribute of the resident travel behavior concept, and the station surrounding built environment indicator field corresponds to the indicator attribute of the built environment indicator concept.

[0032] When converting spatiotemporal sample sequences into ontology instances, type normalization and encoding normalization are first performed on each data field to ensure that the field values ​​meet the data type and unit constraints. Value mapping is then performed on enumerated fields to ensure that the field values ​​fall within the value set defined by the value validity constraints. Subsequently, ontology instance identifiers are generated using the station identifier and time index as instance primary keys. Based on the corresponding rule file, instances of rail transit stations, pedestrian service areas, passenger flow, resident travel behavior, and built environment indicators are created sequentially. The object relationships between these instances are written into the ontology instance set. Specifically, a spatial dependency relationship is established between pedestrian service area instances and rail transit station instances; a statistical attribution relationship is established between rail transit station passenger flow instances and rail transit station instances; a spatial occurrence relationship is established between resident travel behavior instances and pedestrian service area instances; and a spatial statistical relationship is established between built environment indicator instances and pedestrian service area instances. The time index attribute is also written into these instances to maintain temporal consistency. Finally, a consistency check is performed on the generated set of ontology instances. The check includes at least the uniqueness of site identifiers, the resolvability of time indexes, the consistency of data attribute value types, and the integrity of object relationship pointers. When the check fails, the reason for the failure is recorded, and the source identifier and field name of the corresponding field are traced back to allow for correction of the corresponding rule file or source data normalization rules. Through the above semantic annotation process, a structured transformation of spatiotemporal sample sequences into ontology instances is achieved, providing a unified and reproducible data foundation for subsequent ontology-based semantic fusion and feature sequence extraction.

[0033] Semantic fusion of spatiotemporal sample sequences based on the urban rail transit station-city collaborative ontology includes: merging instances with the same station identifier and the same time index for ontology instances, and determining the values ​​of conflicting data attributes according to preset conflict resolution rules; generating feature sequences based on preset feature sorting rules for rail transit station passenger flow data attributes, resident travel behavior data attributes, and surrounding built environment indicator data attributes.

[0034] In this embodiment, semantic fusion is used to aggregate multi-source ontology instances generated during the semantic annotation stage into consistent instances at the "station identifier-time index" granularity, under the constraints of urban rail transit station-city collaborative ontology. Feature sequences that can be directly input into the prediction model are then extracted from these consistent instances. Specifically, the system constructs a fusion key using station identifiers and time indexes, establishing index mappings for rail transit station passenger flow instances, resident travel behavior instances, and built environment indicator instances, respectively. When multiple instances have the same fusion key, an instance merging process is triggered, unifying the object relationships within the instances to the instances corresponding to the same rail transit station concept and the station pedestrian service area concept, and merging data attributes at the field level. For conflicting data attributes, values ​​are determined according to preset conflict resolution rules. These rules include at least: determining value priority based on data source authority level; filtering based on the consistency of collection time and time index within the same authority level; selecting stable values ​​after consistency checks for numerical attributes and selecting unique coded values ​​after field standard encoding for categorical attributes; and filling missing data attributes according to preset filling rules. These rules include filling with values ​​from the same station at the previous time index or with statistical results from adjacent stations at the same time index, and recording filling marks for subsequent quality control. After fusion, rail transit station passenger flow data attributes, resident travel behavior data attributes, and station surrounding built environment indicator data attributes are extracted from the fused instances according to preset feature sorting rules. These preset feature sorting rules are used to fix the arrangement position of features in the vector, ensuring that the feature vector dimensions are consistent across different stations and time indices. Among these, built environment indicators are reused or updated as slowly changing features according to the time index, while resident travel behavior and station passenger flow are aggregated and generated segment by segment according to the time index as rapidly changing features. Finally, the feature vectors of each station are concatenated in ascending order of time index to form a feature sequence. The feature sequence, along with the corresponding station identifier and time index, is written into the sample storage structure for subsequent training and inference of the multi-task spatiotemporal graph convolutional neural network prediction model.

[0035] The feature sequence is input into the multi-task spatiotemporal graph convolutional neural network prediction model, and the output passenger flow prediction results and built environment index prediction results are generated. In this embodiment, the feature sequence obtained by semantic fusion is divided into model input samples by a sliding window according to the time index, and normalization processing is performed on features of different dimensions. The multi-task spatiotemporal graph convolutional neural network prediction model uses rail transit stations as graph nodes, constructs graph connectivity based on the spatial adjacency relationship between stations and the origin-destination travel relationship between stations, performs graph convolution calculation on node features using a spatial feature extraction network, and models the time series using a temporal feature extraction network. During training, an attention mechanism is introduced to dynamically adjust the weights of graph connectivity, and the inconsistency between the outputs of spatial and temporal branches is reduced based on spatiotemporal consistency constraints. The model terminal is set with a multi-task output structure, which outputs the passenger flow prediction results and the built environment index prediction results for the target prediction period, respectively. The training labels are obtained from historical passenger flow statistics and historical built environment index update values. During inference, the corresponding prediction results are directly generated based on the latest feature sequence.

[0036] The multi-task spatiotemporal graph convolutional neural network prediction model constructs graph nodes using rail transit stations and determines graph connectivity based on the geographical distance between rail transit stations and the origin-destination travel association between stations represented by residents' travel behavior data.

[0037] In this embodiment, the multi-task spatiotemporal graph convolutional neural network prediction model constructs graph nodes using rail transit stations. Each graph node is bound to a unique station identifier and carries the node features of the feature sequence generated by semantic fusion under the corresponding time index. To determine graph connectivity, spatial adjacency relationships based on geographical distance and origin-end travel association relationships based on residents' travel behavior data are constructed respectively, and the two are fused into graph connectivity relationships used for model calculation.

[0038] The process of constructing spatial adjacency relationships based on geographical distance is as follows: Obtain the geographical coordinates or center point coordinates of each rail transit station, and calculate the geographical distance between any two rail transit stations; according to the spatial adjacency determination rules recorded in the configuration file, establish spatial connections between station pairs whose geographical distance meets the spatial adjacency conditions, and mark these spatial connections as undirected connections to characterize the spatial influence range of the stations. The process of constructing origin-destination travel association relationships is as follows: Aggregate resident travel behavior data by time index, count the travel volume from the pedestrian service area of ​​the origin station to the pedestrian service area of ​​the destination station, and form origin-destination travel association relationships between stations; according to the travel association determination rules recorded in the configuration file, establish travel connections between station pairs that meet the travel association conditions, and mark these travel connections as directed connections to characterize the propagation direction of travel links on changes in station passenger flow and surrounding built environment indicators. The fusion process is as follows: For station pairs that have both spatial and travel connections, determine the final connection attribute according to the fusion priority and weight rules recorded in the configuration file; for station pairs that have only one type of connection, retain the corresponding connection attribute. The graph connectivity obtained through the above method reflects both the spatial adjacency structure of rail transit stations and the inter-station association structure driven by residents' travel behavior, enabling the prediction model to learn the combined effects of "spatial adjacency influence" and "travel link influence" in the station-city collaboration scenario under the same graph structure.

[0039] The multi-task spatiotemporal graph convolutional neural network prediction model includes a spatial feature extraction network and a temporal feature extraction network. The spatial feature extraction network generates dynamic adjacency relationships based on a self-attention mechanism and performs graph convolution calculation on the node features corresponding to the graph connection relationships. The temporal feature extraction network performs time series modeling on the feature sequence based on a temporal dimension attention mechanism. During training, the consistency between the output of the spatial feature extraction network and the output of the temporal feature extraction network is constrained based on a spatiotemporal consistency loss function. In addition, a multi-task output structure is set up for passenger flow prediction output and built environment indicator prediction output.

[0040] In this embodiment, the multi-task spatiotemporal graph convolutional neural network prediction model consists of a spatial feature extraction network and a temporal feature extraction network. The spatial feature extraction network constructs graph nodes using rail transit stations as examples. Node features are derived from feature sequences generated by semantic fusion, and spatial information propagation is performed on the node features in conjunction with graph connectivity relationships. To enable the spatial connectivity relationships to adaptively adjust with changes in time and features, the spatial feature extraction network introduces a self-attention mechanism to generate dynamic adjacency weights and weightedly aggregates the influence strength of adjacent nodes. The expression for calculating the attention coefficient is:

[0041] For nodes With nodes Attention coefficient between them; A learnable parameter vector; This is a vector join operation; The parameter matrix is ​​a linear transformation matrix; For nodes The node features. The attention coefficients are normalized to obtain the normalized attention coefficients:

[0042] This is the normalized attention coefficient; For nodes The set of adjacent nodes. The features of adjacent nodes are weighted and aggregated based on normalized attention coefficients to update node features:

[0043] For nodes Updated node features; This is the activation function.

[0044] The temporal feature extraction network is used to perform time series modeling on feature sequences and highlights key time period information based on a time-dimensional attention mechanism. The expression for calculating the time-dimensional attention weights is as follows:

[0045] For time step The hidden state; These are the learnable parameters for the attention layer; The attention layer is used as a learnable bias. The hidden states are weighted and converged to obtain a temporal feature representation:

[0046] This represents the weighted temporal features. During training, to ensure consistency between the outputs of the spatial feature extraction network and the temporal feature extraction network, a spatiotemporal consistency loss term is constructed:

[0047] Output the graph using convolution; The output is the hidden state. The prediction error is then combined with the spatiotemporal consistency loss term to obtain the total loss:

[0048] These are the model's predicted values; This is a real label; This is a hyperparameter.

[0049] In terms of output structure, the model sets up a multi-task output structure for passenger flow prediction output and built environment index prediction output. After fusing spatial feature representation and temporal feature representation, they are input into independent prediction heads to output passenger flow prediction results and built environment index prediction results, so as to support the joint decision-making of passenger flow trend analysis and built environment index evolution prediction in station-city collaborative planning.

[0050] like Figure 2 As shown, to achieve joint prediction of future trends in passenger flow and spatial characteristics, this invention extends the AST-GCN model by employing a multi-task output strategy. By setting independent prediction heads in the model, AST-GCN can simultaneously generate numerical predictions of passenger flow and trend predictions of spatial characteristics. This design fully utilizes the model's ability to model spatiotemporal relationships, going beyond single-task prediction to provide more comprehensive analytical results.

[0051] Specifically, the model's output layer is divided into two independent prediction heads: a passenger flow prediction head, which predicts passenger flow values ​​for future time steps. A fully connected layer maps the hidden states of the LSTM (capturing temporal dynamics) and the output of the GCN (encoding spatial characteristics) to the predicted passenger flow values. This output is a direct numerical prediction, suitable for quantitative analysis. The spatial feature prediction head predicts future trends in spatial attributes. Although the input spatial features (such as station location or type) are static, the model learns the interaction between static spatial features and dynamic temporal features (such as passenger flow changes) to predict potential changes in spatial features (such as station importance or regional traffic distribution trends). Through this multi-task learning strategy, AST-GCN can not only accurately predict future passenger flow but also capture the dynamic trends of spatial features. For example, it can predict how a surge in passenger flow at certain stations during peak hours affects the spatial importance of surrounding stations, thus achieving a comprehensive model of the coupling relationship between the urban rail transit system and urban space. This dual-output design fully utilizes the spatial feature extraction capabilities of GCN, the temporal series modeling capabilities of LSTM, and the dynamic adjustment capabilities of the attention mechanism, ensuring that both tasks are optimized independently while sharing feature representations. The spatiotemporal consistency loss measures the difference between the GCN output and the LSTM hidden state, deriving spatial and temporal consistency from the learning processes of the GCN and LSTM modules. This loss allows the model to better integrate graph structure and sequence information, improving its ability to model spatiotemporal data. The AST-GCN model can not only learn dynamic changes over time but also fully utilize the spatial dependencies within the rail transit system. This enhances the model's ability to predict traffic flow or other rail system behaviors, achieving a modeling effect that couples urban rail transit with urban space. Thanks to the introduction of a multi-task output strategy, AST-GCN not only provides passenger flow predictions but also reveals the evolution trend of spatial characteristics, offering richer decision-making basis for urban planning and traffic management.

[0052] Model training hyperparameters, optimization strategies, and comparative experiments were conducted for validation. The training setup included: optimizer: AdamW (lr=0.001, weight_decay=1e-5); learning rate scheduling: CosineAnnealingLR + EarlyStopping (patience=15); batch size: 32; sequence length: T=96; prediction of four future time periods (1 hour each); data partitioning: training in 2022-2023 (80%), validation in January-March 2024, and testing in April-June 2024; training converged after 200 epochs, with a single epoch taking 47 seconds.

[0053] Comparison experiments with benchmark models (LSTM, ST-GCN): LSTM model (3 layers, hidden=128, dropout=0.3): pure sequence model, no spatial information, MAE=58.62, RMSE=89.47, MAPE=13.41%; ST-GCN model (3 blocks, no attention, fixed adjacency matrix): MAE=51.29, RMSE=77.35, MAPE=11.63%; AST-GCN of this invention: MAE=42.37, RMSE=68.14, MAPE=9.28%, an improvement of 27.7% over LSTM and 17.4% over ST-GCN; Ablation experiments show that: removing only self-attention → MAPE increases by 14.3%; removing only spatiotemporal consistency loss → MAPE increases by 21.7%; removing both → MAPE=12.05%, verifying that both innovations are indispensable.

[0054] This step is the first to improve the self-attention mechanism in station-city coupled scenarios by using a dual algorithm of "dynamic spatial weights + temporal weighted summarization." It also uniquely creates a spatiotemporal consistency loss function to force constraints on the output of GCN-LSTM, completely solving the problems of insufficient capture of coupling mechanisms and unreliable predictions in complex scenarios caused by the "fixed graph + no physical constraints" of existing models. It possesses the most prominent substantive features, creativity, and significant progress, constituting the core invention point that distinguishes it from existing technologies. Objective verification through comparison of LSTM-1 / LSTM-2 / AST-GCN systems shows that simply adding features (LSTM-2) already outperforms LSTM-1, but adding the three innovations of this invention further improves accuracy by 21.7%-36.8%, especially in complex spatiotemporal scenarios at midday, where the advantage is an order of magnitude, fully demonstrating the invention's ability to capture nonlinear coupling relationships. The model directly embeds 54-dimensional features of bidirectional interaction between "behavior + space," with richness far exceeding existing technologies, achieving perfect loop closure. This invention provides a highly reliable foundation for GA-SHAP interpretability analysis, which is the core innovation point that distinguishes it from existing technologies.

[0055] Feature contribution quantification and key element identification based on GA-SHAP. This step directly follows the trained AST-GCN model and proposes and implements for the first time the Genetic Algorithm optimized SHapley Additive ex Planations (GA-SHAP). It accurately quantifies the global / local contributions and analyzes interaction effects of 54-dimensional "behavioral + spatial" features, completely solving the technical bottleneck of insufficient interpretability caused by the "black box" problem of existing deep learning models. This represents a fundamental breakthrough in making the station-city coupling and coordination mechanism "unknowable" to "quantifiable, visualized, and plannable for intervention." In an empirical study on Taiyuan Metro Line 2, it successfully identified the top 10 key elements with the greatest impact on coupling and coordination, explaining 97.6% of the variance, providing a scientific and operable indicator system for planning decision support. The AST-GCN model is a complex nonlinear spatiotemporal deep network; traditional TreeSHAP cannot be directly applied (it's a non-tree model), so KernelSHAP must be used. However, KernelSHAP has a computational complexity of O(2^F × N_samples). When F = 54-dimensional features, calculating the complete Shapley value for a single sample requires 2^54 forward propagations, which is practically infeasible. To obtain high-precision SHAP values ​​within a reasonable timeframe, this invention uniquely introduces a genetic algorithm for global optimization selection of the background sample set. In SHAP libraries, GradientExplainer and DeepExplainer are commonly used tools for interpreting deep learning models. However, this invention chooses to manually calculate the gradient (GradientSHAP, G-SHAP) based on GradientExplainer to calculate the SHAP value. The AST-GCN model integrates GNN and LSTM, and its internal structure is complex, involving operations such as graph convolution and temporal convolution. GradientExplainer and DeepExplainer perform well when processing traditional feedforward neural networks or standard LSTM models, but for complex models like AST-GCN, they have been verified to have compatibility issues and cannot accurately capture the internal graph structure and spatiotemporal dependencies of the model. In highly nonlinear models, gradients may vanish or explode, leading to unreliable interpretations. Interpreters assume a finite degree of nonlinearity in the model, resulting in significant approximation errors for complex nonlinear models. The G-SHAP method offers greater control over the computation process, allowing adjustments based on model characteristics and requirements. Since the model constructed in this invention outputs multidimensional data, SHAP values ​​need to be calculated dimension by dimension. G-SHAP can handle this situation more flexibly, while interpreters may be limited when dealing with multiple output dimensions. Gradient-based SHAP value calculation methods share similarities with integrated gradient methods.The integral gradient method quantifies the importance of features by calculating the gradient integral between the input features and the zero vector to the actual input.

[0056] In this invention, due to computational complexity and practical considerations, an approximation method is adopted: directly calculating the gradient of the output with respect to the input without integration. This method assumes that the change between the input and the reference point is linear, and the gradient remains constant within this interval. Although this approximation may introduce some errors, it has proven effective in capturing the importance of features in practice. Building upon the gradient-based SHAP value calculation method, this invention further introduces an attention mechanism to enhance the characterization of feature contribution. The attention mechanism models the weights of different features, ensuring that the importance of a feature depends not only on gradient information but also on the model's attention at a specific time step or in a specific context.

[0057] The original AST-GCN and LSTM models are encapsulated to facilitate input and output control, and to simplify gradient calculation and feature importance analysis. In the encapsulated URTUSISM model designed in this invention, a function to return attention weights is added to support the use of the attention mechanism. When `return_attention=True`, the model output includes not only the prediction result but also the attention weights. These attention weights reflect the model's focus in the spatiotemporal sequence, providing further information for calculating the G-SHAP value. The weighted G-SHAP calculation method, which combines gradients with attention weights, is called GA-SHAP. First, attention weights are obtained during the model's forward propagation. Then, the gradient values ​​of each feature are weighted and multiplied by their corresponding attention weights. Finally, the weighted gradients are accumulated and averaged to obtain the weighted SHAP value.

[0058] The feature contribution is calculated based on the Shapley additive explanatory value, and the background sample set used to calculate the Shapley additive explanatory value is optimized and selected based on the genetic algorithm. The set of key factors is determined based on the feature contribution, and the planning indicator thresholds and intervention parameters are generated based on the set of key factors.

[0059] In one possible implementation, when performing interpretability analysis on a trained spatiotemporal prediction model, candidate background samples are extracted from historical samples. These background samples are used to construct a reference distribution when calculating the Shapley additive interpretation value. The selection results of the candidate background samples are encoded into individuals using a genetic algorithm. Fitness is constructed based on the stability of the interpretation results and sample coverage. Iterative selection, crossover, and mutation are performed to obtain an optimized background sample set. Subsequently, based on GradientExplainer, the SHAP values ​​of features are calculated on the optimized background sample set, and attention weights are used to correct the feature contribution distribution. The feature contribution values ​​are then summarized, and several features with the highest contribution values ​​are selected as the set of key contributing elements. Furthermore, the key contributing elements are mapped to planning indicators, and the threshold of the planning indicators is determined by combining historical distribution and planning constraints. When the planning indicators exceed the threshold, intervention measures parameters are output according to preset intervention rules for adjustment and optimization of the planning scheme.

[0060] The calculation of feature contribution based on Shapley additive interpretation values ​​includes: calculating the contribution of each feature in the feature sequence to the prediction results for both passenger flow prediction results and built environment indicator prediction results, and determining the set of key contributing factors based on feature contribution.

[0061] In this embodiment, when calculating feature contribution based on Shapley additive interpretation values, the feature sequence generated by semantic fusion is used as the interpretation object, and the trained multi-task spatiotemporal graph convolutional neural network prediction model is used as the interpreted model. To ensure consistency in interpretability between passenger flow prediction results and built environment indicator prediction results, the feature sorting rules and time index window rules of the feature sequence are first fixed to ensure that the input features of the same station identifier under the same time index remain consistent in the position of the model input tensor. Subsequently, a background sample set is extracted from the historical spatiotemporal sample sequence, and the background sample set is optimized and selected according to a preset background sample set optimization process to ensure that the background sample set covers different station types, different time periods, and different built environment indicator states.

[0062] When calculating the feature contribution, Shapley additive explanatory values ​​are calculated for both passenger flow prediction results and built environment indicator prediction results: For passenger flow prediction results, the corresponding passenger flow prediction output in the multi-task output structure is used as the explanatory target. A background sample set is used to replace or mask individual features in the feature sequence, forming a set of perturbation inputs. For each perturbation input, the change in passenger flow prediction output is recorded, and the marginal contribution is calculated on a preset feature addition order set. The summations yield the Shapley additive explanatory value of each feature for the passenger flow prediction result. For built environment indicator prediction results, the corresponding built environment indicator prediction output in the multi-task output structure is selected as the explanatory target. A background sample set and perturbation rules consistent with passenger flow prediction are used to calculate the Shapley additive explanatory value of each feature for the built environment indicator prediction result. To adapt to the time series structure of the feature sequence, the Shapley additive explanatory values ​​generated by the same feature under different time indices are summarized according to preset aggregation rules to obtain the overall contribution of the feature within the target prediction period, forming a passenger flow prediction contribution set and a built environment indicator prediction contribution set, respectively.

[0063] When determining the set of key elements based on feature contribution, the contribution sets for passenger flow prediction and built environment indicator prediction are sorted and filtered separately: features with high absolute contribution values ​​that remain stable across multiple time index windows are selected as candidate key elements, and contribution signs are recorded to characterize the direction of the feature's influence on the prediction results. Subsequently, a dual-task consistency screening is performed, prioritizing features that simultaneously contribute highly to both passenger flow prediction and built environment indicator prediction results into the set of key elements; features that only show high contribution in a single task are supplemented according to preset task weight rules and planning target priority rules. The final output set of key elements is consistent with the feature sorting rules of the feature sequence, and each key element is associated with its corresponding data attribute source, station pedestrian service area statistical scope, and time index window scope. This allows the set of key elements to be directly used for the generation of subsequent planning indicator thresholds and intervention measure parameters, and ensures the traceability and reproducibility of the interpretation results in engineering implementation.

[0064] The optimization selection of the background sample set used to calculate the Shapley additive interpretation value based on the genetic algorithm includes: performing selection, crossover, and mutation in the background sample set to determine the target background sample set; generating planning indicator thresholds and intervention parameters includes: transforming the set of key factors into planning indicators, and establishing the correspondence between factors, indicators, thresholds, and interventions based on the planning indicators to output planning indicator thresholds and intervention parameters.

[0065] In this embodiment, to reduce the sensitivity of the Shapley additive explanatory value to the selection of the background sample set and to ensure that the feature contribution can stably reflect the common driving factors of passenger flow prediction and built environment indicator prediction, a genetic algorithm is used to optimize the selection of the background sample set used to calculate the Shapley additive explanatory value. First, candidate background samples are extracted from the historical spatiotemporal sample sequence according to station identifier and time index. The candidate background samples cover different passenger flow states and different built environment indicator states, and maintain the same feature sorting rules as the feature sequence. Then, the selection scheme of the candidate background samples is encoded into a genetic algorithm individual, which indicates the set of sample identifiers selected into the background sample set from the candidate background samples. The Shapley additive explanatory value of the passenger flow prediction output and the Shapley additive explanatory value of the built environment indicator prediction output are calculated using the background sample set corresponding to the individual. An explanatory stability evaluation is generated based on the consistency of multiple repeated calculations, and a prediction retention evaluation is generated based on the degree of retention of the original prediction output by the background sample set. The explanatory stability evaluation and the prediction retention evaluation are combined to form a fitness evaluation. The genetic algorithm iteratively performs selection, crossover, and mutation on the candidate background sample set: selection is used to retain individuals with high fitness evaluation, crossover is used to exchange sample identifiers between different individuals to form new background sample set combinations, and mutation is used to randomly replace some sample identifiers to expand the search space; when the fitness evaluation converges or reaches the preset iteration termination condition, the target background sample set is determined and used for subsequent feature contribution calculation.

[0066] When generating planning indicator thresholds and intervention parameters, the first step is to determine a set of key contributing elements based on their feature contribution, and then transform this set into planning indicators. The transformation rules are determined by the semantic mapping between the elements and the built environment indicator data attributes, ensuring that each key contributing element corresponds to a unique planning indicator name and statistical scope. The calculation scope of the planning indicators is consistent with the pedestrian service area of ​​the stations and aligned with the time index. Subsequently, a correspondence is established between elements, indicators, thresholds, and intervention measures based on the planning indicators. This correspondence is stored in a rule base, which includes at least threshold generation rules for planning indicators and intervention parameter generation rules. The threshold generation rules determine the threshold boundaries based on historical distribution, planning constraints, and predicted trends, and record the applicable station types and time index ranges. The intervention parameter generation rules are used to output intervention parameters when planning indicators exceed or fall below the threshold. Intervention parameters include at least the spatial unit identifier of the intervention object, the name of the intervention indicator, the intervention direction, and the intervention intensity level, and the intervention intensity level is mapped to a set of executable planning adjustment instructions. The final output planning indicator thresholds and intervention measure parameters correspond one-to-one with the set of key elements, and retains the association information with station identification, station pedestrian service area scope and time index scope, enabling planners to transform the interpretation results into implementable station-city collaborative planning decision inputs.

[0067] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.

[0068] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. An ontology-based station-city collaborative data integration and planning prediction method, characterized in that, The method comprises the following steps: obtaining rail transit station passenger flow data, resident travel behavior data and station surrounding built environment index data, and forming a space-time sample sequence with station identification and time index according to a unified space-time granularity of a station walking service area; constructing a city rail transit station-city coordination ontology, performing semantic annotation and semantic fusion on the space-time sample sequence based on the city rail transit station-city coordination ontology, and generating a feature sequence; inputting the feature sequence into a multi-task space-time graph convolutional neural network prediction model to output passenger flow prediction results and built environment index prediction results; calculating feature contribution degree based on Shapley additive explanation value, optimizing and selecting a background sample set used for calculating Shapley additive explanation value based on a genetic algorithm, determining a key role element set based on the feature contribution degree, and generating planning index threshold and intervention measure parameters based on the key role element set.

2. The method of claim 1, wherein, The station surrounding built environment index data includes point of interest density index, land use mixing degree index, land use intensity index, building density index, road network density index, road network accessibility index, residential population density index, employment density index, green coverage index and public transport connection index.

3. The method of claim 1, wherein, Forming a space-time sample sequence with station identification and time index according to a unified space-time granularity of a station walking service area comprises: performing spatial matching on rail transit station passenger flow data, resident travel behavior data and station surrounding built environment index data to associate to a station walking service area, and performing time index alignment on the spatially matched data to generate a space-time sample sequence.

4. The method of claim 1, wherein, Constructing a city rail transit station-city coordination ontology comprises: constructing a top-level ontology, a domain ontology, a task ontology and an application ontology, and defining rail transit station concepts, station walking service area concepts, resident travel behavior concepts and built environment index concepts in the city rail transit station-city coordination ontology, and defining interaction relationship constraints between resident travel behavior concepts and built environment index concepts and association relationship constraints between resident travel behavior concepts and rail transit station concepts.

5. The method of claim 1, wherein, Performing semantic annotation on the space-time sample sequence based on the city rail transit station-city coordination ontology comprises: establishing a corresponding rule between data fields and concepts and data attributes in the city rail transit station-city coordination ontology, and converting the space-time sample sequence into ontology instances according to the corresponding rule.

6. The method of claim 1, wherein, Performing semantic fusion on the space-time sample sequence based on the city rail transit station-city coordination ontology comprises: performing instance merging of the same station identification and the same time index on the ontology instances, and determining the values according to the preset conflict resolution rules for the conflicting data attributes; generating a feature sequence according to a preset feature sorting rule of rail transit station passenger flow data attributes, resident travel behavior data attributes and station surrounding built environment index data attributes.

7. The method of claim 1, wherein, The multi-task space-time graph convolutional neural network prediction model constructs a rail transit station graph node, and determines a graph connection relationship based on geographical distances between rail transit stations and station-to-station trip association relationships represented by resident travel behavior data.

8. The method of claim 1, wherein, The multi-task spatio-temporal graph convolutional neural network prediction model comprises a spatial feature extraction network and a temporal feature extraction network, the spatial feature extraction network generates dynamic adjacency relationships based on a self-attention mechanism and performs graph convolution calculation on node features corresponding to graph connection relationships, the temporal feature extraction network performs time series modeling on feature sequences based on a time dimension attention mechanism, and the consistency of the output of the spatial feature extraction network and the output of the temporal feature extraction network is constrained based on a spatio-temporal consistency loss function during the training process, and a multi-task output structure of passenger flow prediction output and built environment index prediction output is set.

9. The method of claim 1, wherein, The feature contribution degree based on Shapley additive explanation value calculation comprises: calculating the contribution degree of each feature in the feature sequence to the prediction result for the passenger flow prediction result and the built environment index prediction result respectively, and determining a key role element set based on the feature contribution degree.

10. The method of claim 1, wherein, The optimization selection of the background sample set for calculating the Shapley additive explanation value based on the genetic algorithm comprises: performing selection, crossover and mutation in the background sample set to determine a target background sample set; and the generation of the planning index threshold and the intervention measure parameter comprises: converting the key role element set into a planning index, and establishing a corresponding relationship of element-index-threshold-intervention measure based on the planning index to output the planning index threshold and the intervention measure parameter.

Citation Information

Patent Citations

  • Ontological semantic extension and collaborative filter weighting fused process knowledge retrieval method

    CN107885749A

  • Rail transit passenger flow prediction method considering dynamic space-time correlation

    CN113298314A

  • Rail transit short-time passenger flow prediction method and device based on multi-scale convolution and storage medium

    CN120124800A

  • Urban rail transit short-time origin-destination passenger flow prediction method, device and medium

    CN120875118A

  • Short-term subway passenger flow prediction method, system and electronic device

    WO2021098619A1

Cited By

  • Method, system, medium and equipment for predicting railway passenger flow of large city

    CN122114296A

  • A method, system, medium and device for predicting rail passenger flow in a large city

    CN122114296B

  • Intelligent screening method and system for water and soil environment assessment indexes in antimony mine area

    CN122310066A