Rail transit cross-modal prediction method based on search-enhanced reasoning and spatio-temporal knowledge graph
Patent Information
- Application Number
- CN202610855374.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-06-15
AI Technical Summary
[0008]为了解决现有技术中的上述问题,即长尾突发事件预测精度低、语义对齐不足及预测缺乏可解释性的问题,本发明提供了一种基于检索增强推理与时空知识图谱的轨道交通跨模态预测方法,包括:
Smart Images

Figure CN122388520B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent prediction for rail transit, and specifically relates to a cross-modal prediction method for rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graph. Background Technology
[0002] With the continuous expansion of urban rail transit network operations, passenger flow forecasting has become a core technological means to support dynamic capacity allocation, improve operational efficiency, and enhance passenger travel experience. Currently, technical solutions in this field are evolving from traditional statistical methods such as autoregressive integral moving average models to spatiotemporal prediction models based on deep learning. Typical models include Long Short-Term Memory (LSTM) networks, Transformers and their variants, and spatiotemporal graph convolutional networks. These models learn the global statistical patterns of passenger flow changes over time and space from large-scale historical operational data, achieving good predictive results in routine daily passenger flow scenarios.
[0003] Regarding the further integration of external influencing factor data, existing research has begun to explore incorporating heterogeneous data such as meteorological information, holiday markers, and large-scale event notices as auxiliary features into predictive models. Some cutting-edge solutions also utilize large language models combined with retrieval-enhanced generative architectures to process unstructured operational notices or information texts, and manage station topology, line transfer relationships, and operational rules by constructing a spatiotemporal knowledge graph for rail transit, thereby providing domain knowledge support for prediction and scheduling decisions.
[0004] Existing rail transit passenger flow forecasting technologies still have some shortcomings in practical applications. Current forecasting models mainly rely on global statistical regularities for inference. For sudden events that account for a very small proportion of the training data but have a significant impact, such as ground traffic paralysis caused by extreme rainstorms or a sudden increase in rail transit connection pressure caused by large-scale airport flight delays, the models have difficulty drawing on historical experience in handling similar situations, and often give forecast results that deviate significantly from the actual situation.
[0005] Existing methods, when introducing external information such as weather and events, often convert it into discrete labels or simple numerical vectors and then concatenate it with passenger flow features. They lack the ability to explore the deep causal relationship between passenger flow value sequences and event semantic descriptions, resulting in weak generalization ability of the model when converting unstructured information or sudden event notifications into quantitative passenger flow impacts.
[0006] The widely used deep time series models are essentially end-to-end black-box mappings. When outputting predicted values for future periods, they cannot simultaneously provide the historical experience or knowledge logic on which the prediction is based. Schedulers find it difficult to judge the reliability of the predicted values, and the system's auxiliary decision-making utility is limited in critical decision-making scenarios involving traffic adjustments.
[0007] Existing knowledge management systems are mostly built in one go, making it difficult to automatically expand and update existing knowledge items or historical case libraries based on new experiences and patterns accumulated during actual operation. This results in poor adaptability of the system to dynamic changes in the rail transit operating environment. Summary of the Invention
[0008] To address the aforementioned problems in existing technologies, namely low prediction accuracy, insufficient semantic alignment, and lack of interpretability in predictions of long-tail emergencies, this invention provides a cross-modal prediction method for rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graphs, comprising:
[0009] Historical passenger flow data and text data describing environmental conditions are obtained, and passenger flow temporal and semantic features are extracted. Contrastive learning is used to map the passenger flow temporal and semantic features to a unified vector space to form a vector index library. Obtain real-time passenger flow fluctuation sequence and real-time semantic tags, extract real-time passenger flow temporal features and real-time semantic features, generate query vectors, sort the vector index library according to similarity from high to low, select a preset number of historical context segments with the highest ranking, each of the historical context segments contains historical passenger flow temporal features, historical semantic features and corresponding original passenger flow data and text data; The selected historical context fragments are associated with corresponding entities in the spatiotemporal knowledge graph. Based on the causal logic in the graph, diffusion reasoning is performed to obtain logical constraint information. A time-series prediction model is used to make a preliminary prediction of the current passenger flow, and a preliminary prediction value sequence is obtained. The preliminary prediction value sequence, the selected historical context fragments and the logical constraint information are used to construct a prompt message, which is then input into a large language model to output the optimized prediction value and natural language explanation. Based on the optimized predicted values, a capacity adjustment strategy is generated, and the feature pairs consisting of the real-time passenger flow time-series features and the real-time semantic features are stored in the vector index library.
[0010] Furthermore, historical passenger flow data and textual data describing environmental conditions are acquired, and temporal and semantic features of passenger flow are extracted. Contrastive learning is then used to map these temporal and semantic features to a unified vector space, forming a vector index library, including: The acquired historical passenger flow data is divided into multiple patches according to time windows, and the temporal characteristics of passenger flow in each patch are extracted using a spatiotemporal encoder. Semantic features are obtained by semantically encoding text data for the corresponding time period using a large language model; A contrastive loss function is constructed based on noise contrastive estimation. Passenger flow temporal features and semantic features under the same context are used as positive sample pairs, and passenger flow temporal features and semantic features under different contexts are used as negative sample pairs to calculate the contrastive loss. The parameters of the spatiotemporal encoder and the large language model are updated by backpropagating the contrastive loss. The aligned passenger flow time sequence features and semantic features are combined to form feature pairs, which are then stored in the vector index library.
[0011] Furthermore, real-time passenger flow fluctuation sequences and real-time semantic tags are obtained, real-time passenger flow temporal features and real-time semantic features are extracted, and query vectors are generated. These vectors are then sorted in the vector index library according to similarity from high to low, and a predetermined number of historical context segments are selected from the top-ranked segments, including: The real-time passenger flow fluctuation sequence is input into the spatiotemporal encoder to obtain real-time passenger flow time-series features. The real-time semantic label is input into the large language model to obtain real-time semantic features. The real-time passenger flow time-series features and real-time semantic features are concatenated or weighted and fused to generate a query vector. An approximate nearest neighbor search is performed in the vector index library, returning a preset number of historical context fragments that rank highly similar to the query vector. Each historical context fragment contains aligned historical passenger flow time-series features and historical semantic features, as well as their corresponding original passenger flow data and text data.
[0012] Furthermore, the selected historical context fragments are associated with corresponding entities in the spatiotemporal knowledge graph. Based on the causal logic in the graph, diffusion reasoning is performed to obtain logical constraint information, including: Based on the event types and stations identified in the historical context fragments, trigger nodes are located in the spatiotemporal knowledge graph, and the corresponding station nodes and event nodes are activated. Starting from the trigger node, multi-hop diffusion is performed along the predefined operation rule edges and event causal propagation edges in the spatiotemporal knowledge graph. In each hop, the influence state of the current node is passed to the adjacent node along the edge. The adjacent node aggregates the received influence state and updates its own influence state. The passenger flow influence weight of the adjacent station is iteratively calculated to generate logical constraint information containing the affected area and the distribution of passenger flow influence.
[0013] Furthermore, a time-series prediction model is used to make a preliminary prediction of the current passenger flow, obtaining a preliminary predicted value sequence. This preliminary predicted value sequence, selected historical context segments, and the logical constraint information are then used to construct a prompt message, which is input into a large language model. The model outputs optimized predicted values and a natural language explanation, including: The current passenger flow is predicted using a patch-based time-series forecasting model, and the preliminary forecast value sequence for the future preset time period is output. The preliminary predicted value sequence is segmented and aggregated to approximate it, and then converted into a text description sequence containing trends and values. The text description sequence, the historical passenger flow actual value text and event description text in the historical context fragment, and the textual representation of the logical constraint information are assembled into an enhanced prompt word according to the hierarchical structure of instruction layer, example layer, and constraint layer. The enhanced prompt word contains instructions from the constraint language model to correct the preliminary predicted value sequence based on the historical context fragment and the logical constraint information. The enhanced prompt words are input into the large language model, and the corrected prediction values, the historical case identifiers on which the corresponding correction points are based, and the reasons for the corrections are output.
[0014] Furthermore, a capacity adjustment strategy is generated based on the optimized predicted values, and the feature pairs consisting of the real-time passenger flow time-series features and the real-time semantic features are stored in the vector index library, including: The optimized predicted value is compared with the preset scheduling trigger threshold. When the predicted value exceeds the threshold, a capacity adjustment strategy including the departure interval compression ratio or the number of standby vehicle formations is generated based on the magnitude of the exceedance and the time period. Extract the real-time passenger flow time-series features and real-time semantic features used when generating the query vector, form feature pairs, and insert them into the vector index library.
[0015] Furthermore, the contrastive loss function is the normalized temperature cross-entropy loss, which is calculated based on the similarity of positive sample pairs and the similarity of negative sample pairs. Positive sample pairs consist of passenger flow temporal features and semantic features under the same context, while negative sample pairs consist of passenger flow temporal features and semantic features under different contexts. The parameters of the spatiotemporal encoder and the large language model are updated by backpropagating the contrastive loss function.
[0016] Furthermore, the spatiotemporal knowledge graph is constructed through the following process: Station nodes, line edges, event nodes, and their attributes are obtained from rail transit operation records and event notices using named entity recognition and relation extraction. Define event causal propagation edges and spatial adjacency edges according to domain rules, assign propagation weights to the event causal propagation edges and spatial adjacency edges respectively, and form a directed graph structure; In the multi-hop diffusion, the graph attention mechanism calculates the attention coefficient of the source node to the neighbor node at each hop. The neighbor node aggregates the influence from different source nodes according to the attention coefficient and updates the state of the neighbor node. After multiple iterations, the degree of passenger flow influence of each node is output.
[0017] Furthermore, the enhanced prompt words also include a set of trigger instructions, which are used to enable the large language model to extract historical case identifiers from the historical context fragments when generating corrected prediction values, and to generate correction basis based on the passenger flow evolution pattern of the historical cases and the logical constraint information, and to output the historical case identifiers in the output natural language explanation.
[0018] Furthermore, the step of storing the feature pairs composed of the real-time passenger flow time-series features and the real-time semantic features into the vector index library also includes: Calculate the similarity between the newly inserted feature pair and existing feature pairs in the vector index library. If there is an existing feature pair with a similarity exceeding a preset threshold, delete the existing feature pair and write the newly inserted feature pair into the vector index library.
[0019] The beneficial effects of this invention are: This invention maps historical passenger flow time-series features and corresponding semantic features to a unified vector space through contrastive learning, and constructs a vector index library, making data from different modalities comparable within the same metric space. It can support cross-modal similar historical context retrieval based on real-time passenger flow and semantic information, and provide historical reference fragments that are highly relevant to the current context for subsequent prediction.
[0020] This invention generates query vectors based on real-time passenger flow fluctuation sequences and real-time semantic tags, and selects a preset number of historical context fragments from the vector index library in order of similarity. It can retrieve similar cases with reference value from historical data in the case of long-tail sudden events or complex operational scenarios, making up for the lack of experience of pure parameterized models for low-frequency but significant events.
[0021] This invention associates retrieved historical context fragments with corresponding entities in a spatiotemporal knowledge graph, and uses operational rule edges and event causal propagation edges in the graph to perform multi-hop diffusion reasoning, obtaining logical constraint information that includes the distribution of affected areas and the degree of passenger flow impact. This allows the local triggering impact of events to propagate and quantify along the road network topology, providing explicit constraints that conform to the physical laws of the domain for prediction.
[0022] This invention employs a time-series prediction model to obtain a preliminary predicted value sequence. This sequence, along with historical context fragments and logical constraint information, is then used to construct a prompt message that is input into a large language model. The model outputs optimized predicted values and natural language explanations, enabling the final prediction result to integrate data-driven time-series trends, retrieved historical experience, and explicit causal logic from the knowledge graph. Furthermore, it provides historical case identifiers and reasons for corrections, enhancing the transparency of the prediction process and the traceability of the results.
[0023] This invention generates a capacity adjustment strategy based on optimized forecast values, which includes the departure interval reduction ratio or the number of spare vehicle formations. It can directly convert the forecast results into executable scheduling suggestions, reducing the manual analysis steps from numerical forecasting to operational decision-making.
[0024] This invention combines real-time passenger flow time-series features and real-time semantic features used to generate query vectors into feature pairs and stores them in a vector index library. This enables the self-guided expansion of the historical experience library, allowing the system to continuously accumulate new event-passenger flow correspondences during actual operation. It can adapt to the dynamic changes in the operating environment without frequent retraining of the core model.
[0025] In constructing the vector index library, this invention employs a normalized temperature cross-entropy loss function based on noise contrast estimation. Feature pairs in the same context are used as positive samples, and feature pairs in different contexts are used as negative samples. This updates the parameters of the spatiotemporal encoder and the large language model, which helps improve the accuracy of cross-modal feature alignment and thus enhances the matching degree between the retrieval results and the current context.
[0026] The spatiotemporal knowledge graph of this invention extracts station nodes, line edges, event nodes and their attributes from operation records and event announcements, defines event causal propagation edges and spatial adjacency edges and transmission weights, and then calculates attention coefficients and aggregates influences in multi-hop diffusion using graph attention mechanisms. This can meticulously characterize the cascading impact range and degree of emergencies in the rail transit network, and provide structured domain knowledge support for the generation of logical constraint information.
[0027] When assembling the preliminary predicted value sequence, historical context fragments, and logical constraint information into enhanced prompt words, this invention adopts a hierarchical structure of instruction layer, example layer, and constraint layer, and includes trigger instructions. This forces the large language model to extract historical case identifiers from historical context fragments as a basis for correction, ensuring that the output natural language explanation clearly indicates the historical cases cited, thereby providing schedulers with verifiable decision-making references.
[0028] During the update process of the vector index library, this invention calculates the similarity between newly inserted feature pairs and existing feature pairs, and deletes existing feature pairs and writes new feature pairs when the similarity exceeds a preset threshold. This can control the redundancy in the vector index library, maintain retrieval efficiency, and preserve the timeliness of the memory. Attached Figure Description
[0029] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart of a cross-modal prediction method for rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graphs according to the present invention; Figure 2 This is a structural diagram of a rail transit cross-modal prediction system based on retrieval-enhanced reasoning and spatiotemporal knowledge graph according to the present invention; Figure 3 This is a schematic diagram of the structure of a computer system used to implement the methods, systems, and electronic devices of this application. Detailed Implementation
[0030] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0031] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0032] The first embodiment of the present invention provides a method for cross-modal prediction of rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graphs, comprising: Step S10: Obtain historical passenger flow data and text data describing environmental conditions, extract passenger flow temporal features and semantic features, and use contrastive learning to map passenger flow temporal features and semantic features to a unified vector space to form a vector index library. Step S20: Obtain real-time passenger flow fluctuation sequence and real-time semantic tags, extract real-time passenger flow temporal features and real-time semantic features, generate query vectors, sort the vector index library according to similarity from high to low, select a preset number of historical context segments with the highest ranking, each of the historical context segments contains historical passenger flow temporal features, historical semantic features and corresponding original passenger flow data and text data; Step S30: Associate the selected historical context fragments with the corresponding entities in the spatiotemporal knowledge graph, and perform diffusion reasoning based on the causal logic in the graph to obtain logical constraint information; Step S40: Use a time-series prediction model to make a preliminary prediction of the current passenger flow, obtain a preliminary prediction value sequence, construct a prompt message from the preliminary prediction value sequence, the selected historical context fragments and the logical constraint information, input the prompt message into the large language model, and output the optimized prediction value and natural language explanation. Step S50: Generate a capacity adjustment strategy based on the optimized predicted values, and store the feature pair consisting of the real-time passenger flow time-series features and the real-time semantic features into the vector index library.
[0033] To more clearly explain the cross-modal prediction method for rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graphs of this invention, the following will be combined with... Figure 1The steps in the embodiments of the present invention are described in detail below: Step S10: Obtain historical passenger flow data and text data describing environmental conditions, extract passenger flow temporal features and semantic features, and use contrastive learning to map passenger flow temporal features and semantic features to a unified vector space to form a vector index library. In this embodiment, step S10 includes: Step S11: Divide the acquired historical passenger flow data into multiple patches according to the time window, and use a spatiotemporal encoder to extract the passenger flow time sequence features of each patch. Step S13: Use a large language model to perform semantic encoding on the text data of the corresponding time period to obtain semantic features; Step S14: Construct a contrastive loss function based on noise contrastive estimation, using passenger flow temporal features and semantic features under the same context as positive sample pairs and passenger flow temporal features and semantic features under different contexts as negative sample pairs, and calculate the contrastive loss; update the parameters of the spatiotemporal encoder and the large language model by backpropagating the contrastive loss. Step S15: Store the feature pairs consisting of the aligned passenger flow time sequence features and semantic features into the vector index library.
[0034] The contrast loss function is a normalized temperature cross-entropy loss, which is calculated based on the similarity between positive sample pairs and negative sample pairs. Positive sample pairs consist of passenger flow temporal and semantic features under the same context, while negative sample pairs consist of passenger flow temporal and semantic features under different contexts. The parameters of the spatiotemporal encoder and the large language model are updated by backpropagating the contrast loss function.
[0035] In one specific embodiment of the present invention, step S10 is used to construct a vector index library to support subsequent enhanced retrieval prediction, mapping heterogeneous historical passenger flow data and text data to a unified vector space, so that numerical trends and semantic descriptions are comparable in this space, laying the foundation for subsequent similar context retrieval.
[0036] The specific implementation process of step S10 includes steps S11 to S15.
[0037] In step S11, historical passenger flow data is acquired. The data sources include the rail transit automatic fare collection system, the clearing center system, or the network command center database. The data granularity uses entry and exit volume or cross-sectional passenger flow at time intervals of 5 minutes, 15 minutes, or 1 hour. In this embodiment, the historical passenger flow data is divided into multiple patches according to a preset time window, such as by day or by hour. Partial overlap between adjacent patches is allowed to enhance temporal continuity. For each patch, a spatiotemporal encoder is used to extract passenger flow temporal features. This spatiotemporal encoder employs a Transformer architecture based on a self-attention mechanism, or a hybrid architecture combining graph convolutional networks and gated recurrent units. By simultaneously capturing the variation patterns of passenger flow data along the time axis and the spatial correlations along the network topology, it outputs a fixed-dimensional passenger flow temporal feature vector.
[0038] In step S12, text data corresponding to passenger flow data is obtained. The text data comes from weather conditions and forecasts released by meteorological departments, operation adjustment notices released by traffic management departments, public event information released on social media platforms, and service prompts released on station electronic screens. The text data is organized and matched according to the same time window as the passenger flow data to ensure that each text record is strictly aligned with the passenger flow data of the time period it describes.
[0039] In step S13, a large language model is used to semantically encode the text data for the corresponding time period to obtain semantic features. This large language model employs a pre-trained model with contextual understanding capabilities, such as open-source large language models like Llama3, ChatGLM3, or Qwen2, to convert each piece of text data into a high-dimensional, dense semantic feature vector. During the encoding process, specific prompt words can be added to the beginning of the text to guide the large language model to focus on semantic elements related to rail transit passenger flow, such as "road flooding caused by heavy rain" or "crowd evacuation after stadium closing," thereby improving the correlation between semantic features and passenger flow trends.
[0040] After obtaining the temporal and semantic features of passenger flow, step S14 is executed to construct a contrastive loss function based on noise contrast estimation, achieving cross-modal feature alignment. Specifically, the temporal features of passenger flow extracted from passenger flow patches within the same time window are paired with the semantic features extracted from the text data corresponding to that time window to form a positive sample pair; the temporal features and semantic features of passenger flow from different time windows are cross-paired to form negative sample pairs. The contrastive loss function uses normalized temperature cross-entropy loss, which is calculated based on the similarity of positive and negative sample pairs. For a training batch containing N sample pairs, the formula for calculating normalized temperature cross-entropy loss is: ; In the formula, Indicates the comparative loss value; Indicates the number of sample pairs in the training batch; Indicates the first A time-series feature vector of passenger flow; Indicates and Semantic feature vectors belonging to the same context; Indicates the first One semantic feature vector; express and The preferred similarity metric function between them is cosine similarity. The temperature parameter controls the sharpness of the similarity distribution in the feature space. By backpropagating this contrastive loss function, the parameters of both the spatiotemporal encoder and the large language model are updated simultaneously. This allows the spatiotemporal encoder to learn and extract passenger flow temporal features that match the semantic description, while the semantic feature vector generated by the large language model tends to align with the passenger flow temporal feature vector of the corresponding context in terms of direction. After training, passenger flow temporal features and semantic features in the same context become closer in distance within a unified vector space, while feature pairs in different contexts become farther apart, achieving cross-modal feature alignment.
[0041] In step S15, the aligned passenger flow temporal features and semantic features are combined to form feature pairs, which are then stored in a vector index library. The vector index library is constructed based on an approximate nearest neighbor search algorithm, such as a hierarchical navigable small-world graph or an inverted file index structure, to support subsequent fast similarity retrieval in high-dimensional space. During storage, each feature pair is saved as an index record. This record not only contains the passenger flow temporal feature vector and the semantic feature vector, but also associates references or identifiers pointing to the original passenger flow data fragments and text data, so that complete historical context fragments can be retrieved during the retrieval stage.
[0042] In this embodiment, the step of storing the feature pair composed of the real-time passenger flow time-series features and the real-time semantic features into the vector index library further includes: Calculate the similarity between the newly inserted feature pair and existing feature pairs in the vector index library. If there is an existing feature pair with a similarity exceeding a preset threshold, delete the existing feature pair and write the newly inserted feature pair into the vector index library.
[0043] This implementation also covers a dynamic update mechanism for the vector index library. During actual operation, the system continuously acquires new real-time passenger flow data and corresponding real-time semantic tags, extracting real-time passenger flow temporal and semantic features through the same processing flow described above. When it is necessary to store feature pairs of the current context into the vector index library, the similarity between the newly inserted feature pair and existing feature pairs in the vector index library is calculated. The similarity metric can be cosine similarity or Euclidean distance. When an existing feature pair with a similarity exceeding a preset threshold is detected, it indicates that the library already stores historical records highly similar to the current context. At this point, the existing feature pair is deleted, and the newly inserted feature pair is written into the vector index library.
[0044] This replacement strategy controls the growth of the vector index library, avoiding a decrease in retrieval efficiency due to the accumulation of a large number of highly similar records. Simultaneously, it ensures that the feature pairs stored in the library always reflect the latest operational environment characteristics, guaranteeing the timeliness of historical experience. When no existing feature pairs with similarity exceeding the threshold exist, the newly inserted feature pairs are directly written into the vector index library, thereby continuously enriching the context types covered by the vector index library. Preferably, the preset threshold is set to 0.95 or 0.98 to adapt to the deduplication requirements under different operational environments.
[0045] Step S20: Obtain real-time passenger flow fluctuation sequence and real-time semantic tags, extract real-time passenger flow temporal features and real-time semantic features, generate query vectors, sort the vector index library according to similarity from high to low, select a preset number of historical context segments with the highest ranking, each of the historical context segments contains historical passenger flow temporal features, historical semantic features and corresponding original passenger flow data and text data; In this embodiment, real-time passenger flow fluctuation sequences and real-time semantic tags are obtained, real-time passenger flow temporal features and real-time semantic features are extracted, and a query vector is generated. The vector index is then sorted from high to low similarity, and a predetermined number of historical context segments are selected from the top-ranked segments, including: Step S21: Input the real-time passenger flow fluctuation sequence into the spatiotemporal encoder to obtain real-time passenger flow time-series features; input the real-time semantic label into the large language model to obtain real-time semantic features; and concatenate or weightedly fuse the real-time passenger flow time-series features and real-time semantic features to generate a query vector. Step S22: Perform an approximate nearest neighbor search in the vector index library to return a preset number of historical context fragments that rank highly similar to the query vector. Each historical context fragment contains aligned historical passenger flow time-series features and historical semantic features, as well as their corresponding original passenger flow data and text data.
[0046] In one specific embodiment of the present invention, step S20 is used to retrieve the historical context fragment most similar to the current real-time context from the constructed vector index library, providing an empirical reference for subsequent prediction correction and interpretable reasoning. The specific implementation process of step S20 includes steps S21 to S22.
[0047] In step S21, the system acquires the real-time passenger flow fluctuation sequence and real-time semantic tags for the current moment. The real-time passenger flow fluctuation sequence is collected from the same source as historical data, which can be obtained from the rail transit automatic fare collection system or the real-time data interface of the network command center. The data format is a passenger flow time series with the current moment as the endpoint and a preset time length back, such as taking 5-minute granular entry and exit data from 4 hours or 6 hours ago, forming a fluctuation sequence reflecting the recent passenger flow trend. The real-time semantic tags are textual descriptive information related to the current time period, which comes from weather warnings pushed in real time by the meteorological early warning system, real-time operation announcements issued by the transportation management department, emergency event messages instantly disseminated on social media platforms, or operational anomaly records reported on-site by the station. For example, the real-time semantic tags can be text content such as "Airport Line initiates emergency response due to large-scale flight delays" or "Temporary traffic control implemented on roads around the Sports Center Station due to flooding caused by heavy rain."
[0048] After obtaining the real-time passenger flow fluctuation sequence, it is input into the spatiotemporal encoder used in step S11. This spatiotemporal encoder has been aligned with the semantic feature space through contrastive learning training in step S14. Forward inference is then performed on the real-time passenger flow fluctuation sequence to extract the real-time passenger flow temporal features. These real-time passenger flow temporal features are fixed-dimensional vectors, with dimensions consistent with the historical passenger flow temporal feature vectors stored in the vector index library. Simultaneously, real-time semantic labels are input into the large language model used in step S13. This large language model has also undergone contrastive learning training to semantically encode the real-time semantic labels, obtaining real-time semantic features. The large language model selected is one of the open-source large language models such as Llama 3, ChatGLM 3, or Qwen 2, and its output semantic feature vector dimension is consistent with the historical semantic feature vector dimension stored in the vector index library.
[0049] After obtaining real-time passenger flow time-series features and real-time semantic features, the two are fused to generate a query vector. The fusion method can be feature concatenation or weighted fusion. When using feature concatenation, the real-time passenger flow time-series feature vector and the real-time semantic feature vector are concatenated end-to-end along their dimensional direction to form a query vector with a length equal to the sum of their dimensions. When using weighted fusion, weight coefficients are assigned to the real-time passenger flow time-series features and the real-time semantic features respectively. For example, the weight of the real-time passenger flow time-series features is set to 0.6, and the weight of the real-time semantic features is set to 0.4. The two are then summed according to their weights to generate a query vector with the same dimensions as the original features. Alternatively, a learnable linear transformation matrix can be used to map the two to a unified query space. In this embodiment, feature concatenation is preferred to maximize the preservation of complete information from both types of features and improve the discriminative power of subsequent retrieval.
[0050] In step S22, the query vector generated in step S21 is used as the retrieval input to perform an approximate nearest neighbor search in the vector index library. The vector index library is constructed using data structures that support fast retrieval of high-dimensional vectors, such as hierarchical navigable small-world graphs or inverted file indexes. The search process calculates the similarity between the query vector and the historical context feature pairs stored in each index record of the vector index library, using cosine similarity as the similarity measure. After the search is completed, the historical context fragments are sorted from high to low similarity and returned to the top-ranked preset number. The preset number is 3 or 5 to accommodate the balance between the diversity and relevance of reference cases required by the prediction task. Each returned historical context fragment contains aligned historical passenger flow time-series feature vectors and historical semantic feature vectors, as well as associated original passenger flow data fragments and text data. The original passenger flow data fragments are the actual passenger flow time series under the historical context, and the text data are complete text records such as weather descriptions and event announcements corresponding to the historical context. This complete historical information provides factual basis for subsequent reasoning and interpretation by the large language model.
[0051] Step S30: Associate the selected historical context fragments with the corresponding entities in the spatiotemporal knowledge graph, and perform diffusion reasoning based on the causal logic in the graph to obtain logical constraint information; In this embodiment, step S30 includes: Step S31: Based on the event type and the station involved identified in the historical context fragment, locate the trigger node in the spatiotemporal knowledge graph and activate the corresponding station node and event node; Step S32: Starting from the trigger node, perform multi-hop diffusion along the predefined operation rule edges and event causal propagation edges in the spatiotemporal knowledge graph. In each hop, the influence state of the current node is transmitted to the adjacent node along the edge. The adjacent node aggregates the received influence state and updates its own influence state. Iteratively calculate the passenger flow influence weight of the adjacent station and generate logical constraint information containing the affected area and the distribution of passenger flow influence degree.
[0052] The spatiotemporal knowledge graph is constructed through the following process: Station nodes, line edges, event nodes, and their attributes are obtained from rail transit operation records and event notices using named entity recognition and relation extraction. Define event causal propagation edges and spatial adjacency edges according to domain rules, assign propagation weights to the event causal propagation edges and spatial adjacency edges respectively, and form a directed graph structure; In the multi-hop diffusion, the graph attention mechanism calculates the attention coefficient of the source node to the neighbor node at each hop. The neighbor node aggregates the influence from different source nodes according to the attention coefficient and updates the state of the neighbor node. After multiple iterations, the degree of passenger flow influence of each node is output.
[0053] In one specific embodiment of the present invention, step S30 is used to associate the retrieved historical context fragments with domain structured knowledge, and to perform reasoning based on road network topology and event propagation mechanisms to generate logical constraint information for constraining the prediction model. The specific implementation process of step S30 includes steps S31 and S32, and the spatiotemporal knowledge graph on which it depends needs to be pre-constructed.
[0054] The construction of the spatiotemporal knowledge graph is independent of the online prediction process and is completed offline. Corpus data is collected from rail transit operation records and event announcements. Operation records include historical dispatch logs, station passenger flow control records, and temporary line adjustment notices. Event announcements include weather warnings, large-scale event reports, and emergency briefings. Named entity recognition technology is used to extract station entities, line entities, and event entities from the corpus. Station entities include attributes such as station name, line affiliation, and transfer attributes; line entities include line name, direction (up or down), and sequence of stations passed through; and event entities include event type, occurrence time, and scope of impact. Simultaneously, relation extraction technology is used to extract relationships between entities from the text, including the attribution relationship between stations and lines, the spatial adjacency relationship between adjacent stations, and the influence relationship between events and stations. After extraction, station nodes and event nodes are established in the knowledge graph. Station nodes are connected by line edges and spatial adjacency edges, and event nodes are connected to affected station nodes by event influence edges.
[0055] Based on domain rules, event causal propagation edges and spatial adjacency edges are defined. Event causal propagation edges describe how an event occurring at one station affects other stations through passenger flow propagation paths in the network. For example, the event "a surge in passenger flow at Dongzhimen Station on the Airport Line due to flight delays" can be connected to downstream transfer stations such as "Sanyuanqiao Station" and "Dongsi Shitiao Station" via causal propagation edges. Spatial adjacency edges describe the physical adjacency between stations; they are directed edges, with the direction consistent with the train's running direction or the main passenger flow direction. Transmission weights are assigned to both event causal propagation edges and spatial adjacency edges, representing the attenuation or amplification coefficient of the influence propagating along that edge. The weights of event causal propagation edges are statistically obtained based on the proportion of actual passenger flow changes at each adjacent station in historical similar events, while the weights of spatial adjacency edges are comprehensively set based on station spacing, travel time, and transfer coefficients. Taking the spatial adjacency edge between Sanyuanqiao Station and Dongzhimen Station as an example, its propagation weight can be set to 0.8, indicating that 80% of the passenger flow impact from Dongzhimen Station will be propagated to Sanyuanqiao Station. Meanwhile, the causal propagation edge weight of the "rainstorm causing flooding" event on Dongzhimen Station of the Airport Line can be set to 1.5, indicating that the impact of this event on passenger flow at Dongzhimen Station will be amplified by 1.5 times. The resulting directed graph structure provides a topological foundation for subsequent multi-hop diffusion inference.
[0056] During step S31, for each historical context fragment retrieved in step S20, the event type and involved stations are identified from the event description text of that fragment. The identification process employs an information extraction method combined with a large language model. For example, the description text "Heavy rain on a certain day caused flooding at the entrance of a certain rail transit station" is input into ChatGLM 3. The large language model outputs a structured event type field, such as "heavy rain and flooding," and a list of involved stations, such as "a certain station." In the spatiotemporal knowledge graph, event nodes matching the event type are located, and station nodes matching the station are also located. These nodes are used as trigger nodes for diffusion inference, activating the corresponding station nodes and event nodes. The activated state is recorded as the initial influence state.
[0057] Step S32 starts from the trigger node and performs multi-hop diffusion reasoning along predefined operational rule edges and event causal propagation edges. In each hop, the current node transmits its influence state to neighboring nodes along all outgoing edges. Neighboring nodes aggregate the influence states from each source node and update their own influence state. The aggregation and update process introduces a graph attention mechanism to adaptively learn the relative importance of the influence of different neighboring nodes on the current node. In each hop iteration, for node i and one of its neighboring nodes j, the attention coefficient of node j on node i is calculated. The calculation method is to convert the influence state vector of node i. The influence state vector of node j After concatenation, a learnable weight vector is used. A linear transformation is performed, followed by the LeakyReLU activation function, and finally, the normalized attention coefficients are obtained by normalizing them over all neighboring nodes using the softmax function. The formula is expressed as: ; in, and These represent the current influence state vectors of node i and node j, respectively. A shared linear transformation matrix is used for feature mapping of node states; This represents a vector concatenation operation; This is a learnable attention weight vector. Transpose it; Represents the set of all neighboring nodes of node i; The function is a linear rectified activation function with leakage. After calculating the attention coefficients, the influence states of neighboring node j are aggregated to node i according to the product of the attention coefficients and the propagation weights. The updated influence states of node i are then... The weighted summation of the influence states passed from all neighboring nodes is given by the following formula: ; in, The transit weight assigned to the edge from node j to node i. The activation function is non-linear, using ReLU or sigmoid. After one round of updates, the states of all affected nodes are refreshed, completing a one-hop diffusion. Then, using the updated node states as the input for the next hop, the above process is repeated. After a preset T-hop iteration, such as T=2 or T=3, the diffusion process covers the adjacent stations of the triggering node and even more distant stations. After iteration, the passenger flow impact weight is calculated based on the final impact state of each station node. The impact weight can be directly taken as the magnitude of the impact state vector or the value of a certain dimension. The passenger flow impact weight of each station is associated with the station nodes to generate logical constraint information. The logical constraint information includes a list of affected areas and a distribution of passenger flow impact levels, for example, "Dongzhimen Station impact level 0.9, Sanyuanqiao Station impact level 0.7, Dongsi Shitiao Station impact level 0.5". This logical constraint information serves as an explicit domain knowledge constraint, used in subsequent step S40 to guide the large language model in correcting the initial prediction values.
[0058] Step S40: Use a time-series prediction model to make a preliminary prediction of the current passenger flow, obtain a preliminary prediction value sequence, construct a prompt message from the preliminary prediction value sequence, the selected historical context fragments and the logical constraint information, input the prompt message into the large language model, and output the optimized prediction value and natural language explanation. In this embodiment, step S40 includes: Step S41: Use a patch-based time-series prediction model to predict the current passenger flow and output a preliminary prediction sequence for the future preset time period. Step S42: The preliminary predicted value sequence is segmented and aggregated to approximate the data, and converted into a text description sequence containing trends and values. The text description sequence, the historical passenger flow actual value text and event description text in the historical context fragment, and the textual representation of the logical constraint information are assembled into an enhanced prompt word according to the hierarchical structure of instruction layer, example layer, and constraint layer. The enhanced prompt word contains instructions from the constraint language model to correct the preliminary predicted value sequence based on the historical context fragment and the logical constraint information. Step S43: Input the enhanced prompt words into the large language model, and output the corrected prediction value and the historical case identifier and correction reason based on the corresponding correction point.
[0059] The enhanced prompt words also include a set of trigger instructions. These trigger instructions are used to enable the large language model to extract historical case identifiers from the historical context fragments when generating corrected prediction values, and to generate correction criteria based on the passenger flow evolution patterns of the historical cases and the logical constraint information. The historical case identifiers are also output in the output natural language explanation.
[0060] In one specific embodiment of the present invention, step S40 is used to fuse the preliminary prediction results driven by data, the historical experience enhanced by retrieval, and the logical constraints derived from the knowledge graph, and generate optimized prediction values and interpretable correction basis through the reasoning ability of the large language model. The specific implementation process of step S40 includes steps S41 to S43.
[0061] In step S41, a patch-based temporal prediction model is used to make a preliminary prediction of the current passenger flow, outputting a preliminary prediction value sequence for a preset future time period. This temporal prediction model can be either the PatchTST model or the iTransformer model. The PatchTST model divides the input time series into multiple fixed-length, overlapping patches. Each patch serves as an input token for the Transformer encoder, capturing temporal dependencies within and between patches through a self-attention mechanism. In this embodiment, the input current passenger flow sequence is 5-minute granularity entry and exit volume data, tracing back 192 time steps from the current time t. The patch length is set to 16 time steps, the step size to 8 time steps, the encoder has 3 layers, and the number of attention heads is 16. The output is a preliminary prediction value sequence for the next 96 time steps, or 8 hours. The iTransformer model inverts the attention mechanism of the traditional Transformer, treating passenger flow variables from multiple stations at the same time step as tokens. It aggregates cross-channel information through a feedforward network and captures cross-temporal dependencies through self-attention. In this embodiment, the PatchTST model is selected as the preferred preliminary prediction model, and the prediction output is denoted as... Where H represents the prediction step size, in this embodiment H=96, and each element is the predicted passenger flow for the corresponding time step, i.e. y t+H This represents the predicted passenger flow at the H-th time step after the current time t.
[0062] In step S42, the preliminary predicted value sequence undergoes a segmented aggregation approximation transformation, and enhanced prompt words are constructed. First, the preliminary predicted value sequence... A segmented aggregation approximation method is used, dividing the continuous prediction time steps into several equal-length time intervals. Within each time interval, the average value and trend of the predicted values are calculated. For example, 96 time steps are divided into 8 time intervals, each containing 12 time steps (1 hour). The average passenger flow and slope of change are calculated for each time interval, and the trend is determined to be upward, downward, or stable. The resulting text description sequence is described using natural language, for example, "The average passenger flow in the first hour is 12,000, an increase of 15% compared to the current time interval; the average passenger flow in the second hour is 13,500, continuing to rise; the peak of 14,500 passengers is reached in the third hour, and then gradually declines." This text description sequence is used as the prediction trend information to be corrected.
[0063] The enhanced prompts are assembled according to a hierarchical structure of instruction layer, example layer, and constraint layer. The instruction layer contains task descriptions and output format requirements, with an example content of: "You are a rail transit operation and scheduling expert. Based on the provided historical similar scenario cases and network diffusion impact analysis, please revise the preliminary predicted passenger flow sequence point by point, and output the revised predicted value and the reason for the revision. When revising, you must explicitly cite the historical case number in the reason and explain how the passenger flow evolution law applies to the current scenario." The example layer embeds historical scenario fragments retrieved in step S20. Each historical scenario fragment contains the identifier of the historical scenario, date, original event description, and changes in the actual historical passenger flow value. The constraint layer embeds the logical constraint information generated in step S30, which is textually represented as a list of affected stations and their degree of passenger flow impact, such as "Dongzhimen Station impact degree 0.9, Sanyuanqiao Station impact degree 0.7, Dongsi Shitiao Station impact degree 0.5." In addition, the enhanced prompts also include a set of trigger instructions, which are embedded using formatted tags, such as "[REQUIRE_CITATION]" and "[OUTPUT_EXPLANATION]". These instructions instruct the large language model to extract historical case identifiers from historical context fragments in the example layer when generating corrected predictions, and to generate correction criteria based on the passenger flow evolution patterns of the case and the logical constraint information in the constraint layer. The referenced historical case identifiers are also output in the natural language explanation.
[0064] An exemplary enhanced prompt structure is as follows: The instruction layer begins with "Task Instruction:", followed by the task description above; the example layer begins with "Historical Similar Context Cases:", listing retrieved historical context fragments one by one, such as "[Case Example] A heavy rain caused flooding at a rail transit station, and the station's passenger flow decreased significantly in a short period after the event, for example, a decrease of 80% within 2 hours (this value is only an example); adjacent stations also experienced a decrease due to reduced transfer passenger flow, with an example decrease of 30%; the constraint layer begins with "Road Network Diffusion Impact Constraints:", listing the textual representation of logical constraint information; finally, the trigger instruction "[REQUIRE_CITATION] Please cite the specific case number in the correction reason". The entire enhanced prompt serves as the input text for the large language model.
[0065] In step S43, the enhanced prompt words assembled in step S42 are input into the large language model. The large language model is one of the open-source large language models such as Llama 3, ChatGLM 3, or Qwen 2, and generates output in an autoregressive manner. The output contains two parts: the corrected prediction value sequence and the natural language explanation. The corrected prediction values are given in the form of a structured list of numbers, such as "[14000, 14800, 15200, ...]", corresponding to the corrected passenger flow prediction values for each future time period. In the natural language explanation, the large language model forcibly references historical case identifiers according to trigger instructions, such as "Referring to a certain rainstorm case, the decrease in the number of passengers entering the station due to water accumulation at a certain station will be transmitted to the adjacent station through the transfer node. The prediction value of the adjacent station in the second hour will be reduced by 20%. This value is only an example." Each correction point in the output indicates the historical case identifier and the reason for the correction, thereby realizing the transparency and traceability of the prediction process. Finally, the corrected prediction value output in step S43 is used as the optimized prediction value and passed to step S50 for scheduling decisions. Simultaneously, the corresponding natural language explanation is presented to the operations scheduling personnel. In step S50, a capacity adjustment strategy is generated based on the optimized prediction value, and the feature pair consisting of the real-time passenger flow time-series features and the real-time semantic features is stored in the vector index library.
[0066] In this embodiment, step S50 includes: Step S51: Compare the optimized predicted value with the preset scheduling trigger threshold. When the predicted value exceeds the threshold, generate a capacity adjustment strategy that includes the departure interval compression ratio or the number of spare vehicles based on the excess magnitude and time period. Step S52: Extract the real-time passenger flow time-series features and real-time semantic features used when generating the query vector, form feature pairs, and insert them into the vector index library.
[0067] In one specific embodiment of the present invention, step S50 is used to transform the optimized prediction results into an executable operation scheduling strategy and complete the self-guided update of the vector index library, thereby realizing the system's continuous learning and dynamic adaptation capabilities. The specific implementation process of step S50 includes steps S51 and S52.
[0068] In step S51, the system compares the optimized predicted value output in step S43 with the preset scheduling trigger threshold to determine whether the current capacity configuration can meet the upcoming passenger flow demand. The preset scheduling trigger threshold is pre-set based on the line's designed capacity and historical operating experience, including an upper limit trigger threshold and a lower limit trigger threshold. The upper limit trigger threshold is used to detect the risk of passenger flow oversaturation; when the predicted passenger flow exceeds the line's or station's designed capacity, capacity replenishment measures are triggered. The lower limit trigger threshold is used to detect capacity redundancy; when the predicted passenger flow is significantly lower than the capacity corresponding to the current departure interval, capacity reduction measures are triggered to save operating resources. In this embodiment, the upper limit trigger threshold is set between 85% and 95% of the line's designed capacity, preferably 90%; the lower limit trigger threshold is set between 40% and 60% of the line's designed capacity, preferably 50%. The line's designed capacity is determined by the train's passenger capacity, the current departure interval, and the number of train formations. The train's passenger capacity can be obtained from the vehicle's technical specifications, and the current departure interval can be obtained from the operating timetable.
[0069] When the optimized forecast exceeds the upper limit trigger threshold, the system generates a capacity replenishment strategy based on the magnitude and duration of the exceedance. The magnitude of the exceedance is defined as the percentage by which the predicted passenger flow exceeds the threshold, categorized into three levels: mild exceedance (0% to 15%), moderate exceedance (15% to 30%), and severe exceedance (over 30%). The duration is defined as the length of time during which the predicted passenger flow continuously exceeds the threshold. For mild exceedances with a duration not exceeding 2 hours, the generated capacity adjustment strategy is to reduce the departure interval, with the reduction ratio linearly determined by the magnitude of the exceedance. For example, a 10% exceedance results in a 10% reduction in the departure interval, shortening it from the current 6 minutes to 5.4 minutes. For moderate exceedances with a duration between 2 and 4 hours, the generated capacity adjustment strategy is to reduce the departure interval in conjunction with flow control measures at some stations, with the reduction ratio directly proportional to the magnitude of the exceedance. For cases where the excess passenger flow exceeds 4 hours or the duration exceeds 4 hours, the generated capacity adjustment strategy is to increase the number of spare train formations on the basis of reducing the departure interval. The spare trains are transferred from the depot or storage line and put into operation on the main line. The number of formations is determined by rounding up the ratio of the excess passenger flow to the passenger capacity of a single train.
[0070] When the optimized predicted value falls below the lower limit trigger threshold and the duration exceeds the preset duration, the system generates a capacity reduction strategy, including appropriately lengthening the departure interval or reducing the number of train formations online, in order to reduce traction energy consumption and vehicle wear while ensuring service levels. In this embodiment, the preset duration is set to 2 hours. If the duration is less than this, the current capacity configuration remains unchanged to avoid frequent adjustments to the timetable due to short-term passenger flow fluctuations.
[0071] Step S52 is executed in parallel with or sequentially with step S51. The system extracts the real-time passenger flow time-series features and real-time semantic features used in step S21 to generate the query vector, and combines these two types of features into a feature pair. The structure of this feature pair is consistent with the historical feature pairs stored in the vector index library in step S15, including passenger flow time-series feature vectors and semantic feature vectors. When inserting this feature pair into the vector index library, a deduplication operation is performed according to the dynamic update mechanism of the vector index library. The cosine similarity between the newly inserted feature pair and the existing feature pairs in the vector index library is calculated. When an existing feature pair with a similarity exceeding a preset threshold of 0.95 is detected, the existing feature pair is deleted and the new feature pair is written into the vector index library; when no existing feature pair with a similarity exceeding the threshold is found, the new feature pair is directly written into the vector index library. Through this mechanism, the relationship between passenger flow features and semantic features in the current operating context is recorded as new historical experience, which can be retrieved and called by step S20 in future prediction cycles, forming a complete closed loop from perception, prediction, decision-making to experience accumulation, realizing the continuous expansion and gradual optimization of the vector index library as the operating environment changes.
[0072] See Figure 2 A second embodiment of the present invention provides a rail transit cross-modal prediction system based on retrieval-enhanced reasoning and spatiotemporal knowledge graphs, used to execute the method of the first embodiment. The system includes: The vector index library construction module is used to acquire historical passenger flow data and text data describing environmental conditions, extract passenger flow temporal features and semantic features, and use contrastive learning to map passenger flow temporal features and semantic features to a unified vector space to form a vector index library; The retrieval module is used to obtain real-time passenger flow fluctuation sequences and real-time semantic tags, extract real-time passenger flow temporal features and real-time semantic features, generate query vectors, sort the vector index library according to similarity from high to low, select a preset number of historical context segments with the highest ranking, and each historical context segment contains historical passenger flow temporal features, historical semantic features and corresponding original passenger flow data and text data. The diffusion reasoning module is used to associate selected historical context fragments with corresponding entities in the spatiotemporal knowledge graph, and perform diffusion reasoning based on the causal logic in the graph to obtain logical constraint information; The prediction optimization module is used to make a preliminary prediction of the current passenger flow using a time series prediction model, obtain a preliminary prediction value sequence, construct a prompt message from the preliminary prediction value sequence, the selected historical context segment and the logical constraint information, input the prompt message into the large language model, and output the optimized prediction value and natural language explanation. The scheduling suggestion and update module is used to generate a capacity adjustment strategy based on the optimized prediction value, and store the feature pair consisting of the real-time passenger flow time-series feature and the real-time semantic feature into the vector index library.
[0073] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process and related explanations of the method described above can be found in the corresponding process in the foregoing system embodiments, and will not be repeated here.
[0074] It should be noted that the above embodiment of the rail transit cross-modal prediction system based on retrieval-enhanced reasoning and spatiotemporal knowledge graph is only an example illustrating the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be merged into one module, or further divided into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing the various modules or steps and are not considered as an improper limitation of the present invention.
[0075] A device according to a third embodiment of the present invention includes: At least one processor; and a memory communicatively connected to at least one of the processors; The memory stores instructions that can be executed by the processor to implement the above-described method for cross-modal prediction of rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graph.
[0076] A computer-readable storage medium according to a fourth embodiment of the present invention stores computer instructions, which are executed by the computer to implement the above-described method for cross-modal prediction of rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graph.
[0077] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process and related descriptions of the storage device and processing device described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0078] The following is for reference. Figure 3 It shows a schematic diagram of the structure of a computer system for implementing embodiments of the systems, methods, and electronic devices of this application. Figure 3 The server shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0079] like Figure 3As shown, the computer system includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes based on programs stored in Read Only Memory (ROM) 302 or programs loaded from storage section 308 into Random Access Memory (RAM) 303. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.
[0080] The following components are connected to I / O interface 305: an input section 306 including a keyboard, mouse, etc.; an output section 307 including a cathode ray tube (CRT), liquid crystal display (LCD), and speakers, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card and a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.
[0081] Specifically, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the functions defined in the methods of this application. It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0082] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0083] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0084] The terms “first”, “second”, etc., are used to distinguish similar objects, not to describe or indicate a specific order or sequence.
[0085] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent in such process, method, article, or apparatus / device.
[0086] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A method for cross-modal prediction of rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graphs, characterized in that, include: Historical passenger flow data and text data describing environmental conditions are obtained, and passenger flow temporal and semantic features are extracted. Contrastive learning is used to map the passenger flow temporal and semantic features to a unified vector space to form a vector index library. Obtain real-time passenger flow fluctuation sequence and real-time semantic tags, extract real-time passenger flow temporal features and real-time semantic features, generate query vectors, sort the vector index library according to similarity from high to low, select a preset number of historical context segments with the highest ranking, each of the historical context segments contains historical passenger flow temporal features, historical semantic features and corresponding original passenger flow data and text data; The selected historical context fragments are associated with corresponding entities in the spatiotemporal knowledge graph. Based on the causal logic in the graph, diffusion reasoning is performed to obtain logical constraint information. A time-series prediction model is used to make a preliminary prediction of the current passenger flow, obtaining a preliminary predicted value sequence. This preliminary predicted value sequence, selected historical context segments, and the logical constraint information are then used to construct a prompt message, which is input into a large language model. The model outputs optimized predicted values and a natural language explanation. The current passenger flow is predicted using a patch-based time-series forecasting model, and the preliminary forecast value sequence for the future preset time period is output. The preliminary predicted value sequence is segmented and aggregated to approximate it, and then converted into a text description sequence containing trends and values. The text description sequence, the historical passenger flow actual value text and event description text in the historical context fragment, and the textual representation of the logical constraint information are assembled into an enhanced prompt word according to the hierarchical structure of instruction layer, example layer, and constraint layer. The enhanced prompt word contains instructions from the constraint language model to correct the preliminary predicted value sequence based on the historical context fragment and the logical constraint information. The enhanced prompt words are input into the large language model, and the corrected prediction value, the historical case identifiers on which the corresponding correction points are based, and the reasons for the correction are output. The enhanced prompt words also include a set of trigger instructions. The trigger instructions are used to enable the large language model to extract historical case identifiers from the historical context fragments when generating corrected prediction values, and to generate correction basis based on the passenger flow evolution pattern of the historical cases and the logical constraint information, and to output the historical case identifiers in the output natural language explanation. Based on the optimized predicted values, a capacity adjustment strategy is generated, and the feature pairs consisting of the real-time passenger flow time-series features and the real-time semantic features are stored in the vector index library.
2. The method for cross-modal prediction of rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graph as described in claim 1, characterized in that, Historical passenger flow data and textual data describing environmental conditions are acquired. Temporal and semantic features of passenger flow are extracted. Contrastive learning is used to map these features to a unified vector space, forming a vector index library, including: The acquired historical passenger flow data is divided into multiple patches according to time windows, and the temporal characteristics of passenger flow in each patch are extracted using a spatiotemporal encoder. We use a large language model to semantically encode text data for the corresponding time period to obtain semantic features; A contrastive loss function is constructed based on noise contrastive estimation. Passenger flow temporal features and semantic features under the same context are used as positive sample pairs, and passenger flow temporal features and semantic features under different contexts are used as negative sample pairs to calculate the contrastive loss. The parameters of the spatiotemporal encoder and the large language model are updated by backpropagating the contrastive loss. The aligned passenger flow time sequence features and semantic features are combined to form feature pairs, which are then stored in the vector index library.
3. The method for cross-modal prediction of rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graphs according to claim 2, characterized in that, Obtain real-time passenger flow fluctuation sequences and real-time semantic tags, extract real-time passenger flow temporal features and real-time semantic features, generate query vectors, sort the vector indexes from high to low similarity, and select a predetermined number of historical context segments from the top of the sorted list, including: The real-time passenger flow fluctuation sequence is input into the spatiotemporal encoder to obtain real-time passenger flow time-series features. The real-time semantic label is input into the large language model to obtain real-time semantic features. The real-time passenger flow time-series features and real-time semantic features are concatenated or weighted and fused to generate a query vector. An approximate nearest neighbor search is performed in the vector index library, returning a preset number of historical context fragments that rank highly similar to the query vector. Each historical context fragment contains aligned historical passenger flow time-series features and historical semantic features, as well as their corresponding original passenger flow data and text data.
4. The method for cross-modal prediction of rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graph as described in claim 1, characterized in that, The selected historical context fragments are associated with corresponding entities in a spatiotemporal knowledge graph. Based on the causal logic in the graph, diffusion reasoning is performed to obtain logical constraint information, including: Based on the event types and stations identified in the historical context fragments, trigger nodes are located in the spatiotemporal knowledge graph, and the corresponding station nodes and event nodes are activated. Starting from the trigger node, multi-hop diffusion is performed along the predefined operation rule edges and event causal propagation edges in the spatiotemporal knowledge graph. In each hop, the influence state of the current node is passed to the adjacent node along the edge. The adjacent node aggregates the received influence state and updates its own influence state. The passenger flow influence weight of the adjacent station is iteratively calculated to generate logical constraint information containing the affected area and the distribution of passenger flow influence.
5. The method for cross-modal prediction of rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graphs according to claim 1, characterized in that, Based on the optimized predicted values, a capacity adjustment strategy is generated, and the feature pairs consisting of the real-time passenger flow time-series features and the real-time semantic features are stored in the vector index library, including: The optimized predicted value is compared with the preset scheduling trigger threshold. When the predicted value exceeds the threshold, a capacity adjustment strategy including the departure interval compression ratio or the number of standby vehicle formations is generated based on the magnitude of the exceedance and the time period. Extract the real-time passenger flow time-series features and real-time semantic features used when generating the query vector, form feature pairs, and insert them into the vector index library.
6. The method for cross-modal prediction of rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graphs according to claim 2, characterized in that, The contrastive loss function is the normalized temperature cross-entropy loss, which is calculated based on the similarity of positive sample pairs and negative sample pairs. Positive sample pairs consist of passenger flow temporal and semantic features under the same context, while negative sample pairs consist of passenger flow temporal and semantic features under different contexts. The parameters of the spatiotemporal encoder and the large language model are updated by backpropagating the contrastive loss function.
7. The method for cross-modal prediction of rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graphs according to claim 4, characterized in that, The spatiotemporal knowledge graph is constructed through the following process: Station nodes, line edges, event nodes, and their attributes are obtained from rail transit operation records and event notices using named entity recognition and relation extraction. Define event causal propagation edges and spatial adjacency edges according to domain rules, assign propagation weights to the event causal propagation edges and spatial adjacency edges respectively, and form a directed graph structure; In the multi-hop diffusion, the graph attention mechanism calculates the attention coefficient of the source node to the neighbor node at each hop. The neighbor node aggregates the influence from different source nodes according to the attention coefficient and updates the state of the neighbor node. After multiple iterations, the degree of passenger flow influence of each node is output.
8. The method for cross-modal prediction of rail transit based on retrieval-enhanced reasoning and spatiotemporal knowledge graphs according to claim 1, characterized in that, The step of storing the feature pair composed of the real-time passenger flow time-series features and the real-time semantic features into the vector index library further includes: Calculate the similarity between the newly inserted feature pair and existing feature pairs in the vector index library. If there is an existing feature pair with a similarity exceeding a preset threshold, delete the existing feature pair and write the newly inserted feature pair into the vector index library.
Citation Information
Patent Citations
Rail transit passenger flow prediction method and system based on traffic interpretable large model
CN122155011A