Subway station multi-modal data fusion method, computer device and program product
Patent Information
- Application Number
- CN202610802100.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-06-04
AI Technical Summary
[0005]但是,对于所接入的多个数据来源设备而言,由于不同设备在数据采集频率、数据生成机制以及数据表达形式等方面均存在差异,导致不同模态数据之间往往难以实现精准对齐,从而并无法反映地铁车站的运行状态变化
[0011]本申请实施例对地铁车站内来自多种设备的数据,通过执行模态内事件检测、跨模态的事件关联计算,以及跨模态事件关联所得车站事件实体为中心的站点运行状态描述数据结构化聚合处理,得到能够准确反映地铁车站真实运行状态的结构化事件表示数据,在所得事件表示数据的作用下,能够全面准确且真实地刻画车站运行状态,提升车站异常事件识别能力和运营管理效率,为车站运行分析、异常监测和运营决策提供可靠的数据基础。
Smart Images

Figure CN122332993B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent rail transit technology, specifically to a multimodal data fusion method for subway stations, computer equipment, and computer program products. Background Technology
[0002] As urban rail transit networks continue to expand, the operating environment of subway stations is becoming increasingly complex. To ensure station safety and improve operational efficiency, various types of equipment are typically deployed inside stations, such as video surveillance equipment, passenger flow detection equipment, turnstiles, environmental monitoring equipment, and equipment operation status monitoring systems. These devices collectively support the daily operation of subway stations, and the data generated by each device allows for real-time perception of the station's operational status from different perspectives.
[0003] In practical applications, the data generated by these devices exhibits significant multimodal characteristics, meaning that data from different sources can describe changes in the operational status of subway stations from different dimensions. For example, video surveillance equipment can reflect pedestrian behavior and passenger flow distribution within the station area, turnstiles can reflect changes in passenger flow entering and exiting the station, passenger flow detection equipment can reflect changes in regional passenger density, and equipment operation status monitoring systems can reflect the operational status of station equipment. Therefore, theoretically, this multimodal data can provide a more comprehensive reflection of the actual operational status of subway stations.
[0004] However, given the vast amount of diverse data from various devices in subway stations, analysis is typically conducted independently based on a single data source or a limited number of data source devices, or data from multiple devices is simply overlaid. Even when multiple data source devices are accessed, the resulting multimodal data is often only simply aligned along the time dimension to provide a rough characterization of the subway station's operational status.
[0005] However, for the multiple data source devices that are connected, due to the differences in data acquisition frequency, data generation mechanism and data expression form, it is often difficult to achieve accurate alignment between different modal data, thus failing to reflect the changes in the operating status of the subway station. Summary of the Invention
[0006] One objective of this application is to achieve the fusion processing of multimodal data generated by subway stations, so that the resulting data can accurately reflect changes in the station's operational status.
[0007] According to one aspect of the embodiments of this application, a multimodal data fusion method for subway stations is disclosed, the method comprising:
[0008] By pulling real-time multimodal data from subway stations, we continuously acquire multimodal data of the actual data environment of the stations. The multimodal data includes station operation status description data of at least one modality, and the modality corresponds to different types of equipment in the station. Intramodal event detection is performed on the multimodal data, and corresponding candidate events are obtained from the site operation status description data of each modality. The candidate events are the changes and / or abnormal patterns of site operation status represented by the site operation status description data of the corresponding modality. For candidate events corresponding to the operational status description data of each modal station, cross-modal event association calculation is performed based on the data of each modal, and candidate events belonging to the same real station event are merged to obtain the corresponding station event entity; Using the station event entity as the fusion center, the station operation status description data corresponding to each candidate event mapped to the station event entity is subjected to structured aggregation processing to form the event representation data of the station event entity.
[0009] According to one aspect of the embodiments of this application, a computer device is disclosed, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method as described in any of the preceding claims.
[0010] According to one aspect of the embodiments of this application, a computer program product is disclosed, including a computer program that, when executed by a processor, implements the steps of the method as described in any of the preceding claims.
[0011] This application embodiment processes data from various devices within a subway station by performing intramodal event detection, cross-modal event correlation calculation, and structured aggregation of station operation status description data centered on station event entities obtained from cross-modal event correlation. This results in structured event representation data that accurately reflects the true operation status of the subway station. Under the influence of the obtained event representation data, the station operation status can be comprehensively, accurately, and realistically depicted, improving the station's ability to identify abnormal events and operational management efficiency, and providing a reliable data foundation for station operation analysis, anomaly monitoring, and operational decision-making.
[0012] Through the embodiments of this application, targeting the numerous and diverse equipment data in subway stations, i.e., the pulled multimodal data, intramodal event detection and cross-modal event association are performed with events as the center. This enables data from different sources to jointly represent the station's operating status from multiple dimensions, improving the comprehensiveness of the station's operating status perception. It also enables the identification of different candidate events mapped by each modal data as belonging to the same real station event, avoiding misjudgments caused by a single data source.
[0013] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0014] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit this application. Attached Figure Description
[0015] The above and other objectives, features and advantages of this application will become more apparent from a detailed description of exemplary embodiments thereof with reference to the accompanying drawings.
[0016] Figure 1 This is a flowchart illustrating a method for multimodal data fusion of subway stations according to an embodiment.
[0017] Figure 2 It is based on Figure 1 The flowchart shown in the corresponding embodiment describes the steps of performing intra-modal event detection on multimodal data and obtaining corresponding candidate events from the site operation status description data of each modality.
[0018] Figure 3 It is based on Figure 1 The flowchart of the method described in the corresponding embodiment describes the steps of candidate events corresponding to the operational status description data of each modal station, cross-modal event association calculation based on each modal data, merging candidate events belonging to the same real station event, and obtaining the corresponding station event entity.
[0019] Figure 4 It is based on Figure 3 The corresponding embodiment shows a flowchart describing the steps of matching time windows for candidate events based on time features in the event feature vector.
[0020] Figure 5 It is based on Figure 3 The corresponding embodiment shows a method flowchart in which several temporally related candidate events are correlated using spatial features, semantic features, and confidence features to determine candidate events that are aggregated into the same event cluster.
[0021] Figure 6 It is based on Figure 3 The method flowchart described in another embodiment is used to analyze the steps of determining candidate events that are aggregated into the same event cluster by performing association calculations on several temporally related candidate events through spatial features, semantic features, and confidence features, as shown in the corresponding embodiment.
[0022] Figure 7 It is based on Figure 1The corresponding embodiment shows a flowchart describing the steps of forming representational data of station event entities by performing structured aggregation processing on the station operation status description data corresponding to each candidate event mapped to the station event entity, with the station event entity as the fusion center.
[0023] Figure 8 This is a flowchart illustrating a method for multimodal data fusion of subway stations, according to another exemplary embodiment.
[0024] Figure 9 It is based on Figure 8 The corresponding embodiment shows a flowchart describing the steps of calculating the propagation path of a station event entity in the spatial topology and the propagation speed from the event occurrence spatial node to each node on the propagation path, based on the spatial topology mapped according to the spatial structure of the station and the event occurrence spatial node of the station event entity.
[0025] Figure 10 This is a flowchart illustrating a multimodal data fusion method for subway stations according to another exemplary embodiment.
[0026] Figure 11 This is a schematic diagram of the interface of the equipment deployed inside a subway station on the Xicheng Line, according to an exemplary embodiment.
[0027] Figure 12 yes Figure 11 An enlarged view of the selected area on the left side of the interface shown. Detailed Implementation
[0028] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided to make the description of this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art. The drawings are merely illustrative of this application and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.
[0029] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more exemplary embodiments. Numerous specific details are provided in the following description to give a full understanding of exemplary embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced with one or more of the specific details omitted, or other methods, components, steps, etc., can be employed. In other instances, well-known structures, methods, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0030] Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0031] The embodiments of this application are used to automate the modeling of pitched roof buildings in digital space, thereby enabling the acquisition of high-precision pitched roof buildings in digital space, enhancing geometric modeling capabilities, and providing support for the vertical application of digital twin scenarios.
[0032] See Figure 1 , Figure 1 This is a flowchart illustrating a method for multimodal data fusion of subway stations according to an embodiment. The multimodal data fusion method for subway stations provided in this application includes: Step S110: By pulling real-time multimodal data from the subway station, multimodal data of the actual data environment of the station is continuously obtained. The multimodal data includes station operation status description data of at least one mode, and the mode corresponds to different types of equipment in the station. Step S120: Perform intra-modal event detection of multimodal data, and obtain corresponding candidate events from the site operation status description data of each modality. The candidate events are the changes and / or abnormal patterns of site operation status represented by the corresponding site operation status description data within the modality. Step S130: For the candidate events corresponding to the operational status description data of each modal station, perform cross-modal event association calculation based on the data of each modal, merge the candidate events belonging to the same real station event, and obtain the corresponding station event entity. Step S140: Using the station event entity as the fusion center, the station operation status description data corresponding to each candidate event mapped to the station event entity is processed by structured aggregation to form the event representation data of the station event entity.
[0033] These steps are explained in detail below.
[0034] Subway stations are equipped with a variety of devices, such as video surveillance equipment, passenger flow detection equipment, turnstiles, environmental monitoring equipment, and equipment operation status monitoring systems. These devices will be connected as data source devices to enable real-time multimodal data streaming of the subway station.
[0035] The aforementioned different devices continuously generate data to describe the station's operational status during station operation. Data from one type of device constitutes a mode of station operational status description data.
[0036] Specifically, in step S110, a station data access interface is established for the subway station to uniformly access the data streams of various devices, and a real-time data pull mechanism is used to obtain data streams from different devices.
[0037] The equipment deployment adapted to subway stations yields multimodal data, including at least one modality of station operation status description data, where different modalities correspond to data generated by different types of equipment in the station.
[0038] For example, data modalities include video data modalities corresponding to video surveillance equipment, passage record data modalities corresponding to turnstile equipment, passenger flow density data modalities corresponding to passenger flow detection equipment, and equipment status data modalities corresponding to equipment operation monitoring systems. Correspondingly, multimodal data includes various modal data such as video data, passage record data, passenger flow density data, and equipment status data, which also serve as descriptive data of the station's operational status in each modality. Each modal data describes the station's operational status from different dimensions, thus forming multimodal data within the real-world data environment of the station.
[0039] In an exemplary embodiment, step S110 is further described as follows: when acquiring a mode of station operation status description data, the corresponding spatial node is obtained according to the mapping of the data source device in the station spatial structure, and is attached to the station operation status description data as spatial sensing data. The spatial node represented by the spatial sensing data indicates the location where the station operation status description data occurs in the station spatial structure, supporting cross-modal event association of subsequent candidate events and spatial positioning of associated station event entities.
[0040] For a pulled data stream, namely a modal description of the station's operational status, the data source device to which it belongs has a mapping relationship in the station's spatial structure, and this mapping relationship indicates the corresponding spatial node.
[0041] The station spatial structure is a pre-constructed structured spatial model of a subway station containing multiple spatial nodes. These spatial nodes represent different physical or functional areas within the station, including but not limited to the concourse area, platform area, passageway area, entrance / exit area, and equipment layout. The spatial nodes are connected through topological relationships, reflecting the connectivity of the station's internal space.
[0042] For different types of data source devices, a mapping relationship is established between the data source device providing the station operation status description data and at least one spatial node in the station's spatial structure, based on the device's installation location or coverage area. This allows for the automatic parsing of the spatial node to which the data belongs when acquiring station operation status description data for each modality, thereby achieving corresponding spatial sensing data. For example, the obtained spatial sensing data may include at least one of the following: spatial node identifier, spatial region type, and spatial coordinate information.
[0043] Furthermore, for devices with dynamically changing coverage or spatial uncertainties, the obtained site operation status description data will be used for positioning and correction to improve the accuracy of spatial positioning perception.
[0044] In this way, the original site operation status description data is accompanied by clear spatial semantic information when it is acquired, thus forming a multi-dimensional data expression that includes both temporal and spatial information.
[0045] With the addition of spatially aware data to the site operation status description data, subsequent intramodal event detection and cross-modal event association will be able to obtain spatial semantics, thereby avoiding the uncertainty caused by relying solely on time alignment and significantly improving the accuracy and stability of multimodal event association.
[0046] After obtaining multimodal data, step S120 first performs event detection within each modality to identify candidate events that can characterize changes or anomalies in the station's operating status within the respective modality.
[0047] With the streaming of multimodal data, intramodal event detection is oriented towards each data modality. Within each data modality, it identifies and characterizes the changes or anomalies in the station's operational status based on the corresponding modality's station operation status description data. Then, it abstracts the identified status patterns into the execution process of candidate events.
[0048] As mentioned earlier, a data modality refers to a data category comprised of data generated by the same type of equipment. During intra-modal event detection, only the station events corresponding to the data features presented by the station operation status description data within the same modality are identified as candidate events. Continuous station status description data within the same modality form a stable data stream sequence over time, reflecting the dynamic evolution of equipment operation status and thus the station operation status. Based on this, through intra-modal event detection, identification mechanisms and pattern recognition mechanisms are constructed for the data features present in each modality to capture potential abnormal states or key state changes, thereby achieving the identification of candidate events.
[0049] Intramodal event detection, being oriented towards a single modality and a continuous data sequence, significantly enhances robustness and accuracy compared to the diverse and complex equipment and data obtained from subway stations, avoiding errors caused by noise in data from a single moment. Furthermore, obtaining continuous station state description data within the same modality provides reliable temporal context, feature evolution basis, and anomaly discrimination foundation for intramodal event detection, thus reliably supporting the effective acquisition of candidate events.
[0050] To further explain, the candidate events obtained from intra-modal event detection are event description units corresponding to their respective data modal and characterizing changes or abnormal patterns in the current station's operational status. In other words, candidate events are event representations generated by abstracting data segments with significant change or abnormal characteristics from the station operational status description data of their respective data modal.
[0051] For example, each candidate event includes the following event description information: event occurrence time, event occurrence spatial node, event type, event confidence level, and event source modality. The event occurrence time characterizes the time when the event occurs or is detected, and even its duration; the event occurrence spatial node characterizes the station area or equipment corresponding to the event; the event type characterizes the type of state change corresponding to the event, such as abnormal passenger flow, equipment malfunction, or crowd gathering; the event confidence level characterizes the reliability of the event detection result; and the event source modality identifies the data modality type from which the candidate event originates.
[0052] It should be understood that candidate events only reflect the state changes detected in a certain modality of data. Therefore, different modalities may generate multiple candidate events, and these candidate events may correspond to the same real station event or belong to different events.
[0053] For example, when an abnormal passenger flow occurs in the concourse area of a subway station, the currently retrieved multimodal data may include video surveillance data that may detect candidate events of crowd gathering, passenger flow detection equipment that may detect candidate events of abnormal passenger flow density, and turnstile equipment that may detect candidate events of a sudden increase in entry flow. Although these candidate events come from different modalities, they may all reflect the same real station event.
[0054] In step S120, continuous site operation status description data is transformed into discrete candidate events by detecting submodal events within a modality. This enables subsequent steps, such as step S130, to uniformly associate different modal data at the event level, without relying on direct alignment at the original data level. This reduces the impact of differences in the acquisition frequency of different modal data on event recognition and provides a foundation for subsequent cross-modal events and event representation at the data level.
[0055] In an exemplary embodiment, multimodal data-oriented intramodal event detection for each data modality involves numerically transforming site operation status description data, identifying event units corresponding to time segments, and finally attaching spatial nodes to obtain candidate events with temporal and spatial semantic attributes.
[0056] Numerical transformation enables the acquisition of calculable data representations, thereby lowering the threshold for data understanding. Event unit identification oriented towards time segments discretizes continuous data streams into event representation units with clear time boundaries, transforming the event detection process from "instantaneous state judgment" to "dynamic evolution analysis based on time segments." This allows for more accurate capture of key turning points in the state change process, improving the temporal sensitivity and accuracy of event detection.
[0057] Furthermore, by adding spatial node information to the event unit, the original event representation, which only had temporal characteristics, is expanded into candidate events that simultaneously possess temporal and spatial semantic attributes, thus achieving spatiotemporal integrated modeling of events. This approach not only accurately locates the time and spatial position of events but also provides a unified semantic carrier for subsequent cross-modal event association, event evolution analysis, and global situational awareness.
[0058] Furthermore, since intramodal event detection is performed independently within each data modality, the processing procedures between different modalities are decoupled, thereby improving the system's scalability and modularity while ensuring detection accuracy. When a new data modality is added, it only needs to be independently modeled and its events detected to be seamlessly integrated into the overall system, avoiding any impact on existing processing flows.
[0059] Therefore, this exemplary embodiment not only improves the accuracy of candidate event identification and spatiotemporal representation, but also enhances the uniformity, scalability, and supportability of multimodal data processing for subsequent complex event analysis.
[0060] See Figure 2 , Figure 2 It is based on Figure 1 The flowchart shown in the corresponding embodiment describes the steps of performing intra-modal event detection on multimodal data and obtaining corresponding candidate events from the site operation status description data of each modality.
[0061] The step S120 of this application embodiment for intra-modal event detection of multimodal data, which obtains corresponding candidate events from the site operation status description data of each modality, includes: Step S121: Perform numerical conversion of the station operation status of each mode on the station operation status description data of the respective mode, and obtain the numerical description sequence of the subway station by mode. Step S122: Perform sliding window detection of the numerical description sequence to obtain event units corresponding to continuous time segments; Step S123: For the event unit, generate candidate events with temporal and spatial semantic attributes by configuring the spatial nodes attached to the site operation status description data.
[0062] These steps are explained in detail below.
[0063] In step S121, the station operation status description data for each modality is numerically processed to obtain a numerical description sequence of the subway station by modality. Specifically, different numerical conversion methods will be used to adapt to the data modality for the station operation status description data of different modalities.
[0064] For example, for video surveillance data, values such as target quantity, personnel density, movement speed, and area occupancy rate are calculated and extracted to convert them into time-series numerical data. For turnstile access data, the number of people entering and exiting the station per unit time is statistically analyzed into discrete time series. For passenger flow detection equipment data, the output passenger flow density or flow rate numerical series is directly obtained. For equipment operation status data, the equipment status (normal, abnormal, and offline) can be converted into discrete numerical codes or status indicator sequences, which will not be listed here.
[0065] After the numerical transformation is completed, a numerical description sequence that changes over time is obtained for each data mode. This numerical description sequence is used to characterize the dynamic change process of the station's operating status under the corresponding data mode.
[0066] In step S122, after obtaining the numerical description sequence corresponding to each mode, the numerical sequence is segmented and analyzed through a sliding window mechanism to identify time segments with significant state change characteristics.
[0067] Specifically, a time-sliding window can be constructed on the numerical description series. The sliding window slides along the time axis according to the configured window length, and the series data is analyzed and processed within each window. For example, statistical characteristics within the window can be calculated, including mean, rate of change, fluctuation range, or trend changes.
[0068] When the numerical features within the sliding window meet preset conditions, it is determined that the sequence segment corresponding to the sliding window has a state change or an abnormal trend, thereby marking the sequence segment as an event unit. In an exemplary embodiment, the preset conditions may be: the numerical change exceeds a preset threshold, the numerical change trend undergoes a sudden change, the fluctuation amplitude exceeds the normal range, or an abnormal pattern identified based on model inference.
[0069] Furthermore, for the obtained event units, the event units will be merged when the corresponding sliding windows are adjacent or overlapped to obtain time-continuous event units, thereby avoiding the event being segmented by the process.
[0070] In summary, in step S122, the sliding window detection mechanism can transform a continuous numerical sequence into several event units with time boundaries, thereby realizing the transformation from "continuous data stream" to "discrete event fragment".
[0071] After obtaining the event unit in step S123, the spatial location corresponding to the event unit is determined by combining the spatial nodes in the site operation status description data attached in step S110, thereby forming the spatial semantic attributes of the candidate event.
[0072] Building upon this foundation, further information on event type is provided for candidate events. For example, based on numerical sequence variation characteristics or detection rules, event units can be classified into passenger flow anomaly events, equipment anomaly events, or behavioral anomaly events, thereby forming a complete description of candidate events.
[0073] Thus, each candidate event becomes an event object that is structured and contains at least specific descriptive information. For example, the specific descriptive information includes: temporal semantic attributes (the time interval of the event occurrence), spatial semantic attributes (the spatial node corresponding to the event), event type, source modality identifier, and confidence or intensity value.
[0074] This exemplary embodiment realizes the processing steps of numericalization, sliding window temporal decomposition, and spatiotemporal semantic enhancement, achieving an abstract transformation from the data layer to the event layer. Addressing the issues of data alignment and independent analysis in existing implementations, this embodiment constructs standardized event representation units within each modality, enabling subsequent association of data from different modalities within a unified spatiotemporal semantic framework. This effectively reduces the impact of differences in acquisition frequency and representation format among multimodal data, improving the accuracy and consistency of event detection and fusion.
[0075] After completing intramodal event detection in step S120 to obtain candidate events for each modality as they evolve over time, cross-modal event association calculation will be performed across all modalities in step S130 to fuse the candidate events from each modality to obtain the real station events.
[0076] In step S130, for candidate events of different modalities, the correlation between events of different modalities is identified to determine which candidate events belong to the same real station event. In other words, candidate events are aggregated across modalities to obtain an event cluster composed of several cross-modal candidate events to describe a real station event.
[0077] It should be noted that cross-modal event association calculation refers to the process of calculating the association between candidate events based on candidate events from different data modalities in terms of time, space and semantic dimensions, in order to determine whether multiple candidate events belong to the same real station event, and to aggregate candidate events that meet the association conditions.
[0078] Specifically, in an exemplary embodiment, cross-modal event association computation includes at least the feature construction of each candidate event, and the association computation in time, space, semantics and confidence dimensions based on the constructed features.
[0079] For example, the correlation calculations performed in each dimension are based on a step-by-step screening mechanism for candidate events in each dimension, thereby completing the step-by-step screening of candidate events at the time, space, semantic and confidence levels, and obtaining an event cluster that represents a real station event across modalities.
[0080] To further explain, in an exemplary embodiment, the candidate events at the confidence level are screened step by step by performing a weighted evaluation of the spatial consistency value, semantic similarity, and confidence features of the spatial distance mapping calculated for the currently retained candidate events, in order to determine the candidate events belonging to the same event cluster.
[0081] Therefore, in the execution of step S130, three-dimensional association is implemented based on spatial structure, event semantics, and reliability to complete the screening of cross-modal candidate events, so as to finally align them to the same real station event. Specifically, this includes: First, in the time dimension, an event association time window is constructed for each candidate event to screen out a set of candidate events that satisfy overlapping or adjacent relationships in time; Second, in the spatial dimension, based on the station spatial structure model, the consistency or proximity of the spatial nodes corresponding to the candidate events is judged to screen out candidate events that belong to the same area or have spatial association relationships; Further, in the semantic dimension, the event type of the candidate events and the state change characteristics they represent are semantically matched to screen out candidate events that are semantically consistent or have associations; After completing the multi-dimensional screening of time, space, and semantics, the reliability of the candidate events is introduced to evaluate the association credibility of the screened candidate events, and the association relationship between the candidate events is determined based on preset association rules or thresholds. Finally, candidate events that meet the association conditions in terms of time, space and semantics and whose association credibility reaches the preset requirements will be aggregated to form a set of candidate events corresponding to the same real station event, and a unified station event entity will be constructed accordingly.
[0082] Through the aforementioned three-dimensional association mechanism based on spatial structure, event semantics, and reliability, the layer-by-layer screening and precise alignment of candidate events of different modalities are achieved, thereby effectively avoiding the mismatch problem caused by relying solely on a single-dimensional association and improving the accuracy and stability of cross-modal event fusion.
[0083] Candidate events belonging to the same real station event are merged into a unified station event entity. To elaborate further, a real station event refers to a physical or operational state change process that objectively occurs during the actual operation of a subway station and affects the station's operational status. Real station events originate from the actual operating environment of the station, possess objective existence, and do not depend on specific data representation forms. Examples include a passenger flow gathering, a sudden increase in passenger flow entering the station during a certain period, an abnormality or malfunction of a certain piece of equipment, or abnormal behavior or operational status changes in a certain area.
[0084] Because the various devices deployed in subway stations employ different sensing methods, real-world station events typically manifest in multiple data modalities. For example, regarding abnormal passenger flow events, video surveillance equipment might detect crowd gathering, passenger flow detection equipment might detect increased passenger density, and turnstiles might detect a sudden surge in throughput. A single real-world station event can be perceived by multiple modalities, manifesting as multiple data changes across different modalities. These changes, in essence, correspond to the same real-world station event.
[0085] The station event entity corresponds to the obtained real station event and is a unified expression of the real station event at the system level. That is to say, with the cross-modal event association calculation performed in step S130, on the one hand, the real station events corresponding to the changes reflected by multiple modes are determined, and on the other hand, with the fusion of candidate events belonging to this real station event, a unified expression of this real station event at the system level is obtained, that is, the station event entity of this real station event.
[0086] In one exemplary embodiment, the station event entity includes the event occurrence time, the event occurrence spatial node, the event type, and a set of associated candidate events. Within the station event entity, each associated candidate event serves as a component, used to collectively describe the corresponding real station event from different modalities, thereby achieving a multidimensional expression of the real station event.
[0087] In summary, in the implementation of this application embodiment, objective events in the station operating environment, i.e., real station events, are used as several candidate events as the local perception results of their respective modalities, and the station event entity is a unified data representation of real station events. In other words, the station event entity is a data-layer abstract mapping of real station events, used to uniformly express and process real station events in the system implementation.
[0088] By linking cross-modal events in step S130, candidate events scattered in different modalities are aligned and fused to restore real station events and to uniformly model them in the form of station event entities.
[0089] By introducing real station events and station event entities, the multimodal data processing process is upgraded from data-level fusion to event-level fusion, realizing the transformation from changes in multi-source data to a unified event expression. This effectively solves the problems of difficulty in aligning different modal data and difficulty in uniformly representing the same station event, and significantly improves the accuracy and consistency of station operation status modeling.
[0090] See Figure 3 , Figure 3 It is based on Figure 1 The flowchart of the method described in the corresponding embodiment describes the steps of candidate events corresponding to the operational status description data of each modal station, cross-modal event association calculation based on each modal data, merging candidate events belonging to the same real station event, and obtaining the corresponding station event entity.
[0091] The embodiment of this application provides a step S130, which involves performing cross-modal event association calculations based on the modal data to identify candidate events corresponding to the operational status description data of each modal station, fusing candidate events belonging to the same real station event, and obtaining the corresponding station event entity. This step includes: Step S131: For the obtained candidate events, construct event feature vectors by extracting features from the site operation status description data. The event feature vectors include time features, spatial features, semantic features, and confidence features. Step S132: Match the time windows of candidate events based on the time features in the event feature vector to determine several candidate events that are related in time. Step S133: For several candidate events that are temporally related, associate them using spatial features, semantic features, and confidence features to determine candidate events that are aggregated into the same event cluster; Step S134: Construct station event entities across modalities for event clusters. The station event entity includes event occurrence time, event occurrence spatial node, event type, and associated candidate event set.
[0092] These steps are explained in detail below.
[0093] By executing the steps in this exemplary embodiment, candidate events obtained from event detection within a multimodal data modality are identified, and the actual station events to which they belong are determined through cross-modal event association calculation. This results in candidate events of actual station events scattered across different modalities, and then the station event entity is uniformly constructed by fusing these candidate events.
[0094] In step S131, for the candidate events obtained from intramodal event detection, feature extraction is performed on the corresponding site operation status description data to construct an event feature vector that uniformly expresses the candidate events.
[0095] For example, an event feature vector includes at least temporal features, spatial features, semantic features, and confidence features. Temporal features characterize the time information of the candidate event, such as timestamps, durations, or time intervals. Spatial features characterize the spatial structure of the station and the spatial location of the candidate event, such as by station number, equipment node, or spatial coordinates. Semantic features describe the type or semantic content of the candidate event, such as "passenger flow gathering," "equipment malfunction," and "turnstile congestion." Confidence features characterize the detection reliability of the candidate event and are derived from the output probability of in-modal event detection.
[0096] The construction of event feature vectors performed in step S131 enables candidate events from different modalities to be represented in a unified feature space, providing a foundation for subsequent cross-modal association.
[0097] With the feature construction of candidate events corresponding to the obtained site operation status description data, the obtained event feature vectors are matched with time windows in step S132 to first determine the candidate events that are associated in time.
[0098] In step S132, preliminary screening of candidate events is performed based on the temporal features in the event feature vector. Time window matching is then applied to all candidate events based on these temporal features. Specifically, a candidate event (not selected from any event cluster) is used as the event anchor point. Other candidate events are then selected based on their temporal features within the associated time window range. At this point, the candidate event used as the event anchor point has a temporal correlation with the other selected candidate events, thereby effectively reducing the search space for subsequent correlation calculations and improving computational efficiency.
[0099] As candidate events serving as event anchors, the time features in their event feature vectors indicate the time window range corresponding to the current round of time window matching. This time window range is adapted to the window width dynamically determined by the time feature. For example, the window width is the time feature ± Δt (set value), but it can also be determined by the time feature, the set value Δt, and the compensated modal response delay; this is not limited here.
[0100] For multimodal candidate events, each round of time window matching will be performed, and a step-by-step screening will be carried out on this basis to achieve time consistency constraint screening of cross-modal candidate events and event clustering association through multi-feature fusion.
[0101] In step S133, for several temporally related candidate events obtained in step S132 through a round of time window matching, multidimensional association calculation is further performed based on spatial features, semantic features, and confidence features to determine candidate events that can be aggregated into the same event cluster.
[0102] It should be noted that the correlation calculation performed at the spatial feature level is the execution process of spatial consistency determination, that is, detecting whether candidate events have spatial proximity relationships, such as occurring in the same or adjacent areas of the same site.
[0103] The association calculation performed at the semantic feature level determines the consistency or association of candidate events in terms of event type by calculating the semantic proximity of semantic features. At the confidence feature level, the association relationship is evaluated by weighting the spatial consistency value obtained from the spatial consistency judgment, semantic similarity, and confidence features, so as to finally determine the candidate events that can be aggregated into the same event cluster, thereby reducing the interference of low-reliability events on clustering.
[0104] This approach implements multi-feature constraints, enabling the aggregation of candidate events to form event clusters, each corresponding to a real station event. Compared to methods based solely on a single modality, step S133 achieves multi-dimensional joint modeling of time, space, semantics, and reliability, significantly improving the accuracy and robustness of cross-modal event fusion.
[0105] Finally, for the candidate events that have been identified and aggregated into the same event cluster, the station event entities are constructed under the action of S134 to realize the mapping from multimodal candidate events to unified semantic entities, so that the original scattered, multi-source heterogeneous event descriptions are transformed into structured and computable station event representations.
[0106] In summary, by constructing a unified event feature vector encompassing temporal, spatial, semantic, and confidence features, isomorphic representations of multimodal candidate events within the same feature space are achieved, effectively addressing the challenge of directly fusing heterogeneous data. Furthermore, a phased association mechanism combining initial screening based on time windows with multidimensional fine-grained association based on spatial, semantic, and confidence dimensions significantly improves the accuracy of cross-modal event matching while reducing computational complexity. Moreover, the introduction of semantic similarity measurement and confidence weighting mechanisms weakens the interference of low-reliability candidate events on the fusion results, enhancing the robustness of overall event recognition. Finally, by fusing event clusters to construct structured station event entities, closed-loop modeling from dispersed candidate events to unified semantic events is achieved, providing a standardized and computable data foundation for subsequent event evolution analysis and intelligent decision-making.
[0107] See Figure 4 , Figure 4 It is based on Figure 3 The corresponding embodiment shows a flowchart describing the steps of matching time windows for candidate events based on time features in the event feature vector.
[0108] The step S132 of matching candidate events based on time features in the event feature vector provided in this application embodiment includes: Step S1321: For candidate events that are not involved in time association, select event anchor points to obtain the corresponding event anchor points for the current round of time window matching. The event anchor point is the earliest or the candidate event with the highest confidence. Step S1322: Based on the time characteristics of the event anchor point, and with the modal response delay corresponding to the data source device as compensation, adaptively determine the time window for the current round of matching; Step S1323: For the event anchor point, perform a cross-modal candidate event association search within a defined time window to obtain candidate events that are temporally associated with the event anchor point across modalities. Temporally associated candidate events are locked and will not be included in subsequent time window matching.
[0109] These steps are explained in detail below.
[0110] In this exemplary embodiment, for all candidate events that are not involved in time association, that is, taking one of the candidate events as the event anchor point, the temporal correlation between other candidate events and this event anchor point is calculated based on the time features in the event feature vector. This is the implementation of sliding time window matching, thereby realizing an event-driven time window adaptive matching mechanism.
[0111] First, the selection of event anchor points is achieved through the execution of step S1321. For all candidate events that have not yet participated in time association, the earliest or highest confidence candidate event is selected as the event anchor point for the current round of time window matching based on the time feature or confidence feature in the event feature vector.
[0112] By selecting event anchor points, each time window matching revolves around a representative candidate event, avoiding unordered traversal and improving the stability and efficiency of the matching process.
[0113] After determining the event anchor point, step S1322 determines the window width of the time window used in the current round based on the time characteristics of the event anchor point and the modal response delay of the data source device to which it belongs, so as to achieve adaptive adjustment of the time window matching the applicable time window for each round.
[0114] In step S1322, the modal response delay is used to characterize the time lag value of the data modality between the occurrence of the event and its detection. It corresponds to the data source device, such as video detection delay, sensor sampling delay, etc.
[0115] For example, the start and end boundaries of the time window are compensated and corrected based on the timestamp of the event anchor point and the response delay parameter of its mode. In an exemplary embodiment, the start and end boundaries of the time window are determined by a set value Δt, thereby adaptively determining the set value Δt and the delay distribution of different modes.
[0116] By introducing a modal response delay compensation mechanism, observations of the same real station event under different modes can be aligned in time, and the problem of correlation omission caused by acquisition or processing delays can be avoided.
[0117] Based on the adaptive time window in step S1322, cross-modal association and locking in time can be achieved through the execution of step S1323.
[0118] In step S1323, with the event anchor point as the center, within the time window determined in step S1322, candidate events from different modalities are searched to obtain candidate events that are related in time.
[0119] Specifically, in an exemplary embodiment, the execution process of step S1323 includes: performing a traversal on all candidate events that are not involved in time association outside the event anchor point to filter out candidate events whose time features fall within the time window range; establishing a time association relationship between the filtered candidate events and the event anchor point to obtain candidate events for the current round of time window matching; and locking these candidate events so that they no longer participate in the time window matching of subsequent rounds.
[0120] Ultimately, this locking mechanism ensures that each candidate event participates in the time association process only once, avoiding duplicate matching or multiple attribution issues and guaranteeing the consistency of subsequent event clustering.
[0121] In summary, by executing this exemplary embodiment, the matching process is organized starting from the event anchor point, and the time window is dynamically adjusted based on modal differences. This allows for cross-modal search and locking within the window, avoiding computational redundancy and latency sensitivity issues caused by global sliding windows.
[0122] It should be understood that, unlike indiscriminate sliding window scanning of all candidate events, this approach uses event anchors as the core to achieve round-by-round matching, effectively reducing computational complexity and improving processing orderliness. During the time window determination process, response delay parameters for different data modalities are introduced to achieve adaptive alignment across modal time axes, effectively solving the event mismatch problem caused by inconsistent acquisition or processing delays in multi-source data. The time window range is dynamically adjusted based on the event anchors and their modal characteristics, which, compared to a fixed window approach, can more accurately cover the time distribution of real station events and improve the accuracy of time association. By locking candidate events that have already participated in time association, it ensures that each candidate event belongs to only one time association relationship, avoiding duplicate calculations and ambiguous attribution, and providing stable input for subsequent event cluster construction.
[0123] See Figure 5 , Figure 5 It is based on Figure 3 The corresponding embodiment shows a method flowchart in which several temporally related candidate events are correlated using spatial features, semantic features, and confidence features to determine candidate events that are aggregated into the same event cluster.
[0124] The step S133 provided in this application embodiment, which involves performing association calculations on several temporally related candidate events using spatial features, semantic features, and confidence features to determine candidate events aggregated into the same event cluster, includes: Step S201: For several candidate events that are temporally related, calculate the spatial distance between the candidate events and the event anchor point based on the spatial characteristics and the topological structure indicated by the spatial structure of the station. Step S202: Based on the spatial proximity relationship determined by the spatial distance for candidate events, candidate events that occur in the same or adjacent areas of the event anchor point are identified as candidate events of the same event cluster.
[0125] The following is a detailed explanation of these two steps.
[0126] In this exemplary embodiment, for the several candidate events that are temporally correlated obtained in step S132, spatial correlation calculation is performed on the candidate events through spatial features and the topological structure indicated by the station spatial structure, so as to determine the candidate events of the same event cluster at the spatial level.
[0127] In step S201, for temporally correlated candidate events, the spatial distance of each candidate event relative to the event anchor point is calculated with reference to the event anchor point. Unlike calculation methods based solely on geometric coordinates or Euclidean distance, in this exemplary embodiment, the spatial distance is calculated based on the spatial characteristics of the candidate event and the topological structure indicated by the spatial structure of the station to which it belongs.
[0128] In an exemplary embodiment, the execution process of step S201 includes: obtaining spatial node information corresponding to the candidate event, wherein the spatial node may be a station hall, platform, passage, turnstile area or equipment node, etc.; based on a pre-constructed station spatial topology model, abstracting each spatial node into a node in a graph structure, wherein the connection relationship between nodes represents the actual travel path or spatial connectivity relationship; in the topology graph, taking the spatial node where the event anchor point is located as the starting point, determining the topological distance, i.e., the spatial distance, between the node where each candidate event is located and the event anchor point through path search or graph distance calculation (e.g., shortest path length, number of hops, etc.).
[0129] This execution process elevates spatial-level correlation calculations from "geometric coordinates" to structured spatial relationships, enabling the resulting spatial distances to reflect the actual scenario, namely, accessibility and functional zoning within subway stations.
[0130] In step S202, the spatial proximity relationship between the candidate event and the event anchor point is determined based on the spatial distance between the candidate event and the event anchor point, thereby determining whether they belong to the same event cluster.
[0131] Specifically, the execution process of step S202 includes: when the topological distance between a candidate event and an event anchor point is less than a preset threshold, or when they are located in the same spatial node or directly adjacent nodes, they are determined to satisfy the spatial proximity relationship; the candidate events that satisfy the spatial proximity relationship are selected as candidate events spatially associated with the event anchor point; and the candidate events of this type are grouped together with the event anchor point into the same event cluster.
[0132] In an exemplary embodiment, the determination of spatial proximity relationships is not only based on a preset threshold, but also dynamically adjusted in combination with the functional attributes of the area to which the spatial structure belongs in the station. That is, the preset threshold is dynamically adjusted according to the functional attributes of the area to which it belongs. For example, a stricter threshold is set for the platform area, and a larger spatial expansion is allowed for the passage area, thereby adapting to the spatial propagation characteristics of different candidate events.
[0133] It should be understood that by dividing the functional areas in the spatial structure of the station, the station hall, platform, passage and other functional areas are incorporated into the spatial nodes of the subway station. This makes the calculated spatial distance not only reflect distance but also regional attributes. It is not only suitable for the operation of the real environment of the subway station, but also improves the accuracy of the semantic understanding of events. This makes the spatial association process have excellent adaptability and scalability.
[0134] In another exemplary embodiment, step S133, which involves performing association calculations on several temporally correlated candidate events using spatial features, semantic features, and confidence features to determine candidate events aggregated into the same event cluster, further includes: For the event clusters formed by the event anchor and the candidate events that have spatial proximity, the event anchor is used as the benchmark to further filter the candidate events based on semantic features, and the semantically related candidate events are retained in this event cluster.
[0135] After completing steps S201 and S202 to obtain the event clusters constructed based on spatial proximity, a similarity filtering mechanism based on semantic features is introduced to constrain the semantic consistency of the event clusters based on the spatial association results.
[0136] The execution process of this semantic similarity screening includes: for an event cluster formed by an event anchor point and candidate events that have a spatial proximity relationship with it, continue to use the event anchor point as a benchmark to perform semantic similarity calculation on each candidate event in the event cluster, and screen out candidate events that are semantically related to the event anchor point.
[0137] This embodiment introduces a semantic similarity screening mechanism centered on event anchors at the semantic association level. Based on the candidate event set that has completed temporal association and spatial proximity constraints, semantic consistency is judged for each candidate event, thereby constructing a multi-dimensional association process that converges layer by layer in "time-space-semantics". By using event anchors as a unified semantic reference center, not only is the instability caused by centerless semantic clustering avoided, but the accuracy of semantic matching is also improved while reducing computational complexity. At the same time, under the action of semantic features, the system has a stronger semantic characterization ability while ensuring interpretability. This effectively avoids the problem of mis-aggregation when different types of events overlap in time and space, and significantly improves the accuracy and reliability of cross-modal event fusion.
[0138] See Figure 6 , Figure 6 It is based on Figure 3 The corresponding embodiment shows a method flowchart in which several temporally related candidate events are correlated using spatial features, semantic features, and confidence features to determine candidate events that are aggregated into the same event cluster.
[0139] The step S133 provided in this application embodiment, which involves performing association calculations on several temporally related candidate events using spatial features, semantic features, and confidence features to determine candidate events aggregated into the same event cluster, includes: Step S301: Perform a weighted evaluation of the association relationship of candidate events based on the spatial consistency value, semantic similarity, and confidence features of the spatial distance mapping calculated from the candidate events in the event cluster. Step S302: Based on the weighted score of the obtained association relationship, determine the candidate events that are aggregated into the same event cluster. The determined candidate events and the event anchor point constitute an event cluster.
[0140] The following is a detailed explanation of these two steps.
[0141] In this exemplary embodiment, a weighted evaluation mechanism based on spatial consistency, semantic similarity, and confidence features is introduced to comprehensively determine whether candidate events belong to the same event cluster.
[0142] Specifically, based on the confidence features contained in the event feature vector, the spatial consistency value and semantic similarity calculated above are weighted to obtain a score for the association between the candidate event and the event anchor. In this way, the originally scattered multidimensional feature vector is transformed into a numerically comparable weighted score for the association, realizing a quantitative expression of the association between candidate events.
[0143] Based on the weighted score of the obtained correlation, it is determined whether the candidate event has a strong correlation with the event anchor point, and in this way, candidate events with a strong correlation with the event anchor point are selected to form the event cluster used in the end.
[0144] Candidate events that do not have a strong correlation with the event anchor will not be included in the event cluster, but can participate in the correlation calculation in subsequent rounds.
[0145] As mentioned above, after the station event entity is constructed in step S130, the station event entity is further used as the fusion center to perform structured aggregation processing on the station operation status description data corresponding to each candidate event mapped to the station event entity, forming unified event representation data.
[0146] In step S140, for each station event entity, the station operation status description data corresponding to each candidate event mapped to it is obtained. Based on the mapping relationship between the candidate events and the station event entity, the station operation status description data is aggregated under the corresponding station event entity to form a multimodal data set for the event.
[0147] Thus, the data aggregated in the multimodal dataset is structured to construct representation data of station event entities, i.e., event representation data.
[0148] In one exemplary embodiment, the resulting representation data includes different semantic fields depending on the event semantic framework. For example, the representation data includes event identification information (such as event ID), event occurrence and duration, event occurrence spatial node, event type, and associated candidate event index information.
[0149] See Figure 7 , Figure 7 It is based on Figure 1 The corresponding embodiment shows a flowchart describing the steps of forming representational data of station event entities by performing structured aggregation processing on the station operation status description data corresponding to each candidate event mapped to the station event entity, with the station event entity as the fusion center.
[0150] The embodiment of this application provides a step S140, which uses station event entities as the fusion center and performs structured aggregation processing on the station operation status description data corresponding to each candidate event mapped to the station event entity to form the representation data of the station event entity. This step includes: Step S141: Based on the pre-built event semantic framework, extract event attributes for each candidate event of the station event entity, and map the event attributes to the semantic fields corresponding to the event semantic framework. Step S142: Consistently merge multimodal event attributes mapped to the same semantic field, and resolve attribute conflicts when there are conflicting attributes to obtain the target value of the multimodal event attribute; Step S143: Generate representation data of station event entities based on each semantic field.
[0151] These steps are explained in detail below.
[0152] In this exemplary embodiment, the station event entity serves as the fusion center. The station operation status description data corresponding to each candidate event mapped to that station event entity undergoes structured aggregation processing. This is not a simple data summary, but rather a process based on a pre-built event semantic framework. It involves field-level alignment, consistency merging, and conflict resolution of multimodal event attributes. Specifically, this includes the following steps: First, attribute extraction and field mapping based on the event semantic framework are performed. For each candidate event associated with the station event entity, event attributes are extracted according to the pre-built event semantic framework, and the extracted event attributes are mapped to corresponding semantic fields. For example, the event semantic framework is a predefined structured template used to uniformly describe station events, containing multiple semantic fields, such as event time, event location, event type, impact range, status indicators, and alarm levels. For candidate events of different modalities, their corresponding event attributes are extracted, such as passenger flow density and personnel behavior characteristics extracted for video modalities; and numerical indicators (such as temperature and equipment status values) extracted for sensor modalities.
[0153] The extracted event attributes are mapped to corresponding fields in the event semantic framework according to their semantic meaning, realizing the unified alignment of multimodal attributes and unifying the mapping of event descriptions from different sources and with different structures to standard semantic fields, providing a structural foundation for subsequent fusion.
[0154] Then, for multimodal event attributes mapped to the same semantic field, a consistency merging process is performed; in the case of attribute conflicts, further conflict identification and resolution are carried out to determine the target value of the field.
[0155] In an exemplary embodiment, the conflict identification process includes: for numerical attributes, when the difference between multiple attribute values corresponding to the same semantic field exceeds a preset threshold (such as a deviation range), it is determined to be a conflict; for temporal attributes, when the time corresponding to the attribute is inconsistent or there is a significant temporal offset, it is determined to be a potential conflict.
[0156] Instead of simply selecting a specific attribute for the identified conflict attributes, in an exemplary embodiment, the attribute value from the source of high confidence is selected based on the confidence feature of the corresponding candidate event in the event feature vector.
[0157] After obtaining the non-conflicting multimodal event attributes, their target values can be consistently merged to obtain the target attribute values of the semantic fields to which they belong, and finally, the representation data can be generated.
[0158] See Figure 8 , Figure 8 This is a flowchart illustrating a method for multimodal data fusion of subway stations, according to another exemplary embodiment.
[0159] Following step S140, the multimodal data fusion method for subway stations provided in this application embodiment further includes: Step S410: Based on the spatial topology mapped by the station spatial structure of the station and the event occurrence spatial node of the station event entity, calculate the propagation path of the station event entity in the spatial topology and the propagation intensity from the event occurrence spatial node to each node on the propagation path. Step S420: Update the operating status of the station space nodes according to the propagation path and the propagation intensity of each node on the propagation path. The updated status values of the station space nodes are used for dynamic description of the operating status.
[0160] The following is a detailed explanation of these two steps.
[0161] In this exemplary embodiment, after the construction of the representation data of the station event entity is completed in step S140, the propagation process of the station event entity in space is further modeled based on the station spatial topology, and the operating status of each spatial node is updated accordingly to achieve a dynamic depiction of the station's operating status.
[0162] In step S410, the event propagation path and the propagation intensity of each node on the propagation path are calculated based on the spatial topology. For each station event entity, the propagation path of the event in the spatial topology is calculated according to the spatial topology mapped by the spatial node where the event occurs and the station spatial structure of the station to which it belongs, and the propagation intensity of the event at each node on the path is determined.
[0163] Subway stations all have their corresponding spatial structures. That is, each station along the train line has a corresponding spatial structure. The station spatial structure abstracts elements such as the concourse, platform, passageways, entrances / exits, and equipment areas into topological nodes, forming a spatial topology. The connections between these nodes represent traversable paths or functional coupling relationships. Therefore, for a defined station event entity, based on the spatial perception data added to the station's operational status description data, its location within the station spatial structure can be determined. This location is then mapped to a topological node, which becomes the event's spatial node.
[0164] Starting from the spatial node where the event occurs, path expansion is performed in the spatial topology to generate propagation paths that the station event entity may influence. In one exemplary embodiment, the path expansion is performed layer by layer based on the topological adjacency relationships existing in the spatial topology to generate candidate propagation paths; then, the generated candidate propagation paths are adapted to their respective event types and semantic attributes to determine the propagation path of the station event entity in the spatial topology.
[0165] For example, for the event type of "passenger gathering", the path that people can reach is selected as the propagation path for the generated candidate propagation path; for the event type of "equipment failure", the path that the equipment is associated with is selected as the propagation path for the generated candidate propagation path. In short, for the several candidate propagation paths generated, semantic adaptation is performed based on the event type and semantic attributes to make the resulting propagation path semantically adaptable, thereby enhancing the accuracy and reliability of describing event propagation.
[0166] For a given propagation path, the propagation intensity of the event from the source node (the node in the event occurrence space) to each node along the propagation path is calculated. In an exemplary embodiment, the propagation intensity corresponding to a node is calculated based on the topological distance from the source node, the path weight, the severity of the event itself, and node attributes. Furthermore, the calculation of the propagation intensity of the event to each node along the propagation path involves constructing a decay function based on factors such as the topological distance from the source node, the path weight, the severity of the event itself, and node attributes. This decay function characterizes the gradual weakening of the propagation intensity along the propagation path and also incorporates the factors used to facilitate rapid calculation.
[0167] Therefore, through the calculations performed, the propagation path centered on the spatial node where the event occurs and covering multiple nodes, as well as the corresponding propagation intensity distribution, can be obtained.
[0168] Based on the obtained propagation paths and the propagation intensity of each node, the operational status of the station spatial nodes is updated to achieve a dynamic description of the overall operational status.
[0169] Specifically, the process of updating the operational status of station space nodes includes: obtaining the current operational status value of the station space nodes. Since the data source devices corresponding to the station space nodes are different, their corresponding operational status values also differ. For example, operational status values can represent passenger flow density, equipment health status, congestion level, etc., and are not limited here.
[0170] For each node along the propagation path, the event's impact is applied to its own operational state value based on the corresponding propagation intensity, thereby updating its own operational state. Specifically, when a node is affected by a single event, the corresponding propagation intensity is used as an influencing factor to incrementally update the original operational state value; when multiple events act on the same node simultaneously, the propagation intensities of each event are superimposed or weighted and merged, and then the original operational state value is incrementally updated based on the resulting values.
[0171] By updating the obtained operational status values, the operational status of station spatial nodes is updated, thereby updating the operational situation of the nodes accordingly. It should be understood that different event types have different node status update strategies. The operational situation update of station spatial nodes should be implemented based on the obtained operational status values and the corresponding node status update strategies to dynamically describe the current operational situation.
[0172] Through this exemplary embodiment, the static representation of the event, which corresponds to the representation data of the station event entity, is transformed into spatial dynamic impact modeling, achieving event-driven spatial state evolution modeling. Furthermore, since it is a spatial topology implementation based on the mapping of the station spatial structure, the propagation and description of the event impact conform to the actual spatial relationship, achieving a quantitative description of the gradual decay of the event impact in space that conforms to the actual spatial relationship.
[0173] See Figure 9 , Figure 9 It is based on Figure 8 The corresponding embodiment shows a flowchart describing the steps of calculating the propagation path of a station event entity in the spatial topology and the propagation speed from the event occurrence spatial node to each node on the propagation path, based on the spatial topology mapped according to the spatial structure of the station and the event occurrence spatial node of the station event entity.
[0174] The step S410 provided in this application embodiment, which calculates the propagation path of a station event entity in the spatial topology and the propagation intensity from the event occurrence spatial node to each node on the propagation path based on the spatial topology mapped from the station's spatial structure and the event occurrence spatial node of the station event entity, includes: Step S411: Taking the spatial node where the event occurs as the event source node, and according to the spatial topology mapped by the spatial structure of the station, search from the event source node to the adjacent nodes to obtain a propagation path consisting of several nodes. Step S412: For nodes on the propagation path, determine the propagation intensity of the station event entity at the node based on the topological distance between the node and the event source node.
[0175] The following is a detailed explanation of these two steps.
[0176] In this exemplary embodiment, the propagation process of the influence of station event entities in space is not a simple spatial diffusion method, but a topological model based on the spatial structure mapping of the station is used to extend the structured path with the spatial node where the event occurs as the source point, and the propagation intensity is quantitatively modeled by combining the topological distance.
[0177] Using the event occurrence spatial node of the station event entity as the event source node, a neighbor node search is performed on the station spatial topology of the station to generate the event propagation path.
[0178] The generation of propagation paths is an execution process based on the acquisition of topological adjacency relationships and the progressive search of adjacent nodes implemented by the event source node. Specifically, firstly, the spatial node where the event occurs in the station event entity is taken as the event source node and marked in the spatial topology; then, based on the station spatial structure, the nodes formed by abstracting various spatial areas (such as station hall, platform, passage, equipment area, etc.) and the connectivity relationships between nodes represented by edges are determined to establish the topological adjacency relationships; starting from the event source node, the search is expanded layer by layer according to the topological adjacency relationships to obtain the set of nodes directly or indirectly connected to the event source node; finally, the nodes in the node set are organized into one or more paths to obtain candidate propagation paths, each of which represents a possible spatial trajectory of the event propagating outward from the event source node.
[0179] Given the candidate propagation paths, constraints can be selected based on event type and / or semantic attributes to obtain the applicable propagation path.
[0180] After obtaining the propagation path, calculate the propagation strength of the station event entity at each node on the path.
[0181] As mentioned above, by constructing the decay function, the initial influence intensity of the event source node is mapped to each node on the propagation path, and the propagation intensity of the event at each node is obtained. Through the execution process of the embodiment of this application, the continuous propagation and intensity distribution modeling of the event influence in the spatial topology are realized, and the evolution from event location to event spatial influence distribution is realized.
[0182] In one exemplary embodiment, the specific execution process of step S410 further includes: For each node along the propagation path, the spatial influence range of the station event entity is determined based on the propagation intensity, and the station spatial nodes located within the spatial influence range should update their operational status.
[0183] In this exemplary embodiment, for each node on the propagation path, the spatial influence range of the station event entity is determined according to the propagation intensity, and the operating status of the station spatial nodes is selectively updated based on the spatial influence range.
[0184] Specifically, the spatial influence range of an event is determined based on the magnitude of its propagation intensity. This spatial influence range is represented in the spatial topology as a set of discrete but connected node subgraphs. The spatial influence range includes the nodes effectively affected by the event, allowing for selective updates to the operational state, avoiding globally indiscriminate updates, improving efficiency, and enhancing the accuracy of the results.
[0185] This exemplary embodiment constructs a propagation intensity-driven adaptive spatial range determination mechanism, which differs from existing methods for defining the influence range based on a fixed radius or predefined area. It dynamically determines the spatial influence range through propagation intensity, enabling the influence range to adapt to changes in event characteristics.
[0186] See Figure 10 , Figure 10 This is a flowchart illustrating a multimodal data fusion method for subway stations according to another exemplary embodiment.
[0187] The multimodal data fusion method for subway stations provided in this application includes: Step S510: For the constructible station event entity, initialize the event representation data of each event to the event node to form a node set, the node set includes the event nodes mapped to the station event entity; Step S520: Using event nodes as graph nodes, calculate the evolutionary relationship between nodes to determine the node pairs where the evolutionary relationship holds true; Step S530: Construct evolutionary edges between node pairs and assign evolutionary weights to the constructed evolutionary edges by calculating the evolutionary weights between them to form an event evolution graph. The event evolution graph is used to represent the propagation and evolutionary relationships of station event entities.
[0188] In one exemplary embodiment, evolutionary relation computation includes consistency computation of the associated candidate event set, as well as temporal, spatial propagation, and event semantic relation computation.
[0189] In this exemplary embodiment, the evolutionary relationships between multiple station event entities are further modeled using the representational data of the station event entities, and the propagation paths and evolutionary processes between events are described by constructing an event evolution diagram.
[0190] First, for the constructed station event entities, each event representation data is mapped to event nodes in a graph structure, forming a node set. Specifically, each station event entity is abstracted into an event node, which at least includes information such as event time, spatial location, event type, state characteristics, and a set of associated candidate events. The event representation data is structured and encapsulated so that it can participate in subsequent relation calculations as a graph node. All event nodes are aggregated to form a node set, providing a foundation for subsequent evolutionary relation calculations. Through this process, the transformation from "event entity representation" to "graph node representation" is achieved.
[0191] Using event nodes as graph nodes, perform evolutionary relationship calculations on any pair of nodes in the node set to determine whether there are any node pairs with valid evolutionary relationships.
[0192] In an exemplary embodiment, the calculation of the evolution relationship includes: calculating the set similarity between nodes for the set of associated candidate events contained in the station event entity, and determining that the candidate sets of two station event entities are consistent based on the calculated set similarity.
[0193] Once node pairs with evolutionary relationships are identified, evolutionary edges between the nodes are constructed, and evolutionary weights are calculated to form a complete event evolution graph. Specifically, for node pairs that satisfy evolutionary relationships, directed edges are established between the two nodes, with the direction of the edge representing the direction of event evolution (e.g., from the earlier event to the later event); then, the weights of the evolutionary edges are calculated based on the multidimensional relationship characteristics between the node pairs. In one embodiment, the factors determining the weights include: time interval (the closer the time, the higher the weight); spatial propagation strength (the stronger the propagation influence, the higher the weight); semantic association degree; and the consistency degree of the candidate event set.
[0194] Finally, all event nodes and their corresponding evolution edges are integrated to form an event evolution graph, which is used to describe the propagation paths and evolutionary relationships between multiple station events.
[0195] Based on the constructed event evolution graph, the subway station multimodal data fusion method of this application embodiment further includes: The event evolution path is obtained by searching the event evolution graph and execution graph in the event evolution graph, based on the current station event entity and corresponding event representation data obtained from the current real-time multimodal data stream. The event evolution relationship indicated by the event evolution path drives the prediction of future events and generates subway station intervention decisions under the predicted event trends.
[0196] In other words, after completing the construction of the event evolution graph, we further conduct graph search and path reasoning on the event evolution graph based on the current station events driven by real-time multimodal data, thereby realizing the prediction of future event trends and the generation of intervention decisions.
[0197] For the station event entities and their corresponding event representation data obtained from the current real-time multimodal data stream, node matching and localization are first performed in the constructed event evolution graph. Specifically, the current station event entity is converted into the corresponding event node representation; based on event time, spatial nodes, semantic features, and state features, the most similar or corresponding historical event nodes are matched in the event evolution graph; in this implementation, node alignment can be completed through multi-feature similarity calculation, thereby embedding the current event into the existing evolution graph structure.
[0198] After locating the current event node, a graph search is performed in the event evolution graph, starting from that node, to obtain possible event evolution paths; based on the obtained event evolution paths, future possible events are predicted.
[0199] Specifically, subsequent nodes in the path following the current node are designated as predicted event nodes; the time, spatial location, and event type information corresponding to the predicted event nodes are extracted; in one implementation, the evolution paths of different events are probabilistically processed by combining path weights, thereby obtaining multiple possible events and their probabilities of occurrence. Through this process, the deduction from the current event to future events is realized.
[0200] Finally, based on the predicted event trends, corresponding subway station intervention decisions are generated. These decisions are then used to guide station operation and management by analyzing different event evolution paths predicted for future events.
[0201] Therefore, unlike prediction methods based on a single time series or statistical model, this embodiment uses an event evolution graph for path-level reasoning, enabling the prediction results to reflect the structural relationships between events; and dynamically embeds real-time events into the existing evolution graph to achieve the fusion of historical experience and current state, thereby improving prediction accuracy.
[0202] In a specific instance, such as Figure 11 The interface shows the stations along the Xicheng Line. The subway stations are equipped with various devices, specifically as follows: Figure 11 The right side of the interface shows the turnstile equipment, and the left side shows the selected equipment types. Figure 11 The part highlighted on the left, such as Figure 12 As shown, various access control devices, video surveillance devices, perimeter alarm devices, station door devices, warehousing and logistics devices, elevator devices, etc. are deployed inside subway stations. Therefore, through the embodiments of this application, multimodal data fusion is achieved on data from multiple devices in the subway station to obtain structured event representation data that integrates data from multiple devices, thereby comprehensively, accurately and realistically depicting the station's operating status dynamically.
[0203] In one exemplary embodiment, this application also provides a computer device including a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to implement the steps of the method as described above.
[0204] In one exemplary embodiment, this application also provides a computer program product including a computer program that, when executed by a processor, implements the steps of the method as described above.
[0205] In one exemplary embodiment, this application also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method as described above.
[0206] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this application.
[0207] In an exemplary embodiment of this application, a computer-readable storage medium is also provided, on which computer-readable instructions are stored, which, when executed by a computer's processor, cause the computer to perform the methods described in the above method embodiments.
[0208] According to one embodiment of this application, a program product for implementing the methods in the above-described method embodiments is also provided. This product may employ a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of this invention is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0209] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0210] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0211] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, or any suitable combination thereof.
[0212] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0213] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0214] Furthermore, although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0215] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, and includes several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the method according to the embodiments of this application.
[0216] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the appended claims.
Claims
1. A method for multimodal data fusion in subway stations, characterized in that, The method includes: By pulling real-time multimodal data from subway stations, we continuously acquire multimodal data of the actual data environment of the stations. The multimodal data includes station operation status description data of at least one modality, and the modality corresponds to different types of equipment in the station. Intramodal event detection is performed on the multimodal data, and corresponding candidate events are obtained from the site operation status description data of each modality. The candidate events are the changes and / or abnormal patterns of site operation status represented by the site operation status description data of the corresponding modality. For the obtained candidate events, an event feature vector is constructed by extracting features from the site operation status description data. The event feature vector includes time features, spatial features, semantic features, and confidence features. Based on the time features in the event feature vector, time window matching is performed on candidate events to determine several time-related candidate events. The matching step includes: for candidate events not involved in time association, selecting event anchors to obtain the event anchors corresponding to the current round of time window matching; the event anchors are the earliest or highest-confidence candidate events; based on the time features of the event anchors, and with compensation based on the modal response latency corresponding to the data source device, adaptively determining the current round of matching time window; for the event anchors, performing cross-modal candidate event association search within the determined time window to obtain candidate events that are time-related across modalities with the event anchors; the time-related candidate events are locked and will not enter subsequent time window matching. For several candidate events that are temporally related, association calculations are performed using spatial features, semantic features, and confidence features to determine candidate events that are aggregated into the same event cluster; For the event cluster, a station event entity is constructed across modalities. The station event entity includes the event occurrence time, the spatial node of the event occurrence, the event type, and a set of associated candidate events. Using the station event entity as the fusion center, the station operation status description data corresponding to each candidate event mapped to the station event entity is subjected to structured aggregation processing to form the event representation data of the station event entity.
2. The method according to claim 1, characterized in that, The method of continuously acquiring multimodal data of the actual data environment of the subway station by pulling real-time multimodal data streams includes: When acquiring a mode of station operation status description data, the corresponding spatial node is obtained according to the mapping of the data source device in the station spatial structure, and is attached to the station operation status description data as spatial sensing data. The spatial node represented by the spatial sensing data indicates the location of the station operation status description data in the station spatial structure, supporting cross-modal event association of subsequent candidate events and spatial positioning of associated station event entities.
3. The method according to claim 1, characterized in that, The process of performing intra-modal event detection on the multimodal data involves obtaining corresponding candidate events from the site operation status description data of each modality, including: For the station operation status description data of each mode, perform station operation status numerical conversion of the corresponding mode, and obtain the numerical description sequence of the subway station by mode; Perform sliding window detection on the numerical description sequence to obtain event units corresponding to continuous time segments; For the event unit, candidate events with temporal and spatial semantic attributes are generated by configuring additional spatial nodes in the site operation status description data.
4. The method according to claim 1, characterized in that, The process involves performing association calculations on several temporally correlated candidate events using spatial features, semantic features, and confidence features to determine candidate events that aggregate into the same event cluster, including: For several candidate events that are temporally correlated, the spatial distance between the candidate events and the event anchor point is calculated based on spatial characteristics and the topological structure indicated by the spatial structure of the station. Based on the spatial proximity relationship determined by the spatial distance for the candidate events, candidate events occurring in the same or adjacent regions of the event anchor points are identified as candidate events of the same event cluster.
5. The method according to claim 4, characterized in that, The process of determining candidate events clustered into the same event cluster by performing association calculations on several temporally correlated candidate events using spatial features, semantic features, and confidence features, further includes: For the event clusters formed by the event anchor and the candidate events that have spatial proximity, the semantic similarity of the candidate events is further filtered based on the semantic features, using the event anchor as a benchmark, and the semantically related candidate events are retained in the event clusters.
6. The method according to claim 5, characterized in that, The process of determining candidate events clustered into the same event cluster by performing association calculations on several temporally correlated candidate events using spatial features, semantic features, and confidence features, further includes: The association relationship of the candidate events is weighted and evaluated based on the spatial consistency value, semantic similarity, and confidence features of the spatial distance mapping calculated from the candidate events in the event cluster. Based on the weighted scores obtained from the correlation relationships, candidate events that are aggregated into the same event cluster are determined, and the determined candidate events and the event anchor point constitute an event cluster.
7. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-6.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Emergency event handling auxiliary method and system for urban rail transit driving
CN120450209A
System and method for identifying related events in a resource network monitoring system
US8341106B1