Hydro-junction multi-source heterogeneous video intelligent analysis and situation awareness method and system

By constructing a semantic knowledge graph and a multimodal feature fusion network in the field of water conservancy hubs, the problem of identification and early warning of multi-source heterogeneous video data of water conservancy hubs was solved, and efficient target detection and situational awareness were achieved.

CN121838019AInactive Publication Date: 2026-04-10BEIJING KAIDAO ENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING KAIDAO ENG TECH CO LTD
Filing Date
2026-01-16
Publication Date
2026-04-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies lack deep integration with expertise in the field of water conservancy projects, making it difficult to accurately identify specific targets under complex lighting and severe weather conditions using multi-source heterogeneous video data. Furthermore, they lack the ability to correlate and analyze information across video sources, affecting the accuracy and foresight of early warnings.

Method used

We construct a semantic knowledge graph in the field of water conservancy hubs, perform semantic parsing and spatiotemporal benchmark alignment of multi-source heterogeneous video stream data, use a multimodal feature fusion network with attention mechanism to perform cross-modal feature extraction and target detection, and combine causal link analysis to generate situational awareness results.

Benefits of technology

It achieves high-accuracy identification and stability detection of targets within the monitoring area of ​​water conservancy hubs, improves the continuity of target tracking and the accuracy of early warning in complex environments, and dynamically adjusts the video stream weights to adapt to different conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838019A_ABST
    Figure CN121838019A_ABST
Patent Text Reader

Abstract

The invention provides a hydro-junction multi-source heterogeneous video intelligent analysis and situation awareness method and system, and relates to the technical field of hydro-junction safety monitoring, and the method comprises the steps: carrying out the semantic analysis of a video stream through constructing a semantic knowledge graph, achieving the time-space reference alignment, extracting a cross-modal feature through employing a multi-modal feature fusion network of an attention mechanism, and carrying out the monitoring of a hydro-junction multi-source heterogeneous video. And performing target detection and semantic segmentation, and performing causal link analysis based on the knowledge graph enhanced target entity library. According to the method, the collaborative analysis capability of the multi-source video can be improved, the identification accuracy of the monitored object in a complex scene is enhanced, and accurate perception and prediction of the hydro-junction situation are realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of water conservancy hub safety monitoring, and in particular to a water conservancy hub multi-source heterogeneous video intelligent analysis and situation awareness method and system. BACKGROUND

[0002] Water conservancy hub is an important national infrastructure, which has strategic significance for flood control, drought resistance, water resource allocation and water and electricity energy supply. With the wide application of video monitoring technology, a large amount of multi-source heterogeneous video data is generated in the water conservancy hub monitoring system, including different types of video sources such as ordinary optical cameras, infrared thermal imaging, unmanned aerial vehicle aerial photography and underwater cameras. These video data have the characteristics of strong real-time, large data volume and various formats, which provide a data basis for comprehensive monitoring of the safe operation state of the water conservancy hub. Traditional video monitoring mainly relies on manual duty, which is difficult to efficiently process massive video data and cannot realize intelligent situation awareness and risk warning. With the development of artificial intelligence technology, video intelligent analysis technology based on deep learning provides a new technical path for the safety monitoring of water conservancy hub.

[0003] The prior art lacks deep integration of professional knowledge in the field of water conservancy hub, resulting in low accuracy of identifying specific targets such as professional equipment and abnormal states, especially under complex lighting and severe weather conditions, it is difficult to accurately analyze the professional semantics of video content. Multi-source heterogeneous video data lacks a unified reference alignment method in time and space, and the data collected by different video sources are not synchronized in time and unified in spatial reference system, which makes it difficult to realize cross-video source information correlation analysis. The existing situation awareness methods mostly use rule-based judgment logic, lack deep mining and analysis ability of event causal relationship, and cannot effectively identify potential risk factors and their evolution paths in complex situations, thereby affecting the accuracy and forward-looking of the warning. SUMMARY

[0004] The embodiments of the present application provide a water conservancy hub multi-source heterogeneous video intelligent analysis and situation awareness method and system, which can solve the problems in the prior art.

[0005] In a first aspect, the embodiments of the present application provide a water conservancy hub multi-source heterogeneous video intelligent analysis and situation awareness method, comprising:

[0006] Constructing a semantic knowledge graph in the field of water conservancy hub, and performing semantic analysis on multi-source heterogeneous video stream data, mapping the visual features in the video frames to the semantic knowledge graph to obtain semantic annotated video stream data;

[0007] Performing spatio-temporal reference alignment on the semantic annotated video stream data, establishing a unified spatio-temporal coordinate system based on the spatial position relationship and timestamp information of each video stream to obtain spatio-temporally aligned semantic video stream data;

[0008] The multi-modal feature fusion network based on an attention mechanism performs cross-modal feature extraction on the spatio-temporally aligned semantic video stream data, learns the contribution of different video streams to target identification through adaptive learning, generates multi-scale fusion features, and combines the target feature set;

[0009] Target detection and semantic segmentation are performed according to the target feature set, the monitoring objects and their spatial position information in the water conservancy hub monitoring area are identified, and are reversely updated to the semantic knowledge graph to obtain a knowledge graph enhanced target entity library;

[0010] Based on the causal relationship edges in the knowledge graph enhanced target entity library and the semantic knowledge graph, the event sequence in the water conservancy hub monitoring area is analyzed for causal link, the causal driving factors of situation evolution are identified, and the prediction path of potential risk scenarios is generated through counterfactual reasoning to obtain a causal correlation situation awareness result.

[0011] A semantic knowledge graph in the field of water conservancy hubs is constructed, and multi-source heterogeneous video stream data is semantically analyzed, visual features in video frames are mapped to the semantic knowledge graph, and semantic annotated video stream data is obtained, including:

[0012] The professional knowledge documents and operation rule texts in the field of water conservancy hubs are subjected to entity extraction, relation extraction and attribute extraction, and are respectively organized into entity node layers, attribute node layers and relation edge layers to obtain a hierarchical semantic knowledge graph;

[0013] Local visual features and global context features of the multi-source heterogeneous video stream data are obtained through multi-scale convolutional feature encoding, and feature dimension alignment and semantic space mapping are performed to convert the visual feature space to the embedding space of the semantic knowledge graph to obtain visual semantic alignment features;

[0014] The semantic similarity between the visual semantic alignment features and the entity node embedding vectors of the semantic knowledge graph is calculated, the water conservancy visual objects detected in the video frames are mapped to the corresponding entity nodes of the semantic knowledge graph, and semantic labels are assigned to the entity nodes according to the mapping confidence to obtain an entity node mapping result;

[0015] According to the entity node mapping result, the relation edges and attribute nodes of the mapped entity nodes in the semantic knowledge graph are extracted, the semantic information of the relation edges and the attribute nodes is attached to the visual objects of the corresponding video frames to form composite annotation information containing entity semantics, attribute semantics and relation semantics, and the semantic annotated video stream data is obtained.

[0016] spatial and temporal reference alignment is performed on the semantic annotated video stream data, a unified spatial and temporal coordinate system is established based on the spatial position relationship and timestamp information of each video stream, and semantic video stream data after spatial and temporal alignment is obtained, including:

[0017] Based on the camera spatial position parameters and field of view angle parameters of each video stream in the semantic annotated video stream data, a video stream spatial relationship description is constructed.

[0018] Based on the video stream spatial relationship description, overlapping spatial regions with field of view overlap between different video streams are identified, spatial reference markers commonly visible in the overlapping spatial regions are extracted, and projection transformation matrices between pixel coordinates of the spatial reference markers in each video stream image coordinate system and world coordinates in the world coordinate system are calculated, to obtain spatial alignment transformation parameters.

[0019] The clock deviation and frame rate difference of the timestamp information in the semantic annotated video stream data are analyzed, a time synchronization mapping relationship is constructed, the local timestamp of each video stream is mapped to a unified time reference, and each video stream is time-aligned and interpolated based on the unified time reference, to obtain time alignment parameters.

[0020] Based on the spatial alignment transformation parameters and the time alignment parameters, a unified spatial and temporal coordinate system is established, based on the unified spatial and temporal coordinate system, the video frames of the semantic annotated video stream data are converted to the world coordinate system through the projection transformation matrix, and are aligned to the unified time reference through the time synchronization mapping relationship, to obtain the semantic video stream data after spatial and temporal alignment.

[0021] Based on the attention mechanism based multi-modal feature fusion network, cross-modal feature extraction is performed on the semantic video stream data after spatial and temporal alignment, the contribution of different video streams to target identification is adaptively learned, multi-scale fusion features are generated, and a target feature set is obtained, including:

[0022] Each video stream in the semantic video stream data after spatial and temporal alignment is respectively feature-encoded, visual features of each video stream at different spatial scales are extracted through a convolution feature extraction network, and multi-scale visual features are obtained.

[0023] The semantic annotation information in the semantic video stream data after spatial and temporal alignment is encoded into a semantic embedding vector, and is spliced and fused with the multi-scale visual features in the feature dimension, to obtain multi-scale features enhanced by semantics.

[0024] In the multimodal feature fusion network, attention weights between features at different spatial scales within a single video stream are calculated, and the semantically enhanced multi-scale features are weighted and fused to obtain intramodal fusion features; the contribution of different video streams to target recognition is calculated, cross-modal attention weights are assigned according to the contribution, and the intramodal fusion features of each video stream are adaptively weighted and aggregated to obtain cross-modal fusion features;

[0025] A multi-scale feature pyramid is constructed for the cross-modal fusion features. Feature levels of different resolutions are generated through upsampling and downsampling operations. Local features and global features are extracted at each feature level. The local features and global features are combined to form multi-scale fusion features, thereby obtaining the target feature set.

[0026] Based on the target feature set, target detection and semantic segmentation are performed to identify the monitoring objects and their spatial location information within the water conservancy hub monitoring area. This information is then updated back to the semantic knowledge graph to obtain a knowledge graph-enhanced target entity library, including:

[0027] A target detection network is constructed based on the target feature set. The target detection network generates candidate target regions at each feature level of the multi-scale feature pyramid. Boundary box regression and target classification are performed on the candidate target regions to identify the target detection results within the water conservancy hub monitoring area.

[0028] A semantic segmentation network is constructed based on the target feature set. The semantic segmentation network is used to perform pixel-level semantic annotation on the monitored objects in the target detection results. The pixels of the monitored objects are assigned to the corresponding semantic categories to obtain a semantic segmentation mask.

[0029] By combining the target detection results with the semantic segmentation mask, the spatial location information and timestamp information of the monitored object are obtained, and the spatiotemporally located monitored object entity is obtained. By semantically matching the spatiotemporally located monitored object entity with the semantic knowledge graph and calculating the semantic similarity, the category to which the monitored object entity belongs in the semantic knowledge graph is determined.

[0030] Based on the classification, the association between the monitored object entity and the existing entities in the semantic knowledge graph is constructed. The monitored object entity is inserted into the semantic knowledge graph as a new entity node and relation edges are created. The topology of the semantic knowledge graph is updated to obtain the target entity library enhanced by the knowledge graph.

[0031] Based on the enhanced target entity database and the causal relationship edges in the semantic knowledge graph, causal link analysis of event sequences within the monitoring area of ​​the water conservancy hub includes:

[0032] sequencing the monitoring object entities in chronological order to build a time sequence of entities;

[0033] based on the time interval and spatial position relationship between adjacent monitoring object entities in the time sequence of entities, identifying a combination of monitoring object entities with spatio-temporal correlation, mapping the monitoring object entities in the combination of monitoring object entities and their semantic annotation information into event representations to obtain an event sequence;

[0034] retrieving event entity nodes corresponding to events in the event sequence in the semantic knowledge graph, traversing the causal relationship edges between the event entity nodes, judging the causal dependency relationship between events according to the directionality and semantic type of the causal relationship edges, and building a causal link graph of the event sequence;

[0035] calculating the in-degree and out-degree of each event node in the causal link graph, identifying an event node with zero in-degree and greater than zero out-degree as a starting event of a causal link, identifying an event node with zero out-degree and greater than zero in-degree as a result event of a causal link, and identifying event nodes on a causal propagation path from the starting event to the result event through backtracking.

[0036] identifying causal driving factors of situation evolution, and generating a prediction path of a potential risk scenario through counterfactual reasoning to obtain a situation awareness result of causal correlation, including:

[0037] based on the causal relationship edges in the causal link graph and the event nodes on the causal propagation path, building a causal reasoning rule set of situation evolution;

[0038] for the causal driving factors in the causal reasoning rule set, constructing a counterfactual condition, substituting the counterfactual condition into the causal reasoning rule set for reasoning calculation, propagating the influence of the counterfactual condition on subsequent events along the causal relationship edges, and obtaining a potential risk scenario prediction path of counterfactual reasoning;

[0039] according to the risk attribute annotation of each event in the event sequence in the semantic knowledge graph and the causal action intensity of the causal relationship edges between events, calculating the comprehensive risk score of the potential risk scenario prediction path, and identifying a potential risk scenario prediction path with a comprehensive risk score exceeding a preset threshold as a high-risk prediction path;

[0040] associating and organizing the causal driving factors, the causal link graph, the potential risk scenario prediction path, and the high-risk prediction path, building a situation awareness knowledge structure containing the current situation state, causal evolution mechanism, and future risk prediction, and obtaining the situation awareness result of causal correlation in the multi-source heterogeneous video of the water conservancy hub.

[0041] In a second aspect of the embodiment of the present application, a water conservancy hub multi-source heterogeneous video intelligent analysis and situation awareness system is provided, comprising:

[0042] A first unit is configured to construct a semantic knowledge graph in the field of water conservancy hubs, and perform semantic analysis on multi-source heterogeneous video stream data, map visual features in the video frames to the semantic knowledge graph, and obtain semantic annotated video stream data;

[0043] A second unit is configured to perform spatio-temporal reference alignment on the semantic annotated video stream data, establish a unified spatio-temporal coordinate system based on the spatial position relationship and timestamp information of each video stream, and obtain spatio-temporally aligned semantic video stream data;

[0044] A third unit is configured to perform cross-modal feature extraction on the spatio-temporally aligned semantic video stream data based on a multi-modal feature fusion network of an attention mechanism, learn the contribution of different video streams to target identification through adaptive learning, generate multi-scale fusion features, and combine the multi-scale fusion features to obtain a target feature set;

[0045] A fourth unit is configured to perform target detection and semantic segmentation based on the target feature set, identify monitoring objects and their spatial position information in the monitoring area of the water conservancy hub, and update the monitoring objects and their spatial position information to the semantic knowledge graph in reverse to obtain a knowledge graph enhanced target entity library;

[0046] A fifth unit is configured to perform causal link analysis on an event sequence in the monitoring area of the water conservancy hub based on a causal relationship edge in the knowledge graph enhanced target entity library and the semantic knowledge graph, identify causal driving factors of situation evolution, and generate a prediction path of a potential risk scenario through counterfactual reasoning to obtain a situation awareness result of causal correlation.

[0047] In a third aspect of the embodiment of the present application,

[0048] An electronic device is provided, comprising:

[0049] a processor;

[0050] a memory for storing processor executable instructions;

[0051] The processor is configured to invoke the instructions stored in the memory to execute the method described above.

[0052] In a fourth aspect of the embodiment of the present application,

[0053] A computer readable storage medium is provided, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the method described above.

[0054] The beneficial effects of the present application are as follows:

[0055] By constructing a semantic knowledge graph special for the field of water conservancy hubs and mapping video features to the knowledge graph, deep semantic understanding of the video content is achieved, overcoming the problem of insufficient dependence on domain knowledge in traditional video analysis methods and improving the recognition accuracy of visual objects in a specific field. Through the spatiotemporal reference alignment technology, a unified spatiotemporal coordinate system is established, solving the problem of heterogeneity between different monitoring cameras and different types of sensor video streams, effectively integrating multi-angle and multi-view information, and improving the continuity and accuracy of target tracking. The multi-modal feature fusion network based on the attention mechanism can automatically learn the contribution of different video streams to target recognition, dynamically adjust the weights of each video stream, and improve the target detection performance in complex environments, especially the recognition stability under adverse conditions such as changes in light and weather. BRIEF DESCRIPTION OF DRAWINGS

[0056] Figure 1 A flowchart of the water conservancy hub multi-source heterogeneous video intelligent analysis and situation awareness method of the embodiment of the present application is shown in

[0057] Figure 2 A flowchart of the semantic knowledge graph construction and video semantic annotation is shown in DETAILED DESCRIPTION

[0058] To make the purpose, technical scheme and advantages of the embodiment of the present application clearer, the technical scheme in the embodiment of the present application will be described clearly and completely below in combination with the drawings in the embodiment of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0059] The technical scheme of the present application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in some embodiments.

[0060] Figure 1 A flowchart of the water conservancy hub multi-source heterogeneous video intelligent analysis and situation awareness method of the embodiment of the present application is shown in Figure 1 The method comprises:

[0061] A semantic knowledge graph for the field of water conservancy hubs is constructed, and multi-source heterogeneous video stream data is semantically analyzed, mapping visual features in the video frames to the semantic knowledge graph to obtain semantically annotated video stream data;

[0062] aligning the semantic annotated video stream data in time and space, establishing a unified time and space coordinate system based on the spatial position relationship and timestamp information of each video stream, and obtaining semantic video stream data after time and space alignment;

[0063] a multi-modal feature fusion network based on an attention mechanism, cross-modal feature extraction on the semantic video stream data after time and space alignment, adaptive learning of the contribution of different video streams to target identification, generation of multi-scale fusion features, and combination to obtain a target feature set;

[0064] target detection and semantic segmentation according to the target feature set, identification of monitoring objects and their spatial position information in the water conservancy hub monitoring area, and reverse updating to the semantic knowledge graph to obtain a knowledge graph enhanced target entity library;

[0065] causal link analysis of event sequences in the water conservancy hub monitoring area based on the causal relationship edges in the knowledge graph enhanced target entity library and the semantic knowledge graph, identification of causal driving factors of situation evolution, generation of prediction paths of potential risk scenarios through counterfactual reasoning, and obtaining of situation awareness results of causal association.

[0066] In an optional implementation, a semantic knowledge graph in the field of water conservancy hubs is constructed, and multi-source heterogeneous video stream data is semantically analyzed, visual features in video frames are mapped to the semantic knowledge graph, and semantic annotated video stream data is obtained, including:

[0067] entity extraction, relationship extraction, and attribute extraction are performed on professional knowledge documents and operation rule texts in the field of water conservancy hubs, and are respectively organized into an entity node layer, an attribute node layer, and a relationship edge layer, to obtain a hierarchical semantic knowledge graph;

[0068] local visual features and global context features of the multi-source heterogeneous video stream data are obtained through multi-scale convolution feature encoding, and feature dimension alignment and semantic space mapping are performed, the visual feature space is converted to the embedding space of the semantic knowledge graph, and visual semantic alignment features are obtained;

[0069] the semantic similarity between the visual semantic alignment features and the entity node embedding vectors of the semantic knowledge graph is calculated, the water conservancy visual objects detected in the video frames are mapped to the corresponding entity nodes of the semantic knowledge graph, and the entity nodes are assigned semantic labels according to the mapping confidence, to obtain an entity node mapping result;

[0070] According to the entity node mapping result, a relationship edge and an attribute node of the mapped entity node in the semantic knowledge graph are extracted, semantic information of the relationship edge and the attribute node is attached to a visual object of a corresponding video frame, composite annotation information containing entity semantics, attribute semantics and relationship semantics is formed, and the semantic annotated video stream data is obtained.

[0071] As shown in Figure 2 the method comprises:

[0072] Natural language processing technology is used for entity extraction, relationship extraction and attribute extraction. Professional knowledge documents and operation rule texts in the water conservancy hub field are processed. Entity extraction uses a model combining bidirectional long short-term memory network (Bi-LSTM) and conditional random field (CRF) to identify entities such as professional terms, device names and parameter indicators in the text. Relationship extraction extracts semantic relationships such as "component", "control relationship" and "monitoring indicator" between entities through a pattern matching method based on dependency syntax analysis combined with remote supervision learning. Attribute extraction extracts entity characteristics, numerical parameters and state descriptions through a combination of rule templates and deep learning.

[0073] The extracted entities are organized into an entity node layer, including device entities such as gates, water level gauges and flow monitoring stations, and index entities such as water level, flow and gate opening. Then an attribute node layer is constructed to store attribute information such as device specifications, monitoring thresholds and operation parameters. Finally, a relationship edge layer is established to represent the subordinate relationship, monitoring relationship and control relationship between entities. A graph database is used to store the three layers of structure, and an ontology model is used to define the category hierarchy and relationship constraints to form a semantic knowledge graph in the water conservancy hub field.

[0074] Video frames are preprocessed by target detection algorithms such as YOLOv5 to detect key devices such as water gates, water level gauges and overflow weirs. Multi-scale convolutional neural networks are used to encode the features of the detected target regions to extract local visual features. At the same time, global context features are captured through a Transformer structure, and multi-level features are aggregated through an attention mechanism. For video streams from different sources, the features are normalized through a cross-domain adaptation layer to achieve a unified representation of multi-source video features.

[0075] Visual features are converted to a feature space matching the embedding dimension of the knowledge graph through a fully connected layer. The embedding of the knowledge graph is pre-calculated using methods such as TransE or ComplEx. Through a contrastive learning framework, the distance between visual features and corresponding entity embeddings is minimized, while the distance with unrelated entities is maximized. The visual-semantic alignment network is trained to achieve mapping and conversion of visual features to semantic space.

[0076] The cosine similarity between the visual semantic alignment feature and the embedding vector of each entity node in the knowledge graph is calculated, the top K entity nodes with the highest similarity are selected as the candidate mapping targets, the conditional random field model is used to optimize the overall mapping result based on the temporal consistency constraint and the visual context relationship, and the mapping accuracy is improved. For each visual object, a semantic label is assigned according to the mapping confidence. The mapping with a confidence higher than a threshold (such as 0.85) is confirmed as an effective mapping, and the entity node mapping result is formed.

[0077] The one-hop relationship edge of the mapped entity node in the knowledge graph is queried to obtain associated other entity nodes; the attribute node information of the mapped entity node, such as device state parameters and operation specifications, is extracted; and the semantic information of these relationship edges and attribute nodes is associated with the corresponding visual objects in the video frame to generate composite annotation information containing entity semantics (such as "No. 1 gate"), attribute semantics (such as "opening degree 30%"), and relationship semantics (such as "control downstream water level").

[0078] The composite annotation information is attached to the original video stream to form semantic annotated video stream data. For each frame in the video stream, the original video data, the detected visual object coordinates, the corresponding semantic labels and their confidences, the associated attribute information and relationship information are saved, and these annotation data are organized in JSON format. Such semantic annotated video stream data can be used in water conservancy hub operation state monitoring, abnormal event detection, intelligent decision support, and other application scenarios.

[0079] In actual application, when processing the video stream of a certain reservoir dam video monitoring system, the gate device in the video can be automatically identified and mapped to the "No. 3 flood discharge gate" entity node in the knowledge graph, while the "maximum discharge capacity" attribute and the "control downstream river water level" relationship are associated. This allows monitoring personnel to intuitively understand the semantic information of the device and its role in the entire water conservancy system, thereby more efficiently monitoring and decision-making.

[0080] In an optional implementation, the semantic annotated video stream data is spatio-temporally aligned, a unified spatio-temporal coordinate system is established based on the spatial position relationship and timestamp information of each video stream, and the spatio-temporally aligned semantic video stream data includes:

[0081] Based on the camera spatial position parameters and the field of view angle parameters of each video stream in the semantic annotated video stream data, a video stream spatial relationship description is constructed;

[0082] Based on the spatial relationship description of the video streams, overlapping spatial regions with field-of-view overlap between different video streams are identified, and spatial fiducial markers commonly visible to multiple video streams are extracted within the overlapping spatial regions, and a projection transformation matrix between pixel coordinates of the spatial fiducial markers in an image coordinate system of each video stream and world coordinates in a world coordinate system is calculated to obtain spatial alignment transformation parameters;

[0083] Clock skew and frame rate difference of timestamp information in the semantically annotated video stream data are analyzed, a time synchronization mapping relationship is constructed to map local timestamps of each video stream to a unified time reference, and each video stream is time-aligned and interpolated based on the unified time reference to obtain time alignment parameters;

[0084] Based on the spatial alignment transformation parameters and the time alignment parameters, a unified space-time coordinate system is established, and based on the unified space-time coordinate system, video frames of the semantically annotated video stream data are converted to the world coordinate system through the projection transformation matrix and aligned to the unified time reference through the time synchronization mapping relationship to obtain the space-time aligned semantically annotated video stream data.

[0085] The intrinsic matrix of each camera is obtained, including focal length, principal point coordinates, and distortion coefficients; at the same time, the extrinsic matrix is obtained, including rotation matrix and translation vector, which describes the position and orientation of the camera in the world coordinate system. In addition, the field of view (FOV) information of the camera is also recorded, including horizontal FOV and vertical FOV, to determine the visual range of the camera. Based on the above parameters, the space coverage area model of the video stream is constructed by calculating the view cone of each camera and projecting it into the world coordinate system. For each video stream, a data structure is established to store its ID, spatial position coordinates, field-of-view angle range, and geometric description of the field-of-view coverage area, forming a complete spatial relationship description of the video stream.

[0086] Based on the spatial relationship description of the video streams, overlapping spatial regions with field-of-view overlap between different video streams are identified, and the intersection of multiple view cones is calculated to determine the overlapping regions. In the overlapping region detection, ray projection method or geometric intersection calculation is used to obtain the three-dimensional geometric representation of the potential overlapping spatial region. In the determined overlapping region, spatial fiducial markers commonly visible to multiple video streams are extracted. Feature detection algorithms such as SIFT, SURF, or ORB are used to extract feature points of these markers from each video stream image, and the correspondence of the same marker in different video streams is confirmed through descriptor matching.

[0087] To calculate the spatial alignment transformation parameters, the pixel coordinates of each marker in the image coordinate system of each video stream are determined, and a correspondence between the pixel coordinates and the world coordinates is established by the known positions of the markers in the world coordinate system (which can be obtained by measurement or determined by multi-view geometry reconstruction). Based on these pairs of corresponding points, a projection transformation matrix is calculated, usually solved by the DLT (Direct Linear Transformation) algorithm, and the accuracy of the transformation matrix is improved by the RANSAC method to eliminate outliers. The final projection transformation matrix contains the mapping relationship from the image coordinates to the world coordinates, which is the core parameter of spatial alignment.

[0088] For each video stream, its timestamp sequence is extracted, and the clock bias and frame rate difference are analyzed. By comparing the timestamp differences of different video streams capturing the same event, the clock offset is estimated. In specific implementation, a flash or other synchronization signal can be introduced in the scene as a time reference point to calculate the time difference of each video stream capturing the signal. For the case of inconsistent frame rate, the actual frame rate of each video stream is calculated, and the least common multiple frame rate is determined as the unified sampling reference.

[0089] A video stream is selected as the time reference, and the local timestamps of other video streams are mapped to this reference through linear transformation. The mapping function can be represented as: t_mapped = a x t_local + b, where a is the frame rate proportion coefficient and b is the clock offset. Based on the established mapping relationship, time alignment interpolation processing is performed on each video stream to ensure that there is corresponding frame data at the unified time point. For time points that need to be interpolated, linear interpolation or optical flow method can be used to generate intermediate frames to ensure time continuity and smooth transition.

[0090] Based on the spatial alignment transformation parameters and the time alignment parameters, a unified space-time coordinate system is established, in which the spatial dimension is represented by the world coordinate system and the time dimension is represented by the unified time reference. For each video frame, the semantic annotation elements (such as target bounding box, segmentation mask, etc.) in it are converted to the world coordinate system through the projection transformation matrix, and are aligned to the unified time reference through the time synchronization mapping relationship.

[0091] Through the above steps, the space-time reference alignment of the semantic annotated video stream data is completed, and the semantic video stream data in the unified space-time coordinate system is obtained, laying a foundation for subsequent multi-view scene understanding and analysis.

[0092] In an optional implementation, a multi-modal feature fusion network based on attention mechanism is used to extract cross-modal features from the space-time aligned semantic video stream data, and the contribution of different video streams to target identification is adaptively learned to generate multi-scale fusion features and combine them to obtain a target feature set, including:

[0093] The feature encoding is performed on each video stream in the spatio-temporally aligned semantic video stream data, the visual features of each video stream at different spatial scales are extracted through a convolution feature extraction network, and a multi-scale visual feature is obtained.

[0094] The semantic annotation information in the spatio-temporally aligned semantic video stream data is encoded into a semantic embedding vector, and is spliced and fused with the multi-scale visual features in the feature dimension to obtain a semantic-enhanced multi-scale feature.

[0095] In the multi-modal feature fusion network, the attention weights between different spatial scale features in a single video stream are calculated, and the semantic-enhanced multi-scale features are weighted and fused to obtain intra-modal fusion features; the contribution of different video streams to target recognition is calculated, cross-modal attention weights are assigned according to the contribution, and the intra-modal fusion features of each video stream are adaptively weighted and aggregated to obtain cross-modal fusion features.

[0096] A multi-scale feature pyramid is constructed for the cross-modal fusion features, different resolution feature levels are generated through up-sampling and down-sampling operations, local features and global features are extracted on each feature level, the local features and global features are combined to form multi-scale fusion features, and the target feature set is obtained.

[0097] A multi-layer convolutional neural network is used to construct a feature extraction module to process different modal data such as RGB video streams, optical flow video streams and depth video streams. Specifically, ResNet-50 is used as the backbone network, feature extraction points are set after different convolution layers, and low-level features (conv2_x output), medium-level features (conv3_x output) and high-level features (conv4_x output) are obtained respectively. The low-level features retain the detailed information, the medium-level features express the texture structure, and the high-level features contain the semantic concept. These feature maps have spatial resolutions of 1 / 4, 1 / 8 and 1 / 16 of the original input size, and the channel numbers are 256, 512 and 1024 respectively. Through the above operations, a multi-scale visual feature set is generated for each video stream.

[0098] The semantic annotation information in the spatio-temporally aligned semantic video stream data is encoded into semantic embedding vectors, including target class, action type, scene environment, etc. The text annotation is converted into a fixed-dimensional (such as 300-dimensional) semantic vector using a pre-trained word embedding model (such as Word2Vec or GloVe). For class information, the semantic vector is further mapped to a feature space compatible with visual features through a fully connected layer. The obtained semantic embedding vectors are expanded to the same spatial dimension as the visual features through a spatial broadcast operation, and are concatenated with the multi-scale visual features in the channel dimension to form semantic-enhanced multi-scale features. For example, for a visual feature with a resolution of HxW and a channel number of C, the dimension of the concatenated semantic-enhanced feature is HxWx(C+D), where D is the dimension of the semantic vector.

[0099] In the multi-modal feature fusion network, the attention weights between different spatial scale features within a single video stream are calculated, and a channel attention mechanism is used. The features of each scale are first subjected to global average pooling and maximum pooling to obtain two channel descriptors, which are then added after being processed by a shared multi-layer perceptron (two fully connected layers with ReLU activation in between) and then passing through a Sigmoid function to obtain the channel weights. At the same time, a spatial attention mechanism is used to perform average pooling and maximum pooling on the channel dimension, and then a 7x7 convolution layer and a Sigmoid function are used to obtain the spatial weight map. The channel weights and spatial weights are multiplied and applied to the original features to achieve feature enhancement. The features of different scales are upsampled or downsampled to have the same spatial resolution, and are fused by weighted summation, with the weights being dynamically assigned by learnable parameters to obtain intra-modal fusion features.

[0100] An adaptive contribution evaluation module is designed to receive intra-modal fusion features of each video stream as input. Through global context modeling, the feature map is converted into a channel descriptor using global average pooling, and a normalized contribution score is generated through a two-layer fully connected network (with ReLU activation in between). Cross-modal attention weights are assigned according to these scores, with the weight values reflecting the importance of each modality in the current scene. For example, in a well-lit scene, the RGB stream obtains a higher weight; while in a low-light environment, the weight of the depth stream increases. These weights are applied to the intra-modal fusion features of each video stream to obtain cross-modal fusion features through weighted summation.

[0101] A multi-scale feature pyramid is constructed for the cross-modal fusion features, and a feature pyramid network (FPN) structure is adopted to generate feature levels with different resolutions through a top-down path and lateral connections. In the specific implementation, five levels (P2 to P6) are designed, and the resolution gradually decreases from 1 / 4 of the original input to 1 / 64. At each feature level, a 3x3 convolution is used to extract local features, and an attention pooling operation is used to extract global context features. The local features and global features are combined through a residual connection to form fusion features with multi-scale receptive fields. Finally, these multi-level fusion features collectively constitute the target feature set, providing rich feature representations for subsequent target detection, recognition, or segmentation tasks.

[0102] In an optional implementation, target detection and semantic segmentation are performed based on the target feature set, the spatial location information of the monitoring objects in the water conservancy hub monitoring area is identified, and is updated to the semantic knowledge graph in reverse to obtain a target entity library enhanced by the knowledge graph, including:

[0103] A target detection network is constructed based on the target feature set, candidate target regions are generated on each feature level of the multi-scale feature pyramid through the target detection network, and boundary box regression and target classification are performed on the candidate target regions to identify the target detection results in the water conservancy hub monitoring area;

[0104] A semantic segmentation network is constructed based on the target feature set, and the monitoring objects in the target detection results are pixel-level semantic labeled through the semantic segmentation network, the pixels of the monitoring objects are assigned to the corresponding semantic categories, and a semantic segmentation mask is obtained;

[0105] The spatial location information and timestamp information of the monitoring objects are obtained by combining the target detection results and the semantic segmentation mask, the spatiotemporal positioning monitoring object entity is obtained, the semantic matching and semantic similarity calculation are performed between the spatiotemporal positioning monitoring object entity and the semantic knowledge graph to determine the belonging category of the monitoring object entity in the semantic knowledge graph;

[0106] Based on the belonging category, the association relationship between the monitoring object entity and the existing entities in the semantic knowledge graph is constructed, the monitoring object entity is inserted into the semantic knowledge graph as a new entity node and a relationship edge is created, the topology structure of the semantic knowledge graph is updated, and the target entity library enhanced by the knowledge graph is obtained.

[0107] The target detection network is constructed based on the target feature set, ResNet-50 is used as the backbone network, and multi-level features are extracted from the input image through the feature extraction module. On different levels of the multi-scale feature pyramid, an anchor generator is set to generate anchor boxes with different sizes and aspect ratios to adapt to different sizes of monitoring objects. For special targets in the water conservancy hub environment, such as water gates, ships, personnel, etc., anchor box parameters are customized to improve detection accuracy. The target detection network generates candidate regions on each feature level through the region proposal network (RPN), and then filters overlapping boxes through non-maximum suppression (NMS) with a threshold of 0.5. The retained candidate regions are converted to fixed-size feature maps through the ROI pooling layer, and then through the fully connected layer for bounding box regression and target classification, outputting detection results including bounding box coordinates, classification labels and confidence.

[0108] The semantic segmentation network is constructed based on the target feature set, FCN (Fully Convolutional Network) architecture is used, and the receptive field is increased by combining with the dilated convolution to retain more spatial detail information. In the encoder part, ResNet-50 with shared weights is used as a feature extractor to maintain consistency with the features of the target detection network, facilitating subsequent fusion. In the decoder part, the spatial resolution is gradually restored through transposed convolution, and the jump connection is used to fuse features at different levels to alleviate the problem of feature loss. For each monitoring object in the target detection result, the corresponding region is extracted for pixel-level semantic labeling, and the pixels are assigned to predefined semantic categories such as "water gate", "spillway", "patrol personnel", etc. The probability of each pixel belonging to each category is calculated through the softmax function to generate a semantic segmentation mask.

[0109] The spatial location information and timestamp information of the monitoring object are obtained by combining the target detection result and the semantic segmentation mask. The spatial location information includes image coordinates (x, y), width and height (w, h), and the outline point set of the segmentation mask. The timestamp information is extracted from the video frame or image metadata, and the format is "YYYY-MM-DD HH:MM:SS". The image coordinates are converted to actual world coordinates through the camera calibration parameters to establish a spatiotemporal positioning model of the monitoring object. The spatiotemporal positioning of the monitoring object entity is semantically matched with the semantic knowledge graph, and the similarity calculation method based on word embedding is used to vectorize the target category and the entity type in the knowledge graph, and the semantic relevance is calculated through the cosine similarity. The similarity threshold is set to 0.75, and the matching result exceeding the threshold is determined as the attribution category of the monitoring object entity in the semantic knowledge graph.

[0110] According to the spatial position relationship, time sequence relationship and functional dependence relationship and other dimensions, relationship templates such as "located in", "adjacent to", "belongs to" and the like are created. For a detected new entity, such as a specific sluice gate, a "component part" relationship with a superior facility (such as a sluice), a "located in" relationship with a monitoring area, and a "being operated" relationship with other entities (such as a patrol personnel) appearing at the same time are constructed. A unique identifier is assigned to each new entity node, attribute information (type, position, time, etc.) of the new entity node is stored, a relationship edge with a related entity is created, and the topology of the semantic knowledge graph is updated. The storage and query of the knowledge graph are realized through a graph database management system (such as Neo4j), and subsequent intelligent analysis and decision-making are supported.

[0111] In an actual application scenario, taking sluice monitoring as an example, when the target detection network identifies an abnormal opening state of a gate, the semantic segmentation network accurately labels the gate contour and opening degree, obtains the real-time position and opening time of the gate, and updates this information as a new entity to the knowledge graph. The knowledge graph correlates historical data and a rule base to infer the correlation between the abnormal opening event and factors such as the upstream water level and gate maintenance records, thereby providing early warning and decision support for management personnel.

[0112] Through the above technical implementation, a knowledge graph enhanced target entity library is constructed, accurate identification, positioning and semantic understanding of a water conservancy hub monitoring object are realized, and strong support is provided for intelligent monitoring and management.

[0113] In an optional implementation, based on the causal relationship edges in the knowledge graph enhanced target entity library and the semantic knowledge graph, the causal link analysis of an event sequence in a water conservancy hub monitoring area includes:

[0114] The monitoring object entities are sorted in chronological order to construct a time sequence entity sequence;

[0115] Based on the time interval and spatial position relationship between adjacent monitoring object entities in the time sequence entity sequence, a monitoring object entity combination with spatiotemporal correlation is identified, the monitoring object entities in the monitoring object entity combination and their semantic annotation information are mapped into event representations, and an event sequence is obtained;

[0116] Event entity nodes corresponding to the events in the event sequence are searched in the semantic knowledge graph, the causal relationship edges between the event entity nodes are traversed, the causal dependence relationship between the events is judged according to the directionality and semantic type of the causal relationship edges, and a causal link graph of the event sequence is constructed;

[0117] In the causal link graph, the in-degree and out-degree of each event node are calculated, the event node with zero in-degree and greater than zero out-degree is identified as the starting event of the causal link, the event node with zero out-degree and greater than zero in-degree is identified as the result event of the causal link, and the event nodes on the causal propagation path from the starting event to the result event are identified by backtracking.

[0118] In practical applications, the water conservancy hub monitoring system obtains target entity information, including personnel, vehicles, equipment, and their activity states, through video monitoring, sensors, and other devices. Each target entity is assigned a timestamp, recording the specific time of its appearance or activity. These entities are sorted according to the timestamps to construct a time-series entity sequence. For example, in a certain reservoir dam monitoring scenario, the following entity activities are recorded in sequence: "9:15 engineering vehicle enters dam area," "9:20 two workers approach the flood discharge gate," "9:25 flood discharge gate opens," and so on, forming a time-series entity sequence in chronological order.

[0119] The spatiotemporal correlation of adjacent entities in the time-series entity sequence is analyzed. The time interval threshold can be set to a reasonable value in a specific scenario, such as 30 minutes in a reservoir flood discharge operation scenario. The spatial position relationship is determined based on the overlap or proximity of entity activity areas, such as the same monitoring area or adjacent functional areas. When the time interval of two entities is less than the threshold and the spatial position is related, it is determined that they have spatiotemporal correlation, and they are combined and mapped as event representations. In the above example, it is determined that the time interval between "workers approach the gate" and "gate opening" is 5 minutes, and they occur in the same spatial area, so they are identified as an associated event combination, which is further mapped to the "manual flood discharge operation" event.

[0120] Event entity nodes are retrieved in the semantic knowledge graph, and a causal link graph is constructed. The semantic knowledge graph pre-stores entity concepts, relationships, and event causal link knowledge in the water conservancy hub field. For the identified event sequence, the corresponding event entity nodes are retrieved in the knowledge graph. For example, "heavy rain weather," "reservoir water level rise," "manual flood discharge operation," and "downstream water level rise" event nodes are found. Subsequently, the causal relationship edges between these event nodes are traversed, the directionality and semantic type (such as "cause," "trigger," "influence," etc.) of the edges are analyzed, and the causal dependency relationship between events is determined. For example, if there is a "cause" type edge from "heavy rain weather" to "reservoir water level rise" in the knowledge graph, the causal relationship between the two events is established. By connecting all related causal relationship edges, a complete event sequence causal link graph is constructed.

[0121] In the causal link graph, identify the key event nodes, calculate the in-degree (number of edges pointing to the node) and out-degree (number of edges from the node) of each event node. The nodes with zero in-degree and greater than zero out-degree are identified as the starting events of the causal link, such as "heavy rain weather"; the nodes with zero out-degree and greater than zero in-degree are identified as the result events of the causal link, such as "downstream water level rise". Through the backtracking algorithm from the starting event to the result event, all event nodes on the causal propagation path are identified, forming a complete causal link. In practical applications, this backtracking process uses depth-first search to access all paths from the starting event, until the result event is reached, and all intermediate event nodes on the path are recorded.

[0122] In the reservoir flood release scenario embodiment, the finally identified causal link is "heavy rain weather → reservoir water level rise → reservoir water level exceeds warning line → manual flood release operation → downstream water level rise", which clearly shows the complete event causal relationship from cause to result, providing a scientific basis for water conservancy management decision-making.

[0123] Through the above causal link analysis method based on knowledge graph enhancement, the water conservancy hub monitoring system can effectively analyze the causal relationship behind the complex event sequence, timely warn potential risks, assist emergency decision-making, and improve the safety operation level of water conservancy hubs.

[0124] In an optional implementation, the causal driving factors of situation evolution are identified, and the prediction path of potential risk scenarios is generated through counterfactual reasoning, and the causal correlation situation awareness result includes:

[0125] Based on the causal relationship edges in the causal link graph and the event nodes on the causal propagation path, a causal reasoning rule set for situation evolution is constructed;

[0126] For the causal driving factors in the causal reasoning rule set, construct a counterfactual condition, substitute the counterfactual condition into the causal reasoning rule set for reasoning calculation, propagate the influence of the counterfactual condition on subsequent events along the causal relationship edges, and obtain the potential risk scenario prediction path of counterfactual reasoning;

[0127] According to the risk attribute annotation of each event in the event sequence in the semantic knowledge graph and the causal action strength of the causal relationship edges between events, calculate the comprehensive risk score of the potential risk scenario prediction path, and identify the potential risk scenario prediction path with a comprehensive risk score exceeding a preset threshold as a high-risk prediction path;

[0128] The causal driving factors, the causal chain link diagram, the potential risk scenario prediction path, and the high-risk prediction path are associated and organized to build a situation awareness knowledge structure containing the current situation state, the causal evolution mechanism, and the future risk prediction, and a situation awareness result of the causal correlation in the multi-source heterogeneous video of the water conservancy hub is obtained.

[0129] The causal relationship edges in the causal chain link diagram and the event nodes on the causal propagation path are extracted. For example, in the multi-source heterogeneous video monitoring of the water conservancy hub, if it is found that there is a causal relationship edge between the "abnormal vibration of the gate" event node and the "abnormal fluctuation of the water level" event node, a rule can be formed: IF abnormal vibration of the gate THEN abnormal fluctuation of the water level (confidence: 0.85). For complex scenarios, the rule can be extended to include multiple conditions: IF abnormal vibration of the gate AND excessive flow rate THEN abnormal fluctuation of the water level AND cause equipment damage (confidence: 0.92). These rules are organized into a directed acyclic graph structure, facilitating reasoning and propagation along the causal relationship.

[0130] For the key causal driving factors in the constructed causal reasoning rule set, a counterfactual condition is constructed for reasoning analysis. The counterfactual condition is a hypothetical change to the real state, for example: "What would happen if the gate vibration frequency increased by 30%?" When constructing the counterfactual condition, the reasonable parameter variation range needs to be determined according to the domain knowledge. For the water conservancy hub scenario, counterfactual conditions can be set for key factors such as "equipment state", "environmental condition", and "operation behavior", such as "abnormal increase in gate opening", "sudden increase in upstream water level", and "monitoring system failure".

[0131] The constructed counterfactual condition is substituted into the causal reasoning rule set, and the influence is propagated step by step along the causal relationship edge through the forward chain reasoning method. For example, when the counterfactual condition of "abnormal increase in gate opening" is input, according to the rule set, the propagation path of "rapid rise of downstream water level" → "increase in flood control pressure" → "possible triggering of downstream regional flood risk" can be derived. During the reasoning process, the state change and the degree of influence at each step are recorded, and finally a complete potential risk scenario prediction path is formed.

[0132] To evaluate the severity of potential risk scenarios, the comprehensive risk score is calculated according to the risk attribute annotation of each event in the event sequence in the semantic knowledge graph and the causal action strength of the causal relationship edge between events. The calculation method is as follows: the risk level (such as low, medium, high, corresponding to quantized values 1, 2, 3 respectively) of each event node on the predicted path is obtained, combined with the action strength (value range 0-1) of the causal relationship edge, and the comprehensive risk score of the path is calculated by using the weighted cumulative method. For example, for the path "gate failure → water flow anomaly → water level rise → embankment pressure increase", if the risk level of each node is [2, 1, 2, 3] and the strength of each causal edge is [0.8, 0.7, 0.9], the comprehensive risk score is calculated as: 2*0.8 + 1*0.7 + 2*0.9 + 3 = 6.9.

[0133] Set an appropriate risk threshold (such as 6.5), and mark the predicted path with a comprehensive risk score exceeding the threshold as a high-risk predicted path. In the water conservancy hub monitoring scene, the high-risk predicted path is usually associated with device serious failure, water level anomaly under extreme weather conditions, key monitoring system failure, etc.

[0134] The identified causal driving factors, causal link graphs, potential risk scenario prediction paths and high-risk prediction paths are associated and organized to construct a situation awareness knowledge structure. Specifically, a multi-level knowledge representation framework is established, including: a basic layer (recording the currently monitored event and state data), a causal association layer (containing causal driving factors and causal link graphs), and a risk prediction layer (containing potential risk scenario prediction paths and high-risk prediction paths). The three layers of information are associated through metadata such as event ID, timestamp, and spatial location to form a complete situation awareness knowledge system.

[0135] Through the above causal association situation awareness method, the potential risks in the multi-source heterogeneous video of the water conservancy hub can be effectively identified, not only showing "what happened", but also revealing "why it happened" and "what consequences it may cause", providing strong support for the safe operation and management of water conservancy facilities.

[0136] The water conservancy hub multi-source heterogeneous video intelligent analysis and situation awareness system of the embodiment of the application comprises:

[0137] A first unit is configured to construct a semantic knowledge graph in the field of water conservancy hubs, and perform semantic analysis on multi-source heterogeneous video stream data, map visual features in video frames to the semantic knowledge graph, and obtain semantic annotated video stream data;

[0138] A second unit is configured to perform spatio-temporal reference alignment on the semantic annotated video stream data, establish a unified spatio-temporal coordinate system based on the spatial position relationship and timestamp information of each video stream, and obtain spatio-temporally aligned semantic video stream data.

[0139] The third unit is configured to perform cross-modal feature extraction on the spatio-temporally aligned semantic video stream data based on an attention mechanism-based multi-modal feature fusion network, to generate a multi-scale fusion feature by adaptively learning the contribution of different video streams to target identification, and to combine the multi-scale fusion feature to obtain a target feature set;

[0140] The fourth unit is configured to perform target detection and semantic segmentation based on the target feature set, to identify a monitoring object and spatial position information of the monitoring object in the water conservancy hub monitoring area, and to update the semantic knowledge graph in reverse to obtain a knowledge graph enhanced target entity library;

[0141] The fifth unit is configured to perform causal link analysis on an event sequence in the water conservancy hub monitoring area based on a causal relationship edge in the knowledge graph enhanced target entity library and the semantic knowledge graph, to identify a causal driving factor of situation evolution, and to generate a prediction path of a potential risk scenario through counterfactual reasoning to obtain a situation awareness result of causal correlation.

[0142] In a third aspect, an electronic device is provided, including:

[0143] a processor;

[0144] a memory for storing processor-executable instructions;

[0145] The processor is configured to invoke the instructions stored in the memory to execute the method described above.

[0146] In a fourth aspect, a computer-readable storage medium is provided, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the method described above.

[0147] The present application can be a method, device, system and / or computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions loaded thereon for implementing various aspects of the present application.

[0148] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for intelligent analysis and situational awareness of multi-source heterogeneous video of water conservancy projects, characterized in that, include: A semantic knowledge graph in the field of water conservancy hubs is constructed, and semantic parsing is performed on multi-source heterogeneous video stream data. The visual features in the video frames are mapped to the semantic knowledge graph to obtain semantically annotated video stream data. The semantically annotated video stream data is aligned with a spatiotemporal reference. Based on the spatial position relationship and timestamp information of each video stream, a unified spatiotemporal coordinate system is established to obtain the spatiotemporally aligned semantic video stream data. A multimodal feature fusion network based on an attention mechanism is used to extract cross-modal features from the spatiotemporally aligned semantic video stream data. By adaptively learning the contribution of different video streams to target recognition, multi-scale fusion features are generated and combined to obtain a target feature set. Target detection and semantic segmentation are performed based on the target feature set to identify the monitoring objects and their spatial location information within the water conservancy hub monitoring area, and the information is then updated in reverse to the semantic knowledge graph to obtain a knowledge graph-enhanced target entity library. Based on the target entity library enhanced by the knowledge graph and the causal relationship edges in the semantic knowledge graph, causal link analysis is performed on the event sequence within the monitoring area of ​​the water conservancy hub to identify the causal driving factors of the situation evolution. Then, through counterfactual reasoning, predictive paths for potential risk scenarios are generated to obtain the situational awareness results of causal association.

2. The method according to claim 1, characterized in that, A semantic knowledge graph for the field of water conservancy projects is constructed, and semantic parsing is performed on multi-source heterogeneous video stream data. Visual features in video frames are mapped to the semantic knowledge graph to obtain semantically annotated video stream data, including: Entity extraction, relation extraction, and attribute extraction are performed on professional knowledge documents and operational rule texts in the field of water conservancy projects, and these are organized into entity node layer, attribute node layer, and relation edge layer, respectively, to obtain a hierarchical semantic knowledge graph. Local visual features and global contextual features of the multi-source heterogeneous video stream data are obtained by multi-scale convolutional feature encoding, and feature dimension alignment and semantic space mapping are performed to transform the visual feature space into the embedding space of the semantic knowledge graph to obtain visual semantic alignment features. By calculating the semantic similarity between the visual semantic alignment features and the entity node embedding vectors of the semantic knowledge graph, the visual objects of the water conservancy hub detected in the video frame are mapped to the corresponding entity nodes of the semantic knowledge graph, and semantic labels are assigned to the entity nodes according to the mapping confidence to obtain the entity node mapping results. Based on the entity node mapping results, the relation edges and attribute nodes of the mapped entity nodes in the semantic knowledge graph are extracted, and the semantic information of the relation edges and attribute nodes is attached to the visual objects of the corresponding video frames to form composite annotation information containing entity semantics, attribute semantics and relation semantics, thereby obtaining the semantically annotated video stream data.

3. The method according to claim 1, characterized in that, The semantically annotated video stream data is aligned to a spatiotemporal reference. Based on the spatial positional relationship and timestamp information of each video stream, a unified spatiotemporal coordinate system is established, resulting in spatiotemporally aligned semantic video stream data, including: Based on the camera spatial position parameters and field of view parameters of each video stream in the semantically annotated video stream data, a description of the spatial relationship of the video stream is constructed. Based on the description of the spatial relationship of the video streams, overlapping spatial regions with overlapping fields of view between different video streams are identified. By extracting spatial reference markers that are commonly visible to multiple video streams within the overlapping spatial regions, and calculating the projection transformation matrix between the pixel coordinates of the spatial reference markers in the image coordinate system of each video stream and their world coordinates in the world coordinate system, spatial alignment transformation parameters are obtained. The clock deviation and frame rate difference of the timestamp information in the semantically annotated video stream data are analyzed to construct a time synchronization mapping relationship, map the local timestamp of each video stream to a unified time reference, and perform time alignment interpolation on each video stream based on the unified time reference to obtain time alignment parameters. Based on the spatial alignment transformation parameters and the temporal alignment parameters, a unified spatiotemporal coordinate system is established. Based on the unified spatiotemporal coordinate system, the video frames of the semantically annotated video stream data are transformed to the world coordinate system through the projection transformation matrix, and aligned to a unified time reference through the time synchronization mapping relationship to obtain the spatiotemporally aligned semantic video stream data.

4. The method according to claim 1, characterized in that, A multimodal feature fusion network based on an attention mechanism extracts cross-modal features from the spatiotemporally aligned semantic video stream data. Through adaptive learning of the contribution of different video streams to target recognition, it generates multi-scale fusion features and combines them to obtain a target feature set, including: Each video stream in the spatiotemporally aligned semantic video stream data is feature-encoded, and the visual features of each video stream at different spatial scales are extracted by a convolutional feature extraction network to obtain multi-scale visual features. The semantic annotation information in the spatiotemporally aligned semantic video stream data is encoded into a semantic embedding vector, and then concatenated and fused with the multi-scale visual features in the feature dimension to obtain semantically enhanced multi-scale features. In the multimodal feature fusion network, attention weights between features at different spatial scales within a single video stream are calculated, and the semantically enhanced multi-scale features are weighted and fused to obtain intramodal fusion features; the contribution of different video streams to target recognition is calculated, cross-modal attention weights are assigned according to the contribution, and the intramodal fusion features of each video stream are adaptively weighted and aggregated to obtain cross-modal fusion features; A multi-scale feature pyramid is constructed for the cross-modal fusion features. Feature levels of different resolutions are generated through upsampling and downsampling operations. Local features and global features are extracted at each feature level. The local features and global features are combined to form multi-scale fusion features, thereby obtaining the target feature set.

5. The method according to claim 4, characterized in that, Based on the target feature set, target detection and semantic segmentation are performed to identify the monitoring objects and their spatial location information within the water conservancy hub monitoring area. This information is then updated back to the semantic knowledge graph to obtain a knowledge graph-enhanced target entity library, including: A target detection network is constructed based on the target feature set. The target detection network generates candidate target regions at each feature level of the multi-scale feature pyramid. Boundary box regression and target classification are performed on the candidate target regions to identify the target detection results within the water conservancy hub monitoring area. A semantic segmentation network is constructed based on the target feature set. The semantic segmentation network is used to perform pixel-level semantic annotation on the monitored objects in the target detection results. The pixels of the monitored objects are assigned to the corresponding semantic categories to obtain a semantic segmentation mask. By combining the target detection results with the semantic segmentation mask, the spatial location information and timestamp information of the monitored object are obtained, and the spatiotemporally located monitored object entity is obtained. By semantically matching the spatiotemporally located monitored object entity with the semantic knowledge graph and calculating the semantic similarity, the category to which the monitored object entity belongs in the semantic knowledge graph is determined. Based on the classification, the association between the monitored object entity and the existing entities in the semantic knowledge graph is constructed. The monitored object entity is inserted into the semantic knowledge graph as a new entity node and relation edges are created. The topology of the semantic knowledge graph is updated to obtain the target entity library enhanced by the knowledge graph.

6. The method according to claim 5, characterized in that, Based on the enhanced target entity database and the causal relationship edges in the semantic knowledge graph, causal link analysis of event sequences within the monitoring area of ​​the water conservancy hub includes: The monitored entities are sorted in chronological order to construct a time-series entity sequence; Based on the time interval and spatial position relationship between adjacent monitored object entities in the time-series entity sequence, a combination of monitored object entities with spatiotemporal correlation is identified, and the monitored object entities and their semantic annotation information in the combination of monitored object entities are mapped into event representations to obtain an event sequence. Retrieve event entity nodes corresponding to events in the event sequence from the semantic knowledge graph, traverse the causal relationship edges between the event entity nodes, determine the causal dependency relationship between events based on the directionality and semantic type of the causal relationship edges, and construct the causal link graph of the event sequence; In the causal link graph, the in-degree and out-degree of each event node are calculated. Event nodes with an in-degree of zero and an out-degree greater than zero are identified as the starting events of the causal link, and event nodes with an out-degree of zero and an in-degree greater than zero are identified as the resulting events of the causal link. Event nodes on the causal propagation path are identified by tracing back the causal propagation path from the starting event to the resulting event.

7. The method according to claim 1, characterized in that, Identifying the causal drivers of situational evolution and generating predictive paths for potential risk scenarios through counterfactual reasoning yields situational awareness results with causal relationships, including: Based on the causal relationship edges in the causal link graph and the event nodes on the causal propagation path, a set of causal reasoning rules for situational evolution is constructed. For the causal driving factors in the causal reasoning rule set, counterfactual conditions are constructed, and the counterfactual conditions are substituted into the causal reasoning rule set for reasoning calculation. The influence of the counterfactual conditions on subsequent events is propagated along the causal relationship edge to obtain the potential risk scenario prediction path of counterfactual reasoning. Based on the causal strength of the causal relationship edge between the risk attribute annotation of each event in the event sequence and the event in the semantic knowledge graph, the comprehensive risk score of the potential risk scenario prediction path is calculated, and the potential risk scenario prediction path with a comprehensive risk score exceeding a preset threshold is identified as a high-risk prediction path. By associating and organizing the causal driving factors, the causal link diagram, the potential risk scenario prediction path, and the high-risk prediction path, a situational awareness knowledge structure containing the current situation status, causal evolution mechanism, and future risk prediction is constructed, and the situational awareness results of the causal associations in the multi-source heterogeneous video of the water conservancy hub are obtained.

8. A multi-source heterogeneous video intelligent analysis and situational awareness system for water conservancy projects, used to implement the method as described in any one of claims 1-7, characterized in that, include: The first unit is used to construct a semantic knowledge graph in the field of water conservancy hubs and to perform semantic parsing on multi-source heterogeneous video stream data, mapping the visual features in the video frames to the semantic knowledge graph to obtain semantically annotated video stream data; The second unit is used to perform spatiotemporal reference alignment on the semantically annotated video stream data. Based on the spatial position relationship and timestamp information of each video stream, a unified spatiotemporal coordinate system is established to obtain the spatiotemporally aligned semantic video stream data. The third unit is used for a multimodal feature fusion network based on an attention mechanism to extract cross-modal features from the spatiotemporally aligned semantic video stream data. By adaptively learning the contribution of different video streams to target recognition, multi-scale fusion features are generated and combined to obtain a target feature set. The fourth unit is used to perform target detection and semantic segmentation based on the target feature set, identify the monitoring objects and their spatial location information within the water conservancy hub monitoring area, and update them in reverse to the semantic knowledge graph to obtain a knowledge graph-enhanced target entity library. The fifth unit is used to perform causal link analysis on the event sequence within the monitoring area of ​​the water conservancy hub based on the target entity library enhanced by the knowledge graph and the causal relationship edges in the semantic knowledge graph, identify the causal driving factors of the situation evolution, and generate prediction paths for potential risk scenarios through counterfactual reasoning to obtain the situational awareness results of causal association.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.