Campus hidden danger identification method and system based on multi-modal fusion and causal reasoning

The campus hazard identification method based on multimodal fusion and causal reasoning integrates visual, auditory, and sensor data to generate a joint representation of campus hazards and perform causal reasoning. This solves the problems of high false alarm rate and lack of causal reasoning in existing technologies, and achieves high-precision identification and early warning of campus safety.

CN121786747APending Publication Date: 2026-04-03CHENGDU XUNDAO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing campus security monitoring systems rely on a single visual modality, resulting in a high false alarm rate, a lack of causal reasoning ability, an inability to effectively distinguish between behaviors that appear similar but are different in nature, and a lack of comprehensive judgment based on the overall context of the scene and multi-dimensional clues.

Method used

By employing a multimodal fusion and causal reasoning approach, visual, auditory, and sensor data are integrated. A joint representation of campus hazards is generated through a multimodal fusion model, a dynamic temporal knowledge graph is constructed, and a hybrid causal discovery algorithm is used for causal reasoning to identify potential conflicts and provide early warnings.

Benefits of technology

It achieves a deep understanding of complex scenarios, reduces false alarm rates, can distinguish superficially similar behaviors from the essence of behavior, provides forward-looking early warning and intervention strategies, and improves the accuracy and predictability of campus safety management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786747A_ABST
    Figure CN121786747A_ABST
Patent Text Reader

Abstract

The invention discloses a campus hidden danger identification method and system based on multi-modal fusion and causal reasoning, and the method comprises the steps: carrying out the comprehensive judgment of the fine-grained motion semantics, such as motion coherence, regional skeleton motion behaviors and inter-regional skeleton motion behaviors, through a multi-modal fusion model, so as to distinguish the surface similar behaviors; false alarms caused by semantic splitting can be reduced essentially; by analyzing sequential logic and an interaction mode between actions, whether the behavior essence is a benign activity or a potential conflict is traced, and is associated to deep motives such as individual social history and environmental inducement, so that the shallow analysis limitation that a traditional model only depends on statistical association is broken through, and cognitive crossing from phenomenon recording to root reasoning is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hazard identification technology, specifically to a method and system for identifying campus hazards based on multimodal fusion and causal reasoning. Background Technology

[0002] Currently, artificial intelligence technologies, especially computer vision and pattern recognition technologies, have been widely applied in the field of campus security monitoring and hazard identification, aiming to improve the efficiency and response speed of campus security management through automation. Existing technical solutions mainly rely on surveillance cameras deployed in key areas of the campus, utilizing deep learning-based video analysis models to perform target detection, behavior recognition, and abnormal event detection on real-time video streams. A typical implementation involves training convolutional neural networks (CNNs), temporal models (such as 3D CNNs, RNNs / LSTMs), or a combination of both to identify specific targets (such as people and vehicles) and predefined behavioral patterns (such as running, falling, and physical contact) in images or video sequences.

[0003] However, this technology approach, centered on a single visual modality, has significant limitations. First, existing systems exhibit highly fragmented semantic understanding of complex dynamic scenes. Systems typically only perceive and analyze visual segments directly presented in video frames, lacking the ability to comprehensively judge the overall context of the scene, the participants' intentions, and multi-dimensional clues. For example, when recognizing physical interactions between students, the system may only detect low-level visual features such as "rapid movement" and "physical contact," triggering an alarm accordingly, but failing to fundamentally distinguish between "playful roughhousing" and "violent conflict"—two behaviors that may appear similar on the surface but are fundamentally different. This lack of semantic understanding directly leads to a high false alarm rate, causing unnecessary alarm interference, reducing security personnel's trust in the system, and potentially leading to "alarm fatigue" in the long run, causing critical warnings to be ignored.

[0004] A deeper deficiency lies in the fact that existing models are essentially pattern recognition systems based on statistical correlations. Their decisions heavily rely on superficial correlations in the training data, generally lacking the ability for causal reasoning and deep correlation analysis. The system can detect spatial events such as "multiple students gathering," but cannot further infer whether this gathering is a benign collective activity (such as normal teaching activities) or a potentially malicious event. More importantly, existing technology is completely incapable of tracing and revealing the underlying causal drivers of events, such as strained social relationships, changes in the psychological state of specific individuals, or previously occurring triggering events. This lack of causal connection means that the system can only act as a passive "event recorder," at best providing delayed event alerts, but unable to achieve proactive, root-cause risk warnings and intervention guidance, severely limiting the predictability and effectiveness of campus safety hazard prevention and control.

[0005] Therefore, there is an urgent need for a new campus hazard identification technology that can break through the limitations of single-modal perception, deeply integrate multi-source information, and have certain scene semantic understanding and causal reasoning capabilities. This would overcome the many shortcomings of existing technologies, such as high false alarm rate, shallow understanding, delayed early warning and lack of interpretability, so as to achieve more intelligent, accurate and forward-looking proactive campus safety protection. Summary of the Invention

[0006] Based on the problems raised in the background technology above, the purpose of this invention is to provide a method and system for identifying campus hazards based on multimodal fusion and causal reasoning, which solves the problem of the difficulty in fundamentally distinguishing between "playing around" and "violent conflict," two behaviors that may appear similar on the surface but are completely different in nature.

[0007] This invention is achieved through the following technical solution: The first aspect of this invention provides a method for identifying campus safety hazards based on multimodal fusion and causal reasoning, comprising the following steps: Step S1: Obtain multi-source campus hazard data and preprocess the multi-source campus hazard data; Step S2: Establish a multimodal fusion model; The multi-module fusion model includes a body movement behavior recognition model, a spatiotemporal extension model, an acoustic event detection model, and a fusion model. Preprocessed multi-source campus hazard data is sequentially input into the body movement behavior recognition model and the spatiotemporal extension model to generate spatiotemporal movement trajectory semantic information. Preprocessed multi-source campus hazard data is also input into the acoustic event detection model to generate acoustic event semantic information. The fusion model then fuses the spatiotemporal movement trajectory semantic information and the acoustic event semantic information to obtain a joint representation of campus hazards. Step S3: Construct a dynamic temporal knowledge graph based on the joint representation of campus hazards, and use the dynamic temporal knowledge graph to generate a structured hazard event graph; Step S4: Use a hybrid causal discovery algorithm to perform causal reasoning on the structured hidden danger event map to obtain the hidden danger event chain; Step S5: Perform campus hazard analysis on the chain of potential hazards and issue an early warning based on the results of the campus hazard analysis.

[0008] In the above technical solution, firstly, by integrating multi-source potential hazard data from campuses, preprocessed visual, auditory, and sensor data are input into a multimodal fusion model. Using cross-modal attention mechanisms and joint representation learning, a joint representation of campus potential hazards is generated that comprehensively depicts the scene context, behavioral relationships, and dynamic evolution. This representation not only integrates multi-dimensional features such as action, sound, and spatial location, but also comprehensively judges fine-grained motion semantics such as action coherence, regional skeletal movement behavior (e.g., the force and trajectory of arm swings), and inter-regional skeletal movement behavior (e.g., the coordination between upper limb force exertion and trunk stability) through the multimodal fusion model. This fundamentally reduces false alarms caused by semantic fragmentation when distinguishing superficially similar behaviors.

[0009] Based on the aforementioned fusion representation, a dynamic temporal knowledge graph is constructed, transforming multimodal information into a structured potential event graph containing entities, relationships, attributes, and evolutionary paths. A hybrid causal discovery algorithm is then used to mine and verify the implicit causal structures within this graph. This method not only identifies surface events but also traces the nature of behavior—whether it is benign activity or potential conflict—by analyzing the temporal logic and interaction patterns between actions, and connects it to deeper motivations such as individual social history and environmental triggers. This overcomes the limitations of traditional models that rely solely on superficial statistical correlation analysis, achieving a cognitive leap from phenomenon recording to root cause reasoning.

[0010] Ultimately, through in-depth analysis and risk simulation of the causal chain of potential hazards, we can proactively identify potential safety hazards and provide early warning and intervention strategies based on causal relationships.

[0011] In one optional embodiment, establishing a multimodal fusion model includes: A body movement and behavior recognition model is established based on the improved EAST-GCN network. The body movement and behavior recognition model is used to extract skeletal data from the image data in the input data and to recognize movement and behavior based on the extracted skeletal data. A spatiotemporal extension model is established, which is used to extend the identified action behavior in spatiotemporal terms and generate spatiotemporal action trajectory semantic information based on the spatiotemporally extended action behavior. An acoustic event detection model is established, which is used to detect acoustic events in the sound data of the input data and generate acoustic event semantic information based on the acoustic event detection results. A fusion model is established to fuse the spatiotemporal action trajectory semantic information with the acoustic event semantic information.

[0012] In one optional embodiment, a body movement recognition model is established based on the improved EAST-GCN network, including: An input layer is established, which is used to adjust the size of the image data; A feature extraction layer is established based on an efficient long-range attention mechanism. This feature extraction layer is used to extract features from the adjusted image data and generate a feature map. Adversarial training is introduced to establish a neural network, which is used to extract skeletal joint points from the feature map; A graph convolutional neural network is established based on a dynamic expansion strategy of adjacency structure. The graph convolutional neural network is used to dynamically expand the region according to the extracted skeletal joints to generate a region-specific expanded adjacency skeletal graph. A multi-branch spatiotemporal convolutional network is established, which is used to extend the extended adjacency skeleton graph of each region into the spatiotemporal graph convolution.

[0013] In one optional embodiment, a graph convolutional neural network is established based on a dynamic expansion strategy of adjacency structure, including: A mapping function is established based on a partitioning strategy. The mapping function is used to partition the skeletal joints for training and to adjust the convolution kernel size according to the partitioning training results. A weighting function is established based on a dynamic balancing mechanism. The weighting function is used to assign weights to the connection relationships. The connection relationship refers to the connection relationship between skeletal joints. An adaptive matrix is ​​established, which is used to dynamically adjust the weights of each skeletal joint.

[0014] In one optional embodiment, a spatiotemporal extension model is established, including: An input layer is established, which is used to normalize the identified actions and behaviors to generate multi-dimensional input data. A spatiotemporally extended convolutional network is established, which includes several convolutional blocks, each of which includes a spatiotemporal fusion module and an extended temporal module; The spatiotemporal fusion module includes parallel spatiotemporal feature extraction channels and time feature extraction channels, as well as a spatiotemporal feature fusion module for fusing the extracted spatiotemporal features and time features. The extended time module includes a parallel expansion channel and a spatiotemporal sensing channel, as well as a splicing module that splices the outputs of the expansion channel and the spatiotemporal sensing channel.

[0015] In one optional embodiment, the preprocessed multi-source campus hazard data is input as input data into the multimodal fusion model for multimodal fusion processing, including the following steps: The preprocessed multi-source campus hazard data includes hazard image data and hazard audio data; The body action behavior recognition model in the multimodal fusion model is used to perform action behavior recognition on the hidden danger image data to obtain regional skeletal action behavior and inter-regional skeletal action behavior; The regional skeletal motion behavior and the inter-regional skeletal motion behavior are respectively input into the spatiotemporal extension model in the multimodal fusion model for spatiotemporal extension to obtain the regional extended motion trajectory and the inter-regional extended motion trajectory. By comprehensively analyzing the regional extended action trajectory and the inter-regional extended action trajectory, spatiotemporal action trajectory semantic information is obtained. The acoustic event detection model in the multimodal fusion model is used to perform acoustic event detection on the hidden danger audio data, and acoustic event semantic information is generated based on the acoustic event detection results; By fusing the spatiotemporal action trajectory semantic information and the acoustic event semantic information, a joint representation of campus hidden dangers is obtained.

[0016] In one optional embodiment, the body motion behavior recognition model in the multimodal fusion processing is used to perform motion behavior recognition on the hazard image data, including the following steps: Feature extraction is performed on the image data of the potential hazards to obtain a feature map of campus personnel; Extract the skeletal joints of the campus personnel from the campus personnel feature map, and divide the skeletal joints into human body regions to obtain several skeletal joint region regions. Based on the analysis of the laws of human motion coordination, the relationships between skeletal joints in several skeletal joint regions are analyzed, and the skeletal joints are dynamically connected based on the relationships to generate a human skeletal structure diagram; the human skeletal structure diagram includes: skeletal joints and the relationships between skeletal joints. Spatiotemporal expansion of the human skeletal structure diagram is performed to obtain regional skeletal motion behaviors and inter-regional skeletal motion behaviors.

[0017] In one optional embodiment, a comprehensive analysis is performed on the region expansion trajectory and the inter-region expansion trajectory, including: Determine a time step, and calculate the first change in the regional expansion trajectory and the second change in the inter-regional expansion trajectory at the time step; Based on the physical collision mechanism, calculate the physical collision value of the first change; The conflict value is determined based on the physical collision value and the second change amount; Semantic analysis is performed on the regional expansion action trajectory and the inter-regional expansion action trajectory to obtain the initial spatiotemporal action trajectory semantic information; The initial spatiotemporal action trajectory semantic information is corrected by using the conflict value to obtain the spatiotemporal action trajectory semantic information.

[0018] In one optional embodiment, a hybrid causal discovery algorithm is used to perform causal reasoning on the structured hazard event map, including the following steps: Entities, relationships, attributes, and temporal connections are extracted from the structured event graph. Based on the entities and relationships, historical potential hazard events are searched to obtain the historical potential hazard events; Perform event matching between the historical potential hazard events and the attributes. If the event matching is successful, extract the historical conflict value from the historical potential hazard events. By combining a constraint-based causal discovery algorithm with a score-based structure learning algorithm, causal structure analysis is performed on the entities, relationships, attributes, and temporal connections to generate a causal chain of potential event hazards with causal direction. Extract the current conflict value from the attribute, and calculate the conflict weight by combining the historical conflict value and the current conflict value; By combining the causal chain of the potential hazards and the conflict weights, a potential hazard chain is generated.

[0019] A second aspect of this invention provides a campus hazard identification system based on multimodal fusion and causal reasoning, comprising: The data acquisition module is used to acquire multi-source campus hazard data and preprocess the multi-source campus hazard data; The multimodal fusion module is used to build a multimodal fusion model; The multi-module fusion model includes a body movement behavior recognition model, a spatiotemporal extension model, an acoustic event detection model, and a fusion model. Preprocessed multi-source campus hazard data is sequentially input into the body movement behavior recognition model and the spatiotemporal extension model to generate spatiotemporal movement trajectory semantic information. Preprocessed multi-source campus hazard data is also input into the acoustic event detection model to generate acoustic event semantic information. The fusion model then fuses the spatiotemporal movement trajectory semantic information and the acoustic event semantic information to obtain a joint representation of campus hazards. The knowledge graph module is used to construct a dynamic temporal knowledge graph based on the joint representation of campus hazards, and to generate a structured hazard event graph using the dynamic temporal knowledge graph; The causal reasoning module is used to perform causal reasoning on the structured hidden danger event map using a hybrid causal discovery algorithm to obtain the hidden danger event chain; The hazard analysis module is used to perform campus hazard analysis on the hazard event chain and issue early warnings based on the campus hazard analysis results.

[0020] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. By using a multimodal fusion model to comprehensively determine fine-grained motion semantics such as action coherence, regional skeletal action behavior, and inter-regional skeletal action behavior, false alarms caused by semantic fragmentation can be fundamentally reduced when distinguishing superficially similar behaviors. 2. By analyzing the temporal logic and interaction patterns between actions, we can trace the nature of behavior to determine whether it is a benign activity or a potential conflict, and link it to deeper motivations such as individual social history and environmental factors. This breaks through the limitations of traditional models that rely solely on superficial statistical correlation analysis, and achieves a cognitive leap from phenomenon recording to root cause reasoning. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 This is a flowchart illustrating the campus hazard identification method based on multimodal fusion and causal reasoning provided in Embodiment 1 of the present invention. Figure 2 This is a schematic diagram of the campus hazard identification system based on multimodal fusion and causal reasoning provided in Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device provided in Embodiment 2 of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention. Example

[0023] Figure 1 This is a flowchart illustrating the campus hazard identification method based on multimodal fusion and causal reasoning provided in Embodiment 1 of the present invention. Figure 1 As shown, the campus hazard identification method based on multimodal fusion and causal reasoning includes the following steps: Step S1: Obtain multi-source campus hazard data and preprocess the multi-source campus hazard data; Step S2: Establish a multimodal fusion model; The multi-module fusion model includes a body movement behavior recognition model, a spatiotemporal extension model, an acoustic event detection model, and a fusion model. Preprocessed multi-source campus hazard data is sequentially input into the body movement behavior recognition model and the spatiotemporal extension model to generate spatiotemporal movement trajectory semantic information. Preprocessed multi-source campus hazard data is also input into the acoustic event detection model to generate acoustic event semantic information. The fusion model then fuses the spatiotemporal movement trajectory semantic information and the acoustic event semantic information to obtain a joint representation of campus hazards. Step S3: Construct a dynamic temporal knowledge graph based on the joint representation of campus hazards, and use the dynamic temporal knowledge graph to generate a structured hazard event graph; Step S4: Use a hybrid causal discovery algorithm to perform causal reasoning on the structured hidden danger event map to obtain the hidden danger event chain; Step S5: Perform campus hazard analysis on the chain of potential hazards and issue an early warning based on the results of the campus hazard analysis.

[0024] It should be emphasized that the multi-source campus hazard data obtained by this method is obtained in a legal and compliant manner, and the multi-source campus hazard data is used for the purpose of maintaining campus public safety.

[0025] It should be noted that, firstly, by integrating multi-source potential hazard data from campuses, pre-processed visual, auditory, and sensor data are input into a multimodal fusion model. Using cross-modal attention mechanisms and joint representation learning, a joint representation of campus potential hazards is generated that comprehensively depicts the scene context, behavioral relationships, and dynamic evolution. This representation not only integrates multi-dimensional features such as action, sound, and spatial location, but also comprehensively judges fine-grained motion semantics such as action coherence, regional skeletal movement behavior (e.g., the force and trajectory of arm swings), and inter-regional skeletal movement behavior (e.g., the coordination between upper limb force exertion and trunk stability) through the multimodal fusion model. This fundamentally reduces false alarms caused by semantic fragmentation when distinguishing superficially similar behaviors.

[0026] Based on the aforementioned fusion representation, a dynamic temporal knowledge graph is constructed, transforming multimodal information into a structured potential event graph containing entities, relationships, attributes, and evolutionary paths. A hybrid causal discovery algorithm is then used to mine and verify the implicit causal structures within this graph. This method not only identifies surface events but also traces the nature of behavior—whether it is benign activity or potential conflict—by analyzing the temporal logic and interaction patterns between actions, and connects it to deeper motivations such as individual social history and environmental triggers. This overcomes the limitations of traditional models that rely solely on superficial statistical correlation analysis, achieving a cognitive leap from phenomenon recording to root cause reasoning.

[0027] Ultimately, through in-depth analysis and risk simulation of the causal chain of potential hazards, we can proactively identify potential safety hazards and provide early warning and intervention strategies based on causal relationships.

[0028] In one optional embodiment, establishing a multimodal fusion model includes: A body movement and behavior recognition model is established based on the improved EAST-GCN network. The body movement and behavior recognition model is used to extract skeletal data from the image data in the input data and to recognize movement and behavior based on the extracted skeletal data. A spatiotemporal extension model is established, which is used to extend the identified action behavior in spatiotemporal terms and generate spatiotemporal action trajectory semantic information based on the spatiotemporally extended action behavior. An acoustic event detection model is established, which is used to detect acoustic events in the sound data of the input data and generate acoustic event semantic information based on the acoustic event detection results. A fusion model is established to fuse the spatiotemporal action trajectory semantic information with the acoustic event semantic information.

[0029] Among them, the acoustic event detection model can adopt the existing YAMNet model, PANNs model, and HEAR benchmark model. The above acoustic event detection models can perform event detection on the sound data in the input data, including the detection of human voices (such as screams, dialogues, etc.), interactive sounds (such as collisions, knocking, etc.), and ambient sounds. All of the above models can detect the above sound events and convert them into corresponding semantic information.

[0030] The fusion model uses clustering to cluster spatiotemporal action trajectory semantic information with acoustic event semantic information.

[0031] In one optional embodiment, a body movement recognition model is established based on the improved EAST-GCN network, including: An input layer is established, which is used to adjust the size of the image data; A feature extraction layer is established based on an efficient long-range attention mechanism. This feature extraction layer is used to extract features from the adjusted image data and generate a feature map. Adversarial training is introduced to establish a neural network, which is used to extract skeletal joint points from the feature map; A graph convolutional neural network is established based on a dynamic expansion strategy of adjacency structure. The graph convolutional neural network is used to dynamically expand the region according to the extracted skeletal joints to generate a region-specific expanded adjacency skeletal graph. A multi-branch spatiotemporal convolutional network is established, which is used to extend the extended adjacency skeleton graph of each region into the spatiotemporal graph convolution.

[0032] It should be noted that the input video frames are first preprocessed to be standardized and cropped to a fixed size to eliminate scale changes caused by differences in camera angle and distance. This ensures the consistency of input data specifications in subsequent processing, thereby effectively reducing the redundancy of system calculations and laying the foundation for low-latency processing.

[0033] In the feature extraction stage, the network introduces a long-range attention mechanism, enabling the model to automatically focus on the human target region in the image, suppressing interference from complex backgrounds, and thus extracting detailed high-dimensional pose features to construct a clear feature representation for subsequent keypoint localization. Addressing challenges such as frequent human occlusion, limb overlap, and image boundary truncation in campus environments, this embodiment specifically designs an adversarial missing point compensation module. This module includes a generator and a discriminator. The generator predicts the reasonable locations of occluded or missing keypoints based on visible keypoint information and contextual features; the discriminator evaluates the generated complete 2D skeleton sequence, assessing its coherence and biomechanical rationality. Through adversarial training, the two components are collaboratively optimized, ultimately enabling the system to output a complete 2D skeleton sequence that conforms to human motion constraints.

[0034] Based on the acquisition of the complete skeleton sequence, this embodiment employs a dynamically expanded graph convolutional network for spatial relationship modeling. This network is no longer limited to fixed adjacency relationships in human physical connections. Instead, it dynamically constructs and expands the adjacency matrix of the graph based on three strategies: "body functional zoning," "left-right symmetric association," and data-driven "adaptive correlation." This incorporates functionally related or motorly coordinated distant joint nodes into the same receptive field, thereby fully exploring long-range dependencies across joints and limbs and more effectively representing complex behaviors.

[0035] Finally, a multi-branch spatiotemporal convolutional module is used to perform parallel temporal modeling of the dynamic graph structure. This module decomposes the expanded skeleton graph into multiple subgraphs, each processed by a different temporal convolutional branch: one branch focuses on subtle motion changes within a short temporal window, another captures the rhythm and pattern of motion over a medium time span, and the third models the global motion trend of the entire behavioral segment. The outputs of each branch are adaptively fused using learnable weights, preserving detailed features while compressing information redundancy, ultimately forming a unified and robust representation of behavioral features. This enables high-precision recognition of various typical campus scene behaviors such as gathering, running, and pushing.

[0036] In one optional embodiment, a graph convolutional neural network is established based on a dynamic expansion strategy of adjacency structure, including: A mapping function is established based on a partitioning strategy. The mapping function is used to partition the skeletal joints for training and to adjust the convolution kernel size according to the partitioning training results. A weighting function is established based on a dynamic balancing mechanism. The weighting function is used to assign weights to the connection relationships. The connection relationship refers to the connection relationship between skeletal joints. An adaptive matrix is ​​established, which is used to dynamically adjust the weights of each skeletal joint.

[0037] It should be noted that the mapping function established based on the partitioning strategy is based on "body functional partitioning". Here, the body's center of gravity and center are determined, and the center is divided into the first skeletal joint region. Skeletal joints that are close to the center and whose distance from the center of gravity is greater than the distance between the center and the center of gravity are divided into the second skeletal joint region. Skeletal joints that are close to the center and whose distance from the center of gravity is less than the distance between the center and the center of gravity are divided into the third skeletal joint region. Skeletal joints that are far from the center and whose distance from the center of gravity is greater than the distance between the center and the center of gravity are divided into the fourth skeletal joint region. Skeletal joints that are far from the center and whose distance from the center of gravity is less than the distance between the center and the center of gravity are divided into the fifth skeletal joint region.

[0038] The above partitioning can be used to construct a mapping function that maps skeletal joints that meet the conditions to the corresponding partitions. The size of the convolution kernel can be adjusted by the above mapping function to satisfy the processing of the corresponding skeletal joints.

[0039] In this method, the core of distinguishing behaviors that may appear similar on the surface but are fundamentally different lies in finding the connections between skeletal joints and setting and dynamically adjusting the weights of these connections. By establishing a weight function based on a dynamic balancing mechanism and creating an adaptive matrix, we can establish the movement behaviors of campus personnel at each skeletal node and the muscle relationships between them, providing a foundation for subsequent analysis of the nature of the movements.

[0040] Specifically, the basis for establishing the weight function based on the dynamic equilibrium mechanism lies in the fact that the human body needs to adjust its core and limbs to maintain a stable center of gravity to complete the movement. Based on this, the interrelationships between skeletal nodes are established. For example, when extending the right hand forward, the left shoulder will move to compensate for the instability caused by the shift in the center of gravity. While both slapping and punching involve extending the right hand, the corresponding movement distance of the left shoulder is different. Establishing these interrelationships and assigning weights accordingly is necessary. In this embodiment, the greater the distance between skeletal joints, the smaller the edge weight; conversely, the greater the distance between skeletal joints, the larger the edge weight.

[0041] The accuracy of weight assignment based on spatial distance is limited, and the motion connection between skeletal joints does not perfectly match spatial distance. Therefore, an adaptive matrix is ​​constructed to dynamically and adaptively adjust the above weights during training.

[0042] In one optional embodiment, a spatiotemporal extension model is established, including: An input layer is established, which is used to normalize the identified actions and behaviors to generate multi-dimensional input data. A spatiotemporally extended convolutional network is established, which includes several convolutional blocks, each of which includes a spatiotemporal fusion module and an extended temporal module; The spatiotemporal fusion module includes parallel spatiotemporal feature extraction channels and time feature extraction channels, as well as a spatiotemporal feature fusion module for fusing the extracted spatiotemporal features and time features. The extended time module includes a parallel expansion channel and a spatiotemporal sensing channel, as well as a splicing module that splices the outputs of the expansion channel and the spatiotemporal sensing channel.

[0043] It should be noted that the input data is first preprocessed by normalization. The received and identified actions are integrated into a tensor with uniform dimensions, including the number of input channels, time frames, and keypoints, thus providing a standardized and information-rich input foundation for subsequent network processing.

[0044] Each convolutional block of the network contains a spatiotemporal fusion module, which employs a dual-channel parallel architecture for collaborative feature extraction. The spatial feature extraction channel, through an adaptive graph convolution mechanism, dynamically constructs various topological structures, including primal connections, global connections, and local connections, based on the actual correlation between key points in specific action samples, thereby accurately capturing spatially adaptive features highly relevant to the action. Simultaneously, the temporal feature extraction channel, through temporally adaptive graph convolution, generates adaptive weights based on the importance differences at different time stages of the action, effectively extracting key temporal dynamic patterns. Finally, a feature fusion mechanism deeply aggregates the features extracted from the spatial and temporal channels, fully exploring the complementary information between them to form a comprehensive and adaptive spatiotemporal joint feature representation.

[0045] Within each convolutional block, a spatiotemporal fusion module is followed by an extended temporal module to further enhance the modeling capability for temporal dynamics. This module employs two parallel processing channels: one dilation channel uses convolutional kernels with different dilation rates to expand the temporal receptive field, capturing long-term dependencies and macroscopic rhythms of actions to adapt to the recognition needs of actions of varying durations; the other spatiotemporal receptive channel combines pooling and normalization operations to strengthen the perception of subtle local temporal changes, compensating for the limitations of a single convolutional kernel in feature capture granularity. Subsequently, a concatenation operation integrates the multi-scale temporal features output from the two channels, and residual connections and channel attention mechanisms are introduced to optimize and enhance the fused features, thereby improving the model's representation capability and robustness to complex action temporal patterns.

[0046] By stacking the aforementioned modules in series, the network can progressively deepen its understanding of multi-stream input data. The entire processing flow begins with multi-dimensional data normalization, then proceeds through the fusion of dynamic spatial relationship modeling and adaptive temporal feature extraction, followed by multi-scale temporal feature expansion and optimization, ultimately forming a high-level feature representation with strong discriminative power for various human movements.

[0047] In one optional embodiment, the preprocessed multi-source campus hazard data is input as input data into the multimodal fusion model for multimodal fusion processing, including the following steps: The preprocessed multi-source campus hazard data includes hazard image data and hazard audio data; The body action behavior recognition model in the multimodal fusion model is used to perform action behavior recognition on the hidden danger image data to obtain regional skeletal action behavior and inter-regional skeletal action behavior; The regional skeletal motion behavior and the inter-regional skeletal motion behavior are respectively input into the spatiotemporal extension model in the multimodal fusion model for spatiotemporal extension to obtain the regional extended motion trajectory and the inter-regional extended motion trajectory. By comprehensively analyzing the regional extended action trajectory and the inter-regional extended action trajectory, spatiotemporal action trajectory semantic information is obtained. The acoustic event detection model in the multimodal fusion model is used to perform acoustic event detection on the hidden danger audio data, and acoustic event semantic information is generated based on the acoustic event detection results; By fusing the spatiotemporal action trajectory semantic information and the acoustic event semantic information, a joint representation of campus hidden dangers is obtained.

[0048] First, feature extraction is performed on the preprocessed multi-source hazard data from the campus. This data includes hazard image data reflecting the on-site situation and hazard audio data recording environmental sounds. In the visual processing path, a motion behavior recognition model is used to process the image data. This model can identify individual skeletal movements within a specific area, as well as correlated skeletal movements between different areas, thus analyzing human dynamics in the visual scene from both spatial relationships and group interaction perspectives.

[0049] Subsequently, the identified regional skeletal motion behaviors and inter-regional skeletal motion behaviors are input into a spatiotemporal extended model for further processing. This model generates extended trajectories describing the evolution of motion within a region and extended trajectories characterizing the association and interaction patterns of motion between regions by performing extended analysis of motion trajectories in both temporal and spatial dimensions. Through comprehensive analysis of these two types of trajectories, the model can extract spatiotemporal motion trajectory information with clear behavioral semantics, such as behavioral patterns with safety hazard characteristics, such as "rapid gathering," "chasing and running," or "suddenly falling to the ground."

[0050] In the auditory processing path, an acoustic event detection model is used to analyze the audio data of potential hazards. This model can identify specific acoustic events in the audio stream, such as the sound of breaking glass, sharp shouts, and abnormal impact sounds, and generate corresponding semantic descriptions of the acoustic events based on the detection results, providing auditory evidence support for the judgment of safety hazards.

[0051] Finally, a feature fusion mechanism is used to deeply fuse the spatiotemporal action trajectory semantic information extracted from the visual path with the acoustic event semantic information generated from the auditory path. This fusion process comprehensively considers the temporal synchronicity and semantic correlation of the two modalities, generating a unified joint representation of campus safety hazards. This joint representation comprehensively reflects both the "seen" abnormal actions and the "heard" abnormal sounds, thereby enabling more accurate and reliable automatic identification and early warning of various campus safety hazards such as fighting, vandalism, and sudden falls.

[0052] In one optional embodiment, the body motion behavior recognition model in the multimodal fusion processing is used to perform motion behavior recognition on the hazard image data, including the following steps: Feature extraction is performed on the image data of the potential hazards to obtain a feature map of campus personnel; Extract the skeletal joints of the campus personnel from the campus personnel feature map, and divide the skeletal joints into human body regions to obtain several skeletal joint region regions. Based on the analysis of the laws of human motion coordination, the relationships between skeletal joints in several skeletal joint regions are analyzed, and the skeletal joints are dynamically connected based on the relationships to generate a human skeletal structure diagram; the human skeletal structure diagram includes: skeletal joints and the relationships between skeletal joints. Spatiotemporal expansion of the human skeletal structure diagram is performed to obtain regional skeletal motion behaviors and inter-regional skeletal motion behaviors.

[0053] It should be noted that, firstly, feature extraction is performed on the hazard image data to obtain a feature map containing key information about campus personnel. Then, the skeletal joints of the campus personnel are extracted from the feature map, and these skeletal joints are divided into human body regions according to the characteristics of human physiological structure, forming several independent skeletal joint region regions, laying the foundation for subsequent refined analysis of movement behavior.

[0054] Building upon this foundation, we conducted an in-depth analysis of the connections between skeletal joints within and between different skeletal joint regions. Based on these connections, we dynamically linked the skeletal joints to construct a human skeletal structure diagram that includes the skeletal joints and their interrelationships. Subsequently, we expanded this human skeletal structure diagram in terms of its spatiotemporal dimensions to further uncover skeletal movement behaviors within and between different regions. This provides detailed movement data support for subsequently distinguishing the nature of movements and identifying potential hazards.

[0055] In one optional embodiment, a comprehensive analysis is performed on the region expansion trajectory and the inter-region expansion trajectory, including: Determine a time step, and calculate the first change in the regional expansion trajectory and the second change in the inter-regional expansion trajectory at the time step; Based on the physical collision mechanism, calculate the physical collision value of the first change; The conflict value is determined based on the physical collision value and the second change amount; Semantic analysis is performed on the regional expansion action trajectory and the inter-regional expansion action trajectory to obtain the initial spatiotemporal action trajectory semantic information; The initial spatiotemporal action trajectory semantic information is corrected by using the conflict value to obtain the spatiotemporal action trajectory semantic information.

[0056] It should be noted that when extending the right hand forward, the left shoulder will move to compensate for the instability caused by the shift in the center of gravity. While both slapping and punching involve extending the right hand, the corresponding distance the left shoulder moves differs. Based on this principle, it is necessary to calculate the regional expansion trajectory and the changes in the inter-regional expansion trajectory connected to that region at a specific time step; the time step can be selected according to specific circumstances.

[0057] The region expansion trajectory is a core element to consider in this embodiment. Its trajectory needs to be analyzed as the main collision detection element. Therefore, it is necessary to calculate the collision energy brought about by the trajectory change of the region expansion trajectory within the time step based on the physical collision mechanism, and use it as the physical collision value.

[0058] The second change amount corresponds to the aforementioned movement distance of the left shoulder. This is used as an adjustment amount and calculated with the physical collision value to determine the conflict value. In this embodiment, the ratio of the second change amount to the maximum movement of the bone joint is calculated as the adjustment value, and this adjustment value is used to adjust the physical collision value as the conflict value.

[0059] Then, semantic analysis is performed on the regional extended action trajectory and the inter-region extended action trajectory. For example, in the above example, semantic information such as the right hand punching and the left shoulder moving significantly can be obtained.

[0060] Conflict labels are generated based on conflict values, such as normal teaching behavior, conflict events, and neutral events. These conflict labels are then used to correct the semantic information of the initial spatiotemporal action trajectory and assign it an event bias.

[0061] By treating campus personnel as entities, campus personnel's events (such as spatiotemporal action trajectory semantic information and acoustic event semantic information) as relations, conflict tags as attributes, and event temporal sequence as temporal connection, a dynamic temporal knowledge graph is constructed. This graph is then structured to generate a structured hidden danger event graph.

[0062] In one optional embodiment, a hybrid causal discovery algorithm is used to perform causal reasoning on the structured hazard event map, including the following steps: Entities, relationships, attributes, and temporal connections are extracted from the structured event graph. Based on the entities and relationships, historical potential hazard events are searched to obtain the historical potential hazard events; Perform event matching between the historical potential hazard events and the attributes. If the event matching is successful, extract the historical conflict value from the historical potential hazard events. By combining a constraint-based causal discovery algorithm with a score-based structure learning algorithm, causal structure analysis is performed on the entities, relationships, attributes, and temporal connections to generate a causal chain of potential event hazards with causal direction. Extract the current conflict value from the attribute, and calculate the conflict weight by combining the historical conflict value and the current conflict value; By combining the causal chain of the potential hazards and the conflict weights, a potential hazard chain is generated.

[0063] It should be noted that the purpose of searching for historical potential incidents based on entities and relationships is to find out whether relevant campus personnel had interactions before, and whether those interactions were conflicts or normal teaching behaviors, in order to determine the nature of the incident.

[0064] If there was prior interaction and that interaction matches the conflict label of the current event, it indicates that there was prior conflict and the current event may be a continuation of the previous event. In this case, historical conflict values ​​should be extracted from historical potential events.

[0065] Combining constraint-based causal discovery algorithms with score-based structure learning algorithms allows for causal structure analysis of structured hazard event graphs, generating causal chains with causal orientations to illustrate the current state of the event. Furthermore, a conflict weight is calculated based on a ratio of current and historical conflict values, serving as the conflict probability for the current event. By integrating these causal chains with causal orientations, a hazard event chain is obtained. Example

[0066] Figure 2 This is a schematic diagram of the campus hazard identification system based on multimodal fusion and causal reasoning provided in Embodiment 2 of the present invention. Figure 2 As shown, the campus hazard identification system based on multimodal fusion and causal reasoning includes: The data acquisition module is used to acquire multi-source campus hazard data and preprocess the multi-source campus hazard data; The multimodal fusion module is used to build a multimodal fusion model; The multi-module fusion model includes a body movement behavior recognition model, a spatiotemporal extension model, an acoustic event detection model, and a fusion model. Preprocessed multi-source campus hazard data is sequentially input into the body movement behavior recognition model and the spatiotemporal extension model to generate spatiotemporal movement trajectory semantic information. Preprocessed multi-source campus hazard data is also input into the acoustic event detection model to generate acoustic event semantic information. The fusion model then fuses the spatiotemporal movement trajectory semantic information and the acoustic event semantic information to obtain a joint representation of campus hazards. The knowledge graph module is used to construct a dynamic temporal knowledge graph based on the joint representation of campus hazards, and to generate a structured hazard event graph using the dynamic temporal knowledge graph; The causal reasoning module is used to perform causal reasoning on the structured hidden danger event map using a hybrid causal discovery algorithm to obtain the hidden danger event chain; The hazard analysis module is used to perform campus hazard analysis on the hazard event chain and issue early warnings based on the campus hazard analysis results.

[0067] In one optional embodiment, the multimodal fusion module includes: The first establishment unit is used to establish a body action behavior recognition model based on the improved EAST-GCN network. The body action behavior recognition model is used to extract skeletal data from the image data in the input data and to perform action behavior recognition based on the extracted skeletal data. The second establishment unit is used to establish a spatiotemporal extension model, which is used to spatiotemporally extend the identified action behavior and generate spatiotemporal action trajectory semantic information based on the spatiotemporally extended action behavior. The third establishment unit is used to establish an acoustic event detection model, which is used to perform acoustic event detection on the sound data in the input data and generate acoustic event semantic information based on the acoustic event detection results. The fourth establishment unit is used to establish a fusion model, which is used to fuse the spatiotemporal action trajectory semantic information with the acoustic event semantic information. Example

[0068] Figure 3 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention, as shown below. Figure 3 As shown, the electronic device includes a processor 21, a memory 22, an input device 23, and an output device 24; the number of processors 21 in the computer device can be one or more. Figure 3 Taking a processor 21 as an example; the processor 21, memory 22, input device 23, and output device 24 in an electronic device can be connected via a bus or other means. Figure 3 For example, the connection is made via a bus; where processor 21 can be a graphics processor.

[0069] The memory 22, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules. The processor 21 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 22, thereby realizing the campus hazard identification method based on multimodal fusion and causal reasoning in Embodiment 1.

[0070] The memory 22 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, the memory 22 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory, or other non-volatile solid-state storage device. In some instances, the memory 22 may further include memory remotely located relative to the processor 21, which can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks and combinations thereof, and industrial internet platforms.

[0071] Input device 23 can be used to receive user input such as ID and password. Output device 24 is used to output the network configuration page. Example

[0072] Embodiment 4 of the present invention also provides a computer-readable storage medium, wherein the computer-executable instructions, when executed by a computer processor, are used to implement the campus hazard identification method based on multimodal fusion and causal reasoning as provided in Embodiment 1.

[0073] The storage medium containing computer-executable instructions provided in the embodiments of the present invention is not limited to the method operation provided in Embodiment 1, but can also perform related operations in the campus hazard identification method based on multimodal fusion and causal reasoning provided in any embodiment of the present invention.

[0074] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A campus hazard identification method based on multimodal fusion and causal reasoning, characterized in that, Includes the following steps: Step S1: Obtain multi-source campus hazard data and preprocess the multi-source campus hazard data; Step S2: Establish a multimodal fusion model; The multi-module fusion model includes a body movement behavior recognition model, a spatiotemporal extension model, an acoustic event detection model, and a fusion model. Preprocessed multi-source campus hazard data is sequentially input into the body movement behavior recognition model and the spatiotemporal extension model to generate spatiotemporal movement trajectory semantic information. Preprocessed multi-source campus hazard data is also input into the acoustic event detection model to generate acoustic event semantic information. The fusion model then fuses the spatiotemporal movement trajectory semantic information and the acoustic event semantic information to obtain a joint representation of campus hazards. Step S3: Construct a dynamic temporal knowledge graph based on the joint representation of campus hazards, and use the dynamic temporal knowledge graph to generate a structured hazard event graph; Step S4: Use a hybrid causal discovery algorithm to perform causal reasoning on the structured hidden danger event map to obtain the hidden danger event chain; Step S5: Perform campus hazard analysis on the chain of potential hazards and issue an early warning based on the results of the campus hazard analysis.

2. The campus hazard identification method based on multimodal fusion and causal reasoning according to claim 1, characterized in that, Establish a multimodal fusion model, including: A body movement and behavior recognition model is established based on the improved EAST-GCN network. The body movement and behavior recognition model is used to extract skeletal data from the image data in the input data and to recognize movement and behavior based on the extracted skeletal data. A spatiotemporal extension model is established, which is used to extend the identified action behavior in spatiotemporal terms and generate spatiotemporal action trajectory semantic information based on the spatiotemporally extended action behavior. An acoustic event detection model is established, which is used to detect acoustic events in the sound data of the input data and generate acoustic event semantic information based on the acoustic event detection results. A fusion model is established to fuse the spatiotemporal action trajectory semantic information with the acoustic event semantic information.

3. The campus hazard identification method based on multimodal fusion and causal reasoning according to claim 2, characterized in that, A body movement recognition model is established based on the improved EAST-GCN network, including: An input layer is established, which is used to adjust the size of the image data; A feature extraction layer is established based on an efficient long-range attention mechanism. This feature extraction layer is used to extract features from the adjusted image data and generate a feature map. Adversarial training is introduced to establish a neural network, which is used to extract skeletal joint points from the feature map; A graph convolutional neural network is established based on a dynamic expansion strategy of adjacency structure. The graph convolutional neural network is used to dynamically expand the region according to the extracted skeletal joints to generate a region-specific expanded adjacency skeletal graph. A multi-branch spatiotemporal convolutional network is established, which is used to extend the extended adjacency skeleton graph of each region into the spatiotemporal graph convolution.

4. The campus hazard identification method based on multimodal fusion and causal reasoning according to claim 3, characterized in that, Graph convolutional neural networks are built based on a dynamic expansion strategy of adjacency structure, including: A mapping function is established based on a partitioning strategy. The mapping function is used to partition the skeletal joints for training and to adjust the convolution kernel size according to the partitioning training results. A weighting function is established based on a dynamic balancing mechanism, which is used to assign weights to the connection relationships; wherein, the connection relationship is the connection relationship between skeletal joints. An adaptive matrix is ​​established, which is used to dynamically adjust the weights of each skeletal joint.

5. The campus hazard identification method based on multimodal fusion and causal reasoning according to claim 2, characterized in that, Establish a spatiotemporal extension model, including: An input layer is established, which is used to normalize the identified actions and behaviors to generate multi-dimensional input data. A spatiotemporally extended convolutional network is established, which includes several convolutional blocks, each of which includes a spatiotemporal fusion module and an extended temporal module; The spatiotemporal fusion module includes parallel spatiotemporal feature extraction channels and time feature extraction channels, as well as a spatiotemporal feature fusion module for fusing the extracted spatiotemporal features and time features. The extended time module includes a parallel expansion channel and a spatiotemporal sensing channel, as well as a splicing module that splices the outputs of the expansion channel and the spatiotemporal sensing channel.

6. The campus hazard identification method based on multimodal fusion and causal reasoning according to claim 1, characterized in that, The preprocessed multi-source campus hazard data is input into the multimodal fusion model for multimodal fusion processing, including the following steps: The preprocessed multi-source campus hazard data includes hazard image data and hazard audio data; The body action behavior recognition model in the multimodal fusion model is used to perform action behavior recognition on the hidden danger image data to obtain regional skeletal action behavior and inter-regional skeletal action behavior; The regional skeletal motion behavior and the inter-regional skeletal motion behavior are respectively input into the spatiotemporal extension model in the multimodal fusion model for spatiotemporal extension to obtain the regional extended motion trajectory and the inter-regional extended motion trajectory. By comprehensively analyzing the regional extended action trajectory and the inter-regional extended action trajectory, spatiotemporal action trajectory semantic information is obtained. The acoustic event detection model in the multimodal fusion model is used to perform acoustic event detection on the hidden danger audio data, and acoustic event semantic information is generated based on the acoustic event detection results; By fusing the spatiotemporal action trajectory semantic information and the acoustic event semantic information, a joint representation of campus hidden dangers is obtained.

7. The campus hazard identification method based on multimodal fusion and causal reasoning according to claim 6, characterized in that, The body action behavior recognition model in the multimodal fusion processing is used to perform action behavior recognition on the hazard image data, including the following steps: Feature extraction is performed on the image data of the potential hazards to obtain a feature map of campus personnel; Extract the skeletal joints of the campus personnel from the campus personnel feature map, and divide the skeletal joints into human body regions to obtain several skeletal joint region regions. Based on the laws of human motion coordination, the relationship between skeletal joints in several skeletal joint regions is analyzed, and the skeletal joints are dynamically connected based on the relationship to generate a human skeletal structure diagram. The human skeletal structure diagram includes: skeletal joints and the connections between skeletal joints; Spatiotemporal expansion of the human skeletal structure diagram is performed to obtain regional skeletal motion behaviors and inter-regional skeletal motion behaviors.

8. The campus hazard identification method based on multimodal fusion and causal reasoning according to claim 7, characterized in that, A comprehensive analysis is performed on the regional expansion trajectory and the inter-regional expansion trajectory, including: Determine a time step, and calculate the first change in the regional expansion trajectory and the second change in the inter-regional expansion trajectory at the time step; Based on the physical collision mechanism, calculate the physical collision value of the first change; The conflict value is determined based on the physical collision value and the second change amount; Semantic analysis is performed on the regional expansion action trajectory and the inter-regional expansion action trajectory to obtain the initial spatiotemporal action trajectory semantic information; The initial spatiotemporal action trajectory semantic information is corrected by using the conflict value to obtain the spatiotemporal action trajectory semantic information.

9. The campus hazard identification method based on multimodal fusion and causal reasoning according to claim 1, characterized in that, The structured hazard event map is subjected to causal inference using a hybrid causal discovery algorithm, including the following steps: Entities, relationships, attributes, and temporal connections are extracted from the structured event graph. Based on the entities and relationships, historical potential hazard events are searched to obtain the historical potential hazard events; Perform event matching between the historical potential hazard events and the attributes. If the event matching is successful, extract the historical conflict value from the historical potential hazard events. By combining a constraint-based causal discovery algorithm with a score-based structure learning algorithm, causal structure analysis is performed on the entities, relationships, attributes, and temporal connections to generate a causal chain of potential event hazards with causal direction. Extract the current conflict value from the attribute, and calculate the conflict weight by combining the historical conflict value and the current conflict value; By combining the causal chain of the potential hazards and the conflict weights, a potential hazard chain is generated.

10. A campus hazard identification system based on multimodal fusion and causal reasoning, wherein the campus hazard identification system is used to implement the campus hazard identification method based on multimodal fusion and causal reasoning as described in any one of claims 1 to 9, characterized in that, The campus hazard identification system includes: The data acquisition module is used to acquire multi-source campus hazard data and preprocess the multi-source campus hazard data; The multimodal fusion module is used to build a multimodal fusion model; The multi-module fusion model includes a body movement behavior recognition model, a spatiotemporal extension model, an acoustic event detection model, and a fusion model. Preprocessed multi-source campus hazard data is sequentially input into the body movement behavior recognition model and the spatiotemporal extension model to generate spatiotemporal movement trajectory semantic information. Preprocessed multi-source campus hazard data is also input into the acoustic event detection model to generate acoustic event semantic information. The fusion model then fuses the spatiotemporal movement trajectory semantic information and the acoustic event semantic information to obtain a joint representation of campus hazards. The knowledge graph module is used to construct a dynamic temporal knowledge graph based on the joint representation of campus hazards, and to generate a structured hazard event graph using the dynamic temporal knowledge graph; The causal reasoning module is used to perform causal reasoning on the structured hidden danger event map using a hybrid causal discovery algorithm to obtain the hidden danger event chain; The hazard analysis module is used to perform campus hazard analysis on the hazard event chain and issue early warnings based on the campus hazard analysis results.