A campus abnormal behavior recognition method and system and a storage medium

CN121723357BActive Publication Date: 2026-05-26BEIJING CHUANGPU TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610222256.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-25
Publication Date
2026-05-26
Estimated Expiration
2046-02-25

Smart Images

  • Figure CN121723357B_ABST
    Figure CN121723357B_ABST
Patent Text Reader

Abstract

This invention provides a method, system, and storage medium for identifying abnormal behavior on campus, relating to the field of image recognition technology. The method includes: collecting multimodal data from a campus monitoring area; processing the multimodal data to determine feature data; constructing a campus security spatiotemporal map based on the campus physical layout and visual overlap areas, combined with the feature data; inputting the campus security spatiotemporal map into a graph convolutional recurrent neural network to output high-dimensional spatiotemporal features; performing aggregation operations on the high-dimensional spatiotemporal features to determine a comprehensive behavioral representation; determining whether a predefined abnormal campus behavior has occurred based on the comprehensive behavioral representation; if so, outputting the corresponding category and confidence level of the abnormal campus behavior and triggering an alarm; otherwise, continuing monitoring; triggering a real-time audible and visual alarm on a local computing device deployed at the network edge and sending alarm information to the security center; and uploading the alarm information and key behavioral analysis data to a cloud-based security management platform for global situation display.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, and in particular to a method, system, and storage medium for identifying abnormal behavior on campus. Background Technology

[0002] With the continuous expansion of campuses and the increasing mobility of personnel, the complexity and uncertainty of campus security management have significantly increased. To improve the security level of public areas on campus, more and more campuses are deploying video surveillance equipment and IoT sensing facilities to continuously monitor personnel behavior, thereby providing a data foundation for the automatic identification of abnormal behavior.

[0003] Campus abnormal behavior recognition technology can automatically analyze and intelligently identify the behavior of people in the campus environment, and can detect potential security risks in the early stages of abnormal events. At the same time, it can reduce the workload of security personnel. It has important practical significance and application value for building an intelligent and refined campus security management system.

[0004] However, existing methods for identifying abnormal behavior on campuses often rely on single visual information or localized analysis based on individual cameras, making it difficult to fully integrate multi-source sensory data and resulting in limited overall understanding of complex behaviors. Furthermore, some methods neglect the spatiotemporal correlation between campus spatial structure and human behavior, making it difficult to reliably identify abnormal behaviors across regions and time periods. They are also susceptible to occlusion, changes in lighting, or transient noise, leading to false alarms or missed alarms, thus limiting their effectiveness in real-world campus scenarios. Summary of the Invention

[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide a method for identifying abnormal behavior on campus, which can solve the technical problem that existing methods for identifying abnormal behavior on campus mostly rely on single visual information or local analysis based on independent cameras, making it difficult to fully integrate multi-source perception data, resulting in limited overall understanding of complex behaviors, thus producing false alarms or missed alarms, and restricting their application effect in actual campus scenarios.

[0006] A first aspect of this invention proposes a method for identifying abnormal behavior on campus, comprising:

[0007] S1: Real-time acquisition of multimodal data from the campus monitoring area, including: visual surveillance video stream data and IoT sensor data stream;

[0008] S2: Extract features from multimodal data to determine the feature data;

[0009] S3: Based on the physical layout of the campus and the visual overlap area, combined with feature data, construct a spatiotemporal map of campus security;

[0010] S4: Input the campus security spatiotemporal graph into a graph convolutional recurrent neural network and output high-dimensional spatiotemporal features;

[0011] S5: Perform factorization aggregation on high-dimensional spatiotemporal features to determine the comprehensive behavioral representation;

[0012] S6: Based on the comprehensive behavioral representation, determine whether a predefined abnormal campus behavior has occurred; if so, output the behavior category and confidence level corresponding to the abnormal campus behavior, and proceed to step S7; otherwise, return to step S1 and continue monitoring.

[0013] S7: Triggers a real-time audible and visual alarm on a local computing device deployed at the network edge and sends an alarm message to the security center containing the time, location, behavior category, and confidence level.

[0014] S8: Upload alarm information and key behavioral analysis data generated during the anomaly identification process to the cloud-based security management platform for global situation display.

[0015] A second aspect of this invention provides a campus abnormal behavior recognition system, comprising: a processor and a memory;

[0016] The memory stores programs or instructions that can run on a processor, and when the programs or instructions are executed by the processor, they implement the steps of the campus abnormal behavior recognition method of the first aspect.

[0017] A third aspect of the present invention provides a readable storage medium on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of the campus abnormal behavior recognition method of the first aspect are implemented.

[0018] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0019] In this embodiment of the invention, graph convolution operations are introduced into the gated unit of the Long Short-Term Memory Neural Network to achieve joint modeling of the structural relationships between nodes and the temporal features of behavior. The high-dimensional spatiotemporal features output are then factorized and aggregated to obtain a global comprehensive behavioral representation that does not rely on original pixel information. This effectively improves the overall accuracy and stability of abnormal behavior recognition in complex campus environments. Simultaneously, by combining a multi-frame stability determination mechanism with a collaborative processing approach of real-time edge-side alarms and cloud-based global situational awareness display, timely detection, reliable identification, and rapid response to abnormal campus behavior are achieved. This reduces the impact of instantaneous noise and local occlusion on the recognition results, enhancing the practicality and engineering feasibility of the campus security system in actual deployment scenarios. Attached Figure Description

[0020] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0021] Figure 1 This is a flowchart illustrating a method for identifying abnormal behavior on campus provided by an embodiment of the present invention.

[0022] Figure 2 This is a schematic diagram of the structure of a campus abnormal behavior recognition system provided in an embodiment of the present invention. Detailed Implementation

[0023] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0024] The campus abnormal behavior identification method provided by the present invention will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0025] Reference manual attached Figure 1 The diagram shows a flowchart of a campus abnormal behavior identification method provided by an embodiment of the present invention.

[0026] This invention provides a method for identifying abnormal behavior on campus, which may include the following steps:

[0027] S1: Real-time acquisition of multimodal data from the campus monitoring area, including visual surveillance video stream data and IoT sensor data stream.

[0028] The campus monitoring area refers to the spatial scope within the campus where monitoring and sensing equipment has been deployed for personnel activity monitoring and security management, including teaching buildings, dormitories, entrances and exits, perimeter fences, public passages, and other key areas. Multimodal data refers to a collection of data from different sources and in different formats that can collectively describe the same scene or behavioral state, used to comprehensively depict the campus environment and personnel behavior from multiple sensory dimensions. Visual surveillance video stream data refers to video data streams continuously collected and transmitted in real time by campus surveillance cameras, used to acquire visual information such as personnel appearance, posture, movement trajectory, and spatial distribution. IoT sensor data streams refer to time-series data collected and transmitted in real time by various IoT sensors deployed within the campus, including but not limited to infrared sensing, access control status, pressure, vibration, or other environmental and behavioral related sensing data.

[0029] It should be noted that by collecting visual surveillance video stream data and IoT sensor data stream in real time within the campus monitoring area, multi-dimensional and continuous perception of personnel behavior is achieved. This enables the system to simultaneously acquire appearance information and environmental status information, providing a complete and reliable data foundation for subsequent multimodal feature fusion and spatiotemporal correlation modeling, thereby improving the robustness and adaptability of abnormal behavior recognition in complex campus scenarios.

[0030] S2: Extract features from multimodal data to determine the feature data.

[0031] Among them, feature data refers to the structured data set obtained after feature extraction, which is used to describe the behavioral attributes of people, such as appearance, movement changes and environmental response. This feature data serves as the basic input for subsequent spatiotemporal modeling and abnormal behavior discrimination.

[0032] In one possible implementation, S2 specifically includes:

[0033] S201: Decode and extract frames from the visual surveillance video stream data to determine the static image frame sequence.

[0034] Among them, the static image frame sequence refers to the set of images arranged in chronological order obtained by extracting frames from the video stream. Each frame is used to depict the state of people in the campus monitoring scene at a certain moment.

[0035] S202: Extract discrete visual feature sequences of multiple channels of human target regions in static image frame sequences using a lightweight model based on factorized convolution.

[0036] Among them, the lightweight model based on factorized convolution refers to a model form that performs structural decomposition of traditional convolution operations through factorized projection matrices and shared basis convolutional filters, achieving effective expression of multi-channel features with fewer parameters, and is suitable for scenarios with limited computing resources. Discrete visual feature sequences refer to multi-channel discrete feature representations extracted from static image frames targeting human target regions and arranged by temporal or spatial sampling.

[0037] It should be noted that by employing a lightweight model based on factorized convolution to extract features from human target regions in static image frame sequences, the traditional high-parameter convolution operation is decomposed into a combination of factorized projection and shared-basis convolutional filters. This significantly reduces the model parameter scale and computational complexity while still maintaining an effective representation of human appearance and motion structure information. This approach is not only suitable for real-time processing on edge devices with limited computing resources, but also can stably represent the temporal variation characteristics of human targets across multiple channels in the form of discrete visual feature sequences, thus providing a more efficient and discriminative feature foundation for subsequent continuous domain modeling and spatiotemporal behavior analysis.

[0038] S203: Map the discrete visual feature sequences of each channel to a continuous space to obtain continuous domain features:

[0039]

[0040] in, Indicates the first d On the feature channel, the first n A discrete visual feature sequence of discrete sampling points N d Indicates the first d The number of discrete sampling points per channel Denotes the interpolation basis function. T This represents the total length of the continuous field. express t Time is determined by discrete visual feature sequences Continuous domain features are constructed through interpolation.

[0041] Among them, continuous domain features refer to the feature representation formed by mapping discrete sampled visual features to a continuous space through interpolation basis functions. Its domain is continuous variables, which are used to support continuous convolution and frequency domain operations.

[0042] in, It is a continuous domain characteristic with a continuous variable t as its domain, and mathematically it is represented as a continuous function.

[0043] This formula maps discretely sampled multi-channel visual features into a continuous domain representation through interpolation basis functions, enabling efficient convolution and frequency domain optimization of features in continuous space. It is the foundation for realizing factorized convolution and lightweight visual feature extraction.

[0044] S204: Perform factorization projection on continuous domain features to determine the projected features.

[0045] S205: Perform convolution operations on the projected features using a base convolution filter to determine the output of the factorized convolution operator:

[0046]

[0047] in, Indicates input features x The output response after applying the factorized convolution operator, i.e., the output of the factorized convolution operator. Pf Represents the projection matrix by factorization P Set of basis convolutional filters f Commonly parameterized factorized convolution operators Representing continuous domain features, P Represents the factorized projection matrix. f This represents a shared set of basis convolutional filters, i.e., a set of basis convolutional filters. Indicates the first c Each basis convolutional filter Indicates the first d The feature channel to the first c The mapping coefficients of a basis convolutional filter, T This indicates the transpose operation. Indicates by the first d Discrete visual feature sequence of one feature channel Continuous domain features are constructed through interpolation.

[0048] It should be noted that those skilled in the art can set the mapping coefficients according to actual needs. The size is not limited in this invention.

[0049] In this context, the base convolutional filter refers to a set of basic convolutional kernels shared by multiple feature channels in the factorized convolutional structure, used to perform convolution operations and enhance feature representation capabilities. The factorized convolutional operator output refers to the output response obtained after continuous domain features are convolved with the base convolutional filter through factorized projection, and is used as the final visual feature representation of human targets.

[0050] This formula maps high-dimensional multi-channel continuous domain features to a low-dimensional basis space through a projection matrix and completes convolution operations using a small number of shared basis filters. This significantly reduces model complexity while maintaining feature expressive power, and is the core computational form for realizing factorized convolution and lightweight visual feature extraction.

[0051] S206: Output the factorized convolution operator as the visual features of the human target region.

[0052] S207: Update the factorization parameters in the output of the factorized convolution operator by minimizing the energy function value:

[0053]

[0054] in, Representing input continuous domain features Frequency domain representation after Fourier transform This represents the frequency domain form of a base-convolution filter. Indicates the expected output response. Represents the square of the L2 norm. C This indicates the total number of base convolution filters. The frequency domain representation of the spatial weighting function. Indicates the first c Frequency domain representation of a basis convolutional filter This represents the regularization weight coefficient. Represents frequency domain convolution. This represents the square of the Frobenius norm of the matrix.

[0055] It should be noted that those skilled in the art can set the size of the regularization weight coefficient according to actual needs, and this invention does not limit it.

[0056] The factorization parameters include: the factorization projection matrix and the basis convolution filter.

[0057] This energy function achieves joint optimization of the parameters of the factorized convolution operator by minimizing the error between the convolution response and the ideal output, and by combining spatial regularization and projection matrix regularization terms. This significantly reduces model complexity and improves stability while ensuring feature representation capabilities.

[0058] S208: Based on the updated factorized convolution operator output, repeat steps S202 to S206 until the visual features of all human target regions in the visual surveillance video stream data are extracted.

[0059] S209: Perform numerical standardization on the IoT sensor data stream to obtain a standard sensor data stream.

[0060] S210: Construct a time series feature vector based on the standard sensor data stream.

[0061] Among them, the time series feature vector refers to the vectorized representation of standardized sensor data arranged in chronological order, used to describe the characteristics of environmental or behavioral states changing over time.

[0062] S211: Combine visual features and time series feature vectors to determine feature data.

[0063] It should be noted that by employing a lightweight model with factorized convolution to extract features from human targets in visual surveillance video streams, and mapping discrete visual features to the continuous domain to support efficient convolution and frequency domain optimization, the model parameter size and computational complexity are significantly reduced while maintaining feature representation capabilities. Furthermore, by combining standardized processing and time-series modeling of IoT sensor data, a unified feature representation of visual and environmental perception information is achieved, thus providing stable, compact, and discriminative multimodal feature data for subsequent spatiotemporal correlation modeling and abnormal behavior discrimination.

[0064] S3: Based on the physical layout of the campus and the visual overlapping areas, combined with feature data, construct a spatiotemporal map of campus security.

[0065] The campus physical layout refers to the distribution of physical spatial structures such as functional areas, buildings, roads, entrances and exits, and fences within the campus, including information such as area affiliation, spatial adjacency, and reachable paths, used to describe the objective spatial location relationships of different monitoring points. The visual overlap area refers to the area in space where the fields of view of multiple surveillance cameras overlap or partially cover each other. Personnel behavior within this area can be simultaneously perceived by multiple cameras, used to characterize the potential relationships between different visual nodes. The campus security spatiotemporal map refers to a graph structure model constructed using camera nodes and sensor nodes in the campus monitoring system as graph nodes, and spatial relationships and temporal evolution relationships between nodes as connecting edges, used to uniformly represent the joint changes in personnel behavior in the spatial and temporal dimensions.

[0066] In one possible implementation, S3 specifically includes:

[0067] S301: Construct a heterogeneous security node set and assign node role labels to each node in the heterogeneous security node set.

[0068] The heterogeneous security node set refers to a collection of nodes composed of different types of security sensing nodes, including camera nodes for acquiring visual information and IoT sensor nodes for acquiring environmental or behavioral data. These nodes differ in function and data format. Node role labels refer to the identification information assigned to distinguish different types of security nodes, indicating the node's category and its processing method in subsequent feature combination and graph modeling processes.

[0069] The heterogeneous security node set includes visual nodes and sensor nodes. Visual nodes correspond to camera deployment points, and sensor nodes correspond to IoT sensor deployment points. Different types of nodes use different feature combination methods in subsequent graph modeling.

[0070] S302: Construct node state features based on node role labels and feature data.

[0071] Among them, node state features refer to the feature representation used to describe the perceived state of a node at a specific time step. Its content adopts visual features, time series features, or a combination of both, depending on the node role.

[0072] Specifically, for visual nodes, visual features are used as node state features; for sensor nodes, time series feature vectors are used as node state features; and for nodes that simultaneously cover human activity areas, a joint node state representation of visual features and sensor features is constructed.

[0073] S303: Based on the campus physical layout information, construct a set of candidate spatial neighborhoods between different nodes.

[0074] Among them, the candidate spatial neighborhood set refers to the set of node pairs that may have spatial relationships, which are pre-selected based on the campus physical layout information, and is used to limit the candidate range of spatial connection edges.

[0075] Specifically, based on the campus physical layout information, candidate spatial edge sets are generated according to regional affiliation, access topology connectivity, spatial geometry / accessibility, and camera view coverage and overlap. These candidate edges are then merged, deduplicated, and labeled with their types to obtain candidate spatial neighborhood sets for each node. The candidate spatial neighborhood sets describe node pairs with potential spatial relationships, including but not limited to: nodes within the same building unit, nodes in adjacent functional areas, and nodes with potential visual coverage relationships.

[0076] S304: Within the candidate spatial neighborhood set, spatial connection edges are dynamically generated based on the consistency of multimodal data.

[0077] Among them, spatial connection edges refer to the set of potential spatial connection edges generated by the above-mentioned various spatial relationships, which serve as the basis for the generation of dynamic spatial connection edges.

[0078] Specifically, "spatial edges are only established when multiple nodes perceive the 'same type of behavior'." By aligning the compact visual features of nodes with the time-series features of sensors, a consistency index of their behavioral responses is calculated. When the consistency index meets preset conditions, spatial connection edges are constructed between corresponding nodes; otherwise, no spatial connection edges are constructed. This allows the spatial connection relationship to adaptively adjust as human behavior changes.

[0079] S305: Based on the node state characteristics of the same node across consecutive time steps, and considering the continuity of human behavior, determine whether to construct a temporal connection edge. If yes, proceed to step S306. Otherwise, use the node state characteristics of the current time step as the new temporal starting state and proceed to step S306.

[0080] Specifically, time connection edges are constructed between corresponding time steps only when the node state characteristics in adjacent time steps meet the preset behavior continuity conditions.

[0081] Specifically, in the time-dimensional modeling process, this step compares the node state characteristics of the same security node across consecutive time steps and combines this with the temporal continuity characteristics of human behavior to determine whether a temporal connection edge is established between the current time step and the previous time step. When the node state characteristics in adjacent time steps meet the preset behavioral continuity condition, a temporal connection edge is constructed between the corresponding time steps to depict the continuous evolution of the node's behavioral state. When the behavioral continuity condition is not met, no temporal connection edge is constructed with the previous time step, and the node state characteristics of the current time step are used as the new temporal starting state to participate in the subsequent spatiotemporal graph construction, thereby avoiding the incorrect association of discontinuous behaviors with the same temporal process.

[0082] S306: Construct a spatiotemporal map of campus security based on node state characteristics, spatial connection edges, and temporal connection edges.

[0083] Specifically, the node state characteristics of each camera node and sensor node at different time steps are used as graph node attributes. Spatial connection edges are used to characterize the spatial relationship between different nodes at the same time step, and temporal connection edges are used to characterize the temporal evolution relationship between the same node at different time steps. By combining and superimposing the spatial subgraphs corresponding to each time step in chronological order, a unified spatiotemporal graph structure is formed to comprehensively characterize the joint evolution characteristics of human behavior in the spatial and temporal dimensions in the campus environment.

[0084] It should be noted that by constructing a heterogeneous security node set including visual nodes and sensor nodes, and dynamically generating spatial and temporal connection edges based on node role labels and multimodal feature data, the campus security spatiotemporal map can adaptively adjust its structure according to changes in human behavior. This effectively portrays the spatial correlation and temporal continuity of human behavior in complex campus environments, avoiding redundant connections and noise propagation problems caused by static topology modeling. This provides a more accurate and stable structured representation foundation for subsequent spatiotemporal feature modeling and abnormal behavior recognition.

[0085] S4: Input the campus security spatiotemporal graph into a graph convolutional recurrent neural network and output high-dimensional spatiotemporal features.

[0086] Among them, graph convolutional recurrent neural networks refer to a deep learning model structure that combines graph convolutional networks and recurrent neural networks to simultaneously model the spatial dependencies in graph structure data and the dynamic changes in time series data. High-dimensional spatiotemporal features refer to the joint feature representation obtained after processing by graph convolutional recurrent neural networks, which simultaneously contains information on the spatial structural relationships between nodes and the dynamic information on the evolution of node states over time, and is used to comprehensively characterize the spatiotemporal characteristics of human behavior in the campus environment.

[0087] Specifically, in each gated unit of a traditional LSTM, linear transformations are replaced with graph convolution operations, thereby effectively overcoming the problem of the separation between spatial and temporal modeling in traditional methods, and providing a more comprehensive and discriminative high-dimensional spatiotemporal feature representation for the accurate identification of subsequent abnormal behaviors.

[0088] It should be noted that by combining graph convolutional networks with recurrent neural networks to construct graph convolutional recurrent neural networks, the feature extraction process can explicitly model the non-Euclidean spatial structural relationships between nodes and the dynamic dependence of features on time changes at each time step. Compared with traditional convolutional networks that only rely on regular grid structures or traditional recurrent neural networks that only characterize time series relationships, this approach can more comprehensively characterize the spatial correlation and temporal continuity of multi-node and multi-region behaviors in complex scenes, thereby extracting high-dimensional spatiotemporal features containing structural and dynamic information, significantly improving the ability to express complex abnormal behaviors and the stability of discrimination.

[0089] In one possible implementation, S4 specifically includes:

[0090] S401: Represent the spatiotemporal map of campus security as a weighted sequence of graphs arranged in chronological order.

[0091] Among them, the weighted graph sequence refers to a set of graph structures arranged in chronological order, where each time step corresponds to a weighted graph, which is used to describe the spatial relationship between campus security nodes and their weight changes at that time step.

[0092] S402: Construct a graph signal vector based on the weighted graph sequence.

[0093] In this context, the graph signal vector refers to the vectorized representation formed by organizing the node state features of all nodes in a fixed order at a certain time step, which serves as the input data for the graph convolutional recurrent neural network at that time step.

[0094] Specifically, for the weighted graph corresponding to a time step in the weighted graph sequence, the node state features of each graph node at that time step are obtained, and the node state features are arranged and combined according to a preset node index order to construct the graph signal vector for that time step. The graph signal vector is used to characterize the overall distribution of the node state features in the spatiotemporal graph of campus security at the current time step on the graph structure. The node state features of each node at the time step are then organized into a graph signal vector. As input to the graph convolutional recurrent neural network at each time step.

[0095] S403: Construct a normalized graph Laplacian matrix based on the weighted adjacency matrix corresponding to each time step.

[0096] The weighted adjacency matrix is ​​a matrix used to represent the connection relationships and strengths between nodes in a graph. Its elements characterize the existence and weights of spatial connections between nodes. The normalized graph Laplacian matrix is ​​a matrix calculated from the weighted adjacency matrix and the degree matrix. It is used to normalize the graph structure to stabilize feature propagation and numerical computation during graph convolution. The formula for calculating the normalized graph Laplacian matrix is ​​a well-established technology and will not be elaborated upon here.

[0097] S404: Input the graph signal vector into the graph convolutional recurrent neural network, wherein the graph convolutional recurrent neural network includes: a graph convolutional network and a long short-term memory neural network.

[0098] Among them, graph convolutional networks refer to neural network modules that use the graph Laplacian matrix to perform neighborhood aggregation and update of node features, and are used to extract spatial correlation features between nodes. Long Short-Term Memory (LSTM) neural networks refer to a recurrent neural network structure that can pass and update hidden states in the time dimension through a gating mechanism, and are used to model the dynamic process of features changing over time.

[0099] S405: Combining the normalized graph Laplacian matrix, a graph convolution operation is performed on the graph signal vector through a graph convolution network to determine the spatial correlation features between different nodes.

[0100] Among them, spatial association features refer to the feature representations obtained through graph convolution operations, which are used to characterize the structural relationships formed between different security nodes due to spatial adjacency or functional association.

[0101] Specifically, firstly, using the normalized graph Laplacian matrix... The node features at the current time step, i.e. the graph signal vector, are weighted and aggregated so that each node's features are simultaneously integrated with the feature information of its neighboring nodes when updating, thereby introducing the spatial structural relationship between nodes. Then, the aggregated features are multiplied by the learnable weight matrix, and the features are linearly transformed and mapped to obtain a high-dimensional node representation containing graph structural information. This result is the output of the graph convolution operator on the input features.

[0102] S406: Use the output of the graph convolutional network, i.e. the spatial association features, as the input features of the long short-term memory neural network at the current time step.

[0103] S407: Extract the hidden states of spatially correlated features at each time step using a Long Short-Term Memory neural network.

[0104]

[0105]

[0106] in, Indicates the current time step t Input gate at time, This represents the Sigmoid activation function. Represents the normalized graph Laplacian matrix. x t Indicates the current time step t The graph signal vector, Represents the input graph signal vector x t The learnable weight matrix to the input gate, Represents the input graph signal vector x t The learnable weight matrix to the forget gate, Represents the input graph signal vector x t Learnable weight matrix for candidate memories, Represents the input graph signal vector x t The learnable weight matrix to the output gate, Indicates the previous time step t The hidden state when -1, Indicates hidden state The weight matrix to the input gate, Indicates hidden state The weight matrix to the forget gate, Indicates hidden state The weight matrix of candidate memories, Indicates hidden state The weight matrix to the output gate, This represents the bias term of the input gate. The bias term representing the forget gate. The bias term representing candidate memories. This represents the bias term of the output gate. f t Indicates the current time step t The Gate of Oblivion Indicates the current time step t Candidate memory states, where tanh() represents the hyperbolic tangent function. c t Indicates the current time step t The state of the memory unit at that time, This represents element-wise multiplication. o t Indicates the current time step t Output gate at time, h t Indicates the current time step t The hidden state at that time.

[0107] The hidden state refers to the internal state representation of the Long Short-Term Memory Neural Network at time step t, which is used to comprehensively reflect the feature information of the current time step and historical time steps.

[0108] These formulas, by introducing the graph Laplacian operator into the gating structure of LSTM, achieve joint modeling of node structural relationships and temporal dynamics, which is the core computational mechanism of graph convolutional recurrent neural networks. By replacing the linear mapping in ordinary LSTM with graph convolution, "inter-node structural relationships and temporal dynamics" are modeled simultaneously at each time step.

[0109] S408: Combine the hidden states to determine the high-dimensional spatiotemporal features.

[0110] It should be noted that by representing the spatiotemporal graph of campus security as a weighted graph sequence arranged in chronological order, and introducing the normalized graph Laplace operator into the gating structure of the Long Short-Term Memory Neural Network, the model can simultaneously integrate the spatial structural relationships between nodes and the dynamic information of behavior changes over time at each time step. This enables joint modeling of the spatiotemporal dependence of human behavior in complex campus scenarios, avoiding the information fragmentation problem caused by the separate modeling of spatial and temporal features in traditional methods, and significantly improving the expressive power and discrimination stability of high-dimensional spatiotemporal features for abnormal behavior.

[0111] S5: Perform factorization aggregation on high-dimensional spatiotemporal features to determine the comprehensive behavioral representation.

[0112] Among them, the comprehensive behavioral representation refers to the global feature representation obtained after factorization aggregation operation, which is used to describe the behavior patterns of people in the campus scene from an overall perspective. This representation no longer contains the original image pixels or local visual feature information, but rather abstract features at the behavioral level.

[0113] The factorization aggregation operation is used to map high-dimensional spatiotemporal features to a low-dimensional compact behavioral basis vector space. The high-dimensional spatiotemporal features are the joint features output by the graph convolutional recurrent neural network on the spatiotemporal graph, which no longer contain the original image pixels or local visual feature information.

[0114] Factorization aggregation is used to structurally integrate behavioral-level spatiotemporal features to obtain a global behavioral representation for abnormal behavior discrimination, rather than to reduce the computational complexity of the visual feature extraction stage.

[0115] It should be noted that by factorizing and aggregating the high-dimensional spatiotemporal features output by the graph convolutional recurrent neural network, the behavioral information scattered across different nodes and time steps is structurally integrated to obtain a compact and globally consistent comprehensive behavioral representation. This effectively reduces the interference of redundant features on anomaly detection and improves the system's ability to stably identify and accurately distinguish abnormal behaviors in complex campus scenarios.

[0116] In one possible implementation, S5 specifically includes:

[0117] S501: Treat high-dimensional spatiotemporal features as graph signals, filter and aggregate the graph signals to determine the initial spectral aggregation features:

[0118]

[0119] in, Indicates the initial spectral aggregation characteristics. This represents the graph signal corresponding to high-dimensional spatiotemporal features. L Represents the graph Laplace matrix. Represents the spectral domain filter kernel (learnable function). U Represents the Laplacian eigenvector matrix. Let represent the diagonal matrix of Laplace eigenvalues.

[0120] Graph signals refer to the signal representation formed by organizing high-dimensional spatiotemporal features according to the order of nodes in a security spatiotemporal graph, enabling node features to undergo filtering and aggregation operations under graph structure constraints. Initial spectral aggregation features refer to the aggregation features obtained after the graph signal is processed by a spectral domain filtering kernel, used to characterize the high-dimensional spatiotemporal information aggregated under graph structure constraints.

[0121] This formula maps the graph signal to the graph spectral domain by performing spectral decomposition on the graph Laplacian operator and applies learnable frequency domain filtering, thereby achieving effective aggregation and dimensionality reduction of high-dimensional spatiotemporal features under graph structure constraints, and obtaining a comprehensive behavioral representation for abnormal behavior discrimination.

[0122] S502: Constructing a factorized spectral domain aggregation operator based on initial spectral aggregation features:

[0123]

[0124] in, Represents the graph convolution operator. This indicates element-wise multiplication.

[0125] Among them, the factorized spectral domain aggregation operator refers to the form of convolution operator implemented in the graph spectral domain by frequency-wise multiplication, which is used to perform learnable factorized aggregation of graph signals.

[0126] This formula defines convolution operations on graph structures by mapping graph signals to the graph spectral domain and performing frequency-wise Hadamard products, thus providing a mathematical basis for learnable factorization aggregation in the graph Laplace spectral domain.

[0127] S503: Chebyshev polynomial low-dimensional parameterization of the spectral domain filter kernel in the factorized spectral domain aggregation operator is performed to construct a compact behavioral basis space.

[0128]

[0129] in, Indicates the first k Learnable coefficients of a Chebyshev polynomial of order 1 This indicates the order of the Chebyshev expansion. Indicates the first k Chebyshev polynomial of order 1 This represents the scaled diagonal matrix of Laplacian eigenvalues.

[0130] It should be noted that those skilled in the art can set it according to actual needs. The size is not limited in this invention.

[0131] The spectral domain filter kernel refers to a learnable function defined in the graph Laplace spectral domain, used to weight and modulate the graph signal at different graph frequency components to control the propagation range and smoothness of information on the graph structure. The compact behavioral basis space refers to the low-dimensional feature space defined by the spectral domain filter kernel parameterized by Chebyshev polynomials, used to characterize high-dimensional spatiotemporal behavioral patterns with a finite number of learnable parameters.

[0132] This formula achieves low-dimensional parameterization of graph frequency domain filtering behavior by representing the spectral domain filter kernel as a linear combination of finite-order Chebyshev polynomials, thereby mapping high-dimensional spatiotemporal features to a compact behavioral basis vector space for subsequent abnormal behavior discrimination.

[0133] Spectral domain filter kernel A learnable frequency response function is defined in the graph Laplace spectral domain to weightedly modulate the graph frequency components corresponding to different Laplace eigenvalues ​​of the graph signal, thereby controlling the smoothness and information propagation range of high-dimensional spatiotemporal features in the graph structure. To avoid explicit graph Laplace eigenvalue decomposition and reduce the parameter scale, the spectral domain filter kernel is approximated in the scaled spectral domain using a finite-order Chebyshev polynomial, making the filter kernel consist of K learnable coefficients. Parameterization.

[0134] S504: Based on the compact behavioral base space, the synthesized behavioral representation is determined by performing a Chebyshev expansion operation on a scaled Laplacian.

[0135]

[0136] in, This represents the scaled graph Laplacian matrix. y This represents a comprehensive behavioral representation.

[0137] It should be noted that by treating high-dimensional spatiotemporal features as graph signals and performing factorization aggregation in the graph Laplacian spectral domain, the integration process of behavioral features is constrained by the graph structure, effectively preserving the spatial correlation information between nodes. Simultaneously, by using Chebyshev polynomials to perform low-dimensional parameterization of the spectral domain filter kernel, computational complexity is significantly reduced while avoiding explicit feature decomposition. This results in a compact, stable, and globally discriminative comprehensive behavioral representation, improving the accuracy and robustness of abnormal behavior recognition in complex campus scenarios.

[0138] S6: Based on the comprehensive behavioral representation, determine whether a predefined abnormal campus behavior has occurred. If yes, output the corresponding behavior category and confidence level, and proceed to step S7. Otherwise, return to step S1 and continue monitoring.

[0139] The predefined abnormal campus behaviors refer to the set of abnormal behavior categories set before system deployment based on campus security management needs, including but not limited to fighting, climbing over fences, people falling to the ground, and tailgating, which are all behaviors with security risks. Confidence level refers to a numerical indicator used to quantify the probability that the current behavior belongs to a certain abnormal behavior category. It is usually obtained from the probability or score output by the discriminant model and is used to reflect the reliability of the identification results.

[0140] It should be noted that by uniformly identifying predefined abnormal campus behaviors based on comprehensive behavioral representations and outputting corresponding behavior categories and confidence levels for the identification results, the system can achieve real-time judgment of abnormal behaviors while ensuring the interpretability of the judgment results. When no abnormal behavior is detected, it automatically returns to continuous monitoring, thereby avoiding unnecessary alarm triggering and improving the stability and practicality of the campus security system during long-term operation.

[0141] In one possible implementation, S6 specifically includes:

[0142] S601: Construct multiple anomaly detection heads by combining a predefined set of abnormal campus behavior categories.

[0143] Specifically, for each type of predefined abnormal behavior, a corresponding anomaly detection head is constructed. The anomaly detection head is used to determine whether the current comprehensive behavior representation belongs to that abnormal behavior.

[0144] The predefined set of abnormal behavior categories refers to the set of abnormal behavior types pre-defined during the system design phase based on campus security management needs. This set clarifies the types of abnormal events that the system needs to identify and distinguish. The anomaly discrimination header refers to the discrimination model module built for a specific abnormal behavior category. It is used to analyze the comprehensive behavioral representation and determine whether the behavior belongs to the corresponding abnormal category.

[0145] S602: Input the comprehensive behavioral representation into each anomaly detection head to determine the probability of occurrence of the corresponding abnormal campus behavior.

[0146] Specifically, the comprehensive behavioral representation is used as a unified input and fed into the anomaly discrimination head constructed for each type of predefined abnormal behavior. In this process, the behavioral representation is linearly or nonlinearly mapped to obtain the discrimination score for the corresponding abnormal behavior. Subsequently, the score is mapped to an interval using the Sigmoid activation function. The current behavior belongs to the first m Probability of occurrence of abnormal behavior This enables parallel probability assessment and confidence calculation for multiple types of abnormal campus behaviors.

[0147] S603: Calculate multi-frame stability based on the occurrence probability:

[0148]

[0149] in, Indicates the current time step t Time m Multi-frame stability of anomalous behaviors Indicates at time step No. mThe probability of occurrence of abnormal behavior. I ( ) indicates an indicator function. Indicates the first m The probability threshold for classifying abnormal behaviors. W Indicates the number of time steps. k Indicates the time step index.

[0150] It should be noted that those skilled in the art can set the probability determination threshold according to actual needs, and this invention does not limit it.

[0151] Specifically, multi-frame stability is used to quantify the degree of persistence of anomalous behavior over a continuous time span, by statistically analyzing near-term trends. W The proportion of anomalies occurring in each time step that exceeds a threshold is used to assess the temporal stability of abnormal behavior, thereby effectively suppressing false alarms caused by transient noise and serving as an important intermediate parameter for anomaly triggering and confidence determination.

[0152] S604: Set the corresponding multi-frame stability trigger threshold for each abnormal campus behavior in the set of abnormal campus behavior categories.

[0153] It should be noted that those skilled in the art can set the size of the multi-frame stability trigger threshold according to actual needs, and this invention does not limit it.

[0154] S605: Based on the multi-frame stability and the multi-frame stability trigger threshold, determine whether a predefined abnormal campus behavior has occurred. If yes, output the behavior category and confidence level corresponding to the abnormal campus behavior, and proceed to step S7. Otherwise, return to step S1 and continue monitoring.

[0155] Specifically, when the multi-frame stability corresponding to any abnormal behavior is not less than its corresponding multi-frame stability trigger threshold, it is determined that the first abnormal behavior has occurred in the current scene. m The system identifies abnormal campus behaviors and outputs the corresponding abnormal behavior category and its confidence level, then proceeds to step S7. When the multi-frame stability of all abnormal behaviors is less than its corresponding multi-frame stability trigger threshold, the current scene is determined to be in a normal state, and the system returns to step S1 to continue monitoring.

[0156] It should be noted that by constructing an independent anomaly detection head for each type of predefined abnormal campus behavior and combining it with a multi-frame stabilization mechanism to evaluate the persistence of abnormal behavior in the time dimension, the system can effectively suppress false alarms caused by instantaneous noise or occasional actions while identifying multiple types of abnormal behavior in parallel. This improves the reliability and stability of the abnormal behavior judgment results and enhances the practicality and robustness of the campus security system in actual long-term operation.

[0157] In one possible implementation, abnormal campus behavior includes: fighting, climbing over fences, people falling to the ground, and following others into the school.

[0158] S7: Triggers a real-time audible and visual alarm on a local computing device deployed at the network edge and sends an alarm message to the security center containing the time, location, behavior category, and confidence level.

[0159] Local computing devices at the network edge refer to computing devices deployed at or near the data source location within the campus monitoring site. These devices perform real-time analysis and response processing near the data generation point, avoiding the transmission of all data to a remote server. Real-time audible and visual alarms refer to alarm methods that immediately alert on-site personnel through sound and light warning devices upon detecting abnormal behavior on campus, attracting the attention of nearby personnel or security staff. Alarm information refers to a structured set of information describing abnormal events, including at least the time, location, and corresponding behavior category of the abnormal behavior, enabling security personnel to quickly understand the situation.

[0160] It should be noted that by directly triggering real-time audible and visual alarms on local computing devices at the network edge and simultaneously sending alarm information including time, location, and behavior type to the security center, abnormal behavior can be responded to on-site and centrally handled as soon as it is discovered. This significantly reduces system alarm latency, reduces dependence on network bandwidth, and improves the response efficiency and reliability of the campus security system in the process of handling emergencies.

[0161] S8: Upload alarm information and key behavioral analysis data generated during the anomaly identification process to the cloud-based security management platform for global situation display.

[0162] Key behavioral analysis data refers to the data generated during the abnormal behavior identification process to assist in analysis and decision-making, including comprehensive behavioral representations, multi-frame stability indicators, anomaly detection results, and related timestamps and node identifiers. The cloud-based security management platform refers to a centralized security management system deployed in a cloud computing environment, used to uniformly receive, store, analyze, and manage alarm information and behavioral analysis data from multiple edge devices. Global situational awareness display refers to a method of summarizing and visualizing abnormal events within the campus based on time and spatial dimensions, used to intuitively reflect the overall security status of the campus and the distribution of abnormal behaviors.

[0163] In one possible implementation, S8 specifically includes:

[0164] S801: Unify the encapsulation of alarm information and key behavior analysis data, and determine the encapsulation information.

[0165] Among them, encapsulated information refers to the data set formed after unified encapsulation, which contains alarm information and analysis data required to fully describe abnormal events, and is used for interaction between different modules of the system.

[0166] The alarm information includes at least: the type of abnormal behavior, the time of occurrence, the spatial location identifier, and the probability of occurrence (confidence level). Key data includes: the comprehensive behavior representation vector, the probability of occurrence of each abnormal behavior, the multi-frame stability index, the final determined type of abnormal behavior, the timestamp of the abnormal occurrence, and the corresponding spatial node or region identifier.

[0167] S802: Uploads the encapsulated information to the cloud-based security management platform via a network communication interface.

[0168] Among them, the network communication interface refers to the communication interface or protocol implementation used for data transmission between local computing devices and cloud security management platforms, which supports the reliable uploading of encapsulated information.

[0169] S803: Summarizes and statistically analyzes encapsulated information according to time and space dimensions to represent the global situation.

[0170] It should be noted that by uniformly encapsulating alarm information and key behavior analysis data and uploading them to the cloud-based security management platform, and by summarizing and statistically analyzing the relevant data in both time and space dimensions, centralized management and overall situational awareness of abnormal campus behavior are achieved. This enables security personnel to grasp the campus security status from a holistic perspective, thereby enhancing the collaborative management and decision support capabilities of the campus security system under large-scale deployment and long-term operation conditions.

[0171] In this embodiment of the invention, graph convolution operations are introduced into the gated unit of the Long Short-Term Memory Neural Network to achieve joint modeling of the structural relationships between nodes and the temporal features of behavior. The high-dimensional spatiotemporal features output are then factorized and aggregated to obtain a global comprehensive behavioral representation that does not rely on original pixel information. This effectively improves the overall accuracy and stability of abnormal behavior recognition in complex campus environments. Simultaneously, by combining a multi-frame stability determination mechanism with a collaborative processing approach of real-time edge-side alarms and cloud-based global situational awareness display, timely detection, reliable identification, and rapid response to abnormal campus behavior are achieved. This reduces the impact of instantaneous noise and local occlusion on the recognition results, enhancing the practicality and engineering feasibility of the campus security system in actual deployment scenarios.

[0172] Reference manual attached Figure 2 The diagram shows a structural schematic of a campus abnormal behavior recognition system provided by an embodiment of the present invention.

[0173] This invention provides a campus abnormal behavior recognition system 20, including: a processor 201 and a memory 202;

[0174] The memory 202 stores programs or instructions that can run on the processor 201. When the program or instructions are executed by the processor 201, they implement the steps of the above-described campus abnormal behavior identification method and achieve the same technical effect. To avoid repetition, the present invention will not elaborate further.

[0175] It should be understood that the processor 201 in this embodiment of the invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0176] It should also be understood that the memory 202 in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DR RAM).

[0177] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0178] It should be understood that, in various embodiments of the present invention, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0179] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0180] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0181] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0182] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0183] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0184] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0185] This invention provides a readable storage medium comprising: storing a program or instructions on the readable storage medium, wherein when the program or instructions are executed by a processor, the program or instructions implement the steps of the above-described campus abnormal behavior identification method and achieve the same technical effect. To avoid repetition, this invention will not elaborate further.

[0186] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the protection scope of the present invention.

Claims

1. A method for identifying abnormal behavior on campus, characterized in that, include: S1: Real-time acquisition of multimodal data from the campus monitoring area, wherein the multimodal data includes: visual monitoring video stream data and IoT sensor data stream; S2: Extract features from the multimodal data to determine the feature data; S3: Based on the physical layout of the campus and the visual overlap area, combined with the aforementioned feature data, construct a spatiotemporal map of campus security; S4: Input the campus security spatiotemporal graph into a graph convolutional recurrent neural network and output high-dimensional spatiotemporal features; S5: Perform factorization aggregation on the high-dimensional spatiotemporal features to determine the comprehensive behavioral representation; S6: Based on the comprehensive behavioral representation, determine whether a predefined abnormal campus behavior has occurred; if so, output the behavior category and confidence level corresponding to the abnormal campus behavior, and proceed to step S7; otherwise, return to step S1 and continue monitoring. S7: Trigger a real-time audible and visual alarm on a local computing device deployed at the network edge, and send alarm information including time, location, behavior category, and confidence level to the security center; S8: Upload the alarm information and key behavioral analysis data generated during the anomaly identification process to the cloud security management platform for global situation display; Specifically, S2 includes: S201: Decode and extract frames from the visual surveillance video stream data to determine the static image frame sequence; S202: Extract discrete visual feature sequences of multiple channels of the human target region in the static image frame sequence using a lightweight model based on factorized convolution; S203: Map the discrete visual feature sequences of each of the channels to a continuous space to obtain continuous domain features; S204: Perform factorization projection on the continuous domain features to determine the projected features; S205: Perform convolution operation on the projected features through the base convolution filter to determine the output of the factorized convolution operator; S206: Output the output of the factorized convolution operator as the visual feature of the human target region; S207: Update the factorization parameters in the output of the factorization convolution operator by minimizing the energy function value; S208: Based on the updated factorized convolution operator output, repeat steps S202 to S206 until the visual features of all human target regions in the visual surveillance video stream data are extracted. S209: Perform numerical standardization on the IoT sensor data stream to obtain a standard sensor data stream; S210: Construct a time series feature vector based on the standard sensor data stream; S211: Combine the visual features and the time series feature vector to determine the feature data.

2. The campus abnormal behavior identification method according to claim 1, characterized in that, S3 specifically includes: S301: Construct a heterogeneous security node set and assign node role labels to each node in the heterogeneous security node set; S302: Based on the node role label and the feature data, construct node state features; S303: Based on the campus physical layout information, construct a set of candidate spatial neighborhoods between different nodes; S304: Within the candidate spatial neighborhood set, spatial connection edges are dynamically generated based on the consistency of the multimodal data; S305: Based on the node state characteristics of the same node in consecutive time steps, and combined with the continuity of human behavior, determine whether to construct a time connection edge; if yes, proceed to step S306; otherwise, take the node state characteristics of the current time step as the new time sequence start state and proceed to step S306. S306: Construct the campus security spatiotemporal graph based on the node state characteristics, the spatial connection edges, and the temporal connection edges.

3. The campus abnormal behavior identification method according to claim 1, characterized in that, S4 specifically includes: S401: Represent the campus security spatiotemporal map as a weighted graph sequence arranged in chronological order; S402: Construct a graph signal vector based on the weighted graph sequence; S403: Construct a normalized graph Laplacian matrix based on the weighted adjacency matrix corresponding to each time step; S404: Input the graph signal vector into the graph convolutional recurrent neural network, wherein the graph convolutional recurrent neural network includes: a graph convolutional network and a long short-term memory neural network; S405: Combining the normalized graph Laplacian matrix, the graph signal vector is subjected to graph convolution operation through the graph convolution network to determine the spatial correlation features between different nodes; S406: Use the output of the graph convolutional network, i.e. the spatial correlation features, as the input features of the long short-term memory neural network at the current time step; S407: Extract the hidden state of the spatial correlation feature at each time step through the long short-term memory neural network; S408: Combine the hidden states to determine the high-dimensional spatiotemporal features.

4. The campus abnormal behavior identification method according to claim 1, characterized in that, S5 specifically includes: S501: Treat the high-dimensional spatiotemporal features as a graph signal, filter and aggregate the graph signal to determine the initial spectral aggregation features; S502: Based on the initial spectral aggregation features, construct a factorized spectral domain aggregation operator; S503: Perform Chebyshev polynomial low-dimensional parameterization on the spectral domain filter kernel in the factorized spectral domain aggregation operator to construct a compact behavioral basis space; S504: Based on the compact behavioral base space, the synthesized behavioral representation is determined by performing a Chebyshev unrolling operation on a scaled Laplacian.

5. The campus abnormal behavior identification method according to claim 1, characterized in that, S6 specifically includes: S601: Construct multiple anomaly detection headers by combining a predefined set of abnormal campus behavior categories; S602: Input the comprehensive behavioral representation into each of the anomaly detection heads to determine the probability of occurrence of the corresponding abnormal campus behavior; S603: Calculate multi-frame stability based on the occurrence probability; S604: Set a corresponding multi-frame stability trigger threshold for each abnormal campus behavior in the set of abnormal campus behavior categories; S605: Based on the multi-frame stability and the multi-frame stability trigger threshold, determine whether a predefined abnormal campus behavior has occurred; if so, output the behavior category and confidence level corresponding to the abnormal campus behavior, and proceed to step S7; otherwise, return to step S1 and continue monitoring.

6. The campus abnormal behavior identification method according to claim 1, characterized in that, The abnormal behavior on campus includes: fighting, climbing over fences, people falling to the ground, and following others into the campus.

7. The campus abnormal behavior identification method according to claim 1, characterized in that, S8 specifically includes: S801: Unify the encapsulation of the alarm information and the key behavior analysis data, and determine the encapsulation information; S802: Upload the encapsulation information to the cloud security management platform via the network communication interface; S803: The encapsulation information is summarized and statistically analyzed according to the time and space dimensions to represent the global situation.

8. A campus abnormal behavior identification system, characterized in that, include: Processor and memory; The memory stores programs or instructions that can run on the processor, which, when executed by the processor, implement the steps of the campus abnormal behavior identification method as described in any one of claims 1 to 7.

9. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the campus abnormal behavior identification method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intelligent campus access control system based on artificial intelligence

    CN121170934A