Real-time abnormal behavior identification method and system for campus security

By adopting lightweight spatiotemporal attention mechanism and asymmetric spatiotemporal model in campus safe real-time abnormal behavior recognition, the limitations of the existing technology in computing efficiency and long-term dependence modeling are solved, and efficient real-time action recognition and accuracy improvement are achieved.

CN119964240APending Publication Date: 2025-05-09JINAN PRESCHOOL TEACHERS COLLEGE +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510047108.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing online action detection methods have limitations in computing efficiency, inter-fragment interaction modeling and long-term dependency modeling, making it difficult to efficiently handle and identify real-time abnormal behaviors.

Method used

The key spatial characteristics of video data are extracted using a lightweight spatiotemporal attention mechanism and divided into long-term and short-term characteristics. The asymmetric spatiotemporal model is fused, and the action categories are finally identified through the classifier.

Benefits of technology

It effectively reduces the dependence on heavy backbone networks, reduces the demand for computing resources, improves the real-time deployment capabilities under limited hardware conditions, and enhances the accuracy and efficiency of action recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964240A_ABST
    Figure CN119964240A_ABST
Patent Text Reader

Abstract

The invention provides a campus security-oriented real-time abnormal behavior identification method and system, and belongs to the technical field of image identification. According to the scheme, the method comprises the following steps: acquiring video data of camera equipment in a campus, dividing the video data into a plurality of discrete time slices, and encoding to obtain a feature representation sequence; key spatial features of the feature representation sequence are extracted by adopting a lightweight spatiotemporal attention mechanism and divided into long-time-sequence features and short-time-sequence features, and the long-time-sequence features and the short-time-sequence features are input into the asymmetric spatiotemporal model to obtain long-time historical features and short-time historical features; and fusing the long-term historical features and the short-term historical features by using a long-term and short-term fusion module, and identifying the type of the action which is occurring by using a classifier. By means of the method, efficient online action detection can be achieved, dependence on a traditional heavy backbone network is effectively reduced, the requirement for computing resources is greatly reduced, and the method is more suitable for real-time deployment under the limited hardware condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and in particular relates to a real-time abnormal behavior recognition method and system for campus safety. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] As campus safety requirements continue to increase, real-time video analysis technology is increasingly used in educational environments. Among them, end-to-end abnormal behavior recognition systems have become a key tool for ensuring campus safety and can be used to promptly detect and prevent potential security threats. The core task of such systems is real-time abnormal behavior detection, which can quickly identify abnormal or dangerous behaviors in the video based only on current and historical information, thereby achieving immediate response and processing of campus safety incidents. Compared with traditional methods, this end-to-end system integrates the entire process from video input to behavior recognition, providing a more complete and efficient security monitoring solution.

[0004] This task faces two main challenges: First, the method must be able to effectively extract contextual information related to the current moment from a large number of historical frames to understand the action currently being performed. Second, the method needs to process video frames at a high frame rate to avoid delays and accumulation of unprocessed frames. Current research has made significant progress in innovative algorithms and network designs, significantly reducing latency and improving accuracy. These existing technologies usually adopt a unified paradigm, using a heavy and frozen video backbone network to process local segments, and then using lightweight temporal modules to achieve interactions between segments. However, these methods have significant limitations in training efficiency and allocation of modeling capabilities. First, the heavy video backbone network consumes a lot of computing resources, resulting in the need to freeze its parameters to reduce the computational burden, which limits its optimization for specific tasks. Second, although the heavy backbone network is suitable for modeling long-term dependencies, previous methods often apply it to local feature extraction, while modeling long-term dependencies through lightweight temporal modules. This leads to the lack of modeling capabilities of these methods for long-term dependencies.

[0005] It can be seen that the existing online action detection methods have limitations in computational efficiency, interaction modeling between clips, and long-term dependency modeling. There is an urgent need for a more efficient solution that can process both short-term and long-term historical features to improve the efficiency and accuracy of intelligent information extraction in online real-time video processing. Summary of the invention

[0006] In order to overcome the deficiencies of the above-mentioned prior art, the present invention provides a real-time abnormal behavior identification method and system for campus security, which effectively reduces the dependence on traditional heavy backbone networks, greatly reduces the demand for computing resources, and is more suitable for real-time deployment under limited hardware conditions.

[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0008] The first aspect of the present invention provides a real-time abnormal behavior identification method for campus safety;

[0009] A real-time abnormal behavior recognition method for campus security, comprising:

[0010] Obtain video data from a camera device on campus, divide the video data into multiple discrete time segments, and encode each discrete time segment to obtain a feature representation sequence;

[0011] A lightweight spatiotemporal attention mechanism is used to extract key spatial features of feature representation sequences;

[0012] The key spatial features are divided into long-term features and short-term features, and input into the asymmetric spatiotemporal model to identify the action category that is taking place;

[0013] Among them, the asymmetric spatiotemporal model uses the long time series branch to process the long time series features, and uses the short time series branch to process the short time series features, so as to obtain long-term historical features and short-term historical features respectively; the long-term historical features and short-term historical features are fused using the long-term and short-term fusion module, and the fused features are sent to the classifier to identify the category of the action occurring.

[0014] As a further technical solution, the video data obtained from the camera equipment on campus is:

[0015] V={v t |t∈[-N,0]};

[0016] In which, each segment v t By frame f t,i is composed of, i∈[0,τ]; N is the number of observed segments, and τ is the number of frames in each segment.

[0017] As a further technical solution, the process of encoding each discrete time segment to obtain a feature representation sequence is:

[0018] Uniformly sample each discrete time segment and divide the discrete time segment into a number of small blocks;

[0019] Projecting the small blocks into block embeddings to obtain feature representations for each discrete time segment;

[0020] A classification tag is added to each discrete time segment, and the feature representation of each discrete time segment is aggregated and encoded into a feature representation sequence.

[0021] As a further technical solution, the lightweight spatiotemporal attention mechanism adopts a spatial multi-head attention module, which is:

[0022]

[0023] in, To extract key spatial features, are independent of each other at different points in time.

[0024] As a further technical solution, the process of extracting key spatial features of the feature representation sequence using a lightweight spatiotemporal attention mechanism also includes:

[0025] The extracted key spatial features are stored in buffer M t In the embodiment, the buffer zone is as follows:

[0026]

[0027] Where T represents the memory length, t represents the timestamp; M t As a first-in, first-out queue, it continuously accepts the latest fragments and discards the oldest fragments.

[0028] As a further technical solution, the process of using long time series branches to process long time series features and obtain long-term historical features is as follows:

[0029] Get long time series features from the buffer in T L Indicates the length of long-term memory;

[0030] The acquired long-term time series features are input into the spatiotemporal multi-head pooling attention module to obtain long-term historical features. The formula is:

[0031]

[0032] In the formula, Represents long-term historical characteristics; T L ′, and represents the resolution of the compressed history representation; sg(·) is the stopping gradient operator.

[0033] As a further technical solution, the process of fusing the long-term historical features and the short-term historical features using the long-term and short-term fusion module is as follows:

[0034] Using short-term historical features as queries and long-term historical features as keys and values; the process of fusion function is as follows:

[0035]

[0036] In the formula, and have the same dimensions, located at In the feature space; Represents short-term historical characteristics; Represents the fused features.

[0037] A second aspect of the present invention provides a real-time abnormal behavior recognition system for campus safety.

[0038] A real-time abnormal behavior recognition system for campus security, comprising:

[0039] The block embedding module is configured to: obtain video data from a camera device on campus and divide the video data into multiple discrete time segments; encode each discrete time segment to obtain a feature representation sequence;

[0040] The fragment buffer module is configured to: use a lightweight spatiotemporal attention mechanism to extract key spatial features of the feature representation sequence;

[0041] The asymmetric spatiotemporal modeling module is configured to: divide the key spatial features into long-term features and short-term features, and input them into the asymmetric spatiotemporal model to identify the action category that is taking place;

[0042] Among them, the asymmetric spatiotemporal model uses the long time series branch to process the long time series features, and uses the short time series branch to process the short time series features, so as to obtain long-term historical features and short-term historical features respectively; the long-term historical features and short-term historical features are fused using the long-term and short-term fusion module, and the fused features are sent to the classifier to identify the category of the action occurring.

[0043] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in a real-time abnormal behavior identification method for campus security as described in the first aspect of the present invention.

[0044] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in a real-time abnormal behavior identification method for campus safety as described in the first aspect of the present invention are implemented.

[0045] One or more of the above technical solutions have the following beneficial effects:

[0046] The present invention uses a spatial multi-head attention module to extract key spatial features of feature representation sequences. The spatial multi-head attention module can enhance the spatial semantic understanding of discrete time segments, and at the same time use the buffer area to cache the segment features enhanced by multi-head self-attention. These features are decoupled in time and semantically rich, which not only avoids repeated calculations, but also reduces the modeling burden of subsequent modules.

[0047] An asymmetric spatiotemporal model is used to obtain long-term and short-term historical features; short-term history provides important information that is closely adjacent to the current moment, which is crucial for understanding ongoing actions. The spatiotemporal multi-head pooling attention module is used to obtain long-term historical features. The spatiotemporal multi-head pooling attention module introduces a 3D convolutional layer after each self-attention layer to downsample redundant long-term historical features, filter noise, and improve the compactness of information. By effectively integrating the recent dynamics of short-term history with the broad evolution trend of long-term history in long-term and short-term fusion, the accuracy of action recognition is enhanced.

[0048] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0050] Figure 1 This is a flow chart of the method of the first embodiment.

[0051] Figure 2 It is a system structure diagram of the second embodiment. DETAILED DESCRIPTION

[0052] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0053] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.

[0054] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0055] The present invention obtains video input from the camera equipment on campus in real time, divides it into multiple discrete time segments and encodes each segment into a compact feature representation sequence. A lightweight spatiotemporal attention mechanism is used to extract key spatial features, and the features are stored in a buffer for subsequent analysis. Finally, the action in the current time period is identified through an asymmetric spatiotemporal model, achieving efficient online action detection.

[0056] Embodiment 1

[0057] This embodiment discloses a real-time abnormal behavior recognition method for campus safety;

[0058] like Figure 1 As shown, a real-time abnormal behavior recognition method for campus security includes:

[0059] Step S1, obtaining video data from a camera device on campus, dividing the video data into multiple discrete time segments, and encoding each discrete time segment to obtain a feature representation sequence.

[0060] In step S1, the video data obtained from the camera equipment on campus is:

[0061] V={v t |t∈[-N,0]};

[0062] In which, each segment v t By frame f t,i is composed of, i∈[0,τ]; N is the number of observed segments, and τ is the number of frames in each segment.

[0063] The action probability y0∈[0,1] of the current segment v0 is calculated in real time by online action detection (OAD) C , C is the total number of action categories. Segments v1, v2, ... for future time are not available.

[0064] In the process of encoding each discrete time segment to obtain a feature representation sequence, we first encode each segment. uniformly sample τ′ frames and divide them into n h ·n w small blocks, each of which has a size of 3×τ ′ ×h×w. Among them, These small pieces are recorded as Then project it into a block embedding so that the feature representation of the entire segment is E t ={e i,j ∣1≤i≤n h ,1≤j≤n w}, where e i,j ∈R D Represents a small block c i,jSubsequently, a classification tag is added to each segment, and the feature representation of each discrete time segment is aggregated and encoded into a feature representation sequence.

[0065] In step S2, a lightweight spatiotemporal attention mechanism is used to extract key spatial features of the feature representation sequence.

[0066] Spatiotemporal modeling is crucial in many action recognition methods. However, in the online action detection (OAD) task, spatiotemporal modeling often introduces computational redundancy, mainly due to the online reasoning characteristics of the task: only one fragment is updated at each time step. Since the intermediate representation generated by spatiotemporal attention is temporally coupled, the new window needs to recalculate the features of all fragments, even if the features of these fragments have been calculated many times before. To solve this problem, in step S2, the key spatial features of the fragments enhanced by multi-head self-attention are cached in a buffer. These features are decoupled in time and semantically rich, which not only avoids repeated calculations but also reduces the modeling burden of subsequent modules.

[0067] Based on this, in step S2, the lightweight spatiotemporal attention mechanism adopts the spatial multi-head self-attention module (Spatial Multi-Head Self-Attention, MHSA s ). MHPA S The module can enhance the spatial semantic understanding of the fragment, thereby reducing the modeling burden of subsequent spatiotemporal exploration. Its formula is:

[0068]

[0069] In the formula, is the key spatial feature to be extracted. are independent of each other at different points in time.

[0070] Furthermore, in this embodiment, the buffer M t The extracted key spatial features are stored for reuse at future times, as follows:

[0071]

[0072] Where T represents the memory length and t represents the timestamp. t As a first-in-first-out queue (FIFO), it continuously receives the latest fragments and discards the oldest fragments. First-in-first-out is a characteristic of the queue data structure, which means that the elements that enter the queue first always leave first. Here, it refers to the video frame feature that enters the queue buffer early, which always leaves first because it is the oldest video feature, corresponding to M t In

[0073] In step S3, the key spatial features are divided into long-term features and short-term features, and input into the asymmetric spatiotemporal model to identify the category of the action taking place.

[0074] Short-term features provide important information closely adjacent to the current moment, which is crucial for understanding ongoing actions. In addition, sufficient computing resources are allocated in the spatiotemporal modeling of features, increasing the complexity and depth of the model in spatiotemporal modeling.

[0075] In step S3, first, from the buffer M t Retrieve short-term history As a spatiotemporal multi-head attention module Input:

[0076]

[0077] To obtain short-term historical characteristics; lie in in feature space.

[0078] Since action recognition relies heavily on temporal context information, long-term time series features provide a broad field of view and are essential for understanding the evolution of actions. However, long-term time series features are weakly correlated with current events and usually contain significant background redundancy. Therefore, in order to improve efficiency, in this embodiment, fewer attention layers are used in acquiring long-term historical features and a large proportion of spatiotemporal downsampling is introduced compared to acquiring short-term historical features.

[0079] Specifically, from the buffer M t Get long-term history in T L Indicates the length of long-term memory. In the back propagation process, this method This separation is necessary because it has been empirically observed that omitting gradient truncation can lead to training convergence problems.

[0080] Subsequently, this method will Input to the Spatial Multi-Head Pooling-Attention module (MHPA) middle. Relative to MHSA st A 3D convolution layer is introduced after each self-attention layer to downsample redundant long-term historical features, filter noise and improve the compactness of information. The formula is as follows:

[0081]

[0082] Here, T L ′ , and represents the resolution of the compressed history representation. sg(·) is the “stop gradient” operator.

[0083] Furthermore, in step S3, the long-term historical features and the short-term historical features are fused using a long-term and short-term fusion module. The long-term and short-term fusion module can effectively integrate the recent dynamics of the short-term history with the broad evolution trend of the long-term history. The module uses the short-term historical features as the query and the long-term historical features as the key and value. The fusion function process is as follows:

[0084]

[0085] In the formula, and have the same dimensions, located at In the feature space; Represents short-term historical characteristics; Represents the fused features.

[0086] Finally, in The classification mark of the current frame is taken and sent to the classifier, which includes a fully connected layer and a Softmax activation function. The classification mark is first passed to a fully connected layer, and the fully connected layer is used to map the features to the feature space. Then, the Softmax activation function is applied to convert the output into a probability distribution, so that each category has a corresponding probability value, thereby completing the classification task and obtaining the category of the action that is occurring.

[0087] Embodiment 2

[0088] This embodiment discloses a real-time abnormal behavior recognition system for campus safety;

[0089] like Figure 2 As shown, a real-time abnormal behavior recognition system for campus security includes:

[0090] The block embedding module is configured to: obtain video data from a camera device on campus and divide the video data into multiple discrete time segments; encode each discrete time segment to obtain a feature representation sequence;

[0091] The fragment buffer module is configured to: use a lightweight spatiotemporal attention mechanism to extract key spatial features of the feature representation sequence;

[0092] The asymmetric spatiotemporal modeling module is configured to: divide the key spatial features into long-term features and short-term features, and input them into the asymmetric spatiotemporal model to identify the action category that is taking place;

[0093] Among them, the asymmetric spatiotemporal model uses the long time series branch to process the long time series features, and uses the short time series branch to process the short time series features, so as to obtain long-term historical features and short-term historical features respectively; the long-term historical features and short-term historical features are fused using the long-term and short-term fusion module, and the fused features are sent to the classifier to identify the category of the action occurring.

[0094] Embodiment 3

[0095] The purpose of this embodiment is to provide a computer-readable storage medium.

[0096] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a real-time abnormal behavior identification method for campus security as described in Example 1 of the present disclosure.

[0097] Embodiment 4

[0098] The purpose of this embodiment is to provide an electronic device.

[0099] An electronic device comprises a memory, a processor and a program stored in the memory and executable on the processor. When the processor executes the program, the steps in a real-time abnormal behavior identification method for campus security as described in Example 1 of the present disclosure are implemented.

[0100] The steps involved in the apparatuses of the above embodiments 2, 3 and 4 correspond to the method embodiment 1, and the specific implementation methods can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0101] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0102] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A real-time abnormal behavior recognition method for campus safety, characterized in that: include: Obtain video data from a camera device on campus, divide the video data into multiple discrete time segments, and encode each discrete time segment to obtain a feature representation sequence; A lightweight spatiotemporal attention mechanism is used to extract key spatial features of feature representation sequences; The key spatial features are divided into long-term features and short-term features, and input into the asymmetric spatiotemporal model to identify the action category that is taking place; Among them, the asymmetric spatiotemporal model uses the long time series branch to process the long time series features, and uses the short time series branch to process the short time series features, so as to obtain long-term historical features and short-term historical features respectively; the long-term historical features and short-term historical features are fused using the long-term and short-term fusion module, and the fused features are sent to the classifier to identify the category of the action occurring.

2. A real-time abnormal behavior identification method for campus safety as claimed in claim 1, characterized in that: The video data of the camera equipment on campus is obtained as follows: V={v t ∣t∈[-N,0]}; In which, each segment v t By frame f t,i is composed of, i∈[0,τ]; N is the number of observed segments, and τ is the number of frames in each segment.

3. A real-time abnormal behavior identification method for campus safety as claimed in claim 1, characterized in that: The process of encoding each discrete time segment to obtain a feature representation sequence is as follows: Uniformly sample each discrete time segment and divide the discrete time segment into a number of small blocks; Projecting the small blocks into block embeddings to obtain feature representations for each discrete time segment; A classification tag is added to each discrete time segment, and the feature representation of each discrete time segment is aggregated and encoded into a feature representation sequence.

4. A real-time abnormal behavior identification method for campus safety as claimed in claim 1, characterized in that: The lightweight spatiotemporal attention mechanism adopts a spatial multi-head attention module, which is: in, To extract key spatial features, are independent of each other at different points in time.

5. A real-time abnormal behavior identification method for campus safety as claimed in claim 1, characterized in that: The process of extracting key spatial features of the feature representation sequence using a lightweight spatiotemporal attention mechanism also includes: The extracted key spatial features are stored in buffer M t In the embodiment, the buffer zone is as follows: Where T represents the memory length, t represents the timestamp; M t As a first-in, first-out queue, it continuously accepts the latest fragments and discards the oldest fragments.

6. A real-time abnormal behavior identification method for campus safety as claimed in claim 1, characterized in that: The process of using long time series branches to process long time series features and obtain long-term historical features is as follows: Get long time series features from the buffer in T L Indicates the length of long-term memory; The acquired long-term time series features are input into the spatiotemporal multi-head pooling attention module to obtain long-term historical features. The formula is: In the formula, Indicates long-term historical characteristics; T L ′ , and represents the resolution of the compressed history representation; sg(·) is the stopping gradient operator.

7. A real-time abnormal behavior identification method for campus safety as claimed in claim 1, characterized in that: The process of fusing the long-term historical features and the short-term historical features using the long-term and short-term fusion module is as follows: Using short-term historical features as queries and long-term historical features as keys and values; the process of fusion function is as follows: In the formula, and have the same dimensions, located In the feature space; Represents short-term historical characteristics; Represents the fused features.

8. A real-time abnormal behavior recognition system for campus safety, characterized in that: include: The block embedding module is configured to: obtain video data from a camera device on campus and divide the video data into a plurality of discrete time segments; encode each discrete time segment to obtain a feature representation sequence; The fragment buffer module is configured to: use a lightweight spatiotemporal attention mechanism to extract key spatial features of the feature representation sequence; The asymmetric spatiotemporal modeling module is configured to: divide the key spatial features into long-term features and short-term features, and input them into the asymmetric spatiotemporal model to identify the action category that is taking place; Among them, the asymmetric spatiotemporal model uses the long time series branch to process the long time series features, and uses the short time series branch to process the short time series features, so as to obtain long-term historical features and short-term historical features respectively; the long-term historical features and short-term historical features are fused using the long-term and short-term fusion module, and the fused features are sent to the classifier to identify the category of the action occurring.

9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps of a real-time abnormal behavior identification method for campus security as described in any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the real-time abnormal behavior identification method for campus security as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Video prediction method based on time sequence correction convolution

    CN114758282A

  • Self-attention video stream compression method and system for online action detection task

    CN116170638A