Behavior event anomaly detection method, apparatus, device, medium, and program product

CN122090360BActive Publication Date: 2026-08-28CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610543533.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-23
Publication Date
2026-08-28
Estimated Expiration
2046-04-23

AI Technical Summary

Technical Problem

[0004]本申请提供一种行为事件异常检测方法、装置、设备、介质及程序产品,用以解决现有技术中在监控过程中出现难以正确判断事件异常的情况,导致漏检率较高的问题

Benefits of technology

[0018]本申请提供的行为事件异常检测方法、装置、设备、介质及程序产品,在静态的场景层面,从时空立方体中提取全局场景特征,能够全面理解场景上下文,在动态的行为层面,从视频片段中捕捉人物的行为特征,为理解事件提供语义指导,将全局场景特征和行为特征进行异构融合,学习异构特征之间的一致性关系,实现全局语境感知能力,强化了对场景依赖性复杂事件的理解能力,提升了复杂异常的识别准确性,可以更精准地识别视频中潜在的异常行为事件。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090360B_ABST
    Figure CN122090360B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and provides a behavior event anomaly detection method, device, equipment, medium and program product, the method comprising the following steps: extracting behavior features of a target object from a video segment, and extracting global scene features from a space-time cube; fusing the behavior features and the global scene features to obtain a heterogeneous feature fusion result; generating a prediction result of a future frame based on the heterogeneous feature fusion result; performing anomaly detection on the prediction result of the future frame to obtain a behavior event anomaly detection result. The behavior event anomaly detection method provided by the application extracts global scene features from a space-time cube, captures behavior features of a person from a video segment, fuses the global scene features and the behavior features, learns consistency between the heterogeneous features, realizes global context awareness, strengthens the understanding ability of scene-dependent complex events, and improves the recognition accuracy of complex anomalies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, medium, and program product for detecting abnormal behavior events. Background Technology

[0002] In the field of intelligent services, monitoring anomaly detection technology has received widespread attention in recent years due to its important value in scenarios such as safety inspection of industrial production lines, dynamic monitoring of warehousing and logistics, and home environment perception of home service robots.

[0003] For example, for robots capable of indoor patrols and providing dynamic perspective monitoring, a significant challenge lies in the fact that the same behavioral event can exhibit different characteristics depending on the context. For instance, a person falling in a hospital corridor and a patient resting in bed require specific contextual information to make a reasonable and accurate assessment of the event. Existing research primarily focuses on extracting surface features from videos, lacking global contextual awareness. This can easily lead to difficulties in correctly identifying anomalies during monitoring, resulting in a high rate of missed detections. Summary of the Invention

[0004] This application provides a method, apparatus, equipment, medium, and program product for detecting abnormal behavior events, in order to solve the problem in the prior art that it is difficult to correctly judge the abnormality of events during the monitoring process, resulting in a high rate of missed detection.

[0005] Firstly, this application provides a method for detecting abnormal behavioral events, including: Behavioral features of the target object are extracted from the video clip, and global scene features are extracted from the spatiotemporal cube; the spatiotemporal cube is constructed based on a sequence of consecutive frame images extracted from the video clip. The behavioral features and the global scene features are fused to obtain a heterogeneous feature fusion result; Based on the heterogeneous feature fusion result, a prediction result for the future frame is generated; the future frame is the next frame immediately following the last frame of the spatiotemporal cube. Anomaly detection is performed on the prediction results of the future frames to obtain behavioral event anomaly detection results.

[0006] In one embodiment, extracting global scene features from the spatiotemporal cube includes: Global feature extraction is performed on the spatiotemporal cube to obtain initial global scene features, and spatiotemporal feature extraction is performed on the spatiotemporal cube to obtain a spatiotemporal attention vector; The initial global scene features and the spatiotemporal attention vector are fused to obtain global scene features.

[0007] In one embodiment, the spatiotemporal cube comprises T frames of images, where T is a positive integer; the global feature extraction of the spatiotemporal cube to obtain initial global scene features includes: The T-frame images in the spacetime cube are merged according to the channel dimension to generate a tensor; Multi-level semantic information extraction is performed on the tensor to obtain initial global scene features.

[0008] In one embodiment, the spatiotemporal cube comprises T frames of images, where T is a positive integer; the spatiotemporal feature extraction of the spatiotemporal cube to obtain a spatiotemporal attention vector includes: Semantic information is extracted from each frame image in the spatiotemporal cube to obtain the features of each frame image; Based on the features of the T-th frame image, query features are constructed, and based on the features of the intermediate frame images, key features and value features are constructed; the intermediate frame images include the first frame image to the (T-1)-th frame image. Based on the query features, the key features, and the value features, a spatiotemporal attention vector is generated.

[0009] In one embodiment, extracting behavioral features of the target object from the video clip includes: Extracting initial behavioral features of target objects from video clips; Based on the behavioral prototype memory pool, the normal prototypes of the initial behavioral features are extracted to obtain the behavioral features of the target object; the behavioral prototype memory pool is used to store the normal prototypes of the behavioral features.

[0010] In one embodiment, the video segment comprises M frames, where M is a positive integer; the extraction of initial behavioral features of the target object from the video segment includes: Target detection is performed on each frame image in the video segment to obtain the bounding box of the target object in each frame image; Based on the bounding box of the target object in each of the frame images, a continuous action sequence of the target object in the M frame images is generated; Behavioral features are extracted from the continuous action sequence to obtain the initial behavioral features of the target object.

[0011] In one embodiment, the behavioral prototype memory pool stores a memory matrix, which consists of J memory items, where J is a positive integer. These memory items represent the normal prototypes of behavioral features. The step of extracting the normal prototypes of the initial behavioral features from the behavioral prototype memory pool to obtain the behavioral features of the target object includes: The initial behavioral features are compared with each memory item to generate a similarity vector. The similarity vectors are normalized to generate a similarity weight matrix; The behavioral features of the target object are obtained by performing a dot product between the similarity weight matrix and the memory matrix.

[0012] In one embodiment, fusing the behavioral features with the global scene features to obtain a heterogeneous feature fusion result includes: Through an attention mechanism, a scene feature attention map of the global scene features and a behavior feature attention map of the behavior features are generated. The scene feature attention map is subjected to contextual reasoning processing to obtain the scene feature structured result of fusing the related information before and after the fusion; the behavior feature attention map is subjected to contextual reasoning processing to obtain the behavior feature structured result of fusing the related information before and after the fusion. The scene feature structuring results and the behavior feature structuring results are fused to obtain heterogeneous feature fusion results.

[0013] In one embodiment, the step of performing contextual reasoning processing on the scene feature attention map to obtain a structurated result of scene features fused with related information includes: The scene feature attention map is input into the recursive network; In the current time step, based on the scene feature attention map and the hidden state of the previous time step, the scalar is updated and calculated; the scalar is used to characterize the relative importance of each pixel position in the scene feature attention map when extracting context points; Based on the hidden state of the previous time step, the gate vector is updated and calculated; the gate vector is used to assign the contribution of each pixel position in the scene feature attention map to contextual reasoning. Based on the scene feature attention map and the hidden state of the previous time step, the hidden state of the current time step is generated through a gating mechanism, and the hidden state of the current time step is passed to the next time step. The next time step is taken as the current time step. The steps of updating the scalar calculation based on the scene feature attention map and the hidden state of the previous time step are executed iteratively in the current time step until the last time step in the preset step size is reached, and the final scalar and the final gating vector in the last time step are obtained. Based on the final scalar, the final gating vector, and the scene feature attention map, a scene feature structured result is generated that integrates the related information before and after fusion.

[0014] Secondly, this application also provides a behavioral event anomaly detection device, comprising: The heterogeneous feature extraction module is used to extract behavioral features of target objects from video clips and global scene features from a spatiotemporal cube; the spatiotemporal cube is constructed based on a sequence of consecutive frame images extracted from the video clips. The heterogeneous feature fusion module is used to fuse the behavioral features with the global scene features to obtain the heterogeneous feature fusion result; The prediction module is used to generate a prediction result for a future frame based on the heterogeneous feature fusion result; the future frame is the next frame immediately following the last frame of the spatiotemporal cube. An anomaly detection module is used to perform anomaly detection on the prediction results of the future frames to obtain behavioral event anomaly detection results.

[0015] Thirdly, this application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the above-described abnormal behavior event detection methods.

[0016] Fourthly, this application also provides a non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the above-described behavioral event anomaly detection methods.

[0017] Fifthly, this application also provides a computer program product, the computer program product including a computer program that can be stored on a non-transitory computer-readable storage medium, and the computer program, when executed by the processor, implements the steps of any of the above-described behavioral event anomaly detection methods.

[0018] The behavioral event anomaly detection method, apparatus, device, medium, and program products provided in this application extract global scene features from a spatiotemporal cube at the static scene level, enabling a comprehensive understanding of the scene context. At the dynamic behavioral level, they capture behavioral features of characters from video clips, providing semantic guidance for understanding events. By heterogeneously fusing global scene features and behavioral features and learning the consistency relationship between heterogeneous features, they achieve global context perception capabilities, enhance the ability to understand scene-dependent complex events, improve the accuracy of identifying complex anomalies, and can more accurately identify potential abnormal behavioral events in videos. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating the abnormal behavior event detection method provided in this application.

[0021] Figure 2 This is a schematic diagram of the overall framework of the behavioral event anomaly detection method provided in this application.

[0022] Figure 3 This is a schematic diagram of the video preprocessing provided in this application.

[0023] Figure 4 This is a schematic diagram of the consistency scenario encoder provided in this application.

[0024] Figure 5 This is a schematic diagram of the cross-fusion decoder provided in this application.

[0025] Figure 6 This is a schematic diagram of the behavioral event anomaly detection device provided in this application.

[0026] Figure 7 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein.

[0029] The following is combined with Figures 1-7 This application describes the methods, apparatus, equipment, media, and program products for detecting abnormal behavioral events.

[0030] The behavioral event anomaly detection method provided in this application embodiment can be implemented based on a behavioral event anomaly detection device. Therefore, this application embodiment uses a behavioral event anomaly detection device as the execution subject to describe the behavioral event anomaly detection method.

[0031] Combination Figure 1 and Figure 2 , Figure 1 This is a flowchart illustrating the behavioral event anomaly detection method provided in this application. Figure 2 This is a schematic diagram of the overall framework of the behavioral event anomaly detection method provided in this application.

[0032] like Figure 1 As shown, the abnormal behavior event detection method includes the following steps: Step 101: Extract behavioral features of the target object from the video clip and extract global scene features from the spatiotemporal cube; the spatiotemporal cube is constructed based on the sequence of consecutive frame images extracted from the video clip.

[0033] Specifically, for a specific time t, a sequence of consecutive frame images over a time range [tT,t] is acquired. A three-dimensional data structure is formed by stacking consecutive frame image sequences in chronological order. This three-dimensional data structure serves as a spatiotemporal cube. The frame image at time t+1 is predicted using T frames in a continuous frame image sequence.

[0034] The length of the spatiotemporal cube is typically not set too long to avoid excessively complex scene feature changes due to a large time span, which could affect the accuracy of feature extraction. In practical applications, the value of T is flexibly adjusted according to the specific scenario and requirements, ensuring that the scene features accurately reflect the scene state at time t. For example, in an experimental setup, the length of the spatiotemporal cube could be set to 5 frames. However, such a spatiotemporal cube has a limited time span, and capturing the complete behavioral information of the target object requires a longer time frame; therefore, it is necessary to simultaneously capture a video segment with a longer time scale. Behavioral features are extracted during the process.

[0035] Each video clip Length is That is, it covers indivual For example, when When the length is 2s, it contains 6. This longer scale can be achieved in a Capture the complete action events. More importantly, all of them... exist They all have the same movement characteristics This ensures that predictions are made over a continuous period of time while focusing on behavioral information within the scene.

[0036] Therefore, for video clips By performing multi-scale partitioning, we can obtain... indivual , which covers indivual The length can be based on The length of the video segment can be divided equally, or it can be divided unequally, depending on the actual situation. The video segment is represented as... Video clip Includes M Frame image, M and T In a multiple relationship, except for one case, .

[0037] Figure 3 This is a schematic diagram of the video preprocessing provided in this application. For example... Figure 3 As shown, video clip Used to extract behavioral features, each include A spacetime cube, obtained Used to extract scene context information.

[0038] When extracting behavioral features of a target object, target detection and behavior analysis are performed on each frame of the video clip to identify and track the target object's position and actions. This process then extracts key feature vectors that describe the target object's behavior, which are used as its behavioral features. These behavioral features not only include the type of action but may also contain information such as the amplitude and speed of the action, providing behavioral semantic guidance for subsequent anomaly detection.

[0039] When extracting scene context information, features are extracted from each frame of an independent spatiotemporal cube, and these features are fused along the temporal dimension to obtain a global feature vector that describes the entire scene state. This global feature vector serves as the global scene feature. The global feature vector reflects the dynamic changes in the scene and provides important contextual information for subsequent anomaly detection.

[0040] Step 102: Fuse the behavioral features with the global scene features to obtain the heterogeneous feature fusion result.

[0041] Specifically, after obtaining behavioral features and global scene features, since these two types of features depict the video content from different dimensions, it is necessary to fuse these two types of features in order to understand the behavioral events in the video more comprehensively and accurately.

[0042] Feature fusion can employ specific fusion algorithms, such as weighted fusion algorithms and deep learning-based fusion algorithms. Weighted fusion algorithms assign different weights to behavioral features and global scene features based on their importance in anomaly detection, then sum them using weighted methods to obtain the fusion result. Deep learning-based fusion algorithms pre-build a deep neural network model, taking behavioral features and global scene features as input. The model automatically learns the complex relationships between the two types of features, thus outputting fused heterogeneous features. Because global scene features and behavioral features exhibit semantic heterogeneity, directly combining them may be detrimental to the efficient learning of automated detection equipment. Therefore, when dealing with semantically heterogeneous features, deep learning-based fusion algorithms are preferred.

[0043] By establishing joint semantic consistency between scene context information and behavioral features based on the heterogeneous feature fusion results obtained by a specific fusion algorithm, the overall state of behavioral events in the video can be reflected more effectively, providing richer and more accurate feature basis for subsequent anomaly detection.

[0044] Step 103: Based on the heterogeneous feature fusion result, generate the prediction result of the future frame; the future frame is the next frame immediately following the last frame of the spatiotemporal cube.

[0045] Specifically, the heterogeneous feature fusion result obtained based on a specific fusion algorithm serves as the final measure of the correlation between action features and scene context, and is expressed as follows: The heterogeneous feature fusion result is an abstraction and reorganization of information from the original input image, resulting in changes in spatial dimensions. To make the fused features better suited to the task requirements, the heterogeneous feature fusion result needs to be passed through a fully connected layer, an activation function, and an upsampling operation in sequence. The fully connected layer can convert the fused features into the output form required by the task; the activation function can introduce nonlinear factors to enhance the model's expressive power; and the upsampling operation can restore the spatial dimension information of the features, making them match the spatial dimension of the original input image.

[0046] After obtaining the features with restored spatial dimensions, the frame image prediction result at time t+1 is generated based on these features. This frame contains semantic information such as the possible actions of the target object at time t+1 and the current scene. The prediction process can be represented by the following formula:

[0047] in, This indicates the prediction result for future frames; This represents the result of heterogeneous feature fusion; This represents a prediction network.

[0048] Furthermore, after obtaining the features with the restored spatial dimension, they can be fused with the global features of the spatiotemporal cube through residual connections. This allows the prediction network to combine high-level context with fine-grained details, improving the plausibility of future frame predictions. The prediction process can be represented by the following equation:

[0049] in, This indicates the prediction result for future frames; This represents the result of heterogeneous feature fusion; express Global features; This represents a prediction network.

[0050] Step 104: Perform anomaly detection on the prediction results of the future frames to obtain the behavioral event anomaly detection results.

[0051] Specifically, the frame image prediction result at time t+1 is obtained. Subsequently, the real frame image at time t+1 is also acquired. .

[0052] Prediction results of frame images With real frame images Comparative analysis is performed by setting an appropriate anomaly threshold or employing a specific anomaly detection algorithm, such as a distance-based method or a statistical model-based method, to determine the degree of difference between the predicted result and the actual frame image. When the degree of difference exceeds the set threshold or meets the anomaly conditions defined by the anomaly detection algorithm, the frame image at time t+1 is determined to be abnormal, thus obtaining the behavioral event anomaly detection result.

[0053] The behavioral event anomaly detection method provided in this application extracts global scene features from a spatiotemporal cube at the static scene level, enabling a comprehensive understanding of the scene context. At the dynamic behavioral level, it captures behavioral features of characters from video clips, providing semantic guidance for understanding events. By heterogeneously fusing global scene features and behavioral features, and learning the consistency relationship between heterogeneous features, it achieves global context perception capabilities, strengthens the understanding of scene-dependent complex events, improves the accuracy of complex anomaly identification, and can more accurately identify potential abnormal behavioral events in videos.

[0054] In one embodiment, based on step 102, the extraction of global scene features from the spatiotemporal cube includes: Global feature extraction is performed on the spatiotemporal cube to obtain initial global scene features, and spatiotemporal feature extraction is performed on the spatiotemporal cube to obtain a spatiotemporal attention vector; The initial global scene features and the spatiotemporal attention vector are fused to obtain global scene features.

[0055] Specifically, to ensure the temporal consistency of global scene features and enhance temporal coherence during the prediction of future frames, this embodiment proposes a consistent scene encoder, which includes two modules: acquiring... Scene feature extractor with global features and extraction Temporal feature extractor for scene consistency features ,in A residual connection structure is adopted. It employs a convolutional neural network and a unique spatiotemporal attention mechanism. For its specific framework, please refer to [reference needed]. Figure 4 , Figure 4 This is a schematic diagram of the consistency scenario encoder provided in this application.

[0056] spacetime cube The data used as input to the consistency scene encoder are then fed into the scene feature extractor. and temporal feature extractor In the middle. Among them, Indicates the number of channels. Indicates the number of frames. and Indicates the length and width of the dimensions. Indicates video clip The first covered A spacetime cube .

[0057] Scene feature extractor By utilizing the residual connection structure, initial global feature information can be effectively extracted from the spatiotemporal cube. This feature information covers the basic scene elements in the spatiotemporal cube.

[0058] Meanwhile, time feature extractor We use convolutional neural networks to extract features from spatiotemporal cubes and obtain time-enhanced spatiotemporal attention vectors through a unique spatiotemporal attention mechanism.

[0059] Furthermore, the scene feature extractor Output of initial global features and temporal feature extractor The output spatiotemporal attention vectors are fused to obtain global scene features that are both globally representative and temporally consistent, providing strong support for accurate prediction of future frames.

[0060] The above process can be represented by the following formula:

[0061] in, This indicates a scene feature extraction operation; This represents the spatiotemporal attention feature extraction operation; Represents the spatiotemporal attention vector. In For the number of channels, For vector dimensions; Represents the initial global features. In For the number of channels, and The length and width are the dimensions; This represents global scene features with spatiotemporal consistency; This indicates a serial operation.

[0062] The embodiments of this application not only effectively capture global scene features in the spatiotemporal cube, but also enhance the feature representation in the time dimension through the spatiotemporal attention mechanism, so that the global scene features remain consistent in time, thereby more accurately reflecting the dynamic changes of behavioral events in the time series, and providing a reliable feature basis for subsequent prediction of future frame images.

[0063] In one embodiment, the step of extracting global features from the spatiotemporal cube to obtain initial global scene features includes: The T-frame images in the spacetime cube are merged according to the channel dimension to generate a tensor; Multi-level semantic information extraction is performed on the tensor to obtain initial global scene features.

[0064] Specifically, scene feature extractor spacetime cube All T-frame images are merged according to channel dimension into A tensor of dimension, where This represents the number of channels in a tensor. and These represent the height and width of the tensor, respectively.

[0065] Furthermore, a multi-layer convolutional neural network is used to extract semantic information from the generated tensor at multiple levels. The output features of the last level are used as initial global scene features for subsequent connections, and are represented as follows: .

[0066] This application embodiment reduces computational difficulty by merging multiple frames of images in the spatiotemporal cube along the channel dimension, and then uses a multi-layer convolutional neural network to extract multi-level semantic information, which can more comprehensively capture global feature information in the scene, including rich scene semantic information.

[0067] In one embodiment, the step of extracting spatiotemporal features from the spatiotemporal cube to obtain a spatiotemporal attention vector includes: Semantic information is extracted from each frame image in the spatiotemporal cube to obtain the features of each frame image; Based on the features of the T-th frame image, query features are constructed, and based on the features of the intermediate frame images, key features and value features are constructed; the intermediate frame images include the first frame image to the (T-1)-th frame image. Based on the query features, the key features, and the value features, a spatiotemporal attention vector is generated.

[0068] Specifically, time feature extractor Focusing on temporal characteristics, we first examine the spacetime cube. Semantic information is extracted from each frame of the image to obtain the features of each frame. .

[0069] Spatiotemporal attention methods are used to obtain temporally enhanced spatiotemporal attention vectors. Features from each frame of the image are then processed. Features categorized as the last frame based on the time dimension Features of intermediate frames Features of the first frame Based on features of the last frame Build query features Features based on intermediate frames Construct key features Sum value characteristics ,in , , These involve constructing weight matrices for query features, key features, and value features. This is a series operation.

[0070] Based on query features, key features, and value features, a spatiotemporal attention vector is generated. This attention calculation method can employ common scaled dot product attention, and can be represented as follows: By querying relevant positions in the previous frame, the temporal consistency of the spatiotemporal cube can be modeled, thereby enhancing attention to the current frame. In the structural diagram of the consistency scene encoder, the combined frame of the previous frame and the first frame represents changes over past time periods. By querying the current frame... Comparing the spatiotemporal cube with previous image frames ensures that the attention of the spatiotemporal cube is focused only on the previous frame. Furthermore, since the extracted frames are not aligned along the channel dimension, multiple copies of the first frame's features need to be created during the acquisition of key and value features. This ensures data dimensionality consistency and enhances the influence of the previous frame on the current frame during attention weight calculation.

[0071] To preserve pixel location information, the fused multi-frame features need to be deformed back to the original input image size, as shown in the following formula:

[0072] in, This represents the reshaping function.

[0073] The embodiments of this application not only effectively integrate temporal and spatial information, but also enhance the influence of the previous frame on the current frame through an attention mechanism, which can more accurately capture the dynamic change patterns of behavioral events in the spatiotemporal dimension.

[0074] In one embodiment, based on step 101, extracting the behavioral features of the target object from the video clip includes: Extracting initial behavioral features of target objects from video clips; Based on the behavioral prototype memory pool, the normal prototypes of the initial behavioral features are extracted to obtain the behavioral features of the target object; the behavioral prototype memory pool is used to store the normal prototypes of the behavioral features.

[0075] Specifically, the incidence of abnormal behavior events is low in the real world, and the proportion of abnormal frame images in the dataset is extremely small, resulting in an imbalance between normal and abnormal samples, which poses a challenge to methods that rely on balanced data for training.

[0076] For example, in anomaly detection during robot vision monitoring, focusing on local dynamic action information is key to achieving high-level event semantic understanding, while the diversity of normal events also presents challenges for modeling. To address this, this embodiment designs a behavior prototype memory module to assist in capturing local behavioral features and obtaining their compact representations. Since anomalous behavioral events are sparse, it is impossible to train a classifier by exhaustively listing normal behavioral events. Therefore, a memory pool method can be used to generate normal feature prototypes to guide anomaly judgment: during training, behavioral features are input into the behavior prototype memory pool, and typical normal prototypes are obtained through weighted similar memory items, forming a memory of normal patterns; during testing, the input features are compared with the memory items and re-integrated, utilizing the difference between high integration errors for anomalous events and low errors for normal events to achieve anomaly judgment. As the depth of the behavior prototype memory pool increases, the same input behavioral features can be extracted and memorized by the behavior prototype memory pool. Ultimately, after training, the behavior prototype memory module can obtain compact prototypes from a large amount of normal action information, thereby effectively filtering and focusing on input action features.

[0077] First, target detection is performed on each frame of the video clip, and spatiotemporal segments of the target object region are cropped out. Then, an action feature extractor trained based on the fast and slow convolutional network concept is used. This allows us to capture the behavioral characteristics of the target object and thus characterize its behavioral patterns within the video clip.

[0078] Furthermore, the behavioral feature extractor The output initial behavioral feature is weighted with the normal feature prototype stored in the behavioral prototype memory pool by similarity memory items, thereby obtaining the typical normal prototype corresponding to the initial behavioral feature, realizing the extraction of a compact prototype of normal behavior, so as to strengthen the semantic guidance of behavior.

[0079] This application utilizes a memory pool method to capture local behavioral features and obtain their compact representations. This method can effectively filter and focus on input action features from a large amount of normal action information, enabling high-level event semantic understanding even in complex scenes and providing a reliable feature foundation for subsequent future frame image prediction.

[0080] In one embodiment, extracting the initial behavioral features of the target object from the video clip includes: Target detection is performed on each frame image in the video segment to obtain the bounding box of the target object in each frame image; Based on the bounding box of the target object in each of the frame images, a continuous action sequence of the target object in the M frame images is generated; Behavioral features are extracted from the continuous action sequence to obtain the initial behavioral features of the target object.

[0081] Specifically, a pre-trained target detector, such as the model provided by Detectron, is used to detect targets in each frame of the video clip, obtain the bounding boxes of the target objects in each frame, and use this bounding box information to accurately crop out the region where the target object is located in each frame, and generate a continuous action sequence of the target object in M ​​frames based on these regions.

[0082] Behavioral feature extractor trained using the concept of fast and slow convolutional networks The generated continuous action sequence is analyzed in depth to extract key features that can characterize the behavior pattern of the target object. The action features are obtained and compressed into feature vectors to obtain the initial behavior features. Where C represents the vector dimension; if multiple target objects' initial behavioral features are identified in a video clip, their features are superimposed to form an N-dimensional behavioral feature vector. This process can be represented as:

[0083] in, This indicates a behavioral feature extraction operation.

[0084] This application embodiment can efficiently and accurately extract the behavioral features of the target object from video clips by detecting the target and extracting key features of the target object's behavioral patterns. These features include rich action information of the target object over a continuous period of time.

[0085] In one embodiment, the step of extracting the normal prototype of the initial behavioral features based on the behavioral prototype memory pool to obtain the behavioral features of the target object includes: The initial behavioral features are compared with each memory item to generate a similarity vector. The similarity vectors are normalized to generate a similarity weight matrix; The behavioral features of the target object are obtained by performing a dot product between the similarity weight matrix and the memory matrix.

[0086] Specifically, to obtain the prototype features of behavior in normal videos, a memory matrix is ​​used as a prototype memory pool to store the prototypes of the behavior features. Considering the dimensionality of the behavior features, a fixed-size memory matrix is ​​initialized, represented as follows: , The row vector of the behavioral prototype memory pool represents the memory depth of the behavioral prototype memory pool. express One of the memory entries is used to ensure dimensional consistency throughout the subsequent computation process. Therefore, it can be understood that the behavioral prototype memory pool stores... Each memory item The dimension is C, and the memory terms are used to represent the normal prototype of behavioral characteristics.

[0087] Computational Behavioral Feature Extractor The initial behavioral characteristics of the output and each memory item The similarity between the two is determined by a similarity algorithm that can be set according to the actual situation. The algorithm generates a similarity vector and then normalizes the obtained similarity vector to a unit weight W, which is used as the similarity weight matrix.

[0088] Taking the cosine similarity algorithm as an example, the initial behavioral features With memory items The similarity calculation and normalization between them are as follows:

[0089] Among them, initial behavioral characteristics Simplified representation as ; Represents the first in the memory matrix One memory item; Represents the th element in the similarity weight matrix The normalized weight value, that is, the in the similarity weight matrix Row vectors.

[0090] Furthermore, by re-integrating the similarity weight matrix W and the memory matrix M, the initial behavioral features can be obtained. address vector This addressing vector is a typical normal prototype corresponding to the initial behavioral features, and it is used as the final behavioral feature output by the behavioral prototype memory module. The final behavioral feature is shown in the following formula:

[0091] Where T represents transpose.

[0092] This application embodiment calculates and normalizes the similarity between memory items in the behavioral prototype memory pool and behavioral features, which can accurately extract representative normal prototypes from behavioral features, thereby extracting compact prototypes of normal behavior and strengthening behavioral semantic guidance.

[0093] In one embodiment, based on step 102, fusing the behavioral features with the global scene features to obtain a heterogeneous feature fusion result includes: Through an attention mechanism, a scene feature attention map of the global scene features and a behavior feature attention map of the behavior features are generated. The scene feature attention map is subjected to contextual reasoning processing to obtain the scene feature structured result of fusing the related information before and after the fusion; the behavior feature attention map is subjected to contextual reasoning processing to obtain the behavior feature structured result of fusing the related information before and after the fusion. The scene feature structuring results and the behavior feature structuring results are fused to obtain heterogeneous feature fusion results.

[0094] Specifically, understanding video events requires effectively combining scene and behavioral features. However, due to their semantic heterogeneity, direct fusion struggles to establish joint semantic consistency, hindering the learning of semantic information from the fused features. Inspired by contextual cues, hierarchical dynamic integration can focus on areas of significant change and promote understanding of the overall context. Therefore, this embodiment proposes a cross-fusion decoder to integrate scene and behavioral features and predict future frame images. The cross-fusion decoder employs a novel cross-attention mechanism to mine the potential relationship between scene and behavior, enhancing scene perception and event understanding capabilities, establishing joint semantic consistency between scene and behavioral features, and improving the accuracy of future frame generation. The specific framework can be found in [reference needed]. Figure 5 , Figure 5 This is a schematic diagram of the cross-fusion decoder provided in this application.

[0095] First, an attention map of global scene features is generated using an attention mechanism. Attention map of behavioral characteristics The algorithmic framework of the attention mechanism is as follows: Figure 5 The diagram shows common attention algorithms, which will not be described in detail here. Then, the scene feature attention map and the behavior feature attention map are passed through their respective feature attention methods. The computational flow structure of each attention map is the same, but the weight parameters are different.

[0096] In the attention graph computation flow, contextual reasoning is performed on both the scene feature attention graph and the behavior feature attention graph. When performing contextual reasoning on the scene feature attention graph, the correlation information of scene features across multiple dimensions such as time and space is comprehensively considered to uncover potential connections and evolutionary patterns between scene features, resulting in a structured result of scene features that integrates contextual and spatial information. Similarly, when performing contextual reasoning on the behavior feature attention graph, the correlation information of behavior features at different times and in different scenarios is analyzed in depth to capture the changing trends and action logic of behavior features, resulting in a structured result of behavior features that integrates contextual and spatial information.

[0097] Finally, a specific fusion algorithm is used to fuse the structured results of scene features and the structured results of behavior features, establish joint semantic consistency between scene features and behavior features, and finally obtain accurate and comprehensive heterogeneous feature fusion results.

[0098] For example, the output of the scene feature attention stream Output results of behavioral feature attention stream Connect the components to form the overall input for predicting future frames. ,in Indicates the connection method.

[0099] This application's embodiments establish a consistency relationship between dynamic behavior and static scenes through a cross-attention mechanism, enabling joint analysis of behavior and scene context. This strengthens the model's understanding of scene-dependent complex events, providing a reliable and comprehensive feature foundation for subsequent future frame image prediction, thereby improving the accuracy of complex anomaly identification. By emphasizing event perception and establishing the association between behavior and scene, the model can understand video content at the event level, effectively coping with the diversity of complex scenes in the real world. It can also improve the robustness of patrol robot vision in cross-scene monitoring anomaly detection capabilities.

[0100] In one embodiment, the step of performing contextual reasoning processing on the scene feature attention map to obtain a structurated result of scene features fused with related information includes: The scene feature attention map is input into the recursive network; In the current time step, based on the scene feature attention map and the hidden state of the previous time step, the scalar is updated and calculated; the scalar is used to characterize the relative importance of each pixel position in the scene feature attention map when extracting context points; Based on the hidden state of the previous time step, the gate vector is updated and calculated; the gate vector is used to assign the contribution of each pixel position in the scene feature attention map to contextual reasoning. Based on the scene feature attention map and the hidden state of the previous time step, the hidden state of the current time step is generated through a gating mechanism, and the hidden state of the current time step is passed to the next time step. The next time step is taken as the current time step. The steps of updating the scalar calculation based on the scene feature attention map and the hidden state of the previous time step are executed iteratively in the current time step until the last time step in the preset step size is reached, and the final scalar and the final gating vector in the last time step are obtained. Based on the final scalar, the final gating vector, and the scene feature attention map, a scene feature structured result is generated that integrates the related information before and after fusion.

[0101] Specifically, the attention graph computation flow structure is the same for both types of features, but the weight parameters are different. The following explanation uses the scene feature attention flow as an example; the computation process for the behavior feature attention flow is similar.

[0102] In the scene feature attention flow, the scene feature attention map is input into a recursive network, such as a Long Short-Term Memory (LSTM) network.

[0103] In an LSTM network, the scene feature attention map is received first. As input, we also obtain the hidden state passed from the previous time step. LSTM processes this information through its unique gating structure, including forget gates, input gates, and output gates. The forget gate determines which information from the cell state at the previous time step needs to be forgotten, the input gate determines which information from the current input needs to be added to the cell state, and the output gate generates the hidden state for the current time step based on the updated cell state.

[0104] In the current time step, a scalar is calculated based on the scene feature attention map and the hidden state from the previous time step. This scalar characterizes the relative importance of each pixel position in the scene feature attention map when extracting context points. The formula for calculating the scalar is as follows:

[0105] in, The first character in the scene feature attention graph represents the... Features of each pixel The number of pixels in the scene feature attention map; This indicates the hidden state of the previous time step; Indicates based on the first The scalar components are calculated from each pixel and the hidden state of the previous time step. That is scalar components after normalization; and These are the weight matrices, where n is the dimension of the weight matrix. The weight matrix is ​​randomly initialized and is continuously updated.

[0106] Since different pixels provide different amounts of information for contextual reasoning, a gating vector is introduced for the hidden state. ,in Let be the weight matrix, where and The dimensions of the row and column vectors of the weight matrix are used to determine the contribution of each position to contextual reasoning.

[0107] Furthermore, based on scene feature attention maps The hidden state of the previous time step The hidden state at the current time step is generated through a gating mechanism, which combines the functions of the input gate, forget gate, and output gate. And pass the hidden state of the current time step to the next time step.

[0108] The above steps are executed iteratively until the preset step size is reached. The last time step in the time step. (The time step size varies.) By continuously adding new weights based on the gating vector, the LSTM hidden state gradually filters out irrelevant information, establishing an effective correlation between the scene and the corresponding action. Finally, the attention map, gating vector, and scalar are combined to obtain the final result. This is represented as follows:

[0109] in, Indicates the first The gated vector components corresponding to each pixel; Indicates the first scalar components corresponding to each pixel; The first character in the scene feature attention graph Features of each pixel; This represents the structured result of scene features.

[0110] Based on the above calculation process of scene feature attention flow, the calculation process of behavior feature attention flow can be deduced, and the structured behavior feature output of behavior feature attention flow can be obtained. Then, the outputs of the scene feature attention stream and the behavior feature attention stream are concatenated to obtain the overall input for predicting future frame images.

[0111] The scene feature attention flow in this embodiment utilizes a recursive network to effectively process the spatiotemporal information in the scene feature attention map. By progressively filtering out irrelevant information through gating vectors, it establishes an effective association between the scene and the corresponding action, generating a structured result of scene features that integrates the information before and after fusion. Similarly, the behavior feature attention flow utilizes a recursive network to effectively process the spatiotemporal information in the behavior feature attention map. By progressively filtering out irrelevant information through gating vectors, it constructs an effective association between the behavior and the corresponding scene, generating a structured result of behavior features that integrates the information before and after fusion. These two results are then concatenated and used as the overall input for future frame image prediction, providing a comprehensive and accurate feature foundation for subsequent image prediction.

[0112] like Figure 2 As shown in the overall framework diagram, one of the main losses in the frame prediction method is the appearance loss between the predicted frame and the ground truth frame. Here, L2 distance is used to calculate the pixel-by-pixel difference between the predicted and ground truth frames. The expression for this loss function is:

[0113] in, Indicates damage to appearance; This represents the pixel at position (x, y) in the predicted frame image; This represents the pixel at position (x, y) in the real frame image.

[0114] Furthermore, in the behavioral prototype memory module, entropy loss from the behavioral prototype memory pool is introduced to eliminate redundant and irrelevant information while retaining the most important information. To encourage sparsity in the weight distribution, sparsity normalization is performed. The expression for this loss function is:

[0115] in, Indicates entropy loss; Represents memory items in the memory matrix Normalized weights for similarity between behavioral features.

[0116] In summary, the total loss function of the overall model structure is defined as:

[0117] One option is to use a weighted approach, which can be set according to the actual situation.

[0118] Combining steps 101, 102, and 103 and their embodiments, the functional modules of these steps can be encapsulated into the prediction model. That is, the prediction model includes a consistent scene encoder, a behavior prototype memory module, and a cross-fusion decoder. After acquiring the video segment and dividing it into spatiotemporal cubes, the video segment and the corresponding spatiotemporal cube are input into the prediction model. The prediction model is used to: extract behavioral features of the target object from the video segment and extract global scene features from the spatiotemporal cube; fuse the behavioral features and the global scene features to obtain a heterogeneous feature fusion result; and generate prediction results for future frames based on the heterogeneous feature fusion result.

[0119] Based on all the above embodiments, this application designs a scene-aware behavior event anomaly detection method. Addressing the issue that the evaluation of subject behavior events is influenced by scene context, this application focuses on improving the ability to capture and understand events, exploring the behavior of local foreground objects and learning their correlation with the scene context. First, at the static scene level, this application designs a consistency scene encoder with spatiotemporal attention, enhancing the model's perception of the global scene context. Second, at the dynamic behavior level, a behavior prototype memory module is designed to capture the behavioral features of target objects and extract prototype features of normal behavior, providing clear semantic guidance for the model to understand events. Next, this application proposes a cross-fusion decoder to integrate behavioral features and scene features and learn the consistency relationship between them. This method extracts scene features and behavioral prototype features from two scales: frame-centric and object-centric. By comprehensively utilizing these heterogeneous features and establishing event-level understanding, the model's ability to predict future frames can be enhanced.

[0120] Figure 6 This is a schematic diagram of the behavioral event anomaly detection device provided in this application.

[0121] like Figure 6 As shown, the behavioral event anomaly detection device includes: The heterogeneous feature extraction module 610 is used to extract behavioral features of the target object from the video clip and extract global scene features from the spatiotemporal cube; the spatiotemporal cube is constructed based on the sequence of consecutive frame images extracted from the video clip. The heterogeneous feature fusion module 620 is used to fuse the behavioral features with the global scene features to obtain a heterogeneous feature fusion result; The prediction module 630 is used to generate a prediction result for a future frame based on the heterogeneous feature fusion result; the future frame is the next frame immediately following the last frame of the spatiotemporal cube. The anomaly detection module 640 is used to perform anomaly detection on the prediction results of the future frame to obtain the behavioral event anomaly detection results.

[0122] The behavioral event anomaly detection device provided in this application extracts global scene features from a spatiotemporal cube at the static scene level, enabling a comprehensive understanding of the scene context. At the dynamic behavioral level, it captures behavioral features of characters from video clips, providing semantic guidance for understanding events. It heterogeneously fuses global scene features and behavioral features, learns the consistency relationship between heterogeneous features, achieves global context perception capability, strengthens the understanding of scene-dependent complex events, improves the accuracy of complex anomaly identification, and can more accurately identify potential abnormal behavioral events in videos.

[0123] In one embodiment, the heterogeneous feature extraction module 610 is further configured to: Global feature extraction is performed on the spatiotemporal cube to obtain initial global scene features, and spatiotemporal feature extraction is performed on the spatiotemporal cube to obtain a spatiotemporal attention vector; The initial global scene features and the spatiotemporal attention vector are fused to obtain global scene features.

[0124] In one embodiment, the spatiotemporal cube comprises T frames of images, where T is a positive integer; the heterogeneous feature extraction module 610 is further configured to: The T-frame images in the spacetime cube are merged according to the channel dimension to generate a tensor; Multi-level semantic information extraction is performed on the tensor to obtain initial global scene features.

[0125] In one embodiment, the spatiotemporal cube comprises T frames of images, where T is a positive integer; the heterogeneous feature extraction module 610 is further configured to: Semantic information is extracted from each frame image in the spatiotemporal cube to obtain the features of each frame image; Based on the features of the T-th frame image, query features are constructed, and based on the features of the intermediate frame images, key features and value features are constructed; the intermediate frame images include the first frame image to the (T-1)-th frame image. Based on the query features, the key features, and the value features, a spatiotemporal attention vector is generated.

[0126] In one embodiment, the heterogeneous feature extraction module 610 is further configured to: Extracting initial behavioral features of target objects from video clips; Based on the behavioral prototype memory pool, the normal prototypes of the initial behavioral features are extracted to obtain the behavioral features of the target object; the behavioral prototype memory pool is used to store the normal prototypes of the behavioral features.

[0127] In one embodiment, the video segment comprises M frames, where M is a positive integer; the heterogeneous feature extraction module 610 is further configured to: The extraction of initial behavioral features of the target object from the video clip includes: Target detection is performed on each frame image in the video segment to obtain the bounding box of the target object in each frame image; Based on the bounding box of the target object in each of the frame images, a continuous action sequence of the target object in the M frame images is generated; Behavioral features are extracted from the continuous action sequence to obtain the initial behavioral features of the target object.

[0128] In one embodiment, the behavioral prototype memory pool stores a memory matrix, which consists of J memory items, where J is a positive integer. These memory items are used to represent normal prototypes of behavioral features. The heterogeneous feature extraction module 610 is further used for: The initial behavioral features are compared with each memory item to generate a similarity vector. The similarity vectors are normalized to generate a similarity weight matrix; The behavioral features of the target object are obtained by performing a dot product between the similarity weight matrix and the memory matrix.

[0129] In one embodiment, the heterogeneous feature fusion module 620 is further configured to: Through an attention mechanism, a scene feature attention map of the global scene features and a behavior feature attention map of the behavior features are generated. The scene feature attention map is subjected to contextual reasoning processing to obtain the scene feature structured result of fusing the related information before and after the fusion; the behavior feature attention map is subjected to contextual reasoning processing to obtain the behavior feature structured result of fusing the related information before and after the fusion. The scene feature structuring results and the behavior feature structuring results are fused to obtain heterogeneous feature fusion results.

[0130] In one embodiment, the heterogeneous feature fusion module 620 is further configured to: The scene feature attention map is input into the recursive network; In the current time step, based on the scene feature attention map and the hidden state of the previous time step, the scalar is updated and calculated; the scalar is used to characterize the relative importance of each pixel position in the scene feature attention map when extracting context points; Based on the hidden state of the previous time step, the gate vector is updated and calculated; the gate vector is used to assign the contribution of each pixel position in the scene feature attention map to contextual reasoning. Based on the scene feature attention map and the hidden state of the previous time step, the hidden state of the current time step is generated through a gating mechanism, and the hidden state of the current time step is passed to the next time step. The next time step is taken as the current time step. The steps of updating the scalar calculation based on the scene feature attention map and the hidden state of the previous time step are executed iteratively in the current time step until the last time step in the preset step size is reached, and the final scalar and the final gating vector in the last time step are obtained. Based on the final scalar, the final gating vector, and the scene feature attention map, a scene feature structured result is generated that integrates the related information before and after fusion.

[0131] It should be noted that the behavioral event anomaly detection device provided in this application can execute the behavioral event anomaly detection method described in any of the above embodiments during actual operation, which will not be elaborated in this embodiment.

[0132] Figure 7 This is a schematic diagram of the structure of the electronic device provided in this application, such as... Figure 7 As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute a behavioral event anomaly detection method. This method includes: extracting behavioral features of a target object from a video clip and extracting global scene features from a spatiotemporal cube; the spatiotemporal cube is constructed based on a sequence of consecutive frame images extracted from the video clip; fusing the behavioral features and the global scene features to obtain a heterogeneous feature fusion result; generating a prediction result for a future frame based on the heterogeneous feature fusion result; the future frame is the frame immediately following the last frame of the spatiotemporal cube; and performing anomaly detection on the prediction result of the future frame to obtain a behavioral event anomaly detection result.

[0133] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0134] On the other hand, this application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer is able to execute the behavior event anomaly detection method provided in the above embodiments. The method includes: extracting behavior features of a target object from a video clip and extracting global scene features from a spatiotemporal cube; the spatiotemporal cube is constructed based on a sequence of consecutive frame images extracted from the video clip; fusing the behavior features and the global scene features to obtain a heterogeneous feature fusion result; generating a prediction result for a future frame based on the heterogeneous feature fusion result; the future frame is the next frame immediately following the last frame of the spatiotemporal cube; and performing anomaly detection on the prediction result of the future frame to obtain a behavior event anomaly detection result.

[0135] In another aspect, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the behavior event anomaly detection method provided in the above embodiments. The method includes: extracting behavior features of a target object from a video clip and extracting global scene features from a spatiotemporal cube; the spatiotemporal cube is constructed based on a sequence of consecutive frame images extracted from the video clip; fusing the behavior features and the global scene features to obtain a heterogeneous feature fusion result; generating a prediction result for a future frame based on the heterogeneous feature fusion result; the future frame is the next frame immediately following the last frame of the spatiotemporal cube; and performing anomaly detection on the prediction result of the future frame to obtain a behavior event anomaly detection result.

[0136] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0137] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for detecting abnormal behavioral events, characterized in that, The abnormal behavior event detection method includes: Behavioral features of the target object are extracted from the video clip, and global scene features are extracted from the spatiotemporal cube; the spatiotemporal cube is constructed based on a sequence of consecutive frame images extracted from the video clip. The behavioral features and the global scene features are fused to obtain a heterogeneous feature fusion result; Based on the heterogeneous feature fusion result, a prediction result for the future frame is generated; the future frame is the next frame immediately following the last frame of the spatiotemporal cube. Anomaly detection is performed on the prediction results of the future frames to obtain behavioral event anomaly detection results; The step of fusing the behavioral features with the global scene features to obtain a heterogeneous feature fusion result includes: Through an attention mechanism, a scene feature attention map of the global scene features and a behavior feature attention map of the behavior features are generated. The scene feature attention map is subjected to contextual reasoning processing to obtain the scene feature structured result of fusing the related information before and after the fusion; the behavior feature attention map is subjected to contextual reasoning processing to obtain the behavior feature structured result of fusing the related information before and after the fusion. The scene feature structuring results and the behavior feature structuring results are fused to obtain heterogeneous feature fusion results.

2. The behavioral event anomaly detection method according to claim 1, characterized in that, The extraction of global scene features from the spatiotemporal cube includes: Global feature extraction is performed on the spatiotemporal cube to obtain initial global scene features, and spatiotemporal feature extraction is performed on the spatiotemporal cube, and a spatiotemporal attention vector is obtained by calculation through a spatiotemporal attention mechanism; The initial global scene features and the spatiotemporal attention vector are fused to obtain global scene features.

3. The behavioral event anomaly detection method according to claim 2, characterized in that, The spatiotemporal cube comprises T frames of images, where T is a positive integer; the global feature extraction of the spatiotemporal cube to obtain initial global scene features includes: The T-frame images in the spacetime cube are merged according to the channel dimension to generate a tensor; Multi-level semantic information extraction is performed on the tensor to obtain initial global scene features.

4. The behavioral event anomaly detection method according to claim 2, characterized in that, The spatiotemporal cube comprises T frames of images, where T is a positive integer; the spatiotemporal feature extraction of the spatiotemporal cube to obtain a spatiotemporal attention vector includes: Semantic information is extracted from each frame image in the spatiotemporal cube to obtain the features of each frame image; Based on the features of the T-th frame image, query features are constructed, and based on the features of the intermediate frame images, key features and value features are constructed; the intermediate frame images include the first frame image to the (T-1)-th frame image. Based on the query features, the key features, and the value features, a spatiotemporal attention vector is generated.

5. The behavioral event anomaly detection method according to claim 1, characterized in that, The extraction of behavioral features of the target object from the video clip includes: Extracting initial behavioral features of target objects from video clips; Based on the behavioral prototype memory pool, the normal prototypes of the initial behavioral features are extracted to obtain the behavioral features of the target object; the behavioral prototype memory pool is used to store the normal prototypes of the behavioral features.

6. The behavioral event anomaly detection method according to claim 5, characterized in that, The video segment comprises M frames, where M is a positive integer; the extraction of initial behavioral features of the target object from the video segment includes: Target detection is performed on each frame image in the video segment to obtain the bounding box of the target object in each frame image; Based on the bounding box of the target object in each of the frame images, a continuous action sequence of the target object in the M frame images is generated; Behavioral features are extracted from the continuous action sequence to obtain the initial behavioral features of the target object.

7. The behavioral event anomaly detection method according to claim 5, characterized in that, The behavioral prototype memory pool stores a memory matrix, which consists of J memory items, where J is a positive integer. The memory items are used to represent the normal prototypes of behavioral features. The process of extracting normal prototypes of the initial behavioral features from the behavioral prototype memory pool to obtain the behavioral features of the target object includes: The initial behavioral features are compared with each memory item to generate a similarity vector. The similarity vectors are normalized to generate a similarity weight matrix; The behavioral features of the target object are obtained by performing a dot product between the similarity weight matrix and the memory matrix.

8. The method for detecting abnormal behavioral events according to claim 1, characterized in that, The contextual reasoning processing of the scene feature attention map to obtain the fused scene feature structured result includes: The scene feature attention map is input into the recursive network; In the current time step, based on the scene feature attention map and the hidden state of the previous time step, the scalar is updated and calculated; the scalar is used to characterize the relative importance of each pixel position in the scene feature attention map when extracting context points; Based on the hidden state of the previous time step, the gate vector is updated and calculated; the gate vector is used to assign the contribution of each pixel position in the scene feature attention map to contextual reasoning. Based on the scene feature attention map and the hidden state of the previous time step, the hidden state of the current time step is generated through a gating mechanism, and the hidden state of the current time step is passed to the next time step. The next time step is taken as the current time step. The steps of updating the scalar calculation based on the scene feature attention map and the hidden state of the previous time step are executed iteratively in the current time step until the last time step in the preset step size is reached, and the final scalar and the final gating vector in the last time step are obtained. Based on the final scalar, the final gating vector, and the scene feature attention map, a scene feature structured result is generated that integrates the related information before and after fusion.

9. A behavioral event anomaly detection device, characterized in that, The behavioral event anomaly detection device includes: The heterogeneous feature extraction module is used to extract behavioral features of target objects from video clips and global scene features from a spatiotemporal cube; the spatiotemporal cube is constructed based on a sequence of consecutive frame images extracted from the video clips. The heterogeneous feature fusion module is used to fuse the behavioral features with the global scene features to obtain the heterogeneous feature fusion result; The prediction module is used to generate a prediction result for a future frame based on the heterogeneous feature fusion result; the future frame is the next frame immediately following the last frame of the spatiotemporal cube. An anomaly detection module is used to perform anomaly detection on the prediction results of the future frames to obtain behavioral event anomaly detection results; The step of fusing the behavioral features with the global scene features to obtain a heterogeneous feature fusion result includes: Through an attention mechanism, a scene feature attention map of the global scene features and a behavior feature attention map of the behavior features are generated. The scene feature attention map is subjected to contextual reasoning processing to obtain the scene feature structured result of fusing the related information before and after the fusion; the behavior feature attention map is subjected to contextual reasoning processing to obtain the behavior feature structured result of fusing the related information before and after the fusion. The scene feature structuring results and the behavior feature structuring results are fused to obtain heterogeneous feature fusion results.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the behavioral event anomaly detection method as described in any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium, wherein a computer program is stored on the non-transitory computer-readable storage medium, characterized in that, When the computer program is executed by a processor, it implements the steps of the behavioral event anomaly detection method as described in any one of claims 1 to 8.

12. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the behavioral event anomaly detection method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Semantic modeling-based unsupervised video monitoring anomaly detection method and system

    CN120673332A

  • Video anomaly detection method based on infrared visible light image feature fusion

    CN120997733A

  • Simple scene equipment operation action rapid processing method based on test video

    CN121861319A