Abnormal behavior detection method and device, electronic equipment and nonvolatile storage medium
By extracting the timing and spatial features of the video stream in the abnormal behavior detection in public places, and combining attention mechanism and feature fusion technology, the problem of insufficient detection accuracy is solved, and the accurate identification of abnormal behavior is achieved.
Patent Information
- Application Number
- CN202510519111.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has insufficient detection accuracy in the detection of abnormal behavior in public places, making it difficult to accurately identify abnormal behavior in complex environments.
By obtaining the video stream to be detected, timing and spatial features are extracted, and combining attention mechanisms and feature fusion technology, the behavior categories of the target object are determined. Specific steps include timing feature extraction, attention weight determination, global feature extraction, spatial feature extraction and feature fusion.
It improves the accurate identification of target object behavior categories, enhances detection accuracy, and is suitable for abnormal behavior detection in public places in complex environments.
Smart Images

Figure CN120220247A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computer vision applications. Specifically, it relates to an abnormal behavior detection method, device, electronic device, and non-volatile storage medium. Background Art
[0002] With the increasing demand for public place safety management, abnormal behavior detection technology has become increasingly important in fields such as urban surveillance, traffic management, and commercial security. Abnormal behavior detection can not only enhance public safety but also provide important data support for urban management, thereby optimizing resource allocation and emergency response strategies.
[0003] Abnormal behavior detection usually relies on video surveillance data and identifies abnormal activities by analyzing the behavior patterns in the video stream. These technologies use various object detection algorithms to achieve real-time processing and analysis of video frames. However, due to interference factors in complex environments (such as lighting changes, occlusion, multi-person interaction, etc.), the detection algorithms in related technologies often have technical problems such as insufficient detection accuracy in the detection of abnormal behaviors in public places.
[0004] For the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of this application provide an abnormal behavior detection method, device, electronic device, and non-volatile storage medium to at least solve the technical problem of insufficient detection accuracy in the detection of abnormal behaviors in public places in related technologies.
[0006] According to one aspect of the embodiments of this application, an abnormal behavior detection method is provided, including: obtaining a video stream to be detected, and performing temporal feature extraction on multiple consecutive image frames in the video stream to be detected to obtain a temporal feature matrix, where the temporal feature matrix contains feature vectors corresponding to multiple consecutive image frames, and each image frame corresponds to one feature vector; determining the attention weights corresponding to the feature vectors in the temporal feature matrix, and determining the global feature corresponding to the temporal feature matrix according to the attention weights; performing spatial feature extraction on the image frames to obtain local features corresponding to the image frames; fusing the global feature and the local features to obtain a target feature, and determining the behavior category of the target object in the video stream to be detected according to the target feature.
[0007] Optionally, a target network model is used for feature extraction of image frames. The training steps of the target network model include: obtaining a training data set, where the training data set includes a plurality of unlabeled video stream data, and the video stream data includes behavior segments of a target object under different environmental conditions; using a contrast loss function to perform contrastive learning training on an initial network model according to the training data set to adjust the model parameters of the initial network model, where the contrastive learning training is used to determine the similarity between the behavior features corresponding to the behavior segments in different video streams in the training data set, aggregate the behavior features with a similarity greater than a first similarity threshold, and separate the behavior features with a similarity less than a second similarity threshold, and the first similarity threshold is greater than the second similarity threshold; and / or using the initial network model to predict the behavior features of the next future frame corresponding to a plurality of historical image frames in the video stream of the training data set, and determining the prediction error value between the behavior features of the future frame predicted by the model and the behavior features of the real future frame in the video stream, and adjusting the model parameters of the initial network model according to the prediction error value; and / or establishing a temporal association graph between the graph nodes corresponding to the behavior features at different time steps in the training data set, and adjusting the model parameters of the initial network model according to the association relationship features between the graph nodes in the temporal association graph.
[0008] Optionally, after obtaining the training data set, the method further includes: using a generator in an adversarial network to generate pseudo-video stream data according to the real video stream data in the training data set, where the behavior categories corresponding to the target objects in the pseudo-video stream data include various abnormal behaviors; using a discriminator in the adversarial network to determine the confidence of the pseudo-video stream data, and where the confidence is used to represent the probability that the video stream data is real data; adding the pseudo-video stream data with a confidence higher than a preset confidence threshold to the training data set.
[0009] Optionally, the target network model is deployed on an edge computing device and a cloud server; the method further includes: determining the network environment state, and determining the processing durations required for detecting abnormal behaviors of the video stream to be detected on the edge computing device and the cloud server respectively in the network environment state; in the case where the first processing duration corresponding to the edge computing device is less than the second processing duration corresponding to the cloud server, using the target network model deployed on the edge computing device to perform abnormal behavior detection, where the first processing duration includes the time required for data preprocessing and model inference on the edge computing device, and the second processing duration includes the time required for data preprocessing and model inference on the cloud server and the time required for data transmission between the cloud server and the edge computing device; in the case where the detected behavior category is an abnormal behavior, performing local alarm and uploading the alarm information to the cloud.
[0010] Optionally, the method further includes: determining the behavior category detected by the target network model as the pseudo label corresponding to the video stream to be detected, and adding the video stream to be detected and the pseudo label as new training data to the training data set; using the updated training data set to continue training the target network model locally deployed on the edge computing device; uploading the model parameters updated during the training of the target network model in each edge computing device to the cloud server; using the cloud server to aggregate and fuse the updated model parameters uploaded by different edge computing devices to obtain a new target network model, and updating the new target network model to each edge computing device.
[0011] Optionally, extracting spatial features from the image frame to obtain the local features corresponding to the image frame includes: determining the gradient amplitude corresponding to each pixel point in the image frame, where the gradient amplitude is used to characterize the edge intensity feature in the image frame; determining the edge information map corresponding to the image frame according to the gradient amplitude corresponding to each pixel point, where the edge information map is at least used to characterize the contour information of the target object in the image frame; using the target network model to extract features from the image frame to obtain an initial feature map, and fusing the edge information map and the initial feature map to obtain the local features corresponding to the image frame.
[0012] Optionally, after obtaining the video stream to be detected, the method further includes: adjusting the size of the image frame to a target size, where the target size is the input size supported by the network model used for feature extraction of the image frame; normalizing the pixel values of each point in the image frame of the target size, where the normalization process is used to scale the pixel values of each point to a preset range.
[0013] According to another aspect of the embodiments of the present application, there is also provided an abnormal behavior detection device, including: a data acquisition and feature extraction module, configured to acquire a video stream to be detected, and perform temporal feature extraction on multiple consecutive image frames in the video stream to be detected to obtain a temporal feature matrix, where the temporal feature matrix contains feature vectors corresponding to multiple consecutive image frames, and each image frame corresponds to one feature vector; a global temporal feature determination module, configured to determine the attention weights corresponding to the feature vectors in the temporal feature matrix, and determine the global feature corresponding to the temporal feature matrix according to the attention weights; a local spatial feature determination module, configured to extract spatial features from the image frame to obtain the local features corresponding to the image frame; an abnormal behavior classification and prediction module, configured to fuse the global feature and the local feature to obtain a target feature, and determine the behavior category of the target object in the video stream to be detected according to the target feature.
[0014] According to yet another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory and a processor, where the processor is configured to run a program stored in the memory, and when the program runs, it executes the abnormal behavior detection method.
[0015] According to another aspect of the embodiments of the present application, a non-volatile storage medium is further provided. The non-volatile storage medium includes a stored computer program. Wherein, the device where the non-volatile storage medium is located executes the abnormal behavior detection method by running the computer program.
[0016] According to another aspect of the embodiments of the present application, a computer program product is further provided, including a computer program. When the computer program is executed by a processor, the steps of the abnormal behavior detection method are implemented.
[0017] In the embodiments of the present application, the method of obtaining the video stream to be detected, extracting the temporal features of multiple consecutive image frames in the video stream to be detected to obtain a temporal feature matrix is adopted. Wherein, the temporal feature matrix contains feature vectors corresponding to multiple consecutive image frames, and each image frame corresponds to one feature vector; determining the attention weights corresponding to the feature vectors in the temporal feature matrix, and determining the global feature corresponding to the temporal feature matrix according to the attention weights; extracting the spatial features of the image frames to obtain the local features corresponding to the image frames; fusing the global features and the local features to obtain the target features, and determining the behavior category of the target object in the video stream to be detected according to the target features. By extracting the temporal features and spatial features in the video stream and combining the attention mechanism and the feature fusion technology, the purpose of accurately identifying the behavior category of the target object is achieved, and further solves the technical problem of insufficient detection accuracy in the detection of abnormal behaviors in public places in the related art. Description of the Drawings
[0018] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:
[0019] Figure 1 is a hardware structure block diagram of a computer terminal (or electronic device) for implementing the abnormal behavior detection method provided by the embodiments of the present application;
[0020] Figure 2 is a schematic diagram of the method flow of an abnormal behavior detection provided by the embodiments of the present application;
[0021] Figure 3 is a schematic diagram of the structure of an abnormal behavior detection device provided by the embodiments of the present application. Detailed Embodiments
[0022] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope of protection of this application.
[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0024] Due to interference factors in complex environments (such as light changes, occlusion, multi-person interaction, etc.), there are still many problems in the behavior recognition technology in the related art in terms of feature extraction, target recognition, and behavior detection, as follows:
[0025] 1) Insufficient detection accuracy: In the detection of abnormal behaviors in public places in the related art, there is often a problem of insufficient detection accuracy. This is mainly because the manifestations of abnormal behaviors are diverse and may be similar to normal behaviors. Especially in complex backgrounds and multi-target situations, it may lead to false detections or missed detections. This situation affects the accurate recognition of abnormal behaviors and may result in the failure to detect and handle potential safety hazards in a timely manner.
[0026] 2) Real-time issue: In public place monitoring, real-time is a key factor. There are certain deficiencies in the processing speed of the related art. Especially in high-resolution video streams or large-scale video surveillance systems, the deep learning models in the related art usually perform inferences in the cloud, resulting in relatively high data transmission and processing delays, which are not applicable to application scenarios with extremely high real-time requirements. This may lead to delayed detection of abnormal behaviors and affect the effect of rapid response and emergency handling.
[0027] 3) Insufficient model generalization ability: Existing behavior detection models rely on large-scale labeled data for training. However, due to the scarcity of abnormal behavior samples, the adaptability of the models to new types of abnormal behaviors is relatively weak.
[0028] To solve the above problems, relevant solutions are provided in the embodiments of the present application, which are described in detail below.
[0029] According to an embodiment of the present application, an embodiment of a method for detecting abnormal behavior is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0030] The method embodiment provided by the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or electronic device) for implementing the abnormal behavior detection method is shown. As Figure 1 shown, the computer terminal 10 (or electronic device) may include one or more processors 102 (shown as 102a, 102b,..., 102n in the figure) (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than those Figure 1 shown, or have a different configuration from that Figure 1 shown.
[0031] It should be noted that the above one or more processors 102 and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or electronic device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).
[0032] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the abnormal behavior detection method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned abnormal behavior detection method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0033] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0034] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computer terminal 10 (or electronic device).
[0035] Under the above operating environment, the embodiments of the present application provide an abnormal behavior detection method. Figure 2 It is a schematic diagram of a method flow for abnormal behavior detection provided according to the embodiments of the present application, as Figure 2 shown. The method includes the following steps:
[0036] Step S202, obtain a video stream to be detected, and perform temporal feature extraction on multiple consecutive image frames in the video stream to be detected to obtain a temporal feature matrix, where the temporal feature matrix includes feature vectors corresponding to multiple consecutive image frames, and each image frame corresponds to one feature vector;
[0037] Step S204, determine the attention weights corresponding to the feature vectors in the temporal feature matrix, and determine the global feature corresponding to the temporal feature matrix according to the attention weights;
[0038] Step S206, perform spatial feature extraction on the image frames to obtain local features corresponding to the image frames;
[0039] Step S208: Integrate the global features and local features to obtain the target features, and determine the behavior category of the target object in the video stream to be detected based on the target features.
[0040] Through the above steps, by extracting the temporal features and spatial features in the video stream and combining the attention mechanism and feature fusion technology, the purpose of accurately identifying the behavior category of the target object is achieved, thereby solving the technical problem of insufficient detection accuracy in the detection of abnormal behaviors in public places in the related art.
[0041] Next, the abnormal behavior detection method in steps S202 to S208 of the present application embodiment will be further introduced.
[0042] In the embodiment of the present application, the above method process can be executed and used in network element devices. Among them, the network element devices can include: embedded intelligent cameras, mobile embedded systems, communication modules, etc. Connections can be established between each network element device through the communication module to achieve data transmission and interaction; the intelligent camera can be used to obtain the video stream to be detected, and the mobile embedded system can be used to run algorithms to achieve target detection and prediction. By deploying the algorithm model corresponding to pedestrian fall detection on the mobile embedded system, image processing, target detection, and communication functions can be realized through software.
[0043] The abnormal behavior detection method in the embodiment of the present application is particularly suitable for the detection of abnormal behaviors in public places, especially for abnormal behaviors such as smoking and falling that are difficult to accurately identify, and can effectively improve the detection accuracy and real-time performance. The following is a specific introduction.
[0044] First, an intelligent camera can be used to obtain the video stream to be detected. For example, the intelligent camera captures a real-time video stream through the OpenCV library to obtain continuous image frames. This process ensures that the system can monitor the status of pedestrians in real time. As shown in the following formula:
[0045] Frame t = OpenCV.capture(t)
[0046] Among them, Frame_t represents the image frame captured at time t.
[0047] In this embodiment, data enhancement and optimization can be performed on the image frames in the video stream to be detected. The specific steps are as follows.
[0048] In some embodiments of the present application, after obtaining the video stream to be detected, the method further includes the following steps: adjusting the size of the image frame to a target size, where the target size is the input size supported by the network model used for feature extraction of the image frame; performing normalization processing on the pixel values of each point in the image frame with the target size, where the normalization processing is used to scale the pixel values of each point to a preset range interval.
[0049] Specifically, preprocessing each frame of the image may include resizing and normalization. In the preprocessing stage, the image is adjusted to a size acceptable to the model (i.e., the target size, which can be, for example, 640*640 pixels), as shown in the following formula:
[0050] Image resized =Resize(Frame t ,640*640)
[0051] Next, after normalization processing, the pixel values are scaled between 0 and 1 to facilitate network processing, as shown in the following formula:
[0052]
[0053] Such a preprocessing process ensures the consistency and standardization of the image, enabling better extraction of image features subsequently.
[0054] After that, feature extraction can be carried out. When performing behavior recognition in complex scenarios, in related technologies, models based on convolutional neural networks perform excellently in spatial feature extraction, but there are obvious limitations in capturing the temporal features of behaviors, especially long-term dependencies. Therefore, the embodiments of the present application propose an optimized method for object recognition that combines the Transformer architecture with multi-scale feature fusion, aiming to improve the model's ability to understand long-temporal behaviors and the fusion expression ability of multi-scale spatial features, thereby significantly improving the accuracy and robustness of behavior recognition, as follows.
[0055] In the embodiments of the present application, the self-attention mechanism can be used to model the features at different time points in the behavior sequence, enhancing the model's sensitivity to cross-frame target changes. At the same time, to avoid ignoring local spatial details, a convolutional neural network is introduced to extract local features and fuse them with the global temporal features to form a richer multi-scale feature expression.
[0056] Specifically, first, temporal feature extraction is performed on multiple consecutive image frames in the video stream to be detected, obtaining a set of temporal feature matrices X = {x1, x2,..., x T}, and each frame of the image x tEncoded into a vector form by CNN. For example, a continuous sequence of a fixed number of frames (e.g., 16 frames) can be extracted from a video sequence. Each frame image extracts spatial features via a convolutional neural network (such as ResNet) to obtain an initial temporal feature matrix where T is the number of frames and D is the feature dimension.
[0057] After that, X can be input into the Transformer encoder, and the multi-head attention mechanism is used for global modeling. Specifically, in order to capture the global dependencies between different moments, the input features are mapped into queries (Q), keys (K), and values (V):
[0058] Q = XW Q , K = XW K , V = XW V
[0059] where, W Q , W K , W V are linear transformation matrices respectively.
[0060] Then, based on the self-attention mechanism, the attention weights are calculated, that is, the attention weights corresponding to the feature vectors in the temporal feature matrix are determined, as shown in the following formula:
[0061]
[0062] where, d k is the scaling factor of the vector dimension, used to prevent the gradient from being too large; the Softmax operation is used to ensure the normalization of the attention weights.
[0063] The multi-head attention mechanism (Multi-Head Attention) is introduced to capture multiple behavioral relationships simultaneously:
[0064] MultiHead(Q, K, V) = Concat(head1,..., head h )W O
[0065] where each head represents an independent self-attention calculation process, which helps to learn behavioral dependencies from multiple subspaces.
[0066] Finally, the global temporal feature F global is obtained according to the attention weights, reflecting the global dependencies of behaviors in the time dimension.
[0067] On the other hand, in order to enhance the perception ability of the local space, local spatial features F local (focusing on the spatial structure information of the image) can also be extracted, and then combined with the global feature F globalPerform weighted fusion to obtain the final target feature F final = αF local + βF global , where α and β are fusion weight parameters satisfying α + β = 1, which can adjust the contributions of the two types of features to the final recognition result; F final represents the final fused feature (target feature) for target behavior classification and recognition.
[0068] Specifically, in the embodiments of this application, the specific steps for extracting the spatial features of the image frame are as follows.
[0069] In some embodiments of this application, extracting the spatial features of the image frame to obtain the local features corresponding to the image frame includes the following steps: determining the gradient magnitude corresponding to each pixel point in the image frame, where the gradient magnitude is used to characterize the edge intensity feature in the image frame; determining the edge information map corresponding to the image frame according to the gradient magnitude corresponding to each pixel point, where the edge information map is at least used to characterize the contour information of the target object in the image frame; using the target network model to extract features from the image frame to obtain the initial feature map, and fusing the edge information map and the initial feature map to obtain the local features corresponding to the image frame.
[0070] Specifically, first use the target operator (for example, Sobel operator) to calculate the gradient magnitude of the image. The Sobel operator calculates the gradients of the image in these two directions by applying convolution kernels in the horizontal and vertical directions on the image. The calculation formula of the gradient is as follows:
[0071] Horizontal gradient (G x ):
[0072]
[0073] where represents the gradient of the image I in the horizontal (x direction).
[0074] Vertical gradient (G y ):
[0075]
[0076] where represents the gradient of the image I in the vertical (y direction).
[0077] Then, by combining the horizontal and vertical gradients, calculate the gradient magnitude of each pixel point, which helps to determine the edge intensity in the image. The calculation formula of the gradient magnitude is as follows:
[0078]
[0079] Among them, G represents the gradient magnitude, G x and G y are the gradient values in the horizontal and vertical directions respectively. Based on the gradient magnitude corresponding to each pixel point, the edge information map (gradient magnitude map) corresponding to the image frame is determined. This map can highlight the edge features in the image, especially those areas with obvious intensity changes, which is very important for the recognition of contours and contact points in fall detection.
[0080] After that, the edge information map generated through the above steps can be used as part of the feature map to be fused with the initial feature map originally extracted from the image frame to obtain local features.
[0081] The embodiments of this application effectively integrate the global temporal modeling ability and the local spatial perception ability, and are applicable to multi-scenario and multi-scale behavior detection tasks. By using the self-attention mechanism to model the long-term dependencies between behaviors and leveraging CNN to make up for the lack of spatial features, feature-level optimal fusion is achieved, which has extremely high engineering promotion value and technological innovation in practical scenarios.
[0082] Next, the target network model and its training process used for feature extraction in the embodiments of this application will be further introduced.
[0083] In the behavior detection task, feature extraction is the core link of the entire detection process, which determines the accuracy of subsequent target recognition and classification. However, the network models for feature extraction in related technologies need to rely on a large amount of labeled data for training, and the labeling process is usually time-consuming and laborious. Especially in behavior detection tasks, the behaviors in different scenarios vary greatly, resulting in low generalization ability of manual labeling. Therefore, the embodiments of this application propose a feature extraction method based on self-supervised learning (SSL) for model training and application, which can autonomously learn behavior features in an environment without labeled data, thereby reducing the dependence on manual labeling, as follows.
[0084] In some embodiments of the present application, a target network model is used for feature extraction of image frames. The training steps of the target network model include: obtaining a training data set, where the training data set includes a plurality of unlabeled video stream data, and the video stream data contains behavior segments of a target object under different environmental conditions; using a contrastive loss function, and performing contrastive learning training on an initial network model according to the training data set to adjust the model parameters of the initial network model, where the contrastive learning training is used to determine the similarity between the behavior features corresponding to the behavior segments in different video streams in the training data set, aggregating the behavior features with a similarity greater than a first similarity threshold, and separating the behavior features with a similarity less than a second similarity threshold, and the first similarity threshold is greater than the second similarity threshold; and / or, using the initial network model, predicting the behavior features of the next future frame corresponding to a plurality of historical image frames in the video stream of the training data set, determining the prediction error value between the behavior features of the future frame predicted by the model and the behavior features of the real future frame in the video stream, and adjusting the model parameters of the initial network model according to the prediction error value; and / or, establishing a temporal association graph between the graph nodes corresponding to the behavior features at different time steps in the training data set, and adjusting the model parameters of the initial network model according to the association relationship features between the graph nodes in the temporal association graph.
[0085] Specifically, a large amount of unlabeled video data is first collected as the training data set, including behavior segments indoors, outdoors, at different angles, and under different lighting conditions. In this embodiment, techniques such as inter-frame difference and optical flow analysis can be used for data augmentation to ensure data diversity.
[0086] After that, the data in the training data set can be used to perform contrastive learning training on the model.
[0087] Specifically, contrastive learning learns effective feature representations by maximizing the similarity between positive samples and minimizing the similarity between negative samples. In this embodiment, given a pair of behavior video segments x i and x j , if they belong to the same behavior category, they are used as a positive sample pair, otherwise they are used as a negative sample pair.
[0088] The contrastive loss function is defined as follows:
[0089]
[0090] where z i , z jis the feature vector of the input video segment; sim(·) represents the cosine similarity function; τ is the temperature coefficient used to adjust the gradient magnitude of the contrast loss; the denominator term represents all possible sample pairs to ensure the distinguishability of the feature distribution.
[0091] The contrast loss function is used for feature extraction training. Through contrastive learning, the model can autonomously construct a behavioral feature representation space, enabling the aggregation of features of similar behaviors and the separation of features of different behaviors. The contrast loss is optimized by adjusting the temperature parameter to improve the model's discrimination ability for behavioral features and enhance the feature expression ability.
[0092] In this embodiment, the model can also be further trained through future frame prediction for time series modeling.
[0093] Specifically, behavior detection usually involves time series data. To further optimize feature expression, the embodiments of this application can introduce a future frame prediction task. This method can be based on a set of input video frames X t , and predict the frame features at the future time X t+1 , thereby enhancing the model's ability to model temporal information.
[0094] The future frame prediction task can be optimized by minimizing the following loss function:
[0095]
[0096] where F(X t ) is the feature of the next frame predicted by the model; X t+1 is the true feature of the next frame; |·|| 2 represents the Euclidean distance, which is used to measure the prediction error.
[0097] Through future frame prediction, the model can learn temporal correlation information and improve its ability to understand behavioral features. By optimizing the prediction ability with the mean squared error loss, the model's inference of future behaviors becomes more accurate.
[0098] In addition, the relationship between behavioral features can be modeled by combining a graph convolutional network (GCN) to further train the model.
[0099] Specifically, due to the possible complex spatio-temporal correlations among different features in behavior detection, the embodiments of this application can further introduce a graph convolutional network (Graph Convolutional Network, GCN) to model the relationship between behavioral features using a graph structure.
[0100] The calculation process of GCN is shown in the following formula:
[0101] Z ′ =σ(WZ + b)
[0102] Among them, Z is the original feature representation, and Z ′ is the optimized feature representation; W and b are trainable parameters; σ(·)σ is a non-linear activation function (such as ReLU).
[0103] Taking the behavioral features at different time steps as graph nodes, a temporal correlation graph structure is established, and GCN is used to extract the correlations between features. The model can automatically learn the correlations between different behavioral features, thereby improving the expressive ability of the behavioral features of complex behavioral patterns.
[0104] On the other hand, in actual scenarios, for monitoring tasks of illegal or dangerous behaviors such as falling, smoking, and abnormal staying, there is often a problem of scarce samples. Since these behaviors are random, hidden, and the cost of collecting and annotating in the real environment is relatively high, it is usually difficult to obtain a sufficient number of abnormal samples in the training set. This sample imbalance will cause the model to overfit to normal behaviors during training, thereby reducing the recognition accuracy and robustness of abnormal behaviors.
[0105] To solve the above problems, the embodiments of the present application can introduce a generative adversarial network (GAN, Generative Adversarial Network) for data augmentation, construct forged behavior samples, and expand the training set. Under the adversarial training mechanism, the generator of the GAN learns the real sample distribution and generates "pseudo-samples" with high similarity, while the discriminator continuously optimizes its ability to distinguish between true and false samples. During the game process between the two, the quality of the generated samples is continuously improved, which can effectively make up for the shortage of abnormal sample quantity and improve the generalization ability and stability of the model in actual deployment. Specifically as follows.
[0106] In some embodiments of the present application, after obtaining the training data set, the method further includes the following steps: using the generator in the adversarial network to generate pseudo-video stream data based on the real video stream data in the training data set, where the behavioral categories corresponding to the target objects in the pseudo-video stream data include various abnormal behaviors; using the discriminator in the adversarial network to determine the confidence of the pseudo-video stream data, and using the confidence to represent the probability that the video stream data is real data; adding the pseudo-video stream data with a confidence higher than the preset confidence threshold to the training data set.
[0107] Specifically, the adversarial network consists of two parts: a generator and a discriminator. Among them, the generator (Generator, G) is used to sample from the noise distribution and generate forged samples; the discriminator (Discriminator, D) is used to determine whether the input sample is real data or generated data.
[0108] In this embodiment, the optimization objective function of the GAN is a minimax problem:
[0109]
[0110] where x ~ P data represents the true sample distribution; z ~ P z represents the noise distribution (such as Gaussian distribution or uniform distribution); G(z) is the pseudo-sample output by the generator; D(x) is the confidence output of the discriminator on whether the input is a true sample. The goal is to minimize this loss function by the generator G (i.e., making the discriminator unable to distinguish between true and false), and to maximize this loss function by the discriminator D (i.e., correctly classifying true samples and pseudo-samples).
[0111] To enhance the authenticity and usability of behavior samples, the embodiments of this application can introduce the following strategies: 1) Conditional Constrained Generation (Conditional GAN): Attach a behavior label y to the input noise vector z to guide the generation of specific types of behaviors, such as smoking, falling, running fast, etc.; 2) Pseudo-sample Screening Mechanism: Use the trained discriminator to filter the generated samples, and only retain samples with high confidence (such as D(G(z)) > 0.8) for training. 3) Real Data Mixed Training: Combine high-quality pseudo-samples and original real samples to form an extended dataset, and input it into the subsequent detection model to improve robustness.
[0112] Specifically, in the adversarial training stage of data augmentation, the discriminator can be trained using true samples, and the generator can be trained using random noise; as the training progresses, the generator learns to approximate the feature distribution of abnormal behavior samples. Then, high-confidence pseudo-samples are screened out from the generator; these samples are mixed with the original abnormal behavior samples to form a new training set. Use the enhanced dataset to retrain the behavior abnormal behavior detection model.
[0113] The embodiments of this application implement effective expansion of scarce abnormal samples by introducing a behavior pseudo-sample generation mechanism based on a generative adversarial network, significantly improving the recognition accuracy and robustness of the detection model in small-sample scenarios, being applicable to various complex scenarios such as abnormal detection and behavior recognition, and having strong adaptability and broad engineering application prospects.
[0114] In the embodiments of this application, real-time detection optimization can also be performed based on edge computing, sinking model inference and behavior recognition tasks to the edge side for processing, reducing the cloud computing pressure, reducing network transmission latency, and improving the overall real-time response ability of the system. At the same time, design a lightweight neural network structure, combined with a distributed inference scheduling strategy, to effectively reduce the occupancy of computing resources while ensuring detection accuracy, and achieve the deployment of an efficient, fast, and intelligent detection system. For example, it can be applied to terminal devices such as intelligent cameras, embedded chips, and edge servers, as follows.
[0115] In some embodiments of the present application, the target network model is deployed on edge computing devices and cloud servers; the method further includes: determining the network environment status, and determining the processing durations required for detecting abnormal behaviors in the to-be-detected video stream on the edge computing device and the cloud server respectively under the network environment status; in the case where the first processing duration corresponding to the edge computing device is less than the second processing duration corresponding to the cloud server, using the target network model deployed in the edge computing device to perform abnormal behavior detection, where the first processing duration includes: the time required for data preprocessing and model inference on the edge computing device, and the second processing duration includes: the time required for data preprocessing and model inference on the cloud server, and the time required for data transmission between the cloud server and the edge computing device; in the case where the detected behavior category is an abnormal behavior, performing local alarm and uploading the alarm information to the cloud.
[0116] Specifically, in order to adapt to the computing power limitations of edge computing devices, model pruning and quantization techniques can be used to reduce the size of the target network model, and inference graph optimization (Graph Fusion) and intermediate tensor reuse are introduced to improve the model execution efficiency, so that the target network model is converted into a deployment format adapted to edge computing devices, ensuring its smooth operation on resource-constrained terminals. After that, the compressed target network model can be deployed to terminal devices such as intelligent cameras and edge AI boxes on construction sites or in nursing homes to achieve in-situ video analysis and inference.
[0117] In the process of edge-cloud collaborative task scheduling, the system can dynamically select the edge side or the cloud for inference by analyzing the current network condition and computing load:
[0118] T total =min(T edge ,T cloud )
[0119] where T edge is the first processing duration, and T cloud is the second processing duration; if T edge <T cloud , the task is independently processed by the edge computing device; otherwise, it is sent to the cloud server for processing.
[0120] The edge computing device processes the video stream in real time. When continuous fall behavior characteristics are detected, an alarm is immediately triggered locally (such as sound and light alarm, SMS push, etc.), without relying on the cloud for judgment. At the same time as the alarm, the device can upload relevant segments and recognition results to the management platform for further review or event archiving, realizing edge-cloud collaborative management.
[0121] Embodiments of the present application achieve efficient behavior detection capabilities on devices with limited computing power by introducing an edge computing architecture and combining lightweight model design with a dynamic scheduling mechanism. It not only meets the high real-time monitoring requirements but also significantly improves the system's adaptability and response speed.
[0122] In addition, the potential knowledge in unlabeled data can be mined through the pseudo-label mechanism, and the sharing and fusion of model knowledge among edge nodes can be achieved in combination with federated learning without exposing the original data, thereby ensuring privacy security and enhancing the robustness and adaptability of the overall detection system, as follows.
[0123] In some embodiments of the present application, the method further includes: determining the behavior categories detected by the target network model as the pseudo-labels corresponding to the video stream to be detected, and adding the video stream to be detected and the pseudo-labels as new training data to the training dataset; using the updated training dataset to continue training the target network model locally deployed on the edge computing device; uploading the model parameters updated by the target network model in each edge computing device during the training process to the cloud server; using the cloud server to aggregate and fuse the updated model parameters uploaded by different edge computing devices to obtain a new target network model, and updating the new target network model to each edge computing device.
[0124] Specifically, in unlabeled video data, the current model is used to generate prediction labels as "pseudo-labels" and incorporate them into the subsequent training process to continuously enhance the model's adaptability to the new environment. At the same time, to avoid directly transmitting video image data, the system adopts a federated learning strategy to perform local model training at different deployment points (such as different buildings, regions, or terminal devices) and regularly aggregates model parameters to the central server to form a new global model through aggregation and optimization.
[0125] Among them, the federated average (FedAvg) update formula is as follows:
[0126]
[0127] Where: w t represents the model parameters at the t-th iteration; K represents the number of clients participating in the training; n i represents the number of samples of client i; N is the total number of samples; represents the local gradient of the i-th client; η is the learning rate. This method avoids transmitting sensitive original data and only exchanges encrypted gradients or model weights, taking into account both data security and model collaborative training.
[0128] In this embodiment, lightweight target network models can be respectively deployed in multiple different monitoring environments (such as nursing homes, shopping malls, offices); the unlabeled video data collected by each deployment terminal (edge computing device) generates pseudo-labels through an existing model, and continues to train the model locally in combination with a small amount of labeled samples; every fixed period, each edge node uploads the updated model weights during the training process to the central (cloud) server, without transmitting the original video. The central server aggregates the weight information from each node, aggregates it into a new global model using the FedAvg algorithm, and distributes it to all devices after updating. Each edge node deploys and runs the new model, and continues with pseudo-label learning and local fine-tuning to form a closed-loop adaptive optimization mechanism. Through this mechanism, the system can quickly adapt and continuously optimize in different scenarios, with the detection accuracy increased by more than 10% on average, and the robustness significantly enhanced especially in complex lighting or occlusion environments.
[0129] To further enhance the privacy protection ability of the system during the multi-terminal collaborative learning process, as an optional implementation manner, a differential privacy mechanism can also be introduced. Before the edge node uploads the model parameters or gradients, an appropriate amount of noise is added, so that even if an attacker obtains the intermediate parameter information, the original data content cannot be deduced. This solution can strengthen the practicability of the system in scenarios with extremely high data privacy requirements such as government and enterprise, medical care, and elderly care, while taking into account the continuous optimization ability of the detection model.
[0130] At the same time, aiming at the problem of the decline in the model generalization ability in the new environment, a fast adaptation module based on federated transfer learning can be further constructed. When the system is deployed to a brand-new scenario (such as a new building, a new camera angle), local terminals are allowed to obtain "transfer templates" from the central server, and efficient transfer can be achieved with only a small amount of target domain data. By freezing some general layers and only fine-tuning specific recognition layers, the new environment can be quickly adapted without destroying the original recognition ability, improving the deployment efficiency and practicability.
[0131] In addition, to cope with the real-world monitoring environment of multi-sensor fusion, as an optional implementation manner, a pseudo-label generation mechanism supporting multi-modal inputs such as image, audio, and IMU sensor data can be designed. For example, by recognizing sudden sounds through audio, capturing fall impacts through vibration sensors, and detecting action behaviors through cameras, the three are combined to generate more credible pseudo-labels, thereby improving the quality of the model's pseudo-supervised learning. This solution can greatly improve the model's perception ability and stability under conditions such as occlusion, poor light, and limited viewing angles.
[0132] According to the embodiments of the present application, an embodiment of an abnormal behavior detection device is also provided. Figure 3 It is a schematic structural diagram of an abnormal behavior detection device provided according to the embodiments of the present application. AsFigure 3 As shown, the device includes:
[0133] A data acquisition and feature extraction module 30, configured to acquire a video stream to be detected, and perform temporal feature extraction on multiple consecutive image frames in the video stream to be detected, so as to obtain a temporal feature matrix, where the temporal feature matrix includes feature vectors corresponding to multiple consecutive image frames, and each image frame corresponds to one feature vector;
[0134] A global temporal feature determination module 32, configured to determine the attention weights corresponding to the feature vectors in the temporal feature matrix, and determine the global features corresponding to the temporal feature matrix according to the attention weights;
[0135] A local spatial feature determination module 34, configured to perform spatial feature extraction on the image frames to obtain local features corresponding to the image frames;
[0136] An abnormal behavior classification and prediction module 36, configured to fuse the global features and the local features to obtain target features, and determine the behavior category of the target object in the video stream to be detected according to the target features.
[0137] Optionally, a target network model is used for feature extraction of the image frames. The training steps of the target network model include: acquiring a training data set, where the training data set includes multiple unlabeled video stream data, and the video stream data includes behavior segments of the target object under different environmental conditions; using a contrast loss function, and performing contrast learning training on the initial network model according to the training data set to adjust the model parameters of the initial network model, where the contrast learning training is used to determine the similarity between the behavior features corresponding to the behavior segments in different video streams in the training data set, aggregating the behavior features with a similarity greater than a first similarity threshold, and separating the behavior features with a similarity less than a second similarity threshold, and the first similarity threshold is greater than the second similarity threshold; and / or, using the initial network model, predicting the behavior features of the next future frame corresponding to multiple historical image frames in the video stream of the training data set, and determining the prediction error value between the behavior features of the future frame predicted by the model and the behavior features of the real future frame in the video stream, and adjusting the model parameters of the initial network model according to the prediction error value; and / or, establishing a temporal association graph between the graph nodes corresponding to the behavior features at different time steps in the training data set, and adjusting the model parameters of the initial network model according to the association relationship features between the graph nodes in the temporal association graph.
[0138] Optionally, after obtaining the training data set, the abnormal behavior detection device is further configured to: use the generator in the adversarial network to generate pseudo-video stream data based on the real video stream data in the training data set, where the behavior categories corresponding to the target objects in the pseudo-video stream data include various abnormal behaviors; use the discriminator in the adversarial network to determine the confidence of the pseudo-video stream data, and use the confidence to represent the probability that the video stream data is real data; add the pseudo-video stream data with a confidence higher than the preset confidence threshold to the training data set.
[0139] Optionally, the target network model is deployed on the edge computing device and the cloud server; the abnormal behavior detection device is further configured to: determine the network environment state, and determine the processing duration required for detecting abnormal behaviors of the video stream to be detected on the edge computing device and the cloud server respectively in the network environment state; in the case where the first processing duration corresponding to the edge computing device is less than the second processing duration corresponding to the cloud server, use the target network model deployed in the edge computing device to detect abnormal behaviors, where the first processing duration includes the time required for data preprocessing and model inference on the edge computing device, and the second processing duration includes the time required for data preprocessing and model inference on the cloud server and the time required for data transmission between the cloud server and the edge computing device; in the case where the detected behavior category is an abnormal behavior, perform local alarm and upload the alarm information to the cloud.
[0140] Optionally, the abnormal behavior detection device is further configured to: determine the behavior category detected by the target network model as the pseudo-label corresponding to the video stream to be detected, and use the video stream to be detected and the pseudo-label as new training data to be added to the training data set; use the updated training data set to continue training the target network model locally deployed on the edge computing device; upload the model parameters updated during the training of the target network model in each edge computing device to the cloud server; use the cloud server to aggregate and fuse the updated model parameters uploaded by different edge computing devices to obtain a new target network model, and update the new target network model to each edge computing device.
[0141] Optionally, extracting the spatial features of the image frame to obtain the local features corresponding to the image frame includes: determining the gradient amplitude corresponding to each pixel point in the image frame, where the gradient amplitude is used to represent the edge intensity feature in the image frame; determining the edge information map corresponding to the image frame according to the gradient amplitude corresponding to each pixel point, where the edge information map is at least used to represent the contour information of the target object in the image frame; using the target network model to extract features from the image frame to obtain an initial feature map, and fusing the edge information map and the initial feature map to obtain the local features corresponding to the image frame.
[0142] Optionally, after obtaining the video stream to be detected, the data acquisition and feature extraction module 30 is further configured to: adjust the size of the image frame to a target size, where the target size is the input size supported by the network model used for feature extraction of the image frame; perform normalization processing on the pixel values of each point in the image frame of the target size, where the normalization processing is used to scale the pixel values of each point to a preset range interval.
[0143] It should be noted that each module in the above abnormal behavior detection device can be a program module (for example, a set of program instructions that implements a specific function), or a hardware module. For the latter, it can be presented in the following forms, but not limited to: the manifestation form of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.
[0144] It should be noted that the abnormal behavior detection device provided in this embodiment can be used to execute Figure 2 the abnormal behavior detection method shown, so the relevant explanations of the above abnormal behavior detection method also apply to the embodiments of this application, and will not be repeated here.
[0145] The embodiments of this application also provide a non-volatile storage medium, which includes a stored computer program. Among them, the device where the non-volatile storage medium is located executes the following abnormal behavior detection method by running the computer program: obtaining a video stream to be detected, and performing temporal feature extraction on multiple consecutive image frames in the video stream to be detected to obtain a temporal feature matrix, where the temporal feature matrix contains feature vectors corresponding to multiple consecutive image frames, and each image frame corresponds to a feature vector; determining the attention weights corresponding to the feature vectors in the temporal feature matrix, and determining the global feature corresponding to the temporal feature matrix according to the attention weights; performing spatial feature extraction on the image frame to obtain the local feature corresponding to the image frame; fusing the global feature and the local feature to obtain a target feature, and determining the behavior category of the target object in the video stream to be detected according to the target feature.
[0146] The embodiments of the present application also provide a computer program product, including a computer program, which when executed by a processor, implements the steps of the abnormal behavior detection method described in each embodiment of the present application: obtaining a video stream to be detected, and performing temporal feature extraction on multiple consecutive image frames in the video stream to be detected to obtain a temporal feature matrix, where the temporal feature matrix includes feature vectors corresponding to multiple consecutive image frames, and each image frame corresponds to one feature vector; determining the attention weights corresponding to the feature vectors in the temporal feature matrix, and determining the global feature corresponding to the temporal feature matrix according to the attention weights; performing spatial feature extraction on the image frames to obtain local features corresponding to the image frames; fusing the global feature and the local features to obtain a target feature, and determining the behavior category of the target object in the video stream to be detected according to the target feature.
[0147] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0148] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0149] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0150] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0151] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0152] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0153] The foregoing are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A method for detecting abnormal behavior, characterized in that: include: Acquire a video stream to be detected, and perform temporal feature extraction on a plurality of continuous image frames in the video stream to be detected to obtain a temporal feature matrix, wherein the temporal feature matrix contains feature vectors corresponding to the plurality of continuous image frames, and each image frame corresponds to one feature vector; Determine the attention weight corresponding to the feature vector in the time series feature matrix, and determine the global feature corresponding to the time series feature matrix based on the attention weight; Extracting spatial features from the image frame to obtain local features corresponding to the image frame; The global feature and the local feature are integrated to obtain a target feature, and the behavior category of the target object in the video stream to be detected is determined based on the target feature.
2. The abnormal behavior detection method according to claim 1, characterized in that: A target network model is used when extracting features from the image frame, wherein the training step of the target network model includes: Acquire a training data set, wherein the training data set includes a plurality of unlabeled video stream data, and the video stream data includes behavior segments of a target object under different environmental conditions; Using a contrast loss function, based on the training data set, the initial network model is subjected to contrast learning training to adjust the model parameters of the initial network model, wherein the contrast learning training is used to determine the similarity between the behavior features corresponding to the behavior segments in different video streams in the training data set, and to aggregate the behavior features whose similarity is greater than a first similarity threshold, and to distance the behavior features whose similarity is less than a second similarity threshold, wherein the first similarity threshold is greater than the second similarity threshold; And / or, using the initial network model, based on the behavior characteristics of multiple historical image frames in the video stream of the training data set, predicting the behavior characteristics of the next future frame corresponding to the multiple historical image frames, and determining the prediction error value between the behavior characteristics of the future frame predicted by the model and the behavior characteristics of the actual future frame in the video stream, and adjusting the model parameters of the initial network model based on the prediction error value; And / or, establish a time series association graph between graph nodes corresponding to the behavior features at different time steps in the training data set, and adjust the model parameters of the initial network model based on the association relationship characteristics between the graph nodes in the time series association graph.
3. The abnormal behavior detection method according to claim 2, characterized in that: After obtaining the training data set, the method further includes: Using a generator in an adversarial network, based on real video stream data in the training data set, pseudo video stream data is generated, wherein the behavior categories corresponding to the target object in the pseudo video stream data include various abnormal behaviors; Using the discriminator in the adversarial network, determine the confidence of the pseudo video stream data, and use the confidence to characterize the probability that the video stream data is real data; The pseudo video stream data whose confidence level is higher than a preset confidence threshold is added to the training data set.
4. The abnormal behavior detection method according to claim 2, characterized in that: The target network model is deployed on an edge computing device and a cloud server; the method further includes: Determine a network environment state, and determine, under the network environment state, a processing time required for performing abnormal behavior detection on the video stream to be detected on the edge computing device and the cloud server respectively; In a case where a first processing duration corresponding to the edge computing device is less than a second processing duration corresponding to the cloud server, the target network model deployed in the edge computing device is used to perform abnormal behavior detection, wherein the first processing duration includes: the time required for data preprocessing and model reasoning on the edge computing device, and the second processing duration includes: the time required for data preprocessing and model reasoning on the cloud server, and the time required for data transmission between the cloud server and the edge computing device; When the detected behavior category is abnormal behavior, a local alarm is issued and the alarm information is uploaded to the cloud.
5. The abnormal behavior detection method according to claim 4, characterized in that: The method further comprises: Determine the behavior category detected by the target network model as a pseudo label corresponding to the video stream to be detected, and add the video stream to be detected and the pseudo label as new training data to the training data set; Using the updated training data set, continue training the target network model locally deployed on the edge computing device; Uploading the model parameters updated during the training of the target network model in each edge computing device to the cloud server; The cloud server is used to aggregate and merge the updated model parameters uploaded by different edge computing devices to obtain a new target network model, and the new target network model is updated to each edge computing device.
6. The abnormal behavior detection method according to claim 1, characterized in that: Performing spatial feature extraction on the image frame to obtain local features corresponding to the image frame includes: Determining a gradient magnitude corresponding to each pixel point in the image frame, wherein the gradient magnitude is used to characterize an edge intensity feature in the image frame; Determining an edge information map corresponding to the image frame according to the gradient amplitude corresponding to each of the pixel points, wherein the edge information map is at least used to represent contour information of the target object in the image frame; The target network model is used to extract features from the image frame to obtain an initial feature map, and the edge information map and the initial feature map are fused to obtain the local features corresponding to the image frame.
7. The abnormal behavior detection method according to claim 1, characterized in that: After acquiring the video stream to be detected, the method further includes: Adjusting the size of the image frame to a target size, wherein the target size is an input size supported by a network model used for extracting features from the image frame; The pixel value of each point in the image frame of the target size is normalized, wherein the normalization is used to scale the pixel value of each point to a preset range.
8. An abnormal behavior detection device, characterized in that: include: A data acquisition and feature extraction module, used to acquire a video stream to be detected, and perform time series feature extraction on a plurality of continuous image frames in the video stream to be detected to obtain a time series feature matrix, wherein the time series feature matrix contains feature vectors corresponding to the plurality of continuous image frames, and each image frame corresponds to one feature vector; A global temporal feature determination module, used to determine the attention weight corresponding to the feature vector in the temporal feature matrix, and determine the global feature corresponding to the temporal feature matrix based on the attention weight; A local spatial feature determination module, used to extract spatial features from the image frame to obtain local features corresponding to the image frame; The abnormal behavior classification prediction module is used to fuse the global features and the local features to obtain target features, and determine the behavior category of the target object in the video stream to be detected based on the target features.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the abnormal behavior detection method according to any one of claims 1 to 7 is executed when the program is run.
10. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the abnormal behavior detection method according to any one of claims 1 to 7 by running the computer program.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the abnormal behavior detection method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Abnormal behavior detection method and device for video and electronic equipment
CN119763006A
Double-attention dense residual image rain removal method and system for 5G remote control
CN121458587A