Abnormal behavior monitoring detection method and device, computer equipment, storage medium and computer program product

By extracting frames and classifying the video data of infant monitoring, dangerous behaviors are identified and alarms are generated, the problem of insufficient efficiency and accuracy in the existing technology is solved, and efficient and accurate monitoring of abnormal behaviors is achieved.

CN120259971APending Publication Date: 2025-07-04LANTO ELECTRONIC LIMITED
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510352542.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing abnormal behavior detection methods have shortcomings in efficiency and accuracy, and it is difficult to effectively monitor dangerous behaviors such as climbing in infants and young children.

Method used

By obtaining monitoring video data, using the frame extraction algorithm to extract keyframe data, using the target detection model to identify the target area, and using the action classification model to analyze the action categories to generate monitoring alarm information.

Benefits of technology

It improves the efficiency and accuracy of abnormal behavior monitoring and detection, can promptly identify and alert dangerous actions, and enhances the convenience and safety of monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259971A_ABST
    Figure CN120259971A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent image processing, in particular to an abnormal behavior monitoring detection method and device, computer equipment, a storage medium and a computer program product. The abnormal behavior monitoring detection method comprises the following steps: acquiring monitoring data of a target object, wherein the monitoring data comprises monitoring video data; performing frame extraction processing on the monitoring video data based on a preset frame extraction algorithm, and extracting to obtain key frame data; calling a pre-constructed target detection model to process the key frame data, and determining a target area containing the target object in the key frame data; calling a pre-constructed action classification model to analyze the target area, and determining action category information of the target object in the current key frame; and in response to the fact that the action category information of the target object accords with the preset target action type, generating monitoring alarm information. By adopting the method, the monitoring efficiency and accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent image processing technology, and particularly to an abnormal behavior monitoring and detection method, device, computer device, storage medium, and computer program product. Background Art

[0002] During the growth process of infants and young children, climbing is a gross motor skill that users are relatively concerned about. After infants and young children learn to climb, due to curiosity, they may perform dangerous actions such as climbing high, which can cause injury to the child. Guardians may sometimes fail to notice these dangerous behaviors in daily life. If a monitor has the function of detecting the climbing of infants and young children, it can achieve automatic monitoring and warning of dangerous behaviors, and can have a good application prospect.

[0003] In related technologies, the abnormal behavior monitors for infants and young children mainly rely on the rapid development of fields such as sensor technology, Internet of Things (IoT), artificial intelligence (AI), and big data analysis. With the progress of technology, traditional infant care methods have gradually been replaced by intelligent devices, which can monitor the physiological status and behaviors of infants and young children in real time, helping parents or caregivers to detect abnormal situations in a timely manner. First of all, sensor technology is the core foundation of the abnormal behavior monitor for infants and young children. Through a variety of sensors (such as temperature sensors, heart rate sensors, acceleration sensors, sound sensors, etc.), the device can collect data such as the body temperature, heart rate, respiratory rate, body movement, and crying of infants and young children in real time. These sensors are usually integrated into wearable devices (such as smart bracelets, smart clothing) or non-contact devices (such as smart cameras, mattress sensors) to ensure the continuity and accuracy of monitoring. Secondly, Internet of Things technology enables these sensor data to be transmitted to the cloud or mobile terminal in real time. Through Wi-Fi, Bluetooth, or cellular networks, parents can view the status of infants and young children at any time and place through a mobile application, and receive instant reminders when abnormalities occur. This remote monitoring function greatly improves the convenience and efficiency of care. The introduction of artificial intelligence technology further enhances the intelligence level of the abnormal behavior monitor for infants and young children. Through the analysis of a large amount of historical data, AI algorithms can identify the normal behavior patterns of infants and young children, and on this basis, detect abnormal behaviors (such as abnormal crying, apnea, high body temperature, etc.). Machine learning technology can also provide personalized care suggestions according to the individual differences of infants and young children.

[0004] However, the current abnormal behavior detection methods have the following technical problems:

[0005] The current abnormal behavior detection methods need to be optimized in terms of efficiency and accuracy. Summary of the Invention

[0006] Based on this, it is necessary to provide an abnormal behavior monitoring and detection method, device, computer device, computer-readable storage medium, and computer program product that can improve the efficiency and accuracy of abnormal behavior detection for the above technical problems.

[0007] In a first aspect, the present application provides an abnormal behavior monitoring and detection method. The abnormal behavior monitoring and detection method includes:

[0008] Obtain monitoring data of a target object, where the monitoring data includes monitoring video data;

[0009] Perform frame extraction processing on the monitoring video data based on a preset frame extraction algorithm to extract key frame data;

[0010] Call a pre-built target detection model to process the key frame data and determine the target area containing the target object in the key frame data;

[0011] Call a pre-built action classification model to analyze the target area and determine the action category information of the target object in the current key frame;

[0012] In response to the action category information of the target object conforming to a preset target action type, generate a guardianship warning message.

[0013] In one embodiment, performing frame extraction processing on the monitoring video data based on a preset frame extraction algorithm to extract key frame data includes:

[0014] In response to the monitoring setting information of the target device, determine the analysis accuracy and computing resources set for the monitoring data;

[0015] Based on the analysis accuracy and computing resources, select a target frame extraction strategy that matches the analysis accuracy and computing resources from a number of preset candidate frame extraction strategies.

[0016] In one embodiment, performing frame extraction processing on the monitoring video data based on a preset frame extraction algorithm to extract key frame data includes:

[0017] Traverse and process the monitoring video data to obtain the deviation comparison result between adjacent video frames in the monitoring video data;

[0018] Mark video frames with a deviation degree represented by the deviation comparison result lower than a preset difference threshold as redundant video frames and delete the redundant video frames.

[0019] In one embodiment, traversing and processing the monitoring video data to obtain the deviation comparison result between adjacent video frames in the monitoring video data includes:

[0020] Identify significant regions in video frames based on a preset saliency detection algorithm, where the significant regions include target objects and target objects;

[0021] Determine the deviation comparison result based on the deviation comparison processing of the significant regions in the video frames.

[0022] In one embodiment, call a pre - built action classification model to analyze the target region, and determine the action category information of the target object in the current key frame, including:

[0023] Obtain a target region image sequence based on the target region, and input the target region image sequence into a first feature detection path and a second feature detection path respectively. The processing frame rate of the first feature detection path is lower than that of the second feature detection path;

[0024] Process the target region image sequence using the first feature detection path to obtain slow - speed feature information, which is used to characterize the spatial semantic information and long - time - scale motion information in the target region;

[0025] Process the target region image sequence using the second feature detection path to obtain fast - speed feature information, which is used to characterize the fast - motion change information in the target region.

[0026] In one embodiment, the method further includes:

[0027] Construct a generative adversarial network model, and train the first feature detection path and the second feature detection path based on the generative adversarial network model to obtain optimized first and second feature detection paths;

[0028] Process the target region image sequence based on the optimized first and second feature detection paths.

[0029] In a second aspect, the present application also provides an abnormal behavior monitoring and detection device. The abnormal behavior monitoring and detection device includes:

[0030] A monitoring data module, configured to obtain monitoring data of the target object, where the monitoring data includes monitoring video data;

[0031] A frame extraction processing module, configured to perform frame extraction processing on the monitoring video data based on a preset frame extraction algorithm to extract key frame data;

[0032] A target detection module, configured to call a pre - built target detection model to process the key frame data and determine the target region in the key frame data that contains the target object;

[0033] An action analysis module, configured to call a pre-built action classification model to analyze a target area and determine action category information of a target object in a current key frame;

[0034] An abnormal alarm module, configured to generate a guardianship alarm message in response to the action category information of the target object conforming to a preset target action type.

[0035] In one embodiment, the frame extraction processing module includes:

[0036] A monitoring setting module, configured to determine an analysis accuracy and computing resources set for monitoring data in response to monitoring setting information of a target device;

[0037] A target policy module, configured to select a target frame extraction policy matching the analysis accuracy and computing resources from a plurality of preset candidate frame extraction policies based on the analysis accuracy and computing resources.

[0038] In one embodiment, the frame extraction processing module includes:

[0039] A deviation comparison module, configured to traverse and process monitoring video data to obtain a deviation comparison result between adjacent video frames in the monitoring video data;

[0040] A redundancy deletion module, configured to mark video frames with a deviation degree characterized by the deviation comparison result lower than a preset difference threshold as redundant video frames and delete the redundant video frames.

[0041] In one embodiment, the deviation comparison module includes:

[0042] A significant identification module, configured to identify significant regions in a video frame based on a preset saliency detection algorithm, where the significant regions include a target object and a target object;

[0043] A region comparison module, configured to determine a deviation comparison result based on deviation comparison processing of significant regions in a video frame.

[0044] In one embodiment, the action analysis module includes:

[0045] An image sequence module, configured to obtain an image sequence of a target area based on the target area, and input the image sequence of the target area into a first feature detection path and a second feature detection path respectively, where the processing frame rate of the first feature detection path is lower than that of the second feature detection path;

[0046] A first feature detection path module, configured to process the image sequence of the target area by using the first feature detection path to obtain slow feature information, where the slow feature information is used to characterize spatial semantic information and long-time scale motion information in the target area;

[0047] A second feature detection path is used to process the target area image sequence by applying the second feature detection path to obtain fast feature information, and the fast feature information is used to characterize the fast motion change information in the target area.

[0048] In one embodiment, the abnormal behavior monitoring and detection device further includes:

[0049] A model optimization module is used to construct a generative adversarial network model, and based on the generative adversarial network model, train the first feature detection path and the second feature detection path to obtain the optimized first feature detection path and the second feature detection path;

[0050] An optimization application module is used to process the target area image sequence based on the optimized first feature detection path and the second feature detection path.

[0051] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps in an abnormal behavior monitoring and detection method according to any one of the embodiments in the first aspect.

[0052] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in an abnormal behavior monitoring and detection method according to any one of the embodiments in the first aspect.

[0053] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in an abnormal behavior monitoring and detection method according to any one of the embodiments in the first aspect.

[0054] The above abnormal behavior monitoring and detection method, device, computer device, storage medium, and computer program product can achieve the beneficial effects corresponding to the technical problems in the background art through the derivation of the technical features in the claims:

[0055] The present application provides an abnormal behavior monitoring method, which includes obtaining monitoring data of a target object, where the monitoring data includes monitoring video data; performing frame extraction processing on the monitoring video data based on a preset frame extraction algorithm to extract key frame data; calling a pre-constructed target detection model to process the key frame data to determine a target area containing the target object in the key frame data; calling a pre-constructed action classification model to analyze the target area to determine action category information of the target object in the current key frame; and generating a guardianship alarm message in response to the action category information of the target object conforming to a preset target action type. In implementation, the monitoring data of the target object is collected in real time through monitoring devices (such as cameras, infrared sensors, etc.). The monitoring data includes monitoring video data, but is not limited to detection video data, and may also include detection image data. These data can cover static images and dynamic behavior information of the target object, providing a basis for subsequent analysis. The monitoring video data is subjected to frame extraction processing based on a preset frame extraction algorithm to extract key frame data. The purpose of the frame extraction algorithm is to screen out key frames containing important information from the video stream to reduce data redundancy and improve processing efficiency. Common frame extraction algorithms include uniform frame extraction based on time intervals, dynamic frame extraction based on motion detection, etc. A pre-constructed target detection model is called to process the key frame data to determine the target area containing the target object in the key frame data. The target detection model is usually based on deep learning technologies (such as YOLO, Faster R-CNN, etc.) and can accurately identify the position and bounding box of the target object in the image. A pre-constructed action classification model is called to analyze the target area to determine the action category information of the target object in the current key frame. The action classification model usually adopts a convolutional neural network (CNN) or a temporal model (such as LSTM, 3D CNN, etc.) and can identify specific actions of the target object (such as falling, abnormal postures, violent movements, etc.). In response to the action category information of the target object conforming to a preset target action type (such as abnormal behavior or dangerous action), a guardianship alarm message is generated. The alarm message can be transmitted to the guardianship personnel through means such as sound, light, mobile phone push notifications, etc. so as to take intervention measures in time. In summary, by generally sampling the above steps, the purpose of improving the efficiency and accuracy of abnormal behavior guardianship detection is achieved. Description of the Drawings

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for describing the embodiments of the present application or related technologies. Obviously, the following-described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0057] Figure 1It is an application environment diagram of an abnormal behavior monitoring and detection method in an embodiment;

[0058] Figure 2 It is the first process schematic diagram of an abnormal behavior monitoring and detection method in an embodiment;

[0059] Figure 3 It is the second process schematic diagram of an abnormal behavior monitoring and detection method in another embodiment;

[0060] Figure 4 It is the third process schematic diagram of an abnormal behavior monitoring and detection method in another embodiment;

[0061] Figure 5 It is the fourth process schematic diagram of an abnormal behavior monitoring and detection method in another embodiment;

[0062] Figure 6 It is the fifth process schematic diagram of an abnormal behavior monitoring and detection method in another embodiment;

[0063] Figure 7 It is the sixth process schematic diagram of an abnormal behavior monitoring and detection method in another embodiment;

[0064] Figure 8 It is the structural block diagram of an abnormal behavior monitoring and detection device in an embodiment;

[0065] Figure 9 It is the internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0066] In order to make the purpose, technical solutions and advantages of this application clearer, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0067] An abnormal behavior monitoring and detection method provided by an embodiment of this application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed on the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0068] In one embodiment, as Figure 2 shown, an abnormal behavior monitoring and detection method is provided. Taking the terminal in Figure 1 as an example, the method includes the following steps:

[0069] Step 202: Obtain the monitoring data of the target object. The monitoring data includes monitoring video data.

[0070] Step 204: Perform frame extraction processing on the monitoring video data based on a preset frame extraction algorithm to extract key frame data.

[0071] Step 206: Invoke a pre-built target detection model to process the key frame data and determine the target area containing the target object in the key frame data.

[0072] Exemplarily, the terminal can send the key frame data to a target detection model such as an SSD target detector for processing. The SSD target detector is built based on a deep convolutional neural network architecture. Before inputting the key frame, it is necessary to preprocess the image, including normalizing the image pixel values to a specific range (0 to 1) and adjusting the image size to match the input requirements of the SSD model (for example, common input sizes are 300*300 pixels or 512*512 pixels). After the key frame image data is input into the SSD model, feature extraction is performed through multiple convolutional layers. These convolutional layers gradually extract low-level to high-level features in the image, such as edge, texture, shape and other feature information.

[0073] The SSD model uses multi-scale feature maps to detect objects of different sizes. On feature maps of different scales, a series of pre-set prior boxes (anchor boxes) are utilized, and these prior boxes cover different positions and size ranges in the image. By performing classification and regression operations on the features within each prior box, it is determined whether there is an object within the prior box, as well as the category and precise location of the object. For example, for a person object, the model will, based on the learned person feature pattern, determine whether the image region within the prior box is a person, and through regression operations, adjust the position and size of the prior box to tightly enclose the person object region, ultimately obtaining a target region detection result containing the target category label "person" and precise coordinate information (the coordinates of the upper left corner and the lower right corner).

[0074] Step 208: Invoke the pre-built action classification model to analyze the target region and determine the action category information of the target object in the current key frame.

[0075] Step 2010: In response to the action category information of the target object conforming to the preset target action type, generate a guardianship alarm message.

[0076] In the above abnormal behavior guardianship detection method, through reasonable derivation in combination with the technical features in the embodiments, the following beneficial effects of solving the technical problems proposed in the background art can be achieved:

[0077] The present application provides an abnormal behavior monitoring method, which includes obtaining monitoring data of a target object, where the monitoring data includes monitoring video data; performing frame extraction processing on the monitoring video data based on a preset frame extraction algorithm to extract key frame data; calling a pre-constructed target detection model to process the key frame data to determine a target area containing the target object in the key frame data; calling a pre-constructed action classification model to analyze the target area to determine action category information of the target object in the current key frame; and generating a guardianship alarm message in response to the action category information of the target object conforming to a preset target action type. In implementation, monitoring data of the target object is collected in real time through monitoring devices (such as cameras, infrared sensors, etc.), and the monitoring data includes monitoring video data. These data can cover static images and dynamic behavior information of the target object, providing a basis for subsequent analysis. Frame extraction processing is performed on the monitoring video data based on a preset frame extraction algorithm to extract key frame data. The purpose of the frame extraction algorithm is to screen out key frames containing important information from the video stream to reduce data redundancy and improve processing efficiency. Common frame extraction algorithms include uniform frame extraction based on time intervals, dynamic frame extraction based on motion detection, etc. A pre-constructed target detection model is called to process the key frame data to determine a target area containing the target object in the key frame data. The target detection model is usually based on deep learning technologies (such as YOLO, Faster R-CNN, etc.) and can accurately identify the position and bounding box of the target object in the image. A pre-constructed action classification model is called to analyze the target area to determine action category information of the target object in the current key frame. The action classification model usually adopts a convolutional neural network (CNN) or a temporal model (such as LSTM, 3D CNN, etc.) and can identify specific actions of the target object (such as falling, abnormal postures, violent movements, etc.). In response to the action category information of the target object conforming to a preset target action type (such as abnormal behavior or dangerous actions), a guardianship alarm message is generated. The alarm message can be transmitted to the guardianship personnel through means such as sound, light, mobile phone push notifications, etc. so as to take intervention measures in a timely manner. In summary, usually by sampling the above steps, the purpose of improving the efficiency and accuracy of abnormal behavior guardianship detection is achieved.

[0078] In one embodiment, as Figure 3 shown: Step 204 includes:

[0079] Step 302: In response to the monitoring setting information of the target device, determine the analysis accuracy and computing resources set for the monitoring data.

[0080] Step 304: Based on the analysis accuracy and computing resources, select a target frame extraction strategy that matches the analysis accuracy and computing resources from a number of preset candidate frame extraction strategies.

[0081] Exemplarily, the terminal can determine the frame rate information of the video. For example, common video frame rates include 25fps (25 frames per second), 30fps, etc. According to the required analysis accuracy and computing resources, a matching frame extraction strategy is selected. For example, in the uniform frame extraction method, for a video with a frame rate of 30fps, it can be set to extract 1 frame every 5 frames, that is, extract one frame every 166.67 milliseconds. In this way, 6 video pictures will be extracted within one second.

[0082] In this embodiment, during the process of frame extraction for the monitored video data, according to the preset analysis accuracy and computing resources of the monitored data, a suitable target frame extraction strategy is selected, which helps to make the monitoring process meet the application requirements and improves the flexibility of method application.

[0083] In one embodiment, as Figure 4 shown, step 204 includes:

[0084] Step 402: Traverse and process the monitored video data to obtain the deviation comparison result between adjacent video frames in the monitored video data;

[0085] Step 404: Mark the video frames with a deviation degree represented by the deviation comparison result lower than the preset difference threshold as redundant video frames and delete the redundant video frames.

[0086] Exemplarily, during the process of frame extraction for the video, the terminal can traverse each frame of the video, save the pictures with larger differences by comparing the similarity between the front and back frame pictures, and discard the similar pictures, so as to automatically delete the redundant frames and obtain the effective frames. This reduces the subsequent annotation work and obtains more scenes with larger differences. For the specific method of similarity measurement, for example, the perceptual hashing algorithm can be applied. The specific steps are as follows:

[0087] Step 1: Resize the picture to 32 * 32 to facilitate DCT (Discrete Cosine Transform) calculation;

[0088] Step 2: Convert the resized picture into a 256-level grayscale picture;

[0089] Step 3: Convert the grayscale picture to a floating-point type, and then perform DCT transformation, that is, perform discrete cosine transform (DCT) on the image, and usually a coefficient matrix with the same size as the image will be obtained;

[0090] Step 4: Reduce the DCT. The DCT is 32*32, and keep the upper left 8*8, which represents the lowest frequency of the picture, that is, the general outline and texture of the image;

[0091] Step 5: Calculate the average value of all pixel points after reducing the DCT;

[0092] Step 6: Further reduce the DCT. If it is greater than the average value, record it as 1; otherwise, record it as 0.

[0093] Step 7: Combine the 1s and 0s generated in the above steps in order. Combine 64 information bits, and just keep the order consistent randomly.

[0094] Step 8: Calculate the hash values of the two pictures, calculate the Hamming distance (how many times need to be changed from one hash value to another). The greater the Hamming distance, the more inconsistent the pictures are. On the contrary, the smaller the Hamming distance, the more similar the pictures are. When the distance is 0, it means they are exactly the same. (Generally, it is considered that when the distance > 10, the two pictures are completely different).

[0095] In this embodiment, during the process of extracting frames from video frames, redundant video frames are deleted according to the deviation comparison result between adjacent video frames, which helps to improve the efficiency of video frame processing.

[0096] In one of the embodiments, as Figure 5 shown, step 402 includes:

[0097] Step 502: Identify the significant regions in the video frame based on a preset saliency detection algorithm. The significant regions include the target object and the target object.

[0098] Step 504: Determine the deviation comparison result based on the deviation comparison processing of the significant regions in the video frame.

[0099] Exemplarily, the terminal can use the saliency detection algorithm of the image to analyze the significant regions (such as people, key objects, etc.) in the video frame, and preferentially extract the frames with larger changes in the significant regions. When the action amplitude or position of the people in the picture changes beyond a certain threshold, frame extraction is performed. In this way, on the premise of ensuring that key information is not lost, the extraction of redundant frames can be reduced, the subsequent processing efficiency can be improved, and the data storage requirement can be reduced. During the frame extraction process, function interfaces provided by a video processing library (such as OpenCV) are used to perform extraction operations at a set time interval or frame number. By reading the video file stream, positioning to a specific frame position, and extracting and storing the frame image data in a specific image format (such as JPEG, PNG, etc.) in memory or on the hard disk, a series of discrete video picture sequences are formed, and these pictures will be used as the basic data for subsequent object detection.

[0100] In this embodiment, during the process of deviation comparison between video frames, significant regions are pre-identified. By comparing the deviations of the significant regions, the effectiveness of the deviation comparison is enhanced, the deletion efficiency and accuracy of redundant video frames are further improved, and thus the overall implementation efficiency and accuracy of the method are enhanced.

[0101] In one of the embodiments, as Figure 6As shown, step 208 includes:

[0102] Step 602: Obtain a target region image sequence based on the target region, and input the target region image sequence into a first feature detection path and a second feature detection path respectively. The processing frame rate of the first feature detection path is lower than that of the second feature detection path.

[0103] Step 604: Process the target region image sequence using the first feature detection path to obtain slow feature information, which is used to characterize the spatial semantic information and long-term motion information in the target region.

[0104] Step 606: Process the target region image sequence using the second feature detection path to obtain fast feature information, which is used to characterize the fast motion change information in the target region.

[0105] Exemplarily, the terminal can input the target region image sequence into a SlowFast network for processing. In this embodiment, the first feature detection path corresponds to the slow path, and the second feature detection path corresponds to the fast path. The SlowFast network consists of a slow pathway and a fast pathway. First, the target region image sequence detected by SSD is cropped and scaled to meet the input requirements of the SlowFast model. For each target region image sequence, it is input into the Slow pathway and the Fast pathway of SlowFast respectively. The Slow pathway processes the image sequence at a lower frame rate (such as taking 1 frame every 16 frames), focusing on capturing the spatial semantic information and long-term motion information in the image, and performing feature extraction and dimensionality reduction operations through a series of convolutional layers and pooling layers to obtain a slow feature representation. The Fast pathway processes the image sequence at a higher frame rate (such as taking 1 frame every 2 frames), focusing on capturing the fast motion change information in the image. After being processed by convolutional layers and pooling layers in the same way, a fast feature representation is obtained.

[0106] In this embodiment, the target region image sequence is independently processed through two feature monitoring paths, so as to obtain spatial semantic information and long-term motion information respectively, which helps to make the two types of information verify and support each other, and helps to improve the efficiency and accuracy of action category analysis.

[0107] In one embodiment, as Figure 7 shown, the method further includes:

[0108] Step 702: Construct a generative adversarial network model, and train the first feature detection path and the second feature detection path based on the generative adversarial network model to obtain optimized first and second feature detection paths;

[0109] Step 704: Process the target region image sequence based on the optimized first feature detection path and the second feature detection path.

[0110] Exemplarily, the terminal can introduce an adversarial learning mechanism to enhance the generalization ability of the model. Construct an adversarial generative network (GAN) combined with the SlowFast network, where the generator is responsible for generating some virtual target region image sequences that are similar to real climbing behaviors but have slight differences, and the discriminator is used to distinguish between real and generated image sequences. The SlowFast network is trained in this adversarial environment so that it can better learn the essential features of climbing behaviors, without being limited to specific scenarios and styles in the training data, thereby improving the accuracy and robustness of the model in identifying climbing behaviors in different environments and conditions. The results of SlowFast and adversarial learning are fused through lateral connections, enabling the model to comprehensively consider motion information at different spatial, temporal, and frame rates, as well as the generalization improvement brought by adversarial learning. Based on the fused features, action classification operations are performed through fully connected layers, and according to the pre-trained action category model, the action category being executed by the target in the target region is judged. For climbing behaviors, the model will learn feature patterns such as the posture changes of the human body, hand and foot movements, and relative motion relationships with the surrounding environment during the climbing process. When the input target region image sequence conforms to these feature patterns, the model determines it as the climbing behavior category, thus achieving the purpose of climbing behavior recognition. Throughout the process, a large amount of video data labeled with climbing behaviors and other related action categories is required to train the SlowFast model, and the parameters of the model are continuously adjusted through the backpropagation algorithm to improve the recognition accuracy and generalization ability of the model.

[0111] In specific implementation, construct a SlowFast - GAN network architecture:

[0112] The SlowFast network part can include constructing a Slow path:

[0113] Base Model Selection and Modification: Build the Slow path based on ResNet50. ResNet50 has a series of convolutional layers, batch normalization layers, and residual modules. First, remove the last fully connected layer and average pooling layer of ResNet50 because in the SlowFast network, for the Slow path, semantic features of the video need to be extracted, and subsequent operations such as fusion will be performed. Keep the previous convolutional layer (such as conv1), batch normalization layer (bn1), activation function layer (such as relu), and multiple residual modules (layer1 - layer4).

[0114] Input and Feature Extraction Process: The input of the Slow path is video frame data, usually in the format [batch_size, channels, time_steps, height, width]. The data first undergoes a convolutional operation through the conv1 layer, then batch normalization through the bn1 layer, and then passes through the relu activation function. After that, it undergoes a max pooling operation through the maxpool layer and then sequentially passes through the layer1 - layer4 residual modules to gradually extract the semantic features of the video frames.

[0115] Fast Path Construction including Base Model Modification: Similarly, build the Fast path with reference to ResNet50. However, the Fast path is mainly used to capture details such as fast actions in the video, so some adjustments are needed. For example, for the conv1 layer, change its input channel number, kernel size, stride, and padding. The input channel number is generally still corresponding to the channel number of the video frame (such as 3 for RGB channels), but the kernel size is set to (5, 7, 7), the stride is set to (1, 2, 2), and the padding is set to (2, 3, 3). Such settings are to make it more suitable for processing high-frame-rate video data. The subsequent bn1, relu, maxpool layers, and layer1 - layer4 residual modules can then adopt a structure similar to ResNet50 to extract features.

[0116] Feature Extraction Process: The input of the Fast path is also video frame data. It undergoes convolution through the modified conv1 layer, then passes through the bn1, relu, and maxpool layers, and then through the layer1 - layer4 residual modules to extract features that can capture detail information.

[0117] Fusion Layer Construction including Fusion Method Selection: The fusion layer is used to fuse the features extracted by the Slow path and the Fast path. A common fusion method is concatenation.

[0118] Fusion operation details: Assume that after the previous processing of the Slow path and the Fast path, the output feature channel numbers are certain values (for example, both are 2048). In the fusion layer, first, the features output by the Slow path and the Fast path are concatenated in the channel dimension. The concatenated features are then fused and adjusted through a 1x1x1 convolutional layer. The input channel number of this convolutional layer is the sum of the output channel numbers of the Slow and Fast paths, and the output channel number can be set as needed (such as set to 2048 here). Finally, it passes through an activation function (such as relu) to output the fused features.

[0119] The detection head construction includes a location prediction module: Add a module for predicting the target location. This module can be a convolutional layer, and the output channel number is determined according to the number of location parameters to be predicted. For example, if predicting the coordinates of the bounding box (xmin, ymin, xmax, ymax), the output channel number can be set to 4. The kernel size of these convolutional layers can be adjusted according to the size of the feature map and the requirements of detection accuracy. Generally, smaller kernels are used, such as 3x3 or 1x1.

[0120] Category prediction module: Construct a module for predicting the target category information. It can be a convolutional layer, and the output channel number is the number of target categories. For example, if detecting 1 different object, the output channel number is set to 1. Similarly, the kernel size can also be selected according to the actual situation.

[0121] Integrate the prediction results: Integrate the outputs of the location prediction module and the category prediction module to obtain the final detection results. They can be concatenated in the channel dimension to form a feature map containing location and category information. The feature vector at each location represents the information at that location.

[0122] The GAN network part includes the construction of the Generator:

[0123] The generator takes the fused SlowFast features as input and can use multiple fully connected layers or transposed convolutional layers to implement the upsampling operation, gradually mapping the features to a higher-dimensional space, and finally outputting data with the same or similar dimensions as the original video frame (such as [batch_size, channels, new_time_steps, new_height, new_width]). First, it passes through a fully connected layer to expand the input feature dimension, and then gradually restores to the appropriate video frame size through several transposed convolutional layers. Activation functions such as relu are used between each layer to increase the non-linear expression ability.

[0124] The construction of the Discriminator:

[0125] The discriminator takes as input either real video frames or video frames generated by the generator. It extracts features through a series of convolutional layers (such as 3x3x3 convolutional kernels) and pooling layers, and then outputs a scalar through a fully connected layer, representing the probability that the input data is real data. The relu activation function and batch normalization layers can be used between convolutional layers to accelerate training and improve the discriminative ability.

[0126] The combination method and training strategy of SlowFast and GAN

[0127] During the training process, first fix the generator of the GAN network and train the discriminator to distinguish between real video frames and fake video frames obtained by passing the features generated by the SlowFast network through the generator, using the cross-entropy loss function. Then fix the discriminator and train the generator to make the video frames it generates able to deceive the discriminator. At the same time, send the video frames generated by the generator back into the SlowFast network, combine the losses of the detection task (the location loss uses the Smooth L1 loss, the class loss uses the cross-entropy loss, and they are added together with certain weights to obtain the total detection loss), and backpropagate to update the parameters of the SlowFast network and the generator, so that the network can better learn the video feature representation and improve the detection performance.

[0128] In this embodiment, the first feature detection path and the second feature detection path are trained through the generative adversarial network model, so that the performance of the first feature detection path and the second feature detection path can be optimized and enhanced, which helps to further improve the efficiency and accuracy of the method implementation.

[0129] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0130] Based on the same inventive concept, an embodiment of the present application further provides an abnormal behavior monitoring and detection device for implementing the abnormal behavior monitoring and detection method involved above. The solution provided by this device to solve the problem is similar to the solution recorded in the above method. Therefore, the specific limitations in one or more embodiments of the abnormal behavior monitoring and detection device provided below can refer to the limitations on the abnormal behavior monitoring and detection method in the above text, and will not be repeated here.

[0131] In one embodiment, as Figure 8 shown, an abnormal behavior monitoring and detection device is provided, including: a monitoring data module, a frame extraction processing module, a target detection module, an action analysis module, and an abnormal alarm module, where:

[0132] The monitoring data module is used to obtain the monitoring data of the target object, and the monitoring data includes monitoring video data;

[0133] The frame extraction processing module is used to perform frame extraction processing on the monitoring video data based on a preset frame extraction algorithm to extract key frame data;

[0134] The target detection module is used to call a pre-built target detection model to process the key frame data and determine the target area containing the target object in the key frame data;

[0135] The action analysis module is used to call a pre-built action classification model to analyze the target area and determine the action category information of the target object in the current key frame;

[0136] The abnormal alarm module is used to generate a monitoring alarm message in response to the action category information of the target object conforming to a preset target action type.

[0137] In one of the embodiments, the frame extraction processing module includes:

[0138] The monitoring setting module is used to determine the analysis accuracy and computing resources set for the monitoring data in response to the monitoring setting information of the target device;

[0139] The target strategy module is used to select a target frame extraction strategy that matches the analysis accuracy and computing resources from a number of preset candidate frame extraction strategies based on the analysis accuracy and computing resources.

[0140] In one of the embodiments, the frame extraction processing module includes:

[0141] The deviation comparison module is used to traverse and process the monitoring video data to obtain the deviation comparison result between adjacent video frames in the monitoring video data;

[0142] A redundancy deletion module, configured to mark video frames with a deviation degree characterized by the deviation comparison result lower than a preset difference threshold as redundant video frames, and delete the redundant video frames.

[0143] In one embodiment, the deviation comparison module includes:

[0144] A significant identification module, configured to identify significant regions in a video frame based on a preset significance detection algorithm, where the significant regions include a target object and a target object;

[0145] A region comparison module, configured to determine a deviation comparison result based on deviation comparison processing of the significant regions in the video frame.

[0146] In one embodiment, the action analysis module includes:

[0147] An image sequence module, configured to obtain a target region image sequence based on a target region, and input the target region image sequence into a first feature detection path and a second feature detection path respectively, where the processing frame rate of the first feature detection path is lower than that of the second feature detection path;

[0148] A first feature detection path module, configured to process the target region image sequence using the first feature detection path to obtain slow feature information, where the slow feature information is used to characterize spatial semantic information and long-time scale motion information in the target region;

[0149] A second feature detection path, configured to process the target region image sequence using the second feature detection path to obtain fast feature information, where the fast feature information is used to characterize fast motion change information in the target region.

[0150] In one embodiment, the apparatus further includes:

[0151] A model optimization module, configured to construct a generative adversarial network model, and train the first feature detection path and the second feature detection path based on the generative adversarial network model to obtain an optimized first feature detection path and an optimized second feature detection path;

[0152] An optimization application module, configured to process the target region image sequence based on the optimized first feature detection path and the optimized second feature detection path.

[0153] Each module in the above abnormal behavior monitoring and detection device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0154] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structural diagram may be as shown in Figure 9 . The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an abnormal behavior monitoring and detection method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0155] Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0156] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.

[0157] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in the above method embodiments.

[0158] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps in the above method embodiments.

[0159] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0160] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.

[0161] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0162] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. An abnormal behavior monitoring and detection method, characterized in that, The method includes: Obtaining monitoring data of a target object, where the monitoring data includes monitoring video data; Performing frame extraction processing on the monitoring video data based on a preset frame extraction algorithm to extract key frame data; Invoking a pre-constructed target detection model to process the key frame data and determining a target area in the key frame data that contains the target object; Invoking a pre-constructed action classification model to analyze the target area and determining action category information of the target object in the current key frame; Generating a guardianship warning message in response to the action category information of the target object conforming to a preset target action type.

2. The method according to claim 1, characterized in that The performing frame extraction processing on the monitoring video data based on a preset frame extraction algorithm to extract key frame data includes: Determining the analysis accuracy and computing resources set for the monitoring data in response to the monitoring setting information of the target device; Selecting a target frame extraction strategy that matches the analysis accuracy and the computing resources from a number of preset candidate frame extraction strategies based on the analysis accuracy and the computing resources.

3. The method according to claim 1, characterized in that, The performing frame extraction processing on the monitoring video data based on a preset frame extraction algorithm to extract key frame data includes: Traversing and processing the monitoring video data to obtain a deviation comparison result between adjacent video frames in the monitoring video data; Marking video frames with a deviation degree represented by the deviation comparison result lower than a preset difference threshold as redundant video frames and deleting the redundant video frames.

4. The method according to claim 3, characterized in that, The traversing and processing the monitoring video data to obtain a deviation comparison result between adjacent video frames in the monitoring video data includes: Identifying a significant area in the video frame based on a preset saliency detection algorithm, where the significant area includes the target object and target objects; Determining the deviation comparison result based on the deviation comparison processing of the significant area in the video frame.

5. The method according to any one of claims 1 to 4, characterized in that, The invoking a pre-constructed action classification model to analyze the target area and determining action category information of the target object in the current key frame includes: Obtaining a target area image sequence based on the target area, and inputting the target area image sequence into a first feature detection path and a second feature detection path respectively, where the processing frame rate of the first feature detection path is lower than that of the second feature detection path; Processing the target area image sequence using the first feature detection path to obtain slow feature information, where the slow feature information is used to represent spatial semantic information and long-time scale motion information in the target area; Processing the target area image sequence using the second feature detection path to obtain fast feature information, where the fast feature information is used to represent fast motion change information in the target area.

6. The method according to claim 5, wherein The method further includes: Constructing a generative adversarial network model, training the first feature detection path and the second feature detection path based on the generative adversarial network model to obtain optimized first and second feature detection paths; Processing the target area image sequence based on the optimized first and second feature detection paths.

7. An abnormal behavior monitoring and detection device, characterized in that, The device includes: A monitoring data module, configured to obtain monitoring data of a target object, where the monitoring data includes monitoring video data; A frame extraction processing module, configured to perform frame extraction processing on the monitoring video data based on a preset frame extraction algorithm to extract key frame data; A target detection module, configured to call a pre-constructed target detection model to process the key frame data and determine a target area in the key frame data that contains the target object; An action analysis module, configured to call a pre-constructed action classification model to analyze the target area and determine action category information of the target object in the current key frame; An abnormal alarm module, configured to generate a guardianship alarm message in response to the action category information of the target object conforming to a preset target action type.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that, When this computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Video abnormal behavior real-time detection method and system based on deep learning

    CN122244956A

  • Deep learning-based real-time video abnormal behavior detection method and system

    CN122244956B