Abnormal behavior detection method and device, electronic equipment and storage medium
By training a behavior prediction model to process image sets from power self-service terminal equipment, the probability of abnormal behavior is identified, which solves the problem of low recognition accuracy in existing technologies and achieves higher accuracy in abnormal behavior recognition.
Patent Information
- Application Number
- CN202311353480.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-18
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2043-10-18
AI Technical Summary
In existing technologies, the accuracy rate of image recognition for abnormal behavior of power self-service terminal equipment is low, especially in cases where the tool may be engaging in destructive behavior even though it has not touched the equipment.
By using a pre-trained first behavior prediction model, a time-continuous image set is processed to obtain the probability of a target object performing abnormal behavior in multiple image time periods. Features are extracted using convolutional neural networks and recurrent neural networks, and the model parameters are optimized by combining the cross-entropy loss function and an optimizer to identify abnormal behavior.
It improves the accuracy of abnormal behavior recognition results, enabling more accurate identification of actions within multiple image time periods, reducing false judgments, and improving the reliability of equipment damage warning.
Smart Images

Figure CN117218726B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and more specifically, to an abnormal behavior detection method, apparatus, electronic device, and storage medium. Background Technology
[0002] Currently, self-service power terminals can acquire images through cameras and identify the images to determine if there is any abnormal behavior in the images, such as using tools to damage the self-service power terminal.
[0003] However, since the damage may involve not just a single action but a combination of actions, the accuracy of identification using the above method is relatively low. For example, taking the example of using a tool to damage a self-service electricity terminal, the captured images may include images of the tool not touching the terminal and images of the tool touching the terminal. If the image of the tool not touching the terminal is identified, since there was no contact between the tool and the terminal, the identification result for that image may be normal behavior.
[0004] Therefore, improving the accuracy of abnormal behavior identification results has become an urgent problem to be solved. Summary of the Invention
[0005] In view of the above problems, this application is made to provide an abnormal behavior detection method, apparatus, electronic device, and storage medium to improve the accuracy of abnormal behavior identification results. The specific solution is as follows:
[0006] An abnormal behavior detection method, the method comprising:
[0007] Obtain a set of images to be tested corresponding to the target device. The set of images to be tested contains multiple target images, which are temporally continuous and contain target objects.
[0008] The set of images to be tested is processed by the first behavior prediction model to obtain a target value. The target value is the probability that the target object will perform abnormal behavior within the time period to which the multiple target images belong. The abnormal behavior is the behavior that causes damage to the target device.
[0009] The first behavior prediction model is obtained by training a model based on a first training sample set. The first training sample set contains multiple first training samples, including a first input sample and a first output sample. The first input sample is a set of sample images, and each set of sample images contains multiple first sample images. The multiple first sample images are continuous in time, and each first sample image contains a first sample object. The first output sample is a first value, which represents the probability that the first sample object will perform the abnormal behavior within the time period to which the multiple first sample images belong.
[0010] Optionally, the set of images to be tested is processed using a first-row prediction model to obtain the target value, including:
[0011] Based on the set of images to be tested, multiple first feature values are obtained, each first feature value corresponds to one target image, and the first feature value is the probability that the target object performs abnormal behavior at the time of the corresponding target image.
[0012] Based on the set of images to be tested, target feature data is obtained, wherein the target feature data is data associated with the action of the target object within the time period to which the multiple target images belong;
[0013] The target value is obtained based on the target data and the plurality of first feature values.
[0014] Optionally, based on the set of images to be tested, multiple first feature values are obtained, including:
[0015] The second behavior prediction model is used to process each target image to obtain a first feature value corresponding to multiple target images, and one target image corresponds to one first feature value.
[0016] The second behavior prediction model is obtained by training the model based on a second training sample set. The second training sample set contains multiple second training samples, including a second input sample and a second output sample. The second input sample is a second sample image, which contains a second sample object. The second output sample is a second value, which is the probability that the second sample object will perform the abnormal behavior at the time when the second sample image is in the second sample image.
[0017] Optionally, target feature data is obtained based on the set of images to be tested, including:
[0018] Feature extraction is performed on the plurality of target images in the set of images to be tested to obtain a plurality of first feature data. Each target image corresponds to one first feature data. The first feature data is data associated with the action of the target object at the time when the corresponding target image is located.
[0019] Based on the temporal order among the multiple target images, the multiple first feature data are calculated to obtain target feature data.
[0020] Optionally, obtain the set of images to be tested corresponding to the target device, including:
[0021] Obtain the target video corresponding to the target device, wherein the target video contains the target object;
[0022] In the target video, multiple images with consecutive time intervals are extracted to obtain multiple target images, which together form a set of images to be tested.
[0023] Optional, also includes:
[0024] If the target value is greater than or equal to a preset threshold, an alarm message is output, which is used to indicate that the target device has been damaged.
[0025] An abnormal behavior detection device, the device comprising:
[0026] An image set acquisition unit is used to obtain a set of images to be tested corresponding to a target device. The set of images to be tested contains multiple target images, which are temporally continuous images and contain target objects.
[0027] The prediction unit is used to process the set of images to be tested through a first behavior prediction model to obtain a target value, wherein the target value is the probability that the target object performs an abnormal behavior within the time period to which the multiple target images belong, and the abnormal behavior is an behavior that causes damage to the target device.
[0028] The first behavior prediction model is obtained by training a model based on a first training sample set. The first training sample set contains multiple first training samples, including a first input sample and a first output sample. The first input sample is a set of sample images, and each set of sample images contains multiple first sample images. The multiple first sample images are continuous in time, and each first sample image contains a first sample object. The first output sample is a first value, which represents the probability that the first sample object will perform the abnormal behavior within the time period to which the multiple first sample images belong.
[0029] Optionally, the prediction unit includes:
[0030] The feature value acquisition unit is used to obtain multiple first feature values based on the set of images to be tested, wherein one first feature value corresponds to one target image, and the first feature value is the probability that the target object performs abnormal behavior at the time of the corresponding target image.
[0031] The feature data acquisition unit is used to obtain target feature data based on the set of images to be tested, wherein the target feature data is data related to the action of the target object within the time period to which the multiple target images belong;
[0032] The target value acquisition unit is used to obtain the target value based on the target data and the plurality of first feature values.
[0033] An electronic device includes: a memory and a processor;
[0034] The memory is used to store programs;
[0035] The processor is configured to execute the program to: obtain a set of images to be tested corresponding to a target device, the set of images to be tested containing multiple target images, the multiple target images being temporally continuous images, and the target images containing a target object; process the set of images to be tested using a first behavior prediction model to obtain a target value, the target value being the probability that the target object will perform an abnormal behavior within the time period to which the multiple target images belong, the abnormal behavior being an act that causes damage to the target device; wherein, the first behavior prediction model is obtained by training a model based on a first training sample set, the first training sample set containing multiple first training samples, the first training samples including a first input sample and a first output sample, the first input sample being a set of sample images, each set of sample images containing multiple first sample images, the multiple first sample images being temporally continuous, and the first sample images containing a first sample object, the first output sample being a first value, the first value representing the probability that the first sample object will perform the abnormal behavior within the time period to which the multiple first sample images belong.
[0036] A storage medium storing a computer program, which, when executed by a processor, performs the following: obtaining a set of images to be tested corresponding to a target device, the set of images to be tested containing multiple target images, the multiple target images being temporally continuous and each target image containing a target object; processing the set of images to be tested using a first behavior prediction model to obtain a target value, the target value being the probability that the target object will perform an abnormal behavior within the time period to which the multiple target images belong, the abnormal behavior being an act that causes damage to the target device; wherein the first behavior prediction model is obtained by training a model based on a first training sample set, the first training sample set containing multiple first training samples, the first training samples including first input samples and first output samples, the first input samples being a set of sample images, each set of sample images containing multiple first sample images, the multiple first sample images being temporally continuous and each first sample image containing a first sample object, the first output sample being a first value, the first value representing the probability that the first sample object will perform the abnormal behavior within the time period to which the multiple first sample images belong.
[0037] Using the above technical solution, this application obtains a set of images to be tested corresponding to a target device, which includes multiple target images. The multiple target images are temporally continuous and contain target objects. Then, the set of images to be tested is processed by a first behavior prediction model to obtain a target value. The target value is the probability that the target object will perform abnormal behavior within the time period to which the multiple target images belong. Abnormal behavior is behavior that causes damage to the target device. The first behavior prediction model is obtained by training the model based on a first training sample set. The first training sample set contains multiple first training samples, including a first input sample and a first output sample. The first input sample is a set of sample images, and each set of sample images contains multiple first sample images. The multiple first sample images are temporally continuous and contain first sample objects. The first output sample is a first value, which represents the probability that the first sample object will perform abnormal behavior within the time period to which the multiple first sample images belong. Therefore, this application can obtain the probability of a target object performing abnormal behavior within the time period of multiple target images by processing the set of images to be tested through a pre-trained first behavior prediction model. In other words, it can be understood as recognizing the actions performed by the target object within the time period of multiple target images, and comprehensively recognizing multiple actions within the time period of multiple target images to obtain the probability of the target object performing abnormal behavior within the time period of multiple target images, thereby improving the accuracy of abnormal behavior recognition results. Attached Figure Description
[0038] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0039] Figure 1 A flowchart illustrating the abnormal behavior detection method provided in Embodiment 1 of this application;
[0040] Figure 2 A flowchart illustrating the target value acquisition method provided in Embodiment 1 of this application;
[0041] Figure 3 A flowchart illustrating the target feature data acquisition method provided in Embodiment 1 of this application;
[0042] Figure 4 This is a schematic flowchart of a method for acquiring a set of images to be tested provided in Embodiment 1 of this application;
[0043] Figure 5 This is another flowchart illustrating the abnormal behavior detection method provided in Embodiment 1 of this application;
[0044] Figure 6 This is a schematic diagram of the abnormal behavior detection process for a power self-service terminal device provided in Embodiment 1 of this application;
[0045] Figure 7 This is a schematic diagram of an abnormal behavior detection device provided in Embodiment 2 of this application;
[0046] Figure 8 This is a schematic diagram of the structure of a prediction unit provided in Embodiment 2 of this application;
[0047] Figure 9 This is a schematic diagram of another abnormal behavior detection device provided in Embodiment 2 of this application;
[0048] Figure 10 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of this application. Detailed Implementation
[0049] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0050] This application provides an abnormal behavior detection scheme. It pre-trains a first behavior prediction model to detect images that are temporally continuous, obtaining the probability that an object in the images will perform an action that damages a target device within that time period. Based on this, the application obtains a set of test images containing multiple target images corresponding to the target device. These multiple target images are temporally continuous and contain the target object. The first behavior prediction model then processes the set of test images to obtain the probability that the target object will perform an abnormal behavior within the time period encompassed by the multiple target images. Abnormal behavior refers to actions that damage the target device. This can also be understood as identifying the actions performed by the target object within the time period encompassed by the multiple target images. By comprehensively identifying multiple actions within the time period encompassed by the multiple target images, the probability of the target object performing an abnormal behavior within the time period encompassed by the multiple target images is obtained, thus improving the accuracy of the abnormal behavior detection results.
[0051] The proposed solution can be implemented using an electronic device with data processing capabilities, such as a mobile phone, computer, tablet, local server, cloud server, etc.
[0052] Next, combined Figure 1 As shown, the abnormal behavior detection method provided in Embodiment 1 of this application may include the following steps:
[0053] Step 101: Obtain the set of images to be tested corresponding to the target device.
[0054] The target device is a self-service power terminal, such as an ATM. The set of images to be tested contains multiple target images, which are temporally continuous and contain target objects. Target images can be images related to the target object captured by the camera of the self-service power terminal, such as motion images of the target object, including various postures, expressions, and actions. The target object can be a user's face or a tool. For example, if the target device is an ATM and a user is carrying an axe in front of it, then the target objects are the user's face and the axe, and the target image is an image containing the user's face and / or the axe.
[0055] Step 102: The first line of the prediction model is used to process the set of images to be tested in order to obtain the target value.
[0056] The target value is the probability that the target object will perform abnormal behavior within the time period of multiple target images. Abnormal behavior refers to behavior that damages the target device. Abnormal behavior can also be behavior that damages the target device to achieve a certain purpose, such as damaging an ATM and stealing money from it. Abnormal behavior can also be behavior that may damage the target device, such as a user lingering or loitering in front of an ATM multiple times.
[0057] One issue is that some computers cannot process numbers smaller than a certain threshold, so a very small target value may be converted to zero. This problem can be solved by using the logarithm of the probability. For a very small probability, its logarithm will be negative, but not zero. Therefore, using the logarithm of the target value avoids this situation. The first behavior prediction model is trained based on a first training sample set. The first training sample set contains multiple first training samples, including first input samples and first output samples. The first input samples are a set of sample images, each containing multiple first sample images that are temporally consecutive and contain a first sample object. The first output sample is a first value, representing the probability that the first sample object will perform abnormal behavior within the time period encompassed by the multiple first sample images.
[0058] In other words, this application obtains multiple images containing the first sample object in a continuous time as first sample images, thereby obtaining multiple first sample images, and thus obtaining a set of sample images as the first input sample. It also obtains the probability that the first sample object performs abnormal behavior within the time period to which the multiple first sample images belong as the first output sample. The first input sample and the first output sample constitute the first training sample. Multiple first training samples are obtained, and the model is trained based on the multiple first training samples to obtain a first behavior prediction model. This allows the first behavior prediction model to process the set of images to be tested and obtain the probability that the target object performs abnormal behavior within the time period to which the multiple target images belong.
[0059] The first line of the prediction model can be trained based on positive and negative samples. That is, if the sample image set is a positive sample, the first value is 0%, and if the sample image set is a negative sample, the first value is 100%.
[0060] The first-line prediction model includes the following components:
[0061] (1) Feature Extraction Layer: Deep learning models such as convolutional neural networks (CNN) or recurrent neural networks (RNN) are used to extract features from the input data. In CNN, multiple convolutional layers and pooling layers can be used to extract spatial features; in RNN, models such as Long Short Term Memory (LSTM) or Gate Recurrent Unit (GRU) can be used to extract time series features.
[0062] (2) Fully connected layer: The output of the feature extraction layer is classified through a fully connected layer. The softmax function can be used to convert the output into a probability distribution to determine whether abnormal behavior exists.
[0063] (3) Loss function: The cross-entropy loss function is used to calculate the difference between the model prediction results and the true labels, and the backpropagation algorithm is used to update the model parameters.
[0064] (4) Optimizer: Use optimizers such as Adam and Stochastic Gradient Descent (SGD) to optimize model parameters in order to improve the accuracy and generalization ability of the model.
[0065] (5) Dropout layer: To prevent overfitting, a Dropout layer can be added between fully connected layers to randomly drop some neurons, thereby reducing the complexity of the model.
[0066] (6) Batch Normalization layer: To accelerate the training process, a batch normalization layer is added between the feature extraction layer and the fully connected layer to normalize the input data, so as to better adapt to different data distributions. As can be seen from the above scheme, in the abnormal behavior detection method provided in Embodiment 1 of this application, this application processes the set of images to be tested through a pre-trained first behavior prediction model, so as to obtain the probability that the target object performs abnormal behavior in the time period to which multiple target images belong. It can also be understood as recognizing the actions performed by the target object in the time period to which multiple target images belong, and comprehensively recognizing multiple actions in the time period to which multiple target images belong, so as to obtain the probability that the target object performs abnormal behavior in the time period to which multiple target images belong, thereby improving the accuracy of abnormal behavior recognition results.
[0067] In one implementation, step 102, when processing the set of images to be tested, combines... Figure 2 In the first row of the prediction model, the following steps are performed:
[0068] Step 201: Obtain multiple first feature values based on the set of images to be tested.
[0069] In this system, each first feature value corresponds to a target image, and the first feature value represents the probability that the target object will perform an abnormal behavior at the given time in the target image. In other words, by obtaining the first feature data for each target image, multiple first feature values can be obtained. Taking an axe as an example, after obtaining the set of images to be tested, the probability that the axe will perform an abnormal behavior at the given time in each target image is obtained. This abnormal behavior could be destroying an ATM, so the probability that the axe will destroy an ATM at the given time in each target image is obtained.
[0070] Step 202: Obtain target feature data based on the set of images to be tested.
[0071] Among them, target feature data refers to data associated with the action of the target object within the time period to which multiple target images belong, such as the speed of the target object and the acceleration of the target object.
[0072] Step 203: Obtain the target value based on the target data and multiple first feature values.
[0073] Specifically, multiple first feature values and target data of the image set to be tested are input into the algorithm of the first behavior prediction model for calculation. For example, multiple first feature values and target data can be input into Euclidean distance for calculation.
[0074] In this embodiment of the application, multiple first feature values and target feature data are obtained based on the set of images to be tested, and then the target value can be obtained based on the multiple first feature values and target feature data.
[0075] In one implementation, step 201, when obtaining multiple first feature values, can be achieved in the following way:
[0076] The second-line prediction model processes each target image to obtain the first feature values corresponding to multiple target images.
[0077] In this context, each target image corresponds to a first feature value.
[0078] Specifically, each target image is processed by the second-line prediction model to obtain a first feature value corresponding to each target image. The first feature values corresponding to all target images are combined to form multiple first feature values.
[0079] The second behavior prediction model is obtained by training the model based on the second training sample set. The second training sample set contains multiple second training samples, including a second input sample and a second output sample. The second input sample is a second sample image, which contains a second sample object. The second output sample is a second value, which is the probability that the second sample object will perform abnormal behavior at the time of the second sample image.
[0080] In other words, this application obtains an image containing a second sample object as a second input sample, and obtains the probability that the second sample object performs abnormal behavior at the time of the second sample image as a second output sample. The second input sample and the second output sample form a second training sample. Multiple second training samples are obtained, and the model is trained based on the multiple second training samples to obtain a second behavior prediction model, so that by processing each target image through the second behavior prediction model, the first feature value corresponding to each target image can be obtained.
[0081] In this embodiment of the application, each target image is processed by a pre-trained second behavior prediction model to obtain a first feature value corresponding to each target image. After processing all target images, multiple first feature values can be obtained.
[0082] In one implementation, step 202, when obtaining the target feature data, combines... Figure 3 This may include the following steps:
[0083] Step 301: Extract features from multiple target images in the image set to be tested to obtain multiple first feature data.
[0084] Specifically, multiple target images can be distinguished by foreground and background to obtain the region where the target object is located, and then feature extraction can be performed in the region where the target object is located.
[0085] In this context, each target image corresponds to a first feature data. The first feature data is data associated with the action of the target object at the time of the corresponding target image. Specifically, the first feature data may include feature data such as motion trajectory feature data and spatiotemporal feature data. For example, the position of the target object in the first target image and the position in the second target image.
[0086] Step 302: Calculate multiple first feature data according to the time sequence between multiple target images to obtain target feature data.
[0087] For example, based on the position of the target object in the first target image and its position in the second target image, the displacement of the target object can be calculated. By combining the time difference between the first and second target images, the velocity, acceleration, and other feature data of the target object can be calculated.
[0088] In this embodiment of the application, feature extraction is performed on multiple target images in the image set to be tested to obtain data related to the action of the target object at the time of the corresponding target image, which is the first feature data. Then, based on multiple first feature data and the time order between multiple target images, target feature data can be obtained.
[0089] In one implementation, step 101, when obtaining the set of images to be tested corresponding to the target device, combines... Figure 4 Specifically, it may include the following steps:
[0090] Step 401: Obtain the target video corresponding to the target device.
[0091] The target video includes the target object. Taking the target object as the user's face and the target device as the power self-service terminal as an example, the target video can be a video containing the user's face obtained by the camera of the power self-service terminal.
[0092] Step 402: Extract multiple images from the target video that are consecutive in time to obtain multiple target images.
[0093] The set of images to be tested consists of multiple target images.
[0094] For example, by capturing two images per second from a video, multiple target images can be obtained, thus creating a set of images to be tested.
[0095] In this embodiment of the application, a target video corresponding to the target device is obtained, and multiple images with consecutive time are extracted from the target video to obtain multiple target images. The multiple target images form a set of images to be tested.
[0096] In one implementation, after processing the set of images to be tested in step 102, the following steps are combined: Figure 5 The embodiments of this application may further include the following steps:
[0097] Step 103: Determine whether the target value is greater than or equal to the preset threshold. If the target value is greater than or equal to the preset threshold, proceed to step 104. If the target value is less than the preset threshold, no action is taken.
[0098] Specifically, a target value greater than or equal to a preset threshold indicates that the behavior performed by the target object in the image set under test is abnormal, while a target value less than the preset threshold indicates that the behavior performed by the target object in the image set under test is normal. Taking a preset threshold of 95% as an example, if the target value is greater than or equal to 95%, it indicates that the behavior performed by the target object in the image set under test is abnormal, and if the target value is less than 95%, it indicates that the behavior performed by the target object in the image set under test is normal.
[0099] Step 104: Output alarm information.
[0100] The alarm message is used to alert the target device that it has been damaged; for example, it outputs the voice message "The target device is being damaged."
[0101] In other words, if the target value is greater than or equal to the preset threshold, and the behavior performed by the target object is abnormal, an alarm message needs to be output to indicate that the target device has been damaged.
[0102] In this embodiment of the application, when the target value is greater than or equal to a preset threshold, an alarm message is output to indicate that the target device has been damaged, so that staff can handle the target object.
[0103] In specific implementation, combined with Figure 6 The diagram shown is a schematic of the abnormal behavior detection process for the power self-service terminal equipment provided in this application. This solution specifically includes the following parts:
[0104] 1. Image Acquisition
[0105] This section corresponds to steps 401 and 402 above. It obtains the target video containing the target object captured by the camera of the power self-service terminal equipment. Using libraries such as OpenCV, it processes the images in the target video, extracting keyframes, etc. The images corresponding to the keyframes are used as target images, resulting in a set of multiple target images to be tested. 2. Abnormal Behavior Detection
[0106] The content implemented in this part based on the model corresponds to step 102 above. The first behavior prediction model is used to process the set of images to be tested to obtain the probability that the target object will perform abnormal behavior within the time period to which the set of images to be tested belongs.
[0107] Specifically, it can include the following parts:
[0108] (1) Behavioral representation of foreground and moving target detection and feature extraction
[0109] This section is equivalent to step 301 above, distinguishing between the target object and the foreground / background, and extracting the target object's primary feature data from the target object's region. This includes features such as the target object's motion trajectory, skeletal information, appearance features, facial features, spatiotemporal features, and overall appearance and motion characteristics. Appearance features primarily refer to static features, such as facial expressions and occlusion. Motion features refer to the skeletal posture of the tested object. Spatiotemporal features refer to the changes in the tested object over time, such as body displacement.
[0110] (2) Abnormal behavior identification and detection
[0111] Based on multiple first feature data, behavior recognition is performed on each target image to obtain a first feature value, equivalent to step 201. Based on the first feature data of multiple target images and the temporal order among multiple target images, target feature data of the set of images to be tested can be obtained, equivalent to step 302. The target data is equivalent to the behavior sequence in the figure. Based on the first feature value and the target data, it is determined whether it belongs to abnormal behavior features. At the same time, this application can also model the normal behavior pattern and compare it with the normal behavior pattern to determine whether the difference from normal behavior is too large. If the difference from normal behavior is not large and it does not belong to abnormal behavior features, there is no abnormal behavior (i.e., normal in the figure) in the multiple target images in the image set to be tested. If it belongs to abnormal behavior features, the specific abnormal behavior can be further identified. If the difference from normal behavior is too large, an abnormal alarm is issued, which is equivalent to step 103. If the difference from normal behavior is not large and it does not belong to abnormal behavior features, it is equivalent to the probability that the target object performs abnormal behavior within the time period to which the multiple target images in the image set to be tested belong is less than a preset threshold. If it belongs to abnormal behavior features and / or the difference from normal behavior is too large, it is equivalent to the probability that the target object performs abnormal behavior within the time period to which the multiple target images in the image set to be tested belong is greater than or equal to the preset threshold. The abnormal alarm is equivalent to the output alarm information in step 104.
[0112] The abnormal behavior detection device provided in Embodiment 2 of this application is described below. The abnormal behavior detection device described below can be referred to in correspondence with the abnormal behavior detection method described above.
[0113] See Figure 7 , Figure 7 This is a schematic diagram of an abnormal behavior detection device disclosed in Embodiment 2 of this application.
[0114] like Figure 7 As shown, the device may include:
[0115] The image set acquisition unit 701 is used to obtain the image set to be tested corresponding to the target device. The image set to be tested contains multiple target images, which are time-continuous images and contain target objects.
[0116] The prediction unit 702 is used to process the set of images to be tested through the first behavior prediction model to obtain a target value. The target value is the probability that the target object will perform abnormal behavior within the time period of multiple target images. The abnormal behavior is the behavior that causes damage to the target device.
[0117] The first behavior prediction model is obtained by training the model based on the first training sample set. The first training sample set contains multiple first training samples, including first input samples and first output samples. The first input samples are sample image sets, each sample image set contains multiple first sample images, the multiple first sample images are continuous in time, and the first sample images contain first sample objects. The first output sample is a first value, which represents the probability that the first sample object performs abnormal behavior within the time period to which the multiple first sample images belong.
[0118] As can be seen from the above scheme, the abnormal behavior detection device provided in Embodiment 2 of this application processes the set of images to be tested through a pre-trained first behavior prediction model, and can obtain the probability that the target object performs abnormal behavior within the time period to which multiple target images belong. It can also be understood as recognizing the actions performed by the target object within the time period to which multiple target images belong, and comprehensively recognizing multiple actions within the time period to which multiple target images belong, so as to obtain the probability that the target object performs abnormal behavior within the time period to which multiple target images belong, thereby improving the accuracy of the abnormal behavior recognition result.
[0119] In one implementation, combining Figure 8 The prediction unit 702 provided in this application includes:
[0120] The feature value acquisition unit 801 is used to obtain multiple first feature values based on the set of images to be tested. Each first feature value corresponds to a target image, and the first feature value is the probability that the target object will perform abnormal behavior at the time of the corresponding target image.
[0121] The feature data acquisition unit 802 is used to obtain target feature data based on the set of images to be tested. The target feature data is data related to the action of the target object within the time period to which multiple target images belong.
[0122] The target value acquisition unit 803 is used to obtain the target value based on the target data and multiple first feature values.
[0123] In one implementation, the feature value acquisition unit 801 is further configured to process each target image through a second behavior prediction model to obtain a first feature value corresponding to multiple target images, wherein one target image corresponds to one first feature value; wherein the second behavior prediction model is obtained by training the model based on a second training sample set, the second training sample set contains multiple second training samples, the second training samples include a second input sample and a second output sample, the second input sample is a second sample image, the second sample image contains a second sample object, the second output sample is a second value, and the second value is the probability that the second sample object performs abnormal behavior at the time when the second sample image is located.
[0124] In one implementation, the target value acquisition unit 803 is further used to extract features from multiple target images in the set of images to be tested to obtain multiple first feature data. One target image corresponds to one first feature data. The first feature data is data associated with the action of the target object at the time when the corresponding target image is located. According to the time order between multiple target images, the multiple first feature data are calculated to obtain target feature data.
[0125] In one implementation, the image set acquisition unit 701 is further used to acquire a target video corresponding to the target device, the target video containing a target object; in the target video, multiple images with consecutive time are extracted to obtain multiple target images, and the multiple target images form a set of images to be tested.
[0126] In one implementation, such as Figure 9 The schematic diagram of the device shown includes:
[0127] The judgment unit 703 is used to output alarm information when the target value is greater than or equal to a preset threshold. The alarm information is used to indicate that the target device has been damaged.
[0128] The abnormal behavior detection device provided in this application embodiment can be applied to electronic devices, such as mobile phones, tablets, computers, local servers, cloud servers, etc. Optionally, Figure 10 This diagram illustrates a hardware structure block diagram of an electronic device according to Embodiment 3 of this application. (Refer to...) Figure 10 The hardware structure of the electronic device may include: at least one processor 1001, at least one communication interface 1002, at least one memory 1003 and at least one communication bus 1004.
[0129] In this embodiment of the application, the number of processor 1001, communication interface 1002, memory 1003 and communication bus 1004 is at least one, and processor 1001, communication interface 1002 and memory 1003 communicate with each other through communication bus 1004.
[0130] The processor 1001 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0131] The memory 1003 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;
[0132] The memory stores programs, and the processor can call these programs. The programs are used for:
[0133] Obtain the set of images to be tested corresponding to the target device. The set of images to be tested contains multiple target images, which are time-continuous and contain target objects.
[0134] The first behavior prediction model is used to process the set of images to be tested to obtain the target value. The target value is the probability that the target object will perform abnormal behavior within the time period of multiple target images. Abnormal behavior is behavior that causes damage to the target device.
[0135] The first behavior prediction model is obtained by training the model based on the first training sample set. The first training sample set contains multiple first training samples, including first input samples and first output samples. The first input samples are sample image sets, each sample image set contains multiple first sample images, the multiple first sample images are continuous in time, and the first sample images contain first sample objects. The first output sample is a first value, which represents the probability that the first sample object performs abnormal behavior within the time period to which the multiple first sample images belong.
[0136] Optionally, the program's refined and extended functions can be found in the description above.
[0137] Embodiment 4 of this application also provides a storage medium that can store a program suitable for execution by a processor, the program being used for:
[0138] Obtain the set of images to be tested corresponding to the target device. The set of images to be tested contains multiple target images, which are time-continuous and contain target objects.
[0139] The first behavior prediction model is used to process the set of images to be tested to obtain the target value. The target value is the probability that the target object will perform abnormal behavior within the time period of multiple target images. Abnormal behavior is behavior that causes damage to the target device.
[0140] The first behavior prediction model is obtained by training the model based on the first training sample set. The first training sample set contains multiple first training samples, including first input samples and first output samples. The first input samples are sample image sets, each sample image set contains multiple first sample images, the multiple first sample images are continuous in time, and the first sample images contain first sample objects. The first output sample is a first value, which represents the probability that the first sample object performs abnormal behavior within the time period to which the multiple first sample images belong.
[0141] Optionally, the program's refined and extended functions can be found in the description above.
[0142] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0143] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
[0144] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for detecting abnormal behavior, characterized in that, The method includes: Obtain a set of images to be tested corresponding to the target device. The set of images to be tested contains multiple target images, which are temporally continuous and contain target objects. The second behavior prediction model processes each target image to obtain multiple first feature values corresponding to the target images. Each target image corresponds to one first feature value, which represents the probability that the target object will perform abnormal behavior at the time specified in the corresponding target image. The model distinguishes between foreground and background in the multiple target images in the test image set, and extracts features from the region where the target object is located to obtain multiple first feature data. Each target image corresponds to one first feature data, which is data associated with the target object's action at the time specified in the corresponding target image. Based on the temporal order of the multiple target images, the multiple first feature data are calculated to obtain target feature data. The target feature data includes the velocity and acceleration of the target object within the time period of the multiple target images. The target feature data is then... The data and the multiple first feature values are fed into a first behavior prediction model for calculation to obtain a target value. Based on the first feature values and the target feature data, it is determined whether it belongs to an abnormal behavior feature. At the same time, by modeling a normal behavior pattern and comparing it with the normal behavior pattern, it is determined whether the difference from normal behavior is too large. If it belongs to an abnormal behavior feature and / or the difference from normal behavior is too large, it is determined that the target value is greater than or equal to a preset threshold. If the target value is greater than or equal to the preset threshold, an alarm message is output, which is used to indicate that the target device has been damaged. If the difference from normal behavior is not large and it does not belong to an abnormal behavior feature, it is determined that there is no abnormal behavior in the multiple target images in the set of images to be tested. The target value is the probability that the target object performs abnormal behavior within the time period to which the multiple target images belong, and the abnormal behavior is the behavior that causes damage to the target device. The first behavior prediction model is obtained by training a model based on a first training sample set. The first training sample set contains multiple first training samples, including a first input sample and a first output sample. The first input sample is a set of sample images, and each set of sample images contains multiple first sample images. The multiple first sample images are continuous in time, and each first sample image contains a first sample object. The first output sample is a first value, which represents the probability that the first sample object will perform the abnormal behavior within the time period to which the multiple first sample images belong.
2. The abnormal behavior detection method according to claim 1, characterized in that, The second behavior prediction model is obtained by training the model based on a second training sample set. The second training sample set contains multiple second training samples, including a second input sample and a second output sample. The second input sample is a second sample image, which contains a second sample object. The second output sample is a second value, which is the probability that the second sample object will perform the abnormal behavior at the time when the second sample image is in the second sample image.
3. The abnormal behavior detection method according to claim 1, characterized in that, Obtain the set of images to be tested corresponding to the target device, including: Obtain the target video corresponding to the target device, wherein the target video contains the target object; In the target video, multiple images with consecutive time intervals are extracted to obtain multiple target images, which together form a set of images to be tested.
4. An abnormal behavior detection device, characterized in that, The device includes: An image set acquisition unit is used to obtain a set of images to be tested corresponding to a target device. The set of images to be tested contains multiple target images, which are temporally continuous images and contain target objects. The prediction unit is used to process the set of images to be tested through a first behavior prediction model to obtain a target value, wherein the target value is the probability that the target object performs an abnormal behavior within the time period to which the multiple target images belong, and the abnormal behavior is an behavior that causes damage to the target device. The first behavior prediction model is obtained by training a model based on a first training sample set. The first training sample set contains multiple first training samples, including a first input sample and a first output sample. The first input sample is a set of sample images, and each set of sample images contains multiple first sample images. The multiple first sample images are continuous in time, and each first sample image contains a first sample object. The first output sample is a first value, which represents the probability that the first sample object performs the abnormal behavior within the time period to which the multiple first sample images belong. The prediction unit includes: The feature value acquisition unit is used to process each of the target images through the second behavior prediction model to obtain the first feature values corresponding to multiple target images. One target image corresponds to one first feature value, and the first feature value is the probability that the target object performs abnormal behavior at the time of the corresponding target image. The feature data acquisition unit is used to distinguish the foreground and background of the plurality of target images in the set of images to be tested, and to extract features in the region where the target object is located to obtain a plurality of first feature data. Each target image corresponds to one first feature data, and the first feature data is data associated with the action of the target object at the time of the corresponding target image. The unit calculates the plurality of first feature data according to the time order between the plurality of target images to obtain target feature data. The target feature data includes the velocity and acceleration of the target object within the time period to which the plurality of target images belong. The target value acquisition unit is used to input the target feature data and the plurality of first feature values into the first behavior prediction model for calculation to obtain the target value; The prediction unit is further configured to: determine whether the feature belongs to abnormal behavior based on the first feature value and the target feature data; simultaneously, by modeling a normal behavior pattern and comparing it with the normal behavior pattern, determine whether the difference from normal behavior is too large; if the feature belongs to abnormal behavior and / or the difference from normal behavior is too large, determine that the target value is greater than or equal to a preset threshold; if the target value is greater than or equal to the preset threshold, output an alarm message, the alarm message being used to indicate that the target device has been damaged; if the difference from normal behavior is not too large and the feature does not belong to abnormal behavior, determine that there is no abnormal behavior in the plurality of target images in the set of images to be tested.
5. An electronic device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is configured to execute the program to: obtain a set of images to be tested corresponding to the target device, the set of images to be tested containing multiple target images, the multiple target images being temporally consecutive images, and the target images containing a target object; process each target image using a second behavior prediction model to obtain first feature values corresponding to multiple target images, one target image corresponding to one first feature value, the first feature value being the probability that the target object performs abnormal behavior at the time of the corresponding target image; distinguish foreground and background in the multiple target images in the set of images to be tested; extract features in the area where the target object is located to obtain multiple first feature data, one target image corresponding to one first feature data, the first feature data being data associated with the action of the target object at the time of the corresponding target image; and calculate the multiple first feature data according to the temporal order between the multiple target images to obtain target feature data. The target feature data includes the velocity and acceleration of the target object within the time period of the multiple target images; the target feature data and the multiple first feature values are input into the first behavior prediction model for calculation to obtain the target value; based on the first feature value and the target feature data, it is determined whether it belongs to abnormal behavior features; at the same time, by modeling the normal behavior pattern and comparing it with the normal behavior pattern, it is determined whether the difference from the normal behavior is too large. If the behavior is abnormal and / or differs significantly from normal behavior, the target value is determined to be greater than or equal to a preset threshold. If the target value is greater than or equal to the preset threshold, an alarm message is output to indicate that the target device has been damaged. If the behavior is not significantly different from normal behavior and does not belong to abnormal behavior characteristics, it is determined that there is no abnormal behavior in the multiple target images in the set of images to be tested; the target value is the probability that the target object performs abnormal behavior within the time period to which the multiple target images belong, and the abnormal behavior is the behavior that causes damage to the target device; wherein, the first behavior prediction model is obtained by training the model based on a first training sample set, the first training sample set contains multiple first training samples, the first training samples include a first input sample and a first output sample, the first input sample is a set of sample images, each set of sample images contains multiple first sample images, the multiple first sample images are continuous in time, and the first sample images contain a first sample object, the first output sample is a first value, and the first value represents the probability that the first sample object performs the abnormal behavior within the time period to which the multiple first sample images belong.
6. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it performs the following: obtaining a set of images to be tested corresponding to the target device, the set of images to be tested containing multiple target images, the multiple target images being temporally consecutive images, and the target images containing a target object; processing each target image using a second behavior prediction model to obtain first feature values corresponding to multiple target images, one target image corresponding to one first feature value, the first feature value being the probability that the target object performs abnormal behavior at the time of the corresponding target image; distinguishing foreground and background in the multiple target images in the set of images to be tested; extracting features in the area where the target object is located to obtain multiple first feature data, one target image corresponding to one first feature data, the first feature data being data associated with the action of the target object at the time of the corresponding target image; and calculating the multiple first feature data according to the temporal order between the multiple target images to obtain target feature data. The target feature data includes the velocity and acceleration of the target object within the time period of the multiple target images; the target feature data and the multiple first feature values are input into the first behavior prediction model for calculation to obtain the target value; based on the first feature value and the target feature data, it is determined whether it belongs to abnormal behavior features; at the same time, by modeling the normal behavior pattern and comparing it with the normal behavior pattern, it is determined whether the difference from the normal behavior is too large. If the behavior is abnormal and / or differs significantly from normal behavior, the target value is determined to be greater than or equal to a preset threshold. If the target value is greater than or equal to the preset threshold, an alarm message is output to indicate that the target device has been damaged. If the behavior is not significantly different from normal behavior and does not belong to abnormal behavior characteristics, it is determined that there is no abnormal behavior in the multiple target images in the set of images to be tested; the target value is the probability that the target object performs abnormal behavior within the time period to which the multiple target images belong, and the abnormal behavior is the behavior that causes damage to the target device; wherein, the first behavior prediction model is obtained by training the model based on a first training sample set, the first training sample set contains multiple first training samples, the first training samples include a first input sample and a first output sample, the first input sample is a set of sample images, each set of sample images contains multiple first sample images, the multiple first sample images are continuous in time, and the first sample images contain a first sample object, the first output sample is a first value, and the first value represents the probability that the first sample object performs the abnormal behavior within the time period to which the multiple first sample images belong.
Citation Information
Patent Citations
Method and device for processing action behaviors in video
CN110096938A
Behavior analysis method, equipment and device
CN111401296A
Abnormal event detection method and device and model training method and device
CN112528801A