A violence detection method and system based on edge computing

By conducting prospect detection and keyframe screening on the monitoring device end, and combining the deep learning model on the edge server for end-to-end inference, the problem of computing resources and network bandwidth occupation in the existing technology is solved, and efficient brute-force behavior detection is achieved.

CN115346150BActive Publication Date: 2025-06-27INNER MONGOLIA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210845310.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-06-27
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

The existing technology has problems with computing resources and network bandwidth usage in violent behavior detection, which makes it difficult to widely deploy deep learning methods in monitoring terminals, and the cloud-based data summary method is poor in economics.

Method used

The brute-action detection method based on edge computing is adopted to reduce redundant information transmission by performing prospect detection and keyframe screening on the monitoring device side, and call deep learning models on the edge server for end-to-end inference.

Benefits of technology

It effectively reduces the computing resource consumption and network bandwidth usage in the brute-force behavior detection process, improves detection accuracy, and reduces server load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115346150B_ABST
    Figure CN115346150B_ABST
Patent Text Reader

Abstract

A violent behavior detection method based on edge computing, constructing and training a deep learning model and a reinforcement learning method for violent behavior detection on a cloud server; a monitoring device performs foreground detection on a video frame, obtains a region-of-interest frame and uploads it to an edge server, and the edge server performs object detection, and feeds back the result of the region where people exist in the frame to the monitoring device; the monitoring device judges whether the number of people in the region where people exist exceeds a threshold, establishes a video frame buffer and calls a reinforcement learning method to screen key frames for the video frames, stores the key frames in the buffer, and if the buffer is full, uploads the video frames in the buffer as a group to the edge server, and the edge server calls the deep learning model to perform end-to-end inference on the group of video frames to obtain the probability of the existence of violent behavior in the group of video frames; the present invention can effectively reduce the consumption of computing resources and the occupation of network bandwidth in the entire process of violent behavior detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of public security monitoring, and particularly relates to a violent behavior detection method and system based on edge computing. Background Art

[0002] Monitoring violent behavior through video surveillance is one of its important values. When a violent behavior occurs, the parties involved usually cannot call the police immediately when facing a strong external impact. The method of manual duty monitoring is also difficult to process a large amount of data all-weather and without dead angles. Transmitting video data to a computing unit and using computer algorithms to detect it in real time and issue a warning to the security forces in the relevant area is a better solution.

[0003] In the prior art, the detection of violent behavior is mostly limited to the innovation of the detection method itself, but there are many problems in its actual implementation and deployment.

[0004] Common deployment schemes include direct deployment at the terminal and cloud data aggregation. For direct deployment at the terminal, limited by computing resources and manufacturing costs, it is difficult to widely deploy deep learning methods with high accuracy in existing monitoring terminals. Cloud data aggregation is to deploy the algorithm in the cloud to receive all video data frame by frame, but this causes too much unnecessary load on the backbone network and cloud servers. Since violent behavior is an occasional event, this method has poor economy. Summary of the Invention

[0005] In order to overcome the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a violent behavior detection method and system based on edge computing, which can effectively filter redundant video information at the device end, ensure the detection accuracy, and effectively reduce the network load and server load at the same time.

[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0007] A violent behavior detection method based on edge computing includes the following steps:

[0008] Step 1: Construct and train a deep learning model for violent behavior detection on a cloud server, and construct and train a reinforcement learning method; the single input of the deep learning model is a group of video frames, and the output is the probability of the existence of violent behavior; the input of the reinforcement learning method is frame-by-frame video data. After selecting a group of video frames, the selected video frames are input into the deep learning model, and the parameters in the reinforcement learning method are iteratively updated according to a preset reward rule;

[0009] Step 2: The monitoring device end receives video data and reads the video frames in the video data in real time;

[0010] Step 3: The monitoring device uses a foreground detection algorithm to perform foreground detection on the video frame. It makes a judgment based on the foreground area characteristics. If the preset conditions are met, it further calculates the region of interest and crops the frame to obtain the region of interest frame, then proceeds to Step 4; if not, it repeats Step 3.

[0011] Step 4: Upload the region of interest frame to the edge server. The edge server uses an object detection algorithm to perform object detection and feeds back the result of the area where people exist in the frame to the monitoring device.

[0012] Step 5: The monitoring device uses the result of the area where people exist to correct the relevant parameters of the foreground detection algorithm, and determines whether the number of people in the area where people exist exceeds the threshold. If it exceeds, it proceeds to Step 6; otherwise, it returns to Step 3.

[0013] Step 6: Establish a video frame buffer with a maximum capacity of a fixed number of frames on the monitoring device and call a reinforcement learning method to screen key frames from the video frames, and store the key frames in the buffer.

[0014] Step 7: Determine the latency of the video frames in the buffer. If the latency is greater than the set threshold, discard the earliest video frame that entered the buffer. If the number of video frames in the buffer is equal to the maximum capacity of the buffer, i.e., the buffer is full, upload the video frames in the buffer as a group to the edge server to execute Step 8; then, discard a set proportion of video frames in the order of the time they were stored in the buffer. Repeat Steps 6 and 7 when the buffer is not full; when the duration of the non-full state of the buffer reaches the threshold, return to Step 3, and each time the buffer is full, start recording the duration again.

[0015] Step 8: The edge server calls a deep learning model to perform end-to-end inference on this group of video frames to obtain the probability of violent behavior existing in this group of video frames.

[0016] Step 9: Issue a warning level, the related video frame, and the location of the monitoring device according to the probability value.

[0017] In one embodiment, the deep learning model is a long short-term memory convolutional neural network, the reinforcement learning method is the Q-learning method, the foreground detection algorithm is the Vibe algorithm, and the object detection algorithm is the Yolo algorithm. Other mature networks and algorithms are also applicable to the present invention.

[0018] In one embodiment, the preset condition means that there is a connected region with an area larger than the preset threshold in the foreground of the frame, and the threshold is selected as the minimum value of the area of the frame region where humans can be normally recognized in the environment where the monitoring device is located.

[0019] In one embodiment, in step 5, using the result of the occupied area, compare it with the result of the foreground detection algorithm, update the misdetected foreground in the foreground detection algorithm to the background, and at the same time, use the complementary filtering algorithm to update the foreground connectivity area threshold with the minimum value in each area area.

[0020] In one embodiment, in step 6, the method for screening key frames from video frames by the reinforcement learning method is as follows:

[0021] Step 61: Calculate the inter-frame difference between the frame to be screened and the frame that finally enters the buffer as the state input of the reinforcement learning method;

[0022] Step 62: Using the state, query the Q-value table to obtain the action with the maximum expected return value, that is, the action value with the maximum return. The action value is 1 or 0. 1 represents selecting the current frame to be selected as the key frame, and 0 represents discarding the current frame to be selected. The Q-value table is obtained by reinforcement learning training;

[0023] Step 63: Execute the screening action according to the action value and retain the key frame.

[0024] In one embodiment, in step 7, calculate the average distance between the generation time of each video frame in the buffer and the current time. When this distance is greater than the hysteresis threshold, it is considered that the data is too lagged.

[0025] In one embodiment, step 8 further includes the following sub-steps:

[0026] Step 81: Use a convolutional neural network to extract features from a single-frame image to obtain a set of feature sets;

[0027] Step 82: Use a long short-term memory network to perform spatio-temporal feature extraction and inference on this set of features and obtain the final result.

[0028] The present invention also provides a violent behavior detection system based on edge computing, including a training subsystem, a pre-detection subsystem, and an edge computing subsystem;

[0029] The training subsystem is deployed on a cloud server and includes a data set construction module, a data set preprocessing module, a detection model training module, and a reinforcement learning training module; the data set construction module converts the video data set with violent labels into a standard form available for training; the data set preprocessing module preprocesses the video data set and constructs a video frame data set with violent labels for training the deep learning module and a video frame data set with frame importance labels for training the reinforcement learning method respectively; the detection model training module inputs the data into the deep learning model and obtains an end-to-end model available for inference through iterative training; the reinforcement learning training module updates its own parameters based on the feedback of the detection model inference result to obtain a model available for frame screening;

[0030] The pre-detection subsystem is deployed at the monitoring device end and includes a foreground detection module, a key frame screening module, and a computing offloading module; the foreground detection module performs foreground detection on the video frame to obtain the picture of the region of interest. The foreground detection module is the longest-running module and only wakes up other modules of the pre-detection subsystem to execute when an effective foreground is obtained; the key frame screening module extracts key information from the video information to reduce the number of times of waking up the edge computing system; the computing offloading module is responsible for offloading the screened video frames to the edge computing subsystem after the key frame screening module meets the preset screening conditions.

[0031] The edge computing subsystem is deployed on the edge computing server and includes a video preprocessing module, an object detection module, a violence detection module, and a warning prompt module; the video preprocessing module preprocesses the video frames unloaded to the edge computing subsystem to achieve the standardization of data input; the object detection module assists and enhances the foreground detection module in the pre-detection subsystem, identifies the picture sent by it, and returns the position information of the people in the picture; the violence detection module performs inference calculation on the input video frames to obtain the possibility of violence occurrence; the warning prompt module infers the warning level of violence occurrence based on the result of the violence detection module and displays the relevant video frames for the user to view.

[0032] Compared with the prior art, the beneficial effects of the present invention are: effectively reducing the consumption of computing resources and the occupancy of network bandwidth in the whole process of violence behavior detection. Brief Description of the Drawings

[0033] Figure 1 is the flowchart of the violence behavior detection method based on edge computing according to the embodiment of the present invention.

[0034] Figure 2 is the framework diagram of the violence behavior detection system based on edge computing according to the embodiment of the present invention. Detailed Embodiment

[0035] The following will describe the embodiments of the present invention in detail with reference to the drawings and embodiments.

[0036] As Figure 1 shown, the violence behavior detection method based on edge computing according to the embodiment of the present invention includes steps 1-step 9.

[0037] Step 1: Construct and train a deep learning model for violence behavior detection, and on the basis of the availability of the deep learning model, construct and train a reinforcement learning method.

[0038] In the present invention, both the deep learning model and the reinforcement learning method are constructed and trained on the cloud server. Among them, the deep learning model can be a conventional model such as a long short-term memory convolutional neural network, whose single input is a group of video frames, and the output is the probability of the presence of violent behavior. The data set used for training can be a public data set such as Hockey Fight, Movies and RWF-2000. The reinforcement learning method can be a Q learning method, whose input is a continuous video frame, and a group of key video frames are selected and input into the deep learning model, and the parameters in the reinforcement learning method are iteratively updated according to the preset reward rules. The training rule of this embodiment is SARSA, and the reward rule can be based on the comparison of the test results on the positive sample with the results obtained by the fixed interval screening scheme. If the accuracy of the result after the reinforcement learning screening exceeds the result, a positive reward is given, otherwise a negative reward is given. The fixed interval of this embodiment is 3.

[0039] Step 2: The monitoring device receives the video data and reads each video frame in the video data in real time. In this embodiment, the video screen size collected by the monitoring device can be 1920*1080, the typical frame rate is 25Fps, and the video encoding format is H265.

[0040] Step 3: The monitoring device uses a foreground detection algorithm to detect the foreground of the video image and makes a judgment based on the characteristics of the foreground area. If it meets the preset conditions, it further calculates the area of ​​interest and crops the image to obtain the area of ​​interest, and then proceeds to step 4; if it does not meet the conditions, repeat step 3.

[0041] The foreground detection algorithm of this embodiment can adopt mature algorithms such as the Vibe algorithm, or its optimized variant algorithm. Specifically, the area of ​​the foreground connected area is counted, the area size of the connected area is sorted, and it is determined whether the largest connected area exceeds the threshold T Area If the threshold value T is exceeded Area , and obtain the maximum external matrix of all connected areas as the area of ​​interest. Area It is the minimum area of ​​the screen where the monitoring equipment can normally identify humans.

[0042] Step 4: Upload the cropped image of the area of ​​interest to the edge server, and use the target detection algorithm to perform target detection, and feed back the result of the area with people in the image to the monitoring device.

[0043] This embodiment uses the Yolo algorithm for target detection. The algorithm model uses a publicly downloadable model trained with public data such as Coco. Other end-to-end target detection algorithms can also be used. Target detection is used to obtain information about the presence of people in the image, including the position parameters of each person in the image: x, y, w, h. The row coordinates, column coordinates, rectangle width, and rectangle height of the upper left vertex of the circumscribed rectangle of the person are represented in turn. An unordered sequence {[x1, y1, w1, h1], [x2, y2, w2, h2], …, [x n ,y n ,w n ,h n ]}where n is the total number of people, and the sequence is returned to the monitoring device.

[0044] Step 5: The monitoring device uses the results of the manned area to correct the relevant parameters of the foreground detection algorithm to determine whether the number of people in the manned area exceeds the threshold. If so, proceed to step 6, otherwise return to step 3.

[0045] At the same time, this step can also use the results of the human area to compare with the results of the foreground detection algorithm, update the misdetected foreground in the foreground detection algorithm to the background, and use the complementary filtering algorithm to update the foreground connected area threshold with the minimum value of the area of ​​each area.

[0046] Specifically, the monitoring device variable sequence is obtained as w i *h i Minimum Area min , and use this value to update T Area , that is, update the threshold value of the area of ​​the connected region that can be filtered. Calculate the non-intersection area between the area contained in the sequence and the area obtained in step 3, and set the pixels of this part of the area as background pixels to quickly eliminate the ghost problem in Vibe. Count the number of elements n in the sequence. If n is greater than 1, go to step 6, otherwise return to step 3.

[0047] Step 6: Establish a video frame buffer with a maximum capacity of a fixed number of frames on the monitoring device, filter the key frames of the video frames through the reinforcement learning method, and store the key frames in the buffer.

[0048] Specifically, in this step, a size of S is initialized buffer The buffer is used to save the filtered video frames, S buffer The value of is equal to the number of video frames required for a single detection by the edge server. In this embodiment, the value is 24. Then, the video data is read frame by frame, and the video frames are screened using reinforcement learning. The selected frames are placed in the buffer. The method is:

[0049] Step 61: Calculate the inter-frame difference between the frame to be screened and the last frame that entered the buffer, and use it as the state input for the reinforcement learning method. The calculation method of the inter-frame difference can be to divide the picture into a grid of 16*16, calculate the pixel transformation ratio for each grid through the frame difference method, and use the 16*16 difference matrix as the state input for the reinforcement learning method.

[0050] Step 62: Use the state to calculate the action value with the maximum reward. The action value is 1 or 0. 1 represents selecting the current frame to be selected as the key frame, and 0 represents discarding the current frame to be selected. The method of calculating the reward adopted in this embodiment is the Q-table method. By querying the Q-value table, the action with the maximum expected reward value is obtained, where the Q-value table is obtained through reinforcement learning training.

[0051] Step 63: Execute the screening action according to the action value. The screening action can be discarding or selecting, so as to retain the key frames.

[0052] The foreground detection mentioned in the foregoing step 3 will be executed synchronously with step 6. If the requirements are not met, it will return to the mode of only executing step 3.

[0053] Step 7: Judge the lag of the video frames in the buffer. If the lag is greater than the set threshold, discard the earliest video frame that entered the buffer. If the number of video frames in the buffer is equal to the maximum capacity of the buffer, that is, the buffer is full, upload the video frames in the buffer as a group to the edge server to execute step 8; then, discard a set proportion of video frames in the order of the time of entry into the buffer; repeat steps 6 and 7 in the state where the buffer is not full. When the duration of the non-full state of the buffer reaches the threshold, return to step 3, and start recording the duration again every time the buffer is full;

[0054] Exemplarily, in this step, the average acquisition time t avg of all frames in the current buffer is calculated in real time cur , and the difference is made with the current time t diff to obtain the average lag time t cur =(t avg -t diff ), that is, the average distance between the generation time of each video frame in the buffer and the current time. When t delay is greater than the lag threshold T

[0055] is greater than the lag threshold T delay , it is considered that the data is too lagged, and the earliest added frame in the buffer will be removed. In this embodiment, the threshold value is 3 seconds. When the buffer is full of frames, the frames in the buffer will be sent to the edge server and the first 50% of the frames in the current buffer will be removed, and step 8 will be executed. Repeat steps 6 and 7 in the state where the buffer is not full.Step 8: The edge server invokes the deep learning model to perform end-to-end inference on the received video frames, and obtains the probability of the existence of violent behavior in this set of video frames.

[0056] Specifically, this step specifically includes:

[0057] Step 81: Use a convolutional neural network to extract features from a single-frame image to obtain a set of feature sets. The backbone network of the convolutional neural network adopted in this embodiment is MobileNet.

[0058] Step 82: Use a long short-term memory network to perform spatio-temporal feature extraction and inference on this set of features and obtain the final result. The specific long short-term memory network in this embodiment is a convolutional long short-term memory network, and the network length is 24.

[0059] Step 9: Issue the warning level, the involved video frames, and the device location according to the probability value.

[0060] Meanwhile, as Figure 2 shown, the present invention also provides a violent behavior detection system based on edge computing. The system includes a model training subsystem, a pre-detection subsystem, and an edge computing subsystem.

[0061] The training subsystem is deployed on the cloud server and includes a dataset construction module, a dataset preprocessing module, a detection model training module, and a reinforcement learning training module. The dataset construction module converts video datasets with violent labels of different types (such as RWF-2000, Movies, Hockey, etc.) into a standard form available for training. The dataset preprocessing module preprocesses the dataset by means of data augmentation such as scaling, mirroring, and translation, and respectively constructs sets available for training by two methods, that is, a set of video frame data with violent labels for training the deep learning module and a set of video frame data with frame importance labels for training the reinforcement learning method. The detection model training module inputs the data into the deep learning model and obtains an end-to-end model available for inference through iterative training. The reinforcement learning training module updates its own parameters based on the feedback of the detection model inference result to obtain a model available for frame screening.

[0062] The pre-detection subsystem is deployed on the monitoring device side and includes a foreground detection module, a key frame screening module, and a computing offloading module. The foreground detection module executes the foreground detection algorithm function. As the longest-running module, it ensures the low-power operation of the entire system in the absence of foreground with the operating characteristic of low resource consumption. At the same time, it wakes up other modules of this subsystem to execute when an effective foreground is obtained. The key frame screening module extracts key information from the video information, reduces the wake-up times of the edge computing system, and alleviates the network bandwidth pressure. The computing offloading module is responsible for offloading the screened video data to the edge computing subsystem after the key frame screening module meets the preset screening conditions.

[0063] Deployed on the edge computing server, it includes a video preprocessing module, an object detection module, a violence detection module, and a warning prompt module. The video prediction module preprocesses the video data unloaded to the edge computing subsystem to standardize the data input and meet the requirements of the violence detection module. The object detection module is responsible for assisting and enhancing the foreground detection module in the pre-detection subsystem, identifying the sent images, and returning the position information of the people in the images. The violence detection module performs inference calculations on the input video data to obtain the possibility of violence occurrence. The warning prompt module infers the warning level of violence occurrence based on the results of the violence detection module and displays the relevant video data for users to view.

[0064] In this embodiment, the warning levels are divided into no warning, secondary warning, and primary warning. The corresponding probability result ranges for no warning are 0 - 0.3, for secondary warning are 0.3 - 0.6, and for primary warning are 0.6 - 1. And the probability results need to go through sliding filter processing.

[0065] In a typical public area monitoring scenario, the deployment can be divided into three levels. A single monitoring device, that is, a single monitoring camera, is responsible for processing the images generated by itself; a sub-monitoring center, which consists of several monitoring devices that are physically close and an edge server, is responsible for processing all the connected monitoring devices. Taking a school as an example, sub-monitoring centers can be deployed in areas such as the library and the cafeteria respectively; the main monitoring center, which consists of a cloud server or a large local server, is responsible for processing all the sub-monitoring centers within the deployment unit. Taking a school as an example, at least one main monitoring center is deployed.

[0066] The single monitoring device uses a CPU with an ARM architecture as the computing unit. The computing resources are the scarcest among the devices at the three levels and the cost is the lowest. By using it to run the pre-detection subsystem with relatively low computing resource requirements, it can filter video frames in unmanned scenarios and low-information-density scenarios, avoid transmitting such video frames to the edge server, and run algorithms that consume high computing resources for inference calculations. By consuming less computing resources at the monitoring device end, the overall computing power requirement is saved. At the same time, since the non-violence scenarios are filtered out, it will not affect the accuracy of the final result.

[0067] The edge server in the sub-monitoring center uses a low-power GPU as the computing unit. The typical product in the industry is the NVIDIA Jetson series, with medium computing resources and costs. It has the computing resources to support the inference calculation of the violence detection model. It receives the key frames of the video to be detected uploaded by the monitoring devices in the responsible area and conducts detection. Violence detection based on deep neural networks is currently the solution with the highest detection accuracy in the technical solutions, and the accuracy of the final output of the system can be guaranteed to be at the current advanced level. The one-to-many deployment combined with the calculation of video frames in a non-full-volume and non-full-time manner realizes the reduction of the overall deployment cost.

[0068] The general monitoring center is responsible for collecting the early warning information of the sub-monitoring centers under its jurisdiction and forwarding the early warning information to users through preset fast channels such as display large screens, phones or text messages. At the same time, it is responsible for running the training subsystem, using high-computing-resource GPU clusters, etc., to relatively quickly train the models used in the deployment process and distribute them to each device in the responsible area.

Claims

1. A violent behavior detection method based on edge computing, characterized in that, The steps include: Step 1: construct and train a deep learning model for violent behavior detection on a cloud server, and construct and train a reinforcement learning method; the deep learning model takes a set of video frames as a single input, and outputs the probability of violent behavior; the reinforcement learning method takes frame-by-frame video data as input, selects a set of video frames and inputs them into the deep learning model, and iteratively updates the parameters in the reinforcement learning method according to a preset reward rule; Step 2: The monitoring device receives the video data and reads the video frames in the video data in real time; Step 3: The monitoring device uses the foreground detection algorithm to detect the foreground of the video image, and makes a judgment based on the characteristics of the foreground area. If it meets the preset conditions, it further calculates the area of ​​interest and cuts the image to obtain the area of ​​interest, and then proceeds to step 4; if it does not meet the conditions, repeat step 3; Step 4: Upload the image of the area of ​​interest to the edge server, which uses the target detection algorithm to detect the target and feeds back the result of the area with people in the image to the monitoring device; Step 5: The monitoring device uses the result of the manned area to correct the parameters of the foreground detection algorithm and determine whether the number of people in the manned area exceeds the threshold. If so, proceed to step 6, otherwise return to step 3; Step 6: On the monitoring device side, a video frame buffer with a maximum capacity of a fixed number of frames is established and a reinforcement learning method is called to filter key frames of the video frames and store the key frames in the buffer; Step 7: Determine the hysteresis of the video frames in the buffer. If the hysteresis is greater than the set threshold, discard the earliest video frame that enters the buffer. If the number of video frames in the buffer is equal to the maximum capacity of the buffer, that is, the buffer is full, upload the video frames in the buffer as a group to the edge server to execute step 8; then, discard the set proportion of video frames in the order of time they are stored in the buffer; repeat steps 6 and 7 when the buffer is not full; when the duration of the non-full state of the buffer reaches the threshold, return to step 3, and restart the recording of the duration each time the buffer is full; Step 8: The edge server calls the deep learning model to perform end-to-end reasoning on the group of video frames to obtain the probability of violent behavior in the group of video frames; Step 9: Issue the warning level and the involved video images and monitoring equipment locations based on the probability value.

2. The violent behavior detection method based on edge computing according to claim 1, wherein The deep learning model is a long short-term memory convolutional neural network.

3. The violent behavior detection method based on edge computing according to claim 1, wherein, The reinforcement learning method is a Q learning method.

4. The violent behavior detection method based on edge computing according to claim 1, wherein In step 3, the foreground detection algorithm is the Vibe algorithm; The preset condition refers to that there is a connected area in the foreground of the picture whose area is larger than a preset threshold, and the threshold is selected as the minimum value of the area of ​​the picture area in which the monitoring device can normally identify humans in the environment.

5. The violent behavior detection method based on edge computing according to claim 1, characterized in that In step 4, the target detection algorithm is the Yolo algorithm.

6. The method for detecting violent behavior based on edge computing according to claim 1, characterized in that In step 5, the result of the human area is compared with the result of the foreground detection algorithm, and the misdetected foreground in the foreground detection algorithm is updated to the background. At the same time, the minimum value of the area of ​​each area is used to update the foreground connected area threshold by using the complementary filtering algorithm.

7. The method for detecting violent behavior based on edge computing according to claim 1, characterized in that, In step 6, the method for performing key frame screening on the video frame by using the reinforcement learning method is as follows: Step 61: Calculate the inter-frame difference between the frame to be screened and the last frame entering the buffer as the state input of the reinforcement learning method; Step 62: Use the state to obtain the action with the maximum expected return value by querying the Q-value table, that is, obtain the action value with the maximum return. The action value is 1 or 0. 1 represents selecting the current frame to be selected as the key frame, and 0 represents discarding the current frame to be selected. The Q-value table is obtained through reinforcement learning training; Step 63: Execute the screening action according to the action value and retain the key frames.

8. The violent behavior detection method based on edge computing according to claim 1, characterized in that In step 7, calculate the average distance between the generation time of each video frame in the buffer and the current time. When this distance is greater than the hysteresis threshold, it is considered that the data is too lagged.

9. The method for detecting violent behavior based on edge computing according to claim 1, wherein Step 8 also includes the following sub-steps: Step 81: Use a convolutional neural network to extract features from a single-frame image to obtain a set of feature sets; Step 82: Use a long short-term memory network to perform spatio-temporal feature extraction and inference on this set of features and obtain the final result.

10. A violent behavior detection system based on edge computing, characterized in that, It includes a training subsystem, a pre-detection subsystem, and an edge computing subsystem; The training subsystem is deployed on the cloud server and includes a dataset construction module, a dataset preprocessing module, a detection model training module, and a reinforcement learning training module; The dataset construction module converts the video dataset with violence labels into a standard form for training; The dataset preprocessing module preprocesses the video dataset and constructs a video frame data set with violence labels for training the deep learning module and a video frame data set with frame importance labels for training the reinforcement learning method respectively; The detection model training module inputs the data into the deep learning model and obtains an end-to-end model that can be used for inference through iterative training; The reinforcement learning training module updates its own parameters based on the feedback of the detection model inference result to obtain a model that can be used for frame screening; The pre-detection subsystem is deployed on the monitoring device side and includes a foreground detection module, a key frame screening module, and a computing offloading module; The foreground detection module performs foreground detection on the video image to obtain the picture of the region of interest. The foreground detection module is the longest-running module and only wakes up other modules of the pre-detection subsystem to execute when an effective foreground is obtained; The key frame screening module extracts key information in the video information to reduce the wake-up times of the edge computing system; The computing offloading module is responsible for offloading the screened video frames to the edge computing subsystem after the key frame screening module meets the preset screening conditions; The edge computing subsystem is deployed on the edge computing server and includes a video preprocessing module, a target detection module, a violence detection module, and a warning prompt module; The video preprocessing module preprocesses the video frames unloaded to the edge computing subsystem to achieve the standardization of data input; The target detection module assists and enhances the foreground detection module in the pre-detection subsystem, identifies the picture sent by it, and returns the position information of the people in the picture; The violence detection module performs inference calculations on the input video frames to obtain the possibility of violence occurrence; The warning prompt module infers the warning level of violence occurrence based on the result of the violence detection module and displays the relevant video frames for the user to view.

Citation Information

Patent Citations

  • Violence behavior monitoring method and device, storage medium and terminal equipment

    CN110348343A

  • End-to-end printed Mongolian recognition translation method based on spatial transformation network

    CN112329760A