Abnormity identification method and system for realizing intelligent security and protection based on deep learning
By adopting deep learning-based methods in the intelligent security system, the video area is divided and path simulation is performed, and the residence time is determined, the path exception determination problems are solved and complexity problems in the existing system are achieved, and more efficient exception path recognition and judgment are achieved.
Patent Information
- Application Number
- CN202411995237.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
When existing intelligent security systems perform abnormal behavior detection, they are likely to lead to path abnormality determination and are complicated and troublesome.
Using a deep learning method, the video area is divided, the paths over a long time are simulated by combining the algorithm, and the residence time is determined to accurately determine the abnormal path of the target.
It improves the accuracy of security detection, can more accurately identify and determine abnormal paths, and reduces the dependence and misjudgment rate of manual monitoring.
Smart Images

Figure CN120014319A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent security technology, and more specifically, it particularly relates to an abnormality recognition method for intelligent security based on deep learning. At the same time, the present invention also relates to an abnormality recognition system for intelligent security based on deep learning. Background Art
[0002] With the continuous development of science and technology, intelligent security systems play an increasingly important role in social life. Intelligent security systems provide people with a more secure living environment with their efficient and accurate security protection capabilities. In intelligent security systems, the research on abnormal behavior detection and recognition algorithms is one of the important areas. It can monitor and analyze abnormal behaviors in the scene in real time, so as to effectively warn and prevent potential threats.
[0003] Video anomaly detection is a technology that aims to automate surveillance video analysis. Its core goal is to use computer vision systems to monitor surveillance camera images and automatically detect abnormal or unconventional activities. With the widespread use of surveillance cameras in various occasions, manual monitoring has become impractical because the task is monotonous and time-consuming. In addition, the rapid growth of surveillance equipment has made it increasingly difficult to effectively monitor a large number of cameras manually.
[0004] Abnormal detection and early warning technology in intelligent security systems can help the system detect and respond to abnormal situations in a timely manner, so as to take timely measures for security management.
[0005] Traditional intrusion detection technology mainly relies on manual monitoring, but this method has some defects, such as limited monitoring range and susceptibility to human interference. Image recognition algorithms are used to analyze monitoring images and detect abnormal movement patterns in the images to determine whether an intrusion event has occurred.
[0006] However, there are some problems with the existing technology: when the existing intelligent security detects abnormal behavior, it generally determines the target's actions or matches the abnormal behavior. There are fewer determinations of some path anomalies, which easily leads to the loss of path anomaly determination. The determination of path anomalies requires the combination of long-term video data for determination, which is relatively complicated and troublesome. Therefore, we propose an abnormality recognition method and system for intelligent security based on deep learning. Summary of the invention
[0007] In view of the problems existing in the prior art, the purpose of the present invention is to provide an abnormality identification method and system for intelligent security based on deep learning. By dividing the video area and combining the algorithm to simulate the path over a long period of time, and combining the residence time for judgment, the abnormal path of the target can be accurately determined, thereby improving the accuracy of security detection.
[0008] To achieve the above object, the present invention provides the following technical solution: a method for realizing abnormality recognition of intelligent security based on deep learning, comprising the following steps:
[0009] S1. Obtain video data: Obtain video data through the smart security camera, and upload the video data to the edge cloud through the communication module;
[0010] S2. The edge cloud pre-processes the video data: the edge cloud performs video stream transcoding and compression, data filtering and screening, edge cloud storage and caching, and real-time video analysis on the video data;
[0011] S3, real-time video analysis: real-time video analysis divides the video data into frames to extract frame images, and the convolutional neural network performs target detection on the frame images to obtain the position information of the target in the frame images;
[0012] S4, regional warning: after detecting the target in the frame image, determine whether the target is in the warning area through the position information. If the target is in the warning area, the first warning is triggered; if the target is not in the warning area, the first warning is not triggered;
[0013] S5. Extracting target features in the warning area: obtaining target features by performing feature recognition on the target of the first warning, and detecting targets on the previous and next frame images by using the target features, and recording the position information in the frame images respectively;
[0014] S6. Target path recognition: The position information in all sub-frame images is recorded through the recurrent neural network and the long short-term memory network, and the target's motion trajectory is generated, and it is determined whether the target intentionally enters the warning area. If the motion trajectory changes and enters the warning area, and wanders in the warning area, a second warning is triggered. If the motion trajectory does not change and enters the warning area, and directly walks out of the warning area, the second warning is not triggered;
[0015] S7, target stay time determination: when the target enters the warning area for the first time, the entry time is recorded, and when the target stays in the warning area, the target is continuously tracked and the stay time is increased. When the target leaves the warning area, the exit time is recorded; the stay time is calculated as the difference between the exit time and the entry time, and the third warning is triggered by determining the stay time. If the stay time is 3-5 minutes, the third warning is triggered. If the stay time is less than 3 minutes, the third warning is not triggered;
[0016] S8, the edge cloud uploads the calculated data to the platform: the edge cloud uploads the calculated frame map, the first warning, the second warning and the third warning to the platform, and the platform determines the abnormal behavior of the target;
[0017] The first warning is triggered, but the second and third warnings are not triggered, which is considered normal behavior;
[0018] The first and third warnings are triggered, but the second warning is not triggered, which is considered abnormal behavior;
[0019] The first and second warnings are triggered, but the third warning is not triggered, which is considered abnormal behavior;
[0020] Abnormal behavior is recorded and a sound alarm is triggered.
[0021] Specifically, the video stream transcoding and compression in S2 is used to transcode and compress the transmitted video stream;
[0022] Data filtering and screening is used to screen and filter video data to retain only useful or relevant video clips;
[0023] Edge cloud storage and caching is used to store video data locally in the edge cloud first, and upload it to the cloud for deeper processing or storage when needed;
[0024] Real-time video analysis is used to perform real-time analysis and processing on the edge cloud, to extract frame images from video data in real time, and to use convolutional neural networks to perform target detection on the frame images.
[0025] Specifically, the calculation formula of the convolutional neural network in S3 is as follows:
[0026] Convolution layer calculation: frame image I and convolution kernel K. The formula of convolution operation is as follows:
[0027]
[0028] Among them, (I*K) i,j is the pixel value of the output frame image, I i+m,j+n is the pixel value on the input frame image or feature map, K m,n is the convolution kernel, is the total weight value;
[0029] Pooling layer calculation:
[0030]
[0031] Among them, I i,j is an area on the input feature map, and MaxPooling(I) is the pixel value on the output feature map;
[0032] Activation function:
[0033] ReLU(x)=max(0,x),
[0034] Here, x is the value input to the ReLU function, and ReLU(x) is the output after being processed by the activation function.
[0035] Specifically, the calculation of the position information in the frame map in S3 is as follows:
[0036] The convolutional neural network performs target localization and calculates the coordinates of the bounding box in the following way:
[0037] Center coordinates (c x ,c y ): In target detection, the center coordinates of the target are used to represent the position of the target:
[0038] c x =σ(W x )×w+x anchor
[0039] c y =σ(W y )×h+y anchor ,
[0040] Among them, σ is the sigmoid activation function, ensuring that c x and c y In the image range, W x , W y is the offset of the network output; x anchor ,y anchor are the coordinates of the bounding box anchor point, w and h are the width and height of the bounding box;
[0041] Calculation of width w and height h:
[0042] w=exp(W w )×w anchor
[0043] w=exp(W h )×h anchor ,
[0044] Where: W w , W h is the offset of the width and height of the network output; w anchor 、h anchor are the width and height of the anchor point;
[0045] (c x ,c y ) as the target location information, through (c x ,c y ) determines whether the target is in the warning area. If it is in the warning area, the first warning is triggered. If it is not in the warning area, the first warning is not triggered.
[0046] Specifically, the calculation of the warning area in S4 is as follows:
[0047] Assume that there is a source point s and a sink point t in the graph G = (V, E). The minimum cut is obtained by solving the minimum flow. The value of the minimum cut is equal to the maximum flow of the source point s and the sink point t.
[0048]
[0049] Among them, E cut is the set of edges cut from the source point s to the sink point t, w(u,v) is the weight of the edge;
[0050] Maximum Flow-Minimum Cut Theorem:
[0051] MaxFlow(s,t)=MinCut(s,t),
[0052] The image area is represented by a graph, and the frame image is divided into meaningful areas, which are set as warning areas to detect whether the target enters the warning area.
[0053] Specifically, the calculation formula of the recurrent neural network in S6 is:
[0054] Hidden state calculation:
[0055] h t =f(W h ·x t +U h ·h t-1 +b h ),
[0056] Among them, h t is the hidden state at time t, x t is the input at time t, W h , U h and b h are the parameters that need to be learned;
[0057] Output calculation:
[0058] y t =f(W y ·h t +b y ),
[0059] Among them, y t is the predicted output at time t, and f is the activation function;
[0060] Recurrent neural networks are used to perform anomaly detection by identifying abnormal behaviors that are inconsistent with safe trajectories.
[0061] Specifically, the calculation of the long short-term memory network in S6 is as follows:
[0062] Forget gate: f t =σ(W f x t +U f h t-1 +b f ),
[0063] Input gate: i t =tanh(W i x t +U i h t-1 +b i ),
[0064] Candidate memory cells:
[0065] Memory unit update:
[0066] Output gate: o t =σ(W o x t +U o h t-1 +b o ),
[0067] Hide status update:h t =o t tanh(C t ),
[0068] Where: W f , W i , W c , W o are the weight matrices of the input gate, forget gate, candidate memory unit, and output gate; U f , U i , U c , U o are the weight matrices from the hidden layer to each gate; b f , b i , b c , b o are the bias terms of each gate; C t is the state of the memory unit;
[0069] When the target's motion trajectory does not match the safe trajectory, the abnormality of the trajectory is detected through the recurrent neural network and the long short-term memory network;
[0070] If the motion trajectory changes and enters the warning area, and wanders in the warning area, a second warning will be triggered. If the motion trajectory does not change and enters the warning area, and directly walks out of the warning area, the second warning will not be triggered.
[0071] Anomaly recognition system for intelligent security based on deep learning, including intelligent security cameras, communication modules, edge cloud and security platform;
[0072] The intelligent security camera is used to realize video monitoring in the security area and collect video data in the security area;
[0073] The communication module is used to transmit the video data collected by the smart security camera to the edge cloud;
[0074] The edge cloud is used to process the collected video data and analyze the video data to facilitate the acquisition of target abnormal behaviors in the video data;
[0075] The security platform is used to receive the frame image, the first warning, the second warning and the third warning transmitted by the edge cloud, and to make abnormal behavior judgments based on the frame image, the first warning, the second warning and the third warning.
[0076] Specifically, the edge cloud includes a preprocessing module, a framing module, a target detection module, a partitioning module, a path identification module and a timing module;
[0077] The preprocessing module is used to implement video stream transcoding and compression, data filtering and screening, edge cloud storage and caching, and real-time video analysis of video data;
[0078] The framing module is used to implement framing processing on the video data, so as to facilitate subsequent target detection on the frame images;
[0079] The target detection module is used to detect the target in the video data and obtain the position information of the target;
[0080] The partition module is used to realize the division of the screen of the video data collected by the intelligent security camera into warning areas, and detect the target entering the warning area, and determine whether the target is in the warning area through the position information. If it is in the warning area, the first warning is triggered, and if it is not in the warning area, the first warning is not triggered;
[0081] The path recognition module is used to record the position information in all the sub-frame images through a recurrent neural network and a long short-term memory network, and generate a motion trajectory of the target. If the motion trajectory changes and enters the warning area, and wanders in the warning area, a second warning is triggered. If the motion trajectory does not change and enters the warning area, and directly walks out of the warning area, the second warning is not triggered.
[0082] The timing module is used to calculate the time when the target enters the warning area, and trigger the third warning by determining the length of stay. If the length of stay is 3-5 minutes, the third warning is triggered, and if the length of stay is less than 3 minutes, the third warning is not triggered.
[0083] Specifically, the processing of the framing module is as follows:
[0084] The edge cloud reads the video data through video processing tools;
[0085] Read each frame of video data one by one through a loop;
[0086] Save each frame as a sub-frame image, or further process it to identify the target in the sub-frame image;
[0087] By performing target recognition on all sub-frame images in the video data, whether there is a target in the sub-frame image is used as a judgment criterion. If there is a target in the sub-frame image, frame extraction is performed. If there is no target in the sub-frame image, frame extraction is stopped.
[0088] Technical effects and advantages of the present invention:
[0089] The present invention uses edge cloud to pre-process the collected security video data, reduce the volume of video data, thereby reducing bandwidth usage, reducing video data transmission delay and dependence on cloud resources, reducing delays, and responding quickly. The edge cloud can provide real-time processing capabilities, quickly extract key events or objects from the video stream, and avoid uploading a large amount of original video data to the security platform;
[0090] The present invention performs frame processing on security video data to obtain a frame image, which is convenient for detecting a target in the frame image, and determining the position information of the target in the frame image, and dividing the security video data into warning areas, which is convenient for detecting whether the target is in the warning area. In order to improve the detection of abnormal paths, the abnormal path is determined through three-level warnings to improve the accuracy of determining the abnormal path, that is, the first determination is made through the warning area, the second determination is made through the moving path, and the third determination is made in accordance with the length of stay in the warning area, which can improve the accuracy of determining the path abnormality;
[0091] The present invention performs target detection on the frame image through a convolutional neural network, obtains the position information of the target in the frame image, clearly obtains the position information of the target, determines whether the target is in the warning area, and records the position information of all the frame images through a recurrent neural network and a long short-term memory network, and generates a motion trajectory of the target, so as to facilitate the determination of whether the target intentionally enters the warning area according to the path, and transmits the data processed by the edge cloud to the security platform. The security platform makes a judgment through three warnings, which is convenient for improving the judgment accuracy of path abnormalities, reducing the computing pressure of the security platform, improving the processing efficiency, and being able to improve the response rate of the judgment reaction.
[0092] Further features and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the present invention with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1 It is a schematic diagram of the steps provided by the present invention;
[0094] Figure 2 It is a schematic diagram of the system structure provided by the present invention;
[0095] Figure 3 It is a schematic diagram of the system structure of the edge cloud analysis and processing provided by the present invention. DETAILED DESCRIPTION
[0096] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0097] like Figure 1 As shown, the abnormality recognition method for realizing intelligent security based on deep learning provided by the embodiment of the present invention includes the following steps:
[0098] S1. Obtain video data: Obtain video data through the smart security camera, and upload the video data to the edge cloud through the communication module;
[0099] S2. The edge cloud pre-processes the video data: the edge cloud performs video stream transcoding and compression, data filtering and screening, edge cloud storage and caching, and real-time video analysis on the video data;
[0100] S3, real-time video analysis: real-time video analysis divides the video data into frames to extract frame images, and the convolutional neural network performs target detection on the frame images to obtain the position information of the target in the frame images;
[0101] S4, regional warning: after detecting the target in the frame image, determine whether the target is in the warning area through the position information. If the target is in the warning area, the first warning is triggered; if the target is not in the warning area, the first warning is not triggered;
[0102] S5. Extracting target features in the warning area: obtaining target features by performing feature recognition on the target of the first warning, and detecting targets on the previous and next frame images by using the target features, and recording the position information in the frame images respectively;
[0103] S6. Target path recognition: The position information in all sub-frame images is recorded through the recurrent neural network and the long short-term memory network, and the target's motion trajectory is generated, and it is determined whether the target intentionally enters the warning area. If the motion trajectory changes and enters the warning area, and wanders in the warning area, a second warning is triggered. If the motion trajectory does not change and enters the warning area, and directly walks out of the warning area, the second warning is not triggered;
[0104] S7, target stay time determination: when the target enters the warning area for the first time, the entry time is recorded, and when the target stays in the warning area, the target is continuously tracked and the stay time is increased. When the target leaves the warning area, the exit time is recorded; the stay time is calculated as the difference between the exit time and the entry time, and the third warning is triggered by determining the stay time. If the stay time is 3-5 minutes, the third warning is triggered. If the stay time is less than 3 minutes, the third warning is not triggered;
[0105] S8, the edge cloud uploads the calculated data to the platform: the edge cloud uploads the calculated frame map, the first warning, the second warning and the third warning to the platform, and the platform determines the abnormal behavior of the target;
[0106] The first warning is triggered, but the second and third warnings are not triggered, which is considered normal behavior;
[0107] The first and third warnings are triggered, but the second warning is not triggered, which is considered abnormal behavior;
[0108] The first and second warnings are triggered, but the third warning is not triggered, which is considered abnormal behavior;
[0109] Abnormal behavior is recorded and a sound alarm is triggered.
[0110] In this embodiment, preferably, the video stream transcoding and compression in S2 is used to transcode and compress the transmitted video stream;
[0111] Data filtering and screening is used to screen and filter video data to retain only useful or relevant video clips;
[0112] Edge cloud storage and caching is used to store video data locally in the edge cloud first, and upload it to the cloud for deeper processing or storage when needed;
[0113] Real-time video analysis is used to perform real-time analysis and processing on the edge cloud. It extracts frame images from video data in real time, and uses convolutional neural networks to detect targets on the frame images.
[0114] It should be noted that transcoding and compressing the video stream can reduce the volume of video data, thereby reducing bandwidth usage, and through compression, the amount of data that needs to be transmitted to the data center or cloud can be reduced, and only important information or processed data can be sent to the cloud; screening and filtering the video data to retain only useful or relevant video clips can significantly reduce the invalid video data uploaded to the cloud, or discard frames when static and inactive, which can save storage space and optimize bandwidth usage; video data can be stored locally first and uploaded to the cloud for deeper processing or storage only when needed, which helps to reduce the transmission delay of video data and dependence on cloud resources, and can also cache some commonly used processing results or video data to quickly respond to subsequent requests and improve the efficiency of the overall system; on the edge node, the video stream can be analyzed and processed in real time, thereby reducing latency and responding quickly. The edge cloud can provide real-time processing capabilities to quickly extract key events or objects from the video stream and avoid uploading large amounts of raw video data to centralized data centers.
[0115] In this embodiment, preferably, the calculation formula of the convolutional neural network in S3 is as follows:
[0116] Convolution layer calculation: frame image I and convolution kernel K. The formula of convolution operation is as follows:
[0117]
[0118] Among them, (I*K) i,j is the pixel value of the output frame image, I i+m,j+n is the pixel value on the input frame image or feature map, K m,n is the convolution kernel, is the total weight value;
[0119] Pooling layer calculation:
[0120]
[0121] Among them, I i,jis an area on the input feature map, and MaxPooling(I) is the pixel value on the output feature map;
[0122] Activation function:
[0123] ReLU(x)=max(0,x),
[0124] Where x is the value input to the ReLU function, and ReLU(x) is the output after being processed by the activation function;
[0125] It should be noted that by using deep learning technology in spatial feature extraction and anomaly detection, and by automatically learning features, CNN can detect abnormal events that are significantly different from normal behaviors from videos, and can automatically learn abnormal behavior patterns, thereby effectively improving security protection capabilities.
[0126] In this embodiment, preferably, the position information in the frame graph in S3 is calculated as follows:
[0127] The convolutional neural network performs target localization and calculates the coordinates of the bounding box in the following way:
[0128] Center coordinates (c x ,c y ): In target detection, the center coordinates of the target are used to represent the position of the target:
[0129] c x =σ(W x )×w+x anchor
[0130] c y =σ(W y )×h+y anchor ,
[0131] Among them, σ is the sigmoid activation function, ensuring that c x and c y In the image range, W x , W y is the offset of the network output; x anchor ,y anchor are the coordinates of the bounding box anchor point, w and h are the width and height of the bounding box;
[0132] Calculation of width w and height h:
[0133] w=exp(W w )×w anchor
[0134] w=exp(W h )×h anchor ,
[0135] Where: W w , W h is the offset of the width and height of the network output; w anchor 、h anchor are the width and height of the anchor point;
[0136] (c x ,c y ) as the target location information, through (c x ,c y ) determining whether the target is within the warning area, and if so, triggering the first warning; if not, not triggering the first warning;
[0137] It should be noted that determining the position of the target in the frame map facilitates determining the position information of the target, and connecting the target positions in the frame maps through the target positions of multiple frame maps facilitates determining the path information of the target.
[0138] In this embodiment, preferably, the calculation of the warning area in S4 is as follows:
[0139] Assume that there is a source point s and a sink point t in the graph G = (V, E). The minimum cut is obtained by solving the minimum flow. The value of the minimum cut is equal to the maximum flow of the source point s and the sink point t.
[0140]
[0141] Among them, E cut is the set of edges cut from the source point s to the sink point t, w(u,v) is the weight of the edge;
[0142] Maximum Flow-Minimum Cut Theorem:
[0143] MaxFlow(s,t)=MinCut(s,t),
[0144] The image area is represented by a graph, and the frame image is divided into meaningful areas, which are set as warning areas to detect whether the target enters the warning area;
[0145] It should be noted that by dividing the area map of security video collection into areas using a graph segmentation algorithm, such as a minimum cut algorithm, it is convenient to divide the warning area and monitor the warning area.
[0146] In this embodiment, preferably, the calculation formula of the recurrent neural network in S6 is:
[0147] Hidden state calculation:
[0148] h t =f(Wh ·x t +U h ·h t-1 +b h ),
[0149] Among them, h t is the hidden state at time t, x t is the input at time t, W h , U h and b h are the parameters that need to be learned;
[0150] Output calculation:
[0151] y t =f(W y ·h t +b y ),
[0152] Among them, y t is the predicted output at time t, and f is the activation function;
[0153] Recurrent neural networks are used to identify abnormal behaviors that are inconsistent with safe trajectories, thereby performing anomaly detection;
[0154] It should be noted that the recurrent neural network enables information to flow in time steps through the recurrent connections between its hidden layers, and can capture the time dependency in the sequence. The output of the recurrent neural network at each moment depends not only on the current input, but also on the hidden state at the previous moment, making it suitable for processing time series data and facilitating the acquisition of the target's motion trajectory.
[0155] In this embodiment, preferably, the calculation of the long short-term memory network in S6 is as follows:
[0156] Forget gate: f t =σ(W f x t +U f h t-1 +b f ),
[0157] Input gate: i t =tanh(W i x t +U i h t-1 +b i ),
[0158] Candidate memory cells:
[0159] Memory unit update:
[0160] Output gate: o t =σ(W o x t +U o h t-1 +b o ),
[0161] Hide status update:h t =o t tanh(C t ),
[0162] Where: W f , W i , W c , W o are the weight matrices of the input gate, forget gate, candidate memory unit, and output gate; U f , U i , U c , U o are the weight matrices from the hidden layer to each gate; b f , b i , b c , b o are the bias terms of each gate; C t is the state of the memory unit;
[0163] When the target's motion trajectory does not match the safe trajectory, the abnormality of the trajectory is detected through the recurrent neural network and the long short-term memory network;
[0164] If the motion trajectory changes and enters the warning area, and wanders in the warning area, the second warning will be triggered. If the motion trajectory does not change and enters the warning area, and directly walks out of the warning area, the second warning will not be triggered;
[0165] It should be noted that the long short-term memory network solves the gradient vanishing and gradient exploding problems encountered by ordinary RNN when processing long sequences. By introducing memory units and gating mechanisms, the long short-term memory network can retain or forget information for a long time and perform better when dealing with long-term dependency problems. When the target's motion trajectory does not match the safe trajectory, the RNN or LSTM detects the abnormality of the trajectory and detects that the target suddenly changes its speed or direction, which may mean that the target intentionally enters the warning area.
[0166] refer to Figure 2-3 , an abnormality recognition system for intelligent security based on deep learning, including intelligent security cameras, communication modules, edge cloud and security platform;
[0167] The intelligent security camera is used to realize video monitoring in the security area and collect video data in the security area;
[0168] The communication module is used to transmit the video data collected by the smart security camera to the edge cloud;
[0169] The edge cloud is used to process the collected video data and analyze the video data to facilitate the acquisition of target abnormal behaviors in the video data;
[0170] The security platform is used to receive the frame image, the first warning, the second warning and the third warning transmitted by the edge cloud, and to make abnormal behavior judgments based on the frame image, the first warning, the second warning and the third warning.
[0171] In this embodiment, preferably, the edge cloud includes a preprocessing module, a framing module, a target detection module, a partitioning module, a path identification module and a timing module;
[0172] The preprocessing module is used to implement video stream transcoding and compression, data filtering and screening, edge cloud storage and caching, and real-time video analysis of video data;
[0173] The framing module is used to implement framing processing on the video data, so as to facilitate subsequent target detection on the frame images;
[0174] The target detection module is used to detect the target in the video data and obtain the position information of the target;
[0175] The partition module is used to realize the division of the screen of the video data collected by the intelligent security camera into warning areas, and detect the target entering the warning area, and determine whether the target is in the warning area through the position information. If it is in the warning area, the first warning is triggered, and if it is not in the warning area, the first warning is not triggered;
[0176] The path recognition module is used to record the position information in all the sub-frame images through a recurrent neural network and a long short-term memory network, and generate a motion trajectory of the target. If the motion trajectory changes and enters the warning area, and wanders in the warning area, a second warning is triggered. If the motion trajectory does not change and enters the warning area, and directly walks out of the warning area, the second warning is not triggered.
[0177] The timing module is used to calculate the time when the target enters the warning area, and trigger the third warning by determining the length of stay. If the length of stay is 3-5 minutes, the third warning is triggered, and if the length of stay is less than 3 minutes, the third warning is not triggered;
[0178] It should be noted that the edge cloud can be used to pre-process the collected security video data, reduce the volume of video data, thereby reducing bandwidth usage, reducing the transmission delay of video data and dependence on cloud resources, reducing delays, and responding quickly. The edge cloud can provide real-time processing capabilities, quickly extract key events or objects from the video stream, and avoid uploading large amounts of raw video data to the security platform.
[0179] In this embodiment, preferably, the processing of the framing module is as follows:
[0180] The edge cloud reads the video data through video processing tools;
[0181] Read each frame of video data one by one through a loop;
[0182] Save each frame as a sub-frame image, or further process it to identify the target in the sub-frame image;
[0183] By performing target recognition on all sub-frame images in the video data, whether there is a target in the sub-frame image is used as a judgment standard. If there is a target in the sub-frame image, frame extraction is performed. If there is no target in the sub-frame image, frame extraction is stopped.
[0184] It should be noted that by performing frame processing on the video data, the video data can be formed into frame images, which is convenient for target recognition on the frame images and for effective and rapid analysis and processing of the frame images.
[0185] The specific operation process of this application:
[0186] Step 1: Obtain video data: Obtain video data through the smart security camera, and upload the video data to the edge cloud through the communication module;
[0187] Step 2: The edge cloud pre-processes the video data: the edge cloud performs video stream transcoding and compression, data filtering and screening, edge cloud storage and caching, and real-time video analysis on the video data;
[0188] Step 3: Real-time video analysis: Real-time video analysis divides the video data into frames to extract frame images, and a convolutional neural network performs target detection on the frame images to obtain the location information of the target in the frame images;
[0189] The calculation formula of the convolutional neural network is as follows:
[0190] Convolution layer calculation: frame image I and convolution kernel K, the formula of convolution operation is as follows:
[0191]
[0192] Among them, (I*K)i,j is the pixel value of the output frame image, I i+m,j+n is the pixel value on the input frame image or feature map, K m,n is the convolution kernel, is the total weight value;
[0193] Pooling layer calculation:
[0194]
[0195] Among them, I i,j is an area on the input feature map, and MaxPooling(I) is the pixel value on the output feature map;
[0196] Activation function:
[0197] ReLU(x)=max(0,x),
[0198] Where x is the value input to the ReLU function, and ReLU(x) is the output after being processed by the activation function;
[0199] Step 4: Area warning: After detecting the target in the frame image, determine whether the target is in the warning area based on the position information. If the target is in the warning area, the first warning is triggered. If the target is not in the warning area, the first warning is not triggered.
[0200] The warning area is calculated as follows:
[0201] Assume that there is a source point s and a sink point t in the graph G = (V, E). The minimum cut is obtained by solving the minimum flow. The value of the minimum cut is equal to the maximum flow of the source point s and the sink point t.
[0202]
[0203] Among them, E cut is the set of edges cut from the source point s to the sink point t, w(u,v) is the weight of the edge;
[0204] Maximum Flow-Minimum Cut Theorem:
[0205] MaxFlow(s,t)=MinCut(s,t),
[0206] The image area is represented by a graph, and the frame image is divided into meaningful areas, which are set as warning areas to detect whether the target enters the warning area;
[0207] The calculation of the position information in the frame map is as follows:
[0208] The convolutional neural network performs target localization and calculates the coordinates of the bounding box in the following way:
[0209] Center coordinates (cx ,c y ): In target detection, the center coordinates of the target are used to represent the position of the target:
[0210]
[0211] Among them, σ is the sigmoid activation function, ensuring that c x and c y In the image range, W x , W y is the offset of the network output; x anchor ,y anchor are the coordinates of the bounding box anchor point, w and h are the width and height of the bounding box;
[0212] Calculation of width w and height h:
[0213]
[0214] Where: W w , W h is the offset of the width and height of the network output; w anchor 、h anchor is the width and height of the anchor point;
[0215] (c x ,c y ) as the target location information, through (c x ,c y ) determining whether the target is within the warning area, and if so, triggering the first warning; if not, not triggering the first warning;
[0216] Step 5: Extract target features in the warning area: Obtain target features by performing feature recognition on the target of the first warning, and use the target features to detect targets on the previous and next frame images, and record the position information in the frame images respectively;
[0217] Step 6: Target path identification: The position information in all sub-frame images is recorded through the recurrent neural network and the long short-term memory network, and the target's motion trajectory is generated, and it is determined whether the target intentionally enters the warning area. If the motion trajectory changes and enters the warning area, and wanders in the warning area, a second warning is triggered. If the motion trajectory does not change and enters the warning area, and directly walks out of the warning area, the second warning is not triggered;
[0218] The calculation formula of the recurrent neural network is:
[0219] Hidden state calculation:
[0220] h t =f(Wh ·x t +U h ·h t-1 +b h ),
[0221] Among them, h t is the hidden state at time t, x t is the input at time t, W h , U h and b h are the parameters that need to be learned;
[0222] Output calculation:
[0223] y t =f(W y ·h t +b y ),
[0224] Among them, y t is the predicted output at time t, and f is the activation function;
[0225] Recurrent neural networks are used to identify abnormal behaviors that are inconsistent with safe trajectories, thereby performing anomaly detection;
[0226] The calculation of the long short-term memory network is as follows:
[0227] Forget gate: f t =σ(W f x t +U f h t-1 +b f ),
[0228] Input gate: i t =tanh(W i x t +U i h t-1 +b i ),
[0229] Candidate memory cells:
[0230] Memory unit update:
[0231] Output gate: o t =σ(W o x t +U o h t-1 +b o ),
[0232] Hide status update:h t =o ttanh(C t ),
[0233] Where: W f , W i , W c , W o are the weight matrices of the input gate, forget gate, candidate memory unit, and output gate; U f , U i , U c , U o are the weight matrices from the hidden layer to each gate; b f , b i , b c , b o are the bias terms of each gate; C t is the state of the memory unit;
[0234] When the target's motion trajectory does not match the safe trajectory, the abnormality of the trajectory is detected through the recurrent neural network and the long short-term memory network;
[0235] If the motion trajectory changes and enters the warning area, and wanders in the warning area, the second warning will be triggered. If the motion trajectory does not change and enters the warning area, and directly walks out of the warning area, the second warning will not be triggered;
[0236] Step 7: Determine the target's stay time: When the target enters the warning area for the first time, record the entry time, and when the target stays in the warning area, continue to track the target and increase the stay time. When the target leaves the warning area, record the exit time. Calculate the stay time as the difference between the exit time and the entry time. Trigger the third warning by determining the stay time. If the stay time is 3-5 minutes, the third warning is triggered. If the stay time is less than 3 minutes, the third warning is not triggered.
[0237] Step 8: The edge cloud uploads the calculated data to the platform: The edge cloud uploads the calculated frame image, the first warning, the second warning, and the third warning to the platform, and the platform determines the abnormal behavior of the target;
[0238] The first warning is triggered, but the second and third warnings are not triggered, which is considered normal behavior;
[0239] The first and third warnings are triggered, but the second warning is not triggered, which is considered abnormal behavior;
[0240] The first and second warnings are triggered, but the third warning is not triggered, which is considered abnormal behavior;
[0241] Abnormal behavior is recorded and a sound alarm is triggered.
[0242] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. An abnormality identification method for intelligent security based on deep learning, characterized in that: The steps include: S1. Obtain video data: Obtain video data through the smart security camera, and upload the video data to the edge cloud through the communication module; S2. The edge cloud pre-processes the video data: the edge cloud performs video stream transcoding and compression, data filtering and screening, edge cloud storage and caching, and real-time video analysis on the video data; S3, real-time video analysis: real-time video analysis divides the video data into frames to extract frame images, and the convolutional neural network performs target detection on the frame images to obtain the position information of the target in the frame images; S4, regional warning: after performing target detection on the frame image, determine whether the target is in the warning area through the position information. If it is in the warning area, the first warning is triggered; if it is not in the warning area, the first warning is not triggered; S5. Extracting target features in the warning area: obtaining target features by performing feature recognition on the target of the first warning, and detecting targets on the previous and next frame images by using the target features, and recording the position information in the frame images respectively; S6. Target path recognition: The position information in all sub-frame images is recorded through the recurrent neural network and the long short-term memory network, and the target's motion trajectory is generated, and it is determined whether the target intentionally enters the warning area. If the motion trajectory changes and enters the warning area, and wanders in the warning area, a second warning is triggered. If the motion trajectory does not change and enters the warning area, and directly walks out of the warning area, the second warning is not triggered; S7, target stay time determination: when the target enters the warning area for the first time, the entry time is recorded, and when the target stays in the warning area, the target is continuously tracked and the stay time is increased, and when the target leaves the warning area, the exit time is recorded; The length of stay is calculated as the difference between the exit time and the entry time. The third warning is triggered by determining the length of stay. If the length of stay is 3-5 minutes, the third warning is triggered. If the length of stay is less than 3 minutes, the third warning is not triggered. S8, the edge cloud uploads the calculated data to the platform: the edge cloud uploads the calculated frame map, the first warning, the second warning and the third warning to the platform, and the platform determines the abnormal behavior of the target; The first warning is triggered, but the second and third warnings are not triggered, which is considered normal behavior; The first and third warnings are triggered, but the second warning is not triggered, which is considered abnormal behavior; The first and second warnings are triggered, but the third warning is not triggered, which is considered abnormal behavior; Abnormal behavior is recorded and a sound alarm is triggered.
2. The method for realizing abnormality identification of intelligent security based on deep learning according to claim 1 is characterized in that: The video stream transcoding and compression in S2 is used to transcode and compress the transmitted video stream; Data filtering and screening is used to screen and filter video data to retain only useful or relevant video clips; Edge cloud storage and caching is used to store video data locally in the edge cloud first, and upload it to the cloud for deeper processing or storage when needed; Real-time video analysis is used to perform real-time analysis and processing on the edge cloud, to extract frame images from video data in real time, and to use convolutional neural networks to perform target detection on the frame images.
3. The method for abnormality identification based on deep learning to realize intelligent security according to claim 1, characterized in that: The calculation formula of the convolutional neural network in S3 is as follows: Convolution layer calculation: frame image I and convolution kernel K. The formula of convolution operation is as follows: Among them, (I*K) i,j is the pixel value of the output frame image, Ii + m,j + n is the pixel value on the input frame image or feature map, K m,n is the convolution kernel, is the total weight value; Pooling layer calculation: Among them, I i,j is an area on the input feature map, and MaxPooling(I) is the pixel value on the output feature map; Activation function: ReLU(x)=max(0,x), Here, x is the value input to the ReLU function, and ReLU(x) is the output after being processed by the activation function.
4. The method for abnormality identification based on deep learning to realize intelligent security according to claim 1, characterized in that: The calculation of the position information in the frame map in S3 is as follows: The convolutional neural network performs target localization and calculates the coordinates of the bounding box in the following way: Center coordinates (c x ,c y ): In target detection, the center coordinates of the target are used to represent the position of the target: Among them, σ is the sigmoid activation function, ensuring that c x and c y In the image range, W x , W y is the offset of the network output; x anchor ,y anchor are the coordinates of the bounding box anchor point, w and h are the width and height of the bounding box; Calculation of width w and height h: Where: W w , W h is the offset of the width and height of the network output; w anchor 、h anchor are the width and height of the anchor point; Take (cx, cy) as the target location information, and use (c x ,c y ) determines whether the target is in the warning area. If it is in the warning area, the first warning is triggered. If it is not in the warning area, the first warning is not triggered.
5. The method for abnormality identification based on deep learning to realize intelligent security according to claim 1, characterized in that: The calculation of the warning area in S4 is as follows: Assume that there is a source point s and a sink point t in the graph G = (V, E). The minimum cut is obtained by solving the minimum flow. The value of the minimum cut is equal to the maximum flow of the source point s and the sink point t. Among them, E cut is the set of edges cut from the source point s to the sink point t, w(u,v) is the weight of the edge; Maximum Flow-Minimum Cut Theorem: MaxFlow(s,t)=MinCut(s,t), The image area is represented by a graph, and the frame image is divided into meaningful areas, which are set as warning areas to detect whether the target enters the warning area.
6. The method for abnormality identification based on deep learning to realize intelligent security according to claim 1, characterized in that: The calculation formula of the recurrent neural network in S6 is: Hidden state calculation: h t =f(W h ·x t +U h ·h t-1 +b h ), Among them, h t is the hidden state at time t, x t is the input at time t, W h , U h and b h are the parameters that need to be learned; Output calculation: y t =f(W y ·h t +b y ), Among them, y t is the predicted output at time t, and f is the activation function; Recurrent neural networks are used to perform anomaly detection by identifying abnormal behaviors that are inconsistent with safe trajectories.
7. The method for abnormality identification based on deep learning to realize intelligent security according to claim 1, characterized in that: The calculation of the long short-term memory network in S6 is as follows: Forget gate: f t =σ(W f x t +U f h t-1 +b f ), Input gate: i t =tanh(W i x t +U i h t-1 +b i ), Candidate memory cells: Memory unit update: Output gate: o t =σ(W o x t +U o h t-1 +b o ), Hide status update:h t =o t tanh(C t ), Where: W f , W i , W c , W o are the weight matrices of the input gate, forget gate, candidate memory unit, and output gate; U f , U i , U c , U o are the weight matrices from the hidden layer to each gate; b f , b i , b c , b o are the bias terms of each gate; C t is the state of the memory unit; When the target's motion trajectory does not match the safe trajectory, the abnormality of the trajectory is detected through the recurrent neural network and the long short-term memory network; If the motion trajectory changes and enters the warning area, and wanders in the warning area, a second warning will be triggered. If the motion trajectory does not change and enters the warning area, and directly walks out of the warning area, the second warning will not be triggered.
8. An abnormality recognition system for intelligent security based on deep learning, used to implement the steps of claims 1-7 above; characterized in that: Includes smart security cameras, communication modules, edge cloud and security platform; The intelligent security camera is used to realize video monitoring in the security area and collect video data in the security area; The communication module is used to transmit the video data collected by the smart security camera to the edge cloud; The edge cloud is used to process the collected video data and analyze the video data to facilitate the acquisition of target abnormal behaviors in the video data; The security platform is used to receive the frame image, the first warning, the second warning and the third warning transmitted by the edge cloud, and to make abnormal behavior judgments based on the frame image, the first warning, the second warning and the third warning.
9. The abnormality recognition system for realizing intelligent security based on deep learning according to claim 8, characterized in that: The edge cloud includes a preprocessing module, a framing module, a target detection module, a partitioning module, a path identification module and a timing module; The preprocessing module is used to implement video stream transcoding and compression, data filtering and screening, edge cloud storage and caching, and real-time video analysis of video data; The framing module is used to implement framing processing on the video data, so as to facilitate subsequent target detection on the frame images; The target detection module is used to detect the target in the video data and obtain the position information of the target; The partition module is used to realize the division of the screen of the video data collected by the intelligent security camera into warning areas, and detect the target entering the warning area, and determine whether the target is in the warning area through the position information. If it is in the warning area, the first warning is triggered, and if it is not in the warning area, the first warning is not triggered; The path recognition module is used to record the position information in all the sub-frame images through a recurrent neural network and a long short-term memory network, and generate a motion trajectory of the target. If the motion trajectory changes and enters the warning area, and wanders in the warning area, a second warning is triggered. If the motion trajectory does not change and enters the warning area, and directly walks out of the warning area, the second warning is not triggered. The timing module is used to calculate the time when the target enters the warning area, and trigger the third warning by determining the length of stay. If the length of stay is 3-5 minutes, the third warning is triggered, and if the length of stay is less than 3 minutes, the third warning is not triggered.
10. The abnormality recognition system for realizing intelligent security based on deep learning according to claim 9, characterized in that: The processing of the framing module is as follows: The edge cloud reads the video data through video processing tools; Read each frame of video data one by one through a loop; Save each frame as a sub-frame image, or further process it to identify the target in the sub-frame image; By performing target recognition on all sub-frame images in the video data, whether there is a target in the sub-frame image is used as a judgment criterion. If there is a target in the sub-frame image, frame extraction is performed. If there is no target in the sub-frame image, frame extraction is stopped.
Citation Information
Cited By
Behavior pattern abnormity identification method under multi-dimensional security data fusion
CN120873889A
Method for behavior pattern anomaly recognition under multi-dimensional security data fusion
CN120873889B