Leakage behavior identification method based on deep learning

By using the AlphaPose framework and Dynamic R-CNN and SLATracker algorithms in military airports, the node and skeleton features are extracted, combined with the deep learning method of CNN convolutional neural network and attention module, the problems of low recognition efficiency and poor accuracy in traditional methods are solved, real-time and accurate identification and alarm of leaked behavior are achieved.

CN115861877BActive Publication Date: 2025-08-12CHINA CONSTR EIGHT ENG DIV CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211466524.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-22
Publication Date
2025-08-12
Estimated Expiration
2042-11-22

AI Technical Summary

Technical Problem

The prior art is difficult to accurately and in real time to identify leaks in military airports. Traditional surveillance systems rely on manpower to analyze inefficiently and in poor accuracy. The existing deep learning methods such as 3D convolution, dual-stream networks and LSTMs are still insufficient in recognition accuracy.

Method used

The AlphaPose framework is used to extract the joint nodes and skeleton features, combine Dynamic R-CNN and SLATracker algorithms for object detection and tracking, build a CNN convolutional neural network and introduce channel and spatial attention modules to identify leak behaviors through deep learning recognition models.

Benefits of technology

It improves the accuracy and efficiency of identifying illegal shooting behaviors in military airports, reduces costs, and realizes real-time and accurate identification and alarm of leaks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115861877B_ABST
    Figure CN115861877B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying leaking behaviors based on deep learning, comprising the following steps: 1. obtaining a video of a military airport captured by surveillance equipment; 2. preprocessing the surveillance video, pre-training the AlphaPose framework, and extracting the joint points and skeleton features of the people appearing in the surveillance video; 3. calculating the motion joint feature vectors of the people in the surveillance video based on the joint points and skeleton features, and obtaining the motion joint features of each person in the surveillance video; 4. establishing a deep learning recognition model, classifying the motion joint features of each person, and identifying target motion features; 5. marking the people corresponding to the target motion features classified as leaking behaviors as abnormal, obtaining the target abnormal object attributes and their motion trajectory, and reporting the abnormal target objects to the security department for processing. The present invention can accurately and real-timely identify suspected illegal filming leaking behaviors at military airports.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a behavior recognition method, and in particular to a leakage behavior recognition method based on deep learning. Background Art

[0002] Leakage refers to the intentional or negligent disclosure of state and military secrets by entities or individuals responsible for safeguarding state and military secrets to unauthorized individuals or external parties. It also refers to the partial or complete loss of confidentiality of trade secrets due to an enterprise's own intentional or negligent conduct, accidents, intentional infringement by others, or other reasons, resulting in the loss of competitive advantage or disadvantage to the enterprise. However, human behavior is difficult to predict, especially some leaks, which are highly concealed. Failure to promptly detect and prevent these leaks could seriously harm national and corporate interests.

[0003] Common leaks at military airports include voyeurism and illegal copying, recording, and storage of confidential information. Traditional surveillance systems rely mainly on manpower to analyze massive amounts of video information to identify leaks. This not only consumes a lot of manpower and financial resources, is inefficient, but also cannot guarantee the accuracy of identification.

[0004] At present, the main behavior recognition methods based on deep learning are: 3D convolution method, two-stream network method, and LSTM (Long short-term memory) method.

[0005] The advantage of 3D convolution is that it is simple and direct, and can directly obtain the spatiotemporal characteristics of the video. However, 3D convolution simply adds the time dimension to the convolutional network. Although it has made great progress compared to 2D convolution, its recognition accuracy is still poor and it cannot accurately identify the filming and leaking behavior in military airports.

[0006] The advantage of a two-stream network is that it can extract spatial and temporal features separately using two network channels. This network design perfectly aligns with the inherent spatiotemporal characteristics of video, resulting in a more elegant design and more efficient feature extraction compared to independent features. However, a disadvantage of a two-stream network is that it struggles to achieve perfect interaction between spatial and temporal features. The separation of time and space can affect the accuracy of behavior recognition, and the precise layer at which spatiotemporal features should be integrated is also difficult to determine. The accuracy of the two-stream network in identifying leaks recorded at military airports still needs to be improved.

[0007] The LSTM algorithm exploits temporal features, making better use of temporal attributes and resolving the difficulty in effectively utilizing temporal features in videos. The current mainstream approach to spatial feature extraction is to use convolutional neural networks to generate feature sequences. However, the LSTM algorithm requires a well-designed convolutional neural network to generate these feature sequences. Otherwise, temporal information is lost, preventing the RNN (recurrent neural network) from extracting effective temporal features and hindering the effective identification of leaks captured at military airports.

[0008] Furthermore, both recurrent neural networks and convolutional neural networks require specialized knowledge of human motion for modeling, making the modeling process complex and demanding. Therefore, a deep learning-based method for identifying leaks is needed that can accurately and in real time identify suspected illegal filming leaks at military airports. Summary of the Invention

[0009] The purpose of the present invention is to provide a method for identifying leaking behaviors based on deep learning, which can accurately and in real time identify suspected illegal filming leaking behaviors at military airports.

[0010] The present invention is achieved in that:

[0011] A method for identifying leak behavior based on deep learning, comprising the following steps:

[0012] Step 1: Obtain the video of the military airport captured by the surveillance equipment;

[0013] Step 2: Pre-process the surveillance video, pre-train the AlphaPose framework, and use the AlphaPose framework to extract the joint points and skeleton features of the people appearing in the surveillance video;

[0014] Step 3: Based on the extracted joint points and skeleton features of the personnel, the action joint feature vectors of the personnel in the surveillance video are calculated to obtain the action joint features of each person in the surveillance video;

[0015] Step 4: Import the motion joint features into the model, optimize and establish a deep learning recognition model, and use the deep learning recognition model to classify the motion joint features of each person and identify the target motion features;

[0016] Step 5: Mark the person corresponding to the target action feature classified as a leak behavior as abnormal, obtain the target abnormal object attributes and its movement trajectory, and report the abnormal target object to the security department for processing.

[0017] The preprocessing method of the monitoring video screen is: constructing the target motion trajectory and extracting the target features, and screening the tracking target, that is, the person appearing in the monitoring video screen, based on the target feature comparison of the previous and next frame monitoring video frame sequences.

[0018] When screening out tracking targets, the target detection algorithm is used to detect the target, and then the SLATracker tracking algorithm is used to track multiple targets for a long time in complex scenes to screen out tracking targets.

[0019] The method for extracting the joint points and skeleton features of people appearing in the surveillance video screen by the AlphaPose framework is as follows: extracting activity frames with people in the surveillance video screen, using the AlphaPose framework to extract the joint points of people in the activity frames, and connecting the skeletons of adjacent joints.

[0020] The joints of the person include nose, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, hip, left hip, right hip, left knee, right knee, left ankle and right ankle.

[0021] In step 3, the "local pelvic divergence method" is used to draw several left and right action feature joint vectors with the hip joint point as the starting point and the shoulder, elbow, and wrist joint points as the end points, and calculate the linear displacement and vector angle of the action feature in the active frame. The changes in the linear displacement and vector angle are used to determine whether the action feature joint vector is consistent with the target action feature.

[0022] The target action features include single-arm shooting, double-arm shooting, single-arm calling for help, and double-arm calling for help.

[0023] Described step 4 comprises the following sub-steps:

[0024] Step 4.1: Construct the model structure;

[0025] Step 4.2: Initialize the parameters in the model structure, including weights, biases, reduction rates, and learning rates;

[0026] Step 4.3: Based on the joint points and skeleton features, the model structure is trained using forward learning and back propagation;

[0027] Step 4.4: When the end conditions are met, model training is completed, the model parameters are optimized, and a deep learning recognition model is obtained that can meet the requirements of leak behavior identification;

[0028] Step 4.5: Use the deep learning recognition model to identify and classify the motion joint features of each person and identify the target motion features of the leak behavior.

[0029] The model structure consists of a CNN convolutional neural network, a pooling module, an attention module and an activation function;

[0030] Step 4.1 includes the following sub-steps:

[0031] Step 4.1.1: Add a channel attention module to the CNN convolutional neural network and use global average pooling and maximum pooling to utilize different information respectively to summarize spatial features;

[0032] Step 4.1.2: Introduce the residual block into the model structure;

[0033] The method of introducing the residual block is:

[0034] Assuming X is the input of the residual network block and F(X) is the residual mapping function, the original mapping can be expressed as Y=F(X)+X;

[0035] For feature F, global average pooling and maximum pooling are performed in a space to obtain two 1×1×C channel descriptions. The two channel descriptions are then fed into a two-layer neural network. The number of neurons in the first layer is C / r, the activation function is Relu, and the number of neurons in the second layer is C, to obtain two features. This two-layer neural network is shared. The two features are then added together and passed through a Sigmoid activation function to obtain the weight coefficient Mc. Finally, the weight coefficient Mc is multiplied by the original feature F to obtain the scaled new feature.

[0036] Step 4.1.3: After adding the channel attention module, introduce the spatial attention module to focus on meaningful action joint features;

[0037] The introduction method of the spatial attention module is:

[0038] Given a H×W×C feature F', we first perform average pooling and maximum pooling on each channel dimension to obtain two H×W×1 channel descriptions, and then concatenate these two channel descriptions together according to the channel. Then, we pass it through an X×X convolution layer with a Sigmoid activation function to obtain the weight coefficient Ms of spatial attention. Finally, we multiply the weight coefficient Ms by the feature F' to obtain the scaled new feature.

[0039] If the channel attention module teaches the model structure what to pay attention to, the spatial attention module will allow the model structure to know where to pay attention.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] 1. Since the present invention adopts the target detection algorithm Dynamic R-CNN to detect targets, it can dynamically change the training strategy as the data changes when training the target detector. It also uses the SLATracker tracking algorithm to track multiple targets for a long time in complex scenarios, which is used to accurately extract people appearing in surveillance video images, which is conducive to improving the accuracy of subsequent target action feature recognition, thereby accurately identifying illegal filming behaviors in military airports and avoiding military secret leaks.

[0042] 2. The present invention detects people in the picture through the AlphaPose framework, extracts key points and skeleton features, and obtains a two-dimensional human skeleton diagram containing 15 human joints and the bone lines connected to them. It also uses the "pelvic divergence method" to establish a key feature vector of the action with the hip joint as the starting point, identifies linear displacement and vector angle changes, and uses them to determine whether they belong to the target action feature. The accuracy and efficiency of identifying the target action feature are significantly improved.

[0043] 3. The present invention introduces channel (time) and spatial attention modules into the CNN convolutional neural network and deploys them in a deep residual network to reduce the number of convolution stacking times, improve the recognition accuracy of target motion features, and avoid the problem of saturation of recognition accuracy due to increased network depth, thereby ensuring accurate recognition and alarm of leaks such as illegal photography in military airports. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is a flow chart of the disclosure behavior identification method based on deep learning of the present invention. DETAILED DESCRIPTION

[0045] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0046] Please see the attached Figure 1 , a method for identifying leakage behavior based on deep learning, comprising the following steps:

[0047] Step 1: Obtain video footage of the military airport captured by surveillance equipment.

[0048] Step 2: Preprocess the surveillance video, pre-train the AlphaPose (multi-person pose recognition) framework, and use the AlphaPose framework to extract the joint points and skeleton features of the people appearing in the surveillance video.

[0049] The preprocessing method of the monitoring video screen is: constructing the target motion trajectory and extracting the target features, and screening the tracking target, that is, the person appearing in the monitoring video screen, based on the target feature comparison of the previous and next frame monitoring video frame sequences.

[0050] When constructing the target's motion trajectory, tracking errors may occur due to the following reasons: ① When external conditions change, such as: the target person is blocked by obstacles, the background of the target person is too complex, the light intensity changes, etc.; ② The target moves too fast, or the tracking instrument has performance problems for a long time, resulting in the target being lost; ③ The monitoring equipment is interfered with by external factors, resulting in deformation of the target's size, aspect ratio, rotation, etc. In order to ensure accurate tracking of the target, that is, to effectively and accurately extract the people appearing in the surveillance video, it is preferred to use a target detection algorithm (Dynamic R-CNN) to detect the target. This target detection algorithm can dynamically change the training strategy as the data changes when training the target detector, and then use the SLATracker (location-aware multi-target tracking) tracking algorithm to track multiple targets for a long time in complex scenes, thereby screening out the tracking target, that is, the people appearing in the surveillance video.

[0051] Traditional image recognition methods rely on 3D depth-sensing cameras, requiring the relocation of dedicated cameras, increasing costs and making it difficult to cover the entire military airport scene. This invention, based on the AlphaPose framework, extracts a person's joints and skeletal features. Compared to traditional image recognition methods, this method eliminates the need for additional or relocated dedicated cameras and reduces costs.

[0052] At present, the main frameworks for extracting the joint points and skeleton features of people in the picture are Openpose (pose estimation) and AlphaPose. The extraction process of the Openpose framework is: first obtain the joint point position, and then obtain the skeleton, so the calculation amount will not increase significantly due to the increase of people in the picture, but it is easy to make mistakes when there is a dense crowd or two people are too close. The extraction process of the AlphaPose framework is: first detect the human body, and then obtain the joint points and skeleton features. It is a top-down algorithm, and will not arbitrarily obtain the joint points of the obscured part of the human body, but only display the visible part. The accuracy and Ap value (Ap value is an evaluation coefficient) are higher than Openpose. Therefore, the present invention adopts the AlphaPose framework to extract the joint points and skeleton features of people in the surveillance video screen.

[0053] The method for extracting the joint points and skeleton features of people appearing in the surveillance video screen by the AlphaPose framework is as follows: extracting activity frames with people in the surveillance video screen, using the AlphaPose framework to extract the joint points of people in the activity frames, and connecting the skeletons of adjacent joints.

[0054] Preferably, the joints of the person include nose, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, buttocks, left hip, right hip, left knee, right knee, left ankle and right ankle, a total of 15 joints.

[0055] Step 3: Based on the extracted joint points and skeleton features of the personnel, the action joint feature vectors of the personnel in the surveillance video are calculated to obtain the action joint features of each person in the surveillance video.

[0056] Specifically, the "local pelvic divergence method" is used to draw several left and right action feature joint vectors with the hip joint point as the starting point and the shoulder, elbow, and wrist joint points as the end points, and the linear displacement and vector angle of the action feature in the active frame are calculated. The changes in the linear displacement and vector angle are used to determine whether the action feature joint vector is consistent with the target action feature.

[0057] Step 4: Import the motion joint features into the model, optimize and establish a deep learning recognition model, and use the deep learning recognition model to classify the motion joint features of each person and identify the target motion features.

[0058] In military airports, common leaking behaviors include shooting and calling for help, and the target action features include single-arm shooting, double-arm shooting, single-arm calling for help, and double-arm calling for help.

[0059] Taking left-arm photography as an example: the neck, left shoulder, left elbow, and left wrist are connected to form a U-shaped action feature joint vector. If the linear displacement and vector angle changes of the action joint feature are consistent with the U-shaped action feature joint vector, the action joint feature is identified as a target action feature and classified as a leaking behavior. If the linear displacement and vector angle changes of the action joint feature are inconsistent with the U-shaped action feature joint vector, the action joint feature is identified as a non-target action feature and classified as a non-leaking behavior. Standard action feature joint vectors are set for single-arm (left and right arm) photography, double-arm photography, single-arm (left and right arm) calling for help, and double-arm calling for help, respectively, for comparison with the action joint features.

[0060] The relative distance between the hip and the shoulder, elbow, and wrist is relatively large. By establishing a motion joint feature vector with the hip joint as the starting point and the shoulder, elbow, and wrist as the midpoint, the feature vector angle changes greatly, which can effectively change the characteristics of small changes in the angles of adjacent joint feature vectors and low recognition efficiency, and is conducive to improving the training effect of the model structure.

[0061] Described step 4 comprises the following sub-steps:

[0062] Step 4.1: Build the model structure, which consists of CNN convolutional neural network, pooling module, attention module and activation function.

[0063] The pooling module includes average pooling and maximum pooling, and the attention module includes channel (temporal) attention and spatial attention.

[0064] The pooling module (also called subsampling or downsampling) is mainly used to reduce the dimension of each feature map. It can reduce the size of the parameter matrix, thereby reducing the number of final outputs, but retaining the most important information. It is often used in neural convolutional networks.

[0065] The step 4.1 includes the following sub-steps:

[0066] Step 4.1.1: Add a channel (time) attention module to the CNN convolutional neural network, and use global average pooling and maximum pooling to utilize different information respectively to summarize spatial features. Average pooling can compress the spatial dimension of the input and learn the range of the joint features of the person's movements; maximum pooling may collect some clues about the joint features of the unique objects' movements.

[0067] In a CNN (convolutional neural network) model, convolutional layers are used to detect local connections between features in the previous layer, while pooling layers are used to fuse similar features to reliably identify leaks. Max pooling and average pooling are two typical pooling methods. In max pooling, the maximum value of a local block is calculated using the pooling units in the feature map. Average pooling calculates the average value of a local block, which reduces the dimensionality of the data representation and creates invariance to small distortions and shifts. Multiple stages of convolution, nonlinearity, and pooling are stacked, followed by convolutional and fully connected layers.

[0068] However, as the network depth increases, the accuracy may saturate and decrease, so simply stacking convolutional layers does not produce good results.

[0069] Step 4.1.2: Introduce the residual block into the model structure to effectively solve the problem of accuracy saturation caused by increasing network depth.

[0070] The method of introducing the residual block is:

[0071] Assuming X is the input of the residual network block and F(X) is the residual mapping function, the original mapping can be expressed as Y = F(X) + X. The residual mapping function can be implemented by stacking convolution, ReLU, channel (time) attention module and spatial attention module, which can approximate the desired function.

[0072] A channel (temporal) attention module is added to fully capture the dependencies between channels. For feature F, a global average pooling and maximum pooling are first performed to obtain two 1×1×C channel descriptions. Next, the two channel descriptions are sent to a two-layer neural network. The number of neurons in the first layer is C / r, the activation function is Relu, and the number of neurons in the second layer is C to obtain two features. This two-layer neural network is shared. Then, the two features are added together and passed through a Sigmoid activation function to obtain the weight coefficient Mc. Finally, the weight coefficient Mc is multiplied by the original feature F to obtain the scaled new feature.

[0073] The Sigmoid function, i.e. f(x)=1 / (1+ex), is the nonlinear action function of a neuron.

[0074] Step 4.1.3: After adding the channel (time) attention module, the spatial attention module is introduced to focus on meaningful action joint features.

[0075] Given an H×W×C feature F', we first perform average pooling and max pooling on each channel dimension to obtain two H×W×1 channel descriptions, which are then concatenated channel-wise. Next, we pass this through an X×X convolutional layer with a Sigmoid activation function to obtain the spatial attention weight coefficient Ms. Finally, we multiply the weight coefficient Ms by the feature F' to obtain the scaled new feature.

[0076] If the channel (time) attention module teaches the model structure what to pay attention to, the spatial attention module will allow the model structure to know where to pay attention. The spatial attention module is a complement to the channel (time) attention module.

[0077] Step 4.2: Initialize the parameters in the model structure, including weights, biases, reduction rates, and learning rates.

[0078] Step 4.3: Based on the joint points and skeleton features, the model structure is trained using forward learning and back propagation.

[0079] In backpropagation, the parameters are updated using the stochastic gradient descent optimization strategy.

[0080] Step 4.4: When the end conditions are met, for example, the number of iterations is sufficient and the loss value is small, and when the set threshold is reached, the model training is completed, making the model parameters optimal, and a deep learning recognition model is obtained that can meet the requirements of leakage behavior identification.

[0081] Step 4.5: Use the deep learning recognition model to identify and classify the motion joint features of each person and identify the target motion features of the leak behavior.

[0082] Step 5: Mark the person corresponding to the target action feature classified as a leak behavior as abnormal, obtain the target abnormal object attributes and its movement trajectory, and report the abnormal target object to the security department for processing.

[0083] The present invention adds an attention module to the CNN convolutional neural network, links time and space, establishes a deep learning recognition model, extracts the joint points of people in the activity frame through the AlphaPose framework, and connects the skeletons of adjacent joints to obtain behavioral information, thereby using the deep learning recognition model to identify leaking behaviors, improving the real-time, accuracy and robustness of leaking behaviors such as illegal photography in military airports. It can also accurately identify leaking behaviors in different scenarios based on different target action characteristics.

[0084] The above are only preferred embodiments of the present invention and are not intended to limit the scope of protection of the invention. Therefore, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for identifying leaks based on deep learning, characterized by: The following steps are involved: Step 1: Obtain the video of the military airport captured by the surveillance equipment; Step 2: Pre-process the surveillance video, pre-train the AlphaPose framework, and use the AlphaPose framework to extract the joint points and skeleton features of the people appearing in the surveillance video; Step 3: Based on the extracted joint points and skeleton features of the personnel, the action joint feature vectors of the personnel in the surveillance video are calculated to obtain the action joint features of each person in the surveillance video; Step 4: Import the motion joint features into the model, optimize and establish a deep learning recognition model, and use the deep learning recognition model to classify the motion joint features of each person and identify the target motion features; Step 5: Mark the person whose target action features are classified as leaking behavior as abnormal, obtain the abnormal target object attributes and movement trajectory, and report the abnormal target object to the security department for processing; Described step 4 comprises the following sub-steps: Step 4.1: Construct the model structure; Step 4.2: Initialize the parameters in the model structure, including weights, biases, reduction rates, and learning rates; Step 4.3: Based on the joint points and skeleton features, the model structure is trained using forward learning and back propagation; Step 4.4: When the end conditions are met, model training is completed, the model parameters are optimized, and a deep learning recognition model is obtained that can meet the requirements of leak behavior identification; Step 4.5: Use the deep learning recognition model to identify and classify the joint features of each person's movements, and identify the target movement features of the leak behavior; The model structure consists of a CNN convolutional neural network, a pooling module, an attention module and an activation function; Step 4.1 includes the following sub-steps: Step 4.1.1: Add a channel attention module to the CNN convolutional neural network and use global average pooling and maximum pooling to utilize different information respectively to summarize spatial features; Step 4.1.2: Introduce the residual block into the model structure; The method of introducing the residual block is: Assuming X is the input of the residual network block and F(X) is the residual mapping function, the original mapping is expressed as Y=F(X)+X; For feature F, global average pooling and maximum pooling are performed in a space to obtain two 1×1×C channel descriptions. The two channel descriptions are then fed into a two-layer neural network. The number of neurons in the first layer is C / r, the activation function is Relu, and the number of neurons in the second layer is C, to obtain two features. This two-layer neural network is shared. The two features are then added together and passed through a Sigmoid activation function to obtain the weight coefficient Mc. Finally, the weight coefficient Mc is multiplied by the original feature F to obtain the scaled new feature. Step 4.1.3: After adding the channel attention module, introduce the spatial attention module to focus on meaningful action joint features; The introduction method of the spatial attention module is: Given a H×W×C feature F', we first perform average pooling and maximum pooling on each channel dimension to obtain two H×W×1 channel descriptions, and then concatenate these two channel descriptions together according to the channel. Then, we pass it through an X×X convolution layer with a Sigmoid activation function to obtain the weight coefficient Ms of spatial attention. Finally, we multiply the weight coefficient Ms by the feature F' to obtain the scaled new feature. If the channel attention module teaches the model structure what to pay attention to, the spatial attention module will allow the model structure to know where to pay attention.

2. The method for identifying leaks based on deep learning according to claim 1, characterized in that: The preprocessing method of the monitoring video screen is: constructing the target motion trajectory and extracting the target features, and screening the tracking target, that is, the person appearing in the monitoring video screen, based on the target feature comparison of the previous and next frame monitoring video frame sequences.

3. The method for identifying leaks based on deep learning according to claim 2 is characterized in that: When screening out tracking targets, the target detection algorithm is used to detect the target, and then the SLATracker tracking algorithm is used to track multiple targets for a long time in complex scenes to screen out tracking targets.

4. The method for identifying leaks based on deep learning according to claim 1, wherein: The method for extracting the joint points and skeleton features of people appearing in the surveillance video screen by the AlphaPose framework is as follows: extracting activity frames with people in the surveillance video screen, using the AlphaPose framework to extract the joint points of people in the activity frames, and connecting the skeletons of adjacent joints.

5. The method for identifying leaks based on deep learning according to claim 4, characterized in that: The joints of the person include nose, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, hip, left hip, right hip, left knee, right knee, left ankle and right ankle.

6. The method for identifying leaks based on deep learning according to claim 1, wherein: In step 3, the "local pelvic divergence method" is used to draw several left and right action feature joint vectors with the hip joint point as the starting point and the shoulder, elbow, and wrist joint points as the end points, and the linear displacement and vector angle of the action feature in the active frame are calculated. The changes in the linear displacement and vector angle are used to determine whether the action feature joint vector is consistent with the target action feature.

7. The method for identifying leaks based on deep learning according to claim 1, wherein: The target action features include single-arm shooting, double-arm shooting, single-arm calling for help, and double-arm calling for help.

Citation Information

Patent Citations

  • Face detection and recognition method on security system

    CN111860393A

  • Abnormal behavior identification method and system based on skeleton extraction

    CN113688797A