Model construction method, video monitoring method and system for multi-granularity behavior recognition
By introducing stacked hourglass networks and convolutional neural networks into the video surveillance system, a multi-grained behavior recognition model is built, which solves the problem of insufficient manual monitoring of existing video surveillance systems in security, real-time identification and alarm of unsafe behaviors within the camera perspective range, and improves the security and efficiency of the surveillance system.
Patent Information
- Application Number
- CN202210749923.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-06-29
AI Technical Summary
In the security of existing video surveillance systems, there are problems such as limited manual monitoring duration, visual fatigue, untimely processing of manual monitoring results, high labor costs, and low alarm accuracy, making it difficult to realize real-time monitoring of unsafe behaviors within the camera's perspective range.
Intelligent analysis technology is introduced, and personnel posture information in video time series is extracted based on stacked hourglass network, combined with the calculation error of convolutional neural network, and multi-grained behavior recognition model is constructed. By fusing the skeleton information, full-body image and head area image recognition model, abnormal judgment of personnel behavior is achieved.
Real-time identification and alarm of unsafe behaviors within the camera's perspective range is realized, and the active monitoring capabilities of the video surveillance system are improved, and potential accident hazards are discovered in a timely manner, ensuring safety and prevention.
Smart Images

Figure CN115188069B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of security monitoring, and more particularly, to a method for constructing a multi-granularity behavior recognition model, a video monitoring method, and a system. Background Art
[0002] With the development and progress of communication technologies and hardware, the development of video monitoring systems has been very rapid, and now a relatively complete monitoring system has been established. Currently, the widely used network video monitoring system consists of five major parts: camera, transmission, control, display, and recording registration. The camera collects video data, which is transmitted through the network to the control host for video processing, display, and recording, and finally the collected video is saved to the storage device.
[0003] The current human-centered video monitoring mode has problems such as limited duration of manual monitoring for the monitoring of security unsafe behaviors, decreased attention due to visual fatigue, untimely processing of manual monitoring results, high labor costs, and low alarm accuracy.
[0004] Therefore, it is necessary to develop a method for constructing a multi-granularity behavior recognition model, a video monitoring method, and a system to break through the time and space limitations, understand the real-time security within the camera's perspective in real time, and realize the monitoring of unsafe behaviors within the camera's perspective at any time and anywhere.
[0005] The information disclosed in the background art section of the present invention is only intended to deepen the understanding of the general background art of the present invention, and should not be regarded as an admission or any form of suggestion that this information constitutes the prior art known to those skilled in the art. Summary of the Invention
[0006] The present invention proposes a method for constructing a multi-granularity behavior recognition model, a video monitoring method, and a system. By introducing intelligent analysis technology into the traditional video monitoring system, the stacked hourglass network is used to extract the pose information of people in the video time series; then, the convolutional neural network calculation error of the detailed images of the local regions of the human body in the video frame is combined to jointly determine whether the behavior actions of people within the camera range are abnormal.
[0007] In a first aspect, an embodiment of the present disclosure provides a method for constructing a multi-granularity behavior recognition model, including:
[0008] Setting multiple unsafe behaviors and their corresponding action classifications and multi-granularity features of human postures to obtain a training data set, where the multi-granularity features include skeleton information, full-body images of human actions, and head region images;
[0009] Training a network according to the human posture to obtain a human posture detection network model, and further obtaining a human posture recognition line drawing result;
[0010] Based on the line drawing results of the human body posture recognition, network training is respectively performed on the multi-granularity features in the training dataset to respectively obtain a skeleton information recognition model, a whole body image recognition model, and a head region image recognition model;
[0011] Fuse the skeleton information recognition model, the whole body image recognition model, and the head region image recognition model to obtain a multi-granularity behavior recognition model.
[0012] Preferably, network training is performed according to the human body posture to obtain a human body posture detection network model. Furthermore, obtaining the human body posture recognition line drawing results includes:
[0013] Detect the human body region from video frame I, input it into the symmetric spatial transformation network S, obtain the human body region box and save it in the personnel whole body image I b ;
[0014] Input the human body region box into the single-person pose estimator P to obtain a generated heatmap and the ground truth heatmap H for each skeleton point kn , calculate the error L between the generated heatmap and the ground truth heatmap H kn ; m ;
[0015] Extract the center of the joint points blurred by Gaussian in the generated heatmap to obtain a set of reconstructed joint point coordinates
[0016] Input the human body region box into the fixed-weight single-person pose estimator P' to calculate the generated heatmap and the set of reconstructed joint point coordinates ; s ;
[0017] Accumulate the errors L m and L s to obtain the error L of the symmetric spatial transformation network S G , and optimize the symmetric spatial transformation network S by the gradient descent method;
[0018] Remap the human body posture back to the original image through the spatial inverse transformation network D for the optimized set of joint point coordinates to obtain the human body posture recognition line drawing results;
[0019] Optimize the spatial inverse transformation network D by the gradient descent method to obtain a human body posture detection network model based on a stacked hourglass network.
[0020] Preferably, network training is performed on the skeleton information in the training dataset to obtain a skeleton information recognition model, including:
[0021] According to the set of joint point coordinates Calculate the relative positions in the video frame and input them into the bidirectional long short-term memory neural network L;
[0022] Take the output vector of the bidirectional long short-term memory neural network L as the input of the convolutional neural network and output the predicted action classification r1 of the human behavior;
[0023] Calculate the action prediction error L between r1 and the actual action classification r1 ;
[0024] Accumulate the action prediction error L r1 , obtaining the accumulated error L CNN-LSTM , and optimize the neural network by the gradient descent method to obtain the skeleton information recognition model.
[0025] Preferably, network training is performed on the full-body images of human actions in the training dataset to obtain a full-body image recognition model, including:
[0026] Input the full-body image I of the human action b into the image transformer S1 for transformation processing, and the result is I b ';
[0027] Input I b ' into the convolutional neural network C1 to obtain the predicted action classification r2;
[0028] Calculate the error L between the predicted action classification r2 and the actual action classification r2 ;
[0029] Accumulate the error L r2 to obtain the error L of the convolutional neural network C1 img1 , and optimize it by the gradient descent method to obtain the full-body image recognition model.
[0030] Preferably, network training is performed on the head region images in the training dataset to obtain a head region image recognition model, including:
[0031] According to the skeleton point information, intercept the raincloth region image I from the video frame I of the unsafe behavior through OpenCV a ;
[0032] Input I a into the image transformer T for transformation processing, and the result is I a ';
[0033] Input I aInput the convolutional neural network W to obtain the predicted action classification r3;
[0034] Calculate the error Lr3 between the predicted action classification r3 and the actual action classification;
[0035] Accumulate the error L r3 Obtain the error L of the convolutional neural network W img2 , and optimize it by the gradient descent method to obtain the head region image recognition model.
[0036] In a second aspect, the embodiments of the present disclosure further provide a video surveillance method based on multi-granularity behaviors, which is characterized by including:
[0037] Obtain the human body pose recognition drawing result based on the human body pose detection network model;
[0038] According to the human body pose recognition drawing result, perform action classification recognition and error calculation respectively through the skeleton information recognition model, the whole body image recognition model, and the head region image recognition model to obtain 3 action categories and their corresponding errors;
[0039] Judge whether the 3 action categories are the same. If they are the same, use this action category as the action classification recognition result. If they are different, use the action category with the smallest error as the action classification recognition result.
[0040] Preferably, performing action classification recognition and error calculation through the skeleton information recognition model includes:
[0041] According to the human body pose recognition drawing result, obtain the joint point skeleton information;
[0042] Input the joint point skeleton information into the skeleton information recognition model for error calculation to obtain the action category based on the joint point skeleton information.
[0043] Preferably, performing action classification recognition and error calculation through the whole body image recognition model includes:
[0044] Intercept the human body whole body region image according to the human body region to obtain the whole body image;
[0045] Input the whole body image into the whole body image recognition model for error calculation to obtain the action category based on the whole body image.
[0046] Preferably, performing action classification recognition and error calculation through the head region image recognition model includes:
[0047] Intercept the head region image according to the joint point skeleton information;
[0048] Input the head region image into the head region image recognition model for error calculation to obtain the action category based on the head region image.
[0049] In a third aspect, an embodiment of the present disclosure further provides a video surveillance system based on multi-granularity behavior, including a network camera, a client, and a server. Among them, the client is communicatively connected to the network camera and the server.
[0050] The client obtains the streaming video information of the network camera and sends it to the server.
[0051] The server obtains the human pose recognition drawing result based on the human pose detection network model and returns it to the client.
[0052] The client reads the human pose recognition drawing result and converts it into a playable video frame for local display on the client.
[0053] According to the human pose recognition drawing result, the server identifies the action classification recognition result based on the multi-granularity behavior recognition model and returns it to the client.
[0054] The client reads the action classification recognition result, converts it into an integer array, and outputs a prompt for the corresponding action classification to warn of unsafe behaviors.
[0055] Its beneficial effects are as follows: The present invention realizes the recognition of personnel actions without being restricted by time and location through the means of remote server connection and artificial intelligence, improves the active monitoring ability of the video surveillance system for unsafe behaviors within the monitoring range of the camera, discovers potential accident hazards in a timely manner, identifies the unsafe behaviors of personnel based on the human pose information, accurately discovers and issues alarm prompts, and takes precautions to ensure safety prevention.
[0056] The methods and devices of the present invention have other characteristics and advantages, which will be obvious from the accompanying drawings incorporated herein and the subsequent detailed implementation manners, or will be described in detail in the accompanying drawings incorporated herein and the subsequent detailed implementation manners. These accompanying drawings and detailed implementation manners are jointly used to explain the specific principles of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] By describing the exemplary embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present invention will become more obvious. Among them, in the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.
[0058] Figure 1 The flowchart showing the steps of a model construction method for multi-granularity behavior recognition according to an embodiment of the present invention.
[0059] Figure 2 Shows a schematic diagram of a human pose detection network model based on a stacked hourglass network according to an embodiment of the present invention.
[0060] Figure 3 Shows a schematic structural diagram of a stacked hourglass according to an embodiment of the present invention.
[0061] Figure 4 Shows a flowchart of the steps of a video surveillance method based on multi-granularity behavior according to an embodiment of the present invention.
[0062] Figure 5 Shows a schematic diagram of identifying an action classification recognition result according to a multi-granularity behavior recognition model according to an embodiment of the present invention.
[0063] Figure 6 Shows a monitoring flowchart of a video surveillance system based on multi-granularity behavior according to an embodiment of the present invention.
[0064] Figure 7 Shows a schematic diagram of the network interaction process of a video surveillance system based on multi-granularity behavior according to an embodiment of the present invention. Detailed implementation manners
[0065] The preferred embodiments of the present invention will be described in more detail below. Although the preferred embodiments of the present invention are described below, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein.
[0066] To facilitate understanding of the solutions and effects of the embodiments of the present invention, three specific application examples are given below. Those skilled in the art should understand that this example is only for facilitating the understanding of the present invention, and any specific details are not intended to limit the present invention in any way.
[0067] Example 1
[0068] Figure 1 Shows a flowchart of the steps of a model construction method for multi-granularity behavior recognition according to an embodiment of the present invention.
[0069] As Figure 1As shown in the figure, the method for constructing a multi-granularity behavior recognition model includes: Step 101, setting multiple unsafe behaviors and their corresponding action classifications and multi-granularity features of human postures to obtain a training data set. The multi-granularity features include skeleton information, full-body images of human actions, and head region images; Step 102, performing network training based on human postures to obtain a human posture detection network model, and further obtaining a human posture recognition drawing result; Step 103, based on the human posture recognition drawing result, performing network training on the multi-granularity features in the training data set respectively to obtain a skeleton information recognition model, a full-body image recognition model, and a head region image recognition model; Step 104, fusing the skeleton information recognition model, the full-body image recognition model, and the head region image recognition model to obtain a multi-granularity behavior recognition model.
[0070] In one example, performing network training based on human postures to obtain a human posture detection network model, and further obtaining a human posture recognition drawing result includes:
[0071] Detecting the human body region from video frame I and inputting it into the symmetric space transformation network S to obtain a human body region box and saving it in the full-body image I of the person b ;
[0072] Inputting the human body region box into the single-person pose estimator P to obtain a generated heat map and the ground truth heat map H for each skeleton point kn , calculating the error L between the generated heat map kn and the ground truth heat map H m ;
[0073] Extracting the center of the joint points blurred by Gaussian in the generated heat map to obtain a set of reconstructed joint point coordinates
[0074] Inputting the human body region box into the fixed-weight single-person pose estimator P' to calculate the generated heat map and the set of reconstructed joint point coordinates and calculating the error L s ;
[0075] Accumulating the errors L m and L s to obtain the error L of the symmetric space transformation network S G , and optimizing the symmetric space transformation network S by the gradient descent method;
[0076] Remapping the human posture back to the original image through the space inverse transformation network D for the optimized set of joint point coordinates to obtain a human posture recognition drawing result;
[0077] Optimize the spatial inverse transformation network D by the gradient descent method to obtain a human pose detection network model based on the stacked hourglass network.
[0078] In one example, network training is performed on the skeleton information in the training dataset, and obtaining the skeleton information recognition model includes:
[0079] According to the set of joint point coordinates Calculate the relative positions in the video frame and input them into the bidirectional long short-term memory neural network L;
[0080] Take the output vector of the bidirectional long short-term memory neural network L as the input of the convolutional neural network and output the predicted action classification r1 of the human behavior;
[0081] Calculate the action prediction error L between r1 and the actual action classification r1 ;
[0082] Accumulate the action prediction error L r1 to obtain the accumulated error L CNN-LSTM and optimize the neural network by the gradient descent method to obtain the skeleton information recognition model.
[0083] In one example, network training is performed on the full-body images of human actions in the training dataset, and obtaining the full-body image recognition model includes:
[0084] Input the full-body image I of the human action b into the image transformer S1 for transformation processing, and the result is I b ';
[0085] Input I b ' into the convolutional neural network C1 to obtain the predicted action classification r2;
[0086] Calculate the error L between the predicted action classification r2 and the actual action classification r2 ;
[0087] Accumulate the error L r2 to obtain the error L of the convolutional neural network C1 img1 and optimize it by the gradient descent method to obtain the full-body image recognition model.
[0088] In one example, network training is performed on the head region images in the training dataset, and obtaining the head region image recognition model includes:
[0089] According to the skeleton point information, intercept the raincloth region image I from the video frame I of the unsafe behavior through OpenCV a ;
[0090] Input I a into the image transformer T for transformation processing, and the result is Ia ’;
[0091] Input I a ’ into the convolutional neural network W to obtain the predicted action classification r3;
[0092] Calculate the error Lr3 between the predicted action classification r3 and the actual action classification;
[0093] Accumulate the error L r3 to obtain the error L of the convolutional neural network W, and optimize it by the gradient descent method to obtain the head region image recognition model. img2 Specifically, set a variety of unsafe behaviors, and then determine the action classification corresponding to the unsafe behaviors and the multi-granularity features of the human body posture to obtain a training dataset. The multi-granularity features include skeleton information, the full-body image of the human action, and the head region image;
[0094] The human body posture detection network model includes a symmetric spatial transformation network S, a single-person posture estimator P, and a spatial inverse transformation network D. The symmetric spatial transformation network acts as a generator to extract high-quality human body region boxes in the video frame. The single-person posture estimator P acts as a discriminator to identify the human body posture in the human body region box. The spatial inverse transformation network is used to map the estimated human body posture back to the original image coordinates.
[0095] Figure 2
[0096] Figure 2 Shows a schematic diagram of a human body posture detection network model based on a stacked hourglass network according to an embodiment of the present invention.
[0097] Figure 3 Shows a schematic diagram of the structure of a stacked hourglass according to an embodiment of the present invention.
[0098] As Figure 2 、 Figure 3 shown, obtain the video frame I of unsafe actions such as making a phone call and smoking, and the skeleton points marked by the personnel target in the sample as the input of the stacked hourglass network; use YOLO to detect the human body region in the video frame I and input it into the symmetric spatial transformation network S. The points in the human body region are transformed by the vector operation parameters [γ1γ2γ3] generated by the two-dimensional spatial positioning network to obtain the coordinate points of the joint points in the human body region box to make it have translational invariance, where is:
[0099]
[0100] Obtain the human body region box and save it in the full-body image I of the personnel b ; input the human body region box into the single-person posture estimator P to obtain the generated heat map and generate a ground-truth heatmap H for each skeleton point based on the marked skeleton point information X kn , where n represents the nth hourglass unit and k represents the serial number of the joint point in the human body.
[0101] Calculate and generate the heatmap and the ground-truth heatmap H kn to obtain the error L m :
[0102]
[0103] where s ∈ [1, N], and N is the number of joint points 18, indicating the error accumulation for 18 joint points.
[0104] Extract the center of the joint point blurred by Gaussian in the generated heatmap to obtain the reconstructed joint point coordinate set i represents the abscissa of the joint point, and j represents the ordinate of the joint point.
[0105] Add a parallel single-person pose estimator P', input the human body region box, and the weights of this parallel single-person pose estimator P' are fixed. Its purpose is to backpropagate the pose error of center localization into the symmetric spatial transformation network S. Calculate the generated heatmap for each unit in the stacked hourglass and the reconstructed joint point coordinate set to obtain the error:
[0106]
[0107] Accumulate the error L m and L s to obtain the error L of the symmetric spatial transformation network S G , and optimize the symmetric spatial transformation network S by the gradient descent method. Map the optimized joint point coordinate set back to the original image through the spatial inverse transformation network D. The transformation process is as follows:
[0108]
[0109] where, is the image coordinate before transformation, is the image coordinate after transformation, [α1α2α3] is a vector in the two-dimensional space, and [α1α2] = [γ1γ2] -1 , α3 = -1 × [α1α2]γ3;
[0110] Optimize it by the gradient descent method to obtain a human pose detection network model based on the stacked hourglass network, and further obtain the human pose recognition line drawing result.
[0111] Construct a neural network for human pose skeleton information based on convolutional neural network and bidirectional long short-term memory network, including a convolutional neural network C and a bidirectional long short-term memory network L. The convolutional neural network C is used to extract the spatial information of motion, and the branches of the bidirectional long short-term memory network L are used to extract the temporal information of motion.
[0112] Use video frames I of unsafe behaviors such as making phone calls and smoking and the set of joint coordinates as the input; calculate the relative positions in the video frames according to the set of joint coordinates :
[0113]
[0114] where represents the output vector at the relative position s connected to the k-th level H corresponding to the k-th joint point k ; represents the output vector in the relative rate Q connected to the k-th level H corresponding to the k-th joint point k .
[0115] Divide the joint points into global, middle-level, and local ones and input them into the bidirectional long short-term memory neural network L respectively to obtain rich context information of unsafe behaviors. Combine the outputs of the three types of joint points as the final output:
[0116]
[0117] where v K represents the final output vector of the k-th joint point on the bidirectional long short-term memory neural network.
[0118] Use the output vector v K of the bidirectional long short-term memory neural network as the input of the convolutional neural network to extract the spatial information of the human pose estimation joint points and output the predicted action classification r1 of human behavior; calculate the action prediction error L r1 :
[0119]
[0120] where m is the total number of samples, t is the number of action classifications, y is the true labeled action classification, represents the distribution of the true labels, represents the predicted label distribution of the trained model.
[0121] Accumulate the action prediction error L r1 to obtain the accumulated error L CNN-LSTM, optimize the neural network by gradient descent method to obtain a skeleton information recognition model based on convolutional neural network and bidirectional long short-term memory network.
[0122] Construct a full-body image action recognition model based on convolutional neural network, which includes an image transformer S1 and a convolutional neural network C1. The image transformer S1 is used to perform operations such as scaling, rotating, and shearing on the input image, and the convolutional neural network C1 is used to identify the action category to which the image belongs.
[0123] Input the full-body image I of human action b into the image transformer S1 for random processing such as scaling, rotating, and shearing, and the result is I b ’. Input I b ’ into the convolutional neural network C1. After operations such as convolution, max pooling, and ReLU, obtain the predicted action classification r2; calculate the error L between the predicted action classification r2 and the actual action classification r2 :
[0124]
[0125] where m is the total number of samples, t is the number of action classifications, y is the action classification of the true label, represents the distribution of the true label, represents the predicted label distribution of the trained model.
[0126] Accumulate the error L r2 to obtain the error L of the convolutional neural network C1 img1 , and optimize it by gradient descent method to obtain a full-body image recognition model based on convolutional neural network.
[0127] Construct a head region image recognition model based on convolutional neural network, which includes an image transformer T and a convolutional neural network W. The image transformer T is used as a generator to perform operations such as scaling, rotating, and shearing on the input image, and the convolutional neural network W is used as a discriminator to identify the action category to which the image belongs.
[0128] According to the skeleton point information, intercept the rain cloth area image I from the video frame I of unsafe behavior through OpenCV a :
[0129]
[0130] where is the rightmost side of the image intercepted based on the position of the right ear joint point of the person in the frame, is the leftmost side of the image intercepted based on the position of the left ear joint point of the person in the frame, The uppermost side of the image intercepted based on the position of the nose joint point of the person in the frame, the lowermost side of the image intercepted based on the position of the neck joint point of the person in the frame; Input the image transformer T for transformation processing, and the result is I a '; Input I a ' into the convolutional neural network W. After convolution, max pooling, and ReLU operations, the predicted action classification r3 is obtained; Calculate the error Lr3 between the predicted action classification r3 and the actual action classification: a where m is the total number of samples, t is the number of action classifications, y is the action classification of the true label,
[0131]
[0132] represents the distribution of the true label, and represents the predicted label distribution of the trained model.
[0133] Accumulate the error L r3 to obtain the error L img2 of the convolutional neural network W, and optimize it by the gradient descent method to obtain a head region image recognition model based on the convolutional neural network.
[0134] Fuse the skeleton information recognition model, the whole body image recognition model, and the head region image recognition model to obtain a multi-granularity behavior recognition model.
[0135] Example 2
[0136] Figure 4 shows a monitoring flow chart of a video monitoring system based on multi-granularity behavior according to an embodiment of the present invention.
[0137] As Figure 4 shown, the video monitoring method based on multi-granularity behavior is characterized by including:
[0138] Step 201, obtain the human pose recognition drawing result based on the human pose detection network model;
[0139] Step 202, according to the human pose recognition drawing result, perform action classification recognition and error calculation respectively through the skeleton information recognition model, the whole body image recognition model, and the head region image recognition model to obtain 3 action categories and their corresponding errors;
[0140] Step 203, determine whether the 3 action categories are the same. If they are the same, use this action category as the action classification recognition result. If they are different, use the action category with the smallest error as the action classification recognition result.
[0141] In one example, the action classification recognition and error calculation through the skeleton information recognition model include:
[0142] Obtain the joint point skeleton information according to the human pose recognition drawing result;
[0143] Input the joint point skeleton information into the skeleton information recognition model, perform error calculation, and obtain the action category based on the joint point skeleton information.
[0144] In one example, the action classification recognition and error calculation through the whole body image recognition model include:
[0145] Intercept the human whole body area image according to the human body area to obtain the whole body image;
[0146] Input the whole body image into the whole body image recognition model, perform error calculation, and obtain the action category based on the whole body image.
[0147] In one example, the action classification recognition and error calculation through the head area image recognition model include:
[0148] Intercept the head area image according to the joint point skeleton information;
[0149] Input the head area image into the head area image recognition model, perform error calculation, and obtain the action category based on the head area image.
[0150] Specifically, perform pose estimation on the time series video, measure each human detection box in the video using a top-down method, perform human pose recognition on the video frames based on the human pose detection network model, and obtain the human pose recognition drawing result.
[0151] Figure 5 The schematic diagram shows the action classification recognition result according to the multi-granularity behavior recognition model according to an embodiment of the present invention.
[0152] As Figure 5 shown, according to the human pose recognition drawing result, the remote server-side splits the received video stream into a time series in units of M frames, inputs the time series into the multi-granularity behavior recognition model, and performs action classification recognition and error calculation respectively through the skeleton information recognition model, the whole body image recognition model, and the head area image recognition model of the multi-granularity behavior recognition model to obtain 3 action categories and their corresponding errors. The recognition result is a one-dimensional array; specifically:
[0153] Obtain the joint point skeleton information according to the human pose recognition drawing result; input the joint point skeleton information into the skeleton information recognition model, perform error calculation, and obtain the judgment score of the action category based on the joint point skeleton information;
[0154] The whole body region image of the human body is intercepted according to the human body region to obtain a whole body RGB image; the whole body RGB image is input into the whole body image recognition model, and the error is calculated to obtain a judgment score of the action category based on the whole body RGB image;
[0155] The head region image is captured according to the skeleton information of the joint points; the head region image is input into the head region image recognition model, and the error is calculated to obtain a judgment score of the action category based on the head region image.
[0156] Determine whether the three action categories are the same. If they are the same, then the action category is used as the action classification recognition result. If they are different, then the action category with the smallest error is used as the action classification recognition result.
[0157] Example 3
[0158] Figure 6 A monitoring flow chart of a video monitoring system based on multi-granularity behaviors according to an embodiment of the present invention is shown.
[0159] like Figure 6 As shown, the video surveillance system based on multi-granularity behavior includes a network camera, a client and a server, wherein the client and the network camera, and the client and the server are all connected in communication.
[0160] The client obtains the streaming video information of the network camera and sends it to the server;
[0161] The server obtains the human posture recognition line drawing result based on the human posture detection network model and returns it to the client;
[0162] The client reads the human posture recognition line drawing results and converts them into playable video frames for local display on the client;
[0163] According to the human posture recognition and line drawing results, the server recognizes the action classification results based on the multi-granularity behavior recognition model and returns them to the client;
[0164] The client reads the action classification recognition results, converts them into an integer array, and then outputs the prompt of the action classification corresponding to the index, warning of unsafe behavior.
[0165] In one example, based on the human body posture recognition line drawing result, the server-side recognizes the action classification recognition result based on the multi-granularity behavior recognition model, including:
[0166] According to the human posture recognition and line drawing results, the skeleton information recognition model, the whole body image recognition model, and the head area image recognition model are used to perform action classification and error calculation respectively, and three action categories and their corresponding errors are obtained;
[0167] Determine whether the three action categories are the same. If they are the same, use this action category as the action classification recognition result. If they are different, use the action category with the smallest error as the action classification recognition result.
[0168] In one example, action classification recognition and error calculation through the skeleton information recognition model include:
[0169] Obtain joint point skeleton information based on the human pose recognition line drawing result;
[0170] Input the joint point skeleton information into the skeleton information recognition model for error calculation to obtain the action category based on the joint point skeleton information.
[0171] In one example, action classification recognition and error calculation through the full body image recognition model include:
[0172] Intercept the human full body area image according to the joint point skeleton information to obtain the full body image;
[0173] Input the full body image into the full body image recognition model for error calculation to obtain the action category based on the full body image.
[0174] In one example, action classification recognition and error calculation through the head area image recognition model include:
[0175] Intercept the head area image according to the joint point skeleton information;
[0176] Input the head area image into the head area image recognition model for error calculation to obtain the action category based on the head area image.
[0177] Specifically, after constructing the human pose detection network model and the multi-granularity behavior recognition model, a client UI interface with a friendly and convenient operation is pre-designed through PyQt5, and the corresponding interfaces are pre-set.
[0178] The client sends a request to obtain the streaming media video. The camera video source uses the RTSP protocol to transmit the H.264 encoded real-time video stream to a computer under the same local area network. After reading the video frame through OpenCV and converting it into a numpy array, the received video stream is completely sent to the remote server with a public network IP through socket communication for waiting to be processed.
[0179] Figure 7 Shows a schematic diagram of the network interaction of the video surveillance system based on multi-granularity behavior according to an embodiment of the present invention.
[0180] As Figure 7 shown, remote transmission of the video stream can break through the spatial limitation, remotely monitor the real-time video and analyze the human actions therein to achieve intelligent security. The specific steps are as follows:
[0181] When the server side turns on the OpenSSH service, the client must know the IP address of the server side, and the communication port is determined by both parties; the client host uses RTSP (Real Time Streaming Protocol, a protocol for real-time streaming, a protocol for streaming media services) to quickly obtain the real-time video stream of the network camera within the same local area network, and reads the frame (frame_src) in the video stream through OpenCV. The frame size is (F w ,F h ,F d ), converts each frame of video signal into a numpy array, and then constructs a python byte (frame_byte) based on the original data in this array.
[0182] The client creates a socket to initiate a connection request to the server side and waits for the server side to respond. The server side creates a socket and binds the socket to the local IP and port, starts listening for connections, and the server side enters a loop to continuously accept the connection requests from the client.
[0183] After the two sides establish a connection, the client sends the python byte array (frame_byte) containing the original data bytes of the video stream to the server side completely. The server side receives the data sent by the client, and the size should be B bytes. The received data is put into a buffer, all the data in the buffer is read and converted into a one-dimensional array by using numpy, and then the one-dimensional array is changed to the size of the original camera video frame without changing the data content, that is, (F w ,F h ,F d ). At this time, the video frame can be used to play the video stream through OpenCV.
[0184] The server side pushes the received video frame into a public queue, starts a pose recognition thread to perform pose estimation on the time-series video, measures each human detection box in the video by using a top-down method, performs human pose recognition on the video frame based on the human pose detection network model, obtains the human pose recognition drawing result and returns it to the client through socket communication; the client receives the human pose recognition drawing result in the UI interface and converts it into a playable video frame for local display on the client.
[0185] Based on the human pose recognition drawing result, the server side recognizes the action classification recognition result based on the multi-granularity behavior recognition model, sends the action classification recognition result to the client through socket communication, and sends the one-dimensional array containing the list index corresponding to the action result after converting it into a python byte. After receiving it in the UI interface of the client, it is converted into an integer array and then the Chinese expression of the index corresponding action is output to warn of unsafe behaviors.
[0186] The client and the server close the socket, shut down all threads, empty and release the buffer.
[0187] The present invention realizes the identification of human actions without being restricted by time and location through the means of remote server connection and artificial intelligence, improves the active monitoring ability of the video monitoring system for unsafe behaviors within the monitoring range of the camera, discovers potential accident hazards in a timely manner, identifies unsafe behaviors of personnel through human body posture information, accurately discovers and issues alarm prompts, takes precautions before they happen, and ensures safety prevention.
[0188] Those skilled in the art should understand that the purpose of the above description of the embodiments of the present invention is only to exemplarily illustrate the beneficial effects of the embodiments of the present invention, and is not intended to limit the embodiments of the present invention to any of the examples given.
[0189] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
Claims
1. A method for constructing a multi-granularity behavior recognition model, characterized in that, Including: Setting multiple unsafe behaviors, their corresponding action classifications, and multi-granularity features of human postures to obtain a training dataset, where the multi-granularity features include skeleton information, full-body images of human actions, and head region images; Training a network according to the human posture to obtain a human posture detection network model, and further obtaining a human posture recognition line-drawing result; Based on the human posture recognition line-drawing result, respectively training the multi-granularity features in the training dataset through network training to respectively obtain a skeleton information recognition model, a full-body image recognition model, and a head region image recognition model; Fusing the skeleton information recognition model, the full-body image recognition model, and the head region image recognition model to obtain a multi-granularity behavior recognition model; Among them, training a network according to the human posture to obtain a human posture detection network model, and further obtaining a human posture recognition line-drawing result includes: Detect the human body region based on the video frame I, input it into the symmetric spatial transformation network S, obtain the human body region bounding box, and save it in the full-body image I of the person b ; Input the human body region box into a single-person pose estimator P to obtain a generated heatmap and the ground-truth heatmap H for each skeleton point kn , and calculate the error L between the generated heatmap kn and the ground-truth heatmap H m ; Extract the center of the joint points blurred by Gaussian in the generated heat map to obtain a set of reconstructed joint point coordinates Input the human body region box into the fixed-weight single-person pose estimator P’ to calculate and generate a heat map and the set of reconstructed joint point coordinates The error L s ; Accumulated error L m and L s to obtain the error L of the symmetric space transformation network S G , and optimize the symmetric space transformation network S by the gradient descent method; The optimized set of keypoint coordinates The human pose is remapped back to the original image through the spatial inverse transformation network D to obtain the human pose recognition line drawing result; Optimizing the spatial inverse transformation network D through the gradient descent method to obtain a human posture detection network model based on a stacked hourglass network.
2. The method for constructing a multi-granularity behavior recognition model according to claim 1, wherein, Training the skeleton information in the training dataset through network training to obtain a skeleton information recognition model includes: According to the set of joint point coordinates calculate the relative positions in the video frame and input them into the bidirectional long short-term memory neural network L; Taking the output vector of the bidirectional long short-term memory neural network L as the input of the convolutional neural network, and outputting the predicted action classification r1 of the human behavior; Calculate the action prediction error L between r1 and the actual action classification r1 ; Accumulated action prediction error L r1 , to obtain the accumulated error L CNN-LSTM , optimize the neural network by the gradient descent method to obtain the skeleton information recognition model.
3. The method for constructing a multi-granularity behavior recognition model according to claim 1, wherein, Training the full-body images of human actions in the training dataset through network training to obtain a full-body image recognition model includes: Input the full-body human motion image I b into the image transformer S1 for transformation processing, and the result is I b ’; Input I b into the convolutional neural network C1 to obtain the predicted action classification r2; Calculate the error L between the predicted action classification r2 and the actual action classification r2 ; Accumulated error L r2 Obtain the error L of the convolutional neural network C1 img1 and optimize it by the gradient descent method to obtain the full-body image recognition model.
4. The method for constructing a multi-granularity behavior recognition model according to claim 1, wherein, Training the head region images in the training dataset through network training to obtain a head region image recognition model includes: Intercept the image I of the raincloth area from the video frame I of the unsafe behavior according to the skeleton point information through OpenCV a ; Input I a to the input image transformer T for transformation processing, and the result is I a '; Input I a into the convolutional neural network W to obtain the predicted action classification r3; Calculating the error Lr3 between the predicted action classification r3 and the actual action classification; Accumulated error L r3 Obtain the error L of the convolutional neural network W img2 , and optimize it by the gradient descent method to obtain the head region image recognition model.
5. A video surveillance method based on multi-granularity behavior using the method for constructing a multi-granularity behavior recognition model according to any one of claims 1-4, characterized in that, Including: Obtaining a human posture recognition line-drawing result based on the human posture detection network model; According to the human posture recognition line-drawing result, respectively performing action classification recognition and error calculation through the skeleton information recognition model, the full-body image recognition model, and the head region image recognition model to obtain 3 action categories and their corresponding errors; Judging whether the 3 action categories are the same. If they are the same, using this action category as the result of the action classification recognition. If they are different, using the action category with the smallest error as the result of the action classification recognition.
6. The video monitoring method based on multi-granularity behavior according to claim 5, wherein, Performing action classification recognition and error calculation through the skeleton information recognition model includes: According to the human posture recognition line-drawing result, obtaining joint point skeleton information; Inputting the joint point skeleton information into the skeleton information recognition model for error calculation to obtain an action category based on the joint point skeleton information.
7. The video surveillance method based on multi-granularity behavior according to claim 5, wherein, Performing action classification recognition and error calculation through the full-body image recognition model includes: Intercepting a full-body region image of the human body according to the human region to obtain a full-body image; Inputting the full-body image into the full-body image recognition model for error calculation to obtain an action category based on the full-body image.
8. The video surveillance method based on multi-granularity behavior according to claim 6, wherein, Performing action classification recognition and error calculation through the head region image recognition model includes: Intercepting a head region image according to the joint point skeleton information; Inputting the head region image into the head region image recognition model for error calculation to obtain an action category based on the head region image.
9. A video surveillance system based on multi-granularity behavior using the method for constructing a multi-granularity behavior recognition model according to any one of claims 1-4, characterized in that It includes a webcam, a client, and a server. Among them, the client is communicatively connected to the webcam and the server respectively. The client obtains the streaming video information of the webcam and sends it to the server. The server obtains the human pose recognition drawing result based on the human pose detection network model and returns it to the client. The client reads the human pose recognition drawing result and converts it into a playable video frame for local display on the client. According to the human pose recognition drawing result, the server recognizes the action classification recognition result based on the multi-granularity behavior recognition model and returns it to the client. The client reads the action classification recognition result, converts it into an integer array, and outputs a prompt for the action classification corresponding to the index to warn of unsafe behaviors.