Lightweight Behavior Recognition Method and Device for Gas Station Attendants Based on Skeletal Joints
By adopting a lightweight method based on bone joints in the behavior recognition of refuelers, using human object detection network and spatial pyramid pooling structure, the problem of high computing power requirements in traditional deep learning methods is solved, and efficient behavior recognition on embedded devices is achieved.
Patent Information
- Application Number
- CN202211555546.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-06
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-12-06
AI Technical Summary
Traditional fueling staff behavior recognition methods based on deep learning have high requirements for computing power and cannot be effectively deployed on embedded devices with weak computing power.
The lightweight fueling crew behavior recognition method based on bone joints is adopted, and basic feature extraction and anchor frame structure are performed through the human object detection network, combining the spatial pyramid pooling structure and point-by-point convolution to reduce the computational complexity, and the overlap rate is used to determine the sample type for network training.
It realizes efficient refueler behavior recognition on embedded devices with weak computing power, reduces the training and detection inference time of deep learning model, while maintaining high accuracy and real-time.
Smart Images

Figure CN115862136B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image vision deep learning, and particularly relates to a lightweight behavior recognition method and device for fuel attendants based on skeletal joints. Background Art
[0002] Currently, airport behavior supervision mainly adopts two methods: manual supervision or video monitoring. Manual supervision is time-consuming and laborious to handle in practice. At the same time, it is difficult to meet the requirements of real-time and all-weather, and it is also difficult to effectively supervise the behavior of staff, which may lead to relatively large potential safety hazards.
[0003] In video monitoring, video-based behavior recognition algorithms are mainly divided into two categories: traditional algorithms and deep learning methods. Traditional behavior recognition methods require manual design of features to represent behaviors. They are simple to implement but vulnerable to experience, and their accuracy and robustness are average. Deep learning video behavior recognition methods consume a large amount of computing power because they need to extract temporal and spatial features, and are not suitable for deployment on embedded devices with weak computing power. Summary of the Invention
[0004] Therefore, the present invention provides a lightweight behavior recognition method and device for fuel attendants based on skeletal joints, which solves the problems that traditional deep learning-based solutions have high requirements for computing power, consume a large amount of computing power, and are not suitable for deployment on embedded devices with weak computing power.
[0005] To achieve the above object, the present invention provides the following technical solution: A lightweight behavior recognition method for fuel attendants based on skeletal joints, comprising:
[0006] Collecting behavior images to be recognized, decoding the behavior images, and performing size preprocessing on the decoded behavior images;
[0007] Constructing a human target detection network, using the human target detection network to extract basic features from the preprocessed behavior images, and constructing anchor boxes of a preset scale centered on each pixel point on the convolutional feature blocks;
[0008] Traversing the pixel points of the behavior images to obtain predicted anchor boxes, obtaining the overlapping part between the predicted anchor boxes and the real human anchor boxes, determining the sample type according to the overlap rate, and training the human target detection network using the determined sample type;
[0009] Performing detection target association on the anchor boxes corresponding to the previous and next frames obtained by the human target detection network to track the human target and obtain a human motion trajectory detection box;
[0010] Send the human body movement trajectory detection box into the human body target detection network to obtain human body target pose information, where the human body target pose information includes human bone key points, and use the human bone key points to identify the behavior of the fuel dispenser.
[0011] As a preferred solution of the lightweight fuel dispenser behavior recognition method based on bone joints, the human body target detection network includes a Block unit and a SandGlass unit;
[0012] The Block unit expands the dimension through pointwise convolution, and the Block unit extracts channel features through depth convolution;
[0013] The SandGlass unit performs dimension scaling using two depth convolutions and two layers of pointwise convolutions.
[0014] As a preferred solution of the lightweight fuel dispenser behavior recognition method based on bone joints, a spatial pyramid pooling structure is introduced into the human body target detection network. The input feature map passes through max-pooling layers with three preset sizes respectively, and the input feature map and the three pooling outputs are dimensionally concatenated through a shortcut path, and then a convolutional layer is used to fuse and learn the feature information of four different scales.
[0015] As a preferred solution of the lightweight fuel dispenser behavior recognition method based on bone joints, anchor boxes with an overlap rate greater than or equal to 35% are determined as positive samples, and anchor boxes with an overlap rate less than 35% are determined as negative samples, and the determined positive samples and negative samples are used to train the human body target detection network.
[0016] As a preferred solution of the lightweight fuel dispenser behavior recognition method based on bone joints, calculate the overlap degree IOU of the front and rear two-frame anchor boxes a and b obtained by using the human body target detection network. The overlap degree IOU calculation formula is:
[0017] IOU = (Area(a) ∩ Area(b)) / (Area(a) ∪ Area(b))
[0018] In the formula, Area(a) is the area of the region occupied by anchor box a, and Area(b) is the area of the region occupied by anchor box b.
[0019] As a preferred solution of the lightweight fuel dispenser behavior recognition method based on bone joints, the steps of tracking the human body target to obtain the human body movement trajectory detection box include:
[0020] For the current frame detection set D f , for each trajectory t a in the active trajectory set T i, select the anchor box information of the last added trajectory, and calculate the overlap degree IOU between the current position information and all detection boxes in the current frame detection set in turn. If the maximum IOU(d best ,t i ) is greater than or equal to the preset threshold, it is determined that the current detection box belongs to the corresponding added trajectory, and the current detection box is deleted from the current frame detection set D f .
[0021] As an optimal solution of the lightweight behavior recognition method for fueling workers based on skeletal joints, if the maximum IOU(d best ,t i ) is not greater than or equal to the preset threshold, calculate the similarity S between the current detection box and the color histogram of the trajectory box. If S is greater than the preset value, it is determined that the current detection box belongs to the corresponding added trajectory.
[0022] As an optimal solution of the lightweight behavior recognition method for fueling workers based on skeletal joints, all the remaining detection boxes in the current frame detection set D f are inserted into the active trajectory set T a as the start of a new trajectory;
[0023] After the detection is completed, for each active trajectory t a in the active trajectory set T i , judge whether the condition for tracking completion is satisfied. If the condition for tracking completion is satisfied, transfer it to the trajectory set T f of the end of tracking, and use the trajectory set T f as the detected box of the extracted human motion trajectory.
[0024] As an optimal solution of the lightweight behavior recognition method for fueling workers based on skeletal joints, use the similarity OKS to measure the similarity of human skeletal key points, and update the human pose and skeletal key point information;
[0025] During the update process, assign a tracking ID to the human detection box in each frame, and calculate the similarity OKS of the human skeletal key points between two adjacent frames. The similarity OKS calculation formula is:
[0026]
[0027] In the formula, p represents the label of the person, i represents the label of the skeletal point, represents the Euclidean distance between the labeled joint point and the predicted joint point, is the standard deviation, represents the normalization factor for the i-th skeletal key point; δ(v pi =1) indicates that the i-th skeletal key point of the p-th human body is visible.
[0028] The present invention also provides a lightweight behavior recognition device for gas station attendants based on skeletal joints, which adopts the above-mentioned lightweight behavior recognition method for gas station attendants based on skeletal joints, and includes:
[0029] An image acquisition and processing module, configured to acquire a behavior image to be recognized, decode the behavior image, and perform size preprocessing on the decoded behavior image;
[0030] A model construction and processing module, configured to construct a human target detection network, use the human target detection network to extract basic features from the preprocessed behavior image, and construct anchor boxes with a preset scale centered on each pixel point on the convolutional feature block;
[0031] A model training module, configured to traverse the pixel points of the behavior image to obtain predicted anchor boxes, obtain the overlapping part of the predicted anchor boxes and the real human anchor boxes, determine the sample type according to the overlap rate, and use the determined sample type to train the human target detection network;
[0032] A human target tracking module, configured to perform detection target association on the anchor boxes corresponding to the previous and next frames obtained by the human target detection network to track the human target and obtain a human motion trajectory detection box;
[0033] A target behavior recognition module, configured to send the human motion trajectory detection box into the human target detection network to obtain human target pose information, where the human target pose information includes human skeletal key points, and use the human skeletal key points to recognize the behavior of the gas station attendant.
[0034] The present invention has the following advantages: By acquiring the behavior image to be recognized, decoding the behavior image, and performing size preprocessing on the decoded behavior image; constructing a human target detection network, using the human target detection network to extract basic features from the preprocessed behavior image, and constructing anchor boxes with a preset scale centered on each pixel point on the convolutional feature block; traversing the pixel points of the behavior image to obtain predicted anchor boxes, obtaining the overlapping part of the predicted anchor boxes and the real human anchor boxes, determining the sample type according to the overlap rate, and using the determined sample type to train the human target detection network; performing detection target association on the anchor boxes corresponding to the previous and next frames obtained by the human target detection network to track the human target and obtain a human motion trajectory detection box; sending the human motion trajectory detection box into the human target detection network to obtain human target pose information, where the human target pose information includes human skeletal key points, and using the human skeletal key points to recognize the behavior of the gas station attendant. The present invention greatly reduces the time required for deep learning model training and detection inference, and at the same time maintains excellent characteristics such as high accuracy and good real-time performance, and is suitable for application in deep learning embedded terminals with higher cost performance and simpler deployment and installation. Brief Description of the Drawings
[0035] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.
[0036] Figure 1 Schematic flow chart of the lightweight behavior recognition method for fueling workers based on skeletal joints provided in Embodiment 1 of the present invention;
[0037] Figure 2 Schematic diagram of the human target detection network structure in the lightweight behavior recognition method for fueling workers based on skeletal joints provided in Embodiment 1 of the present invention;
[0038] Figure 3 Human detection result diagram in the lightweight behavior recognition method for fueling workers based on skeletal joints provided in Embodiment 1 of the present invention;
[0039] Figure 4 Behavior recognition result diagram of the application scenario in the lightweight behavior recognition method for fueling workers based on skeletal joints provided in Embodiment 1 of the present invention;
[0040] Figure 5 Another behavior recognition result diagram of the application scenario in the lightweight behavior recognition method for fueling workers based on skeletal joints provided in Embodiment 1 of the present invention;
[0041] Figure 6 Schematic diagram of the lightweight behavior recognition device for fueling workers based on skeletal joints provided in Embodiment 2 of the present invention. Specific implementation manners
[0042] The following specific embodiments illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0043] Embodiment 1
[0044] Refer to Figure 1 and Figure 2 , Embodiment 1 of the present invention provides a lightweight behavior recognition method for fueling workers based on skeletal joints, including the following steps:
[0045] S1. Collect the behavior images to be recognized, decode the behavior images, and perform size preprocessing on the decoded behavior images;
[0046] S2, constructing a human target detection network, using the human target detection network to extract basic features of the preprocessed behavior image, and constructing an anchor frame of a preset scale with each pixel point on the convolution feature block as the center;
[0047] S3, traversing the pixel points of the behavior image to obtain a predicted anchor frame, obtaining the overlapping part of the predicted anchor frame and the real human anchor frame, determining the sample type according to the overlap rate, and using the determined sample type to train the human target detection network;
[0048] S4, performing detection target association on the anchor frames corresponding to the two frames before and after acquired by the human target detection network, so as to track the human target and obtain a human motion trajectory detection frame;
[0049] S5. Send the human motion trajectory detection frame to the human target detection network to obtain human target posture information, wherein the human target posture information includes human skeleton key points, and the human skeleton key points are used to perform gas station attendant behavior recognition.
[0050] In this embodiment, in step S1, the behavior image to be identified is collected by the camera, and then decoded by the edge device, and the decoded image is resized to obtain a 608×608 pre-processed behavior image. The resize operation itself is a function in the OpenCV library, which can play the role of scaling the image.
[0051] In this embodiment, in step S2, the human target detection network includes a Block unit and a SandGlass unit; the Block unit performs dimensionality expansion through point-by-point convolution, and the Block unit extracts channel features through deep convolution; the SandGlass unit uses two deep convolutions and two layers of point-by-point convolution for dimensionality scaling.
[0052] Due to the limited computing power and storage resources of embedded devices, edge devices, etc., it is difficult to achieve real-time target detection and effective deployment. It is necessary to sacrifice some performance loss in exchange for faster reasoning speed. Therefore, the human target detection network structure is lightweight optimized:
[0053] Among them, the input size of the human target detection network is 608×608×3, which is then transformed into a 32-channel 304×304 output after convolution, batch normalization, and ReLU activation. It then passes through several feature extraction networks composed of Block units and SandGlass units to obtain high-dimensional image information.
[0054] Specifically, the Block unit first uses pointwise convolution for dimensional expansion to avoid information loss caused by direct dimensional reduction. Then, it uses depthwise convolution to extract features of each channel. Finally, the pointwise convolution adopts the Linear activation function to reduce information loss caused by using the ReLU activation function. By using pointwise convolution and depthwise convolution instead of traditional convolution, the computational complexity is reduced.
[0055] Specifically, the SandGlass unit makes full use of the lightweight feature of depthwise convolution. It uses depthwise convolution twice at the beginning and end, and scales the dimension through two layers of pointwise convolution in the middle, so as to retain more spatial information and improve the classification performance. Finally, high-dimensional information of the image on 128 channels is obtained.
[0056] In this embodiment, a spatial pyramid pooling structure is introduced into the human target detection network. The input feature map passes through max pooling layers with three preset sizes respectively, and the input feature map is concatenated with the three pooling outputs in dimension through a shortcut path. Then, a convolutional layer is used to fuse and learn the feature information of four different scales.
[0057] Specifically, since the input size of the human target detection network is fixed, image distortion and feature information distortion are caused by processing the size of the original behavior image using methods such as cropping and scaling. A spatial pyramid pooling structure is introduced after the feature extractor of the human target detection network. The input feature map passes through max pooling layers with sizes of 3, 5, and 7 respectively, and the input feature map is concatenated with the three pooling outputs in dimension through a short-cut path. Finally, a convolutional layer is used to fuse and learn the feature information of four different scales.
[0058] Among them, considering that the human target is large and relatively simple, in order to reduce the computational complexity of the human target detection network and improve the inference speed, only the feature information of the two layers at the middle 76×76×32 and the end 19×19×128 of the human target detection network is fused. The high-dimensional small-size image information is concatenated with the low-dimensional large-size image information from bottom to top through upsampling, while the low-dimensional large-size image information is concatenated with the high-dimensional small-size image information from top to bottom through downsampling, so that the multi-scale features are repeatedly fused and enhanced with each other, fully understanding the high-level semantic information and low-dimensional information of the feature extraction network, and improving the network expression ability.
[0059] In this embodiment, anchor boxes with an overlap rate greater than or equal to 35% are determined as positive samples, and anchor boxes with an overlap rate less than 35% are determined as negative samples. The determined positive and negative samples are used to train the human target detection network.
[0060] Specifically, by using a human target detection network to extract basic features from the preprocessed behavior images, six anchor boxes with different scales are constructed centered on each pixel point on the convolutional feature block. All the predicted anchor boxes are obtained by traversing all the pixel points of the entire image. The overlapping part between the predicted anchor boxes and the true human anchor boxes is calculated. The anchor boxes with an overlap rate greater than or equal to 35% are determined as positive samples, and the anchor boxes with an overlap rate less than 35% are determined as negative samples to train the human target detection network.
[0061] In this embodiment, in step S4, since the human target detection network can only find the position of the human body in each frame of image, the overlap degree IOU is calculated using the anchor boxes a and b of the previous and next frames obtained by the human target detection network. The calculation formula for the overlap degree IOU is as follows:
[0062] IOU = (Area(a) ∩ Area(b)) / (Area(a) ∪ Area(b))
[0063] In the formula, Area(a) is the area of the region occupied by anchor box a, and Area(b) is the area of the region occupied by anchor box b.
[0064] In this embodiment, the steps for obtaining the detection frame of the human body movement trajectory by tracking the human target are as follows:
[0065] Among them, let D0, D1, D... F-1 respectively represent the detection images of the 0th, 1st,..., F - 1th frames, and d0, d1, d... N-1 respectively represent N human targets in each frame of the detection image; T a represents the set of active trajectories, which is composed of the anchor boxes of the human targets that are still being tracked; T f represents the set of object trajectory sets that have been tracked, which is composed of the human target boxes that have been tracked.
[0066] Specifically, for the current frame detection set D f , for each trajectory t a in the set of active trajectories T i , the anchor box information of the last added trajectory is selected, and the overlap degree IOU between the current position information and all the detection boxes in the current frame detection set is calculated in turn. If the maximum IOU(d best , t i ) is greater than or equal to the preset threshold σ IOU (0.25), it is determined that the current detection box belongs to the corresponding added trajectory, and the current detection box is deleted from the current frame detection set D f .
[0067] Since relying solely on the size of the overlap degree IOU cannot well handle complex situations, in order to avoid the situation of missed detection due to setting a fixed threshold σ IOU :
[0068] If the maximum IOU(d best , t i ) is not greater than or equal to the preset threshold, calculate the similarity S between the current detection box and the color histogram of the trajectory box. If S is greater than the preset value of 0.3, it is determined that the current detection box belongs to the corresponding added trajectory.
[0069] When neither of the above two points is satisfied, then determine whether the highest score at the historical position in this trajectory is greater than the threshold σ h and whether the appearance time of this trajectory is greater than t min (conditions for tracking completion), determine whether the object has been tracked. If so, move the trajectory t i from T a to Tf.
[0070] Specifically, for all the remaining detection boxes in the current frame detection set D f , insert them as the start of new trajectories into the active trajectory set T a ; after all detections are completed, for each active trajectory t a in the active trajectory set T i , determine whether the conditions for tracking completion are met. If the conditions for tracking completion are met, transfer it to the trajectory set T f for tracked completion, and use the trajectory set T f for the detected bounding boxes of the human body movement trajectories extracted.
[0071] In this embodiment, for the detected bounding boxes T f of the human body movement trajectories extracted, it is necessary to send them into a lightweight human target detection network to obtain pose information. The human target detection network uses a top-down method to supplement the human detection bounding boxes using optical flow estimation to reduce missed detections, and then after cropping the detected human target regions, input them into the pose estimation network for two-dimensional pose estimation to obtain the skeletal key points of the human body.
[0072] Specifically, use the NMS method to unify the human detection bounding boxes output by human tracking and the human detection bounding boxes based on optical flow estimation. Then perform appropriate cropping and Resize operations on the bounding boxes to retain as little irrelevant background information as possible, and then perform pose estimation.
[0073] Specifically, to reduce the attribution of detection bounding boxes of different people to the same trajectory, especially when the movement routes of two humans cross and overlap, this kind of misdetection will occur. Use the similarity OKS to measure the similarity degree of human key points and continuously update the human pose and skeletal key point information.
[0074] Among them, the update strategy is as follows: First, assign a unique ID to the human detection box in each frame to facilitate tracking, calculate the similarity OKS of human bone points between two adjacent frames, and those with large similarity correspond to the same ID. The definition of OKS is as follows: Use the similarity OKS to measure the similarity of human bone key points and update the human pose and bone key point information;
[0075] During the update process, assign a tracking ID to the human detection box in each frame, calculate the similarity OKS of human bone key points between two adjacent frames, and the calculation formula of the similarity OKS is:
[0076]
[0077] In the formula, p represents the label of the person, and i represents the label of the bone point. represents the Euclidean distance between the labeled joint point and the predicted joint point. is the standard deviation. represents the normalization factor for the i-th bone key point; δ(v pi =1) indicates that the i-th bone key point of the p-th human body is visible.
[0078] Among them, the larger σ is, the more difficult it is to label the key point. OKS is in the range of [0, 1], and the closer it is to 1, the more similar the two are.
[0079] See Figure 3 、 Figure 4 and Figure 5 , in this embodiment, the human target detection network is mainly responsible for extracting the position information of 17 key points such as the two eyes, nose, two ears, two shoulders, two elbows, two wrists, two thighs, two knees, and two ankles of the human body. Then, taking a 10-frame sequence as a unit, input it into the final classification network. For the command scenario of airport refueling, a series of actions are divided into three categories: fingers, bending over, and others. Taking 10 frames as a unit, the position information (x, y) of all joint points is spanned into a one-dimensional feature vector 1*(10*17*2), a total of 340 features F = [f1, f2.f 340 , input it into the fully connected classification network of bone key points, perform dimensional amplification, extract detailed features, and finally divide them into three categories: fingers, bending, and others.
[0080] In summary, the present invention collects behavior images to be recognized, decodes the behavior images, and performs size preprocessing on the decoded behavior images; constructs a human target detection network, uses the human target detection network to extract basic features from the preprocessed behavior images, and constructs anchor boxes of a preset scale centered on each pixel point on the convolutional feature block; traverses the pixel points of the behavior images to obtain predicted anchor boxes, obtains the overlapping part between the predicted anchor boxes and the real human anchor boxes, determines the sample type according to the overlap rate, and uses the determined sample type to train the human target detection network; performs detection target association on the anchor boxes corresponding to the previous and current frames obtained by the human target detection network to track the human target and obtain a human motion trajectory detection box; sends the human motion trajectory detection box into the human target detection network to obtain human target pose information, where the human target pose information includes human skeleton key points, and uses the human skeleton key points to recognize the behavior of the fuel dispenser. For the current frame detection set D f , for each trajectory t a in the active trajectory set T i , select the anchor box information of the last added trajectory, and calculate the overlap degree IOU between the current position information and all detection boxes in the current frame detection set in turn. If the maximum IOU(d best , t i ) is greater than or equal to the preset threshold σ IOU (0.25), it is determined that the current detection box belongs to the corresponding added trajectory, and the current detection box is deleted from the current frame detection set D f . To avoid the situation of missed detection due to setting a fixed threshold σ IOU : If the maximum IOU(d best , t i ) is not greater than or equal to the preset threshold, calculate the similarity S between the current detection box and the color histogram of the trajectory box. If S is greater than the preset value 0.3, it is determined that the current detection box belongs to the corresponding added trajectory. When neither of the above two points is satisfied, it is judged whether the highest score at the historical position in the trajectory is greater than the threshold σ h and whether the appearance time of the trajectory is greater than t min (tracking completion condition), determine whether the object has been tracked. If so, move the trajectory t i from T a to T f . For all the remaining detection boxes in the current frame detection set D f , they are inserted as the start of a new trajectory into the active trajectory set T a ; when all detections are completed, for each active trajectory t a in the active trajectory set T i , judge whether the tracking completion condition is satisfied. If the tracking completion condition is satisfied, transfer it to the tracked end trajectory set T fIn it, the set of trajectories T where the tracking ends f is used as the detected bounding box of the human motion trajectory extracted. To reduce the attribution of detected bounding boxes of different persons to the same trajectory, especially when the motion routes of two humans cross and overlap, such misdetection will occur. The similarity OKS is used to measure the similarity degree of human key points, and the human pose and bone key point information are continuously updated. The present invention greatly reduces the time required for training the deep learning model and detecting and inferring, and at the same time maintains excellent characteristics such as high precision and good real-time performance, and is suitable for application in a deep learning embedded terminal with higher cost performance and simpler deployment and installation.
[0081] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server, etc. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0082] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0083] Embodiment 2
[0084] See Figure 6 , Embodiment 2 of the present invention also provides a lightweight behavior recognition device for fueling workers based on bone joints, adopting the lightweight behavior recognition method for fueling workers based on bone joints in the above embodiment, including:
[0085] An image acquisition and processing module 1, configured to acquire a behavior image to be recognized, decode the behavior image, and perform size preprocessing on the decoded behavior image;
[0086] A model construction and processing module 2, configured to construct a human target detection network, use the human target detection network to perform basic feature extraction on the preprocessed behavior image, and construct an anchor box with a preset scale centered on each pixel point on the convolutional feature block;
[0087] The model training module 3 is used to traverse the pixel points of the behavior image to obtain predicted anchor boxes, obtain the overlapping part of the predicted anchor boxes and the real human body anchor boxes, determine the sample type according to the overlap rate, and use the determined sample type to train the human target detection network;
[0088] The human target tracking module 4 is used to perform detection target association on the anchor boxes corresponding to the front and rear frames obtained by the human target detection network to track the human target and obtain the human motion trajectory detection box;
[0089] The target behavior recognition module 5 is used to send the human motion trajectory detection box into the human target detection network to obtain human target pose information, where the human target pose information includes human bone key points, and use the human bone key points to recognize the behavior of the fuel dispenser.
[0090] In this embodiment, in the model construction processing module 2, the human target detection network includes a Block unit and a SandGlass unit;
[0091] The Block unit expands the dimension through pointwise convolution, and the Block unit extracts channel features through depth convolution;
[0092] The SandGlass unit performs dimension scaling using two depth convolutions and two layers of pointwise convolutions.
[0093] In this embodiment, in the model construction processing module 2, a spatial pyramid pooling structure is introduced into the human target detection network. The input feature map passes through maximum pooling layers with three preset sizes respectively, and the input feature map and the three pooling outputs are dimensionally concatenated through a shortcut path, and then a convolutional layer is used to fuse and learn the feature information of four different scales.
[0094] In this embodiment, in the model training module 3, the anchor boxes with an overlap rate greater than or equal to 35% are determined as positive samples, and the anchor boxes with an overlap rate less than 35% are determined as negative samples, and the determined positive samples and negative samples are used to train the human target detection network.
[0095] In this embodiment, in the human target tracking module 4, the overlap degree IOU is calculated using the anchor boxes a and b of the front and rear frames obtained by the human target detection network. The formula for the overlap degree IOU is:
[0096] IOU = (Area(a) ∩ Area(b)) / (Area(a) ∪ Area(b))
[0097] In the formula, Area(a) is the area of the region occupied by the anchor box a, and Area(b) is the area of the region occupied by the anchor box b.
[0098] In this embodiment, in the human target tracking module 4, obtaining the human motion trajectory detection frame by tracking the human target includes:
[0099] For the current frame detection set D f , for each trajectory t a in the active trajectory set T i , select the anchor box information of the last added trajectory, and calculate the overlap degree IOU between the current position information and all detection boxes in the current frame detection set in sequence. If the maximum IOU(d best , t i ) is greater than or equal to the preset threshold, it is determined that the current detection box belongs to the corresponding added trajectory, and the current detection box is deleted from the current frame detection set D f .
[0100] If the maximum IOU(d best , t i ) is not greater than or equal to the preset threshold, calculate the similarity S between the current detection box and the color histogram of the trajectory box. If S is greater than the preset value, it is determined that the current detection box belongs to the corresponding added trajectory.
[0101] All the remaining detection boxes in the current frame detection set D f are inserted into the active trajectory set T a as the start of a new trajectory;
[0102] After the detection is completed, for each active trajectory t a in the active trajectory set T i , determine whether the condition for tracking completion is satisfied. If the condition for tracking completion is satisfied, transfer it to the trajectory set T f for tracking completion, and use the trajectory set T f as the extracted human motion trajectory detection frame.
[0103] In this embodiment, in the target behavior recognition module 5, the similarity OKS is used to measure the similarity of human skeleton key points, and the human pose and skeleton key point information are updated;
[0104] During the update process, a tracking ID is assigned to the human detection box in each frame, and the similarity OKS of the human skeleton key points between two adjacent frames is calculated. The similarity OKS calculation formula is:
[0105]
[0106] In the formula, p represents the label of the person, i represents the label of the skeleton point, represents the Euclidean distance between the labeled joint point and the predicted joint point, is the standard deviation, represents the normalization factor for the i-th skeleton key point; 6(vpi = 1) indicates that the i-th skeletal key point of the p-th human body is visible.
[0107] It should be noted that for the information interaction, execution process, etc. among the above-mentioned device modules, since they are based on the same concept as the method embodiment in Embodiment 1 of the present application, the technical effects brought by them are the same as those of the method embodiment of the present application. For specific content, reference can be made to the description in the method embodiment shown above in the present application, and details will not be elaborated here.
[0108] Embodiment 3
[0109] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which program codes of a lightweight behavior recognition method for fueling workers based on skeletal joints are stored. The program codes include instructions for executing the lightweight behavior recognition method for fueling workers based on skeletal joints in Embodiment 1 or any possible implementation manner thereof.
[0110] The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center integrating one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state disk (SSD)), etc.
[0111] Embodiment 4
[0112] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;
[0113] The processor and the memory complete communication with each other through a bus; the memory stores program instructions executable by the processor, and the processor can execute the lightweight behavior recognition method for fueling workers based on skeletal joints in Embodiment 1 or any possible implementation manner thereof by invoking the program instructions.
[0114] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software codes stored in the memory. The memory can be integrated in the processor or can exist independently outside the processor.
[0115] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.).
[0116] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program code executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.
[0117] Although the present invention has been described in detail above with general descriptions and specific embodiments, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A lightweight behavior recognition method for fueling workers based on skeletal joints, characterized in that, Including: Collect the behavior images to be recognized, decode the behavior images, and perform size preprocessing on the decoded behavior images; Construct a human target detection network, use the human target detection network to extract basic features from the preprocessed behavior images, and construct anchor boxes with a preset scale centered on each pixel point on the convolutional feature block; Traverse the pixel points of the behavior image to obtain predicted anchor boxes, obtain the overlapping part of the predicted anchor boxes and the real human anchor boxes, determine the sample type according to the overlap rate, and use the determined sample type to train the human target detection network; Perform detection target association on the anchor boxes corresponding to the previous and next frames obtained by the human target detection network to track the human target and obtain a human motion trajectory detection box; Send the human motion trajectory detection box into the human target detection network to obtain human target pose information, where the human target pose information includes human skeleton key points, and use the human skeleton key points to identify the behavior of the fuel dispenser.
2. The lightweight behavior recognition method of a fuel dispenser based on skeletal joints according to claim 1, characterized in that, The human target detection network includes a Block unit and a SandGlass unit; The Block unit expands the dimension through pointwise convolution, and the Block unit extracts channel features through depth convolution; The SandGlass unit performs dimension scaling using two depth convolutions and two layers of pointwise convolutions.
3. The lightweight behavior recognition method of a fuel dispenser based on skeletal joints according to claim 2, wherein, A spatial pyramid pooling structure is introduced into the human target detection network. The input feature map passes through three maximum pooling layers with preset sizes respectively, and the input feature map and the three pooling outputs are dimensionally concatenated through a shortcut path, and then a convolutional layer is used to fuse and learn the feature information of four different scales.
4. The lightweight behavior recognition method for fuel dispensers based on skeletal joints according to claim 1, characterized in that, Determine the anchor boxes with an overlap rate greater than or equal to 35% as positive samples, and determine the anchor boxes with an overlap rate less than 35% as negative samples, and use the determined positive samples and negative samples to train the human target detection network.
5. The lightweight behavior recognition method of a fuel dispenser based on skeletal joints according to claim 1, characterized in that Calculate the overlap degree IOU of the anchor box a and the anchor box b of the previous and next frames obtained by the human target detection network. The overlap degree IOU calculation formula is: IOU = (Area(a) ∩ Area(b)) / (Area(a) ∪ Area(b)) In the formula, Area(a) is the area of the region occupied by the anchor box a, and Area(b) is the area of the region occupied by the anchor box b.
6. The lightweight behavior recognition method of a fuel dispenser based on skeletal joints according to claim 5, characterized in that The steps of tracking the human target to obtain the human motion trajectory detection box include: For the current frame detection set D f , for each trajectory t a in the active trajectory set T i , select the anchor box information of the last added trajectory, and calculate the overlap degree IOU between the current position information and all detection boxes in the current frame detection set in turn. If the maximum IOU(d best , t i ) is greater than or equal to the preset threshold, it is determined that the current detection box belongs to the corresponding added trajectory, and the current detection box is deleted from the current frame detection set D f .
7. The lightweight behavior recognition method for fuel dispensers based on skeletal joints according to claim 6, characterized in that If the maximum IOU(d best ,t i ) is less than the preset threshold, calculate the similarity S between the current detection box and the color histogram of the trajectory box. If S is greater than the preset value, it is determined that the current detection box belongs to the corresponding added trajectory.
8. The lightweight behavior recognition method of a fuel dispenser based on skeletal joints according to claim 7, characterized in that, Insert all the remaining detection boxes in the current frame detection set D f as the start of new tracks into the active track set T a ; After the detection is completed, for the set of active trajectories T a for each active trajectory t i , determine whether the condition for the end of tracking is satisfied. If the condition for the end of tracking is satisfied, transfer it to the set of trajectories T f in which the set of trajectories T f is used as the detected bounding box of the human motion trajectory to be extracted.
9. The lightweight behavior recognition method for fuel dispensers based on skeletal joints according to claim 8, characterized in that, Use the similarity OKS to measure the similarity of the human skeleton key points, and update the human pose and skeleton key point information; During the update process, assign a tracking ID to the human detection box in each frame, and calculate the similarity OKS of the human skeleton key points between adjacent frames. The similarity OKS calculation formula is: Where p represents the label of a person and i represents the label of a skeletal point, represents the Euclidean distance between the labeled joint point and the predicted joint point, is the standard deviation, represents the normalization factor for the i-th skeletal key point; δ(v pi = 1) indicates that the i-th skeletal key point of the p-th human body is visible.
10. A lightweight fuel dispenser behavior recognition device based on skeletal joints, which adopts the lightweight fuel dispenser behavior recognition method according to any one of claims 1 to 9, characterized in that, Including: An image acquisition and processing module for collecting the behavior images to be recognized, decoding the behavior images, and performing size preprocessing on the decoded behavior images; A model construction and processing module for constructing a human target detection network, using the human target detection network to extract basic features from the preprocessed behavior images, and constructing anchor boxes with a preset scale centered on each pixel point on the convolutional feature block; A model training module, which is used to traverse the pixel points of the behavior image to obtain predicted anchor boxes, acquire the overlapping part of the predicted anchor boxes and the true human body anchor boxes, determine the sample type according to the overlap rate, and use the determined sample type to train the human target detection network; A human target tracking module, which is used to perform detection target association on the anchor boxes corresponding to the previous and next frames obtained by the human target detection network, so as to track the human target to obtain a human motion trajectory detection box; A target behavior recognition module, which is used to send the human motion trajectory detection box into the human target detection network to obtain human target pose information, where the human target pose information includes human bone key points, and use the human bone key points to recognize the behavior of the fuel dispenser.
Citation Information
Patent Citations
Twin network target tracking method based on inverse residual error
CN113436227A
Information processing method and terminal device
EP3709224A1