Lightweight Behavior Recognition Method and Device for Fuel Attendants Based on Sequence Diagrams
Through a lightweight refueler behavior recognition method based on sequence diagrams, using human object detection network and feature fusion technology, the problem of insufficient real-time and accuracy in traditional methods is solved, and real-time behavior recognition and accurate supervision on embedded devices is realized.
Patent Information
- Application Number
- CN202211555543.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-06
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-12-06
AI Technical Summary
The existing technology is difficult to achieve a balance of real-time and accuracy in airport behavior supervision, especially when mobile device resources are limited, traditional methods are time-consuming and labor-intensive and difficult to supervise the behavior of refueling personnel around the clock, resulting in safety hazards.
A lightweight cheerleader behavior recognition method based on sequence diagram is adopted, high-dimensional information is extracted through the human object detection network, spatial pyramid pooling and feature fusion are performed, and the time fusion feature map is extracted in combination with the convolutional layer and the maximum pooling layer, and behavior classification is used to use the fully connected network and Softmax.
It realizes real-time and accurate identification of the behavior of the fueling crew on embedded devices, reduces calculation consumption and maintenance costs, and is suitable for all-weather supervision.
Smart Images

Figure CN115713810B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image vision deep learning, and particularly relates to a lightweight fuel dispenser behavior recognition method and device based on sequence diagrams. Background Art
[0002] At present, airport behavior supervision mainly adopts two methods: manual supervision or video monitoring. Manual supervision is time-consuming and laborious to handle in practice. At the same time, it is difficult to meet the requirements of real-time and all-weather, and it is also difficult to supervise the behavior of staff in place, which may lead to relatively large potential safety hazards. The command behaviors of airport fueling personnel mainly include pointing upwards with fingers, bending over, etc. It is necessary to identify their behaviors to determine whether the action instructions are completed to reduce the occurrence of accidents.
[0003] At this stage, the computing power and storage resources of mobile devices, edge devices, etc. are limited. Traditional technical solutions are difficult to balance real-time performance and accuracy. Moreover, for the simple case of detecting a relatively regular human body, complex structures and a large number of parameters will cause an increase in cost and power consumption, which is not conducive to 24-hour all-weather supervision. Among them, the sequence-based behavior classification algorithm is mainly divided into two steps. First, intermediate features such as contour maps and skeletal key points are extracted from RGB images, and then the extracted features are classified. However, in actual applications, extracting contour maps will lose spatial fine-grained information, and there is also noise interference from other visual information, resulting in a slow inference speed. Summary of the Invention
[0004] Therefore, the present invention provides a lightweight fuel dispenser behavior recognition method and device based on sequence diagrams to solve the problems of poor accuracy and real-time performance of traditional solutions.
[0005] To achieve the above object, the present invention provides the following technical solutions: A lightweight fuel dispenser behavior recognition method based on sequence diagrams, including:
[0006] Collect the behavior video stream to be recognized, decode the behavior video stream, decompose the decoded behavior video stream into continuous frames, and perform size preprocessing on the decomposed continuous frame images;
[0007] Construct a human target detection network, use the human target detection network to extract high-dimensional information from the preprocessed continuous frame images, perform spatial pyramid pooling after high-dimensional information extraction, use a shortcut path to splice the input feature map and the pooling output in dimension, and then use a convolutional layer to fuse feature information;
[0008] Detect the target association of the anchor boxes corresponding to the previous and next frames obtained by the human target detection network to track the human target and obtain a set of human motion trajectory sequence diagrams;
[0009] Using the set of human motion trajectory sequence diagrams obtained by tracking as the input, the convolutional layer and the max pooling layer are used to extract features from the input set of human motion trajectory sequence diagrams and stack them to obtain a time fusion feature map;
[0010] The time fusion feature map is unfolded into a one-dimensional feature vector fusion space information, and the feature vector is classified using a fully connected network and Softmax to obtain a predetermined behavior type of the fuel dispenser.
[0011] As an optimal solution of the lightweight fuel dispenser behavior recognition method based on sequence diagrams, for the preprocessed continuous frame images, high-dimensional feature images are obtained through two-dimensional convolution, batch normalization, and ReLu activation, and then high-dimensional information is extracted through a feature extractor composed of Block units and SandGlass units for several times.
[0012] As an optimal solution of the lightweight fuel dispenser behavior recognition method based on sequence diagrams, in the feature information fusion process, mutual feature fusion enhancement is performed on the high-dimensional feature information of 20×20×256 and the low-dimensional feature information of 80×80×64, and the output is sent to two detection heads;
[0013] Decouple the detection head to eliminate the conflict in the classification and regression features in object detection; use a 1×1 convolutional layer to reduce the channel dimension, decouple the prediction branch, and then use two 3×3 convolutional layer parallel branches for classification and regression tasks respectively.
[0014] As an optimal solution of the lightweight fuel dispenser behavior recognition method based on sequence diagrams, calculate the overlap degree IOU between the front and rear two frame anchor boxes a and b obtained by using the human object detection network, and the overlap degree IOU calculation formula is:
[0015] IOU = (Area(a) ∩ Area(b)) / (Area(a) ∪ Area(b))
[0016] In the formula, Area(a) is the area of the region occupied by the anchor box a, and Area(b) is the area of the region occupied by the anchor box b.
[0017] As an optimal solution of the lightweight fuel dispenser behavior recognition method based on sequence diagrams, the steps of tracking the human target to obtain the set of human motion trajectory sequence diagrams include:
[0018] Adopt a threshold σ l Filter out the human detection frames with low confidence;
[0019] For the current frame detection set D f , for each trajectory t a in the active trajectory set T i, select the human detection box information of the last added trajectory, and calculate the overlap degree IOU between the current position information and all detection boxes in the current frame detection set in sequence. If the maximum IOU(d best ,t i ) is greater than or equal to the preset threshold, it is determined that the current detection box belongs to the corresponding added trajectory, and the current detection box is deleted from the current frame detection set D f .
[0020] As an optimal solution of the lightweight behavior recognition method of fueling workers based on sequence diagrams, if the maximum IOU(d best ,t i ) is not greater than or equal to the preset threshold, use the aspect ratio and confidence score of the human detection box to calculate the similarity Sim between different human detection boxes in the previous and current frames for re-matching:
[0021] Sim = 1 - 0.5 * abs[(R a +R b ) / R a *R b )] + 0.5 * abs(C a -C b )
[0022] In the formula, R a , R b represent the ratio of the length to the width of the detection box respectively; C a , C b represent the difference in the confidence of the detection box respectively.
[0023] As an optimal solution of the lightweight behavior recognition method of fueling workers based on sequence diagrams, when Sim is less than the preset value, judge whether the highest score of the historical position in the corresponding trajectory is greater than the threshold σ h , and whether the appearance time of the corresponding trajectory is greater than the tracking completion time t min . If the conditions are met, it is determined that the tracking is completed, and the corresponding trajectory t i is moved from the active trajectory set T a to the set T f of the human trajectories that have been tracked.
[0024] The present invention also provides a lightweight behavior recognition device for fueling workers based on sequence diagrams, including:
[0025] A video acquisition and processing module, configured to acquire a behavior video stream to be recognized, decode the behavior video stream, decompose the decoded behavior video stream into consecutive frames, and perform size preprocessing on the decomposed consecutive frame images;
[0026] The model construction processing module is used to construct a human target detection network, extract high-dimensional information from the preprocessed consecutive frame images using the human target detection network, perform spatial pyramid pooling after high-dimensional information extraction, splice the dimensions of the input feature map and the pooling output using a shortcut path, and then perform feature information fusion using a convolutional layer;
[0027] The human target tracking module is used to associate the detection targets of the anchor boxes corresponding to the previous and next frames obtained by the human target detection network to track the human target and obtain a set of human motion trajectory sequence diagrams;
[0028] The feature fusion module is used to take the set of human motion trajectory sequence diagrams obtained by tracking as the input, extract features from the input set of human motion trajectory sequence diagrams using a convolutional layer and a max pooling layer and stack them to obtain a temporal fusion feature map;
[0029] The feature classification processing module is used to expand the temporal fusion feature map into a one-dimensional feature vector to fuse spatial information, classify the feature vector using a fully connected network and Softmax, and obtain a predetermined behavior type of the fuel dispenser.
[0030] As a preferred solution of the lightweight fuel dispenser behavior recognition device based on sequence diagrams, in the model construction processing module, for the preprocessed consecutive frame images, after two-dimensional convolution, batch normalization, and ReLu activation, a high-dimensional feature image is obtained, and then high-dimensional information is extracted through a feature extractor composed of Block units and SandGlass units for several times;
[0031] In the model construction processing module, in the process of feature information fusion, mutual feature fusion enhancement is performed on the high-dimensional feature information of 20×20×256 and the low-dimensional feature information of 80×80×64, and the output is sent to two detection heads;
[0032] In the model construction processing module, decouple the detection heads to eliminate the conflict in the classification and regression features in target detection; use a 1×1 convolutional layer to reduce the channel dimension, decouple the prediction branches, and then use two 3×3 convolutional layer parallel branches for classification and regression tasks respectively.
[0033] Calculate the overlap degree IOU of the previous and next frame anchor boxes a and anchor boxes b obtained by using the human target detection network. The overlap degree IOU calculation formula is:
[0034] IOU = (Area(a) ∩ Area(b)) / (Area(a) ∪ Area(b))
[0035] In the formula, Area(a) is the area of the region occupied by the anchor box a, and Area(b) is the area of the region occupied by the anchor box b.
[0036] As an optimal solution for the lightweight behavior recognition device of fuel dispensers based on sequence diagrams, the human target tracking module includes:
[0037] A filtering sub-module for filtering out human detection frames with low confidence using a threshold σ l ;
[0038] A trajectory processing sub-module for, for the current frame detection set D f , for each trajectory t a in the active trajectory set T i , selecting the human detection frame information of the last added trajectory, calculating the overlap degree IOU between the current position information and all detection frames in the current frame detection set in sequence. If the maximum IOU(d best ,t i ) is greater than or equal to the preset threshold, it is determined that the current detection frame belongs to the corresponding added trajectory, and the current detection frame is deleted from the current frame detection set D f ;
[0039] A similarity calculation sub-module for, if the maximum IOU(d best ,t i ) is not greater than or equal to the preset threshold, calculating the similarity Sim between different human detection frames of the front and rear frames using the aspect ratio and confidence score of the human detection frame for re-matching:
[0040] Sim = 1 - 0.5 * abs[(R a +R b ) / R a *R b )] + 0.5 * abs(C a -C b )
[0041] wherein, R a , R b respectively represent the ratio of the length to the width of the detection frame; C a , C b respectively represent the difference in the confidence of the detection frame;
[0042] A tracking judgment sub-module for, when Sim is less than the preset value, judging whether the highest score of the historical position in the corresponding trajectory is greater than the threshold σ h , and whether the appearance time of the corresponding trajectory is greater than the tracking completion time t min . If the conditions are met, it is determined that the tracking is completed, and the corresponding trajectory t i is moved from the active trajectory set T a to the set T f of human trajectories that have been tracked.
[0043] The present invention has the following advantages: By collecting the behavior video stream to be recognized, decoding the behavior video stream, decomposing the decoded behavior video stream into consecutive frames, and performing size preprocessing on the decomposed consecutive frame images; constructing a human target detection network, using the human target detection network to extract high-dimensional information from the preprocessed consecutive frame images, performing spatial pyramid pooling after high-dimensional information extraction, using a shortcut path to splice the dimensions of the input feature map and the pooling output, and then using a convolutional layer to perform feature information fusion; detecting the target association of the anchor boxes corresponding to the previous and next frames obtained by the human target detection network to track the human target and obtain a set of human motion trajectory sequence diagrams; using the set of human motion trajectory sequence diagrams obtained by tracking as the input, using a convolutional layer and a max pooling layer to extract features from the input set of human motion trajectory sequence diagrams and superimposing them to obtain a time fusion feature map; expanding the time fusion feature map into a one-dimensional feature vector to fuse spatial information, and classifying the feature vector using a fully connected network and Softmax to obtain a predetermined behavior type of a fuel dispenser. The present invention has low computational consumption, a simple structure, is suitable for being deployed on an embedded device, can ensure real-time performance, has good accuracy at the same time, and can reduce the operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained according to the provided drawings.
[0045] Figure 1 Schematic flow diagram of a lightweight fuel dispenser behavior recognition method based on sequence diagrams provided in Embodiment 1 of the present invention;
[0046] Figure 2 Schematic diagram of a human target detection network in a lightweight fuel dispenser behavior recognition method based on sequence diagrams provided in Embodiment 1 of the present invention;
[0047] Figure 3 Decoupled head diagram of a human target detection network in a lightweight fuel dispenser behavior recognition method based on sequence diagrams provided in Embodiment 1 of the present invention;
[0048] Figure 4 Structural diagram of a lightweight behavior recognition network for fusing spatio-temporal features in a lightweight fuel dispenser behavior recognition method based on sequence diagrams provided in Embodiment 1 of the present invention;
[0049] Figure 5 Human target test result diagram in a lightweight fuel dispenser behavior recognition method based on sequence diagrams provided in Embodiment 1 of the present invention;
[0050] Figure 6 This is the behavior test result graph in the lightweight fuel dispenser behavior recognition method based on sequence diagrams provided in Embodiment 1 of the present invention;
[0051] Figure 7 This is another behavior test result graph in the lightweight fuel dispenser behavior recognition method based on sequence diagrams provided in Embodiment 1 of the present invention;
[0052] Figure 8 This is the schematic diagram of the lightweight fuel dispenser behavior recognition device based on sequence diagrams provided in Embodiment 2 of the present invention. Detailed implementation manners
[0053] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0054] Embodiment 1
[0055] Refer to Figure 1 , Figure 2 , Figure 3 and Figure 4 , Embodiment 1 of the present invention provides a lightweight fuel dispenser behavior recognition method based on sequence diagrams, including the following steps:
[0056] S1. Collect the behavior video stream to be recognized, decode the behavior video stream, decompose the decoded behavior video stream into continuous frames, and perform size preprocessing on the decomposed continuous frame images;
[0057] S2. Construct a human target detection network, use the human target detection network to extract high-dimensional information from the preprocessed continuous frame images, perform spatial pyramid pooling after high-dimensional information extraction, use a shortcut path to splice the dimensions of the input feature map and the pooling output, and then use a convolutional layer for feature information fusion;
[0058] S3. Detect the target association of the anchor boxes corresponding to the previous and next frames obtained by the human target detection network to track the human target and obtain a set of human motion trajectory sequence diagrams;
[0059] S4. Use the set of human motion trajectory sequence diagrams obtained by tracking as the input, and use a convolutional layer and a max pooling layer to extract features from the input set of human motion trajectory sequence diagrams and stack them to obtain a temporal fusion feature map;
[0060] S5. Unfold the time-fused feature map into a one-dimensional feature vector to fuse spatial information, and use a fully-connected network and Softmax to classify the feature vector to obtain a predetermined behavior type of the fueler.
[0061] In this embodiment, in step S1, a network camera is used to collect the behavior video stream of the airport fueler in real time, the edge device is used to decode the video stream address, the video stream is divided into a series of consecutive frames, and the frames are resized to consecutive frame images of 640×640. Among them, the resize operation itself is a function in the OpenCV library, which can scale the picture.
[0062] In this embodiment, in step S2, for the preprocessed consecutive frame images, after two-dimensional convolution, batch normalization, and ReLu activation, a high-dimensional feature image is obtained, and then high-dimensional information is extracted through a feature extractor composed of Block units and SandGlass units for several times.
[0063] Auxiliary Figure 2 、 Figure 3 and Figure 4 , specifically, by constructing a lightweight human target detection network, the decoded and preprocessed consecutive frame images (640×640×3) are lifted to a high-dimensional feature image with 32 channels and a size of 320×320 after two-dimensional convolution, batch normalization, and ReLu activation. Then, after 5 times of feature extraction by a feature extractor composed of Block units and SandGlass units to continue extracting high-dimensional information, after obtaining high-dimensional information with 256 channels and a size of 20×20, spatial pyramid pooling is performed, and the input feature map is dimensionally concatenated with the pooling outputs of three maximum values (3, 7, 11) through a shortcut path, and a convolutional layer is used to fuse the feature information of these four different scales.
[0064] In step S2, in the feature information fusion process, the high-dimensional feature information of 20×20×256 and the low-dimensional feature information of 80×80×64 are mutually feature-fused and enhanced, and output to two detection heads;
[0065] Decouple the detection heads to eliminate the conflict in the classification and regression features in object detection; use a 1×1 convolutional layer to reduce the channel dimension, decouple the prediction branch, and then use two 3×3 convolutional layer parallel branches for classification and regression tasks respectively.
[0066] Specifically, for simple human targets, in this embodiment, only the high-dimensional feature information of 20×20×256 and the low-dimensional feature information of 80×80×64 are mutually feature-fused and enhanced, and output to two detection heads.
[0067] Among them, for the output part, the original detection head would output information such as classification, regression, and scores simultaneously. The detection head was decoupled to eliminate the conflicts in the classification and regression features in object detection. By adjusting the original detection head, a 1×1 convolutional layer was used to reduce the channel dimension, and at the same time, the prediction branches were decoupled. Then, two parallel branches of 3×3 convolutional layers were respectively used for the classification and regression tasks.
[0068] Since in working scenarios such as airport refueling trucks, there are usually more than one staff member to be detected and recognized. After the human object detection finds the specific positions of humans in each frame of the image, it cannot give the human objects corresponding to the detection boxes between adjacent frames.
[0069] Specifically, this embodiment uses the intersection over union object tracking algorithm. The core idea is that in slow motion, the probability that two detection boxes with a relatively large intersection over union ratio between two adjacent frames belong to the same object is relatively large. Usually, the command movements of airport refueling personnel only involve slow and stable gestures and posture changes, and often can meet the requirement that the same human anchor box has a relatively high overlap degree IOU in adjacent consecutive frame detections.
[0070] In this embodiment, in step S3, the overlap degree IOU is calculated using the previous and next frame anchor boxes a and anchor box b obtained by the human object detection network. The formula for the overlap degree IOU is:
[0071] IOU = (Area(a) ∩ Area(b)) / (Area(a) ∪ Area(b))
[0072] In the formula, Area(a) is the area of the region occupied by anchor box a, and Area(b) is the area of the region occupied by anchor box b.
[0073] In this embodiment, in step S3, the steps of obtaining the human motion trajectory sequence diagram set by tracking the human object are as follows:
[0074] Among them, let D0, D1... D F-1 respectively represent the detection sets of the 0th, 1st,..., F - 1th frames, and d0, d1... d N-1 respectively represent the N human objects to be detected; T a represents the set of active trajectories, and T f represents the set of human trajectories that have been tracked;
[0075] First, a threshold σ l is used to filter out the human detection boxes with low confidence;
[0076] Then, for the current frame detection set D f , for each trajectory t a in the set of active trajectories T i, select the human detection box information of the last added trajectory, and calculate the overlap degree IOU between the current position information and all detection boxes in the current frame detection set in sequence. If the maximum IOU(d best ,t i ) is greater than or equal to the preset threshold, it is determined that the current detection box belongs to the corresponding added trajectory, and the current detection box is deleted from the current frame detection set D f .
[0077] Since, when there are occlusions and overlaps in the movement trajectories of different human targets, or when a human makes a sudden large movement, it is easy to have missed detections and false detections only relying on the overlap degree IOU. Therefore, if the maximum IOU(d best ,t i ) is not greater than or equal to the preset threshold, use the aspect ratio and confidence score of the human detection box to calculate the similarity Sim between different human detection boxes in the previous and current frames for re-matching:
[0078] Sim = 1 - 0.5 * abs[(R a +R b ) / R a *R b )] + 0.5 * abs(C a -C b )
[0079] In the formula, R a , R b respectively represent the ratio of the length to the width of the detection box; C a , C b respectively represent the difference in the confidence of the detection box.
[0080] Among them, Sim reflects the similarity degree of two detection boxes. The closer Sim is to 1, the more similar the two detection boxes are, and the more likely they belong to the same target. When Sim is less than the preset value of 0.5, it is judged whether the highest score of the historical position in the corresponding trajectory is greater than the threshold σ h , and whether the appearance time of the corresponding trajectory is greater than the tracking completion time t min . If the conditions are met, it is determined that the tracking is completed, and the corresponding trajectory t i is moved from the active trajectory set T a to the set T f of human trajectories that have been tracked.
[0081] Among them, if all the remaining detection boxes in the current frame detection set D f do not match the detection boxes of any active trajectories, they are inserted into T a and considered as a new trajectory. When all detections are completed, for each active trajectory t a in T i, determine whether the condition for the tracking to be completed is satisfied. If so, transfer to the set T of the human body trajectories that have been tracked to completion f and finally return the set T of the human body trajectories that have been tracked to completion f .
[0082] In this embodiment, in steps S4 and S5, the sequence diagrams (with a fixed size of 64×64) corresponding to the continuously tracked 30-frame human body trajectory set are used as the input for the action recognition process. Then, through an 8-layer convolutional neural network, convolution and max pooling layers are used to extract features from the input 30 sequence diagrams and stack them. After feature extraction, 30 feature maps with a size of 16×16 and 128 channels are obtained, and then the spatio-temporal features are fused.
[0083] Auxiliary Figure 5 , Figure 6 and Figure 7 , the 30 feature maps are stacked into a time fusion feature map of 128×16×16 through the max pooling layer, and then the time fusion feature map of 128×16×16 is unfolded into a one-dimensional feature vector to fuse the spatial information, so that the obtained one-dimensional feature vector contains both the time domain and the spatial domain feature information. Finally, a fully connected network and Softmax are used to complete the classification of the feature vector: finger, bending down, other actions.
[0084] In summary, the present invention captures the behavior video stream to be recognized, decodes the behavior video stream, decomposes the decoded behavior video stream into consecutive frames, and preprocesses the dimensions of the decomposed consecutive frame images; constructs a human target detection network, uses the human target detection network to extract high-dimensional information from the preprocessed consecutive frame images, performs spatial pyramid pooling after high-dimensional information extraction, uses a shortcut path to splice the dimensions of the input feature map and the pooling output, and then uses a convolutional layer to fuse the feature information; associates the anchor boxes corresponding to the previous and next frames obtained by the human target detection network to detect the targets and tracks the human target to obtain a set of human motion trajectory sequence diagrams; uses the set of human motion trajectory sequence diagrams obtained by tracking as the input, and uses a convolutional layer and a max pooling layer to extract features from the input set of human motion trajectory sequence diagrams and stack them to obtain a time fusion feature map; unfolds the time fusion feature map into a one-dimensional feature vector to fuse the spatial information, and classifies the feature vector using a fully connected network and Softmax to obtain the predetermined behavior types of fuel dispensers. By constructing a lightweight human target detection network, the decoded and preprocessed consecutive frame images (640×640×3) are lifted to a high-dimensional feature image with 32 channels and a size of 320×320 through two-dimensional convolution, batch normalization, and ReLu activation, and then continue to extract high-dimensional information through a feature extractor composed of 5 Block units and SandGlass units. After obtaining high-dimensional information with 256 channels and a size of 20×20, spatial pyramid pooling is performed, and the dimensions of the input feature map are spliced with the pooling outputs of three maximum values (3, 7, 11) through a shortcut path, and a convolutional layer is used to fuse the feature information of these four different scales. In the process of feature information fusion, the high-dimensional feature information of 20×20×256 and the low-dimensional feature information of 80×80×64 are mutually fused and enhanced in features, and output to two detection heads; the detection heads are decoupled to eliminate the conflict between classification and regression features in target detection; a 1×1 convolutional layer is used to reduce the channel dimension, and the prediction branch is decoupled, and then two 3×3 convolutional layer parallel branches are respectively used for classification and regression tasks. For simple human targets, only the high-dimensional feature information of 20×20×256 and the low-dimensional feature information of 80×80×64 are mutually fused and enhanced in features, and output to two detection heads. For the output part, the original detection head will output information such as classification, regression, and score at the same time. The detection head is decoupled to eliminate the conflict between classification and regression features in target detection. By adjusting the original detection head, a 1×1 convolutional layer is used to reduce the channel dimension, and at the same time the prediction branch is decoupled, and then two 3×3 convolutional layer parallel branches are respectively used for classification and regression tasks. The present invention has low computational consumption, a simple structure, is suitable for being deployed on embedded devices, can ensure real-time performance, has good accuracy at the same time, and can reduce the operation and maintenance costs.
[0085] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by the cooperation of multiple devices. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0086] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0087] Embodiment 2
[0088] See Figure 8 , Embodiment 2 of the present invention also provides a lightweight fuel dispenser behavior recognition device based on a sequence diagram, including:
[0089] A video acquisition and processing module 1, configured to acquire a behavior video stream to be recognized, decode the behavior video stream, decompose the decoded behavior video stream into consecutive frames, and perform size preprocessing on the decomposed consecutive frame images;
[0090] A model construction and processing module 2, configured to construct a human target detection network, use the human target detection network to extract high-dimensional information from the preprocessed consecutive frame images, perform spatial pyramid pooling after high-dimensional information extraction, use a shortcut path to splice the dimensions of the input feature map and the pooling output, and then use a convolutional layer to perform feature information fusion;
[0091] A human target tracking module 3, configured to detect the target association of the anchor boxes corresponding to the front and rear two frames obtained by the human target detection network to track the human target and obtain a set of human motion trajectory sequence diagrams;
[0092] A feature fusion module 4, configured to use the set of human motion trajectory sequence diagrams obtained by tracking as input, and use a convolutional layer and a max pooling layer to extract features from the input set of human motion trajectory sequence diagrams and stack them to obtain a time fusion feature map;
[0093] A feature classification and processing module 5, configured to expand the time fusion feature map into a one-dimensional feature vector to fuse spatial information, and use a fully connected network and Softmax to classify the feature vector to obtain a predetermined fuel dispenser behavior type.
[0094] In this embodiment, in the model construction processing module 2, for the preprocessed consecutive frame images, after two-dimensional convolution, batch normalization, and ReLu activation, a high-dimensional feature image is obtained, and then high-dimensional information is extracted through a feature extractor composed of a Block unit and a SandGlass unit for several times;
[0095] In the model construction processing module 2, in the feature information fusion process, mutual feature fusion enhancement is performed on the high-dimensional feature information of 20×20×256 and the low-dimensional feature information of 80×80×64, and the result is output to two detection heads;
[0096] In the model construction processing module 2, the detection heads are decoupled to eliminate the conflict between classification and regression features in object detection; a 1×1 convolutional layer is used to reduce the channel dimension, and the prediction branches are decoupled, and then two 3×3 convolutional layer parallel branches are respectively used for classification and regression tasks.
[0097] The overlap degree IOU is calculated using the front and rear two-frame anchor boxes a and anchor box b obtained by the human target detection network. The calculation formula for the overlap degree IOU is:
[0098] IOU = (Area(a) ∩ Area(b)) / (Area(a) ∪ Area(b))
[0099] In the formula, Area(a) is the area of the region occupied by the anchor box a, and Area(b) is the area of the region occupied by the anchor box b.
[0100] In this embodiment, the human target tracking module 3 includes:
[0101] A filtering sub-module 31, which is used to use a threshold σ l to filter out the human detection frames with low confidence;
[0102] A trajectory processing sub-module 32, which is used for the current frame detection set D f , for each trajectory t a in the active trajectory set T i , select the human detection frame information of the last added trajectory, and calculate the overlap degree IOU between the current position information and all detection frames in the current frame detection set in turn. If the maximum IOU(d best , t i ) is greater than or equal to the preset threshold, it is determined that the current detection frame belongs to the corresponding added trajectory, and the current detection frame is deleted from the current frame detection set D f ;
[0103] A similarity calculation sub-module 33, which is used if the maximum IOU(d best , t i)Greater than or equal to the preset threshold, using the aspect ratio and confidence score of the human detection box, calculate the similarity Sim between different human detection boxes in the front and rear frames for re - matching:
[0104] Sim = 1 - 0.5 * abs[(R a +R b ) / R a *R b )]+0.5 * abs(C a -C b )
[0105] In the formula, R a and R b represent the ratio of the length to the width of the detection box respectively; C a and C b represent the difference in the confidence of the detection box respectively;
[0106] The tracking judgment sub - module 34 is used to judge whether the highest score of the historical position in the corresponding trajectory is greater than the threshold σ h when Sim is less than the preset value, and whether the appearance time of the corresponding trajectory is greater than the tracking completion time t min . If the conditions are met, it is determined that the tracking is completed, and the corresponding trajectory t i is moved from the active trajectory set T a to the set T f of the human trajectories that have been tracked.
[0107] It should be noted that for the information interaction, execution process, etc. between the above - mentioned device modules, since they are based on the same concept as the method embodiment in Embodiment 1 of the present application, the technical effects brought by them are the same as those of the method embodiment of the present application. For the specific content, reference can be made to the description in the method embodiment shown above in the present application, and details will not be repeated here.
[0108] Embodiment 3
[0109] Embodiment 3 of the present invention provides a non - transient computer - readable storage medium, in which program codes for the lightweight behavior recognition method of fuel dispensers based on sequence diagrams are stored. The program codes include instructions for executing the lightweight behavior recognition method of fuel dispensers based on sequence diagrams in Embodiment 1 or any possible implementation manner thereof.
[0110] The computer - readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center integrating one or more available media. The available media can be magnetic media (for example, floppy disks, hard disks, magnetic tapes), optical media (for example, DVDs), or semiconductor media (for example, solid - state drives (Solid State Disk, SSD)), etc.
[0111] Example 4
[0112] Example 4 of the present invention provides an electronic device, including: a memory and a processor;
[0113] The processor and the memory complete communication with each other through a bus; the memory stores program instructions executable by the processor, and the processor can execute the lightweight behavior recognition method of fuel dispensers based on sequence diagrams in Example 1 or any possible implementation thereof by invoking the program instructions.
[0114] Specifically, the processor can be implemented by hardware or software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software code stored in the memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.
[0115] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.).
[0116] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0117] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made thereto based on the present invention, which will be obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of the present invention claimed.
Claims
1. A lightweight behavior recognition method for fuel dispensers based on sequence diagrams, characterized in that Including: Collect the behavior video stream to be recognized, decode the behavior video stream, decompose the decoded behavior video stream into consecutive frames, and perform size preprocessing on the decomposed consecutive frame images; Construct a human target detection network, use the human target detection network to extract high-dimensional information from the preprocessed consecutive frame images, perform spatial pyramid pooling after high-dimensional information extraction, use a shortcut path to splice the dimensions of the input feature map and the pooling output, and then use a convolutional layer for feature information fusion; Detect the target association of the anchor boxes corresponding to the previous and next frames obtained by the human target detection network to track the human target and obtain a set of human motion trajectory sequence diagrams; Use the set of human motion trajectory sequence diagrams obtained by tracking as the input, and use a convolutional layer and a max pooling layer to extract features from the input set of human motion trajectory sequence diagrams and stack them to obtain a temporal fusion feature map; Unfold the temporal fusion feature map into a one-dimensional feature vector to fuse spatial information, and use a fully connected network and Softmax to classify the feature vector to obtain the predetermined behavior types of fuel dispensers; Calculate the overlap IOU between the anchor box a and the anchor box b of the previous and next frames obtained by the human target detection network. The calculation formula for the overlap IOU is: IOU = (Area(a) ∩ Area(b)) / (Area(a) ∪ Area(b)) In the formula, Area(a) is the area occupied by the anchor box a, and Area(b) is the area occupied by the anchor box b; The steps of tracking the human target to obtain a set of human motion trajectory sequence diagrams include: Adopt the threshold σ l Filter out the human detection boxes with low confidence; For the current frame detection set D f , for each track t a in the active track set T i , select the human detection box information of the last added track, and calculate the overlap degree IOU between the current position information and all detection boxes in the current frame detection set in sequence. If the maximum IOU(d best , t i ) is greater than or equal to the preset threshold, it is determined that the current detection box belongs to the corresponding added track, and the current detection box is deleted from the current frame detection set D f ; If the maximum IOU(d best ,t i ) is not greater than or equal to the preset threshold, the similarity Sim between different human detection boxes in the front and rear frames is calculated using the aspect ratio and confidence score of the human detection box for re-matching: Sim = 1 - 0.5 * abs[(R a + R b ) / (R a * R b )] + 0.5 * abs(C a - C b ) Wherein, R a and R b respectively represent the ratio of the length to the width of the detection box; C a and C b respectively represent the difference in the confidence levels of the detection boxes.
2. The lightweight behavior recognition method for fuel dispensers based on sequence diagrams according to claim 1, characterized in that For the preprocessed consecutive frame images, obtain high-dimensional feature images after two-dimensional convolution, batch normalization, and ReLu activation, and then extract high-dimensional information through a feature extractor composed of Block units and SandGlass units for several times.
3. The lightweight behavior recognition method for fuel dispensers based on sequence diagrams according to claim 2, wherein In the process of feature information fusion, mutually enhance the feature fusion of the high-dimensional feature information of 20×20×256 and the low-dimensional feature information of 80×80×64, and output to two detection heads; Decouple the detection heads to eliminate the conflicts in the classification and regression features in target detection; use a 1×1 convolutional layer to reduce the channel dimension, decouple the prediction branch, and then use two 3×3 convolutional layer parallel branches for classification and regression tasks respectively.
4. The lightweight behavior recognition method for fuel dispensers based on sequence diagrams according to claim 1, characterized in that, When Sim is less than the preset value, it is determined whether the highest score of the historical positions in the corresponding trajectory is greater than the threshold σ h , and whether the appearance time of the corresponding trajectory is greater than the tracking completion time t min , if the conditions are met, it is determined that the tracking is completed, and the corresponding trajectory t i is moved from the active trajectory set T a to the set T of human trajectories that have been tracked f in.
5. A lightweight behavior recognition device for fuel dispensers based on sequence diagrams, characterized in that Including: A video acquisition and processing module, which is used to collect the behavior video stream to be recognized, decode the behavior video stream, decompose the decoded behavior video stream into consecutive frames, and perform size preprocessing on the decomposed consecutive frame images; A model construction and processing module, which is used to construct a human target detection network, use the human target detection network to extract high-dimensional information from the preprocessed consecutive frame images, perform spatial pyramid pooling after high-dimensional information extraction, use a shortcut path to splice the dimensions of the input feature map and the pooling output, and then use a convolutional layer for feature information fusion; A human target tracking module, which is used to detect the target association of the anchor boxes corresponding to the previous and next frames obtained by the human target detection network to track the human target and obtain a set of human motion trajectory sequence diagrams; A feature fusion module, which is used to take the set of human motion trajectory sequence diagrams obtained by tracking as input, and use convolutional layers and max-pooling layers to extract features from the input set of human motion trajectory sequence diagrams and stack them to obtain a time fusion feature map; A feature classification processing module, which is used to expand the time fusion feature map into a one-dimensional feature vector fusion space information, and use a fully connected network and Softmax to classify the feature vector to obtain a predetermined behavior type of a fuel dispenser; In the model construction processing module, the overlap degree IOU is calculated by using the front and rear two-frame anchor boxes a and b obtained by the human target detection network. The overlap degree IOU calculation formula is: IOU = (Area(a) ∩ Area(b)) / (Area(a) ∪ Area(b)) In the formula, Area(a) is the area of the region occupied by the anchor box a, and Area(b) is the area of the region occupied by the anchor box b; The human target tracking module includes: Filter sub-module, which is used to adopt a threshold value σ l to filter out the human detection frames with low confidence; A trajectory processing sub-module, used for the current frame detection set D f , for each trajectory t a in the active trajectory set T i , select the human detection box information of the last added trajectory, and calculate the overlap degree IOU between the current position information and all detection boxes in the current frame detection set in sequence. If the maximum IOU(d best , t i ) is greater than or equal to the preset threshold, it is determined that the current detection box belongs to the corresponding added trajectory, and the current detection box is deleted from the current frame detection set D f ; The similarity calculation sub-module is used to calculate the similarity Sim between different human detection boxes in the front and rear frames for re-matching by using the aspect ratio and confidence score of the human detection box if the maximum IOU(d best ,t i ) is less than the preset threshold: Sim = 1 - 0.5 * abs[(R a + R b ) / (R a * R b )] + 0.5 * abs(C a - C b ) Wherein, R a and R b respectively represent the ratio of the length to the width of the detection box; C a and C b respectively represent the difference in the confidence levels of the detection boxes.
6. The lightweight fuel dispenser behavior recognition device based on the sequence diagram according to claim 5, wherein, In the model construction processing module, for the preprocessed consecutive frame images, after two-dimensional convolution, batch normalization, and ReLu activation, a high-dimensional feature image is obtained, and then high-dimensional information is extracted through a feature extractor composed of Block units and SandGlass units for several times; In the model construction processing module, in the feature information fusion process, mutual feature fusion enhancement is performed on the high-dimensional feature information of 20×20×256 and the low-dimensional feature information of 80×80×64, and the output is sent to two detection heads; In the model construction processing module, the detection heads are decoupled to eliminate the conflict in the classification and regression features in target detection; a 1×1 convolutional layer is used to reduce the channel dimension, and the prediction branches are decoupled, and then two 3×3 convolutional layer parallel branches are respectively used for classification and regression tasks.
7. The lightweight fuel dispenser behavior recognition device based on the sequence diagram according to claim 6, wherein The human target tracking module includes: A tracking judgment sub-module, which is used to judge whether the highest score of the historical position in the corresponding trajectory is greater than the threshold σ when Sim is less than a preset value h , and whether the occurrence time of the corresponding trajectory is greater than the tracking completion time t min , if the condition is met, it is determined that the tracking is completed, and the corresponding trajectory t i is moved from the active trajectory set T a to the set of human trajectories T that have been tracked f .
Citation Information
Patent Citations
Driver behavior recognition method based on deep hybrid encoding and decoding neural network
CN111695435A
Behavior detection method and device, electronic equipment and computer readable storage medium
CN114511930A