Computer Vision-Based Theft Behavior Recognition Method and System
By introducing multi-scale feature extraction and micro-expression analysis in computer vision technology, combining deep learning decision classification model and theft behavior scoring mechanism analysis engine, the problem of low accuracy of theft behavior recognition in complex scenarios in the existing technology is solved, and a high accuracy and practical theft behavior recognition system is achieved.
Patent Information
- Application Number
- CN202510142818.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The accuracy of theft behavior recognition in the prior art in complex scenarios is low, especially in environments with dense people and severe light changes, misjudgment and misjudgment are prone to occur. The existing systems lack in-depth analysis of behavioral characteristics and multi-character fusion mechanisms, making it difficult to fully and accurately judge the true intention of behavior.
Using a computer vision-based method, multi-scale features and micro-expression features of behavioral targets are extracted through preprocessing of video data, YOLOv11 deep learning object detection algorithm, feature pyramid network, Openpose algorithm, partial affinity field technology, wavelet transformation and space-time attention mechanism, and behavior judgment is performed through deep learning decision classification model and theft behavior scoring mechanism analysis engine.
It significantly improves the comprehensiveness and accuracy of target recognition capabilities and behavioral analysis in complex scenarios, can accurately identify theft behavior and trigger alarm signals, and improves the practicality and reliability of the system.
Smart Images

Figure CN119649467B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to machine vision technology, and particularly to a method and system for identifying theft behavior based on computer vision. Background Art
[0002] With the rapid development of social economy, the security management problems in public places such as shopping malls, supermarkets, and exhibition halls have become increasingly prominent, among which the prevention and timely detection of theft behavior are particularly important. Traditional security monitoring systems mainly rely on manual monitoring and simple video surveillance, which have problems such as high labor costs and low monitoring efficiency. In recent years, with the development of computer vision and deep learning technologies, behavior recognition systems based on intelligent vision analysis have gradually become a research hotspot in the security field. Currently, the intelligent vision analysis systems mainly achieve automatic recognition and early warning of suspicious behaviors through technologies such as video image processing, target detection, and behavior analysis.
[0003] The existing theft behavior recognition technologies mainly have the following deficiencies:
[0004] The recognition accuracy of the existing technologies in complex scenarios is relatively low. Especially in environments with dense crowds and drastic light changes, false positives and false negatives are likely to occur. This is mainly due to the insufficient adaptability of the existing algorithms to environmental interference factors and the lack of in-depth analysis of behavior characteristics.
[0005] The current recognition systems often only focus on single behavior characteristics, such as human body postures or movement trajectories, and ignore the analysis of detailed characteristics such as micro-expressions and hand movements during the theft behavior process, resulting in the system being unable to comprehensively and accurately judge the true intention of the behavior.
[0006] The existing technologies lack an effective multi-feature fusion mechanism and fail to organically combine multi-dimensional information such as human body postures, facial expressions, and behavior time series. At the same time, they lack a reasonable confidence evaluation system and it is difficult to accurately classify and timely warn of suspicious behaviors at different levels. Summary of the Invention
[0007] The embodiments of the present invention provide a method and system for identifying theft behavior based on computer vision, which can solve the problems in the existing technologies.
[0008] In the first aspect of the embodiments of the present invention,
[0009] A method for identifying theft behavior based on computer vision is provided, including:
[0010] Collecting a video data stream, converting the frame rate of the video data stream to obtain a video stream with a target frame rate, applying Gaussian filtering and median filtering to the video stream with the target frame rate to eliminate noise, and using a threshold segmentation algorithm to perform image segmentation on the filtered video stream to obtain preprocessed video frames;
[0011] Use the YOLOv11 deep learning object detection algorithm to perform object detection and localization on the preprocessed video frames to obtain the bounding box information of the behavior objects. Input the bounding box information into the Feature Pyramid Network for multi-scale feature extraction, and fuse the high-level semantic features and low-level detail features through the Path Aggregation Network to obtain the object feature map. Based on the object feature map, use the Openpose algorithm to extract the coordinate data of the hand key points and torso key points of the behavior objects, and use the Part Affinity Fields technique to establish the association relationship between the hand key points and the torso key points to construct the pose skeleton. Use wavelet transform and spatio-temporal attention mechanism to extract the micro-expression features from the facial region of the behavior objects. Input the pose skeleton and the micro-expression features into the deep learning decision classification model, and output the behavior feature vector including the distance data between the hand and the torso, the hand movement trajectory data, and the behavior time window data.
[0012] Input the behavior feature vector into the theft behavior scoring mechanism analysis engine for behavior determination. The theft behavior scoring mechanism analysis engine fuses the convolutional neural network and the Transformer structure, calculates the similarity between the behavior feature vector and the preset behavior template, and generates a confidence score. When the confidence score is greater than 0.8, it is determined as a theft behavior and an alarm signal is triggered. When the confidence score is between 0.5 and 0.8, it is determined as an abnormal behavior and a tracking mark is recorded. When the confidence score is less than 0.5, it is determined as a normal behavior.
[0013] Using the YOLOv11 deep learning object detection algorithm to perform object detection and localization on the preprocessed video frames to obtain the bounding box information of the behavior objects, and inputting the bounding box information into the Feature Pyramid Network for multi-scale feature extraction and fusing the high-level semantic features and low-level detail features through the Path Aggregation Network to obtain the object feature map includes:
[0014] Divide the preprocessed video frames into several grid cells, perform bounding box prediction on the grid cells, obtain the center coordinates, width, height and confidence of the bounding box, and the confidence is calculated by the product of the probability of the object existing in the box and the intersection over union of the predicted box and the ground truth box.
[0015] Based on the bounding box, construct a Feature Pyramid Network. Downsample the preprocessed video frames through convolutional layers to obtain five different-scale bottom feature maps, and the sizes of the five different-scale bottom feature maps are halved in turn. Upsample the high-level feature maps and perform one-dimensional convolution fusion with the same-scale feature maps connected horizontally to obtain the pyramid feature map.
[0016] Input the pyramid feature map into the path aggregation network, construct a top-down enhancement path and a bottom-up enhancement path in the path aggregation network. The top-down enhancement path obtains a top-down feature map through upsampling and feature concatenation. The bottom-up enhancement path obtains a bottom-up feature map through downsampling and feature concatenation. Fuse the top-down feature map and the bottom-up feature map through a learnable fusion weight coefficient to obtain a fused feature map;
[0017] Perform attention enhancement on the fused feature map in both the channel dimension and the spatial dimension. The attention enhancement in the channel dimension obtains channel attention weights by performing average pooling and max pooling on the fused feature map and passing through a multi-layer perceptron. The attention enhancement in the spatial dimension obtains spatial attention weights by performing average pooling and max pooling on the fused feature map and passing through a convolutional operation. Multiply the fused feature map with the channel attention weights and the spatial attention weights to obtain an attention-enhanced feature map;
[0018] Calculate the information entropy of the attention-enhanced feature map at different scales, calculate the dynamic fusion weights at different scales based on the information entropy, and perform weighted summation of the attention-enhanced feature map with the corresponding dynamic fusion weights to obtain a final feature map, which contains multi-scale semantic information and detailed features of the target.
[0019] Based on the target feature map, use the Openpose algorithm to extract the coordinate data of the hand key points and torso key points of the behavior target, and use the part affinity fields technique to establish the association relationship between the hand key points and the torso key points to construct a pose skeleton, including:
[0020] Based on the target feature map, through a two-branch convolutional neural network, generate a confidence map through the first branch and generate part affinity fields through the second branch. Use the Gaussian kernel function to calculate the distance from the pixel points in the target feature map to the true positions of the key points to obtain an initial confidence map of multi-level key points, and determine the candidate positions of the key points of multiple human targets based on the initial confidence map;
[0021] Based on the candidate positions of the key points, obtain the initial coordinates of the wrist key points, perform adaptive cropping of the hand region centered on the initial coordinates of the wrist key points to obtain a hand region of interest, input the hand region of interest into a hand key point extraction network, generate a hand key point heat map through multi-layer convolutional operations, and extract the accurate coordinates of the fine hand key points from the hand key point heat map;
[0022] Adopt a hierarchical detection strategy for the torso key points among the candidate positions of the key points, determine the reference position of the torso main key points based on the maximum response value of the initial confidence map, input the reference position into the position regression network to calculate the spatial offset of the secondary key points relative to the reference position, and obtain the accurate position of the torso secondary key points according to the spatial offset and the reference position;
[0023] Construct a bidirectional part affinity field based on the accurate coordinates of the hand fine key points and the accurate positions of the torso secondary key points, calculate the integral value of the feature vector of the forward part affinity field along the key point connection direction, calculate the integral value of the feature vector of the reverse part affinity field along the reverse direction of the key point connection, and fuse the forward part affinity field and the reverse part affinity field through an adaptive direction weight to obtain an enhanced part affinity field;
[0024] Use the enhanced part affinity field to calculate the connection relationship score between adjacent key points, obtain the optimal skeleton topology structure by maximizing the global cumulative sum of the connection relationship score, apply a temporal smoothing constraint to the optimal skeleton topology structure, and perform an adaptive weight fusion on the current frame pose, the previous frame historical pose, and the motion prediction pose to obtain a spatio-temporally consistent smoothed pose skeleton, and the smoothed pose skeleton includes a stable connection relationship between the hand and torso key points.
[0025] Adopting a hierarchical detection strategy for the torso key points among the candidate positions of the key points, determining the reference position of the torso main key points based on the maximum response value of the initial confidence map, inputting the reference position into the position regression network to calculate the spatial offset of the secondary key points relative to the reference position, and obtaining the accurate position of the torso secondary key points according to the spatial offset and the reference position includes:
[0026] Perform a weighted sum on the heat map responses of multiple channels to obtain a confidence response map, set a response threshold based on the confidence response map to screen and obtain the candidate region of the main key points, and perform Gaussian smoothing processing on the candidate region of the main key points to obtain a smoothed response map;
[0027] Determine the maximum response position in the smoothed response map as the reference position of the main key points, calculate the sub-pixel offset using the response differences in the horizontal and vertical directions, and add the reference position and the sub-pixel offset to obtain the refined coordinates of the main key points;
[0028] Construct a multi-level feature pyramid network, perform upsampling operations and convolution operations on the input feature map to obtain feature maps of different levels, cascade the feature maps of different levels and input them into the position regression network to predict the spatial offset of the secondary key points relative to the refined coordinates of the main key points;
[0029] Based on human anatomical features, bone length constraint conditions and joint angle constraint conditions are established. The refined backbone key point coordinates are added to the spatial offset to obtain the initial position of the secondary key points, and the initial position of the secondary key points is constrained and optimized to obtain the optimized position of the secondary key points that satisfies the bone length constraint conditions and the joint angle constraint conditions;
[0030] A temporal smoothing constraint is introduced. The displacement difference between the secondary key point positions in the current frame and the secondary key point positions in the previous two frames is calculated, and the motion speed difference between the current frame and the previous frame is calculated. The spatial positioning error, displacement difference and speed difference are used as optimization objectives, and the optimized position of the secondary key points is optimized for temporal consistency to obtain the final accurate position of the secondary key points.
[0031] Wavelet transform and spatio-temporal attention mechanism are used to extract micro-expression features from the facial region of the behavior target. The pose skeleton and the micro-expression features are input into a deep learning decision classification model, and the behavior feature vector including the distance data between the hand and the torso, the hand movement trajectory data, and the behavior time window data is output:
[0032] The facial region image of the behavior target is subjected to multi-scale two-dimensional discrete wavelet transform to obtain wavelet coefficients at different scales. Based on the wavelet coefficients, the wavelet energy value at each scale is calculated and the energy ratio is determined to obtain the initial feature map of the facial region;
[0033] The initial feature map is input into a two-stream spatio-temporal attention network. The initial feature map is subjected to average pooling and max pooling processing through the spatial attention branch to obtain the spatial attention weight, and the initial feature map is subjected to three-dimensional convolution operation through the temporal attention branch to obtain the temporal attention weight. The initial feature map is multiplied by the spatial attention weight and the temporal attention weight to obtain an enhanced feature map;
[0034] The Euclidean distance between the hand key points and the torso key points in the pose skeleton is calculated to construct a hand-torso distance matrix, and the speed, acceleration and direction angle of the hand movement are calculated based on the continuous frame position information of the hand key points to obtain a hand movement trajectory descriptor;
[0035] The enhanced feature map is subjected to feature aggregation by using a sliding time window, the temporal correlation of the feature sequence within the time window is calculated, and the feature sequence is weighted and summed with the time weight to obtain the time window feature;
[0036] Construct a multi-modal feature vector by concatenating the enhanced feature map, the hand-trunk distance matrix, the hand movement trajectory descriptor, and the time window feature. Input the multi-modal feature vector into a multi-layer perceptron and a bidirectional long short-term memory network simultaneously for feature fusion. Pass the fused features through a softmax classifier to obtain the probability distribution of behavior categories, and optimize and train the classification model based on the cross-entropy loss function and regularization terms to obtain the final behavior classification result.
[0037] Input the behavior feature vector into the theft behavior scoring mechanism analysis engine for behavior determination. The theft behavior scoring mechanism analysis engine fuses a convolutional neural network and a Transformer structure, and calculates the similarity between the behavior feature vector and a preset behavior template and generates a confidence score, including:
[0038] Recombine the input behavior feature vector into a spatio-temporal feature matrix, and embed sine position encoding and cosine position encoding in the spatio-temporal feature matrix to obtain the encoded spatio-temporal feature matrix;
[0039] Perform parallel convolutional operations on the encoded spatio-temporal feature matrix using convolutional branches with different kernel sizes to obtain multi-scale convolutional features, and perform weighted fusion on the multi-scale convolutional features to obtain fused convolutional features;
[0040] Perform query matrix transformation, key matrix transformation, and value matrix transformation on the encoded spatio-temporal feature matrix respectively to obtain a query matrix, a key matrix, and a value matrix. Calculate the attention weights based on the product of the query matrix and the key matrix, multiply the attention weights by the value matrix to obtain attention features, and concatenate multiple attention features to obtain a multi-head attention output feature;
[0041] Calculate the cosine similarity between the fused convolutional features and the preset behavior template to obtain similarity features, and perform dynamic time warping calculation on the fused convolutional features and the preset behavior template to obtain time warping distance features;
[0042] Perform weighted summation on the fused convolutional features, the multi-head attention output features, the similarity features, and the time warping distance features to obtain an integrated feature, and perform sigmoid activation operation on the integrated feature to obtain an initial confidence score;
[0043] Perform exponential smoothing processing on the initial confidence score to obtain a smoothed confidence score, calculate the temporal consistency of the smoothed confidence score, and use the weighted sum of the smoothed confidence score and the temporal consistency as the final theft behavior determination result.
[0044] In the second aspect of the embodiments of the present invention, a theft behavior recognition system based on computer vision is provided, including:
[0045] The first unit is used to collect a video data stream, convert the frame rate of the video data stream to obtain a video stream with a target frame rate, apply Gaussian filtering and median filtering to the video stream with the target frame rate to eliminate noise, and use a threshold segmentation algorithm to perform image segmentation on the filtered video stream to obtain a preprocessed video frame;
[0046] The second unit is used to use the YOLOv11 deep learning object detection algorithm to perform object detection and localization on the preprocessed video frame to obtain the bounding box information of the behavior target, input the bounding box information into a feature pyramid network for multi-scale feature extraction, and fuse high-level semantic features and low-level detail features through a path aggregation network to obtain a target feature map; based on the target feature map, use the Openpose algorithm to extract the coordinate data of the hand key points and torso key points of the behavior target, and use the part affinity fields technique to establish the association relationship between the hand key points and the torso key points to construct a pose skeleton, use wavelet transform and spatio-temporal attention mechanism to extract the micro-expression features of the facial area of the behavior target, and input the pose skeleton and the micro-expression features into a deep learning decision classification model to output a behavior feature vector including the distance data between the hand and the torso, the hand movement trajectory data, and the behavior time window data;
[0047] The third unit is used to input the behavior feature vector into a theft behavior scoring mechanism analysis engine for behavior determination. The theft behavior scoring mechanism analysis engine fuses a convolutional neural network and a Transformer structure, calculates the similarity between the behavior feature vector and a preset behavior template and generates a confidence score. When the confidence score is greater than 0.8, it is determined as a theft behavior and an alarm signal is triggered. When the confidence score is between 0.5 and 0.8, it is determined as an abnormal behavior and a tracking mark is recorded. When the confidence score is less than 0.5, it is determined as a normal behavior.
[0048] The third aspect of the embodiments of the present invention
[0049] A kind of electronic device is provided, including:
[0050] A processor;
[0051] A memory for storing instructions executable by the processor;
[0052] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0053] The fourth aspect of the embodiments of the present invention,
[0054] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0055] The beneficial effects of this application are as follows:
[0056] In the present invention, Gaussian filtering and median filtering are adopted for noise elimination, and combined with the threshold segmentation algorithm for image preprocessing, significantly improving the accuracy and robustness of subsequent target detection. At the same time, the YOLOv11 deep learning algorithm is used for target detection and positioning, and combined with the feature pyramid network and path aggregation network for multi-scale feature extraction and fusion, effectively enhancing the target recognition ability in complex scenarios.
[0057] The present invention innovatively combines the Openpose algorithm and the part affinity fields technology to accurately extract the hand and torso key points of the behavior target, construct the pose skeleton, and extract the micro-expression features through wavelet transform and spatio-temporal attention mechanism. This method of multi-modal feature fusion greatly enhances the comprehensiveness and accuracy of behavior analysis, and can capture subtle behavior features and emotional changes.
[0058] The present invention designs a theft behavior scoring mechanism analysis engine that integrates a convolutional neural network and a Transformer structure. By calculating the similarity between the behavior feature vector and the preset behavior template, a confidence score is generated, and behavior determination is performed according to different score ranges. This method can not only accurately identify theft behaviors, but also mark and track abnormal behaviors, greatly improving the practicality and reliability of the system, and providing an efficient and intelligent solution for the security monitoring field. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a schematic flowchart of the method for identifying theft behaviors based on computer vision according to an embodiment of the present invention;
[0060] Figure 2 It is a schematic structural diagram of the system for identifying theft behaviors based on computer vision according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0062] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0063] Figure 1 This is a schematic flowchart of the theft behavior recognition method based on computer vision according to an embodiment of the present invention. As Figure 1 shown, the method includes:
[0064] S101. Collect video data streams, perform frame rate conversion on the video data streams to obtain a video stream with a target frame rate, apply Gaussian filtering and median filtering to the video stream with the target frame rate to eliminate noise, and use a threshold segmentation algorithm to perform image segmentation on the filtered video stream to obtain preprocessed video frames;
[0065] S102. Use the YOLOv11 deep learning object detection algorithm to perform object detection and localization on the preprocessed video frames to obtain the bounding box information of the behavior target, input the bounding box information into a feature pyramid network for multi-scale feature extraction, and fuse high-level semantic features and low-level detail features through a path aggregation network to obtain a target feature map; based on the target feature map, use the Openpose algorithm to extract the coordinate data of the hand key points and torso key points of the behavior target, and use the part affinity fields technique to establish the association relationship between the hand key points and the torso key points to construct a pose skeleton, use wavelet transform and spatio-temporal attention mechanism to extract the micro-expression features of the facial area of the behavior target, and input the pose skeleton and the micro-expression features into a deep learning decision classification model to output a behavior feature vector including the distance data between the hand and the torso, the hand movement trajectory data, and the behavior time window data;
[0066] S103. Input the behavior feature vector into a theft behavior scoring mechanism analysis engine for behavior determination. The theft behavior scoring mechanism analysis engine fuses a convolutional neural network and a Transformer structure, calculates the similarity between the behavior feature vector and a preset behavior template, and generates a confidence score. When the confidence score is greater than 0.8, it is determined as a theft behavior and an alarm signal is triggered. When the confidence score is between 0.5 and 0.8, it is determined as an abnormal behavior and a tracking mark is recorded. When the confidence score is less than 0.5, it is determined as a normal behavior.
[0067] In an alternative embodiment, using the YOLOv11 deep learning object detection algorithm to perform object detection and localization on the preprocessed video frames to obtain the bounding box information of the behavior target, and inputting the bounding box information into a feature pyramid network for multi-scale feature extraction and fusing high-level semantic features and low-level detail features through a path aggregation network to obtain a target feature map includes:
[0068] Divide the preprocessed video frames into several grid units, perform bounding box prediction on the grid units, and obtain the center coordinates, width, height, and confidence of the bounding box. The confidence is calculated by the product of the probability of the object existing within the box and the intersection over union (IoU) between the predicted box and the ground truth box;
[0069] Based on the bounding boxes, construct a Feature Pyramid Network (FPN). Downsample the preprocessed video frames through convolutional layers to obtain five bottom - layer feature maps of different scales, where the sizes of the five different - scale bottom - layer feature maps are halved in sequence. Upsample the high - layer feature maps and perform one - dimensional convolutional fusion with the same - scale feature maps connected horizontally to obtain pyramid feature maps;
[0070] Input the pyramid feature maps into a Path Aggregation Network (PAN). In the PAN, construct a top - down enhancement path and a bottom - up enhancement path. The top - down enhancement path obtains the top - down feature map through upsampling and feature concatenation, and the bottom - up enhancement path obtains the bottom - up feature map through downsampling and feature concatenation. Fuse the top - down feature map and the bottom - up feature map using learnable fusion weight coefficients to obtain a fused feature map;
[0071] Perform attention enhancement on the fused feature map in both the channel dimension and the spatial dimension. The attention enhancement in the channel dimension is obtained by performing average pooling and max pooling on the fused feature map and then passing through a multi - layer perceptron to obtain channel attention weights. The attention enhancement in the spatial dimension is obtained by performing average pooling and max pooling on the fused feature map and then through convolutional operations to obtain spatial attention weights. Multiply the fused feature map by the channel attention weights and the spatial attention weights to obtain an attention - enhanced feature map;
[0072] Calculate the information entropy of the attention - enhanced feature map at different scales, calculate the dynamic fusion weights at different scales based on the information entropy, and perform weighted summation of the attention - enhanced feature map with the corresponding dynamic fusion weights to obtain a final feature map, which contains multi - scale semantic information and detailed features of the object.
[0073] A method for extracting video behavior object features based on the YOLOv11 deep - learning object - detection algorithm. First, preprocess the input video frames, including operations such as image normalization and data augmentation. The preprocessed video frames are input into the object - detection module for processing.
[0074] In the object detection stage, the preprocessed video frames are divided into grid cells of size 13×13. Each grid cell is responsible for predicting bounding boxes at three different scales, and each bounding box contains center coordinates, width, and height information. Overlapping boxes are filtered by the non-maximum suppression algorithm with a threshold of 0.5 to improve the detection accuracy. The confidence of the bounding box is obtained by multiplying the intersection over union of the predicted box and the ground truth box with the probability of the object existing inside the box, and the confidence threshold is set to 0.6.
[0075] The Feature Pyramid Network adopts a five-layer structure design. The input feature map is downsampled by a 3×3 convolutional kernel and a max pooling layer with a stride of 2. Taking an input size of 416×416 as an example, the sizes of the feature maps obtained after downsampling are 208×208, 104×104, 52×52, 26×26, and 13×13 respectively. When fusing features, a 1×1 convolutional layer is used to adjust the number of channels to ensure that the number of channels of feature maps at different levels is the same. The upsampling uses the nearest neighbor interpolation method with a sampling rate of 2 to double the size of the feature map.
[0076] In the Path Aggregation Network, the top-down path enhances features through 2x upsampling and feature concatenation. For the top-level feature map of size 13×13, it is upsampled to obtain a 26×26 feature map, which is concatenated with the feature map of the same scale to obtain enhanced features. The bottom-up path uses max pooling with a stride of 2 to downsample features and concatenates them with the feature map of the same scale. The feature fusion weights are adaptively adjusted through trainable parameters, and the initial weights are all set to 0.5.
[0077] The attention mechanism enhancement module is divided into two branches: channel attention and spatial attention. The channel attention branch performs global average pooling and max pooling on the feature map to obtain two one-dimensional vectors, and the weight coefficients are obtained through two fully connected layers. The spatial attention branch also performs dual pooling operations and obtains a two-dimensional attention map through a 7×7 convolutional layer. The enhanced features are obtained by multiplying the feature map with the two attention weights.
[0078] Finally, the information entropy of feature maps at different scales is calculated, and the fusion weights are dynamically allocated according to the entropy values. The larger the information entropy, the richer the feature information, and the corresponding weight is larger. The enhanced feature maps are weighted and summed with the dynamic weights to obtain the final feature map that fuses multi-scale semantic information and detailed features.
[0079] This application can achieve:
[0080] Through the multi-scale feature extraction and fusion strategy, the multi-level feature information of the target is effectively extracted, improving the feature expression ability and detection accuracy. The bidirectional path enhancement and adaptive feature fusion mechanism are adopted to make full use of the complementarity of features at different levels, enhancing the discriminative ability of features. The channel and spatial dual attention mechanisms are introduced to highlight important feature information, suppress irrelevant background interference, and improve the robustness and accuracy of object detection.
[0081] In an optional implementation manner, based on the target feature map, the Openpose algorithm is used to extract the coordinate data of the hand key points and torso key points of the behavior target, and the partial affinity field technology is used to establish the association relationship between the hand key points and the torso key points. Constructing the pose skeleton includes:
[0082] Based on the target feature map, through a two-branch convolutional neural network, a confidence map is generated through the first branch, and a partial affinity field is generated through the second branch. The Gaussian kernel function is used to calculate the distance from the pixel points in the target feature map to the true position of the key points to obtain the initial confidence map of the multi-level key points. Based on the initial confidence map, the candidate positions of the key points of multiple human targets are determined;
[0083] Based on the candidate positions of the key points, the initial coordinates of the wrist key points are obtained, and the hand region is adaptively cropped with the initial coordinates of the wrist key points as the center to obtain the hand region of interest. The hand region of interest is input into the hand key point extraction network, and a hand key point heat map is generated through multi-layer convolution operations, and the accurate coordinates of the hand fine key points are extracted from the hand key point heat map;
[0084] A hierarchical detection strategy is adopted for the torso key points in the candidate positions of the key points. Based on the maximum response value of the initial confidence map, the reference position of the torso main key points is determined. The reference position is input into the position regression network to calculate the spatial offset of the secondary key points relative to the reference position. According to the spatial offset and the reference position, the accurate position of the torso secondary key points is obtained;
[0085] Based on the accurate coordinates of the hand fine key points and the accurate positions of the torso secondary key points, a bidirectional partial affinity field is constructed. The integral value of the feature vector of the forward partial affinity field is calculated along the key point connection direction, and the integral value of the feature vector of the reverse partial affinity field is calculated along the reverse direction of the key point connection. The forward partial affinity field and the reverse partial affinity field are fused through an adaptive direction weight to obtain an enhanced partial affinity field;
[0086] The enhanced partial affinity field is used to calculate the connection relationship scores between adjacent key points, and the optimal skeleton topology structure is obtained by maximizing the global accumulation of the connection relationship scores. A temporal smoothing constraint is applied to the optimal skeleton topology structure, and the current frame posture is adaptively weighted fused with the previous frame historical posture and motion predicted posture to obtain a smoothed posture skeleton that is consistent in time and space. The smoothed posture skeleton contains a stable connection relationship between the hand and torso key points.
[0087] The human posture estimation method based on the target feature map first uses a two-branch convolutional neural network for feature extraction. The first branch network contains multiple convolutional layers and pooling layers, and the output size is a confidence map with a width multiplied by a height multiplied by the number of channels. Each channel of the confidence map corresponds to a key point type. The pixels in the target feature map are processed by the Gaussian kernel function, and the distance from the pixel to the true position of the key point is calculated to generate an initial confidence map. The non-maximum suppression method is used on the initial confidence map to extract the candidate positions of the key points.
[0088] For the precise positioning of the key points of the hand, the hand region is adaptively cropped based on the candidate position of the wrist key point. The size of the cropped area is dynamically adjusted according to the image features of the area around the wrist key point. Generally, a square area centered on the wrist key point is taken, with a side length of 2 to 3 times the width of the wrist. The cropped hand region of interest is input into a dedicated hand key point extraction network, which adopts a stacked hourglass network structure and includes downsampling and upsampling processes. A heat map of the hand key points is generated through multi-layer convolution operations, and each channel corresponds to a hand key point. The sub-pixel positioning method is used on the heat map to extract the precise coordinates of the hand key points.
[0089] The trunk key points are located using a hierarchical detection strategy. First, the trunk key points, including the neck, shoulder and hip key points, are determined by the maximum response value on the initial confidence map. These key points are input into the position regression network as the reference position, which predicts the spatial offset of the secondary key points relative to the reference position. The position regression network adopts a residual structure and outputs the x and y direction components of the offset. The predicted offset is superimposed on the reference position to obtain the precise position of the secondary trunk key points.
[0090] After obtaining the precise positions of the key points of the hands and torso, a bidirectional partial affinity field is constructed to establish the connection relationship between the key points. Along the connection direction between the key point pairs, the feature vector is integrated in this direction to obtain the forward partial affinity field value. Similarly, the integration is performed in the reverse direction to obtain the reverse partial affinity field value. The weights of the forward and reverse fields are adaptively adjusted according to the complexity of the image content, and they are fused to obtain the enhanced partial affinity field.
[0091] Finally, the connection relationship scores between adjacent key points are calculated using the enhanced partial affinity fields. The Hungarian algorithm is employed to maximize the sum of scores for all possible connections, resulting in the optimal skeleton topology. To ensure the temporal consistency of the pose estimation results, the pose of the current frame is weighted and fused with the historical pose of the previous frame and the predicted motion pose based on the historical pose. The weights are adaptively determined according to the confidence levels of the poses of each frame, and finally a temporally and spatially consistent smoothed pose skeleton is obtained.
[0092] This application can achieve:
[0093] By combining a two-branch convolutional neural network with a Gaussian kernel function, the candidate positions of key points are extracted, improving the accuracy and robustness of the initial key point localization. An adaptive cropping strategy is adopted to accurately locate the hand region, effectively solving the scale variation problem in hand key point detection.
[0094] The hierarchical torso key point detection strategy makes full use of the prior knowledge of the human body structure. By predicting the spatial offset through a position regression network, the accuracy of torso key point localization is significantly improved. The construction method of the bidirectional partial affinity fields enhances the expression ability of the connection relationship between key points.
[0095] The temporal smoothing constraint and the adaptive weight fusion mechanism effectively suppress the jitter of the pose estimation results, ensuring the stability of the pose skeleton in continuous video sequences. The overall solution has strong environmental adaptability and high real-time performance in complex scenarios.
[0096] In an optional implementation manner, a hierarchical detection strategy is adopted for the torso key points among the candidate positions of the key points. Based on the maximum response value of the initial confidence map, the reference position of the torso main key point is determined. The reference position is input into a position regression network to calculate the spatial offset of the secondary key point relative to the reference position. The accurate position of the torso secondary key point is obtained according to the spatial offset and the reference position, including:
[0097] The confidence response map is obtained by weighted summation of the heatmap responses of multiple channels. Based on the confidence response map, a response threshold is set to screen and obtain the candidate region of the main key point, and the candidate region of the main key point is subjected to Gaussian smoothing processing to obtain a smoothed response map;
[0098] In the smoothed response map, the maximum response position is determined as the reference position of the main key point, and the sub-pixel offset is calculated using the response differences in the horizontal and vertical directions. The accurate coordinates of the main key point are obtained by adding the reference position and the sub-pixel offset;
[0099] Construct a multi-level feature pyramid network, perform upsampling operations and convolutional operations on the input feature map to obtain feature maps of different levels, concatenate the feature maps of different levels and input them into the position regression network to predict the spatial offset of the secondary key points relative to the coordinates of the refined backbone key points;
[0100] Establish bone length constraint conditions and joint angle constraint conditions based on human anatomical features, add the coordinates of the refined backbone key points and the spatial offset to obtain the initial position of the secondary key points, and perform constraint optimization on the initial position of the secondary key points to obtain the optimized position of the secondary key points that satisfies the bone length constraint conditions and the joint angle constraint conditions;
[0101] Introduce temporal smoothing constraints, calculate the displacement difference between the secondary key point positions in the current frame and the secondary key point positions in the previous two frames, calculate the motion speed difference between the current frame and the previous frame, use the spatial positioning error, displacement difference and speed difference as optimization objectives, and perform temporal consistency optimization on the optimized secondary key point positions to obtain the final accurate positions of the secondary key points.
[0102] First, generate multi-channel heatmaps for the input images. Extract features from the input images through a convolutional neural network to generate heatmaps of 17 channels of key points, with each channel corresponding to a type of human key point. Perform weighted summation on these heatmaps, and the weights can be set according to the importance of the key points. For example, the weight of the trunk part can be set to 0.6, and the weight of the limbs part can be set to 0.4. After weighted summation, obtain the confidence response map, and use a response threshold of 0.3 for screening to obtain candidate regions that may contain backbone key points. Smooth these candidate regions using a Gaussian kernel with a standard deviation of 1.5 to reduce the influence of noise.
[0103] In the smoothed response map, search for the position with the maximum response value as the reference position of the backbone key point. To obtain more accurate coordinates, calculate the response differences in the horizontal and vertical directions respectively. For example, if the response value at a certain position is 0.8, and the response values of its adjacent positions on the left and right are 0.7 and 0.6 respectively, then the sub-pixel level offset in the horizontal direction of this point can be calculated as 0.2 pixels according to the response difference. Similarly, calculate the offset in the vertical direction, and add the reference position and the offset to obtain the refined coordinates of the backbone key point.
[0104] Construct a feature pyramid network for secondary key point detection. This network contains 4 levels, corresponding to 1 / 2, 1 / 4, 1 / 8, and 1 / 16 scales of the input feature map respectively. Each level extracts features through 3x3 convolutions, and then uses bilinear interpolation for 2x upsampling and concatenates with the feature map of the previous level. The multi-scale feature maps obtained in this way are input into the position regression network, which consists of 3 fully connected layers and outputs the spatial offset of the secondary key points relative to the backbone key points.
[0105] Set constraint conditions based on human anatomical features. In terms of bone length constraints, the upper arm length range is set to 20 - 35 cm, and the forearm is 22 - 32 cm. In terms of joint angle constraints, the elbow joint range of motion is 0 - 145 degrees, and the shoulder joint is -40 to 180 degrees horizontally. Add the trunk key point coordinates and the predicted offset to obtain the initial position of the secondary key points, and then iteratively optimize them to meet these constraint conditions.
[0106] Finally, introduce temporal constraints to achieve smoothing. Calculate the position differences between the current frame and the previous two frames. For example, if the coordinates of a certain key point in the current frame are (100, 100), the previous frame is (98, 99), and the frame before the previous frame is (95, 97), then the displacement difference is 5 pixels. At the same time, calculate the speed change. If the speed in the previous frame is 3 pixels / frame and the current frame is 4 pixels / frame, the speed difference is 1 pixel / frame. Optimize by comprehensively considering the spatial positioning error, displacement difference, and speed difference to obtain the final position of the secondary key points.
[0107] This application can achieve:
[0108] Through a hierarchical detection strategy and multi-scale feature extraction, the accuracy of trunk key point detection has been significantly improved, the detection error has been reduced, and the key point positioning has become more accurate and reliable.
[0109] Introducing human anatomical constraints and temporal smoothing constraints effectively suppresses the jumping and drifting of key points, enhances the stability of the detection results, and makes the pose estimation more in line with human motion characteristics.
[0110] Adopting sub-pixel level precise positioning and multi-layer optimization strategies improves the robustness of the algorithm, can adapt to different scenarios and pose changes, and has strong practical value and application prospects for promotion.
[0111] In an alternative embodiment, wavelet transform and spatio-temporal attention mechanism are used to extract micro-expression features from the facial region of the behavior target, and the pose skeleton and the micro-expression features are input into a deep learning decision classification model, and the output behavior feature vector including the distance data between the hand and the trunk, the hand movement trajectory data, and the behavior time window data includes:
[0112] Perform multi-scale two-dimensional discrete wavelet transform on the facial region image of the behavior target to obtain wavelet coefficients at different scales, calculate the wavelet energy value at each scale based on the wavelet coefficients and determine the energy ratio to obtain the initial feature map of the facial region;
[0113] Input the initial feature map into the dual-stream spatio-temporal attention network. Perform average pooling and max pooling on the initial feature map through the spatial attention branch to obtain the spatial attention weights, and perform 3D convolution operations on the initial feature map through the temporal attention branch to obtain the temporal attention weights. Multiply the initial feature map with the spatial attention weights and the temporal attention weights to obtain the enhanced feature map;
[0114] Calculate the Euclidean distance between the hand key points and the torso key points in the pose skeleton to construct the hand-torso distance matrix, and calculate the speed, acceleration, and direction angle of the hand movement based on the continuous frame position information of the hand key points to obtain the hand movement trajectory descriptor;
[0115] Adopt a sliding time window to perform feature aggregation on the enhanced feature map, calculate the temporal correlation of the feature sequence within the time window, and perform weighted summation of the feature sequence and the time weights to obtain the time window feature;
[0116] Concatenate the enhanced feature map, the hand-torso distance matrix, the hand movement trajectory descriptor, and the time window feature to construct a multi-modal feature vector. Input the multi-modal feature vector into a multi-layer perceptron and a bidirectional long short-term memory network simultaneously for feature fusion, and obtain the behavior category probability distribution through a softmax classifier. Optimize and train the classification model based on the cross-entropy loss function and the regularization term to obtain the final behavior classification result.
[0117] The present invention proposes a behavior recognition method based on multi-modal feature fusion, which combines wavelet transform, spatio-temporal attention mechanism, pose skeleton analysis, and deep learning technology to achieve accurate recognition of behavior targets. The specific implementation steps are as follows:
[0118] First, extract features from the facial region of the behavior target. Process the facial region image using multi-scale two-dimensional discrete wavelet transform to obtain wavelet coefficients at different scales. Specifically, Haar wavelet or Daubechies wavelet, etc., can be selected as the basis function, and the image is decomposed into 3 or 4 layers. For each scale, calculate the wavelet energy value and determine the energy proportion. For example, for a 256x256 pixel facial image, after 3-layer wavelet decomposition, 10 sub-band images can be obtained. Calculate the energy value of each sub-band image and normalize it to obtain the energy proportion. Combine these features into an initial feature map, whose dimension may be 10x32x32.
[0119] Next, the initial feature map is input into the dual-stream spatio-temporal attention network for enhancement. In the spatial attention branch, average pooling and max pooling are respectively performed on the initial feature map to obtain two feature maps of 32x32. Then, they are concatenated and passed through a 1x1 convolutional layer to obtain the spatial attention weights. In the temporal attention branch, 3D convolution is used to process the initial feature maps of consecutive multiple frames. For example, a convolutional kernel of 16 frames with a size of 3x3x3 and a stride of 1 is used to obtain the temporal attention weights. The initial feature map is multiplied by these two attention weights to obtain the enhanced feature map, whose dimension remains unchanged, being 10x32x32.
[0120] Meanwhile, the pose skeleton is analyzed. The Euclidean distances between hand key points (such as the wrist, palm center) and torso key points (such as the shoulder, chest center) are calculated to construct a hand-torso distance matrix. For example, for each frame of the image, a 4x2 distance matrix can be obtained, representing the distances between the left and right hands and the shoulders and chest. Based on the position information of the hand key points in consecutive frames, the speed, acceleration, and direction angle of the hand movement are calculated. Specifically, a 5-frame sliding window can be used to calculate the displacement, speed change, and angle change of the center frame relative to the front and back frames, obtaining a 15-dimensional hand movement trajectory descriptor.
[0121] To capture the temporal information, a sliding time window is adopted to perform feature aggregation on the enhanced feature map. 16 frames are selected as the time window size, and the temporal correlation is calculated for the feature sequence within each time window. The self-attention mechanism or the temporal convolutional network can be used to model the temporal relationship. Then, the feature sequence is weighted and summed with the learned temporal weights to obtain the 160-dimensional time window features.
[0122] The enhanced feature map (10x32x32), the hand-torso distance matrix (4x2), the hand movement trajectory descriptor (15-dimensional), and the time window features (160-dimensional) are concatenated to construct a multi-modal feature vector. This feature vector contains multi-faceted information such as facial micro-expressions, body postures, and movement trajectories.
[0123] Finally, the multi-modal feature vector is simultaneously input into a multi-layer perceptron (MLP) and a bidirectional long short-term memory network (Bi-LSTM) for feature fusion. The MLP can contain 3 fully connected layers, with the number of neurons in each layer being 512, 256, and 128 respectively. The Bi-LSTM can use 128 hidden units. The output features of the MLP and the Bi-LSTM are concatenated, reduced to 64 dimensions through a fully connected layer, and then input into a softmax classifier to obtain the probability distribution of the behavior categories.
[0124] During the training phase, the cross-entropy loss function is used as the objective function, and an L2 regularization term is added to prevent overfitting. The Adam optimizer is adopted to update the model parameters, with the initial learning rate set to 0.001 and decaying by 10% every 50 epochs. The training dataset can include various daily behaviors such as walking, running, sitting, standing up, etc., and each type of behavior contains at least 1000 samples. After achieving the best performance on the validation set, the model parameters are saved as the final behavior classification model.
[0125] This application can achieve:
[0126] By combining wavelet transform and spatio-temporal attention mechanism, this method can effectively extract micro-expression features of the facial region, capture subtle expression changes, and improve the accuracy and robustness of behavior recognition.
[0127] By introducing pose skeleton analysis and hand movement trajectory description, this method can comprehensively depict the motion characteristics of the human body, considering not only static postures but also dynamic motion information, thus more accurately distinguishing different types of behaviors.
[0128] Adopting multi-modal feature fusion and deep learning models, this method can adaptively learn the correlation between different features, make full use of various information sources, achieve more intelligent and accurate behavior recognition, and is applicable to complex practical application scenarios.
[0129] In an alternative implementation, the behavior feature vector is input into the theft behavior scoring mechanism analysis engine for behavior determination. The theft behavior scoring mechanism analysis engine fuses a convolutional neural network and a Transformer structure, and calculates the similarity between the behavior feature vector and a preset behavior template and generates a confidence score, including:
[0130] The input behavior feature vector is reorganized into a spatio-temporal feature matrix, and sine position encoding and cosine position encoding are embedded in the spatio-temporal feature matrix to obtain an encoded spatio-temporal feature matrix;
[0131] Parallel convolutional operations are performed on the encoded spatio-temporal feature matrix using convolutional branches with different kernel sizes to obtain multi-scale convolutional features, and the multi-scale convolutional features are weighted and fused to obtain fused convolutional features;
[0132] The encoded spatio-temporal feature matrix is respectively subjected to query matrix transformation, key matrix transformation, and value matrix transformation to obtain a query matrix, a key matrix, and a value matrix. The attention weights are calculated based on the product of the query matrix and the key matrix, the attention weights are multiplied by the value matrix to obtain attention features, and multiple attention features are concatenated to obtain a multi-head attention output feature;
[0133] Calculate the cosine similarity between the fused convolutional feature and the preset behavior template to obtain a similarity feature, and perform dynamic time warping calculation on the fused convolutional feature and the preset behavior template to obtain a time warping distance feature;
[0134] Perform weighted summation on the fused convolutional feature, the multi-head attention output feature, the similarity feature, and the time warping distance feature to obtain an integrated feature, and perform sigmoid activation operation on the integrated feature to obtain an initial confidence score;
[0135] Perform exponential smoothing processing on the initial confidence score to obtain a smoothed confidence score, calculate the temporal consistency of the smoothed confidence score, and use the weighted sum of the smoothed confidence score and the temporal consistency as the final theft behavior determination result.
[0136] The present invention provides a theft behavior scoring mechanism analysis engine, which integrates a convolutional neural network and a Transformer structure, and is used to calculate the similarity between a behavior feature vector and a preset behavior template and generate a confidence score. The specific implementation process is as follows:
[0137] First, reorganize the input behavior feature vector into a spatio-temporal feature matrix. Assume that the input behavior feature vector contains 100 features, which can be reorganized into a 10x10 spatio-temporal feature matrix. Embed sine position encoding and cosine position encoding in this matrix to enhance the model's perception ability of feature positions. Specifically, position encodings can be generated using sine and cosine functions with different frequencies, and then added to the original feature matrix to obtain the encoded spatio-temporal feature matrix.
[0138] Next, use convolutional branches with different kernel sizes to perform parallel convolutional operations on the encoded spatio-temporal feature matrix to obtain multi-scale convolutional features. For example, convolutional operations can be performed on the feature matrix using 3x3, 5x5, and 7x7 convolutional kernels respectively, and the number of output channels of each convolutional kernel can be set to 64. Then, perform weighted fusion on these multi-scale convolutional features to obtain a fused convolutional feature. Weighted fusion can be achieved by learning the weight coefficients of each scale, for example, using a 1x1 convolutional layer to learn the weights.
[0139] At the same time, perform query matrix transformation, key matrix transformation, and value matrix transformation on the encoded spatio-temporal feature matrix respectively to obtain a query matrix, a key matrix, and a value matrix. These transformations can be implemented through fully connected layers, and the output dimension can be set to 64. Calculate the attention weights based on the product of the query matrix and the key matrix, multiply the attention weights by the value matrix to obtain an attention feature. Eight attention heads can be set, with the dimension of each head being 8. Finally, splice multiple attention features to obtain a multi-head attention output feature with a dimension of 64.
[0140] Then, calculate the cosine similarity between the fused convolutional features and the preset behavior templates to obtain the similarity features. The preset behavior templates can be a set of typical theft behavior feature vectors. By calculating the similarity between the current features and these templates, the abnormality degree of the behavior can be judged. At the same time, perform dynamic time warping calculation on the fused convolutional features and the preset behavior templates to obtain the time warping distance features. Dynamic time warping can handle the comparison between sequences of different lengths and can capture local deformations and global trends in time series.
[0141] Perform weighted summation on the fused convolutional features, multi-head attention output features, similarity features, and time warping distance features to obtain the integrated features. The weighting coefficients can be learned during the model training process and can be initially set to equal weights. Perform sigmoid activation operation on the integrated features to obtain the initial confidence score, whose value range is between 0 and 1.
[0142] To improve the stability of the judgment result, perform exponential smoothing on the initial confidence score to obtain the smoothed confidence score. Exponential smoothing can select an appropriate smoothing coefficient, such as 0.8, to balance the influence of historical data and current data. Then calculate the temporal consistency of the smoothed confidence score, which can be measured by calculating the variance of the confidence scores at consecutive multiple time points. Finally, use the weighted sum of the smoothed confidence score and the temporal consistency as the final theft behavior judgment result. The weighting coefficients can be adjusted according to the actual application scenario. For example, the weight of the confidence score can be set to 0.7, and the weight of the temporal consistency can be set to 0.3.
[0143] In practical applications, a threshold can be set, such as 0.8. When the final judgment result exceeds this threshold, the system will issue a theft behavior alarm. At the same time, different response strategies can be set according to different score intervals. For example, when the score is between 0.6 and 0.8, key monitoring is carried out, and no action is taken when the score is lower than 0.6.
[0144] The theft behavior scoring mechanism analysis engine of the present invention can achieve:
[0145] First, by integrating the convolutional neural network and the Transformer structure, the advantages of both models are fully utilized. The convolutional neural network can effectively extract local features and multi-scale information, while the self-attention mechanism of the Transformer can capture long-range dependencies, enabling the model to comprehensively analyze behavior features and improving the accuracy and robustness of theft behavior recognition.
[0146] Secondly, technologies such as dynamic time warping and exponential smoothing are introduced, effectively addressing the issues of inconsistency and noise in time-series data. Dynamic time warping can adapt to behavior sequences of different lengths, while exponential smoothing can smooth short-term fluctuations, enhancing the model's sensitivity to long-term trends and thus strengthening the stability and reliability of theft behavior determination.
[0147] Finally, a multi-feature fusion and weighting strategy is adopted, comprehensively considering convolutional features, attention features, similarity features, and time warping features, and introducing time-series consistency evaluation, making the determination results more comprehensive and credible. This multi-dimensional analysis method significantly reduces the misjudgment rate while improving the model's ability to identify complex theft behaviors, providing a more reliable basis for decision-making in practical applications.
[0148] Figure 2 FIG. is a schematic structural diagram of a theft behavior recognition system based on computer vision according to an embodiment of the present invention, as Figure 2 shown, the system includes:
[0149] A first unit for collecting a video data stream, converting the frame rate of the video data stream to obtain a video stream with a target frame rate, applying Gaussian filtering and median filtering to the video stream with the target frame rate for noise elimination, and using a threshold segmentation algorithm to perform image segmentation on the filtered video stream to obtain preprocessed video frames;
[0150] A second unit for using the YOLOv11 deep learning object detection algorithm to perform object detection and localization on the preprocessed video frames to obtain bounding box information of behavior targets, inputting the bounding box information into a feature pyramid network for multi-scale feature extraction, and fusing high-level semantic features and low-level detail features through a path aggregation network to obtain a target feature map; based on the target feature map, using the Openpose algorithm to extract coordinate data of hand key points and torso key points of the behavior target, and using the part affinity fields technique to establish an association relationship between the hand key points and the torso key points to construct a pose skeleton, using wavelet transform and spatio-temporal attention mechanism to extract micro-expression features from the facial region of the behavior target, and inputting the pose skeleton and the micro-expression features into a deep learning decision classification model to output a behavior feature vector including distance data between the hand and the torso, hand movement trajectory data, and behavior time window data;
[0151] A third unit is configured to input the behavioral feature vector into a theft behavior scoring mechanism analysis engine for behavior determination. The theft behavior scoring mechanism analysis engine integrates a convolutional neural network and a Transformer structure, calculates the similarity between the behavioral feature vector and a preset behavior template, and generates a confidence score. When the confidence score is greater than 0.8, it is determined as a theft behavior and an alarm signal is triggered. When the confidence score is between 0.5 and 0.8, it is determined as an abnormal behavior and a tracking mark is recorded. When the confidence score is less than 0.5, it is determined as a normal behavior.
[0152] In a third aspect of the embodiments of the present invention,
[0153] there is provided an electronic device, comprising:
[0154] a processor;
[0155] a memory for storing instructions executable by the processor;
[0156] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0157] In a fourth aspect of the embodiments of the present invention,
[0158] there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0159] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0160] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A theft behavior recognition method based on computer vision, characterized in that: include: Collecting a video data stream, performing frame rate conversion on the video data stream to obtain a video stream with a target frame rate, applying Gaussian filtering and median filtering to the video stream with the target frame rate to eliminate noise, and performing image segmentation on the filtered video stream using a threshold segmentation algorithm to obtain a preprocessed video frame; Using the YOLOv11 deep learning target detection algorithm to perform target detection and positioning on the preprocessed video frame to obtain bounding box information of the behavioral target, inputting the bounding box information into a feature pyramid network to extract multi-scale features, and fusing high-level semantic features and low-level detail features through a path aggregation network to obtain a target feature map; Based on the target feature map, the Openpose algorithm is used to extract the coordinate data of the hand key points and the trunk key points of the behavior target, and the association relationship between the hand key points and the trunk key points is established by using the partial affinity field technology to construct a posture skeleton, and the wavelet transform and spatiotemporal attention mechanism are used to extract the features of the facial area of the behavior target to obtain micro-expression features, and the posture skeleton and the micro-expression features are input into the deep learning decision classification model, and the behavior feature vector containing the distance data between the hand and the trunk, the hand movement trajectory data, and the behavior time window data is output; The behavior feature vector is input into the theft behavior scoring mechanism analysis engine for behavior determination. The theft behavior scoring mechanism analysis engine integrates the convolutional neural network and the Transformer structure, calculates the similarity between the behavior feature vector and the preset behavior template and generates a confidence score. When the confidence score is greater than 0.8, it is determined to be a theft behavior and triggers an alarm signal. When the confidence score is between 0.5 and 0.8, it is determined to be an abnormal behavior and a tracking mark is recorded. When the confidence score is less than 0.5, it is determined to be a normal behavior; Wavelet transform and spatiotemporal attention mechanism are used to extract features from the facial area of the behavior target to obtain micro-expression features, and the posture skeleton and the micro-expression features are input into a deep learning decision classification model, and the output behavior feature vector including the distance data between the hand and the torso, the hand movement trajectory data, and the behavior time window data includes: Performing a multi-scale two-dimensional discrete wavelet transform on the facial region image of the behavioral target to obtain wavelet coefficients of different scales, calculating the wavelet energy value at each scale based on the wavelet coefficients and determining the energy proportion, thereby obtaining an initial feature map of the facial region; Input the initial feature map into a dual-stream spatiotemporal attention network, perform average pooling and maximum pooling on the initial feature map through the spatial attention branch to obtain a spatial attention weight, perform a three-dimensional convolution operation on the initial feature map through the temporal attention branch to obtain a temporal attention weight, and multiply the initial feature map with the spatial attention weight and the temporal attention weight to obtain an enhanced feature map; The Euclidean distance between the hand key points and the trunk key points in the posture skeleton is calculated to construct the hand-torso distance matrix. The speed, acceleration and direction angle of the hand movement are calculated based on the continuous frame position information of the hand key points to obtain the hand movement trajectory descriptor. A sliding time window is used to perform feature aggregation on the enhanced feature map, the temporal correlation of the feature sequence in the time window is calculated, and the feature sequence is weighted and summed with the time weight to obtain the time window feature; The enhanced feature map, the hand-torso distance matrix, the hand motion trajectory descriptor and the time window feature are concatenated to construct a multimodal feature vector, and the multimodal feature vector is simultaneously input into a multilayer perceptron and a bidirectional long short-term memory network for feature fusion. The fused features are passed through a softmax classifier to obtain the behavior category probability distribution, and the classification model is optimized and trained based on the cross entropy loss function and the regularization term to obtain the final behavior classification result.
2. The method according to claim 1, characterized in that The preprocessed video frames are subjected to target detection and positioning using the YOLOv11 deep learning target detection algorithm to obtain bounding box information of the behavioral target, and the bounding box information is input into the feature pyramid network for multi-scale feature extraction and the target feature map is obtained by fusing high-level semantic features and low-level detail features through the path aggregation network, including: Divide the preprocessed video frame into a number of grid units, perform bounding box prediction on the grid units, and obtain the center coordinates, width, height and confidence of the bounding box, where the confidence is calculated by multiplying the probability of the object existing in the box by the intersection and union ratio of the predicted box and the real box; A feature pyramid network is constructed based on the bounding box, the preprocessed video frame is downsampled by a convolution layer to obtain five bottom-level feature maps of different scales, the sizes of the five bottom-level feature maps of different scales are halved successively, the high-level feature map is upsampled and one-dimensionally convolved with the horizontally connected feature map of the same scale to obtain a pyramid feature map; Input the pyramid feature map into a path aggregation network, construct a top-down enhancement path and a bottom-up enhancement path in the path aggregation network, the top-down enhancement path obtains a top-down feature map through upsampling and feature concatenation, the bottom-up enhancement path obtains a bottom-up feature map through downsampling and feature concatenation, and the top-down feature map and the bottom-up feature map are fused through a learnable fusion weight coefficient to obtain a fused feature map; The fused feature map is respectively subjected to attention enhancement in the channel dimension and the spatial dimension. The attention enhancement in the channel dimension is performed by performing average pooling and maximum pooling on the fused feature map and obtaining a channel attention weight through a multilayer perceptron. The attention enhancement in the spatial dimension is performed by performing average pooling and maximum pooling on the fused feature map and obtaining a spatial attention weight through a convolution operation. The fused feature map is multiplied by the channel attention weight and the spatial attention weight to obtain an attention enhancement feature map. The information entropy of the attention enhanced feature map at different scales is calculated, and the dynamic fusion weights of different scales are calculated based on the information entropy. The attention enhanced feature map and the corresponding dynamic fusion weights are weightedly summed to obtain a final feature map, wherein the final feature map contains multi-scale semantic information and detail features of the target.
3. The method according to claim 1, characterized in that Based on the target feature graph, the coordinate data of the hand key points and the trunk key points of the behavior target are extracted by using the Openpose algorithm, and the association relationship between the hand key points and the trunk key points is established by using the partial affinity field technology to construct the posture skeleton, including: Based on the target feature map, a two-branch convolutional neural network is used to generate a confidence map through the first branch, and a partial affinity field is generated through the second branch. The distance from the pixel point in the target feature map to the real position of the key point is calculated using a Gaussian kernel function to obtain an initial confidence map of multi-level key points, and the candidate positions of the key points of multiple human targets are determined based on the initial confidence map; Based on the candidate positions of the key points, the initial coordinates of the wrist key points are obtained, the hand region is adaptively cropped with the initial coordinates of the wrist key points as the center to obtain the hand region of interest, the hand region of interest is input into the hand key point extraction network, a hand key point heat map is generated through multi-layer convolution operations, and the accurate coordinates of the hand fine key points are extracted from the hand key point heat map; A hierarchical detection strategy is adopted for the trunk key points in the key point candidate positions, a reference position of the trunk main key point is determined based on the maximum response value of the initial confidence map, the reference position is input into a position regression network to calculate the spatial offset of the secondary key point relative to the reference position, and the precise position of the trunk secondary key point is obtained according to the spatial offset and the reference position; A bidirectional partial affinity field is constructed based on the accurate coordinates of the hand fine key points and the precise positions of the trunk secondary key points, a feature vector integral value of the forward partial affinity field is calculated along the key point connection direction, a feature vector integral value of the reverse partial affinity field is calculated along the reverse direction of the key point connection, and the forward partial affinity field and the reverse partial affinity field are fused through adaptive direction weights to obtain an enhanced partial affinity field; The enhanced partial affinity field is used to calculate the connection relationship scores between adjacent key points, and the optimal skeleton topology structure is obtained by maximizing the global accumulation of the connection relationship scores. A temporal smoothing constraint is applied to the optimal skeleton topology structure, and the current frame posture is adaptively weighted fused with the previous frame historical posture and motion predicted posture to obtain a smoothed posture skeleton that is consistent in time and space. The smoothed posture skeleton contains a stable connection relationship between the hand and torso key points.
4. The method according to claim 3, characterized in that A hierarchical detection strategy is adopted for the trunk key points in the key point candidate positions, a reference position of the trunk trunk key points is determined based on the maximum response value of the initial confidence map, the reference position is input into the position regression network to calculate the spatial offset of the secondary key points relative to the reference position, and the precise position of the trunk secondary key points is obtained according to the spatial offset and the reference position, including: Performing weighted summation on the heat map responses of multiple channels to obtain a confidence response map, setting a response threshold based on the confidence response map to screen and obtain a candidate region of a backbone key point, and performing Gaussian smoothing on the candidate region of the backbone key point to obtain a smoothed response map; Determine the maximum response position in the smooth response map as the reference position of the backbone key point, calculate the sub-pixel offset using the response difference in the horizontal direction and the vertical direction, and add the reference position and the sub-pixel offset to obtain the refined backbone key point coordinates; Construct a multi-level feature pyramid network, perform upsampling and convolution operations on the input feature map to obtain feature maps of different levels, cascade the feature maps of different levels into a position regression network, and predict the spatial offset of the secondary key point relative to the refined main key point coordinates; Establishing bone length constraints and joint angle constraints based on human anatomical features, adding the refined trunk key point coordinates to the spatial offset to obtain the secondary key point initial position, and performing constraint optimization on the secondary key point initial position to obtain the optimized secondary key point position that satisfies the bone length constraints and the joint angle constraints; A temporal smoothness constraint is introduced to calculate the displacement difference between the secondary key point position of the current frame and the secondary key point position of the previous two frames, and the movement speed difference between the current frame and the previous frame. The spatial positioning error, displacement difference and speed difference are used as optimization targets, and the timing consistency optimization is performed on the optimized secondary key point position to obtain the final precise position of the secondary key point.
5. The method according to claim 1, characterized in that Inputting the behavior feature vector into the theft behavior scoring mechanism analysis engine for behavior determination, the theft behavior scoring mechanism analysis engine integrates the convolutional neural network and the Transformer structure, calculates the similarity between the behavior feature vector and the preset behavior template and generates a confidence score including: Reorganize the input behavior feature vector into a spatiotemporal feature matrix, embed sine position coding and cosine position coding into the spatiotemporal feature matrix to obtain a coded spatiotemporal feature matrix; Using convolution branches with different convolution kernel sizes to perform parallel convolution operations on the encoded spatiotemporal feature matrix to obtain multi-scale convolution features, and performing weighted fusion on the multi-scale convolution features to obtain fused convolution features; Performing query matrix transformation, key matrix transformation and value matrix transformation on the encoded spatiotemporal feature matrix to obtain a query matrix, a key matrix and a value matrix, calculating an attention weight based on the product of the query matrix and the key matrix, multiplying the attention weight by the value matrix to obtain an attention feature, and concatenating a plurality of the attention features to obtain a multi-head attention output feature; Calculate the cosine similarity between the fused convolution feature and the preset behavior template to obtain a similarity feature, and perform dynamic time warping calculation on the fused convolution feature and the preset behavior template to obtain a time warping distance feature; Performing weighted summation on the fused convolution feature, the multi-head attention output feature, the similarity feature and the time-warped distance feature to obtain an integrated feature, and performing a sigmoid activation operation on the integrated feature to obtain an initial confidence score; The initial confidence score is subjected to exponential smoothing to obtain a smoothed confidence score, the temporal consistency of the smoothed confidence score is calculated, and a weighted sum of the smoothed confidence score and the temporal consistency is taken as a final theft behavior determination result.
6. A theft behavior recognition system based on computer vision, used to implement the method according to any one of claims 1 to 5, characterized in that: include: The first unit is used to collect a video data stream, perform frame rate conversion on the video data stream to obtain a video stream with a target frame rate, apply Gaussian filtering and median filtering to the video stream with the target frame rate to eliminate noise, and use a threshold segmentation algorithm to perform image segmentation on the filtered video stream to obtain a preprocessed video frame; The second unit is used to detect and locate the target on the preprocessed video frame using the YOLOv11 deep learning target detection algorithm to obtain the bounding box information of the behavior target, input the bounding box information into the feature pyramid network to extract multi-scale features, and fuse the high-level semantic features and low-level detail features through the path aggregation network to obtain the target feature map; Based on the target feature map, the Openpose algorithm is used to extract the coordinate data of the hand key points and the trunk key points of the behavior target, and the association relationship between the hand key points and the trunk key points is established by using the partial affinity field technology to construct a posture skeleton, and the wavelet transform and spatiotemporal attention mechanism are used to extract the features of the facial area of the behavior target to obtain micro-expression features, and the posture skeleton and the micro-expression features are input into the deep learning decision classification model, and the behavior feature vector containing the distance data between the hand and the trunk, the hand movement trajectory data, and the behavior time window data is output; The third unit is used to input the behavior feature vector into the theft behavior scoring mechanism analysis engine for behavior judgment. The theft behavior scoring mechanism analysis engine integrates the convolutional neural network and the Transformer structure, calculates the similarity between the behavior feature vector and the preset behavior template and generates a confidence score. When the confidence score is greater than 0.8, it is judged as theft behavior and triggers an alarm signal. When the confidence score is between 0.5 and 0.8, it is judged as abnormal behavior and a tracking mark is recorded. When the confidence score is less than 0.5, it is judged as normal behavior.
7. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 5.
8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Crowd behavior analysis system and method
CN117351405A
Computer automated interactive activity recognition based on keypoint detection
US20220101556A1