Human uncivilized behavior detection method integrating non-uniform sampling and feature enhancement

By introducing non-uniform sampling and YoloX detection in the SlowFast network combined with a cascaded pooling three-dimensional spatial pyramid feature enhancement module, the imbalance between computational complexity and detection accuracy in existing technologies is solved, and the accuracy and speed of identifying uncivilized behavior are improved.

CN117392753BActive Publication Date: 2025-09-09SOUTHWEST PETROLEUM UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202311386768.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-24
Publication Date
2025-09-09
Estimated Expiration
2043-10-24

AI Technical Summary

Technical Problem

Existing methods for recognizing abnormal human movements have difficulty balancing computational complexity and detection accuracy, and uniform sampling methods have problems of false detection and low detection accuracy of local limb behaviors when recognizing similar movements.

Method used

Based on the SlowFast network, combined with non-uniform sampling and the lightweight one-stage target detection network YoloX, it enhances the feature expression ability and improves the recognition accuracy of uncivilized behavior by fusing the cascade pooling three-dimensional spatial pyramid feature enhancement module with shallow features.

Benefits of technology

It effectively improves the detection accuracy of similar behaviors and the speed of network inference, reduces the amount of calculation, and achieves efficient identification of uncivilized behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117392753B_ABST
    Figure CN117392753B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting uncivilized human behavior by integrating non-uniform sampling and feature enhancement, comprising the following steps: constructing an action recognition model based on a SlowFast network to obtain multi-scale spatiotemporal features with a hierarchical structure; proposing a non-uniform sampling method in the video frame acquisition stage to effectively enhance the network's attention to the detailed changing features of similar behaviors; and embedding the proposed cascade pooling three-dimensional spatial pyramid feature enhancement module that integrates shallow features behind the feature extraction network to enhance the applicability of features at different scales, effectively reduce the loss of action detail information in the feature extraction process and reduce the interference of background information, thereby achieving the effect of feature enhancement; and building a human position border detection model based on a one-stage lightweight target detection network YoloX, extracting key frames to realize personnel detection, using ROI extraction to obtain the spatiotemporal features of the corresponding area of ​​the human body and perform action classification to determine whether there is uncivilized behavior in the video. Experiments have shown that this method can effectively improve the spatiotemporal action detection effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of spatiotemporal action detection based on deep learning, and in particular to a neural network based on 3D convolution and an uncivilized behavior detection method. Background Art

[0002] With the widespread use of the internet, the use of various short video software, surveillance equipment, and video recordings has increased. Over time, a vast amount of video and image data has emerged and been disseminated. In recent years, short video software has become a primary platform for people to share their lives and interact with each other. While this enriches their leisure time, it can also easily foster uncivilized behavior. Regarding video information extraction and review, a significant amount of manual effort is required to review videos and determine sensitive content. This is time-consuming, labor-intensive, subjective, and inefficient. Using computer vision and artificial intelligence technologies to detect uncivilized behavior in video content can quickly extract key information, improve review efficiency, and nip the occurrence and spread of uncivilized behavior at the source, thereby enhancing the ability to maintain social harmony.

[0003] Reference 1 (Tan Shuqiu, Tang Guofang, Tu Yuanya, et al. Abnormal student behavior detection system under classroom monitoring [J]. Computer Engineering and Applications, 2022, 58(07): 176-184.) expands the YOLOv3 shallow network to improve the network's attention to image edges or small target objects, but lacks the use of temporal information. Reference 2 (Carreira J, Zisserman A. Quo vadis, action recognition? a new model and the kinetics dataset [C] / / proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 6299-6308.) expands the 2D convolutional neural network pre-training model to a 3D structure, solving the problem of the lack of a good 3D convolutional neural network pre-training model at the time, and combines optical flow feature input to effectively obtain temporal information between video frames. However, due to the relatively dense extraction of optical flow features, the distinction between similar actions is confused. Reference 3 (Yang Le, Li Yifan, Chen Xi, et al. Detection of illegal behaviors in power production environment based on ST-SlowFast [J]. Smart Power, 2023, 51(06): 71-77.) An auxiliary channel based on the spatiotemporal attention mechanism was constructed based on SlowFast to improve the accuracy of illegal behavior detection in power production environment. This method realizes the extraction and enhancement of detailed features between sampling streams with different frame rates, and can better capture the changes in action details when the behavior occurs, but there is a problem of insufficient extraction of high-level and low-level features.

[0004] Patent 1 (Yu Li, Chen Shuaichao, Cao Wenjun, et al. Two-stage, dual-channel distracted driving behavior recognition method based on keypoint detection [P]. Sichuan Province: CN116721405A, 2023-09-08.) uses a keypoint detection network to obtain keypoint features of the driver, uses ResNet-50 as the backbone network, and uses a GCN graph convolutional network to design a binary classifier, improving behavior recognition. However, this method cannot effectively utilize temporal information. Patent 2 (Zhang Fengquan, Cheng Jian, Zhou Feng, Wang Guiling. A multi-stage abnormal human motion detection method based on a residual network [P]. Beijing: CN114202803A, 2022-03-18.) uses a residual network to perform multi-stage abnormal human motion detection, continuously identifying the bounding boxes, positions, and sizes of human objects appearing in surveillance video instances, and then performing a weighted fusion of the anomaly scores for each surveillance video instance. This method can effectively detect surveillance anomalies, but requires continuous processing of multiple videos, resulting in a high computational load. Patent 3 (Hao Zixuan, Liu Mengyu, Tian Shuochen, et al. A method for identifying illegal behaviors in hot work operations [P]. Shandong Province: CN116721383A, 2023-09-08.) describes a method for identifying illegal behaviors in hot work operations. By training a YOLOv5 one-stage detection model that incorporates an attention mechanism and performing inference on key frames, the position information of each worker in each video frame is obtained. The Deepsort algorithm is used to track the workers, and an I3D two-stream network is used as an action recognition model to identify whether the workers have violated the rules. This method uses technology based on a two-stream 3D convolutional neural network, but utilizes multiple algorithms. The patent includes the use of attention mechanisms, YOLOv5, Deepsort, and I3D algorithms. In actual deployment, it faces complex model embedding problems and model computational overhead issues.

[0005] The above-mentioned methods for identifying abnormal human movements do not strike a good balance between computational effort and detection accuracy. In video understanding tasks, uniform sampling is often used to obtain input samples. While this method can effectively sample video frames, it also has certain drawbacks. Recognizing similar movements requires paying special attention to the detailed information between adjacent sampled frames with a shorter duration, thereby capturing the distinct characteristics of various behaviors and enabling discrimination. Because long-term contextual information is crucial for action recognition, the duration of the sample spanning video frames cannot be arbitrarily reduced. Reducing the sampling step size will result in an increase in the number of sampled frames within the same duration, which not only increases computational effort but also may introduce redundant frames, thus affecting the model's detection efficiency. The present invention uses the SlowFast network as the basis to build an uncivilized behavior detection network. In this network, the lightweight one-stage target detection network YoloX is used to realize target detection in the spatial dimension. The spatiotemporal features extracted by the SlowFast model are combined for action detection. The network inference speed is relatively fast. A non-uniform sampling method is first used to collect samples to obtain a data distribution that is more suitable for tasks involving similar behaviors. Then, a cascaded pooling three-dimensional spatial pyramid feature enhancement module that integrates shallow features is used to enhance the feature expression ability of the SlowFast model and improve the recognition accuracy of uncivilized behavior by this method. Summary of the Invention

[0006] In order to solve the problems of false detection of similar behaviors and low accuracy of local limb behavior detection in the spatiotemporal motion detection task of abnormal human behavior, the present invention discloses a method for detecting uncivilized human behavior that integrates non-uniform sampling and feature enhancement, based on a self-made spatiotemporal motion detection dataset of uncivilized human behavior. In the video frame sample acquisition stage, a non-uniform sampling method is used to obtain a data distribution that is more suitable for the task, so that the model can obtain motion change information in a hierarchical manner, effectively improving the network's attention to the detailed change features of similar behaviors. After the feature extraction network SlowFast, a cascaded pooling three-dimensional spatial pyramid structure that integrates shallow features is used to enhance the availability of feature information at different scales, reduce the loss of behavioral detail features, increase attention to motion features, and improve the accuracy of spatiotemporal motion detection.

[0007] The method for detecting uncivilized human behavior by integrating non-uniform sampling and feature enhancement includes the following steps:

[0008] S1. Collect surveillance videos and online video data of uncivilized behavior, convert the video clips into images at 30 frames per second, and annotate the video frames one frame per second. Finally, divide the data into training and test sets in a 4:1 ratio to complete the construction of the uncivilized behavior spatiotemporal action detection dataset.

[0009] S2. Using a non-uniform sampling method on the data set obtained in step S1, obtain an input data distribution that is more suitable for tasks involving similar actions, and obtain T frames of video information and β×T frames of video information as input samples, where β=1 / 4;

[0010] S3. Use the feature extraction network built on SlowFast to extract features from the input data samples obtained in step S2. The Slow channel uses β×T frames of video information as input samples to learn spatial semantic information, and the Fast channel uses T frames of video information as input samples to learn motion information.

[0011] S4: The feature data obtained in S3 is enhanced using a cascaded pooling three-dimensional spatial pyramid that integrates shallow features. The YoloX network is used to detect the location of people in the sample, and the ROI algorithm is used to extract the features of the corresponding key areas. The features are sent to the classification network for classification and to determine whether there is any uncivilized behavior.

[0012] Furthermore, step S2 includes the following steps:

[0013] S21. Use a non-uniform sampling method to perform sample sampling on the dataset. While ensuring that the model computational complexity remains unchanged, a data distribution more suitable for the task is obtained, which improves the network's attention to the changing characteristics of similar action details. This method can be expressed as:

[0014]

[0015] Where Z = {...-3, -2, -1, 0, 1, 2, 3...} represents the number of video frames to be selected; t0 is the sampling time corresponding to the video frame at the current prediction moment; r is the interval coefficient for non-uniform sampling, which defaults to 1; and sgn() represents the sign function. The above description applies to sampling odd-numbered frames. For even-numbered frames, after t0, an additional frame is sampled at an interval of half the current sampling interval. The entire set is arranged in chronological order to represent the sample.

[0016] Furthermore, step S3 includes the following steps:

[0017] S31. The feature extraction backbone network used for the input data samples is constructed by a 3D convolutional neural network. The specific network structure is as follows:

[0018] The convolution modules with kernel sizes of 1*1*1, 1*3*3, and 3*1*1, a batch normalization layer, and a ReLU activation function layer are set in sequence. The process is expressed as follows:

[0019]

[0020] Where x is the input feature, FRL is the ReLU activation function layer, F BN is the batch normalization layer, The convolution kernel sizes are 1*1*1, 1*3*3, and 3*1*1, respectively. The combination is represented as x i ′ is the output feature.

[0021] S32. After obtaining the feature extraction combination described in S31, multiple basic convolution modules are used to combine them to obtain two Res modules, which are expressed as:

[0022]

[0023] Among them F Pool is the pooling layer, N represents the number of residual stacking layers. Except for the last stage, the output features are downsampled in the spatial dimension after each stage to realize the multi-level and multi-scale structure of the network and effectively reduce the amount of network calculation.

[0024] S33. After completing the feature extraction described in S32, multiple Res modules are combined to form two channels, Slow and Fast, which are then fused to obtain the SlowFast module, which can be expressed as:

[0025]

[0026] Among them F Slow (x) and F Fast (x) Both pathways include 4 Res stages. After each stage, the features of the fast pathway are fused with the features of the slow pathway. The formula is as follows:

[0027]

[0028] in and are the output features of the fast channel and slow channel at stage i, respectively. is the output feature after the fusion of the two channels in the i-th stage.

[0029] Furthermore, step S4 includes the following steps:

[0030] S41. For the features obtained in S3, a cascaded pooling three-dimensional spatial pyramid that integrates shallow features is used to perform feature enhancement operations to reduce the loss of detail information and obtain more expressive multi-scale semantic information. The process is expressed as follows:

[0031] x i-1 =Concat c (x,AdaPool w,h (x shallow ))

[0032]

[0033] x enhenced =Concat c (x1,x2,x3,x4)

[0034] where x shallow It is the shallow feature obtained by fusing the fast and slow channels in the first Res stage of the SlowFast module. Indicates maximum pooling in w,h dimensions, with a pooling window size of 2, AdaPool w,h Indicates that adaptive maximum pooling is performed on the w, h dimensions with reference to the feature extraction network output layer feature dimensions. After concatenating the shallow features and the final Res stage features, three cascade pooling operations are performed. As the pooling depth increases, the retained features will have a larger regional receptive field. The shallower feature maps capture the basic features and provide the initial feature representation. Through cross-level pooling connections, the ability of the structure to simultaneously capture information at different scales is further enhanced, effectively reducing the loss of semantic information in the feature extraction process. Subsequently, through Concat c Perform feature concatenation on the channel dimension to obtain the final enhanced feature x enhenced .

[0035] S42. Extract the key frame of the video sample at the current prediction moment and use the one-stage lightweight target detection network YoloX to detect people in the video frame. Combine the enhanced features obtained in S41 with the person location information obtained by detection, and use the ROI algorithm to obtain the features of the area where the corresponding person is located, which is expressed as:

[0036] x′=ROI(x)

[0037] The ROI is the selective extraction of feature areas where target persons exist; these features are sent to the classification network for classification and to determine whether there is uncivilized behavior.

[0038] Beneficial effects:

[0039] 1. This paper proposes a method for detecting uncivilized human behavior that integrates non-uniform sampling and feature enhancement. It uses a 3D fast-slow dual-path convolutional neural network for feature extraction. It replaces the optical flow channel in the traditional two-stream network with the RGB channel to obtain temporal information. This effectively solves the problem of difficulty in extracting optical flow information and improves network inference speed.

[0040] 2. When sampling video frames, the present invention first uses a non-uniform sampling method to obtain input samples to obtain a data distribution that is more suitable for spatiotemporal action detection tasks involving similar behaviors. It then uses a cascaded pooling three-dimensional spatial pyramid that integrates shallow features to perform feature enhancement, reduce information loss, enhance feature expressiveness, and strengthen attention to action features, effectively improving the accuracy of spatiotemporal action detection.

[0041] 3. The present invention implements dual-model independent embedding distribution for the YoloX target detection network and the human uncivilized behavior detection network that integrates non-uniform sampling and feature enhancement, and realizes continuous processing of video frames under the premise of real-time judgment of the spatiotemporal actions of uncivilized behavior. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a diagram showing the overall structure of the method for detecting uncivilized human behavior that integrates non-uniform sampling and feature enhancement, used in an embodiment of the present invention;

[0043] Figure 2 This is a diagram illustrating an implementation of the non-uniform sampling method for detecting uncivilized human behavior that integrates non-uniform sampling and feature enhancement, used in an embodiment of the present invention;

[0044] Figure 3 This is a structural diagram of the cascaded pooling three-dimensional spatial pyramid module that integrates shallow features in the method for detecting uncivilized human behavior that integrates non-uniform sampling and feature enhancement used in an embodiment of the present invention;

[0045] Figure 4 The training method for detecting uncivilized human behavior by integrating non-uniform sampling and feature enhancement used in an embodiment of the present invention is implemented;

[0046] Figure 5 This is a diagram showing the effect of implementing the method for detecting uncivilized human behavior by integrating non-uniform sampling and feature enhancement used in an embodiment of the present invention, showing the recognition effects of yelling, smoking, fighting, stealing, and normal behaviors. DETAILED DESCRIPTION

[0047] To provide a clearer understanding of the technical features, objectives, and beneficial effects of the present invention, an embodiment of the present invention is further described with reference to the accompanying drawings. The embodiment is intended only to further illustrate the present invention and is not to be construed as limiting the scope of protection of the present invention. Non-essential improvements and adjustments made by those skilled in the art based on the contents of the present invention also fall within the scope of protection of the present invention.

[0048] The method for detecting uncivilized human behavior by integrating non-uniform sampling and feature enhancement includes the following steps:

[0049] S1. Collect surveillance videos and online video data of uncivilized behavior, convert the video clips into images at 30 frames per second, and annotate the video frames one frame per second. Finally, divide the data into training and test sets in a 4:1 ratio to complete the construction of the uncivilized behavior spatiotemporal action detection dataset.

[0050] S2. Using a non-uniform sampling method on the data set obtained in step S1, obtain an input data distribution that is more suitable for tasks involving similar actions, and obtain T frames of video information and β×T frames of video information as input samples, where β=1 / 4;

[0051] S3. Use the feature extraction network built on SlowFast to extract features from the input data samples obtained in step S2. The Slow channel uses β×T frames of video information as input samples to learn spatial semantic information, and the Fast channel uses T frames of video information as input samples to learn motion information.

[0052] S4: The feature data obtained in S3 is enhanced using a cascaded pooling three-dimensional spatial pyramid that integrates shallow features. The YoloX network is used to detect the location of people in the sample, and the ROI algorithm is used to extract the features of the corresponding key areas. The features are sent to the classification network for classification and to determine whether there is any uncivilized behavior.

[0053] The step S2 comprises the following steps:

[0054] S21. Use a non-uniform sampling method to perform sample sampling on the dataset. While ensuring that the model computational complexity remains unchanged, a data distribution more suitable for the task is obtained, which improves the network's attention to the changing characteristics of similar action details. This method can be expressed as:

[0055]

[0056] Where Z = {...-3, -2, -1, 0, 1, 2, 3...} represents the number of video frames to be selected; t0 is the sampling time corresponding to the video frame at the current prediction moment; r is the interval coefficient for non-uniform sampling, which defaults to 1; and sgn() represents the sign function. The above description applies to sampling odd-numbered frames. For even-numbered frames, after t0, an additional frame is sampled at an interval of half the current sampling interval. The entire set is arranged in chronological order to represent the sample.

[0057] The step S3 comprises the following steps:

[0058] S31. The feature extraction backbone network used for the input data samples is constructed by a 3D convolutional neural network. The specific network structure is as follows:

[0059] The convolution modules with kernel sizes of 1*1*1, 1*3*3, and 3*1*1, a batch normalization layer, and a ReLU activation function layer are set in sequence. The process is expressed as follows:

[0060]

[0061] Where x is the input feature, F RL is the ReLU activation function layer, F BN is the batch normalization layer, The convolution kernel sizes are 1*1*1, 1*3*3, and 3*1*1, respectively. The combination is represented as x i ′ is the output feature.

[0062] S32. After obtaining the feature extraction combination described in S31, multiple basic convolution modules are used to combine them to obtain two Res modules, which are expressed as:

[0063]

[0064] Among them F Pool is the pooling layer, N represents the number of residual stacking layers. Except for the last stage, the output features are downsampled in the spatial dimension after each stage to realize the multi-level and multi-scale structure of the network and effectively reduce the amount of network calculation.

[0065] S33. After completing the feature extraction described in S32, multiple Res modules are combined to form two channels, Slow and Fast, which are then fused to obtain the SlowFast module, which can be expressed as:

[0066]

[0067] Among them F Slow (x) and F Fast (x) Both pathways include 4 Res stages. After each stage, the features of the fast pathway are fused with the features of the slow pathway. The formula is as follows:

[0068]

[0069] in and are the output features of the fast channel and slow channel at stage i, respectively. is the output feature after the fusion of the two channels in the i-th stage.

[0070] The step S4 comprises the following steps:

[0071] S41. For the features obtained in S3, a cascaded pooling three-dimensional spatial pyramid that integrates shallow features is used to perform feature enhancement operations to reduce the loss of detail information and obtain more expressive multi-scale semantic information. The process is expressed as follows:

[0072] x i-1 =Concat c (x,AdaPool w,h (x shallow ))

[0073]

[0074] x enhenced =Concat c (x1,x2,x3,x4)

[0075] where x shallow It is the shallow feature obtained by fusing the fast and slow channels in the first Res stage of the SlowFast module. Indicates maximum pooling in w,h dimensions, with a pooling window size of 2, AdaPool w,h Indicates that adaptive maximum pooling is performed on the w, h dimensions with reference to the feature extraction network output layer feature dimensions. After concatenating the shallow features and the final Res stage features, three cascade pooling operations are performed. As the pooling depth increases, the retained features will have a larger regional receptive field. The shallower feature maps capture the basic features and provide the initial feature representation. Through cross-level pooling connections, the ability of the structure to simultaneously capture information at different scales is further enhanced, effectively reducing the loss of semantic information in the feature extraction process. Subsequently, through Concat c Perform feature concatenation on the channel dimension to obtain the final enhanced feature x enhenced .

[0076] S42. Extract the key frame of the video sample at the current prediction moment and use the one-stage lightweight target detection network YoloX to detect people in the video frame. Combine the enhanced features obtained in S41 with the person location information obtained by detection, and use the ROI algorithm to obtain the features of the area where the corresponding person is located, which is expressed as:

[0077] x′=ROI(x)

[0078] The ROI is the selective extraction of feature areas where target persons exist; these features are sent to the classification network for classification and to determine whether there is uncivilized behavior.

[0079] Simulation experiment

[0080] Depend on Figure 5This method can effectively detect uncivilized behavior in spatiotemporal space based on video, and is applicable to a variety of environments, including indoors and outdoors, during the day, and at night. To quantitatively evaluate the detection effectiveness of the present invention, the mAP evaluation metric scores are shown in Table 1. The SlowFast method uses the baseline SlowFast network for action recognition, the ACRN method uses the ACRN network for action recognition, and the Ours method is the method described in the present invention. To ensure fairness, all of these methods use the same object detection network to detect person bounding boxes in video frames.

[0081] As can be seen from Table 1, the mAP index is improved to a certain extent by using the human uncivilized behavior detection method that integrates non-uniform sampling and feature enhancement as described in the present invention, and the spatiotemporal action detection effect is better than other comparison models.

[0082] Table 1 Statistics of simulation experiment evaluation indicators

[0083] method mAP SlowFast 54.98 ACRN 57.84 Ours 58.68

[0084] The above simulation experimental results show that compared with the comparison method, the present invention can effectively improve the detection accuracy and enhance the overall performance of the model.

[0085] The above describes the method of the present invention. Those skilled in the art can implement the method of the present invention based on the description of this content. Based on the above content of the present invention, other embodiments obtained by those skilled in the art without making any creative work should fall within the scope of protection of the present invention.

Claims

1. A method for detecting uncivilized human behavior that integrates non-uniform sampling and feature enhancement is characterized by: The following steps are involved: S1. Collect surveillance videos and online video data of uncivilized behavior, convert the video clips into images at 30 frames per second, and annotate the video frames one frame per second. Finally, divide the data into training and test sets in a 4:1 ratio to complete the construction of the uncivilized behavior spatiotemporal action detection dataset. S2. Using a non-uniform sampling method on the data set obtained in step S1, obtain an input data distribution that is more suitable for tasks involving similar actions, and obtain T frames of video information and β×T frames of video information as input samples, where β=1 / 4; S3. Use the feature extraction network built on SlowFast to extract features from the input data samples obtained in step S2. The Slow channel uses β×T frames of video information as input samples to learn spatial semantic information, and the Fast channel uses T frames of video information as input samples to learn motion information. S4: The feature data obtained in S3 is enhanced using a cascaded pooling three-dimensional spatial pyramid that integrates shallow features. The YoloX network is used to detect the location of people in the sample, and the ROI algorithm is used to extract features of the corresponding key areas. The features are then fed into a classification network for classification and to determine whether there is uncivilized behavior. The step S2 comprises the following steps: S21. Use a non-uniform sampling method to perform sample sampling on the dataset. While ensuring that the model computational complexity remains unchanged, a data distribution more suitable for the task is obtained, which improves the network's attention to the changing characteristics of similar action details. This method can be expressed as: Where Z = {...-3, -2, -1, 0, 1, 2, 3...} represents the number of video frames you want to select; t0 is the corresponding sampling time of the video frame at the current prediction moment; r is the interval coefficient of non-uniform sampling, which defaults to 1; sgn() represents the sign function; the above description is applicable to the case of sampling odd frames. If it is an even frame, after t0, one more frame is sampled with an interval of half the current sampling interval, and the entire set is arranged in chronological order to represent the sample.

2. The method for detecting uncivilized human behavior by integrating non-uniform sampling and feature enhancement according to claim 1 is characterized in that: The step S3 comprises the following steps: S31. The backbone network of the feature extraction network used for the input data sample is constructed by a 3D convolutional neural network. The specific network structure is as follows: The convolution modules with kernel sizes of 1*1*1, 1*3*3, and 3*1*1, a batch normalization layer, and a ReLU activation function layer are set in sequence. The process is expressed as follows: Where x is the input feature, F RL is the ReLU activation function layer, F BN is the batch normalization layer, The convolution kernel sizes are 1*1*1, 1*3*3, and 3*1*1, respectively. The combination is represented as x i ′ is the output feature; S32, after obtaining the combination of S31 Finally, multiple basic convolution modules are combined to obtain two Res modules, which are expressed as: Among them F Pool It is a pooling layer, where N represents the number of residual stacking layers. Except for the last stage, the output features are downsampled in the spatial dimension after each stage to achieve a multi-level and multi-scale structure of the network and effectively reduce the amount of network computation. S33. After completing S32, multiple Res modules are combined to form two channels, Slow and Fast, which are then fused to obtain the SlowFast module, which can be expressed as: Among them F Slow (x) and F Fast (x) Both pathways include 4 Res stages. After each stage, the features of the fast pathway are fused with the features of the slow pathway. The formula is as follows: in and are the output features of the fast channel and slow channel at stage i, respectively. is the output feature after the fusion of the two channels in the i-th stage.

3. The method for detecting uncivilized human behavior by integrating non-uniform sampling and feature enhancement according to claim 1 is characterized in that: The step S4 comprises the following steps: S41. For the features obtained in S3, a cascaded pooling three-dimensional spatial pyramid that integrates shallow features is used to perform feature enhancement operations to reduce the loss of detail information and obtain more expressive multi-scale semantic information. The process is expressed as follows: x i-1 =Concat c (x,AdaPool w,h (x shallow )) x enhenced =Concat c (x1,x2,x3,x4) where x shallow It is the shallow feature obtained by fusing the fast and slow channels in the first Res stage of the SlowFast module. Indicates maximum pooling in w,h dimensions, with a pooling window size of 2, AdaPool w,h Indicates that adaptive maximum pooling is performed on the feature dimension of the feature extraction network output layer in the w, h dimension; the shallow features and the final Res stage features are concatenated and then cascaded three times. As the pooling depth increases, the retained features will have a larger regional receptive field, where the shallower feature maps capture the basic features and provide the initial feature representation. Through cross-level pooling connections, the ability of the structure to simultaneously capture information at different scales is further enhanced, effectively reducing the loss of semantic information in the feature extraction process. Subsequently, through Concat c Perform feature concatenation on the channel dimension to obtain the final enhanced feature x enhenced ; S42. Extract the key frame of the video sample at the current prediction moment and use the one-stage lightweight target detection network YoloX to detect people in the video frame. Combine the enhanced features obtained in S41 with the person location information obtained by detection, and use the ROI algorithm to obtain the features of the area where the corresponding person is located, which is expressed as: x′=ROI(x) The ROI is the selective extraction of feature areas where target persons exist; these features are sent to the classification network for classification and to determine whether there is uncivilized behavior.

Citation Information

Patent Citations

  • Multi-stage human body abnormal action detection method based on residual network

    CN114202803A

  • Fire operation violation behavior identification method

    CN116721383A

  • Two-stage two-channel distracted driving behavior identification method based on key point detection

    CN116721405A

  • Behavior detection method and device based on spatio-temporal context, equipment and medium

    CN115359570A