Personnel abnormal behavior detection method

By building the improved S3DD-Yolov11 model, the problems of manpower consumption and false alarms and misreports of traditional monitoring methods are solved, efficient abnormal behavior detection is achieved, and the safety of public places is improved.

CN120299089APending Publication Date: 2025-07-11JIANGSU SECOND NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510480126.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-04-15
Filing Date
2025-04-17
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The traditional manual monitoring method consumes a lot of manpower and material resources, and it is easy to cause false alarms and miss reports when security personnel are not in good health, making it difficult to effectively ensure public safety.

Method used

The S3DD-Yolov11 model is constructed, and personnel abnormal behavior detection is detected through the improved Yolov11 neural network model. Using 3D cavity convolution, SE attention mechanism and deformable convolution technologies, the time feature extraction and feature fusion of video data is achieved, and the detection accuracy and speed are improved.

Benefits of technology

On the COCO-Action dataset, the accuracy of time feature extraction is increased by 14.2%, and the detection speed of abnormal behavior is increased by 41.3%. The error detection rate is reduced in dense crowd scenarios, which enhances the robustness and detection efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299089A_ABST
    Figure CN120299089A_ABST
Patent Text Reader

Abstract

The invention discloses a method for detecting abnormal behaviors of personnel, which comprises the following steps of: taking video data of a monitoring scene, and constructing a data set; the method comprises the following steps: constructing an S3DD-Yolov11 model; establishing a personnel abnormal behavior monitoring system; and target monitoring. The method comprises the following steps: deploying a compression-excitation SE (Squeezeamp; the original 2D convolution of the Yolov11 is replaced by the 3D dilated convolution and the deformable convolution, so that the human body form is better simulated, and the network detection performance is improved; algorithm optimization is carried out for an embedded platform, a large-range pooling is divided into a plurality of small window pooling, the calculation complexity is reduced under the condition of function equivalence, and the monitoring speed is improved; the S3DD-Yolov11-based personnel abnormal behavior monitoring system is constructed, personnel abnormal behavior detection is realized, functions of alarm, real-time information display, historical query and the like are integrated, and the system can be used for community resident security and social security management scenes. And the method plays a crucial role in preventing and timely processing related events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting abnormal behaviors of personnel and belongs to the field of computer vision. Background Art

[0002] In today's society, with the development of the global economy and the increase in population density, the process of urbanization shows an accelerating trend, and the public security issues of society have become increasingly complex. Security issues have always been the focus of concern for managers of public places, such as important places like enterprises, shopping malls, hospitals, and schools. The "Safe City" plan being implemented in some domestic cities includes hundreds of thousands of monitoring and defense facilities. However, due to the huge flow of people, if the traditional method of manual duty is adopted, arranging special personnel to monitor 24 hours a day, not only will it consume a large amount of manpower, material resources, and financial resources, but also it is inevitable to miss potential hazards that have not yet occurred.

[0003] Traditional security measures have high maintenance costs, and in some special cases such as when security personnel are physically unwell, the monitoring intensity and accuracy will inevitably show large fluctuations in the error rate, and even false alarms and missed alarms of abnormal behaviors may occur. Security without timeliness is difficult to effectively protect people's personal safety.

[0004] Therefore, by adopting an intelligent method to timely detect and control personnel holding dangerous objects or performing dangerous actions, the occurrence of safety accidents can be greatly reduced, and people's lives and property can be protected. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a method for detecting abnormal behaviors of personnel. By constructing an S3DD - Yolov11 model, the abnormal behavior features can be accurately located, which can effectively avoid some illegal and criminal behaviors, timely handle relevant behaviors, and ensure public safety.

[0006] The present invention adopts the following technical solutions to solve the above technical problems: A method for detecting abnormal behaviors of personnel, comprising: S1: Construct a video data set including a training set, a validation set, and a test set; S2: Based on Yolov11, construct a neural network model; S3: Use the training set, the validation set, and the test set to train the machine learning model to obtain a model for detecting abnormal behaviors of personnel; S4: Use the model for detecting abnormal behaviors of personnel to realize real - time detection of abnormal behaviors of personnel in the target area.

[0007] Preferably, in the step S2, the neural network model is an improvement of Yolov11, and the improvements include: Replace the first 2D convolutional module in the backbone network with a 3D dilated convolutional module; Introduce an SE attention mechanism module before the 3D dilated convolutional module, and replace the global pooling in the SE attention mechanism module with 3D pooling to form an SE_3D module.

[0008] Preferably, the SE_3D module includes a 3D pooling layer, a first FC fully connected layer, a ReLU activation function, a second FC fully connected layer, a Sigmoid activation function, and a Scale layer.

[0009] Preferably, in step S2, the neural network model is an improvement of Yolov11, and the improvements include: Add an SE attention mechanism to the C3k2 module in the backbone network to form a C3k2_SE module.

[0010] Preferably, in step S2, the neural network model is an improvement of Yolov11, and the improvements include: Replace the last 2D convolutional module in the backbone network with a deformable convolutional DCN module.

[0011] Preferably, the improved backbone network includes an SE_3D module, a 3D dilated convolutional module, a 2D convolutional module, a C3k2_SE module, a 2D convolutional module, a C3k2_SE module, a 2D convolutional module, a C3k2_SE module, a DCN module, a C3k2_SE module, and an SPPF module.

[0012] Preferably, the video data includes the actions, postures, voices, and location information of people.

[0013] Preferably, it also includes preprocessing of rectangular loading and enhancement of the video data.

[0014] The present invention also provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned method for detecting abnormal behaviors of people are implemented.

[0015] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned method for detecting abnormal behaviors of people are implemented.

[0016] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects: (1) Optimization of temporal dimension feature extraction: By replacing the first 2D convolution layer of the backbone network with a 3D dilated convolution module, the accuracy of temporal feature extraction for consecutive video frames is improved by 14.2% on the COCO-Action dataset, and the detection speed of instantaneous abnormal behaviors (such as punching actions during robbery) is increased by 41.3%; (2) Dynamic enhancement of channel features: By integrating the SE attention mechanism in the C3K2 module, the model's attention to key channel features is increased by 27.8%, and the false detection rate of abnormal behaviors (such as theft) is reduced in dense crowd scenarios; (3) Multi-scale feature fusion and calculation acceleration: Adopting a multi-step pooling combination strategy, while maintaining an equivalent receptive field, the computational amount of a single pooling is reduced, and combined with a matrix multiplication accelerator, the overall inference speed is improved; (4) Enhancement of robustness in complex scenarios: Through multi-granularity attention feature fusion, noise interference (such as rain, fog, and shadows) is suppressed, and the consistency modeling of human body contours and dynamic behaviors is enhanced. Description of the Drawings

[0017] Figure 1 is a flowchart of an embodiment of the present application.

[0018] Figure 2 is a schematic diagram of the specific structure of the improved Yolov11.

[0019] Figure 3 is a flowchart of feature fusion in the present invention.

[0020] Figure 4 is a flowchart of the NMS algorithm in the present invention. Detailed Embodiments

[0021] The following details the embodiments of the present invention, and examples of the embodiments are shown in the drawings. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.

[0022] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the technical field to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as such here.

[0023] The following, in combination with the drawings and specific embodiments, further details the technical solutions of the present invention: In one embodiment, as Figure 1As shown in the figure, a method for detecting abnormal human behavior based on S3DD-Yolov11 includes the following steps: Step 1: Obtain surveillance video data from public areas such as actual streets, covering different abnormal public scenarios, such as street robbery, fighting, theft, arson, etc., and different population densities. Divide the video data, with 80% as the training set, 15% as the test set, and 5% as the validation set.

[0024] Preprocess the video data, and the processing process includes the following steps: (1) Load the video data in a rectangle to eliminate redundant data and improve the model training speed; (2) Perform operations such as Random_Perspective enhancement, HSV transformation enhancement, random cutout enhancement, random up / down or left / right flipping enhancement on the video data to retain temporal information.

[0025] Step 2: Improve Yolov11 to form S3DD-Yolov11 as shown in Figure 2 the figure.

[0026] (1) Improve the backbone structure of Yolov11. Use 3D dilated convolution to replace the 2D convolution in the first layer to achieve feature extraction in the time dimension of the video data; use deformable convolution DCN to replace the 2D convolution in the last layer; (2) Improve the C3K2 convolutional neural network module of Yolov11, fuse the SE attention mechanism, process the feature extraction in different stages of the backbone, and extract multi-scale information of the image.

[0027] Among them, as shown in Figure 3 the figure, the backbone network feature extraction of S3DD-Yolov11 mainly includes the following steps: (1) During the convolution process, the convolution kernel slides on the input image at a certain stride. The input data is arranged in matrix form through the im2col operation, and the pixels in each sliding window are weighted and summed. This process can extract the local features of the image. (2) Perform multiple convolution operations through multiple convolution kernels with different parameters to obtain multiple feature maps, which contain various feature information of the image, such as edges, textures, colors, etc. (3) Use the activation function ReLu(x)=max(0,x) after each convolution operation; where ReLu(x) represents taking the larger value between the function input parameter x and 0. The goal of this function is to perform a non-linear transformation on the feature map, thereby introducing non-linear factors and improving the expression ability of the model.

[0028] (4) During the pooling operation, the original large-size sliding window pooling operation is changed to multiple sliding window pooling operations with smaller strides, expanding the receptive field, capturing more extensive input information, and at the same time reducing the pooling window size to decrease the computational amount of each pooling operation and improve the computational speed of the convolutional network. (5) Use a matrix multiplication accelerator to perform matrix operations in parallel to improve the computational speed of the convolutional network. (6) Fuse the attention feature fusion mechanism (MG-RAFA) that aggregates feature information at different scales with multi-granularity so that the model can capture multi-scale information in the image simultaneously.

[0029] Transmit feature information such as human body contours, dynamic behaviors, and the surrounding environment to the neck network for feature fusion and enhancement; transmit the processed feature information to the head network, and use the non-maximum suppression (NMS) algorithm to remove overlapping detection boxes to obtain the final detection result. Among them, as Figure 4 shown, the steps of the NMS algorithm are as follows: S1: Sort in descending order according to the confidence scores of the prediction boxes; S2: Select the prediction box B i with the highest score as the candidate box, and calculate its intersection over union (IOU) with the remaining prediction boxes; S3: If the IOU exceeds the set threshold, suppress (delete or reduce the score) these prediction boxes; S4: Iterate S2 and S3 until all prediction boxes are processed.

[0030] Step 3: Use the preprocessed dataset to pre-train the model to improve the generalization ability of the model. Set the epoch to 50 and set the initial learning rate to 0.01.

[0031] To comprehensively evaluate the performance of the algorithm of the present invention, select evaluation metrics: Recall rate R (Recall), which reflects the ability of the model to find all positive class samples, ; Precision rate P (Precision), which reflects the accuracy of the model's prediction as a positive class, ; Average precision AP (average precision), which refers to the average value of the precisions at all different recall rate levels, ; mAP (mean average precision) is the average precision of all classes, ; Among them, T P (True Positive) is the number of true positives in the test samples; T N(True Negative) is the number of true negatives in the test samples; F P (False Positive) is the number of false positives in the test samples; F N (False Negative) is the number of false negatives in the test samples; r1, r2, …, r n is the P interpolation segment arranged in ascending order; AP i is the average precision of class i.

[0032] In the same environment, using the same data set, the S3DD-Yolov11 algorithm and the existing algorithms are measured. The performance of the S3DD-Yolov11 algorithm has significant superiority.

[0033] The classification decision characteristics of the model are systematically characterized by the above quadruple, and on this framework, three types of core indicators are hierarchically set: The first level is the single-dimensional efficacy index, including the recall rate R and the precision rate P; the average precision AP is introduced in the second level, and the precision change curve is integrated through the measured average value to construct an integral evaluation index in the P-R space, effectively overcoming the limitations of single-point evaluation; in the third level, mAP is used as a comprehensive evaluation index, and by taking the arithmetic mean of the AP values of all classes, a standardized measure of the cross-class generalization performance is achieved.

[0034] In actual application, the S3DD-Yolov11 model is loaded into the embedded device running the monitoring system, and functions such as alarm, real-time display of alarm information, and historical query are integrated on the central control platform. When abnormal behaviors such as street robbery, fighting, theft, arson, etc. are detected, the analysis results and location alarm information are sent to the central platform. The central platform will display the alarm information on the display platform in real time and store the alarm data in the database, providing the function of querying and displaying historical alarm information, so as to realize the timely monitoring of abnormal behaviors of personnel in public areas, which plays a crucial role in the prevention and handling of related social vicious events.

[0035] Based on the same technical solution, the present invention also designs an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned method for detecting abnormal behaviors of personnel are realized.

[0036] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for detecting abnormal behaviors of personnel are realized.

[0037] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0038] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person familiar with the technology can think of transformations or substitutions within the technical scope disclosed by the present invention, and they should all be covered within the scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for detecting abnormal behavior of personnel, characterized in that, Including: S1: Construct a video dataset including a training set, a validation set, and a test set; S2: Based on Yolov11, construct a neural network model; S3: Use the training set, validation set, and test set to train the machine learning model to obtain a personnel abnormal behavior detection model; S4: Use the personnel abnormal behavior detection model to achieve real-time detection of personnel abnormal behavior in the target area.

2. The method according to claim 1, wherein In step S2, the neural network model is an improvement of Yolov11. The improvements include: Replace the first 2D convolution module in the backbone network with a 3D dilated convolution module; Introduce an SE attention mechanism module before the 3D dilated convolution module, and replace the global pooling in the SE attention mechanism module with 3D pooling to form an SE_3D module.

3. The method according to claim 2, wherein The SE_3D module includes a 3D pooling layer, a first FC fully connected layer, a ReLU activation function, a second FC fully connected layer, a Sigmoid activation function, and a Scale layer.

4. The method according to claim 3, characterized in that In step S2, the neural network model is an improvement of Yolov11. The improvements include: Add an SE attention mechanism to the C3k2 module in the backbone network to form a C3k2_SE module.

5. The method according to claim 4, wherein In step S2, the neural network model is an improvement of Yolov11. The improvements include: Replace the last 2D convolution module in the backbone network with a deformable convolution DCN module.

6. The method according to claim 5, characterized in that, The improved backbone network includes an SE_3D module, a 3D dilated convolution module, a 2D convolution module, a C3k2_SE module, a 2D convolution module, a C3k2_SE module, a 2D convolution module, a C3k2_SE module, a DCN module, a C3k2_SE module, and a SPPF module.

7. The method according to claim 1, wherein The video data includes the actions, postures, voices, and location information of personnel.

8. The method according to claim 1, wherein It also includes preprocessing such as rectangular loading and enhancement of the video data.

9. An electronic device, comprising a memory and a processor, characterized in that: The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the method as described in any one of claims 1-8.