Edge calculation-based worker abnormal behavior identification method
By using the YOLOv11n-CBAM network and Retinaface algorithm in the recognition of workers' abnormal behavior, combined with edge computing technology, the problems of degraded model accuracy, insufficient resource limitation and generalization capabilities in the existing technology are solved, efficient and accurate abnormal behavior detection and face recognition are achieved, and the intelligence level of industrial monitoring is improved.
Patent Information
- Application Number
- CN202510124800.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-06-20
AI Technical Summary
The existing deep learning combined with edge computing technology has problems such as decreasing model compression accuracy, limitation of edge equipment resources, insufficient algorithm generalization capabilities, and poor real-time performance and energy consumption management in the recognition of workers' abnormal behavior.
The YOLOv11n-CBAM network is used in combination with the Retinaface algorithm to build basic data sets, image enhancement, model training and face library establishment to realize abnormal behavior detection and face recognition, and deploy it on devices close to the data source through edge computing architecture, reducing data transmission delay and improving real-time processing.
It significantly improves detection efficiency and accuracy, ensures stable performance under complex lighting and occlusion conditions, reduces the calculation and communication costs of traditional centralized methods, and improves monitoring efficiency and intelligence level in industrial scenarios.
Smart Images

Figure CN120183031A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of image recognition technology, and particularly to a method for recognizing abnormal behaviors of workers based on edge computing. Background Art
[0002] With the continuous advancement of the industrialization process, work safety has become one of the key concerns in various industries. However, due to the complex and ever-changing production environment, workers may exhibit abnormal behaviors during work (such as not wearing safety helmets, improper operations, etc.), which pose potential threats to production safety. Traditional monitoring methods usually rely on manual observation or centralized data analysis, making it difficult to identify abnormal behaviors in real time and effectively, especially in scenarios with multiple cameras or large network delays. Edge computing technology, due to its distributed architecture and real-time processing capabilities, has shown significant advantages in the field of work safety. It can complete calculations and analyses at the device end close to the data source, thereby reducing data transmission delays and improving recognition efficiency.
[0003] At the same time, object detection algorithms based on deep learning provide technical support for abnormal behavior recognition. Currently, mainstream object detection algorithms mainly include YOLO (You Only Look Once), Faster R-CNN (Faster Region-based Convolutional Neural Networks), and SSD (Single Shot MultiBox Detector), etc. Among them, YOLO, with its end-to-end structure and fast and accurate detection capabilities, has been widely used in real-time video analysis tasks. Faster R-CNN, with its Region Proposal Network (RPN) as the core, has high detection accuracy and is suitable for scenarios with relatively high requirements for detection accuracy but relatively low real-time requirements. SSD, by adopting multi-scale feature detection and directly regressing the target bounding boxes, achieves a good balance between detection speed and accuracy.
[0004] Combined with edge computing technology, these algorithms can all be optimized and deployed in edge devices. Among them, YOLO is particularly suitable for real-time monitoring tasks due to its lightweight and speed advantages. Edge devices can use neural network models to quickly identify the target features and behavior patterns of workers, realizing real-time monitoring and early warning of abnormal behaviors. This combination of technology based on edge computing and object detection algorithms not only significantly improves the monitoring efficiency and accuracy but also effectively reduces the computational and communication costs of traditional centralized methods, providing an intelligent and scalable solution for work safety.
[0005] The main defects of the existing deep learning combined with edge computing technology include: there is a significant risk of accuracy decline during the model compression process, and the excessive lightweighting of the model often sacrifices key detection performance; the computing resources and storage of edge devices are limited, and complex deep learning models are difficult to deploy efficiently; the algorithm generalization ability for specific scenarios is weak, and it is difficult to adapt to the complex and changeable industrial environment; the real-time performance and energy consumption management still need to be further optimized, especially for the monitoring scenarios with long-term continuous operation. In response to these challenges, future research needs to conduct in-depth exploration and innovation in aspects such as model lightweighting, edge inference efficiency, and algorithm robustness.
[0006] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.
[0007] It should be noted that this part aims to provide background or context for the technical solutions of the present disclosure stated in the claims. The description herein is not admitted to be prior art merely because it is included in this part. Summary of the Invention
[0008] The purpose of the embodiments of the present disclosure is to provide a method for identifying abnormal behaviors of workers based on edge computing, thereby at least to some extent overcoming one or more problems caused by the limitations and defects of related technologies.
[0009] According to the embodiments of the present disclosure, a method for identifying abnormal behaviors of workers based on edge computing is provided. The method includes: Construct a basic data set, and divide the basic data set into a training set, a validation set, and a test set according to a preset ratio; Perform image enhancement on the training set, the validation set, and the test set; Construct a YOLOv11n-CBAM network; wherein, the YOLOv11n-CBAM network includes a backbone network and a detection head; Use the image-enhanced training set, validation set, and test set to train the YOLOv11n-CBAM network to obtain the trained YOLOv11n-CBAM network; Establish a worker face database, and combine the Retinaface algorithm to construct a Retinaface face recognition model; Obtain worker operation images, and input the worker operation images into the trained YOLOv11n-CBAM network and the Retinaface face recognition model respectively to obtain abnormal behavior detection results and face recognition information; Generate an alarm result according to the abnormal behavior detection result and the face recognition information.
[0010] Further, in the step of constructing the basic data set and dividing the basic data set into a training set, a validation set, and a test set according to a preset ratio, the following steps are included: Collect operation images during the workers' work and annotate the operation images to generate a basic data set; wherein, the operation images at least include images when wearing a safety helmet, images when not wearing a safety helmet, images of holding goods by hand, images of hooks, images of suspended objects, and images of people. Randomly divide the basic data set into the training set, the validation set, and the test set according to a preset ratio through the random.shuffle() function; wherein, the preset ratio is 7:2:1.
[0011] Further, in the step of performing image enhancement on the training set, the validation set, and the test set, the following steps are included: Add fog effect to simulate a low visibility environment. Perform blurring to restore motion blur or focal length offset scenarios. Convert to grayscale to enhance the model's adaptability to color changes. Perform gamma correction to optimize the bright-dark contrast effect. Perform mosaic enhancement to generate complex scene samples by randomly splicing multiple images.
[0012] Further, the backbone network includes: A Conv convolutional layer, a first C3k2 layer, a second Conv convolutional layer, a second C3k2 layer, a third Conv convolutional layer, a third C3k2 layer, a fourth Conv convolutional layer, a fourth C3k2 layer, an SPPF layer, and a C2PSA layer connected in sequence; wherein, The outputs of the second C3k2 layer, the third C3k2 layer, and the C2PSA layer are all used as the feature outputs of the backbone network. The backbone network extracts multi-level feature information at five scales of P1 / 2, P2 / 4, P3 / 8, P4 / 16, and P5 / 32 through a hierarchical feature extraction module. The C2PSA layer introduces a Polarized Self-Attention mechanism for optimizing the feature channel and spatial weight distribution.
[0013] Further, the detection head includes: It is composed of a first Upsample layer, a first Concat layer, a fifth C3k2 layer, a second Upsample layer, a second Concat layer, a sixth C3k2 layer, a first downsampling Conv layer, a third Concat layer, a seventh C3k2 layer, a second downsampling Conv layer, a fourth Concat layer, and a Detect layer connected in sequence; wherein, After each of the C3k2 modules, a CBAM module is embedded. The CBAM module includes a sixth C3k2 layer, a seventh C3k2 layer, and an eighth C3k2 layer. The outputs of the sixth C3k2 layer, the seventh C3k2 layer, and the eighth C3k2 layer all participate in feature fusion, and finally, object detection of multi-scale features is achieved through the Detect module.
[0014] Further, in the step of training the YOLOv11n-CBAM network using the enhanced training set, validation set, and test set of the images to obtain the trained YOLOv11n-CBAM network, it includes: Training the YOLOv11n-CBAM network using the enhanced training set of the images until a preset convergence condition is met to obtain the trained YOLOv11n-CBAM network; Validating the trained YOLOv11n-CBAM network using the enhanced validation set of the images to obtain the validated YOLOv11n-CBAM network; Testing the validated YOLOv11n-CBAM network using the enhanced test set of the images to obtain the trained YOLOv11n-CBAM network.
[0015] Further, in the step of establishing a worker face database and constructing a Retinaface face recognition model in combination with the Retinaface algorithm, it includes: Establishing the worker face database according to the face images of the workers; Using the Retinaface algorithm to perform key point detection and feature extraction on the faces of each worker to obtain face data; Matching the face data with the face images in the worker face database to construct the Retinaface face recognition model.
[0016] Further, in the step of obtaining worker operation images and inputting the worker operation images into the trained YOLOv11n-CBAM network and the Retinaface face recognition model respectively to obtain abnormal behavior detection results and face recognition information, it includes: Obtaining the worker operation video captured by the monitoring system and decoding the worker operation video to obtain the worker operation images; Using the multi-thread processing module to input the worker operation images into the trained YOLOv11n-CBAM network and the Retinaface face recognition model respectively; The trained YOLOv11n-CBAM network detects the worker operation image to obtain the abnormal behavior detection result; The etinaface face recognition model performs face recognition on the worker operation image to obtain the face recognition information.
[0017] Further, in the step of generating an alarm result according to the abnormal behavior detection result and the face recognition information, it includes: Using the data collection module to converge the abnormal behavior detection result and the face recognition information to obtain the alarm result; Sending the alarm result to the monitoring service center for early warning disposal and subsequent analysis.
[0018] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: In the embodiments of the present disclosure, through the above-mentioned method for identifying abnormal behaviors of workers based on edge computing, on the one hand, the method introduces the CBAM module into the detection head to optimize feature expression to improve detection efficiency and accuracy; combined with the Retinaface algorithm, it realizes accurate recognition of workers' faces and ensures stable performance under complex lighting and occlusion conditions. On the other hand, the method uses the edge computing architecture and is deployed on edge devices close to the data source, significantly reducing data transmission latency and improving processing real-time performance. Through multi-threaded processing and the data collection module, parallel inference of multiple models for the same video frame is realized. The parallel multi-threads simultaneously input the decoded video frame into the trained YOLOv11n-CBAM network and the Retinaface face recognition model. The data collection module converges the abnormal behavior detection and face recognition results and transmits them to the monitoring service center through the HTTP protocol, providing efficient support for real-time early warning and subsequent analysis. This method comprehensively improves the monitoring efficiency and intelligent level in industrial scenarios. Description of the Drawings
[0019] The drawings here are incorporated into the specification and constitute a part of this specification, showing the embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 A step diagram showing a method for identifying abnormal behaviors of workers based on edge computing in an exemplary embodiment of the present disclosure; Figure 2 A specific flowchart showing the processing of worker operation images in an exemplary embodiment of the present disclosure; Figure 3Shows the P-R curve of the trained YOLOv11-CBAM network in the test set in an exemplary embodiment of the present disclosure; Figure 4 Shows the convergence curve of the trained YOLOv11-CBAM network in the test set in an exemplary embodiment of the present disclosure. Detailed implementation manners
[0021] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.
[0022] In addition, the accompanying drawings are only schematic illustrations of the embodiments of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.
[0023] In this example embodiment, a method for identifying abnormal behaviors of workers based on edge computing is provided. Referring to Figure 1 as shown in, the method for identifying abnormal behaviors of workers based on edge computing may include: step S101 to step S107.
[0024] Step S101: Construct a basic data set and divide the basic data set into a training set, a validation set, and a test set according to a preset ratio; Step S102: Perform image enhancement on the training set, the validation set, and the test set; Step S103: Construct a YOLOv11n-CBAM network; wherein, the YOLOv11n-CBAM network includes a backbone network and a detection head; Step S104: Use the image-enhanced training set, validation set, and test set to train the YOLOv11n-CBAM network to obtain the trained YOLOv11n-CBAM network; Step S105: Establish a worker face database and construct a Retinaface face recognition model in combination with the Retinaface algorithm; Step S106: Obtain worker operation images and input the worker operation images into the trained YOLOv11n-CBAM network and the Retinaface face recognition model respectively to obtain abnormal behavior detection results and face recognition information; Step S107: Generate an alarm result based on the abnormal behavior detection result and the face recognition information.
[0025] Through the above-mentioned method for identifying abnormal behaviors of workers based on edge computing, on the one hand, the method introduces the CBAM module into the detection head to optimize feature representation to improve detection efficiency and accuracy; combined with the Retinaface algorithm, it realizes the accurate recognition of workers' faces and ensures stable performance under complex lighting and occlusion conditions. On the other hand, the method uses the edge computing architecture and is deployed on edge devices close to the data source, significantly reducing data transmission latency and improving processing real-time performance. Through multi-threaded processing and the data collection module, parallel inference of multiple models for the same video frame is achieved. The parallel multi-threads simultaneously input the decoded video frames into the trained YOLOv11n-CBAM network and the Retinaface face recognition model. The data collection module aggregates the abnormal behavior detection and face recognition results and transmits them to the monitoring service center through the HTTP protocol, providing efficient support for real-time warning and subsequent analysis. This method comprehensively improves the monitoring efficiency and intelligent level in industrial scenarios.
[0026] Next, reference will be made to Figures 1 to 4 to describe each step of the above-mentioned method for identifying abnormal behaviors of workers based on edge computing in the present exemplary embodiment in more detail.
[0027] In step S101, a basic data set is constructed, and the basic data set is divided into a training set, a validation set, and a test set according to a preset ratio; image enhancement is performed on the training set, the validation set, and the test set.
[0028] In one embodiment, constructing the basic data set specifically involves collecting images of workers during their work in the workshop, including scenarios related to wearing safety helmets, not wearing safety helmets, holding goods by hand, hooks, suspended objects, and people, as sample images of the basic data set, and using the Labelimg tool to label these images with corresponding tags. The original image size in the data set is 1920×1080, which is uniformly scaled to 640×640 to provide data support for subsequent model training.
[0029] The basic data set is randomly divided into a training set, a validation set, and a test set in a ratio of 7:2:1 through the random.shuffle( ) function.
[0030] In step S102, image enhancement is performed on the training set, the validation set, and the test set.
[0031] In one embodiment, the online data augmentation of the basic data set is specifically as follows: The divided sample set is subjected to image enhancement to improve the generalization ability and robustness of the model. The Albumentations library is introduced in ultralytics / data / augment.py to perform various enhancement operations on the basic dataset, including adding fog effects to simulate low visibility environments, blurring to restore motion blur or focal length offset scenarios, grayscaling to enhance the model's adaptability to color changes, gamma correction to optimize the bright-dark contrast effect, and mosaic augmentation to generate complex scene samples by randomly stitching multiple images.
[0032] In step S103, the YOLOv11n-CBAM network is constructed; among them, the YOLOv11n-CBAM network includes a backbone network and a detection head.
[0033] CBAM (Convolutional Block Attention Module) is a lightweight attention mechanism module designed to enhance the key feature extraction ability of convolutional neural networks. CBAM adaptively optimizes the input features by jointly using channel attention and spatial attention. First, the channel attention module generates a channel-level attention weight distribution through global average pooling and global max pooling operations, thus highlighting important channel features; subsequently, the spatial attention module generates a spatial-level attention distribution by applying convolutional operations in the spatial dimension of the feature map, further focusing on key regions. The introduction of the CBAM module can improve the network's identification ability for target regions while maintaining computational efficiency, and is suitable for resource-constrained scenarios such as edge computing.
[0034] The construction method of the YOLOv11-CBAM network is as follows: The CBAM module is introduced into the head part of the YOLOv11n model, and the feature expression is optimized through its channel attention and spatial attention mechanisms to highlight the weights of key features. Create a yolov11-cbam.yaml file in the ultralytics / cfg / models / 11 directory, and add an attention layer after the three-layer C3K2 module, with the specific format of [-1, 1, CBAM, ], to obtain the improved YOLOv11n network, that is, the YOLOv11n-CBAM network.
[0035] The backbone network of the YOLOv11n-CBAM network includes a Conv convolutional layer, a first C3k2 layer, a second Conv convolutional layer, a second C3k2 layer, a third Conv convolutional layer, a third C3k2 layer, a fourth Conv convolutional layer, a fourth C3k2 layer, an SPPF layer, and a C2PSA layer connected in sequence. Among them, the outputs of the second C3k2 layer, the third C3k2 layer, and the C2PSA layer are all used as the feature outputs of the backbone network. The backbone network extracts multi-level feature information at five scales of P1 / 2, P2 / 4, P3 / 8, P4 / 16, and P5 / 32 through a hierarchical feature extraction module. Among them, the C2PSA layer introduces a Polarized Self-Attention mechanism to optimize the feature channel and spatial weight distribution.
[0036] The detection head of the YOLOv11n-CBAM network is composed of a first Upsample layer, a first Concat layer, a fifth C3k2 layer, a second Upsample layer, a second Concat layer, a sixth C3k2 layer, a first downsampling Conv layer, a third Concat layer, a seventh C3k2 layer, a second downsampling Conv layer, a fourth Concat layer, and a Detect layer. A CBAM module is embedded after each C3k2 module in the detection head, specifically including the sixth C3k2 layer (16), the seventh C3k2 layer (19), and the eighth C3k2 layer (22). The outputs of the sixth C3k2 layer, the seventh C3k2 layer, and the eighth C3k2 layer all participate in feature fusion, and finally, object detection of multi-scale features (P3, P4, P5) is achieved through the Detect module.
[0037] In step S104, the YOLOv11n-CBAM network is trained using the image-augmented training set, validation set, and test set to obtain the trained YOLOv11n-CBAM network.
[0038] For example, the YOLOv11n-CBAM network is trained using the image-augmented training set until a preset convergence condition is met to obtain the trained YOLOv11n-CBAM network.
[0039] The YOLOv11n-CBAM network trained using the image-augmented validation set is validated to obtain the validated YOLOv11n-CBAM network.
[0040] The validated YOLOv11n-CBAM network is tested using the image-augmented test set to obtain the trained YOLOv11n-CBAM network.
[0041] In one embodiment, the yolov11-cbam.yaml script is tested using the test_model.py test script to verify its inference ability and correctness; the YOLOv11n-CBAM detection model is trained using the training set until a preset convergence condition is met, and the hyperparameters of the object detection model, including at least the dropout rate, weight decay rate, and learning rate, are continuously optimized during the training process.
[0042] In step S105, a worker face database is established, and a Retinaface face recognition model is constructed in combination with the Retinaface algorithm.
[0043] In one embodiment, Retinaface is a face detection model based on a single-stage architecture, focusing on high-precision and high-efficiency face detection tasks. Its core design lies in integrating a multi-task learning framework, combining face detection, key point localization, and facial pose estimation in one network. Retinaface uses a Feature Pyramid Network (FPN) to enhance its multi-scale feature extraction ability, thus performing particularly well in small-scale face detection. In addition, Retinaface introduces a pixel-level supervision strategy, further improving the detection accuracy by densely sampling and refining predictions for the face region. Compared with traditional methods, Retinaface balances detection performance and inference speed, having significant advantages in terms of real-time performance and accuracy.
[0044] The worker face images are uploaded to the face database; then, the Retinaface algorithm is used to detect and extract the features of the face key points, and the collected face data is quickly compared with the data in the face database, enabling accurate determination of the worker's identity even under complex lighting or occlusion conditions.
[0045] In steps S106 and S107, the worker operation images are obtained and input into the trained YOLOv11n-CBAM network and the Retinaface face recognition model respectively to obtain abnormal behavior detection results and face recognition information; an alarm result is generated based on the abnormal behavior detection result and the face recognition information.
[0046] Exemplarily, the worker operation video captured by the monitoring system is obtained and decoded to obtain the worker operation images.
[0047] The multi-thread processing module is used to input the worker operation images into the trained YOLOv11n-CBAM network and the Retinaface face recognition model respectively.
[0048] The trained YOLOv11n-CBAM network is used to detect the images of workers' operations to obtain the abnormal behavior detection results.
[0049] The Retinaface face recognition model is used to perform face recognition on the images of workers' operations to obtain the face recognition information.
[0050] The data collection module is used to converge the abnormal behavior detection results and the face recognition information to obtain the alarm results.
[0051] The alarm results are sent to the monitoring service center for early warning disposal and subsequent analysis.
[0052] In one embodiment, as Figure 2 shown, it is the specific flowchart for processing the images of workers' operations.
[0053] The edge intelligent device deploys the YOLOv11n-CBAM detection model and the Retinaface face recognition model. Specifically: First, the edge side pulls data from multiple real-time video streams transmitted by the RTSP (Real Time Streaming Protocol) protocol. After each video is decoded by the decoding module, the video frames are simultaneously allocated to the YOLOv11n-CBAM detection model and the Retinaface face recognition model for parallel inference. During the inference process, the YOLOv11n-CBAM model is used for abnormal behavior detection, and the Retinaface model is used for face recognition. Subsequently, the abnormal behavior detection results and the face recognition results are converged through the collection module, and the JSON format alarm results are sent to the monitoring service center in the HTTP transmission mode for early warning disposal and subsequent analysis.
[0054] In a specific embodiment, the edge intelligent device adopts a model based on MLIR compilation and quantization. The system is of the ARM64 architecture and runs Ubuntu 20.04. The trained YOLOv11-CBAM detection model and the face recognition model are deployed to the edge computing device. By connecting to multiple camera monitoring devices deployed in the production workshop, video data is collected and processed. The edge intelligent device performs real-time algorithm inference on the received video data. When detecting the abnormal behavior of workers, the device transmits the recognition information to the service center by sending an HTTP request to generate an alarm message, thereby triggering the corresponding alarm disposal process.
[0055] In a specific embodiment, to verify the effectiveness of the present application, for the unsafe behaviors of workers in the safety production workshop such as not wearing safety helmets, holding the lifting hook by hand, and lifting items, an 8,000-image self-built sample dataset is collected, and the annotation content is divided into six categories, namely not wearing a safety helmet (No_Helmet), wearing a safety helmet (Helmet), holding the sling by hand (Hand_held_sling), lifting hook (Hanger), lifting item (Lifting_item), and people (People). The sample set is divided into a training set, a validation set, and a test set in a ratio of 7:2:1.
[0056] Figure 3 The P-R curve of the YOLOv11-CBAM network is given. It can be seen from the P-R curve that the model performs excellently in the performance of each detection category. Among them, the curves of categories such as 'People', 'Helmet', and 'Lifting_item' are close to the upper right corner of the coordinate axis, showing a high precision and recall rate. Considering the performance of all categories, the area enclosed by the overall curve (all classes) and the coordinate axis is relatively large, indicating that the model has a balanced detection ability for different categories and a high mean average precision (mAP). This result verifies the powerful performance of the model on this dataset and can efficiently and accurately complete the detection task of workers' abnormal behaviors.
[0057] Figure 4 The convergence curves of the key loss functions and performance indicators of the YOLOv11-CBAM network in the training and validation stages are shown. The upper half shows the bounding box loss (train / box_loss), object loss (train / obj_loss), and class loss (train / cls_loss) of the training set, and the lower half shows the corresponding loss curves of the validation set (val / box_loss, val / obj_loss, val / cls_loss). The right part shows the convergence of the model in performance indicators such as Precision, Recall, mAP@0.5, and mAP@0.5:0.95. The abscissa represents the number of epochs, and the ordinate represents the corresponding index value. The curves show that as the training progresses, the loss gradually decreases, and the performance indicators gradually improve and tend to be stable, indicating that the model shows good convergence and stability on this dataset.
[0058] In summary, the present application can not only meet the requirements of real-time and accuracy for the detection of workers' abnormal behaviors, but also further improve the intelligent level of safety production detection through functions such as alarm data reporting, providing efficient and reliable technical support for safety production.
[0059] Through the above-mentioned method for identifying abnormal behaviors of workers based on edge computing, the CBAM module is introduced by improving the YOLOv11n network to optimize feature representation and enhance detection efficiency and accuracy. In combination with the Retinaface algorithm, accurate identification of workers' faces is achieved to ensure stable performance under complex lighting and occlusion conditions. This method utilizes the edge computing architecture and is deployed on edge devices close to the data source, significantly reducing data transmission latency and enhancing processing real-time performance. Through multi-threaded processing and the data collection module, parallel inference of multiple models for the same video frame is realized. The parallel multi-threads simultaneously input the decoded video frame into the object detection model and the face recognition model. The data collection module aggregates the results of abnormal behavior detection and face recognition and transmits them to the monitoring service center via the HTTP protocol, providing efficient support for real-time warning and subsequent analysis. This application comprehensively improves the monitoring efficiency and intelligent level in industrial scenarios.
[0060] It should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. in the above description is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the embodiments of the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the embodiments of the present disclosure.
[0061] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, the meaning of "plurality" is two or more, unless otherwise specifically defined.
[0062] In the embodiments of the present disclosure, unless otherwise clearly specified and limited, the terms "install", "connect", "couple", "fix", etc. should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present disclosure can be understood according to specific circumstances.
[0063] In the embodiments of the present disclosure, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature may include direct contact between the first and second features, or may include the first and second features not being in direct contact but being in contact through additional features therebetween. Moreover, the first feature being "above", "over" and "on top of" the second feature includes the first feature being directly above and obliquely above the second feature, or merely indicating that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "below" and "beneath" the second feature includes the first feature being directly below and obliquely below the second feature, or merely indicating that the horizontal height of the first feature is less than that of the second feature.
[0064] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example" or "some examples", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0065] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
Claims
1. A method for identifying abnormal behavior of workers based on edge computing, characterized in that: The method includes: Constructing a basic data set, and dividing the basic data set into a training set, a validation set, and a test set according to a preset ratio; Performing image enhancement on the training set, the validation set, and the test set; Construct a YOLOv11n-CBAM network; wherein the YOLOv11n-CBAM network includes a backbone network and a detection head; The YOLOv11n-CBAM network is trained using the image-enhanced training set, the validation set, and the test set to obtain the trained YOLOv11n-CBAM network; Establish a worker face database and combine it with the Retinaface algorithm to build a Retinaface face recognition model; Obtain worker work images, and input the worker work images into the trained YOLOv11n-CBAM network and the Retinaface face recognition model respectively to obtain abnormal behavior detection results and face recognition information; An alarm result is generated according to the abnormal behavior detection result and the face recognition information.
2. The method for identifying abnormal worker behavior based on edge computing according to claim 1 is characterized in that: The step of constructing a basic data set and dividing the basic data set into a training set, a validation set and a test set according to a preset ratio includes: Collecting work images of workers during work and annotating the work images to generate a basic data set; wherein the work images at least include images of workers wearing safety helmets, images of workers not wearing safety helmets, images of hand-held goods, images of hooks, images of hanging objects, and images of people; The basic data set is randomly divided into the training set, the validation set and the test set according to a preset ratio through the random.shuffle() function; wherein the preset ratio is 7:2:
1.
3. The method for identifying abnormal worker behavior based on edge computing according to claim 1 is characterized in that: The step of performing image enhancement on the training set, the validation set and the test set comprises: Add fog effects to simulate low visibility environments; Blur processing to restore motion blur or focus shift scenes; Grayscale to enhance the model's adaptability to color changes; Gamma correction to optimize light-dark contrast; Mosaic enhancement generates complex scene samples by randomly splicing multiple images.
4. The method for identifying abnormal worker behavior based on edge computing according to claim 1 is characterized in that: The backbone network includes: The Conv convolution layer, the first C3k2 layer, the second Conv convolution layer, the second C3k2 layer, the third Conv convolution layer, the third C3k2 layer, the fourth Conv convolution layer, the fourth C3k2 layer, the SPPF layer and the C2PSA layer are connected in sequence; wherein, The outputs of the second C3k2 layer, the third C3k2 layer and the C2PSA layer are all used as feature outputs of the backbone network; The backbone network extracts multi-level feature information at five scales of P1 / 2, P2 / 4, P3 / 8, P4 / 16 and P5 / 32 through a step-by-step feature extraction module; The C2PSA layer introduces a Polarized Self-Attention mechanism to optimize feature channels and spatial weight distribution.
5. The method for identifying abnormal worker behavior based on edge computing according to claim 4 is characterized in that: The detection head comprises: The first Upsample layer, the first Concat layer, the fifth C3k2 layer, the second Upsample layer, the second Concat layer, the sixth C3k2 layer, the first downsampled Conv layer, the third Concat layer, the seventh C3k2 layer, the second downsampled Conv layer, the fourth Concat layer and the Detect layer are sequentially connected; wherein, A CBAM module is embedded after each of the C3k2 modules. The CBAM module includes a sixth C3k2 layer, a seventh C3k2 layer and an eighth C3k2 layer. The output of the sixth C3k2 layer, the output of the seventh C3k2 layer and the output of the eighth C3k2 layer all participate in feature fusion, and finally the target detection of multi-scale features is realized through the Detect module.
6. The method for identifying abnormal worker behavior based on edge computing according to claim 1 is characterized in that: The step of training the YOLOv11n-CBAM network using the image-enhanced training set, the validation set, and the test set to obtain the trained YOLOv11n-CBAM network includes: The YOLOv11n-CBAM network is trained using the training set after image enhancement until a preset convergence condition is met to obtain the trained YOLOv11n-CBAM network; Verify the YOLOv11n-CBAM network trained by the verification set after image enhancement to obtain the verified YOLOv11n-CBAM network; The verified YOLOv11n-CBAM network is tested using the test set after image enhancement to obtain the trained YOLOv11n-CBAM network.
7. The method for identifying abnormal worker behavior based on edge computing according to claim 1 is characterized in that: The steps of establishing a worker face database and combining it with the Retinaface algorithm to build a Retinaface face recognition model include: Establishing the worker face database according to the worker's face images; Using the Retinaface algorithm to perform key point detection and feature extraction on the faces of each worker to obtain face data; The facial data is matched with the facial images in the worker face library to construct the Retinaface facial recognition model.
8. The method for identifying abnormal behavior of workers based on edge computing according to claim 1 is characterized in that: The steps of obtaining worker work images and inputting the worker work images into the trained YOLOv11n-CBAM network and the Retinaface face recognition model respectively to obtain abnormal behavior detection results and face recognition information include: Obtaining a worker's work video captured by a monitoring system, and decoding the worker's work video to obtain the worker's work image; Using a multi-threaded processing module, the worker work image is input into the trained YOLOv11n-CBAM network and the Retinaface face recognition model respectively; The trained YOLOv11n-CBAM network detects the worker's work image to obtain the abnormal behavior detection result; The etinaface face recognition model performs face recognition on the worker's work image to obtain the face recognition information.
9. The method for identifying abnormal worker behavior based on edge computing according to claim 1 is characterized in that: The step of generating an alarm result according to the abnormal behavior detection result and the face recognition information includes: A data collection module is used to collect the abnormal behavior detection result and the face recognition information to obtain the alarm result; The alarm result is sent to the monitoring service center for early warning disposal and subsequent analysis.