A production environment mobile phone illegal use detection method and system based on YOLOv12 optimization
By improving the YOLOv12 architecture and introducing the MLA-C2F, MCAF, and LK-GLU modules, the accuracy and real-time detection issues of illegal mobile phone use in power operation environments are solved, the detection accuracy and efficiency are improved, and the safety supervision needs in complex environments are adapted.
Patent Information
- Application Number
- CN202511198310.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-08-26
AI Technical Summary
Existing deep learning methods for detecting illegal use of mobile phones in power operation environments have problems such as insufficient small target recognition ability, poor adaptability to complex environments, and limited real-time performance, making it difficult to meet safety supervision needs in high-risk scenarios.
An improved YOLOv12 architecture is adopted, the MLA-C2F module is introduced for feature extraction, the MCAF module is combined for feature fusion, and the LK-GLU module is used for feature output to enhance the network's representation ability and adaptability, and improve detection accuracy and efficiency through a multi-scale attention mechanism.
It improves the accuracy and efficiency of detecting illegal mobile phone use, adapts to complex power operation environments, achieves higher recognition accuracy and faster response speed, and meets the safety supervision needs of power production environments.
Smart Images

Figure CN120689760B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization. Background Art
[0002] While mobile phones have become essential tools for daily communication and information processing, their use in the high-risk production environment of the power industry poses serious safety risks, frequently causing accidents and even casualties. To prevent accidents caused by mobile phone use during power operations, such as distraction, electromagnetic interference, or fire, the power system has implemented strict restrictions on mobile phone use in the workplace. The power industry has developed numerous safety regulations, including the "Ten Prohibitions," which explicitly prohibit employees from carrying or using mobile phones during hazardous situations such as operating high-voltage equipment, performing hot work, and working at heights. They also emphasize the centralized storage and management of mobile phones during critical operations, such as dispatching and switching operations. Furthermore, key positions, such as operators and maintenance personnel, are required to turn off or silence their phones, and their non-work use is strictly restricted. If an accident occurs due to illegal mobile phone use, those responsible will face performance deductions, suspension, and even legal prosecution. This practice in the power industry fully embodies the principle of "safety first, prevention first" and is highly consistent with the safety management philosophy of "no work if it's not safe."
[0003] Therefore, driven by the stringent requirements for industrial safety and employee safety in high-risk power system operations, an increasing number of power plants are developing and implementing mobile phone management policies, strictly restricting or prohibiting employee mobile phone use during operations. Currently, traditional monitoring methods primarily fall into two categories: one is to establish manual checkpoints at the entrance to the work area, prohibiting employees from bringing mobile phones into critical areas; the other is to rely on safety officers to conduct continuous patrols to monitor on-site behavior. However, these methods have significant limitations, particularly in high-intensity, high-frequency operations. Traditional manual monitoring is not only inefficient but also prone to oversight, making it difficult to immediately detect and effectively prevent illegal use. Therefore, deep learning technology is being introduced to intelligently analyze video footage from surveillance cameras installed in key operational areas, such as substations and distribution rooms. This technology automatically identifies illegal mobile phone use by employees and issues real-time warnings and records. This not only improves the accuracy and real-time nature of on-site safety monitoring, but also effectively reduces labor input and improves management efficiency.
[0004] With the continuous advancement of deep learning technology, its application in the power industry for behavioral recognition and violation detection is becoming increasingly feasible. Compared to traditional manual inspections and rule-matching methods, deep learning-based detection systems demonstrate significant advantages in recognition accuracy, stability, and processing efficiency. At power operation sites, deep learning can analyze video surveillance streams in real time to automatically identify whether workers engage in irregular behaviors such as illegal mobile phone use and improper wearing of protective equipment, effectively enhancing the intelligence and refinement of on-site safety management. Furthermore, AI systems can automatically record recognition results, creating a traceable data record that provides strong support for accident tracing, hazard investigation, and subsequent management. They can also issue timely warnings, enabling rapid response and intervention to risky behaviors. Deep learning is an advanced algorithm based on artificial neural networks. By building a multi-layered neural network structure to mimic the information processing mechanisms of the human brain, it can automatically extract features from large amounts of data and continuously optimize model weights using a backpropagation algorithm to improve recognition performance. In recent years, technologies such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and generative adversarial networks (GANs) have been widely applied in image and video processing, providing strong support for behavioral detection in power operation scenarios. Deep learning methods, exemplified by real-time object detection algorithms like YOLO, are capable of rapidly processing large amounts of image data, making them particularly well-suited for efficiently identifying small and medium-sized objects (such as handheld mobile phones) in power plants. Furthermore, deep learning models possess excellent generalization and adaptability, enabling them to improve accuracy through continuous training and adapt to diverse environmental conditions and operational contexts. Numerous studies and practices have demonstrated that deep learning-based systems offer significant advantages over traditional visual inspection methods in terms of recognition accuracy, response speed, and stability.
[0005] Berri et al. designed a mobile phone usage detection system for driving environments. The system uses images captured by in-car cameras and combines image feature extraction with a support vector machine (SVM) classifier to determine whether the driver is using their phone. When tested on a self-constructed dataset (containing 100 positive and negative sample images), the system achieved a recognition accuracy of 91.57%. When applied to a 3-second video clip recognition task, the system achieved an accuracy of 87.43%, demonstrating excellent video recognition capabilities. Li Zhiwei et al. used the SSD algorithm to conduct mobile phone detection research. They first enhanced the original image data to improve image quality and increase the sample size, thereby mitigating the negative impact of insufficient data. In terms of model architecture, they introduced a DenseNet as the backbone network and optimized it, significantly improving detection accuracy and bounding box localization capabilities. Experimental results show that this improved SSD model achieves better performance than the original model on the constructed dataset. Dai Teng et al. constructed an end-to-end neural network named OMPDNet based on YOLOv4 to solve the problem of detecting small objects (such as mobile phones) in real-world driving environments. To address the small size and unclear features of mobile phone targets, they redesigned the feature extraction module and used the K-Means clustering algorithm to adaptively adjust the size of the prior box, thereby improving the model's detection capabilities in complex backgrounds. Experimental verification shows that OMPDNet not only has the advantages of high efficiency and high accuracy in detecting mobile phone targets, but also demonstrates good generalization and robustness in other small target detection tasks.
[0006] However, traditional neural network detection methods still have flaws. For example, the method proposed by Berri et al., while achieving good experimental results, relies on handcrafted features and a shallow model, resulting in weak generalization and adaptability to complex environments. It is no longer able to meet the high-precision and robustness requirements of current industrial scenarios. The target detection algorithm proposed by Li Zhiwei et al., based on SSD, is an early representative of one-stage detection algorithms. However, it still has inherent deficiencies in small target detection and has limited adaptability to occlusion, complex backgrounds, and other situations. With the iterative updates of detection technology, its performance has been gradually superseded by more advanced algorithms.
[0007] In summary, while the aforementioned research provides early insights and technical implementations for mobile phone detection, the methods employed are often based on traditional or early deep learning frameworks, resulting in issues such as insufficient small object recognition, poor adaptability to complex industrial environments, and limited real-time performance. With the advancement of deep learning, particularly the Transformer architecture and real-time detection algorithms, the application of next-generation detection technologies (such as RT-DETR, YOLOv11, and YOLOv12) in high-risk scenarios like power generation will become an important direction for improving detection accuracy, stability, and system response efficiency. Summary of the Invention
[0008] The present invention provides a method and system for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization, aiming to solve at least one of the above technical problems.
[0009] To achieve the above objectives, the present invention provides a method for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization, the method comprising the following steps:
[0010] S1: Construct an improved YOLOv12 architecture; wherein the improved YOLOv12 architecture includes: a feature extraction layer configured with an MLA-C2F module, a feature fusion layer configured with an MLA-C2F module, an MCAF module, and an LK-GLU module, and a prediction output layer having three core detection modules;
[0011] In the feature extraction layer and feature fusion layer: the MLA-C2F module is configured to use the multi-branch key feature design of the local attention mechanism to represent the query constructing adaptive cross-scale response and model the interaction behavior between the character and the target;
[0012] In the feature fusion layer: the MCAF module is configured to use multi-channel attention fusion to complementarily fuse shallow features and deep features, and the LK-GLU module is configured to dynamically adjust the fusion feature output of the correlation features between the mobile phone and the surrounding environment and the original features using channel segmentation design;
[0013] S2: Collect several scene images of illegal mobile phone use actions and normal production operation actions in the target production environment, annotate the scene images with illegal mobile phone use actions, convert the annotated images into training annotation data for YOLOv12, and construct a training sample set;
[0014] S3: Use the training sample set to train the constructed improved YOLOv12 architecture to obtain a trained production environment mobile phone misuse detection model;
[0015] S4: Inputting the image to be detected into the production environment mobile phone illegal use detection model to obtain the production environment mobile phone illegal use detection result output by the production environment mobile phone illegal use detection model.
[0016] Optionally, the feature extraction layer specifically includes: a first Conv convolution layer, a second Conv convolution layer, a first C3k2 feature extraction module, a third Conv convolution layer, a second C3k2 feature extraction module, a fourth Conv convolution layer, a first MLA-C2f module, a fifth Conv convolution layer and a second MLA-C2f module connected in sequence.
[0017] Optionally, the feature fusion layer specifically includes: a first MCAF module, a second MCAF module, a third MLA-C2f module, a sixth Conv convolution layer, a third MCAF module, a fourth MLA-C2f module, a seventh Conv convolution layer, a fourth MCAF module and an LK-GLU module connected in sequence.
[0018] Optionally, the output ends of the first MLA-C2f module and the second MLA-C2f module are connected to the input end of the first MCAF module, and the output end of the second C3k2 feature extraction module is connected to the input end of the second MCAF module; the output end of the first MCAF module is connected to the input end of the fourth MCAF module, and the output end of the second MCAF module is connected to the input end of the third MCAF module.
[0019] Optionally, the feature fusion layer specifically includes: a first Detect core detection module connected to the output end of the third MLA-C2f module, a second Detect core detection module connected to the output end of the fourth MLA-C2f module, and a third Detect core detection module connected to the output end of the LK-GLU module.
[0020] Optionally, the MLA-C2F module specifically includes:
[0021] The key branch is configured to perform direct normalization, normalization after extraction with a 5×5 standard convolution kernel, and normalization after extraction with a 7×7 standard convolution kernel on the input and divided features, modeling different levels of semantic relationships between the interaction between the character and the target;
[0022] The query branch is configured to perform dot product calculations on the input and divided features with the three output features of the key branch, and perform weighted fusion of the three dot product calculation results;
[0023] The Value branch is configured to perform attention interaction between the input and divided features and the weighted fusion features of the Query branch and output them.
[0024] Optionally, the MCAF module specifically includes:
[0025] A sampling unit is configured to downsample shallow features in the initial features, upsample deep features in the initial features, and cross-join the upsampled results with the downsampled results;
[0026] A feature extractor configured to extract features output by the sampling unit respectively;
[0027] The dual-path attention module is configured to perceive and enhance the features extracted by the feature extractor, and interacts with the initial features using a point-by-point multiplication method;
[0028] The normalization layer is configured to normalize the information interaction of the two-way attention module and then output it.
[0029] Optionally, the feature extractor is configured to adopt a convolutional multi-layer perceptron architecture in which the activation function is replaced by PReLu.
[0030] Optionally, the LK-GLU module specifically includes:
[0031] The head Conv convolution layer is configured to adjust the channel dimension and compress the number of channels of the input feature map, and perform linear transformation on the output features to extract basic features;
[0032] The initial feature transfer branch is configured to directly transfer the features output by the head Conv convolutional layer;
[0033] The initial feature processing branch is configured to use a large kernel convolution layer and a Gaussian error linear unit to extract features from the output of the head Conv convolution layer and dynamically adjust the phone to the surrounding environment;
[0034] The tail Conv convolution layer is configured to perform point-by-point multiplication of the features of the initial feature transfer branch and the initial feature processing branch and perform channel adjustment and feature fusion on the processed features.
[0035] In addition, to achieve the above objectives, the present invention also provides a production environment mobile phone illegal use detection system based on YOLOv12 optimization, comprising:
[0036] An architecture building module for building an improved YOLOv12 architecture, wherein the improved YOLOv12 architecture includes: a feature extraction layer configured with an MLA-C2F module, a feature fusion layer configured with an MLA-C2F module, an MCAF module, and an LK-GLU module, and a prediction output layer having three core detection modules;
[0037] In the feature extraction layer and feature fusion layer: the MLA-C2F module is configured to use the multi-branch key feature design of the local attention mechanism to represent the query constructing adaptive cross-scale response and model the interaction behavior between the character and the target;
[0038] In the feature fusion layer: the MCAF module is configured to use multi-channel attention fusion to complementarily fuse shallow features and deep features, and the LK-GLU module is configured to dynamically adjust the fusion feature output of the correlation features between the mobile phone and the surrounding environment and the original features using channel segmentation design;
[0039] The sample construction module is used to collect several scene images of illegal mobile phone use actions and normal production operation actions in the target production environment, annotate the scene images with illegal mobile phone use actions, convert the annotated images into training annotation data for YOLOv12, and construct a training sample set;
[0040] The model training module is used to train the constructed improved YOLOv12 architecture using the training sample set to obtain a trained production environment mobile phone misuse detection model;
[0041] The model detection module is used to input the image to be detected into the production environment mobile phone illegal use detection model and obtain the production environment mobile phone illegal use detection result output by the production environment mobile phone illegal use detection model.
[0042] The beneficial effects of the present invention are as follows: a method and system for detecting illegal mobile phone use in production environments based on YOLOv12 optimization is proposed. Based on the YOLOv12 algorithm model, a multi-scale attention mechanism is introduced to enhance the network's representation capabilities by focusing on the importance of the channel and spatial dimensions of the feature graph. An innovative MLA-C2f module is designed to replace the traditional A2C2f module. Finally, an experimental comparison between the proposed YOLOv12-optimized model and the original YOLOv12 model verifies the effectiveness of the improved method in terms of accuracy and detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a flow chart of a method for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization according to an embodiment of the present invention;
[0044] Figure 2 This is a schematic diagram of the structure of a production environment mobile phone illegal use detection system based on YOLOv12 optimization according to an embodiment of the present invention;
[0045] Figure 3 Schematic diagram of the grid architecture of the improved YOLOv12 of the present invention;
[0046] Figure 4 This is a schematic diagram of the structure of the MLA-C2f module proposed in the present invention;
[0047] Figure 5 This is a schematic diagram of the structure of the MCAF module proposed in the present invention;
[0048] Figure 6 This is a schematic diagram of the structure of the LK-GLU module proposed in the present invention;
[0049] Figure 7 This is a schematic diagram of the average accuracy mean curve before and after the algorithm improvement;
[0050] Figure 8 Schematic diagram of the accuracy curve before and after algorithm improvement;
[0051] Figure 9 Schematic diagram of the bounding box loss curve before and after algorithm improvement. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0053] The embodiment of the present invention provides a method for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization, referring to Figure 1 , Figure 1 This is a flow chart of a method for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization according to an embodiment of the present invention.
[0054] In this embodiment, a method for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization specifically includes:
[0055] S1: Construct an improved YOLOv12 architecture; wherein the improved YOLOv12 architecture includes: a feature extraction layer configured with an MLA-C2F module, a feature fusion layer configured with an MLA-C2F module, an MCAF module, and an LK-GLU module, and a prediction output layer having three core detection modules;
[0056] In the feature extraction layer and feature fusion layer: the MLA-C2F module is configured to use the multi-branch key feature design of the local attention mechanism to represent the query constructing adaptive cross-scale response and model the interaction behavior between the character and the target;
[0057] In the feature fusion layer: the MCAF module is configured to use multi-channel attention fusion to complementarily fuse shallow features and deep features, and the LK-GLU module is configured to dynamically adjust the fusion feature output of the correlation features between the mobile phone and the surrounding environment and the original features using channel segmentation design;
[0058] S2: Collect several scene images of illegal mobile phone use actions and normal production operation actions in the target production environment, annotate the scene images with illegal mobile phone use actions, convert the annotated images into training annotation data for YOLOv12, and construct a training sample set;
[0059] S3: Use the training sample set to train the constructed improved YOLOv12 architecture to obtain a trained production environment mobile phone misuse detection model;
[0060] S4: Inputting the image to be detected into the production environment mobile phone illegal use detection model to obtain the production environment mobile phone illegal use detection result output by the production environment mobile phone illegal use detection model.
[0061] In a preferred embodiment, the feature extraction layer specifically includes: a first Conv convolution layer, a second Conv convolution layer, a first C3k2 feature extraction module, a third Conv convolution layer, a second C3k2 feature extraction module, a fourth Conv convolution layer, a first MLA-C2f module, a fifth Conv convolution layer and a second MLA-C2f module connected in sequence.
[0062] In a preferred embodiment, the feature fusion layer specifically includes: a first MCAF module, a second MCAF module, a third MLA-C2f module, a sixth Conv convolution layer, a third MCAF module, a fourth MLA-C2f module, a seventh Conv convolution layer, a fourth MCAF module and an LK-GLU module connected in sequence.
[0063] In a preferred embodiment, the output ends of the first MLA-C2f module and the second MLA-C2f module are connected to the input end of the first MCAF module, and the output end of the second C3k2 feature extraction module is connected to the input end of the second MCAF module; the output end of the first MCAF module is connected to the input end of the fourth MCAF module, and the output end of the second MCAF module is connected to the input end of the third MCAF module.
[0064] In a preferred embodiment, the feature fusion layer specifically includes: a first Detect core detection module connected to the output end of the third MLA-C2f module, a second Detect core detection module connected to the output end of the fourth MLA-C2f module, and a third Detect core detection module connected to the output end of the LK-GLU module.
[0065] In a preferred embodiment, the MLA-C2F module specifically includes:
[0066] The key branch is configured to perform direct normalization, normalization after extraction with a 5×5 standard convolution kernel, and normalization after extraction with a 7×7 standard convolution kernel on the input and divided features, modeling different levels of semantic relationships between the interaction between the character and the target;
[0067] The query branch is configured to perform dot product calculations on the input and divided features with the three output features of the key branch, and perform weighted fusion of the three dot product calculation results;
[0068] The Value branch is configured to perform attention interaction between the input and divided features and the weighted fusion features of the Query branch and output them.
[0069] In a preferred embodiment, the MCAF module specifically includes:
[0070] A sampling unit is configured to downsample shallow features in the initial features, upsample deep features in the initial features, and cross-join the upsampled results with the downsampled results;
[0071] A feature extractor configured to extract features output by the sampling unit respectively;
[0072] The dual-path attention module is configured to perceive and enhance the features extracted by the feature extractor, and interacts with the initial features using a point-by-point multiplication method;
[0073] The normalization layer is configured to normalize the information interaction of the two-way attention module and then output it.
[0074] In a preferred embodiment, the feature extractor is configured to adopt a convolutional multi-layer perceptron architecture with the activation function replaced by PReLu.
[0075] In a preferred embodiment, the LK-GLU module specifically includes:
[0076] The head Conv convolution layer is configured to adjust the channel dimension and compress the number of channels of the input feature map, and perform linear transformation on the output features to extract basic features;
[0077] The initial feature transfer branch is configured to directly transfer the features output by the head Conv convolutional layer;
[0078] The initial feature processing branch is configured to use a large kernel convolution layer and a Gaussian error linear unit to extract features from the output of the head Conv convolution layer and dynamically adjust the phone to the surrounding environment;
[0079] The tail Conv convolution layer is configured to perform point-by-point multiplication of the features of the initial feature transfer branch and the initial feature processing branch and perform channel adjustment and feature fusion on the processed features.
[0080] It should be noted that current deep learning-based mobile phone detection algorithms for production environments lack the ability to identify small targets, and their accuracy and average precision are insufficient to meet actual engineering requirements. This is specifically addressed in two aspects: 1. Real-time performance: In power production environments, mobile phone detection algorithms must achieve millisecond-level real-time response, which is the lifeline for core safety. Illegal use of mobile phones in high-voltage areas can cause fatal arc shocks in as little as 0.1 seconds. Real-time alerts are the last chance to save lives, and any delay could trigger an irreversible chain reaction of disasters. Therefore, ensuring real-time performance is a key factor in ensuring power grid security. 2. Model size: Given that the equipment used in engineering applications often has limited computing resources, lightweight models can better run on resource-constrained devices, ensuring system portability and deployment flexibility. Therefore, miniaturizing the model size is also a key issue.
[0081] In this embodiment, by integrating the MLA-C2f module for multi-level feature aggregation, using the LK-GLU module for large kernel convolution global correlation + cross-layer feature modulation, and introducing the multi-way attention fusion MCAF module, the model's spatial perception ability is improved, which helps to establish a better feature mapping relationship, enhances the model's feature representation ability, and improves the model's feature processing ability, which has obvious advantages over existing algorithms.
[0082] Reference Figure 2 , Figure 2 This is a structural diagram of a production environment mobile phone illegal use detection system based on YOLOv12 optimization according to an embodiment of the present invention.
[0083] like Figure 2 As shown, the production environment mobile phone illegal use detection system based on YOLOv12 optimization proposed in an embodiment of the present invention includes:
[0084] An architecture construction module 10 is used to construct an improved YOLOv12 architecture; wherein the improved YOLOv12 architecture includes: a feature extraction layer configured with an MLA-C2F module, a feature fusion layer configured with an MLA-C2F module, an MCAF module, and an LK-GLU module, and a prediction output layer having three core detection modules;
[0085] In the feature extraction layer and feature fusion layer: the MLA-C2F module is configured to use the multi-branch key feature design of the local attention mechanism to represent the query constructing adaptive cross-scale response and model the interaction behavior between the character and the target;
[0086] In the feature fusion layer: the MCAF module is configured to use multi-channel attention fusion to complementarily fuse shallow features and deep features, and the LK-GLU module is configured to dynamically adjust the fusion feature output of the correlation features between the mobile phone and the surrounding environment and the original features using channel segmentation design;
[0087] The sample construction module 20 is used to collect a number of scene images showing illegal mobile phone use and normal production operation in the target production environment, annotate the scene images with illegal mobile phone use, convert the annotated images into training annotation data for YOLOv12, and construct a training sample set;
[0088] The model training module 30 is used to train the constructed improved YOLOv12 architecture using the training sample set to obtain a trained production environment mobile phone illegal use detection model;
[0089] The model detection module 40 is used to input the image to be detected into the production environment mobile phone illegal use detection model, and obtain the production environment mobile phone illegal use detection result output by the production environment mobile phone illegal use detection model.
[0090] Other embodiments or specific implementations of the production environment mobile phone illegal use detection system based on YOLOv12 optimization of the present invention can refer to the above-mentioned method embodiments and will not be repeated here.
[0091] In order to explain the present invention more clearly, a specific example of detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization is provided below.
[0092] In order to improve the average precision and accuracy of mobile phone detection in actual application scenarios, this paper proposes a mobile phone detection model based on YOLOv12. Among them, the improved YOLOv12 architecture is as follows Figure 3As shown. First, for the feature extraction layer of YOLOv12, the present invention modifies the A2C2f layer, modifies the regional attention therein, adds multiple scale detection branches, and aggregates features of different scales to form an MLA-C2f (Multi-level aggregation C2f) module. The MLA-C2f module improves the spatial perception ability of the model and helps to establish a better feature mapping relationship. Secondly, the present invention designs the MCAF module to efficiently fuse features and enhance the feature representation ability of the model. Finally, the present invention also designs the LK-GLU module to improve the feature processing capability of the model. Finally, the present invention uses a data set of mobile phone detection in an electric power operation environment for model training, and compares the differences in the YOLOv12 algorithm models before and after improvement. The results show that the average detection accuracy and accuracy of the mobile phone detection algorithm for production environment optimized by YOLOv12 for mobile phone recognition of the present invention reached 93.68% and 89.47% respectively, which has obvious advantages over other algorithms and is more suitable for the needs of actual engineering applications. The detailed method of the present invention is described as follows:
[0093] (1) MLA-C2f module:
[0094] In this invention, we designed a new feature extraction module, the MLA-C2f module, to address the problems faced by mobile phone use detection in power production environments, such as difficulty in detecting small targets, strong interference from complex backgrounds, and diverse postures of people. This module combines multi-scale convolution with a local attention mechanism, aiming to extract semantic key features under different spatial receptive fields and enhance attention focusing capabilities through fine spatial guidance. The core idea is to use multi-branch key feature design to construct an adaptive cross-scale response representation for the query, thereby achieving detailed modeling of the interaction between people and targets. The specific structure of MLA-C2f is as follows: Figure 4 shown.
[0095] Specifically, the input features are first partitioned into a query branch and a key-value branch. The key is then further decomposed into three subpaths with different receptive fields and modeling purposes: key_local, key_middle, and key_high. Each subpath models semantic relationships at different levels through customized convolution and normalization operations. The key_local (local perception branch) branch does not perform any convolution operations and directly normalizes the input features. This design preserves the most original pixel-level detail information and is suitable for capturing very local and fine-grained target patterns, such as finger shapes, palm contours, and subtle geometric changes that occur when holding a phone.
[0096] In real-world scenarios, phones are often very small, especially when partially occluded or only partially visible. Only local perception channels that maintain pixel-level sensitivity can provide crucial clues for determining whether a phone is being held. Furthermore, by avoiding additional convolution processing, this path maintains minimal information loss and the highest spatial resolution, enabling the attention mechanism to focus on the target regions at the very edge of the image. The key_middle branch (mid-scale context branch) uses a standard 5×5 convolution kernel to spatially model the local neighborhood and extract region-level contextual semantics. The motivation for designing the mid-scale receptive field is to capture the spatial structural relationships between the phone and the hand, and between the hand and the face. For example, in detection scenarios, "hand close to the face" and "holding a suspected object" are often strong signatures of phone use. These relationships cannot be determined solely through pixel analysis; they must rely on modeling the spatial layout of local regions. Through its moderate receptive field, the key_middle branch effectively models these local structural relationships, providing structural semantic support for the attention mechanism while avoiding background interference introduced by global modeling.
[0097] The Key_high (global semantic guidance) branch uses a large 7×7 convolution kernel to expand the receptive field and capture the overall structure and semantic relationships of the human body over a wider range. Its design goal is to understand a person's posture, movement trends, and possible phone usage patterns from a global perspective. For example, in some scenarios, the phone is not exposed in the hand, but rather placed on the chest or to the side of the head. In these cases, local texture alone may not be sufficient to determine the holding state. A large receptive field helps the model infer potential interactive behaviors from the overall body contour and spatial configuration. Furthermore, this path has a certain suppression effect on complex backgrounds (such as equipment and cables) in power operation environments, and can determine semantic consistency through global context, thereby reducing false positives. Each branch is normalized after convolution, and a learnable positional bias is introduced to enhance the attention's ability to model spatial position information. Subsequently, the query vector is dot-producted with the three key branches to obtain three sets of attention weights. These are fused through a weighted mechanism to integrate semantic responses at each scale to form a cross-scale representation. Finally, the fused features undergo attention interaction with the value branch and are output. The entire modular structure combines C2f's lightweight concept and cross-level information integration capabilities.
[0098] Through the above design, the MLA-C2f module demonstrates significant advantages in several key aspects. First, the multi-scale key branch improves the modeling capabilities of small objects (mobile phones) and key parts of people, especially in complex scenes such as long-range shots and occlusions. Second, the spatial guidance mechanism helps the model actively focus on high-confidence areas, significantly suppressing background interference. Finally, local modeling instead of fully connected attention reduces computational costs, making the module suitable for deployment in edge computing platforms on industrial sites.
[0099] (2) MCAF multi-path attention fusion module:
[0100] In power operation scenarios, identifying whether workers are using mobile phones presents unique challenges. Due to the complex equipment and expansive landscapes in power operation environments, mobile phones, due to their inherent small size, only appear as a very small area in the captured image. Features such as their contours and textures are easily obscured by background information, resulting in insufficient effective features for identification. This requires the feature fusion module to accurately capture both shallow details and deep semantics. Shallow features typically have rich texture details but lack semantic information; deep features have sufficient semantic information but lack some details. The multi-channel attention fusion module (MCAF) uses multi-channel attention fusion to achieve a complementary fusion of the two types of features, providing high-quality features for subsequent precise detection of mobile phone usage behavior.
[0101] like Figure 5 As shown in the figure, MCAF consists of upsampling, downsampling, a feature extractor (FeatureExtractor), a dual attention module (DA), and a normalization layer. To leverage semantic layer complementarity to enhance the details and semantic information of small-target mobile phones and adapt to feature capture in complex environments, the MCAF module first downsamples shallow features and upsamples deep features, then cross-concatenates them to form distinct features with similar characteristics. The feature extractor then extracts each feature separately to obtain an efficient representation. The dual attention module then perceives and enhances the effective features, interacting with the initial features through point-by-point multiplication. Finally, the features are normalized and output.
[0102] Due to the outstanding performance of the Convolutional Multi-Layer Perceptron (CMP), the feature extractor in the MCAF module of this invention also adopts this module. The activation function is replaced with PReLU to introduce learnable parameters, avoiding activation failure caused by the unique lighting conditions of power operations (such as strong light and shadows) and enhancing feature expression. The DA module utilizes multiple parallel branches. Max pooling emphasizes prominent features such as the phone's "hard edges," while Avg pooling preserves overall regional statistics. This processing generates attention weights to suppress background interference, and the fusion branch enhances the differentiation of phone features in complex environments. The two branches work together, with the former combing through basic features and the latter focusing on key information, enhancing the diversity of the fused features.
[0103] The MCAF module addresses the issues of small size of mobile phones and scarce and difficult-to-identify image features. By cross-sampling and up-sampling, shallow features that retain fine-grained details and deep features that contain semantic understanding achieve adaptive complementarity between spatial dimensions and semantic levels. Relying on the rich details of the shallow layer and the clear semantics of the deep layer, it greatly expands the feature dimensions used to identify mobile phones. It also uses multi-dimensional complementarity to accurately capture subtle features of mobile phones and strengthen semantic expression. This allows the model to break through the limitations of a single feature in complex backgrounds, such as screen reflections under strong light or a corner showing in the distance. It enriches features and focuses on small targets, greatly improving the success rate of detecting small-target mobile phones and achieving accurate identification.
[0104] (3) LK-GLU module:
[0105] Power production scenarios often face complex background interference, lighting changes, occlusion and other problems, which require the model to be able to effectively distinguish mobile phones from other objects in the image. To solve this problem, the present invention introduces the LK-GLU (Large-KernelGated Linear Unit) module. This module adopts a channel segmentation design, combined with the advantages of the large kernel convolution structure, replacing the original C3k2 structure at the end of YOLOv12. The overall structure of the LK-GLU module is as follows: Figure 6 .
[0106] Specifically, the 1×1 convolution at the beginning of the module adjusts the number of feature channels. It adjusts the channel dimension of the input feature map, compressing the number of channels to meet subsequent processing requirements. It also uses linear transformations to extract basic features, laying the foundation for subsequent operations such as large-kernel convolution. The features are then split into two parts to reduce the computational overhead of the entire modality. Processing only a subset of features helps mitigate the impact of ambient noise on the model. One part, serving as a comparison branch, directly transfers the initial features, preserving the integrity of the input information and effectively mitigating gradient vanishing. This allows shallow-layer features to be smoothly transferred to deeper layers, providing a raw reference for feature fusion. The other branch sequentially passes through large-kernel convolution and Gaussian error linear units (GELUs). Large-kernel convolution, with its large receptive field, captures broader image context. In the complex background of power scenes, it can accurately locate the target across multiple interference areas, correlate features of the phone with those of the surrounding environment, and accurately locate the target. GELUs introduce nonlinearity, leveraging the characteristics of the Gaussian distribution. Compared to traditional activation functions, they adjust feature responses more smoothly, improving the discrimination of phone features under complex lighting conditions and occlusion. The two branches are fused through point-by-point multiplication. Unlike simple addition, point-by-point multiplication dynamically adjusts the output based on the strength of branch features. The unique features extracted by the large core and activation path are multiplied with the original features to highlight effective features, suppress interference, and strengthen features critical for phone recognition. The final 1×1 convolution further adjusts the channels, integrates the fused features, and outputs a feature map adapted for the subsequent YOLOv12 network, completing the entire module's feature processing flow and achieving a complete transformation from input to optimized feature output.
[0107] The LK-GLU module accurately addresses the pain points of detection in power production scenarios by leveraging the collaborative design of large-kernel convolution global association, cross-layer feature modulation, and adaptive nonlinear activation: the large receptive field (7×7) spans complex background elements such as equipment and cables, and by associating the spatial features of the mobile phone with distant reference objects (such as markings / equipment), the difference information between the target and the background is enhanced in the feature multiplication stage; GELU activation dynamically regulates feature responses based on Gaussian distribution, suppressing overexposure artifacts and retaining the metal details of the mobile phone under strong light, and enhancing effective signals in weak light. Combined with the adaptive capabilities of the 1×1 convolution channel, it ensures recognition robustness under drastic light fluctuations; to address the problem of tool occlusion, the jump connection retains the original high-resolution features, and after being fused with the global semantics of the large-kernel path, the occluded area (such as the outline of a mobile phone half-covered by a helmet) is inferred through multiplication gated reasoning, significantly improving local visibility and enhancing detection effects in power production environments.
[0108] In summary, the network architecture optimized based on YOLOv12 in this paper solves the limitations of the above-mentioned traditional detection methods and the key problems such as insufficient accuracy and precision in detecting mobile phone targets in power production environments. The specific experimental results and analysis are as follows:
[0109] (1) Comparative experimental results and analysis:
[0110] As shown in Table 1, the original YOLOv12 model algorithm has an average accuracy of 89.92%, a detection precision of 89.86%, and a recall rate of 85.35%. The optimized detection algorithm has a recall rate that is 5.88% higher than the original YOLOv12 algorithm, and a detection precision that is 1.16% higher than the original YOLOv12 algorithm. In addition, by observing the average detection precision (mAP) indicator value, it can be seen that the optimized detection algorithm has also improved by 3.43% on the original basis. Therefore, the improved algorithm of the present invention is generally superior to the original YOLOv12 model algorithm and will have more advantages in practical engineering applications.
[0111] Table 1 Experimental comparison before and after network optimization In addition, the performance of the original YOLOv12 model and the improved algorithm proposed in this invention in terms of accuracy and mAP training process. By comparing the average precision of the YOLOv12 model before and after optimization, as well as the comparison curves of accuracy and bounding box loss, the performance difference between the two models during training is intuitively demonstrated. Figure 7 、 8 , as shown in 9.
[0112] It can be clearly seen from the above comparison curve graph that the improved algorithm proposed in the present invention outperforms the original YOLOv12 model in both accuracy and mAP. Throughout the training cycle, the accuracy level of the improved algorithm is significantly higher than that of the original algorithm. Similarly, in terms of the mAP indicator, the optimized algorithm has consistently shown more significant performance throughout the training cycle. At the same time, it can be seen from the figure that the Loss curve of the improved algorithm is smoother than that of the original algorithm, which means that the improved YOLOv12 algorithm has better convergence properties during training. This comparison result fully verifies the effectiveness of the YOLOv12-based optimization algorithm of the present invention in mobile phone detection tasks.
[0113] (2) Ablation experiment and analysis:
[0114] Ablation experiments systematically validate the effectiveness of each core component in the proposed model and clarify their contribution to the final results. To comprehensively evaluate the proposed improvement strategy, we conducted ablation experiments using the modified YOLOv12 model on a test set for mobile phone detection in a power plant environment. The experimental results are shown in Table 2. The first improvement replaces C3k2 with LK-GLU. This addition enables multi-scale feature capture and enhances the model's multi-scale object detection capabilities. In the second improvement, the backbone of the baseline model is replaced from A2C2F to MLA-C2F. Results show that this replacement improves performance across all metrics, increasing mean average precision (mAP50) by 0.97%, demonstrating that MLA-C2F enhances the model's ability to extract object features. In the third improvement, the proposed method replaces the original Upsample+Concat model with the MCAF fusion module. The results show significant improvements across all metrics, including a 2.21% increase in mean average precision and a 3.56% increase in recall, demonstrating its effectiveness in improving detection accuracy.
[0115] Table 2 Effectiveness experiment of improved modules
[0116] (3) Visual analysis:
[0117] The improved YOLOv12 algorithm demonstrates strong performance on benchmark datasets, but its practical application requires in-depth evaluation using visualization techniques. Visual analysis can intuitively demonstrate the model's performance characteristics across various tasks and scenarios, including: 1) differences in recognition capabilities for different object categories; 2) the specific distribution of false positives and missed positives; and 3) the contribution of each improved module to real-world tasks. This analytical approach provides an intuitive basis for verifying the effectiveness of model improvements.
[0118] It should be understood that, in the description of this specification, reference to terms such as "one embodiment," "another embodiment," "other embodiments," or "first to Nth embodiments" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples.
[0119] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.
[0120] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization, characterized in that: The method comprises the following steps: S1: Construct an improved YOLOv12 architecture; wherein the improved YOLOv12 architecture includes: a feature extraction layer configured with an MLA-C2F module, a feature fusion layer configured with an MLA-C2F module, an MCAF module, and an LK-GLU module, and a prediction output layer having three core detection modules; In the feature extraction layer and feature fusion layer: the MLA-C2F module is configured to use the multi-branch key feature design of the local attention mechanism to represent the query constructing adaptive cross-scale response and model the interaction behavior between the character and the target; In the feature fusion layer: the MCAF module is configured to use multi-channel attention fusion to complementarily fuse shallow features and deep features, and the LK-GLU module is configured to dynamically adjust the fusion feature output of the correlation features between the mobile phone and the surrounding environment and the original features using channel segmentation design; The LK-GLU module specifically includes: The head Conv convolution layer is configured to adjust the channel dimension and compress the number of channels of the input feature map, and perform linear transformation on the output features to extract basic features; The initial feature transfer branch is configured to directly transfer the features output by the head Conv convolutional layer; The initial feature processing branch is configured to use a large kernel convolution layer and a Gaussian error linear unit to extract features from the output of the head Conv convolution layer and dynamically adjust the phone to the surrounding environment; The tail Conv convolution layer is configured to perform point-by-point multiplication of the features of the initial feature transfer branch and the initial feature processing branch and perform channel adjustment and feature fusion on the processed features; S2: Collect several scene images of illegal mobile phone use actions and normal production operation actions in the target production environment, annotate the scene images with illegal mobile phone use actions, convert the annotated images into training annotation data for YOLOv12, and construct a training sample set; S3: Use the training sample set to train the constructed improved YOLOv12 architecture to obtain a trained production environment mobile phone misuse detection model; S4: Inputting the image to be detected into the production environment mobile phone illegal use detection model to obtain the production environment mobile phone illegal use detection result output by the production environment mobile phone illegal use detection model.
2. The method for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization according to claim 1, characterized in that: The feature extraction layer specifically includes: a first Conv convolution layer, a second Conv convolution layer, a first C3k2 feature extraction module, a third Conv convolution layer, a second C3k2 feature extraction module, a fourth Conv convolution layer, a first MLA-C2f module, a fifth Conv convolution layer and a second MLA-C2f module connected in sequence.
3. The method for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization according to claim 2, characterized in that: The feature fusion layer specifically includes: a first MCAF module, a second MCAF module, a third MLA-C2f module, a sixth Conv convolution layer, a third MCAF module, a fourth MLA-C2f module, a seventh Conv convolution layer, a fourth MCAF module and an LK-GLU module connected in sequence.
4. The method for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization according to claim 3, characterized in that: The output ends of the first MLA-C2f module and the second MLA-C2f module are connected to the input end of the first MCAF module, and the output end of the second C3k2 feature extraction module is connected to the input end of the second MCAF module; the output end of the first MCAF module is connected to the input end of the fourth MCAF module, and the output end of the second MCAF module is connected to the input end of the third MCAF module.
5. The method for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization according to claim 4, characterized in that: The feature fusion layer specifically includes: a first Detect core detection module connected to the output end of the third MLA-C2f module, a second Detect core detection module connected to the output end of the fourth MLA-C2f module, and a third Detect core detection module connected to the output end of the LK-GLU module.
6. The method for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization according to claim 1, characterized in that: The MLA-C2F module specifically includes: The key branch is configured to perform direct normalization, normalization after extraction with a 5×5 standard convolution kernel, and normalization after extraction with a 7×7 standard convolution kernel on the input and divided features, modeling different levels of semantic relationships between the interaction between the character and the target; The query branch is configured to perform dot product calculations on the input and divided features with the three output features of the key branch, and perform weighted fusion of the three dot product calculation results; The Value branch is configured to perform attention interaction between the input and divided features and the weighted fusion features of the Query branch and output them.
7. The method for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization according to claim 1, characterized in that: The MCAF module specifically includes: A sampling unit is configured to downsample shallow features in the initial features, upsample deep features in the initial features, and cross-join the upsampled results with the downsampled results; A feature extractor configured to extract features output by the sampling unit respectively; The dual-path attention module is configured to perceive and enhance the features extracted by the feature extractor, and interacts with the initial features using a point-by-point multiplication method; The normalization layer is configured to normalize the information interaction of the two-way attention module and then output it.
8. The method for detecting illegal use of mobile phones in a production environment based on YOLOv12 optimization according to claim 7, characterized in that: The feature extractor is configured to adopt a convolutional multi-layer perceptron architecture in which the activation function is replaced by PReLu.
9. A production environment mobile phone illegal use detection system based on YOLOv12 optimization, characterized by: include: An architecture building module for building an improved YOLOv12 architecture, wherein the improved YOLOv12 architecture includes: a feature extraction layer configured with an MLA-C2F module, a feature fusion layer configured with an MLA-C2F module, an MCAF module, and an LK-GLU module, and a prediction output layer having three core detection modules; In the feature extraction layer and feature fusion layer: the MLA-C2F module is configured to use the multi-branch key feature design of the local attention mechanism to represent the query constructing adaptive cross-scale response and model the interaction behavior between the character and the target; In the feature fusion layer: the MCAF module is configured to use multi-channel attention fusion to complementarily fuse shallow features and deep features, and the LK-GLU module is configured to dynamically adjust the fusion feature output of the correlation features between the mobile phone and the surrounding environment and the original features using channel segmentation design; The LK-GLU module specifically includes: The head Conv convolution layer is configured to adjust the channel dimension and compress the number of channels of the input feature map, and perform linear transformation on the output features to extract basic features; The initial feature transfer branch is configured to directly transfer the features output by the head Conv convolutional layer; The initial feature processing branch is configured to use a large kernel convolution layer and a Gaussian error linear unit to extract features from the output of the head Conv convolution layer and dynamically adjust the phone to the surrounding environment; The tail Conv convolution layer is configured to perform point-by-point multiplication of the features of the initial feature transfer branch and the initial feature processing branch and perform channel adjustment and feature fusion on the processed features. The sample construction module is used to collect several scene images of illegal mobile phone use actions and normal production operation actions in the target production environment, annotate the scene images with illegal mobile phone use actions, convert the annotated images into training annotation data for YOLOv12, and construct a training sample set; The model training module is used to train the constructed improved YOLOv12 architecture using the training sample set to obtain a trained production environment mobile phone misuse detection model; The model detection module is used to input the image to be detected into the production environment mobile phone illegal use detection model and obtain the production environment mobile phone illegal use detection result output by the production environment mobile phone illegal use detection model.
Citation Information
Patent Citations
Power transmission line foreign matter identification method and system based on improved YOLOv10 model
CN119559584A
Underwater target detection system and method based on RT-DETR improvement
CN119851108A