Intelligent detection method for illegal wearing of high-altitude power operation safety equipment

By combining a hybrid architecture of CNN and Transformer, a CAS-MTL model was designed to solve the problem of detecting illegal wearing of speed differential self-locking devices during high-altitude operations, achieving high-accuracy intelligent detection. It is suitable for high-risk high-altitude operation scenarios such as power inspection and tower maintenance.

CN120635938APending Publication Date: 2025-09-12ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510611740.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively detect illegal wearing of speed differential self-locking devices in high-altitude work scenarios, especially behaviors such as not being fastened, loose, and hanging low and using high. Traditional convolutional neural networks have limitations in illegal reasoning and cannot effectively model long-range dependencies and global semantic interactions. In addition, difficulties in data collection lead to poor model training results.

Method used

A hybrid architecture based on multi-task learning and spatial semantic guidance is adopted, combining CNN and Transformer. Through a dual-branch collaborative working mechanism, a CAS-MTL model is designed. The multi-head cross-attention module and adaptive weight distributor are used to achieve accurate violation detection of speed difference self-locking devices.

Benefits of technology

The accuracy rate of detecting illegal wearing of speed differential self-locking devices during high-altitude operations reached 94.1%, which is better than existing methods, overcomes the problems of high false alarm rate and low accuracy, and provides a real-time and standardized safety monitoring solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635938A_ABST
    Figure CN120635938A_ABST
Patent Text Reader

Abstract

An intelligent detection method for illegal wearing of high-altitude electric power operation safety equipment comprises the following steps: step 1, collecting high-altitude operation field data image samples, and making a data set which can be used for model training; step 2, constructing a joint detection framework CAS-MTL for detecting non-standard wearing of the high-altitude operation speed difference self-locking device, including three typical illegal behaviors of untying, loosening and low hanging and high use; and step 3, carrying out end-to-end training and deployment application on the CAS-MTL model. According to the method, a hybrid architecture fusing CNN and Transform is innovatively proposed, accurate violation detection of the speed difference self-locking device is realized through a double-branch cooperative work mechanism, and the bidirectional feature enhancement mechanism effectively overcomes the representation limitation of a traditional single-branch classification network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention lies at the intersection of aerial work safety monitoring and computer vision, specifically a method for intelligently detecting the wearing of safety equipment for aerial work based on multi-task learning and a hybrid neural network architecture. This method is particularly suitable for detecting the wearing of speed differential locks (including three typical violations: unfastened, loose, and used at high altitude) in high-risk aerial work scenarios such as power inspections and tower maintenance. It can also be expanded to intelligent compliance checks for safety belts, safety ropes, and other aerial protective equipment. Background Art

[0002] As a critical process in industries such as construction, energy, and communications, aerial work continues to face significant challenges in managing safety risks. The proper use of a speed differential self-locking device, a core piece of fall protection equipment, directly impacts the effectiveness of individual fall protection. Traditional manual inspections suffer from inherent flaws such as high subjectivity and low coverage density, making it difficult to achieve dynamic, all-weather monitoring. Intelligent monitoring technology based on computer vision, by constructing a multi-scale feature fusion network, can accurately identify abnormalities in the wearing state of safety equipment, providing a real-time, standardized technical approach for the safety supervision of aerial work.

[0003] Intelligent detection of violations during aerial work faces significant technical challenges, primarily due to three constraints: First, aerial work scenarios are characterized by complex backgrounds and the variability of safety equipment. Second, the limited working environment makes effective data collection difficult, resulting in a scarcity of training samples. More importantly, violation determination requires a deep understanding of the semantic relationships between human and machine interactions, placing higher demands on the cognitive reasoning capabilities of the detection model. Traditional convolutional neural networks (CNNs), with their local perception properties, can effectively extract the outlines of workers and the morphological features of safety equipment. However, they have significant limitations in violation reasoning tasks: the limited receptive field of a single-layer convolutional kernel makes it difficult to model long-range dependencies, and the layered hierarchical structure lacks an explicit global semantic interaction mechanism, resulting in insufficient contextual understanding. Aerial work scenarios are subject to interference such as occlusion by power grid equipment and sudden changes in lighting. Traditional computer vision methods rely on extracting superficial semantic features and fail to establish cross-layer contextual semantic associations, leading to significant feature degradation. Models such as YOLO and Mask R-CNN are designed for rigid targets such as helmets and seat belts. However, the rope of a speed differential self-locking device has flexible deformation characteristics. Its illegal status (such as hanging low and using high, and slack) requires a joint analysis of the rope winding trajectory and human posture. Existing models lack the ability to model geometric deformation, resulting in the failure of fine-grained detection. Summary of the Invention

[0004] In order to overcome the shortcomings of the existing technical solutions, the present invention proposes an intelligent detection method for illegal wearing of high-altitude power operation safety equipment based on multi-task learning and spatial semantic guidance. It innovatively proposes a hybrid architecture that integrates CNN and Transformer, and realizes accurate illegal detection of speed differential self-locking devices through a dual-branch collaborative working mechanism. This bidirectional feature enhancement mechanism effectively overcomes the representation limitations of traditional single-branch classification networks.

[0005] The present invention solves the technical problem by adopting the following technical solutions:

[0006] An intelligent detection method for illegal wearing of safety equipment for high-altitude power work, comprising the following steps:

[0007] Step 1: Collect image samples of aerial work site data to create a dataset that can be used for model training;

[0008] Step 2: A joint detection framework, CAS-MTL (Classification with Auxiliary Segmentation via Multi-Task Learning), was constructed to detect improper wearing of speed differential self-locking devices during aerial work.

[0009] Step 3: Perform end-to-end training and deployment of the CAS-MTL model.

[0010] Furthermore, the process of step 1 is as follows:

[0011] (1.1) Collect data from the power grid construction site, set the positive and negative sample ratios and the dataset size;

[0012] (1.2) Label the collected work site image samples, including labels for semantic segmentation and image classification;

[0013] (1.3) In order to improve the generalization ability of the model, data augmentation processing is performed on the dataset samples, and measures taken include color jittering and geometric transformation.

[0014] Furthermore, the process of step 2 is as follows:

[0015] (2.1) A multi-task joint detection framework, CAS-MTL, is proposed. By building a multi-task collaborative optimization mechanism, it innovatively solves the problem of detecting illegal wearing of speed differential self-locking devices in high-altitude work scenarios. A dual-path feature encoder is designed, deformable convolution is used to capture device deformation characteristics, and a multi-head cross-attention module is introduced to enhance inter-task interaction. The spatial correlation between segmentation masks and classification features is established in the channel dimension, achieving collaborative optimization of pixel-level supervision of semantic segmentation and high-level semantic representation of classification tasks.

[0016] (2.2) We propose differentiated optimization strategies to address the heterogeneous feature requirements of segmentation and classification tasks. For segmentation, we enhance spatial detail and deformation feature modeling and suppress background noise through multi-scale feature fusion, SE channel attention mechanism, and Deformable Convolution v4 (DCNv4). For classification, we use a pyramid feature refinement process combined with SE modules and residual connections to achieve semantic abstraction and hierarchical representation. We also enhance discriminative feature extraction through spatial compression and cross-layer splicing. These two strategies optimize feature fusion for target geometric adaptability and semantic abstraction requirements, respectively.

[0017] (2.3) Design a cross-task feature guidance module, SEAB (Segmentation Enhancement Attention Block). This module takes the high-resolution spatial features of the semantic segmentation branch as input and generates a spatial attention weight map through a three-level spatial attention mechanism. This weight map is then Hadamard-multiplied with the global features of the classification branch to achieve pixel-level spatial attention adjustment, allowing the classification task to dynamically focus on the key areas of the semantic segmentation markers. Through this hierarchical cross-task feature interaction, SEAB effectively establishes an attention distillation pipeline from the dense prediction task to the global classification task.

[0018] (2.4) By integrating heteroscedastic uncertainty theory with dynamic parameter scaling, we construct an adaptive weight allocator with dual adjustment capabilities. We also construct an adaptive multi-task loss function that includes segmentation loss, classification loss, and regularization terms. This function is automatically optimized to avoid task conflicts. The mathematical expression of the loss function is defined as follows:

[0019]

[0020] In the formula Represents the i-th basic loss (including Dice Loss, Focal Loss and CrossEntropyLoss), α i is the dynamic scaling factor, σ i is the uncertainty weight parameter;

[0021] The segmentation loss uses Dice Loss to enhance the geometric sensitivity of small target segmentation, combined with FocalLoss to alleviate the imbalance of pixels between classes, and jointly optimizes the segmentation task through parameter sharing; the classification loss is optimized using CrossEntropy Loss.

[0022] Furthermore, the process of step 2 also includes:

[0023] (2.5) To verify the effectiveness of the CAS-MTL architecture, a systematic ablation experiment was designed: a modular removal strategy was used to quantitatively analyze the contribution of the multi-task collaboration mechanism, the SEAB module, and the multi-task adaptive loss function. Comparative experiments were conducted with mainstream detection frameworks such as ViT and YOLOv8. The results showed that CAS-MTL achieved the best detection accuracy of 94.1% on the power grid high-altitude operation dataset.

[0024] Furthermore, in (2.1), CAS-MTL uses a hierarchical feature extractor to generate a multi-scale feature pyramid. In the feature optimization stage, the semantic segmentation branch uses an adaptive weighted feature fusion mechanism and deformable convolution to enhance the representation of spatial details, while the classification branch uses channel attention to filter discriminative features. In the feature interaction stage, the two branches implement cross-task knowledge transfer through a cross-task cross-attention module: the high-resolution feature map of the segmentation branch dynamically guides the classification attention distribution, and the global semantic information extracted by the classification branch inversely optimizes the segmentation boundary. Finally, the semantic segmentation head outputs a pixel-level safety equipment positioning mask, and the classification head generates a violation conditional probability matrix based on spatial semantic associations.

[0025] The process of step 3 is as follows:

[0026] (3.1) Use the PyTorch framework to train the model and obtain the training weights of CAS-MTL;

[0027] (3.2) Encapsulate it as a server-side application to facilitate the inference service of the request model.

[0028] The present invention has the following beneficial effects:

[0029] (1) A method for detecting illegal wearing of safety equipment for high-altitude power operations based on multi-task learning and spatial semantic guidance is provided. The method can effectively detect three types of irregular wearing of speed differential self-locking device protective devices (unfastened, loose, and low hanging and high use) with an accuracy rate of 94.1%, which is better than the existing detection methods.

[0030] (2) Overcome the shortcomings of existing technologies such as low accuracy and high false alarm rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a structural diagram of the CAS-MTL model of the multi-task joint detection framework of the present invention.

[0032] Figure 2 It is the semantic segmentation branch feature fusion optimization network diagram used in the present invention.

[0033] Figure 3 It is the image classification branch feature fusion optimization network diagram used in the present invention.

[0034] Figure 4 1 is a block diagram of the SEAB of the present invention. DETAILED DESCRIPTION

[0035] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0036] Reference Figures 1 to 4 , an intelligent detection method for illegal wearing of safety equipment for high-altitude power work, comprising the following steps:

[0037] Step 1: Collect image samples of aerial work site data to create a dataset that can be used for model training;

[0038] The process of step 1 is as follows:

[0039] (1.1) Data was collected from power grid construction sites. Images of high-altitude operations where the speed differential self-locking device was improperly worn (not tied, loose, or hung low and used high) were used as negative samples, and images of high-altitude operations where the speed differential self-locking device was properly worn (normal) were used as positive samples. The number of samples was 1,186, and the positive-negative sample ratio was maintained at 2:1.

[0040] (1.2) Label the collected work site image samples, including pixel-level labeling for semantic segmentation, with the labeled objects being the workers and the speed differential self-locking device of the protective device; and scene judgment semantic labels (normal, untied, loose, and low hanging high use) for image classification.

[0041] (1.3) A multi-level data augmentation strategy is employed to improve model generalization. In the color space dimension, image brightness (Δ = 0.2), contrast (scaling factor [0.5, 1.5]), saturation (scaling factor [0.5, 1.5]), and hue (offset by ±0.1) are randomly perturbed, and channel permutation is introduced to effectively simulate the complex lighting variations and device color differences found in real scenes. At the geometric transformation level, random horizontal and vertical flips (with probability p = 0.5) are implemented to enhance spatially invariant feature learning. To address input scale differences, a dynamic scaling strategy using bilinear interpolation to maintain aspect ratio is employed, supplemented by zero padding to a target size of 512 x 512. Finally, data distribution is normalized using ImageNet statistical parameters (mean: [0.485, 0.456, 0.406], standard deviation: [0.229, 0.224, 0.225]) to ensure network training stability.

[0042] Step 2: In light of the differences and difficulties between speed differential lock detection and conventional safety belt and helmet detection, we combined the advantages of semantic segmentation for accurate feature extraction and used a multi-task learning approach to construct a joint detection framework, CAS-MTL, to detect improper wearing of speed differential locks during aerial work.

[0043] The process of step 2 is as follows:

[0044] (2.1) The intelligent detection framework CAS-MTL of the speed differential self-locking device constructed by the present invention (architecture as shown in FIG. Figure 1 (as shown in Figure 2) uses a hierarchical feature extractor to generate a multi-scale feature pyramid. During the feature optimization phase, the semantic segmentation branch employs an adaptive weighted feature fusion mechanism and deformable convolution to enhance the representation of spatial details, while the classification branch uses channel attention to filter discriminative features. During the feature interaction phase, the two branches implement cross-task knowledge transfer via a cross-task cross-attention module: the segmentation branch's high-resolution feature map dynamically guides the classification attention distribution, while the classification branch extracts global semantic information that inversely optimizes the segmentation boundaries. Ultimately, the semantic segmentation head outputs a pixel-level safety equipment location mask, and the classification head generates a violation conditional probability matrix based on spatial semantic associations.

[0045] (2.2) In response to the heterogeneous feature requirements of segmentation and classification tasks, this study designs a differentiated feature fusion optimization strategy. Segmentation feature optimization focuses on spatial detail preservation and adopts a multi-scale feature fusion architecture: unify the multi-scale feature resolution through bilinear interpolation, and use 1×1 convolution for channel normalization; introduce SE channel attention mechanism to achieve feature recalibration and suppress background noise interference; construct learnable weight parameters to achieve adaptive feature fusion. In response to the problem of insufficient geometric adaptability of traditional standard convolution in feature extraction of complex deformable targets (such as the non-rigid structure of the speed differential self-locking device), the hybrid deformable convolution module DCNv4 is connected to enhance the deformation feature modeling of safety equipment through dynamic offset learning. The network structure is as follows: Figure 2 As shown, the process can be expressed as:

[0046]

[0047] F seg =DCNv4(Conv 3*3 (F weight )) (2-2)In the formula, “↑” represents upsampling.

[0048] Classification feature optimization focuses on semantic abstraction construction and adopts a pyramid feature refinement process: 3×3 convolution is used to achieve bottom-up feature transfer and spatial compression (1 / 2→1 / 8). After cross-layer feature splicing, the SE module is used to filter the discriminative channels and superimpose residual connections to finally form a hierarchical semantic representation. The network structure is as follows: Figure 3 As shown, the process can be expressed as:

[0049]

[0050] In the formula Represents the feature map of the kth scale, and “↓” represents downsampling.

[0051] (2.3) To achieve a balance between multi-task feature interaction and independence, a two-stage cross-attention mechanism is used to achieve task collaboration. This mechanism achieves knowledge sharing through the cascade structure of the feature interaction module (Task Interaction Block) and the task query module (TaskQuery Block):

[0052] In the feature interaction stage, the segmentation feature map and classification feature maps Flattened into sequence features and Compute cross-task correlations through a shared multi-head attention mechanism:

[0053] A fusion =MultiHead(Q=S seg ,K=S cls ,V=S cls ) (2-4) where attention output The global semantic association is encoded, and the gate unit generates the channel weight matrix Γ∈[0,1] through the Sigmoid activated linear layer B×(HW)×C , to achieve dynamic feature selection:

[0054] A′ fusion =Γ⊙A fusion (2-5)

[0055] In the task query phase, each task is obtained from A′ through an independent multi-head attention module. fusion Extract task-related features and use residual connections to preserve the integrity of the original features:

[0056]

[0057] This architecture uses a shared-specific dual-stream design to promote task knowledge transfer while maintaining feature independence, effectively alleviating the feature interference problem in multi-task learning.

[0058] In the semantic segmentation guidance enhancement stage, in order to enhance the spatial perception ability of the classification task, this paper proposes a semantic guidance attention module SEAB, whose structure is as follows Figure 4 This module uses the spatial prior of the segmentation task to guide the classification feature focus. First, the segmentation feature is bilinearly interpolated. Adjusted to categorical features The same space size, we get After two levels of convolutional networks (3×3 convolution reduces the dimension to 64 channels → 1×1 convolution compresses to a single channel), the spatial attention map M is generated with the Sigmoid activation function. attn ∈[0,1]B×1×h×w , M attn Multiply element-wise with the classification features to highlight the discriminative regional responses:

[0059] M attn =σ(Conv 1×1 (ReLU(Conv 3×3 (F′ seg )))) (2-8)

[0060] F′ cls =F cls ⊙M attn (2-9)

[0061] (2.4) The difficulties faced in detecting illegal wearing of speed differential self-locking devices in high-altitude work scenarios include the irregular geometry of the area where the speed differential self-locking device interacts with the human body; the pixel-level classification imbalance between the background human body and the target device; and the susceptibility of classification tasks to local occlusion. By integrating heteroscedastic uncertainty theory and dynamic parameter scaling technology, an adaptive weight allocator with dual adjustment capabilities is constructed. The loss function is defined as follows:

[0062]

[0063] In the formula Represents the i-th basic loss (including Dice Loss, Focal Loss and CrossEntropyLoss), α i is the dynamic scaling factor, σ i is the uncertainty weight parameter.

[0064] The segmentation loss uses Dice Loss to enhance the geometric sensitivity of small target segmentation, combined with FocalLoss to alleviate the imbalance of pixels between classes, and jointly optimizes the segmentation task through parameter sharing; the classification loss is optimized using CrossEntropy Loss. In order to avoid imbalance between tasks, an uncertainty weighted adaptive weighting mechanism is used. The item automatically adjusts the weights of different losses. When a loss fluctuates greatly (σ i When it is large, its weight is automatically reduced; use logσ i Prevent σ i Infinitely increase to ensure numerical stability. Design a learnable scaling factor α i , dynamically calibrate different levels of loss (such as segmentation loss level 10) through gradient back propagation 1 With classification loss level 10 2 ). While ensuring the independent learning of each subtask, this optimization framework achieves dynamic balance of gradients between tasks through the differential characteristics of weight parameters, effectively alleviating the optimization instability problem of traditional fixed weight strategies in complex scenarios.

[0065] (2.5) To verify the effectiveness of the design of each module of the model, the contribution of different components is analyzed through systematic ablation experiments. All experiments are conducted under the same training settings.

[0066] Single-task baseline model (SingleTask-Base): To verify the necessity of the multi-task collaboration mechanism, by removing the semantic segmentation branch and simplifying the task interaction mechanism construction, the cross-task multi-head cross attention TCA is replaced with a self-attention mechanism, and the gate unit Gate and SEAB module are removed, retaining only the cross-entropy loss function of the classification task.

[0067] Dual-task model without SEAB (CAS-MTL w / o SEAB): restores the dual-path architecture but removes the semantic enhancement attention module, and adopts an adaptive multi-task loss function to decouple the independent impact of the SEAB module on the classification task performance.

[0068] Static Weighted Multi-Task Model (StaticWeight-MTL): In order to compare the advantages and disadvantages of adaptive loss and static weighting strategy, the architecture is consistent with CAS-MTL, using a fixed weight multi-task loss

[0069]

[0070] Complete model (CAS-MTL): includes the complete dual-task branch, SEAB module and gating unit, and uses an adaptive multi-task loss function.

[0071] Table 1 shows the comparison of the Top-1 classification accuracy and semantic segmentation mIoU of each control group. The best result is bolded, ↑ means the larger the better, ↓ means the smaller the better;

[0072]

[0073] The results in Table 1 illustrate the effects of the dual-task path architecture, SEAB module, and adaptive loss function on accuracy and segmentation precision.

[0074] To validate the comprehensive performance of the CAS-MTL model, comparative experiments were conducted on a dataset of power grid aerial work, comparing it to eight mainstream models. All methods used the same training / test set split (8:2) and were replicated on an NVIDIA L20 GPU platform. Considering the sensitivity of ViT and Swin Transformer to positional encoding (using ImageNet pre-trained weights), their input dimensions were uniformly resized to 384×384; the remaining models maintained the original 512×512 resolution to balance computational overhead and feature detail. To ensure the integrity of the aerial scene context, data augmentation that destroys spatial structure, such as random cropping, was disabled in the experiment.

[0075] Table 2 shows the comparison of the results of CAS-MTL and other detection models on the power grid high-altitude operation dataset;

[0076]

[0077] The experimental results in Table 2 show that CAS-MTL achieves the best results in both Top-1 accuracy and F1 score indicators.

[0078] Step 3: Perform end-to-end training and deployment of the CAS-MTL model;

[0079] The process for step 3 is as follows:

[0080] (3.1) The model was trained using the annotated high-altitude work dataset. This model was implemented in PyTorch 2.0 and uses the HRNet-W48 feature extraction network. Its multi-scale feature fusion capabilities are suitable for high-resolution image analysis. The network outputs four layers of feature maps (with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image, respectively), which were fed into the CAS-MTL multi-task learning framework. The model training platform used the deep learning framework Torch 2.5.1, CUDA version 12.4, Ubuntu 22.04, an Intel Xeon Platinum 8457C CPU, and an NVIDIA L20 GPU (48GB). The input image size was 512×512, and the training iterations were 10,000. The batch size was set to 2. The Adam optimizer was used with a learning rate of 1e-5. The model was initialized using pre-trained weights from ImageNet.

[0081] (3.2) Encapsulate the model file (.pt or .pth format) obtained through (3.1) into an interface. Use the Flask framework to encapsulate the trained model as a RESTful API service, define input / output data specifications (such as JSON format), and generate interactive documentation through Swagger. Use Docker technology to build a lightweight container image, integrate the model dependency environment (Python library, CUDA version, etc.), and achieve cross-platform one-click deployment. Configure load balancing through Nginx reverse proxy on the cloud server side, and use Gunicorn as the WSGI server to host the Flask application. Provide encrypted API access services based on the HTTPS protocol to the outside world, and use SSL / TLS certificates to ensure data transmission security.

[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent detection method for illegal wearing of safety equipment for high-altitude power work, characterized in that: The following steps are involved: Step 1: Collect image samples of aerial work site data to create a dataset that can be used for model training; Step 2: A joint detection framework, CAS-MTL, was constructed to detect improper wearing of speed differential locks during aerial work, including three typical violations: not fastened, loose, and hanging low and using high. Step 3: Perform end-to-end training and deployment of the CAS-MTL model.

2. The intelligent detection method for illegal wearing of safety equipment for high-altitude power work according to claim 1, characterized in that: The process of step 2 is as follows: (2.1) We propose a multi-task joint detection framework, CAS-MTL. By building a multi-task collaborative optimization mechanism, designing a dual-path feature encoder, using deformable convolution to capture device deformation features, introducing a multi-head cross-attention module to enhance inter-task interaction, and establishing spatial correlation between segmentation masks and classification features in the channel dimension, we achieve collaborative optimization of pixel-level supervision for semantic segmentation and high-level semantic representation for classification tasks. (2.2) We propose differentiated optimization strategies for the heterogeneous feature requirements of segmentation and classification tasks. In terms of segmentation, we use multi-scale feature fusion, SE channel attention mechanism, and deformable convolution DCNv4 to enhance spatial details and deformation feature modeling and suppress background noise. For classification, a pyramid feature refinement process is used in conjunction with the SE module and residual connections to achieve semantic abstraction and hierarchical representation. Discriminative feature extraction is enhanced through spatial compression and cross-layer splicing. These two strategies optimize feature fusion based on target geometric adaptability and semantic abstraction requirements, respectively. (2.3) We designed a cross-task feature guidance module (SEAB). This module takes the high-resolution spatial features of the semantic segmentation branch as input and generates a spatial attention weight map through a three-level spatial attention mechanism. This weight map is then Hadamard-multiplied with the global features of the classification branch to achieve pixel-level spatial attention adjustment, enabling the classification task to dynamically focus on the key areas identified by semantic segmentation. Through this hierarchical cross-task feature interaction, SEAB effectively establishes an attention distillation channel from the dense prediction task to the global classification task. (2.4) By integrating heteroscedastic uncertainty theory with dynamic parameter scaling, we construct an adaptive weight allocator with dual adjustment capabilities. We also construct an adaptive multi-task loss function that includes segmentation loss, classification loss, and regularization terms. This function is automatically optimized to avoid task conflicts. The mathematical expression of the loss function is defined as follows: In the formula represents the i-th basic loss, α i is the dynamic scaling factor, σ i is the uncertainty weight parameter; The segmentation loss uses Dice Loss to enhance the geometric sensitivity of small target segmentation, combined with Focal Loss to alleviate the imbalance of pixels between classes, and jointly optimizes the segmentation task through parameter sharing; The classification loss is optimized using CrossEntropy Loss.

3. The intelligent detection method for illegal wearing of safety equipment for high-altitude power work according to claim 2, characterized in that: The process of step 2 further includes: (2.5) To verify the effectiveness of the CAS-MTL architecture, a systematic ablation experiment is designed: the contribution of the multi-task collaboration mechanism, the SEAB module, and the multi-task adaptive loss function is quantitatively analyzed through a modular removal strategy, and a comparative experiment is conducted with mainstream detection frameworks.

4. A smart detection method for illegal wearing of safety equipment for high-altitude power work according to claim 2 or 3, characterized in that: In (2.1), CAS-MTL uses a hierarchical feature extractor to generate a multi-scale feature pyramid; In the feature optimization stage, the semantic segmentation branch uses an adaptive weighted feature fusion mechanism and deformable convolution to enhance the representation of spatial details, while the classification branch uses channel attention to filter discriminative features. In the feature interaction stage, the two branches achieve cross-task knowledge transfer through a cross-task cross-attention module: the high-resolution feature map of the segmentation branch dynamically guides the classification attention distribution, and the global semantic information extracted by the classification branch reversely optimizes the segmentation boundary; finally, the semantic segmentation head outputs a pixel-level safety equipment positioning mask, and the classification head generates a violation conditional probability matrix based on spatial semantic associations.