Multi-target cooperative identification method and system for safety violation behaviors in unmanned aerial vehicle construction area
Patent Information
- Application Number
- CN202610985026.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-22
AI Technical Summary
[0005]本发明针对现有施工区域安全违规行为识别方法存在的小目标检测精度低、不安全特征缺乏量化标准、多类别违规行为难以协同识别以及缺乏量化评估机制等问题,提供一种面向无人机边缘计算的施工区域安全违规行为多模态协同识别与量化评估方法及系统
[0024] A four-dimensional unsafe feature standard list, constructed using trajectory intersection theory and the word frequency-inverse document frequency algorithm, transforms construction specification texts into quantifiable visual judgment criteria, solving the problem of a lack of unified standards for unsafe features. Fourier Merlin transform is employed for frequency domain registration and correction, combined with a mosaic-hybrid online enhancement strategy, effectively eliminating drone perspective distortion and enhancing the model's robustness to small and occluded targets. A dual-path lightweight architecture combined with deep reparameterizable convolution maintains multi-branch residual topology during training to improve feature extraction capabilities, while converting it to a single-path linear topology during inference to reduce computational load, balancing recognition accuracy with the real-time requirements of edge computing platforms. The fusion of an adversarial transfer learning framework and a channel attention mechanism enables collaborative recognition of binary classification qualitative discrimination and multi-class target localization, improving the detection accuracy of multi-level safety violations. The combination of the risk matrix method and an ordered logistic regression model transforms violation identification results into an actionable severity index and rectification priority, providing a quantitative decision-making basis for construction site safety management.
Smart Images

Figure CN122799313A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent inspection and construction safety management technology of unmanned aerial vehicles (UAVs), specifically involving a multi-target collaborative identification method and system for safety violations in construction areas using UAVs. Background Technology
[0002] In the fields of housing construction and municipal engineering, safety violations in construction areas are one of the main causes of personal injury and property damage. Violations such as not wearing safety belts, not wearing safety helmets, improper crane positioning, and lack of support in foundation pits are frequent, seriously threatening the lives of personnel on construction sites. Traditional safety supervision methods mainly rely on manual inspections, which have problems such as limited coverage, low inspection efficiency, and strong subjectivity, making it difficult to achieve comprehensive real-time monitoring of the construction area.
[0003] In recent years, drone technology has been increasingly applied to safety supervision at construction sites due to its advantages such as maneuverability, wide field of view, and low cost. However, existing methods for identifying safety violations based on drone video still have the following shortcomings: First, drone aerial images suffer from small target scale, complex backgrounds, and perspective distortion, leading to low detection accuracy and high false negative rates for small targets. Second, unsafe features in construction safety regulations have not yet been structured into a quantifiable list, lacking a unified visual judgment basis. Third, existing methods mostly target single violations, making it difficult to achieve collaborative detection of multiple categories of safety violations. Fourth, identified violations lack quantitative severity assessments and rectification priority rankings, failing to provide effective decision support for on-site management.
[0004] Therefore, how to utilize the edge computing platform of drones to integrate multimodal visual information and achieve high-precision collaborative identification and quantitative assessment of safety violations in construction areas is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0005] This invention addresses the problems of low accuracy in small target detection, lack of quantitative standards for unsafe features, difficulty in collaborative identification of multiple types of violations, and lack of quantitative evaluation mechanisms in existing methods for identifying safety violations in construction areas. It provides a multimodal collaborative identification and quantitative evaluation method and system for safety violations in construction areas based on UAV edge computing.
[0006] In a first aspect, embodiments of this application provide a multi-target collaborative identification method for safety violations in unmanned aerial vehicle (UAV) construction areas, the method comprising:
[0007] Based on trajectory intersection theory and word frequency-inverse document frequency algorithm to analyze construction specifications, a four-dimensional list of unsafe features is constructed.
[0008] Based on the aforementioned list, multi-source images were collected, and after Fourier Merlin transform correction and mosaic-mixing enhancement, a scale-invariant training dataset was constructed.
[0009] A dual-path lightweight architecture is constructed and trained using the dataset. The first path extracts dynamic regions of interest through semantic segmentation, and the second path outputs a target localization feature map using depthwise reparameterizable convolution and a weighted intersection-union loss function.
[0010] An adversarial transfer learning framework is constructed, which integrates the binary classification results with the feature map through a channel attention mechanism to output multi-level collaborative identification results of security violations.
[0011] A multidimensional risk vector is constructed using the identified violations as risk factors. The severity index and rectification priority are output by combining the risk matrix method and the ordered logistic regression model.
[0012] Secondly, embodiments of this application provide a multi-target collaborative identification system for safety violations in unmanned aerial vehicle (UAV) construction areas, applied to the method described in the first aspect, the system comprising:
[0013] The feature construction module is used to analyze construction specifications based on trajectory intersection theory and word frequency-inverse document frequency algorithm, and to construct a four-dimensional list of unsafe features.
[0014] The dataset construction module is used to collect multi-source images according to the list, and after Fourier Merlin transform correction and mosaic-mixing enhancement, construct a scale-invariant training dataset.
[0015] The dual-path recognition module is used to construct a lightweight dual-path architecture and train it using the dataset. The first path extracts dynamic regions of interest through semantic segmentation, and the second path outputs a target localization feature map using depthwise reparameterizable convolution and a weighted intersection-over-union loss function.
[0016] The collaborative identification module is used to construct an adversarial transfer learning framework. It integrates the binary classification results with the feature map through a channel attention mechanism and outputs multi-level collaborative identification results of security violations.
[0017] The quantitative assessment module is used to construct a multidimensional risk vector based on the identified violations as risk factors, and outputs a severity index and rectification priority by combining the risk matrix method and the ordered logistic regression model.
[0018] Thirdly, embodiments of this application provide an electronic device, including:
[0019] processor;
[0020] Memory used to store processor-executable instructions;
[0021] The processor is configured to implement the multi-target collaborative identification method for safety violations in UAV construction areas as described in the first aspect when executing the instructions.
[0022] Fourthly, embodiments of this application provide a computer-readable storage medium storing a program that instructs a device to perform the multi-target collaborative identification method for safety violations in UAV construction areas as described in the first aspect.
[0023] Compared with the prior art, the present invention has the following beneficial effects:
[0024] A four-dimensional unsafe feature standard list, constructed using trajectory intersection theory and the word frequency-inverse document frequency algorithm, transforms construction specification texts into quantifiable visual judgment criteria, solving the problem of a lack of unified standards for unsafe features. Fourier Merlin transform is employed for frequency domain registration and correction, combined with a mosaic-hybrid online enhancement strategy, effectively eliminating drone perspective distortion and enhancing the model's robustness to small and occluded targets. A dual-path lightweight architecture combined with deep reparameterizable convolution maintains multi-branch residual topology during training to improve feature extraction capabilities, while converting it to a single-path linear topology during inference to reduce computational load, balancing recognition accuracy with the real-time requirements of edge computing platforms. The fusion of an adversarial transfer learning framework and a channel attention mechanism enables collaborative recognition of binary classification qualitative discrimination and multi-class target localization, improving the detection accuracy of multi-level safety violations. The combination of the risk matrix method and an ordered logistic regression model transforms violation identification results into an actionable severity index and rectification priority, providing a quantitative decision-making basis for construction site safety management. Attached Figure Description
[0025] Figure 1 A schematic flowchart of a multi-target collaborative identification method for safety violations in construction areas using drones, provided as an embodiment of this application.
[0026] Figure 2 This is an architecture diagram of a multi-target collaborative identification system for safety violations in construction areas using drones, provided as an embodiment of this application.
[0027] Figure 3 A schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.
[0029] It should be noted that in the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.
[0030] Based on the embodiments described in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0031] Example 1
[0032] Figure 1 This is a schematic flowchart illustrating a multi-target collaborative identification method for safety violations in a construction area using unmanned aerial vehicles (UAVs), provided as an embodiment of this application. Figure 1 As shown, a multi-target collaborative identification method for safety violations in drone construction areas includes:
[0033] S1. Based on trajectory intersection theory and word frequency-inverse document frequency algorithm, analyze construction specifications and construct a four-dimensional unsafe feature standard list. This step first analyzes the construction safety-related specification texts, using accident causation theory to decompose the causes of safety accidents into human behavioral factors and object state factors. Then, through text analysis algorithms, extract frequently occurring unsafe feature words. These features are classified and organized according to four dimensions: personnel posture, equipment wearing, machinery occupancy, and edge support, forming a standardized feature list that can be recognized by computer vision systems, providing a basis for subsequent image acquisition and target recognition.
[0034] Specifically, in this embodiment, the specific method for constructing the four-dimensional list of insecure features includes:
[0035] S1.1 Based on the trajectory intersection theory, the causes of construction safety accidents are analyzed as unsafe human behavior and unsafe conditions of objects. Unsafe human behavior is mapped to personnel posture characteristics and equipment wearing characteristics, and unsafe conditions of objects are mapped to mechanical occupancy characteristics and edge support characteristics.
[0036] This step first introduces the trajectory intersection theory to analyze the causal mechanism of construction safety accidents. Trajectory intersection theory is one of the fundamental theories in the field of accident causation analysis. This theory posits that the vast majority of safety accidents are the result of the intersection and collision of unsafe human behavior trajectories and unsafe object trajectories under specific spatiotemporal conditions. Based on this theory, this embodiment decomposes the causes of construction safety accidents into two main categories: human factors and object factors. Human factors are further mapped to personnel posture characteristics and equipment wearing characteristics. Personnel posture characteristics include behavioral patterns such as whether workers stand near edges or stay under crane booms. Equipment wearing characteristics mainly include whether safety helmets are worn and safety belts are fastened. Object factors are further mapped to machinery occupancy characteristics and edge support characteristics. Machinery occupancy characteristics include whether large equipment such as cranes and excavators are in unauthorized operating positions. Edge support characteristics include whether guardrails are installed around the excavation pit and whether edge openings are obstructed. Through this theory-driven mapping relationship, abstract accident causes are transformed into concrete, visually perceptible characteristic categories.
[0037] S1.2. Use the term frequency-inverse document frequency algorithm to extract keywords from the construction safety specification text and filter out high-frequency unsafe feature words.
[0038] This step employs the term frequency-inverse document frequency (TNF) algorithm to automate the analysis of construction safety specification texts. The TNF algorithm is a statistical method for evaluating the importance of words in text. Its core idea is that the higher the frequency of a word in the current document and the lower its frequency in the entire document set, the more representative the word is. In this embodiment, firstly, textual materials such as national and industry-issued construction safety specifications and safety production management regulations are collected. These texts are used as input to the algorithm to calculate the TNF weight value of each unsafety-related word. Then, words are sorted from high to low according to their weight values, and words with weights exceeding a preset threshold are selected as high-frequency unsafety feature words. For example, words such as "safety belt," "safety helmet," "edge protection," and "foundation pit support" are automatically identified as key feature words by the algorithm because of their high frequency in the specification texts and their domain specificity.
[0039] S1.3. The extracted high-frequency unsafe feature words are classified and labeled according to the four-dimensional mapping relationship to generate a four-dimensional unsafe feature standard list including abnormal personnel posture, missing safety equipment, illegal occupation of construction machinery and lack of edge protection. The list defines geometric constraints and visual judgment thresholds for each feature that can be analyzed by visible light imaging of UAV.
[0040] This step categorizes and labels the extracted high-frequency unsafe feature words according to the four-dimensional mapping relationship established in S1.1, generating a structured four-dimensional unsafe feature standard list. The four-dimensional mapping relationship refers to using the personnel posture and equipment worn corresponding to unsafe human behavior as the first and second dimensions, and the mechanical occupancy and edge support corresponding to unsafe object conditions as the third and fourth dimensions. This embodiment clearly defines geometric constraints and visual judgment thresholds for each feature dimension, which can be analyzed by UAV visible light imaging: for abnormal personnel posture, the judgment threshold is defined as the angle between the line connecting key points of the human skeleton and the horizontal plane exceeding 45 degrees and located within 2 meters of the edge area. The key points of the human skeleton are extracted in real time using an open-source posture estimation model, including key points of the shoulders, hips, and knees; the edge area is automatically generated by buffering 2 meters inward from the boundary lines of the foundation pit and floor slab output by the semantic segmentation network. For missing safety equipment, the threshold for judgment is defined as the absence of a safety helmet color feature in the head area or the absence of a safety belt reflective strip feature in the torso area; for construction machinery illegally occupying space, the threshold for judgment is defined as the distance between the machinery outline and the edge of the pit or pedestrian walkway being less than a preset safety distance; for missing edge protection, the threshold for judgment is defined as the absence of a continuous guardrail structure within 1.5 meters around the pit. This list serves as the judgment benchmark for subsequent image recognition models.
[0041] S2. Based on the aforementioned list, collect multi-source images, perform Fourier Merlin transform correction and mosaic-hybrid enhancement, and construct a scale-invariant training dataset. This step, based on the feature list established in the first step, acquires construction scene image data from multiple channels such as UAV aerial photography, building information models, and the internet. To address image distortion and jitter issues generated during UAV flight, frequency domain registration technology is used to correct the images. Simultaneously, data augmentation techniques such as random stitching and color gamut transformation are used to simulate the complex occlusion environment of the construction site. Finally, a training dataset with uniform scale and diverse scenes is established to support the training of subsequent deep learning models.
[0042] Specifically, in this embodiment, the specific method for constructing the scale-invariant training dataset includes:
[0043] S2.1 Based on the geometric constraints defined in the four-dimensional unsafe feature standard list, collect drone aerial video streams, building information model virtual simulation images, and open-source images from the Internet as multi-source image data.
[0044] This step first involves collecting multi-source image data from three channels based on the geometric constraints defined in the four-dimensional unsafe feature standard list established in S1.3. The first channel is drone aerial video streams, where drones are deployed at the actual construction site to automatically cruise and capture continuous video frames of the real construction scene along a preset route. The second channel is building information modeling (BIM) virtual simulation images, where BIM software is used to construct a 3D model of the construction site and render virtual scene images from different perspectives and heights to supplement categories lacking in the real dataset. The third channel is open-source images from the internet, where web crawling technology is used to obtain various construction scene images from publicly available datasets to increase the diversity of the data samples. The image data from these three channels complement each other, forming a multi-source image dataset covering multiple construction stages, multiple shooting angles, and multiple lighting conditions.
[0045] S2.2. Fourier Merlin transform is used to perform frequency domain registration and correction on the acquired multi-source images. The translation, rotation angle and scaling factor between adjacent frames are extracted through phase correlation analysis to eliminate perspective distortion and image jitter caused by UAV attitude changes.
[0046] This step employs Fourier-Melin transform to perform frequency domain registration and correction on the acquired multi-source images. Fourier-Melin transform is an image registration technique based on frequency domain analysis. Its basic principle is to transform the image from the spatial domain to the frequency domain and use phase correlation analysis to calculate the translation, rotation angle, and scaling factor between adjacent frames. In practical applications, slight shaking of the drone's fuselage during flight causes positional shifts and angular rotations between consecutively captured video frames. This embodiment performs a Fourier transform on each frame to obtain its frequency domain representation, calculates the phase difference between adjacent frames, and thus deduces the motion parameters between frames. Then, based on these parameters, a reverse transformation is performed on the images, registering all frames to a unified reference coordinate system. This effectively eliminates perspective distortion and image jitter caused by drone attitude changes, allowing subsequent target detection algorithms to operate on a stable image base.
[0047] To achieve sub-pixel level registration accuracy between drone video frames, this embodiment employs the following frequency domain phase correlation model to solve for the displacement parameters. Let two adjacent frames be... and ,and Depend on After translation Rotation The transformation with scaling factor s yields the following relationship in the frequency domain:
[0048] ,
[0049] in, and They are respectively and Fourier transform spectra, and are frequency-domain coordinates with a unit of cycles per pixel, and j is the imaginary unit. is a phase offset term, which is used to describe the linear phase change introduced by translation in the frequency domain.
[0050] After the above modulo operation eliminates the translation component, the amplitude spectrum relationship is obtained:
[0051] ,
[0052] converting the amplitude spectrum from Cartesian coordinates to log-polar coordinates , wherein is the logarithmic radius, is the polar angle, and after conversion the translation transformation is converted into an addition form:
[0053] ,
[0054] wherein, and are representations of the transformed amplitude spectra under log-polar coordinates, respectively. Through the phase correlation method, the peak position is detected in the log-polar coordinate domain, so that the rotation angle and the scaling factor s can be solved simultaneously. After inversely transforming the image according to the obtained and s, the translation amounts are solved again by the translation phase correlation method, thereby realizing synchronous fine correction of three parameters. s is a scaling factor, where s>1 means the image is enlarged, and 0<s<1 means the image is reduced. Wherein, the value ranges of the translation amounts Δx and Δy are [-W,W] and [-H,H], W and H are the width and height of the image respectively, the unit is pixel; the value range of the rotation angle Δθ is [-π,π], the unit is radian; the value range of the scaling factor s is [0.5, 2]. Phase correlation analysis calculates the cross-power spectrum of the Fourier transforms of two images, performs inverse Fourier transform on the cross-power spectrum, and the coordinate corresponding to the peak position is the translation amount.
[0055] S2.3, Introduce a mosaic-mixed online enhancement strategy, randomly splice four corrected images into a single composite image, and simultaneously perform color gamut perturbation and random cropping during the splicing process.
[0056] This step introduces a mosaic-hybrid online enhancement strategy to augment the corrected image. Mosaic-hybrid enhancement is a data augmentation method that stitches multiple images into a single composite image; its name comes from the mosaic-like effect of combining four small images into one large image. In this embodiment, four corrected images are first randomly selected from the dataset, and then local regions of these four images are stitched together into a complete composite image according to randomly positioned cutting lines. During the stitching process, color gamut perturbation is performed on each sub-image simultaneously, i.e., the brightness, contrast, saturation, and hue parameters of the image are randomly adjusted to simulate the imaging effect under different weather and lighting conditions; at the same time, random cropping is performed, i.e., some edge pixels are randomly discarded to simulate the situation where the target is partially occluded in a construction scene. This enhancement strategy enables the model to learn more diverse scene features during training, enhancing its adaptability to complex scenes with high-density occlusion and multi-scale targets.
[0057] To simulate the differences in imaging under different lighting and weather conditions, this embodiment uses the following color gamut transformation model to perturb the stitched sub-image.
[0058] The original image is converted from the RGB color space to the HSV color space, where H represents hue, S represents saturation, and V represents lightness, corresponding to the converted values respectively. , , Independent random perturbations are applied to each of the three HSV channels:
[0059] ,
[0060] in, This is the hue offset, with a value range of [-0.1, 0.1]. Modulo 1 operation ensures that the hue value always falls within the interval [0, 1). This is the saturation scaling factor, with a value range of [1, 1.5]. This is the brightness scaling factor, with a value range of [0.8, 1.2]. The function truncates the values to a specified interval. The perturbed image is then converted back to the RGB color space to obtain an enhanced sub-image. This is achieved by randomizing the values of each sub-image. , , For parameter combinations, the three parameters of each sub-image are independently and randomly sampled, with a uniform sampling distribution. During training, the data in each batch is independently resampled to ensure that the same original image generates different enhancement variants in different training rounds. A single mosaic stitch can generate 4×3×3×3=108 parameter combinations. Combined with different arrangements of the four sub-images, more than 2000 enhancement variants can be generated, significantly improving data diversity.
[0061] S2.4 Divide the corrected and enhanced image data into training set, validation set and test set according to a preset ratio to construct a multi-source heterogeneous training dataset with scale invariance.
[0062] This step divides the corrected and enhanced image data according to a preset ratio to construct a scale-invariant multi-source heterogeneous training dataset. Scale invariance refers to the model's stable recognition ability for the same target at different pixel sizes presented at different imaging distances. In this embodiment, all image data is divided into three subsets according to their purpose in the training, validation, and testing phases. The ratio is typically set as 60% for the training set, 20% for the validation set, and 20% for the testing set. The training set is used for parameter learning of the deep learning model, the validation set is used to monitor the model's generalization ability and adjust hyperparameters in a timely manner during training, and the testing set is used to finally evaluate the model's recognition performance. The images in the three subsets cover the complete scale range from small targets at a distance to large targets at close range, ensuring that the trained model has good recognition performance for targets at different imaging scales.
[0063] S3. Construct a lightweight dual-channel architecture and train it using the dataset. The first channel extracts dynamic regions of interest through semantic segmentation, while the second channel uses depthwise reparameterizable convolution and a weighted intersection-over-union (IoU) loss function to output target localization feature maps. This step constructs a lightweight neural network architecture that works collaboratively with two channels. One channel is responsible for pixel-level segmentation of the image, identifying the areas where construction workers and equipment are located; the other channel is responsible for accurate localization and classification of targets within these areas. A reparameterizable convolution structure is used to reduce computational complexity, and a loss function specifically designed for small targets is introduced to improve detection accuracy. Finally, the class information and location coordinates of each target in the image are output.
[0064] Specifically, in this embodiment, the specific method for constructing the dual-path lightweight architecture and training it using the dataset includes:
[0065] S3.1 Construct a dual-path lightweight recognition architecture. The first path adopts a semantic segmentation network that introduces a spatial pyramid pooling module and a regularized discarding mechanism to extract dynamic regions of interest containing construction workers and construction equipment from the input image.
[0066] This step constructs the first path in the dual-path lightweight recognition architecture. This path employs a semantic segmentation network that incorporates a spatial pyramid pooling module and a regularized dropout mechanism. The spatial pyramid pooling module is a multi-scale feature extraction structure that processes the input feature map in parallel through multiple pooling operations at different scales, then fuses the pooling results, enabling the network to simultaneously capture both large-scale overall structural information and small-scale local detail information. The regularized dropout mechanism refers to randomly setting the output of some neurons to zero with a certain probability during network training to prevent the network from overfitting to the training data. In this embodiment, the semantic segmentation network classifies each pixel in the input image, determining whether it belongs to construction workers, construction machinery, or the background. The final output is a semantic segmentation map of the same size as the input image, from which dynamic regions of interest containing construction workers and machinery are extracted—i.e., key monitoring areas that may contain violations.
[0067] S3.2 The second path adopts a backbone network that integrates deep reparameterizable convolutions and combines a weighted intersection-over-union loss function to output multi-class target localization feature maps.
[0068] This step constructs the second path in the dual-path lightweight recognition architecture. This path employs a backbone network incorporating deep reparameterizable convolutions and combines a weighted intersection-and-conclusion (CICC) loss function to output multi-class target localization feature maps. Deep reparameterizable convolutions are a special type of convolutional structure that uses multi-branch parallel topology to enhance feature representation during training and folds the multi-branch into single-path convolutions through parameter remapping during inference, balancing training accuracy and inference efficiency. The weighted CICC loss function is an evaluation metric used to measure the degree of overlap between predicted and ground truth bounding boxes. By assigning different weight coefficients to targets of different sizes and categories, the model focuses more on the localization accuracy of small and difficult-to-classify targets. In this embodiment, the second path uses the dynamic region of interest extracted by the first path as input to accurately locate and classify targets such as construction workers, safety helmets, safety belts, and cranes within the region, outputting a feature map containing target category labels and bounding box coordinates.
[0069] In view of the characteristics of small targets and imbalanced class samples in aerial images, this embodiment designs the following weighted intersection-union loss function.
[0070] Let the predicted bounding box be The true bounding box is The area of their intersection is The area of the union is Then the standard intersection-union ratio is:
[0071] ,
[0072] This embodiment introduces a category weighting factor. and scale weighting factor Construct a weighted intersection-union loss function:
[0073] ,
[0074] To ensure the loss function's range is [0,1], set the following constraint: Let , in the formula Replace with ω. When If the value exceeds 1, it is counted as 1. When When the value is greater than 1, truncating to 1 means that the excess portion does not participate in gradient calculation, which is equivalent to setting an attention saturation upper limit for small object detection. Experiments show that when α = 32 pixels, approximately 90% of the training samples satisfy the condition. If the value is ≤1, the impact of the truncation operation on the overall gradient stability is negligible.
[0075] Among them, category weighting factor Defined as:
[0076] ,
[0077] in, The total number of target instances in the training set. For the number of instances of the c-th type of target (e.g., not wearing a helmet), This is used to balance the differences in the number of samples in different categories, with rarer categories receiving higher weights.
[0078] Scale weighting factor Defined as:
[0079] ,
[0080] in, and These are the width and height of the actual bounding box, respectively, both in pixels; As the scale reference threshold, this embodiment uses α = 32 pixels. When the target box area is less than 32 × 32 pixels, >1. Smaller targets receive higher loss weights, guiding the model to pay more attention to the localization accuracy of small-scale targets.
[0081] The final loss function is:
[0082] ,
[0083] Where N is the number of positive sample anchor frames in a single batch. β is the regression loss auxiliary term, and β is the balance coefficient, with a value range of [0.1, 0.5]. In this embodiment, β = 0.3 is used.
[0084] S3.3. Use the dynamic region of interest output by the first path as the attention guide for the second path to limit the target detection search range of the second path.
[0085] This step uses the dynamic region of interest output by the first path as the attention guide for the second path, limiting the target detection search range of the second path. Attention guidance refers to using the positional information output by one path to guide the allocation of computational resources for another path, causing the latter to focus its computation on the effective output region of the former. In this embodiment, the first path identifies which pixels in the image belong to construction areas through semantic segmentation. The second path performs target detection only within these pixels marked as construction areas, without searching the entire image. This collaborative working mode significantly reduces the computational load of the second path. For example, in a 1920×1080 aerial image, construction areas typically only occupy about 30% of the image area. By limiting the region through the first path, the second path can reduce the search computation by about 70%, while avoiding the negative impact of background objects such as vehicles and buildings on the detection results.
[0086] The architecture transformation method of the backbone network in the second path that incorporates deep reparameterizable convolutions specifically includes: during the training phase, standard 3×3 convolutional layers, 1×1 convolutional branches, and identity mapping branches are combined in parallel to form a multi-branch residual topology; during the inference phase, the parameters of each branch are folded into equivalent convolutional kernels and stacked in corresponding positions, and the multi-branch residual topology is equivalently transformed into a single-path linear topology; the backbone network in the inference phase after the transformation consists only of alternating stacks of 3×3 convolutional layers and activation functions.
[0087] This step details the architectural transformation of the deep reparameterizable convolutional backbone network from the training to the inference phase. During training, to improve the network's feature extraction capability, this embodiment employs a multi-branch residual topology, specifically combining a standard 3×3 convolutional layer, a 1×1 convolutional branch, and an identity mapping branch in parallel. The 1×1 convolutional branch adjusts the channel dimension and increases non-linear transformation capability, while the identity mapping branch directly passes the input to the output, allowing gradients to bypass the convolutional layers and propagate back through, effectively mitigating the gradient vanishing problem during deep network training. After parallel computation by the three branches, the results are added element-wise to form a rich feature representation. During inference, to reduce computational load and improve the inference speed of UAV edge devices, this embodiment performs an equivalent transformation of the multi-branch structure from the training phase. The transformation process consists of two steps: first, the convolutional layers within each branch are merged with the subsequent batch normalization layers, absorbing the normalization parameters into the convolutional kernel weights; then, the equivalent convolutional kernels of the three branches are stacked point-by-point according to their corresponding positions to generate a single 3×3 convolutional kernel. After this parameter folding operation, the original multi-branch residual topology is equivalently transformed into a single-path linear topology. The backbone network in the inference stage consists only of standard 3×3 convolutional layers and activation functions stacked alternately. There is no need to calculate 1×1 convolutional branches and identity mapping branches, which greatly reduces the amount of floating-point operations and memory accesses, enabling the model to run in real time on drone edge computing platforms with limited computing power.
[0088] S4. Construct an adversarial transfer learning framework. This framework fuses the binary classification results with the feature map using a channel attention mechanism to output multi-level collaborative identification results for safety violations. This step constructs a cross-domain transfer learning framework. First, a pre-trained binary classification model is used to quickly filter out background images that do not contain construction workers and equipment. Then, the filtered image features are fused with the target localization information obtained in step three using an attention mechanism. Different importance weights are assigned to features from different channels. Finally, based on the feature standard list established in step one, it determines whether there are violations in the image such as not wearing a safety belt, not wearing a safety helmet, cranes illegally occupying space, and unsupported foundation pits.
[0089] Specifically, in this embodiment, the specific method for outputting the collaborative identification results of multi-level security violations includes:
[0090] S4.1 Construct a source-target domain adversarial transfer learning framework, where the source domain is a pre-trained binary classification model and the target domain is the scenario for identifying safety violations in construction areas.
[0091] This step constructs a source-target domain adversarial transfer learning framework. Transfer learning is a technique that applies knowledge learned from one task or domain to another related task or domain. In this embodiment, the source domain refers to a binary classification model pre-trained on a large public dataset, which has powerful general image feature extraction capabilities, such as distinguishing whether an image contains a person or other object. The target domain refers to the specific application scenario of this invention: identification of safety violations in construction areas. The adversarial transfer learning framework introduces a domain classifier and a gradient inversion layer. The feature extractor attempts to generate feature representations that the domain classifier cannot distinguish whether features come from the source or target domain, thereby minimizing the feature distribution differences between the two domains and achieving effective knowledge transfer from a general scenario to a construction-specific scenario.
[0092] S4.2. Qualitatively classify the input image using a binary classification model, filter out background noise images, and output the binary classification qualitative classification result.
[0093] This step uses a binary classification model to qualitatively classify the input images. Qualitative classification means the model only determines whether the image contains construction workers and equipment requiring further analysis, without classifying specific violation types. In drone aerial images of actual construction sites, a significant proportion of the images consist of background areas outside the construction site, such as surrounding roads, buildings, and green belts. These images do not contain any construction workers or equipment and do not require subsequent violation identification. This embodiment utilizes the binary classification model built in S4.1 to quickly filter these images, removing background noise images that do not contain the target object and retaining only images containing construction workers or equipment for subsequent recognition, thereby reducing the computational burden on subsequent multi-class classifiers.
[0094] S4.3. Input the binary classification qualitative judgment result and the target localization feature map into the channel attention mechanism module, extract the channel weights of the feature map through global average pooling and global max pooling, and reconstruct the original feature map by weighting to generate a fused feature map.
[0095] This step fuses the binary classification qualitative judgment result with the target localization feature map output from S3.2 into the channel attention mechanism module. Channel attention is a technique that allows the network to automatically learn the importance of each feature channel. Its core idea is that different feature channels correspond to different semantic information in the image; some channels are crucial for identifying violations, while others contribute less and should be assigned different weights. In this embodiment, the channel attention mechanism module uses two pooling methods to extract feature weights in parallel: global average pooling averages all pixels in each feature channel, reflecting the overall response strength of that channel; global max pooling maximizes each feature channel, reflecting the most significant feature in that channel. After the two pooling results are normalized by the activation function, weight coefficients for each channel are generated. These coefficients are then used to weight the original feature map, enhancing the features of important channels and suppressing the features of secondary channels, ultimately generating a fused feature map.
[0096] This embodiment uses the Sigmoid function as the activation function for normalization. The global average pooling result and the global max pooling result are first added element-wise, and then input into the Sigmoid function to generate channel weights. Let the global average pooling output be A and the global max pooling output be M, then the channel weight vector is W = Sigmoid(A + M).
[0097] S4.4. Based on the four-dimensional unsafe feature standard list, perform multi-class classifier decoding on the fused feature map and output the collaborative identification results of multi-level safety violations such as not wearing a safety belt, not wearing a safety helmet, improper crane positioning, and unsupported foundation pit.
[0098] This step, based on the four-dimensional unsafe feature standard list constructed in S1, performs multi-class classifier decoding on the fused feature map generated in S4.3. Multi-class classifier decoding refers to mapping the fused feature map to a predefined violation category space and outputting a probability score for each category. In this embodiment, the predefined violation categories include: not wearing a safety belt (detecting whether the reflective features of the safety belt on the worker's waist and chest are missing); not wearing a safety helmet (detecting whether the worker's head area has the color and shape characteristics of a safety helmet); improper crane positioning (detecting whether there are workers standing under the crane boom or within its swing radius); and unsupported foundation pit (detecting whether there are continuous guardrails or support structures at the edge of the foundation pit). The decoder outputs confidence scores for each type of violation based on the feature response intensity in the fused feature map. When the score exceeds a preset threshold, the violation is determined to exist, and finally, a multi-level collaborative identification result of safety violations is output.
[0099] In this embodiment, the preset threshold is set to 0.5. For violations such as not wearing a seatbelt and not wearing a safety helmet, the threshold can be reduced to 0.45 to improve recall rate because the visual characteristics of missing safety equipment are relatively clear. For violations such as improper crane positioning and unsupported foundation pits, the threshold is increased to 0.55 to ensure accuracy because it involves understanding complex scenarios. This threshold can be dynamically adjusted using a validation set according to the actual application scenario.
[0100] Furthermore, the training and adaptation methods of the source-target domain adversarial transfer learning framework specifically include:
[0101] S4.1.1 Freeze the convolutional layer parameters of the source domain pre-trained model.
[0102] This step freezes the convolutional layer parameters of the source domain pre-trained model. Parameter freezing means that the weights of these convolutional layers are not updated during subsequent training, maintaining their pre-training state. The pre-trained model has already learned rich low-level visual features on large general datasets, such as edge detection, texture recognition, and shape discrimination in images. These features have good generality and can be directly transferred to the task of identifying violations in construction scenes. In this embodiment, all convolutional layer parameters of the pre-trained model are frozen, allowing the model to retain its ability to extract basic features from general images, while avoiding destructive updates to these already trained basic features due to the limited amount of data in the target domain.
[0103] S4.1.2 In the target domain, a domain adaptation layer is added to the top of the pre-trained model, consisting of a global average pooling layer, a dropout layer, and a fully connected layer connected in sequence.
[0104] This step adds a domain adaptation layer on top of the pre-trained model in the target domain. "Top of the model" refers to the position after the last convolutional layer in the pre-trained model. In this embodiment, the added domain adaptation layer consists of three sub-layers connected sequentially: First, a global average pooling layer compresses the two-dimensional feature map output by the convolutional layer into a one-dimensional feature vector, with each feature channel represented by an average value; second, a dropout layer randomly sets the output of some neurons to zero with a certain probability during training to prevent the model from overfitting to the training data; and finally, a fully connected layer maps the one-dimensional feature vector to the class space of the target domain, outputting a binary classification result. These three sub-layers together form the adaptation bridge from general features to construction-specific features.
[0105] S4.1.3 Introducing a domain adversarial training strategy, the feature extractor and the domain classifier form an adversarial game through the gradient inversion layer, minimizing the feature distribution difference between the source domain and the target domain.
[0106] This step introduces a domain adversarial training strategy, using a gradient inversion layer to create an adversarial game between the feature extractor and the domain classifier. Adversarial training is a training method where two networks compete and optimize together. In this embodiment, the feature extractor aims to extract feature representations that cannot be distinguished as originating from the source or target domain, while the domain classifier aims to accurately determine the source domain of the features. The gradient inversion layer is the key component for implementing this adversarial mechanism. During forward propagation, it keeps the input unchanged, and during backward propagation, it multiplies the gradient by a negative constant. This causes the feature extractor to update its parameters in a direction that increases the loss of the domain classifier, while the domain classifier updates its parameters in a direction that decreases its own loss. This adversarial game ultimately enables the feature extractor to learn domain-invariant feature representations, minimizing the feature distribution difference between the source and target domains.
[0107] S4.1.4. Fine-tune the domain adaptation layer and channel attention mechanism module using the target domain labeled dataset, and gradually unfreeze the high-level convolutional layer parameters of the pre-trained model.
[0108] This step uses a target domain labeled dataset to fine-tune the domain adaptation layer and channel attention mechanism module, and gradually unfreezes the parameters of the high-level convolutional layers in the pre-trained model. Fine-tuning refers to continuing to train some network layers using labeled data from the target domain based on the pre-trained model, making the model adapt to a specific task. In this embodiment, the target domain data is first used to train the domain adaptation layer and channel attention mechanism module, enabling these two parts to learn to identify specific violations in the construction scene from general features. As the training rounds increase, the parameters of the high-level convolutional layers near the output end in the pre-trained model are gradually unfrozen, allowing them to participate in the fine-tuning training and adapt to the feature representation specific to the construction scene; while the low-level convolutional layers near the input end remain frozen, retaining basic edge and texture detection capabilities. This phased, hierarchical fine-tuning strategy ensures the model's generalization ability while achieving knowledge transfer and domain adaptation from the source domain to the target domain.
[0109] S5. Construct a multi-dimensional risk vector using identified violations as risk factors. Combine the risk matrix method and ordered logistic regression model to output a severity index and rectification priority. This step uses the various violations identified in step four as risk sources, and conducts a quantitative assessment from three perspectives: potential personnel injury, equipment damage, and impact on construction progress. Statistical methods are used to screen out the most influential risk factors. The violation is then classified into levels using the risk matrix method. Finally, an ordered regression model is used to calculate the comprehensive severity score for each violation, and a rectification priority order is generated according to the score, providing a decision-making reference for on-site management personnel.
[0110] Specifically, in this embodiment, the specific methods for outputting the severity index and rectification priority include:
[0111] S5.1. The identified violations are used as initial risk factors to construct a multi-dimensional risk vector from three dimensions: probability of personnel injury or death, critical value of equipment damage, and impact of the construction environment.
[0112] This step uses the various violations identified in S4 as initial risk factors to construct a multidimensional risk vector from three dimensions. These three dimensions correspond to three possible consequences of the violation: the probability of personnel injury or death, measuring the likelihood of injury or death to construction workers (e.g., not wearing a safety helmet increases the probability of head injury in the event of falling objects); the equipment damage threshold, measuring the severity of potential damage to construction machinery (e.g., improper crane positioning could lead to boom collisions or overturning); and the impact on the construction environment, measuring the scope of the violation's impact on normal construction operations (e.g., unsupported foundation pits could affect safe passage in surrounding areas). Each violation is assigned a quantitative score across these three dimensions, and the combined scores form a three-dimensional vector, serving as the basis for subsequent severity assessment.
[0113] The quantitative scores for all three dimensions range from [0,1]. The score for the probability of personal injury is determined based on historical accident statistics; for example, not wearing a safety helmet scores 0.85, and not wearing a safety belt scores 0.90. The critical value for equipment damage is assessed based on a comprehensive evaluation of equipment maintenance costs and downtime losses; improper crane positioning scores 0.70. The score for the impact on the construction environment is assessed based on a comprehensive evaluation of the affected area and duration; unsupported foundation pits score 0.60. The specific calibration of the scores can be determined by combining expert experience with historical on-site data.
[0114] S5.2. The Pearson chi-square test is used to perform correlation analysis on the factors in the multidimensional risk vector to screen out the key influencing factors.
[0115] This step employs the Pearson chi-square test to analyze the correlation between the factors in the multidimensional risk vector, identifying key influencing factors. The Pearson chi-square test is a method for determining whether a statistical correlation exists between two discrete variables; its basic principle is to compare the degree of difference between the actual observed frequency and the theoretical expected frequency. In this embodiment, historical data of various violations are used as samples to calculate the chi-square statistic between each risk factor and the severity of the accident consequences. A larger statistic indicates a stronger correlation between the factor and the severity. By assessing the significance level of the chi-square statistic, redundant factors without significant correlation to severity are eliminated, retaining key factors that significantly influence the severity of violations, such as the probability of casualties and the expected rectification time, thereby simplifying the complexity of the subsequent assessment model.
[0116] S5.3 Using the expected rectification time and affected construction area of safety violations as indicators, and combining the risk matrix method, key influencing factors are mapped to preset risk level ranges, and divided into three severity levels: low, medium and high.
[0117] This step uses the estimated rectification time and affected construction area area of safety violations as key indicators, combined with the risk matrix method to classify severity levels. The risk matrix method is a management tool that uses the two dimensions of risk—probability of occurrence and severity of consequences—as rows and columns of a matrix, respectively, and determines the risk level through cross-referencing. In this embodiment, the estimated rectification time is used as the horizontal axis indicator; a longer rectification time indicates higher complexity in handling the violation and a greater potential impact. The affected construction area area is used as the vertical axis indicator; a larger affected area indicates a more severe obstacle to construction progress. Each indicator is divided into several intervals, forming a two-dimensional matrix, with different matrix cells corresponding to different risk levels. This embodiment classifies the severity level into three levels: low risk (short rectification time and small affected area), medium risk (moderate rectification time or moderate affected area), and high risk (long rectification time or large affected area, or both at a high level).
[0118] S5.4. Using the risk level as the dependent variable and the selected key influencing factors as independent variables, input them into an ordered logistic regression model to solve for the probability weights of each factor under different severity levels.
[0119] This step uses the risk level determined by the risk matrix method as the dependent variable and the key influencing factors screened by the Pearson chi-square test as independent variables, inputting them into an ordered logistic regression model to solve for probability weights. Ordered logistic regression is a regression analysis method specifically designed to handle dependent variables with a natural order relationship; for example, in this embodiment, there is a clear progressive relationship between the low, medium, and high levels. The core function of this model is to calculate the contribution of each independent variable to the value of the dependent variable, that is, to solve for the probability weight of each factor at different severity levels. In this embodiment, the model output results are as follows: the probability weight of the casualty probability factor for the high-risk level is 0.65, meaning that this factor plays a major role in the high-risk level; the probability weight of the expected rectification time factor for the medium-risk level is 0.48, indicating that the impact of rectification time is relatively significant at this level. These weight values quantify the influence of each factor on the severity.
[0120] S5.5 Calculate the comprehensive severity index based on the probability weights output by the ordered logistic regression model, and generate a rectification priority sequence for each violation by sorting the index from high to low.
[0121] This step calculates the comprehensive severity index based on the probability weights of each factor output by the ordered logistic regression model and generates a rectification priority sequence. The comprehensive severity index is calculated by weighting and summing the weights of each factor according to a preset formula to obtain a value that comprehensively reflects the overall severity of the violation; a higher index indicates a more severe violation. In this embodiment, the comprehensive severity index is calculated for each identified violation, and then the violations are sorted in descending order of index, with the violation having the highest index listed first, indicating that it requires priority handling. For example, if a construction site simultaneously has three violations—unsupported foundation pit, failure to wear a safety helmet, and improper crane positioning—and the calculated comprehensive severity indices for these three violations are 0.87, 0.52, and 0.71 respectively, then the rectification priority sequence is, in order, unsupported foundation pit, improper crane positioning, and failure to wear a safety helmet, providing a clear order of action for on-site management personnel.
[0122] Furthermore, the specific methods for constructing and solving the ordered logistic regression model include:
[0123] S5.5.1 Divide the severity index into three ordered levels: Level 1 is low risk, Level 2 is medium risk, and Level 3 is high risk.
[0124] This step divides the severity index into three ordered levels. Ordered levels refer to a clear hierarchical relationship between these levels; for example, low risk is superior to medium risk, and medium risk is superior to high risk, but the differences between levels are not necessarily equidistant. In this embodiment, based on the risk matrix method in S5.3, the risk intervals corresponding to the severity index are mapped to three levels: Level 1 is low risk, representing a low probability of consequences from the violation, which can be handled through conventional management methods; Level 2 is medium risk, representing a certain safety hazard from the violation, requiring special rectification; and Level 3 is high risk, representing a violation that could lead to a serious safety accident, necessitating immediate work stoppage and rectification. This three-level division ensures both the precision of the assessment and ease of understanding and implementation by on-site management personnel.
[0125] S5.5.2 Establish a cumulative logistic regression equation to establish a logistic function mapping relationship between the probability that the severity of the violation is less than or equal to the current level and the linear combination of each influencing factor.
[0126] This step establishes a cumulative logistic regression equation, mapping the probability of a violation severity level being less than or equal to the current level to a linear combination of various influencing factors using a logistic function. Cumulative logistic regression is a statistical modeling method for handling ordinal dependent variables. Its core idea is to establish cumulative probability models at each level boundary. In this embodiment, for low-risk levels, the model calculates the probability of severity belonging to level one; for medium-risk levels, the model calculates the cumulative probability of severity belonging to level one or two; and for high-risk levels, the cumulative probability is the maximum value. Each cumulative probability corresponds to a logistic function, whose independent variable is a weighted linear combination of various influencing factors. The function's value range is between 0 and 1, converting the results of linear combinations into probability predictions, thereby characterizing the impact of each factor on the cumulative probability of accident severity.
[0127] This embodiment uses the following cumulative logic function to establish the mapping relationship between severity level and impact factor.
[0128] Let the severity level Y take values of 1, 2, and 3, corresponding to low risk, medium risk, and high risk, respectively. The cumulative probability is defined as:
[0129] ,
[0130] Where j=1,2, corresponding to the two cumulative thresholds of low risk and medium risk; when j=3, P(Y≤3|X)=1. K is the total number of key influencing factors, and in this embodiment, K=5; For example, the range of values for the k-th influencing factor is [0,1], and the range of values for the expected rectification time factor is [0,120] minutes. Let be the regression coefficient of the k-th factor, which is dimensionless; For the first The cumulative intercept term satisfies The monotonically increasing constraint.
[0131] The probability of occurrence for each level is derived from the cumulative probability difference:
[0132] ,
[0133] The parameter set is solved using maximum likelihood estimation. The likelihood function is:
[0134]
[0135] Where M is the total number of training samples, and in this embodiment, M=3000; Let i be the influence factor vector of the i-th sample; This is the indicator function, which takes a value of 1 when the true rank is j, and 0 otherwise. It is achieved by maximizing the log-likelihood function. The optimal estimate of the parameters is obtained.
[0136] S5.5.3. The maximum likelihood estimation method is used to solve the model parameters and obtain the regression coefficients and significance levels of each influencing factor.
[0137] This step uses maximum likelihood estimation to solve for the model parameters, obtaining the regression coefficients and significance levels of each influencing factor. Maximum likelihood estimation is a statistical method that solves for unknown parameters by maximizing the probability of occurrence of sample data. In this embodiment, historical violation data is used as training samples, containing the values of each influencing factor for each violation and its actual severity level. The maximum likelihood estimation method continuously adjusts the values of the regression coefficients to make the severity level distribution predicted by the model as consistent as possible with the actual observed level distribution. The coefficient corresponding to the maximum value of the likelihood function is the optimal estimate. After solving, each influencing factor obtains a regression coefficient, reflecting the direction and strength of the factor's influence on severity. The model also outputs the significance level of each coefficient to determine whether the factor is statistically significant.
[0138] S5.5.4. Based on the sign and absolute value of the regression coefficients, determine the promoting or inhibiting effect of each influencing factor on the severity of the violation, and calculate the average marginal effect of each factor at different severity levels.
[0139] This step determines the promoting or inhibiting effect of each influencing factor on the severity of violations based on the sign and absolute value of the regression coefficients, and calculates the average marginal effect of each factor at different severity levels. The sign of the regression coefficient indicates the direction of influence: a positive coefficient indicates that the larger the factor value, the greater the probability of a higher severity of violation, i.e., it has a promoting effect; a negative coefficient indicates that the larger the factor value, the greater the probability of a lower severity of violation, i.e., it has an inhibiting effect. The absolute value of the regression coefficient indicates the strength of the influence: the larger the absolute value, the more significant the factor's influence on severity. The average marginal effect further quantifies the change in the probability of severity being classified into each level when a certain factor changes by one unit. For example, for every unit increase in the expected rectification time, the probability of being classified as high-risk increases by 15%, while the probability of being classified as low-risk decreases by 10%.
[0140] S5.5.5. Use the parallel lines test to evaluate the goodness of fit of the model. When the significance level of the test statistic is greater than the preset threshold, the model is confirmed to satisfy the proportional dominance hypothesis.
[0141] This step uses the parallel lines test to evaluate the model's goodness of fit and confirm whether the model satisfies the proportional dominance assumption. The proportional dominance assumption is a prerequisite for the validity of an ordered logistic regression model. It requires that each explanatory variable has the same cumulative probability effect on different level limits, meaning that the regression coefficients remain consistent across all cumulative logistic regression equations. The parallel lines test is a statistical test specifically used to verify whether the proportional dominance assumption holds. Its null hypothesis is that there are no significant differences between the regression coefficients corresponding to each level limit. In this embodiment, after running the parallel lines test, a significance level value is obtained. When this value is greater than a preset threshold, it indicates that there is insufficient evidence to reject the null hypothesis, meaning the model satisfies the proportional dominance assumption. This indicates that the constructed ordered logistic regression model has statistical validity and can be used for subsequent severity index calculation and rectification priority ranking.
[0142] This embodiment uses the following likelihood ratio statistic to verify the proportional advantage hypothesis.
[0143] Null hypothesis of parallel lines test For the regression coefficients at each level boundary to be equal, i.e. ,in This is the vector of regression coefficients that define the boundary between level 1 and level 2. This is the regression coefficient vector representing the boundary between levels 2 and 3. The test statistic is calculated as follows:
[0144] ,
[0145] in, This represents the log-likelihood function value of the constrained model (assuming that the regression coefficients at each level are equal); This represents the log-likelihood function value for an unconstrained model (allowing different regression coefficients for each level). Under the null hypothesis, the statistic LR follows a sequence with [degrees of freedom]. Chi-square distribution:
[0146] ,
[0147] Degrees of freedom The number of constraints is calculated using the following formula: ,in The number of key influencing factors. This represents the number of severity levels. In this embodiment... =5, =3, therefore =5×1=5. Taking a significance level of γ=0.05, the critical value is obtained from the chi-square distribution table. If the calculated LR value is less than 11.07, the null hypothesis is accepted, confirming that the model satisfies the proportional advantage hypothesis and has statistical validity.
[0148] This embodiment uses the following normalized weighted summation formula to calculate the comprehensive severity index.
[0149] Suppose that the probability weight of the m-th type of violation at the j-th severity level is obtained by solving the ordered logistic regression model. The overall severity index of this type of violation is... Defined as:
[0150] ,
[0151] The numerator is the weighted sum of the grade values for each grade, and the denominator is the sum of the weights, used for normalization. The value range is [1,3], and the larger the value, the more serious the violation.
[0152] When multiple types of violations exist simultaneously, let there be a total of Q types of violations. The rectification priority sequence is then determined by... Sort in descending order to get:
[0153] ,
[0154] Among them, descending indicates descending order, sorted from high to low according to the comprehensive severity index, with the violation with the highest index ranked first in the priority sequence. The function returns a sorted sequence of original indices, where the first element represents the highest-priority violation category. For violations with similar index values, a secondary ranking index is further applied. To distinguish:
[0155] ,
[0156] This represents the dispersion of the weight distribution for each level, with a value range of [0, 2]. When the difference in the overall severity index of multiple types of violations is less than the preset threshold ε = 0.1, Smaller cases are prioritized because their risk distribution is more concentrated and their risk certainty is higher. If and If all violations are equal, they are ranked according to a preset baseline priority for the type of violation. The preset baseline priority, from highest to lowest, is: not wearing a safety belt, not wearing a safety helmet, improper crane positioning, and no support for the foundation pit. This baseline priority is determined based on the average casualty rate caused by various violations in the statistics of construction safety accidents over the years.
[0157] Example 2
[0158] like Figure 2As shown, this application provides a multi-target collaborative identification system architecture for safety violations in drone construction areas, applied to the multi-target collaborative identification method for safety violations in drone construction areas as described in Embodiment 1, including:
[0159] Feature construction module 210 is used to analyze construction specifications based on trajectory intersection theory and word frequency-inverse document frequency algorithm, and to construct a four-dimensional unsafe feature standard list.
[0160] The dataset construction module 220 is used to collect multi-source images according to the list, and after Fourier Merlin transform correction and mosaic-mixing enhancement, construct a scale-invariant training dataset.
[0161] The dual-path recognition module 230 is used to construct a dual-path lightweight architecture and train it using the dataset. The first path extracts dynamic regions of interest through semantic segmentation, and the second path outputs a target localization feature map using depthwise reparameterizable convolution and a weighted intersection-over-union loss function.
[0162] The collaborative recognition module 240 is used to construct an adversarial transfer learning framework, which integrates the binary classification results with the feature map through a channel attention mechanism, and outputs multi-level collaborative recognition results of security violations.
[0163] The quantitative assessment module 250 is used to construct a multidimensional risk vector using identified violations as risk factors, and outputs a severity index and rectification priority by combining the risk matrix method and the ordered logistic regression model.
[0164] Figure 3 This is an electronic device provided in one embodiment of this application. For example... Figure 3 As shown, the electronic device includes at least the following components: processor 301 and memory 300, communication interface 303, and bus 302.
[0165] In this embodiment of the application, memory 300 is used to store executable instructions of processor 301, which, when configured to execute instructions, implements the method as described in the first aspect.
[0166] In embodiments of this application, a computer-readable storage medium includes instructions that instruct a device to perform the method as described in the first aspect. For example, the instructions instruct the device to perform... Figure 1 The method is shown in the process steps.
[0167] In one embodiment of this application, the program operating in the electronic device may be a program that controls a central processing unit (CPU) or similar device to achieve the functions of the above-described embodiments of the present invention (a program that enables the computer to function). Information processed by these systems is then temporarily stored in random access memory (RAM) during processing, and subsequently stored in various ROMs such as read-only memory (FlashROM) and hard disk drives (HDDs), and read, corrected, and written by the CPU as needed.
[0168] It should be noted that a portion of the electronic device described in the above embodiments can also be implemented using a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and the program recorded on the recording medium can be read into the computer and executed.
[0169] It should be noted that the computer mentioned here refers to a computer built into an electronic device, employing hardware including an operating system and peripheral devices. Furthermore, computer-readable recording media refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage systems such as hard drives built into the computer.
[0170] Furthermore, computer-readable recording media can include: media that dynamically stores programs for short periods of time, such as communication lines used when transmitting programs via networks like the Internet or communication lines like telephone lines; and media that store programs for fixed periods of time, such as volatile memory inside a computer that serves as a server or client in this case. In addition, the aforementioned program can be a program used to implement the above-mentioned functions, or it can be a program that can implement the above-mentioned functions by combining them with programs already recorded in the computer.
[0171] Furthermore, the electronic device in the above embodiments can also be implemented as an assembly (system group) composed of multiple systems. Each system constituting the system group can possess some or all of the functions or functional blocks of the electronic device in the above embodiments. As a system group, it is sufficient to have all the functions or functional blocks of the electronic device.
[0172] Those skilled in the art should recognize that the above embodiments are only used to illustrate this application and are not intended to limit this application. Any appropriate changes and variations made to the above embodiments within the essential spirit and scope of this application fall within the scope of protection claimed in this application.
Claims
1. A multi-target collaborative identification method for safety violations in unmanned aerial vehicle (UAV) construction areas, characterized in that, Includes the following steps: Based on trajectory intersection theory and word frequency-inverse document frequency algorithm to analyze construction specifications, a four-dimensional list of unsafe features is constructed. Based on the aforementioned list, multi-source images were collected, and after Fourier Merlin transform correction and mosaic-mixing enhancement, a scale-invariant training dataset was constructed. A dual-path lightweight architecture is constructed and trained using the dataset. The first path extracts dynamic regions of interest through semantic segmentation, and the second path outputs a target localization feature map using depthwise reparameterizable convolution and a weighted intersection-union loss function. An adversarial transfer learning framework is constructed, which integrates the binary classification results with the feature map through a channel attention mechanism to output multi-level collaborative identification results of security violations. A multidimensional risk vector is constructed using the identified violations as risk factors. The severity index and rectification priority are output by combining the risk matrix method and the ordered logistic regression model.
2. The method according to claim 1, characterized in that, The specific methods for constructing the four-dimensional list of unsafe features include: Based on the trajectory intersection theory, the causes of construction safety accidents are analyzed as unsafe human behavior and unsafe conditions of objects. Unsafe human behavior is mapped as personnel posture characteristics and equipment wearing characteristics, and unsafe conditions of objects are mapped as mechanical occupancy characteristics and edge support characteristics. Keyword extraction was performed on the construction safety specification text using the term frequency-inverse document frequency algorithm, and high-frequency unsafe feature words were selected. The extracted high-frequency unsafe feature words are classified and labeled according to the four-dimensional mapping relationship, generating a four-dimensional unsafe feature standard list that includes abnormal personnel posture, missing safety equipment, illegal occupation of construction machinery and lack of edge protection. The list defines geometric constraints and visual judgment thresholds for each feature that can be analyzed by visible light imaging of UAVs.
3. The method according to claim 1, characterized in that, The specific methods for constructing the scale-invariant training dataset include: Based on the geometric constraints defined in the four-dimensional unsafe feature standard list, drone aerial video streams, building information model virtual simulation images, and open-source images from the Internet are collected as multi-source image data; Fourier Merlin transform was used to perform frequency domain registration and correction on the acquired multi-source images. The translation, rotation angle and scaling factor between adjacent frames were extracted by phase correlation analysis to eliminate perspective distortion and image jitter caused by UAV attitude changes. A mosaic-hybrid online enhancement strategy is introduced to randomly stitch four corrected images into a single synthetic image, and color gamut perturbation and random cropping are performed simultaneously during the stitching process; The corrected and enhanced image data is divided into training set, validation set and test set according to a preset ratio to construct a multi-source heterogeneous training dataset with scale invariance.
4. The method according to claim 1, characterized in that, The specific methods for constructing the dual-path lightweight architecture and training it using the dataset include: A dual-path lightweight recognition architecture is constructed. The first path adopts a semantic segmentation network that introduces a spatial pyramid pooling module and a regularized discarding mechanism to extract dynamic regions of interest containing construction workers and construction equipment from the input image. The second path employs a backbone network that integrates deep reparameterizable convolutions and combines a weighted intersection-over-union loss function to output multi-class target localization feature maps. The dynamic region of interest output by the first path is used as the attention guide for the second path, thus limiting the target detection search range of the second path.
5. The method according to claim 4, characterized in that, The specific architecture transformation methods for the backbone network that incorporates depthwise reparameterizable convolutions in the second path include: During the training phase, the standard 3×3 convolutional layer, 1×1 convolutional branch and identity mapping branch are combined in parallel to form a multi-branch residual topology. During the inference phase, the parameters of each branch are folded into equivalent convolutional kernels and stacked at corresponding positions, and the multi-branch residual topology is equivalently converted into a single-path linear topology. The transformed inference stage backbone network consists of only alternating stacks of 3×3 convolutional layers and activation functions.
6. The method according to claim 1, characterized in that, The specific methods for outputting the collaborative identification results of multi-level security violations include: A source-target domain adversarial transfer learning framework is constructed, where the source domain is a pre-trained binary classification model and the target domain is the scenario of identifying safety violations in construction areas. The input image is qualitatively classified using a binary classification model, background noise is filtered out, and the binary classification qualitative classification result is output. The binary classification qualitative judgment result and the target localization feature map are input into the channel attention mechanism module. The channel weights of the feature map are extracted by global average pooling and global max pooling. The original feature map is then reconstructed by weighting to generate a fused feature map. Based on the four-dimensional unsafe feature standard list, a multi-class classifier is used to decode the fused feature map, and the collaborative identification results of multi-level safety violations such as not wearing a safety belt, not wearing a safety helmet, improper crane positioning, and unsupported foundation pit are output.
7. The method according to claim 6, characterized in that, The training and adaptation methods of the source-target domain adversarial transfer learning framework specifically include: Freeze the convolutional layer parameters of the source domain pre-trained model; In the target domain, a domain adaptation layer is added on top of the pre-trained model, consisting of a global average pooling layer, a dropout layer, and a fully connected layer connected in sequence. A domain adversarial training strategy is introduced, which uses a gradient inversion layer to enable the feature extractor and the domain classifier to engage in an adversarial game, minimizing the feature distribution difference between the source domain and the target domain. The domain adaptation layer and channel attention mechanism module are fine-tuned and trained using the target domain labeled dataset, and the parameters of the high-level convolutional layers of the pre-trained model are gradually unfrozen.
8. The method according to claim 1, characterized in that, The specific methods for outputting the severity index and rectification priority include: The identified violations are used as initial risk factors, and a multi-dimensional risk vector is constructed from three dimensions: probability of personnel injury or death, critical value of equipment damage, and impact of the construction environment. The Pearson chi-square test was used to perform correlation analysis on the factors in the multidimensional risk vector to screen out key influencing factors. Using the expected rectification time and affected construction area of safety violations as indicators, and combining the risk matrix method, key influencing factors are mapped to preset risk level ranges, and divided into three severity levels: low, medium and high. The risk level is used as the dependent variable, and the selected key influencing factors are used as independent variables. The results are input into an ordered logistic regression model to solve for the probability weight of each factor under different severity levels. The comprehensive severity index is calculated based on the probability weights output by the ordered logistic regression model, and a rectification priority sequence for each violation is generated by sorting the indexes from high to low.
9. The method according to claim 8, characterized in that, The specific methods for constructing and solving the ordered logistic regression model include: The severity index is divided into three ordered levels: Level 1 is low risk, Level 2 is medium risk, and Level 3 is high risk. Establish a cumulative logistic regression equation to establish a logistic function mapping relationship between the probability that the severity of the violation is less than or equal to the current level and the linear combination of each influencing factor; The maximum likelihood estimation method was used to solve the model parameters, and the regression coefficients and significance levels of each influencing factor were obtained. Based on the sign and absolute value of the regression coefficients, determine the promoting or inhibiting effect of each influencing factor on the severity of the violation, and calculate the average marginal effect of each factor at different severity levels. The goodness of fit of the model is evaluated using the parallel lines test. When the significance level of the test statistic is greater than the preset threshold, the model is confirmed to satisfy the proportional dominance hypothesis.
10. A multi-target collaborative identification system for safety violations in unmanned aerial vehicle (UAV) construction areas, applied to the method described in any one of claims 1 to 9, characterized in that, The system includes: The feature construction module is used to analyze construction specifications based on trajectory intersection theory and word frequency-inverse document frequency algorithm, and to construct a four-dimensional list of unsafe features. The dataset construction module is used to collect multi-source images according to the list, and after Fourier Merlin transform correction and mosaic-mixing enhancement, construct a scale-invariant training dataset. The dual-path recognition module is used to construct a lightweight dual-path architecture and train it using the dataset. The first path extracts dynamic regions of interest through semantic segmentation, and the second path outputs a target localization feature map using depthwise reparameterizable convolution and a weighted intersection-over-union loss function. The collaborative identification module is used to construct an adversarial transfer learning framework. It integrates the binary classification results with the feature map through a channel attention mechanism and outputs multi-level collaborative identification results of security violations. The quantitative assessment module is used to construct a multidimensional risk vector based on the identified violations as risk factors, and outputs a severity index and rectification priority by combining the risk matrix method and the ordered logistic regression model.