Helmet wearing normativity analysis method based on HGNN

Through the standardized analysis method of safety helmet wearing based on hybrid aggregation network and bottom-up decoding network, the problem that existing detection methods cannot accurately judge the standardizedness of safety helmet wearing is solved, automatic and efficient inspection is achieved, and safety supervision capabilities at the construction site are improved.

CN120339939APending Publication Date: 2025-07-18NANTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510341010.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing safety helmet wear detection methods cannot automatically and accurately judge the standardization of safety helmets on construction sites, resulting in time-consuming and labor-intensive manual patrols and inaccurate inspections, making it difficult to achieve real-time and comprehensive monitoring.

Method used

The standardized analysis method of safety helmet wear based on hybrid aggregation network and bottom-up decoding network is adopted. By constructing a single-stage object detection model, a hybrid aggregation network and HGNN are introduced, feature encoding network is constructed, depth information is extracted, and bottom-up decoding network output detection results are constructed.

Benefits of technology

It realizes automatic, efficient and accurate inspection of standardized wear of safety helmets, improves construction site safety supervision capabilities, and reduces safety hazards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339939A_ABST
    Figure CN120339939A_ABST
Patent Text Reader

Abstract

The invention provides a safety helmet wearing normalization analysis method based on HGNN, and belongs to the technical field of artificial intelligence. The technical problems of how to automatically, efficiently and accurately detect the normalization of safety helmet wearing so as to improve the safety supervision capability of a construction site and reduce potential safety hazards are solved. According to the technical scheme, the method comprises the following steps that S1, a safety helmet data set is collected, and a classification data set is made according to wearing normativity; s2, constructing a single-stage target detection model by taking the helmet wearing normativity detection data set as input and taking a wearing normativity detection result as output; s3, introducing a hybrid aggregation network, constructing a feature coding network, and extracting depth information; s4, constructing an HGNN-based information processing network, and efficiently processing depth information; and S5, constructing a Bottom-Up decoding network from bottom to top, and outputting a detection result. The system has the beneficial effects that the normalization of safety helmet wearing can be automatically, efficiently and accurately detected, the construction site safety supervision capability is provided, and the potential safety hazard of the construction site is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method for analyzing the wearing standardization of safety helmets based on HGNN. Background Art

[0002] With the rapid development of the construction industry, construction site safety management has become an important link in ensuring the safety of construction workers and the quality of projects. As one of the most basic personal protective equipment on the construction site, the correct wearing of safety helmets plays a crucial role in preventing the head from being accidentally impacted, injured by falling objects, etc. However, in the actual construction process, the phenomenon of non-standard wearing of safety helmets occurs from time to time. For example, the helmet strap is not fastened, the visor is not worn correctly, the safety helmet is damaged, etc. These situations may pose safety hazards to construction workers and even lead to serious accidents.

[0003] Traditional safety helmet wearing detection methods mainly rely on manual inspections, and this method has many limitations. First of all, manual inspections require a large amount of manpower and time, and it is difficult to achieve real-time and comprehensive monitoring of the construction site. Secondly, manual judgment is subjective and is easily affected by factors such as fatigue and negligence, resulting in inaccurate detection results. In addition, for large construction sites, the coverage of manual inspections is limited, and it is difficult to timely discover all situations of non-standard wearing of safety helmets.

[0004] In recent years, with the rapid development of computer vision technology and deep learning, safety helmet wearing detection methods based on image recognition have gradually attracted attention. These methods collect image or video data of the construction site through cameras, and use deep learning algorithms to identify and analyze the people in the images, so as to judge whether the safety helmet is worn correctly. However, the existing detection methods based on deep learning cannot analyze the wearing standardization of safety helmets, and there is an urgent need for a method for analyzing the wearing standardization of safety helmets. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for analyzing the wearing standardization of safety helmets based on HGNN, which solves the technical problem that the existing detection algorithms cannot judge the wearing standardization of safety helmets. The present invention can automatically, efficiently and accurately detect the wearing standardization of safety helmets, improve the construction site safety supervision ability, and reduce construction site safety hazards.

[0006] In order to achieve the above invention purpose, the technical solution adopted by the present invention is specifically as follows: A method for analyzing the wearing standardization of safety helmets based on HGNN of the present invention includes the following steps:

[0007] Step 1, collect a safety helmet data set and make a classification data set according to the wearing standardization;

[0008] Step 2: Using the helmet wearing compliance inspection dataset as input and the wearing compliance inspection results as output, a single-stage target detection model is constructed;

[0009] Step 3: Introduce a hybrid aggregation network, construct a feature encoding network, and extract depth information;

[0010] Step 4: Construct an information processing network based on HGNN to efficiently process deep information;

[0011] Step 5: Construct a bottom-up decoding network and output the detection results.

[0012] Furthermore, step 1 specifically includes the following steps:

[0013] Step 1.1: Collect monitoring image data of several workers at a construction site, including samples of workers wearing helmets in a standardized and irregular manner;

[0014] Step 1.2: In the labeling stage, the helmet wearing conditions are manually labeled into two types: standard wearing and irregular wearing, and low-quality data images such as blurry and low-light images are eliminated;

[0015] Step 1.3: Based on the annotation information, take the center point of the annotation box as the reference point and randomly move 104 to 312 pixels to the upper left to determine the starting position of the cropped area. Then, using this position as the starting point, expand 614 pixels to the lower right to extract a local image of size 614×614 pixels. Each annotation box will be randomly sampled three times to generate three new training samples. If the generated image area exceeds the boundary of the original image during the cropping process, the sample will be discarded;

[0016] Step 1.4, perform Mosaic data enhancement on the defect image;

[0017] Step 1.5: After dividing the data set into training set, validation set and test set, the optimization algorithm is used to adjust the model parameters to minimize the loss function. In this process, the loss function combines the cross entropy loss function with the positioning loss function, and its formula is:

[0018] L=αL cls +βL loc

[0019] Among them, L cls is the classification loss, L loc is the positioning loss, α and β are weight coefficients used to balance the losses of classification and positioning.

[0020] Furthermore, the step 2 comprises the following steps:

[0021] Step 2.1: Build and train an end-to-end object detection framework. The loss function includes object confidence loss classification loss and bounding box regression loss which are three parts; the total loss is the weighted sum of these three parts, and the expression is:

[0022]

[0023] where the weights of object confidence loss, classification loss and bounding box regression loss are λ1, λ2, λ3 respectively; L conf and L cls are calculated using the binary cross-entropy function, and the calculation formula is as follows:

[0024] L n = -(y n log(δ(x n )) + (1 - y n ) log(1 - δ(x n )))

[0025] In the formula: x n is the score for predicting the nth sample as a positive example, y n represents the label of the nth sample, δ represents the sigmoid function, L reg uses the CloU Loss function. CloU is a method for calculating bounding box regression loss, and the calculation formula is as follows:

[0026]

[0027] In the formula: the parameter A represents the ground truth box, B represents the predicted box, IoU is called the intersection over union, which represents the overlapping degree of the predicted box and the ground truth box, ρ 2 (·) represents the Euclidean distance, b represents the center point coordinates of the predicted box, b gt represents the center point coordinates of the ground truth box, a represents the balance parameter, and v is the consistency of the aspect ratio of the ground truth box and the predicted box.

[0028] Furthermore, Step 3 specifically includes the following steps:

[0029] Step 3.1: In helmet specification detection, traditional methods often have difficulty effectively modeling spatial correlation features. We introduce the idea of hypergraph computing to solve this problem. The hypergraph G = (V, E) is defined by its vertex set V and hyperedge set E. To construct the vertex set V, different from the traditional YOLO architecture, we extract five feature layers in the backbone as the vertices of the hypergraph.

[0030] Step 3.2: To fuse cross - level information and enhance the feature extraction ability of the basic network, we designed a parallel convolutional coding structure. This structure synergistically fuses three typical convolutional variants: 1×1 convolution for channel feature recalibration, deformable convolution (DCNV) for effective spatial correlation information, and retains the C3K2 convolutional module in the network. This enables the model to adjust the feature extraction and fusion strategy for different inputs, achieving systematic multi - level feature modeling from local to global.

[0031] Furthermore, step 4 specifically includes the following steps:

[0032] Step 4.1: To model the neighborhood relationship in the hypergraph module, a hyper - edge set E is constructed through a distance threshold λ. Specifically, for each feature point x u , find all feature points with a distance less than ∈, and form a hyper - edge with them and x v . The hyper - edge e can be expressed as:

[0033] e = {u|||x u - x v ||2 < ∈, u ∈ V}

[0034] where ||·||2 represents the Euclidean norm. All such hyper - edges constitute the hyper - edge set E. The incidence matrix H of the hypergraph G=(V, E) is defined as:

[0035]

[0036] Step 4.2: In the hypergraph convolutional network, as the number of layers increases, information may gradually be lost during propagation. By introducing residual connections, the model can directly pass the input information to the output, thus avoiding information loss. To propagate high - order information on the hypergraph structure, spatial hypergraph convolution with residual connections is adopted, and its calculation process is as follows:

[0037]

[0038] where M V (e) represents the vertex neighborhood of the hyper - edge e, and Θ e is the trainable parameter. For the feature matrix X of a given vertex and the adjacency matrix H of the hypergraph, assuming D v and D e are the diagonal matrices of vertices and hyper - edges respectively, the hypergraph convolution can be expressed as:

[0039]

[0040] where, calculates the normalized adjacency matrix, It represents aggregating vertex features through hyperedges to capture the high-order relationships between vertices. Θ is a learnable parameter matrix used to transform the aggregated features and enhance the expressive power of the model. Then, X is introduced as a residual connection to retain the original feature information and avoid information loss.

[0041] Furthermore, step 5 specifically includes the following steps:

[0042] Step 5.1: To enhance the semantic consistency of the feature pyramid, a decoding network based on bidirectional cross-scale connections is designed. An improved weighted bidirectional feature pyramid network (BiFPN) is introduced, and its mathematical expression is as follows:

[0043]

[0044] Among them, is the input feature of the l-th layer, is upsampling, is downsampling, w k is the learnable weight, and σ is the Sigmoid normalization is channel concatenation, is a 3×3 depthwise separable convolution.

[0045] Step 5.2: Hypergraph-guided cross-layer aggregation:

[0046] Perform cross-modal fusion on the high-order features Hyper output by HGNN and the features of the C3-C5 layers of the backbone network:

[0047]

[0048] Among them, K l = C l W K , V l = C l W V is a linear transformation. DCNv2 is a deformable convolution to enhance spatial adaptability.

[0049] Step 5.3: Adaptive spatial reweighting module: Design a spatial attention weight map to dynamically adjust the feature response:

[0050]

[0051] Among them, W xy is RC 3 is a position-sensitive learnable parameter to achieve pixel-level feature enhancement.

[0052] Step 5.4: Dynamic prototype matching detection head: Construct a classification-regression joint output layer based on prototype learning:

[0053]

[0054] Among them, P+ / P- is the prototype library of positive and negative samples, and Δ is the boundary threshold. The classification robustness is improved by dynamically updating the prototype.

[0055] Step 5.5, Improved non-maximum suppression: "Differentiable NMS (dNMS)" is proposed to solve the non-differentiability of traditional NMS

[0056] Problem:

[0057]

[0058] Among them, IoU(b i ,b j ) is the intersection over union between bounding boxes b i and b j . k is an adjustment parameter, and the prediction results of occluded targets are retained through a soft suppression mechanism.

[0059] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0060] 1. When the C3K2 module used in the original YOLO11 model extracts features, it mainly relies on a single convolution operation, and there are deficiencies in the richness of the information flow. In addition, when C3K2 processes multi-scale features, it mainly relies on a simple feature fusion method, and its ability to integrate cross-level features is limited. The hybrid aggregation network introduced in the present invention enriches the diversity of the information flow through a parallel convolution mechanism, thereby enhancing the feature extraction ability and making full use of semantic information at different levels;

[0061] 2. Construct an information processing network of HGNN, and cleverly apply the idea of the hypergraph convolutional network to the field of object detection. Different visual features of wearing details are constructed as hypergraph vertices, and a neural network is trained using hyperedges. HGNN fuses information at different levels and enhances the ability of the basic network to model global information;

[0062] 3. Use a Bottom-Up decoding network. The feature layer extracted by the Backbone is introduced again in the Neck part of the network, effectively solving the problem of feature loss caused by continuous convolution operations in the original model. Bottom-Up compensates for this loss by transferring high-resolution spatial details from the lower layer back to the higher layer through a bottom-up fusion mechanism, so as to ensure that the large-scale and medium-scale object detection branches can access these key details and obtain higher detection accuracy. Description of the Drawings

[0063] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention.

[0064] Figure 1 This is the overall flowchart of the present invention.

[0065] Figure 2 This is the HGNN convolution design diagram in the present invention.

[0066] Figure 3 This is the parallel convolution encoding structure diagram in the present invention.

[0067] Figure 4 This is the deep learning network structure diagram in the present invention. Detailed implementation manners

[0068] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0069] Embodiment 1

[0070] Refer to Figure 1 and Figure 4 , a method for analyzing the compliance of safety helmet wearing based on HGNN of the present invention includes the following steps:

[0071] S1. Collect a safety helmet data set and make a classification data set according to the wearing compliance:

[0072] (1) At a construction site, collect the monitoring image data of several workers at work, which includes samples of safety helmets worn in compliance and samples of non-compliant wearing.

[0073] (2) In the label annotation stage, manually annotate the wearing conditions of safety helmets into two types: compliant wearing and non-compliant wearing, and eliminate low-quality data images such as blurred and low-light images.

[0074] (3) According to the annotation information, use the center point of the annotation box as a reference point, and randomly move 104 to 312 pixels to the upper left to determine the starting position of the cropping area. Then, starting from this position, expand 614 pixels to the lower right to extract a local image with a size of 614×614 pixels. Each annotation box will be randomly sampled three times to generate three new training samples. If the generated image area exceeds the boundary of the original image during the cropping process, the sample will be discarded;

[0075] (4) Perform Mosaic data augmentation on defective images;

[0076] (5) After dividing the dataset into a training set, a validation set, and a test set, an optimization algorithm is used to adjust the model parameters to minimize the loss function. During this process, the loss function combines the cross-entropy loss function and the localization loss function, and its specific formula is:

[0077] L = αL cls + βL loc

[0078] where L cls is the classification loss, L loc is the localization loss, and α and β are weight coefficients used to balance the classification and localization losses.

[0079] S2. Using the safety helmet wearing compliance detection dataset as the input and the wearing compliance detection result as the output, a single-stage object detection model is constructed:

[0080] (1) Build and train an end-to-end object detection framework. The loss function includes the object confidence loss classification loss and bounding box regression loss These three parts; the total loss is the weighted sum of these three parts, and the expression is:

[0081]

[0082] where the weight values of the object confidence loss, classification loss, and bounding box regression loss are λ1, λ2, and λ3 respectively; L conf and L cls are calculated using the binary cross-entropy function, and the calculation formula is as follows:

[0083] L n = -(y n log(δ(x n )) + (1 - y n ) log(1 - δ(x n )))

[0084] In the formula: x n is the score for predicting the nth sample as a positive example, y n represents the label of the nth sample, δ represents the sigmoid function, and L reg uses the CloU Loss function. CloU is a method for calculating the bounding box regression loss, and the calculation formula is as follows:

[0085]

[0086] In the formula: the parameter A represents the ground truth box, B represents the predicted box, IoU is called the intersection over union, which represents the overlap degree between the predicted box and the ground truth box, ρ 2(·) represents the Euclidean distance, b represents the coordinates of the center point of the predicted bounding box, and b gt represents the coordinates of the center point of the ground truth bounding box, a represents the balance parameter, and v is the consistency of the aspect ratio between the ground truth bounding box and the predicted bounding box.

[0087] S3. Introduce a hybrid aggregation network, construct a feature encoding network, and extract depth information:

[0088] (1) In helmet specification detection, traditional methods often have difficulty effectively modeling spatial correlation features. We introduce the idea of hypergraph computing to solve this problem. A hypergraph G=(V, E) is defined by its vertex set V and hyperedge set E. To construct the vertex set V, different from the traditional YOLO architecture, we extract five feature layers in the backbone as the vertices of the hypergraph.

[0089] (2) To fuse cross-level information and enhance the feature extraction ability of the basic network, we design a parallel convolutional encoding structure. This structure synergistically fuses three typical convolutional variants: a 1×1 convolution for channel feature recalibration, a deformable convolution (DCNV) for effective spatial correlation information, and retains the C3K2 convolutional module in the network. This enables the model to adjust the feature extraction and fusion strategy for different inputs, achieving systematic multi-level feature modeling from local to global.

[0090] S4. Construct an information processing network based on HGNN to efficiently process depth information:

[0091] (1) To model the neighborhood relationship in the hypergraph module, the hyperedge set E is constructed through a distance threshold λ. Specifically, for each feature point x u , find all feature points with a distance less than ∈, and form a hyperedge with them and x v . The hyperedge e can be expressed as:

[0092] e = {u|||x u -x v ||2 < ∈, u ∈ V}

[0093] Among them, ||·||2 represents the Euclidean norm. All such hyperedges constitute the hyperedge set E, and the incidence matrix H of the hypergraph G=(V, E) is defined as:

[0094]

[0095] (2) In the hypergraph convolutional network, as the number of layers increases, information may gradually be lost during the propagation process. By introducing residual connections, the model can directly pass the input information to the output, thus avoiding information loss. To transmit high-order information on the hypergraph structure, a spatial hypergraph convolution with residual connections is adopted, and its calculation process is as follows:

[0096]

[0097] Among which N V (e) represents the vertex neighborhood of the hyperedge e, Θ e is a trainable parameter. For the feature matrix X of a given vertex and the adjacency matrix H of the hypergraph, assume D v and D e are the diagonal matrices of vertices and hyperedges respectively, and the hypergraph convolution can be expressed as:

[0098]

[0099] Among them, calculate the normalized adjacency matrix, represents aggregating vertex features through hyperedges to capture the high-order relationships between vertices. Θ is a learnable parameter matrix used to transform the aggregated features and enhance the expressive power of the model. Then introduce X as a residual connection to retain the original feature information and avoid information loss.

[0100] S5. Construct a bottom-up decoding network to output the detection result:

[0101] (1) To enhance the semantic consistency of the feature pyramid, design a decoding network based on bidirectional cross-scale connections. Introduce an improved weighted bidirectional feature pyramid network (BiFPN), and its mathematical expression is as follows:

[0102]

[0103] Among them, is the input feature of the l-th layer, is upsampling, is downsampling, w k is a learnable weight, σ is Sigmoid normalization is channel concatenation, is a 3×3 depthwise separable convolution.

[0104] (2) Hypergraph-guided cross-layer aggregation:

[0105] Perform cross-modal fusion on the high-order features Hyper output by HGNN and the features of the C3-C5 layers of the backbone network:

[0106]

[0107] Among them, K l = C l W K , V l = C l W Vis a linear transformation. DCNv2 is a deformable convolution that enhances spatial adaptability.

[0108] (3) Adaptive Spatial Reweighting Module: Design a spatial attention weight map to dynamically adjust feature responses:

[0109]

[0110] where W xy is RC 3 is a position-sensitive learning parameter to achieve pixel-level feature enhancement.

[0111] (4) Dynamic Prototype Matching Detection Head: Construct a classification-regression joint output layer based on prototype learning:

[0112]

[0113] where P+ / P- are the positive and negative sample prototype libraries, and Δ is the boundary threshold. The classification robustness is improved by dynamically updating the prototypes.

[0114] (5) Improved Non-Maximum Suppression: Propose "Differentiable NMS (dNMS)" to solve the non-differentiable problem of traditional NMS:

[0115]

[0116] where k is a tuning parameter, and the prediction results of occluded objects are retained through a soft suppression mechanism.

[0117] Example 2

[0118] Train and test on a dataset. The dataset is divided into a training set, a test set, and a validation set in a ratio of 6:2:2.

[0119] For computing resources, use 16GB of memory and the NVIDIA GeForce RTX 4060 along with the CUDA 11.3 GPU acceleration library. Train the end-to-end network. During training, the training batch is set to 200 times, the training batch size is 16, the learning rate is set to 0.937, the momentum is 0.937, the optimizer is selected as ADAM, the batch size is 64, and the input feature size is set to 640×640.

[0120] This example selects mAP@0.5 as the model evaluation metric. mAP@0.5 refers to the mAP when the IoU threshold is 50%. As shown in Table 1 below, the improved algorithm in this example is higher than the RT-DETR algorithm, YOLOv8, and YOLO11 algorithms before improvement in terms of mAP@0.5 and mAP@0.5:0.95 metrics.

[0121] Table 1

[0122]

[0123] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for analyzing the standardization of safety helmet wearing based on HGNN, characterized in that The steps include: Step 1: Collect the helmet data set and create a classification data set according to the wearing standardization; Step 2: Using the helmet wearing compliance inspection dataset as input and the wearing compliance inspection results as output, a single-stage target detection model is constructed; Step 3: Introduce a hybrid aggregation network, construct a feature encoding network, and extract depth information; Step 4: Construct an information processing network based on HGNN to efficiently process deep information; Step 5: Construct a bottom-up decoding network and output the detection results.

2. The method for analyzing the standardization of safety helmet wearing based on HGNN according to claim 1, characterized in that The step 1 comprises the following steps: Step 1.1: Collect monitoring image data of several workers at a construction site, including samples of workers wearing helmets in a standardized and irregular manner; Step 1.2: In the labeling stage, the helmet wearing conditions are manually labeled into two types: standard wearing and irregular wearing, and blurry, low-light and low-quality data images are eliminated; Step 1.3: Based on the annotation information, take the center point of the annotation box as the reference point and randomly move 104 to 312 pixels to the upper left to determine the starting position of the cropping area. Then, take this position as the starting point and expand 614 pixels to the lower right to extract a local image with a size of 614×614 pixels. Each annotation box will be randomly sampled three times to generate three new training samples. If the generated image area exceeds the boundary of the original image during the cropping process, the sample will be discarded. Step 1.4, perform Mosaic data enhancement on the defect image; Step 1.5: After dividing the data set into training set, validation set and test set, the optimization algorithm is used to adjust the model parameters to minimize the loss function. The loss function combines the cross entropy loss function with the positioning loss function. The formula is: L = αL cls + βL loc Among them, L cls is the classification loss, and L loc is the localization loss. α and β are weight coefficients used to balance the classification and localization losses.

3. A method for analyzing the compliance of safety helmet wearing based on HGNN according to claim 1, characterized in that, The step 2 comprises the following steps: Step 2.1: Build and train an end-to-end object detection framework. The loss function includes object confidence loss classification loss and bounding box regression loss These three parts; the total loss is the weighted sum of these three parts, and the expression is: Among them, the weights of the target confidence loss, classification loss, and bounding box regression loss are λ1, λ2, and λ3 respectively; L conf and L cls are calculated using the binary cross-entropy function, and the calculation formula is as follows: L n = -(y n log(δ(x n )) + (1 - y n ) log(1 - δ(x n )) where: x n is the score for predicting that the nth sample is a positive example, y n represents the label of the nth sample, δ represents the sigmoid function, L reg uses the CloU Loss function, and CloU is a method for calculating the bounding box regression loss. The calculation formula is as follows: Where: parameter A represents the ground truth box, B represents the predicted box, IoU is called the intersection over union, which represents the overlapping degree of the predicted box and the ground truth box, and ρ 2 (·) represents the Euclidean distance, b represents the coordinates of the center point of the predicted box, b gt represents the coordinates of the center point of the ground truth box, a represents the balance parameter, and v is the consistency of the aspect ratios of the ground truth box and the predicted box.

4. The helmet wearing standardization analysis method based on HGNN according to claim 1, wherein, The step 3 comprises the following steps: Step 3.1, in the helmet specification detection, the hypergraph calculation is referenced, the hypergraph G = (V, E) is defined by its vertex set V and hyperedge set E, the vertex set V is constructed, and five feature layers are extracted in the backbone to be divided into vertices of the hypergraph; Step 3.2: In order to fuse cross-level information and enhance the feature extraction capability of the basic network, a parallel convolutional coding structure is designed, which synergistically integrates three typical convolution variants: a 1×1 convolution for channel feature recalibration, a deformable convolution DCNV for effective spatial correlation information, and retains the C3K2 convolution module in the network, so that the model adjusts the feature extraction and fusion strategies for different inputs, and realizes systematic modeling of multi-level features from local to global.

5. The method for analyzing the standardization of safety helmet wearing based on HGNN according to claim 1, wherein, The step 4 comprises the following steps: Step 4.1: To model the neighborhood relationship in the hypergraph module, construct a hyperedge set E through a distance threshold λ. For each feature point x u , find all feature points with a distance less than ∈, and form a hyperedge with them and x v . The hyperedge e is represented as: e = {u|||x u -x v ||2 <∈, u ∈ V} Among them, ||·||2 represents the Euclidean norm, all such hyperedges constitute the hyperedge set E, and the incidence matrix H of the hypergraph G=(V,E) is defined as: Step 4.2: In the hypergraph convolutional network, as the number of layers increases, information is gradually lost during the propagation process. By introducing residual connections, the model directly transmits input information to the output, and uses spatial hypergraph convolution with residual connections. The calculation process is as follows: where N V (e) represents the vertex neighborhood of the hyperedge e, Θ e is a trainable parameter. For the feature matrix X of the given vertices and the adjacency matrix H of the hypergraph, assume D v and D e are the diagonal matrices of vertices and hyperedges respectively, and the hypergraph convolution is expressed as: Among them, Calculate the normalized adjacency matrix, indicating aggregating vertex features through hyperedges to capture the high-order relationships between vertices. Θ is a learnable parameter matrix used to transform the aggregated features. Then introduce X as a residual connection to retain the original feature information and avoid information loss.

6. The method for analyzing the standardization of safety helmet wearing based on HGNN according to claim 1, wherein The step 5 comprises the following steps: Step 5.1: Design a decoding network based on bidirectional cross-scale connections, and introduce an improved weighted bidirectional feature pyramid network BiFPN, whose mathematical expression is as follows: Among them, is the input feature of the l-th layer, is upsampling, is downsampling, w k is the learnable weight, σ is the Sigmoid normalization, is channel concatenation, is a 3×3 depthwise separable convolution; Step 5.2: Hypergraph-guided cross-layer aggregation: Perform cross-modal fusion on the high-order features Hyper output by HGNN and the features of the C3-C5 layers of the backbone network: Among them, is a linear transformation, and DCNv2 is a deformable convolution to enhance spatial adaptability; Step 5.3: Adaptive spatial reweighting module: Design a spatial attention weight map to dynamically adjust the feature response: Among which, A spatial (x, y) is the dynamic adjustment feature response of the spatial attention weight map, is the fused feature, W xy is RC 3 is the location-sensitive learning parameter to achieve pixel-level feature enhancement; Step 5.4: Dynamic prototype matching detection head: Construct a classification-regression joint output layer based on prototype learning: Among them, L proto is the loss function of the classification-regression joint output layer for prototype learning, s i ′ is the prediction result after improved non-maximum suppression, P+ / P- is the positive and negative sample prototype libraries, Δ is the boundary threshold, and the classification robustness is improved by dynamically updating the prototypes; Step 5.5: Improved non-maximum suppression: Propose differentiable NMS to solve the non-differentiable problem of traditional NMS: where k is a tuning parameter, and the prediction results of occluded targets are retained through a soft suppression mechanism.

Citation Information

Cited By

  • Fall behavior identification method based on hypergraph depth feature fusion

    CN121236821A