A method for detecting the number of people in a high-altitude working hanging basket based on an improved YOLOv9 model

By improving the YOLOv9 model, building the C2f_SE module, combining the HWD wavelet transform downsampling module, introducing SE and CBAM attention mechanisms, and adding DIoU loss function, the problem of low detection accuracy of the YOLOv9 model and difficulty in meeting the real-time monitoring needs is solved, and high-precision and high-efficiency number of people detection are achieved.

CN118942033BActive Publication Date: 2025-05-16INNER MONGOLIA UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410980835.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2025-05-16
Estimated Expiration
2044-07-19

AI Technical Summary

Technical Problem

In the prior art, the detection accuracy of the YOLOv9 model is low, and traditional monitoring equipment is difficult to meet the real-time monitoring and accuracy requirements for the number of people working at high altitudes.

Method used

By building an improved YOLOv9 model, including building a C2f_SE module, replacing some RepNCSPELAN4 module in the YOLOv9 backbone network, combining the HWD wavelet transform downsampling module, introducing SE and CBAM attention mechanisms, and adding DIoU loss functions to improve the detection accuracy and efficiency of the model.

Benefits of technology

It significantly improves the accuracy, accuracy and average accuracy of the number of people in the high-altitude operation hanging baskets, meets the needs of real-time monitoring, and improves the generalization ability and detection efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118942033B_ABST
    Figure CN118942033B_ABST
Patent Text Reader

Abstract

The invention discloses a method for detecting the number of people in a high-altitude working hanging basket based on an improved YOLOv9 model, belongs to the technical field of target detection, and comprises the following steps: step 1, constructing a data set of the number of people in the hanging basket, preprocessing and data enhancement, and dividing the data into a training set and a test set after manual annotation; step 2, building an improved YOLOv9 model framework, constructing a C2f_SE module and replacing the RepNCSPELAN4 modules of the third layer, the fifth layer and the seventh layer in the YOLOv9 backbone network; replacing the ADown downsampling modules of the sixth layer and the eighth layer of the backbone network with HWD_ADown; introducing an SE attention mechanism into the RepNCSPELAN4 module of the seventeenth layer of the Head part, and introducing a CBAM attention mechanism into the RepNCSPELAN4 module of the twenty-first layer; finally adding a DIoU loss function; step 3, training the improved YOLOv9 model using the training set, and detecting a data picture of the number of people in the hanging basket to be detected; step 4, using accuracy, precision, average precision and recall as evaluation indicators to perform performance evaluation on the detection result of step 3.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection, and in particular relates to a method for detecting the number of people in a high-altitude working hanging basket based on an improved YOLOv9 model. Background Art

[0002] High-altitude work baskets are widely used in high-altitude work tasks such as construction, maintenance, and cleaning. They provide a safe and efficient working platform that enables workers to work at heights. However, some workers have a fluke mentality in order to speed up the progress of construction work and do not follow the rules and regulations to perform illegal operations. In order to ensure the safety of people in the basket, it is necessary to monitor the number of people in the basket in real time to ensure that it is not overloaded and maintain proper working conditions. In traditional methods of monitoring the number of people in high-altitude work baskets, manual counting is often relied on. This method requires additional human resources, is not only inefficient, but also prone to errors. In addition, since high-altitude work baskets are often in a changing environment, traditional monitoring equipment often cannot meet the needs of real-time monitoring and accuracy. With the rapid development of computer vision and deep learning, image and video-based object detection methods have become key technologies for solving the problem of high-altitude work basket number monitoring. Among them, the YOLOv9 (You Only Look Once) model is an advanced object detection model with the advantages of real-time and accuracy. The YOLOv9 model uses an end-to-end training method to detect multiple targets in an image at the same time. It achieves real-time target detection by dividing the image into grids and predicting the location and category of the target in each grid cell, but the detection accuracy of the current YOLOv9 model needs to be further improved. Summary of the invention

[0003] The purpose of the present invention is to provide a method for detecting the number of people in a high-altitude working hanging basket based on an improved YOLOv9 model, so as to solve the problems in the prior art that the traditional YOLOv9 model has low detection accuracy, the monitoring equipment is difficult to meet the requirements of real-time monitoring and has poor accuracy.

[0004] To achieve the above object, the present invention provides a method for detecting the number of people in a high-altitude working hanging basket based on an improved YOLOv9 model, comprising the following steps:

[0005] Step 1: Construct a dataset of the number of people working in a suspended basket, perform preprocessing and data enhancement, and divide it into a training set and a test set after manual annotation;

[0006] Step 2: Build the improved YOLOv9 model framework, construct the C2f_SE module and replace the RepNCSPELAN4 modules in the third, fifth and seventh layers of the YOLOv9 backbone network; replace the ADown downsampling modules in the sixth and eighth layers of the backbone network with HWD_ADown; introduce the SE attention mechanism in the RepNCSPELAN4 module in the seventeenth layer of the Head part, and introduce the CBAM attention mechanism in the RepNCSPELAN4 module in the twenty-first layer; finally, add the DIoU loss function;

[0007] Step 3: Use the training set to train the improved YOLOv9 model, and detect the data image of the number of construction workers on the hanging basket to be detected;

[0008] Step 4: Use accuracy, precision, average precision, and recall as evaluation indicators to evaluate the performance of the detection results of step 3.

[0009] Preferably, the specific process of inputting the data image to be detected into Backbone is: first pass through 2 3×3 basic convolution modules Conv, then pass through a C2f_SE module, then pass through an ADown module, then pass through a C2f_SE module, then pass through a HWD_ADown downsampling module and a C2f_SE module, then pass through a HWD_ADown downsampling module, and finally pass through a RepNCSPELAN4 module for output.

[0010] Preferably, the specific process of step 1 is as follows:

[0011] Step 1.1, collect the data set of the number of people working on the hanging basket;

[0012] Step 1.2, using crawler technology to obtain images of the number of people working on the hanging basket in different scenarios, including images of hanging baskets in different scenarios such as outdoors, construction sites, and factories, as well as images of hanging basket construction;

[0013] Step 1.3: Use global feature extraction, binary feature comparison and the trained SSD model to clean the dataset and remove duplicate, damaged and irrelevant images;

[0014] Step 1.4, use Gaussian filtering, histogram equalization and normalization algorithms to perform image denoising and normalization operations on the data set;

[0015] Step 1.5: Each original image containing a fire extinguisher is expanded using image augmentation. The specific method of expansion is to rotate each image 180 degrees. Image augmentation is a data enhancement technique that can generate more training samples by transforming the image, such as rotating and flipping, which helps to improve the generalization ability of the model.

[0016] Step 1.6: Use LabelImg software to manually annotate the dataset in YOLO format and divide it into training set and test set according to 9:1.

[0017] Preferably, the specific process of constructing the C2f_SE module in step 2 is as follows:

[0018] Add the SE attention mechanism to the dimensions L (scale), S (space), and C (channel) of the feature map U. The function of the SE attention mechanism is expressed as:

[0019]

[0020] In the formula, H represents the height of the feature map, W represents the width of the feature map, and u c represents the feature map of the cth channel, i represents the vertical index of the feature map, j represents the horizontal index of the feature map, and F sq Represents the function of global average pooling of feature maps, generating a 1*1*C vector, where each channel C is represented by a weight vector S. The specific expression is as follows:

[0021] S=F ex (z,α)σ(g(z,α))=σ(α 2 δ(α 1 z))

[0022] In the formula, F ex represents the feature extraction function, z represents the input vector, α represents the weight matrix, σ represents the activation function, g represents the linear transformation before the activation function, α 2 represents the second layer weight matrix, δ represents the nonlinear activation function, α 1 represents the first layer weight matrix;

[0023] The feature map U is weighted by the weight vector S to obtain the final feature map. The specific expression is as follows:

[0024] F scale (u c ,s c )=s c u c

[0025] In the formula, F scale represents the feature map weighting function, s c Represents the c component of the weight vector.

[0026] Preferably, the specific content of the data picture to be detected passing through the C2f_SE module is as follows:

[0027] The data image to be detected is first split into two parts of features through the features processed by CBS. One part of the features is retained without any processing, and the other part is processed by several BottleNecks; each BottleNeck will be divided into two channels, one is to pass the processed features to the next BottleNeck, and the other is retained for the subsequent concat connection, and then all the features are fused after passing through 3 BottleNecks; each BottleNeck module includes a 1×1 convolution to reduce the number of channels and reduce the computational complexity, and then a 3×3 convolution is used to extract features on a smaller number of channels; the features processed by the BottleNeck module will be spliced ​​with the features directly used as output supplements to achieve feature fusion; finally, the fused features are processed by a convolution layer to generate the output feature map of the C2f_SE module, and the output feature map is used as the input of the subsequent layer to continue to be transmitted and processed in the network.

[0028] Preferably, the HWD_ADown module is formed by fusing the HWD wavelet transform to the ADown downsampling module, wherein the Adown downsampling process includes that the input feature map is first subjected to average pooling with a step size of 1, and then evenly divided into two groups, one group undergoes a convolution with a step size of 2; the other group undergoes a maximum pooling and then a convolution; finally, the two groups are stacked through the cat layer and then output to the Head part of the model.

[0029] Preferably, the specific expression of the DIoU loss function is as follows:

[0030] DIOU=1-IOU+d / c 2

[0031] In the formula, IOU represents the degree of overlap between two bounding boxes, d represents the Euclidean distance between the center point of the predicted box and the real box, and c represents the diagonal length of the minimum closed box covering the predicted box and the real box.

[0032] Preferably, in step 4, the specific expression for evaluating the performance of the detection result of step 3 using accuracy, precision, average precision, and recall as evaluation indicators is as follows:

[0033]

[0034]

[0035] Where Accuracy represents accuracy, Precision represents precision, MAP represents average precision, Recall represents recall, Epoch represents round, FP represents the number of negative predictions as positives, TP represents the number of positive predictions as positives, TN represents the number of negative predictions as negatives, and FN represents the number of positive predictions as negatives.

[0036] Therefore, the present invention adopts the above-mentioned method for detecting the number of people in a high-altitude working hanging basket based on the improved YOLOv9 model. First, the SE attention mechanism is integrated with the C2f module to construct a C2f_SE module, and the C2f_SE module is used to replace part of the RepNCSPELAN4 module in the YOLOv9 backbone network. At the same time, the downsampling module ADown in the YOLOv9 Backbone is combined with the HWD wavelet transform downsampling to form a new HWD_Adown downsampling module, thereby improving the target detection capability and efficiency in the backbone network, and integrating the SE attention mechanism and the CBAM attention mechanism into the Head, respectively, so as to increase the representation ability of the model; in addition, in order to improve the reasonable selection rate of the candidate box, the DIoU loss function is added, and the detection model is optimized so that it can better detect the position of the target.

[0037] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is the overall flow chart of the improved YOLOv9 model of the present invention;

[0039] Figure 2 It is a schematic diagram of the structure of the improved YOLOv9 provided by the present invention;

[0040] Figure 3 is a schematic diagram of a C2f module provided by the present invention;

[0041] Figure 4 Schematic diagram of the SE attention mechanism provided by the present invention;

[0042] Figure 5 1 is a comparison diagram of the results of the original YOLOv9 and improved YOLOv9 algorithm models provided by the present invention, wherein (a) is a PR diagram obtained by training the original YOLOv9, and (b) is a PR diagram obtained by training the improved YOLOv9;

[0043] Figure 6 The effect diagram provided by the present invention. DETAILED DESCRIPTION

[0044] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0045] See also Figure 1-6 , a method for detecting the number of people in a high-altitude working hanging basket based on an improved YOLOv9 model, comprising the following steps:

[0046] Step 1: Construct a dataset of the number of people working in the suspended basket, perform preprocessing and data enhancement, and divide it into training and test sets after manual annotation. The specific process is as follows:

[0047] Step 1.1, collect the data set of the number of people working on the hanging basket;

[0048] Step 1.2, using crawler technology to obtain images of the number of people working on the hanging basket in different scenarios, including images of hanging baskets in different scenarios such as outdoors, construction sites, and factories, as well as images of hanging basket construction;

[0049] Step 1.3: Use global feature extraction, binary feature comparison and the trained SSD model to clean the dataset and remove duplicate, damaged and irrelevant images;

[0050] Step 1.4, use Gaussian filtering, histogram equalization and normalization algorithms to perform image denoising and normalization operations on the data set;

[0051] Step 1.5: Each original image containing a fire extinguisher is expanded using image augmentation. The specific method of expansion is to rotate each image 180 degrees. Image augmentation is a data enhancement technique that can generate more training samples by transforming the image, such as rotating and flipping, which helps to improve the generalization ability of the model.

[0052] Step 1.6: Use LabelImg software to manually annotate the dataset in YOLO format and divide it into training set and test set according to 9:1.

[0053] Step 2, build the improved YOLOv9 model framework, construct the C2f_SE module and replace the RepNCSPELAN4 modules in the third, fifth and seventh layers of the YOLOv9 backbone network; replace the ADown downsampling modules in the sixth and eighth layers of the backbone network with HWD_ADown; introduce the SE attention mechanism in the RepNCSPELAN4 module in the seventeenth layer of the Head part, and introduce the CBAM attention mechanism in the RepNCSPELAN4 module in the twenty-first layer; this can improve the feature expression ability, improve the detection accuracy and improve the generalization of the model. Introducing the SE and CBAM attention mechanisms into the Head part of YOLOv9 can significantly improve the model's feature expression ability, detection accuracy, convergence speed and generalization ability, while maintaining the flexibility of the model to adapt to different application requirements; finally, add the DIoU loss function;

[0054] The specific process of building the C2f_SE module is as follows:

[0055] Add the SE attention mechanism to the dimensions L (scale), S (space), and C (channel) of the feature map U. The function of the SE attention mechanism is expressed as:

[0056]

[0057] In the formula, H represents the height of the feature map, W represents the width of the feature map, and u c represents the feature map of the cth channel, i represents the vertical index of the feature map, j represents the horizontal index of the feature map, and F sq Represents the function of global average pooling of feature maps, generating a 1*1*C vector, where each channel C is represented by a weight vector S. The specific expression is as follows:

[0058] S=F ex (z,α)σ(g(z,α))=σ(α 2 δ(α 1 z))

[0059] In the formula, F ex represents the feature extraction function, z represents the input vector, α represents the weight matrix, σ represents the activation function, g represents the linear transformation before the activation function, α 2 represents the second layer weight matrix, δ represents the nonlinear activation function, α 1 represents the first layer weight matrix;

[0060] The feature map U is weighted by the weight vector S to obtain the final feature map. The specific expression is as follows:

[0061] F scale (u c ,s c)=s c u c

[0062] In the formula, F scale represents the feature map weighting function, s c Represents the c component of the weight vector.

[0063] The specific contents of the data image to be detected passing through the C2f_SE module are as follows: the data image to be detected is first split into two parts of features through the features after CBS processing, one part of the features is retained without any processing, and the other part is processed by several BottleNecks; each BottleNeck will be divided into two channels, one is to pass the processed features to the next BottleNeck, and the other is retained for the subsequent concat connection, and then all the features are fused after passing through 3 BottleNecks; each BottleNeck module includes a 1×1 convolution to reduce the number of channels and reduce the computational complexity, and then a 3×3 convolution is used to extract features on a smaller number of channels; the features processed by the BottleNeck module will be spliced ​​with the features directly used as output supplements to achieve feature fusion; finally, the fused features are processed by a convolution layer to generate the output feature map of the C2f_SE module, and the output feature map is used as the input of the subsequent layer to continue to be transmitted and processed in the network.

[0064] The HWD_ADown module is formed by fusing the HWD wavelet transform to the ADown downsampling module, where the Adown downsampling process includes the input feature map first undergoing average pooling with a step size of 1, and then being evenly divided into two groups, one group undergoing a convolution with a step size of 2; the other group undergoing a maximum pooling and then a convolution; finally, the two groups are stacked through the cat layer and then output to the Head part of the model; in YOLOv9, the ADown module originally reduced the spatial resolution of the feature map through convolution and pooling operations. The HWD_ADown module achieves more efficient downsampling by replacing some of the convolution operations with HWD operations. Specifically, the HWD_ADown module retains some structures in the ADown module (such as block processing, splicing, etc.), but uses HWD operations in the downsampling part.

[0065] The specific expression of the DIoU loss function is as follows:

[0066] DIOU=1-IOU+d / c 2

[0067] In the formula, IOU represents the degree of overlap between two bounding boxes, d represents the Euclidean distance between the center point of the predicted box and the real box, and c represents the diagonal length of the minimum closed box covering the predicted box and the real box.

[0068] Step 3. Use the training set to train the improved YOLOv9 model, and detect the data image of the number of construction workers on the hanging basket to be detected; the specific process of inputting the data image to be detected into Backbone is: first pass through 2 3×3 basic convolution modules Conv, then pass through a C2f_SE module, then pass through an ADown module, then pass through a C2f_SE module, then pass through a HWD_ADown downsampling module and a C2f_SE module, then pass through a HWD_ADown downsampling module, and finally pass through a RepNCSPELAN4 module for output.

[0069] Step 4: Use accuracy, precision, average precision, and recall as evaluation indicators to evaluate the performance of the detection results of step 3. The specific expressions are as follows:

[0070]

[0071] Where Accuracy represents accuracy, Precision represents precision, MAP represents average precision, Recall represents recall, Epoch represents round, FP represents the number of negative predictions as positives, TP represents the number of positive predictions as positives, TN represents the number of negative predictions as negatives, and FN represents the number of positive predictions as negatives.

[0072] Example

[0073] The data for experimental training in this paper comes from web crawler pictures and real pictures taken at the construction site. The total amount of picture data set is 2000, 1800 for training set and 200 for test set.

[0074] The comparative experimental results are shown in Table 1:

[0075] Model MAP P R YOLOv9 70.4% 71.3% 78.7% Improved YOLOv9 81.8% 75.0% 83.1%

[0076] Combination Figure 5It can be seen that the average accuracy, precision and recall rate of the improved YOLOv9 model used in the present invention are better than those of the traditional YOLOv9 model; among them, Person 0.577 in Figure (a) indicates that the accuracy of the person is 57.7%. In this data set, if the accuracy is relatively low, it will not be recognized or missed. Person 0.741 in Figure (b) indicates that the accuracy of the person is 74.1%. The accuracy of Figure (b) on this data set is better than that of Figure (a). Aerial workplatform is an aerial work platform, which is a category name. All classes 0.704mAP@0.5 indicates the average accuracy, indicating that the average accuracy of the model in all categories is 0.704, which is calculated under the condition that the IoU threshold is 0.5. The higher this value is, the better the performance of the model is. Obviously, the performance of the improved YOLOv9 model is better than that of the traditional YOLOv9 model.

[0077] Therefore, the present invention adopts the above-mentioned method for detecting the number of people in a high-altitude working hanging basket based on the improved YOLOv9 model. First, the SE attention mechanism is integrated with the C2f module to construct a C2f_SE module, and the C2f_SE module is used to replace part of the RepNCSPELAN4 module in the YOLOv9 backbone network. At the same time, the downsampling module ADown in the YOLOv9 Backbone is combined with the HWD wavelet transform downsampling to form a new HWD_Adown downsampling module, thereby improving the target detection capability and efficiency in the backbone network, and integrating the SE attention mechanism and the CBAM attention mechanism into the Head, respectively, so as to increase the representation ability of the model; in addition, in order to improve the reasonable selection rate of the candidate box, the DIoU loss function is added, and the detection model is optimized so that it can better detect the position of the target.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A method for detecting the number of people in a high-altitude working hanging basket based on an improved YOLOv9 model, characterized in that: The following steps are involved: Step 1: Construct a dataset of the number of people working in a suspended basket, perform preprocessing and data enhancement, and divide it into a training set and a test set after manual annotation; Step 2: Build the improved YOLOv9 model framework, construct the C2f_SE module and replace the RepNCSPELAN4 modules in the third, fifth and seventh layers of the YOLOv9 backbone network; replace the ADown downsampling modules in the sixth and eighth layers of the backbone network with HWD_ADown; introduce the SE attention mechanism in the RepNCSPELAN4 module in the seventeenth layer of the Head part, and introduce the CBAM attention mechanism in the RepNCSPELAN4 module in the twenty-first layer; finally, add the DIoU loss function; Step 3: Use the training set to train the improved YOLOv9 model, and detect the data image of the number of construction workers on the hanging basket to be detected; Step 4: Use accuracy, precision, average precision, and recall as evaluation indicators to evaluate the performance of the detection results of step 3; The specific process of building the C2f_SE module in step 2 is as follows: Add the SE attention mechanism to the dimensions L, S, and C of the feature map U. The function of the SE attention mechanism is expressed as: In the formula, H represents the height of the feature map, W represents the width of the feature map, and u c represents the feature map of the cth channel, i represents the vertical index of the feature map, j represents the horizontal index of the feature map, and F sq Represents the function of global average pooling of feature maps, generating a 1*1*C vector, where each channel C is represented by a weight vector S. The specific expression is as follows: S=F ex (z,α)σ(g(z,α))=σ(α2δ(α1z)) In the formula, F ex represents the feature extraction function, z represents the input vector, α represents the weight matrix, σ represents the activation function, g represents the linear transformation before the activation function, α2 represents the second layer weight matrix, δ represents the nonlinear activation function, and α1 represents the first layer weight matrix; The feature map U is weighted by the weight vector S to obtain the final feature map. The specific expression is as follows: F scale (u c ,s c) =s c u c In the formula, F scale represents the feature map weighting function, s c Represents the c component of the weight vector.

2. According to claim 1, a method for detecting the number of people in a high-altitude working hanging basket based on an improved YOLOv9 model is characterized in that: The specific process of inputting the data image to be detected into Backbone is: first, it passes through two 3×3 basic convolution modules Conv, then through a C2f_SE module, then through an ADown module, then through a C2f_SE module, followed by a HWD_ADown downsampling module and a C2f_SE module, then through a HWD_ADown downsampling module, and finally through a RepNCSPELAN4 module for output.

3. According to claim 2, a method for detecting the number of people in a high-altitude working hanging basket based on an improved YOLOv9 model is characterized in that: The specific process of step 1 is as follows: Step 1.1, collect the data set of the number of people working on the hanging basket; Step 1.2, using crawler technology to obtain images of the number of people working on the hanging basket in different scenarios, including images of hanging baskets in different scenarios such as outdoors, construction sites, and factories, as well as images of hanging basket construction; Step 1.3: Use global feature extraction, binary feature comparison and the trained SSD model to clean the dataset and remove duplicate, damaged and irrelevant images; Step 1.4, use Gaussian filtering, histogram equalization and normalization algorithms to perform image denoising and normalization operations on the data set; Step 1.5: The original image is expanded by using image expansion. The specific expansion method is to rotate each image 180 degrees; Step 1.6: Use LabelImg software to manually annotate the dataset in YOLO format and divide it into training set and test set according to 9:

1.

4. According to claim 3, a method for detecting the number of people in a high-altitude working hanging basket based on an improved YOLOv9 model is characterized in that: The specific content of the data image to be detected passing through the C2f_SE module is as follows: The data image to be detected is first split into two parts of features through the features after CBS processing. One part of the features is retained without any processing, and the other part is processed by several BottleNecks; each BottleNeck will be divided into two channels, one is to pass the processed features to the next BottleNeck, and the other is retained for the subsequent concat connection, and then all the features are fused after passing through 3 BottleNecks; finally, the fused features are processed by a convolution layer to generate the output feature map of the C2f_SE module, and the output feature map is used as the input of the subsequent layer to continue to be transmitted and processed in the network.

5. According to claim 4, a method for detecting the number of people in a high-altitude working hanging basket based on an improved YOLOv9 model is characterized in that: The HWD_ADown module is formed by fusing the HWD wavelet transform to the ADown downsampling module. The Adown downsampling process includes that the input feature map is first subjected to average pooling with a step size of 1, and then evenly divided into two groups, one group undergoes a convolution with a step size of 2; the other group undergoes a maximum pooling and then a convolution; finally, the two groups are stacked through the cat layer and then output to the Head part of the model.

6. The method for detecting the number of people in a high-altitude working hanging basket based on an improved YOLOv9 model according to claim 5 is characterized in that: The specific expression of the DIoU loss function is as follows: DIOU=1-IOU+d / c 2 In the formula, IOU represents the degree of overlap between two bounding boxes, d represents the Euclidean distance between the center point of the predicted box and the real box, and c represents the diagonal length of the minimum closed box covering the predicted box and the real box.

7. The method for detecting the number of people in a high-altitude working hanging basket based on an improved YOLOv9 model according to claim 6 is characterized in that: In step 4, the specific expression for evaluating the performance of the detection results of step 3 using accuracy, precision, average precision, and recall as evaluation indicators is as follows: Where Accuracy represents accuracy, Precision represents precision, MAP represents average precision, Recall represents recall, Epoch represents round, FP represents the number of negative predictions as positives, TP represents the number of positive predictions as positives, TN represents the number of negative predictions as negatives, and FN represents the number of positive predictions as negatives.

Citation Information

Patent Citations

  • Infrared image pedestrian target detection method based on improved YOLOv5

    CN113688723A

  • Target detection method based on improved YOLOv5 model

    CN116597276A