Method, device and equipment for identifying safety helmets in loading and unloading area and storage medium
By annotating and classifying the image samples of the loading and unloading area of the distribution center, and building a RetinaNet network for safety helmet identification, it solves the problem that manual supervision is difficult to cover all aspects, real-time and intelligent monitoring of safety helmet wearing is achieved, and the efficiency and accuracy of safety production management are improved.
Patent Information
- Application Number
- CN202510064662.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-27
AI Technical Summary
In the loading and unloading area of the distribution center, due to the fast working pace and noisy environment, manual supervision is difficult to cover all the time, resulting in blind spots in the supervision of safety helmet wearing.
By annotating and classifying the image samples of the collected distribution center unloading area, a RetinaNet network is built, a hard hat recognition model is trained, and whether there are people in the image are not wearing hard hats are detected in real time, and whether an alarm is issued based on the detection results.
Real-time and intelligent monitoring of the wear of safety helmets in the loading and unloading area of the distribution center is realized, effectively solving the blind spot problem of manual supervision, and improving the efficiency and accuracy of safety production management.
Smart Images

Figure CN120047890A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a method, device, equipment and storage medium for identifying safety helmets in a loading and unloading area. Background Art
[0002] The loading and unloading area of the sorting center is a key link in logistics operations. Due to heavy machinery operations, cargo handling and aerial work involved, the staff face relatively high safety risks. Head protection is an important measure to prevent work injuries. Therefore, it is crucial to ensure that the staff wear safety helmets correctly.
[0003] However, in the loading and unloading area of the sorting center, due to the fast work rhythm and noisy environment, it is difficult for manual supervision to cover comprehensively, resulting in blind spots in the supervision of safety helmet wearing.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The main purpose of the present invention is to solve the problem of blind spots existing in the prior art when manually monitoring the wearing of safety helmets in the loading and unloading area of the sorting center.
[0006] The first aspect of the present invention provides a method for identifying safety helmets in a loading and unloading area, including: performing annotation and classification processing on the collected image samples of the loading and unloading area of the sorting center to obtain an image sample dataset of the loading and unloading area of the sorting center; dividing the image samples of the loading and unloading area of the sorting center into a training set and a test set according to a predetermined ratio, where both the training set and the test set include all the image samples of wearing safety helmets and the image samples of not wearing safety helmets completely; constructing an initial recognition model based on the RetinaNet network; training the initial recognition model through the training set, adjusting the parameters of the initial recognition model, and obtaining a safety helmet recognition model after testing through the test set; inputting the image of the loading and unloading area of the sorting center to be recognized obtained in real time into the safety helmet recognition model, outputting a detection result, and judging whether to issue an alarm according to the detection result.
[0007] Optionally, in the first implementation manner of the first aspect of the present invention, performing annotation and classification processing on the collected image samples of the loading and unloading area of the sorting center to obtain an image sample dataset of the loading and unloading area of the sorting center includes the steps of: regularly collecting image samples through a camera installed in the loading and unloading area of the sorting center to obtain image samples of the loading and unloading area of the sorting center; performing data augmentation processing on the image samples of the loading and unloading area of the sorting center to obtain expanded picture samples; performing annotation and classification processing on the expanded picture samples to obtain an image sample dataset of the loading and unloading area of the sorting center.
[0008] Optionally, in the second implementation manner of the first aspect of the present invention, data augmentation processing is performed on the image samples of the unloading area of the sorting center to obtain augmented picture samples, including the steps of: performing mirroring, rotation, scaling, cropping, translation, or Gaussian noise processing on the image samples of the unloading area of the sorting center to obtain deformed picture samples; performing random cropping processing and / or shear fusion processing on the image samples of the unloading area of the sorting center and the deformed picture samples to obtain augmented picture samples.
[0009] Optionally, in the third implementation manner of the first aspect of the present invention, annotation and classification processing are performed on the augmented picture samples to obtain a dataset of image samples of the unloading area of the sorting center, including the steps of: using the Labelme tool to perform feature annotation on the head areas of sorting personnel in the augmented picture samples to obtain image samples annotated with feature information; classifying the image samples according to the feature information to obtain image samples of all wearing safety helmets and image samples of not all wearing safety helmets, and the image samples of all wearing safety helmets and the image samples of not all wearing safety helmets constitute the dataset of image samples of the unloading area of the sorting center.
[0010] Optionally, in the fourth implementation manner of the first aspect of the present invention, an initial recognition model is constructed based on the RetinaNet network, including the steps of: building a RetinaNet network framework, where the RetinaNet network framework includes a feature extraction layer, a feature pyramid network, and a classification and regression sub-network, and among them, the feature extraction layer is composed of multiple convolutional layers and pooling layers; introducing a Self-Attention mechanism on each layer of the feature pyramid network, and enhancing the expression ability of multi-scale features by calculating the correlation between the features at each position, so as to construct an initial recognition model.
[0011] Optionally, in the fifth implementation manner of the first aspect of the present invention, the initial recognition model is trained by the training set, the parameters of the initial recognition model are adjusted, and after being tested by the test set, a safety helmet recognition model is obtained, including the steps of: using a random initialization method to assign initial values to all parameters in the initial recognition model; randomly selecting picture samples from the training set and inputting them into the initial recognition model for training, calculating the prediction result through forward propagation, and then calculating the loss value according to the Focal Loss function; using the backpropagation algorithm to update the parameters of the network to make the loss value gradually decrease; repeating the above training steps until a predetermined number of training rounds is reached to obtain a trained initial recognition model; using the test set to evaluate the trained initial recognition model, and adjusting the hyperparameters according to the evaluation result to obtain an optimized safety helmet recognition model.
[0012] Optionally, in the sixth implementation manner of the first aspect of the present invention, calculating the loss value according to the Focal Loss function includes the steps of: using focal loss as the classification loss function, and measuring the difference between the class probability predicted by the initial recognition model and the true class through the classification loss function. The calculation formula of Focal loss is as follows: L_cls = -α(1 - p_t)^γ log(p_t), where p_t is the class probability predicted by the initial recognition model, and α and γ are hyperparameters used to adjust the weights of different difficult and easy samples. When p_t is close to 1, (1 - p_t)^γ will become smaller, making the loss value of this sample smaller; conversely, when p_t is close to 0, (1 - p_t)^γ will become larger, making the loss value of this sample larger. Using smooth L1 loss as the regression loss function, and measuring the difference between the predicted bounding box position of the initial recognition model and the true bounding box position through the regression loss function. The calculation formula of Smooth L1 loss is as follows: L_reg = 0.5 * (IoU - 1)^2 * 1 if |IoU - 1| < 1, otherwise L_reg = |IoU - 1| - 0.5, where IoU represents the IOU value between the predicted bounding box and the true bounding box. When the IoU value is close to 1, it means that the predicted box is very close to the true box, and the loss value is smaller at this time; when the IoU value deviates from 1, the loss value will gradually increase.
[0013] The second aspect of the present invention provides a safety helmet recognition device for the loading and unloading area, including: a labeling and classification module for labeling and classifying the collected image samples of the loading and unloading area of the sorting center to obtain an image sample dataset of the loading and unloading area of the sorting center; a data division module for dividing the image samples of the loading and unloading area of the sorting center into a training set and a test set according to a predetermined ratio, and both the training set and the test set include all image samples of wearing safety helmets and image samples of not wearing safety helmets completely; a model construction module for constructing an initial recognition model based on the RetinaNet network; a training module for training the initial recognition model through the training set, adjusting the parameters of the initial recognition model, and obtaining a safety helmet recognition model after testing through the test set; an identification module for inputting the real-time acquired image of the loading and unloading area of the sorting center to be recognized into the safety helmet recognition model, outputting a detection result, and judging whether to issue an alarm according to the detection result.
[0014] Optionally, in the first implementation manner of the second aspect of the present invention, the annotation classification module includes: an acquisition unit, configured to regularly collect image samples through a camera installed in the loading and unloading area of the sorting center to obtain image samples of the loading and unloading area of the sorting center; an amplification unit, configured to perform data amplification processing on the image samples of the loading and unloading area of the sorting center to obtain expanded picture samples; a classification unit, configured to perform annotation and classification processing on the expanded picture samples to obtain an image sample dataset of the loading and unloading area of the sorting center.
[0015] Optionally, in the second implementation manner of the second aspect of the present invention, the amplification unit includes: a deformation sub-unit, configured to perform mirroring, rotation, scaling, cropping, translation or Gaussian noise processing on the image samples of the loading and unloading area of the sorting center to obtain deformed picture samples; a transformation sub-unit, configured to perform random cropping processing and / or shear fusion processing on the image samples of the loading and unloading area of the sorting center and the deformed picture samples to obtain expanded picture samples.
[0016] Optionally, in the third implementation manner of the second aspect of the present invention, the classification unit includes: an annotation sub-unit, configured to perform feature annotation on the head area of the sorting personnel in the expanded picture samples through the Labelme tool to obtain image samples annotated with feature information; a classification sub-unit, configured to classify the image samples according to the feature information to obtain image samples of all wearing safety helmets and image samples of not all wearing safety helmets, and the image samples of all wearing safety helmets and the image samples of not all wearing safety helmets constitute the image sample dataset of the loading and unloading area of the sorting center.
[0017] Optionally, in the fourth implementation manner of the second aspect of the present invention, the model construction module includes: a building unit, configured to build a RetinaNet network framework, where the RetinaNet network framework includes a feature extraction layer, a feature pyramid network, and a classification and regression sub-network, and among them, the feature extraction layer is composed of multiple convolutional layers and pooling layers; an introduction unit, configured to introduce a Self-Attention mechanism on each layer of the feature pyramid network, and enhance the expression ability of multi-scale features by calculating the correlation between feature positions at each position, so as to build an initial recognition model.
[0018] Optionally, in the fifth implementation manner of the second aspect of the present invention, the training module includes: an initialization unit for assigning initial values to all parameters in the initial recognition model using a random initialization method; a calculation unit for randomly selecting picture samples from the training set and inputting them into the initial recognition model for training, calculating the prediction result through forward propagation, and then calculating the loss value according to the Focal Loss function; an update unit for updating the parameters of the network using the backpropagation algorithm to gradually reduce the loss value; a repetition unit for repeating the above training steps until a predetermined number of training rounds is reached to obtain a trained initial recognition model; and an evaluation unit for evaluating the trained initial recognition model using the test set and adjusting the hyperparameters according to the evaluation results to obtain an optimized safety helmet recognition model.
[0019] Optionally, in the sixth implementation manner of the second aspect of the present invention, the calculation unit includes: a classification loss calculation sub-unit for using focal loss as the classification loss function to measure the difference between the class probability predicted by the initial recognition model and the true class through the classification loss function. The calculation formula of Focal loss is as follows: L_cls = -α(1 - p_t)^γlog(p_t), where p_t is the class probability predicted by the initial recognition model, and α and γ are hyperparameters used to adjust the weights of different difficult and easy samples. When p_t is close to 1, (1 - p_t)^γ becomes smaller, making the loss value of this sample smaller; conversely, when p_t is close to 0, (1 - p_t)^γ becomes larger, making the loss value of this sample larger; a regression loss calculation sub-unit for using smooth L1 loss as the regression loss function to measure the difference between the predicted bounding box position of the initial recognition model and the true bounding box position through the regression loss function. The calculation formula of Smooth L1 loss is as follows: L_reg = 0.5*(IoU - 1)^2*1 if |IoU - 1| < 1, otherwise L_reg = |IoU - 1| - 0.5, where IoU represents the IOU value between the predicted bounding box and the true bounding box. When the IoU value is close to 1, it means that the predicted box is very close to the true box, and the loss value is smaller at this time; when the IoU value deviates from 1, the loss value gradually increases.
[0020] The third aspect of the present invention provides a safety helmet recognition device for a loading and unloading area, including: a memory and at least one processor, where computer-readable instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor calls the computer-readable instructions in the memory to enable the safety helmet recognition device for the loading and unloading area to execute each step of the safety helmet recognition method for the loading and unloading area as described above.
[0021] A fourth aspect of the present invention provides a computer-readable storage medium storing computer-readable instructions that, when run on a computer, cause the computer to execute the steps of the above-described safety helmet recognition method for the loading and unloading area.
[0022] Beneficial effects: In the technical solution of the present invention, the collected image samples of the loading and unloading area of the sorting center are labeled and classified to obtain an image sample dataset of the loading and unloading area of the sorting center; the image samples of the loading and unloading area of the sorting center are divided into a training set and a test set according to a predetermined ratio, and both the training set and the test set include all the image samples of workers wearing safety helmets and the image samples of workers not wearing safety helmets completely; an initial recognition model is constructed based on the RetinaNet network; the initial recognition model is trained by the training set, the parameters of the initial recognition model are adjusted, and after being tested by the test set, a safety helmet recognition model is obtained; the image of the loading and unloading area of the sorting center to be recognized obtained in real time is input into the safety helmet recognition model, and a detection result is output, and whether to issue an alarm is judged according to the detection result. The present invention provides a complete solution from sample collection and labeling, model construction and training to actual deployment and application, realizes real-time and intelligent monitoring of the wearing situation of safety helmets in the loading and unloading area of the sorting center, effectively solves the blind area problem of manual supervision, improves the efficiency and accuracy of safety production management, and has stronger practicability and innovation compared with some existing safety helmet detection technologies. Description of the Drawings
[0023] Figure 1 The first flow chart of the safety helmet recognition method for the loading and unloading area provided by the embodiment of the present invention;
[0024] Figure 2 The second flow chart of the safety helmet recognition method for the loading and unloading area provided by the embodiment of the present invention;
[0025] Figure 3 The third flow chart of the safety helmet recognition method for the loading and unloading area provided by the embodiment of the present invention;
[0026] Figure 4 The fourth flow chart of the safety helmet recognition method for the loading and unloading area provided by the embodiment of the present invention;
[0027] Figure 5 The fifth flow chart of the safety helmet recognition method for the loading and unloading area provided by the embodiment of the present invention;
[0028] Figure 6 The sixth flow chart of the safety helmet recognition method for the loading and unloading area provided by the embodiment of the present invention;
[0029] Figure 7 A structural schematic diagram of the safety helmet recognition device for the loading and unloading area provided by the embodiment of the present invention;
[0030] Figure 8 Another structural schematic diagram of the safety helmet recognition device for the loading and unloading area provided by the embodiment of the present invention;
[0031] Figure 9 Structural schematic diagram of the safety helmet recognition device for the loading and unloading area provided by the embodiment of the present invention. Specific implementation manners
[0032] The embodiment of the present invention provides a safety helmet recognition method, device, equipment and storage medium for the loading and unloading area. The image samples of the sorting center's loading and unloading area collected are labeled and classified to obtain the image sample dataset of the sorting center's loading and unloading area; the image samples of the sorting center's loading and unloading area are divided into a training set and a test set according to a predetermined ratio, and both the training set and the test set include all the image samples of workers wearing safety helmets and the image samples of workers not wearing safety helmets completely; an initial recognition model is constructed based on the RetinaNet network; the initial recognition model is trained by the training set, the parameters of the initial recognition model are adjusted, and after being tested by the test set, a safety helmet recognition model is obtained; the image of the to-be-recognized sorting center's loading and unloading area obtained in real time is input into the safety helmet recognition model, and a detection result is output, and whether to issue an alarm is judged according to the detection result. The present invention provides a complete solution from sample collection and labeling, model construction and training to actual deployment and application, realizes real-time and intelligent monitoring of the wearing situation of safety helmets in the loading and unloading area of the sorting center, effectively solves the blind area problem of manual supervision, and improves the efficiency and accuracy of safety production management.
[0033] The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present invention and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.
[0034] For ease of understanding, the specific process of the embodiment of the present invention is described below. Please refer to Figure 1 , the first embodiment of the safety helmet recognition method for the loading and unloading area in the embodiment of the present invention includes:
[0035] S100. Label and classify the collected image samples of the unloading area of the sorting center to obtain the image sample dataset of the unloading area of the sorting center.
[0036] Specifically, in this embodiment, surveillance cameras are reasonably installed in the loading and unloading area of the sorting center to ensure that their positions can fully cover the entire loading and unloading area and avoid blind spots in the viewing angle. For example, the number, height, and angle of the cameras can be determined according to factors such as the layout of the loading and unloading area, the stacking height of the goods, and the operation process. In a loading and unloading area that is 50 meters long, 30 meters wide, and has an average stacking height of 5 meters for the goods, it may be necessary to install 4 - 6 high-definition cameras, which are respectively arranged at the four corners and the middle position of the area to ensure clear monitoring of the operations in each corner and at different height levels. After obtaining the image samples of the unloading area of the sorting center, the collected image samples of the unloading area of the sorting center can be labeled and classified manually or by machine, so as to obtain the image sample dataset of the unloading area of the sorting center, including the original images, annotation information, classification labels, etc., and by ensuring the quality and diversity of the dataset, in order to train a more accurate model subsequently.
[0037] S200. Divide the image samples of the unloading area of the sorting center into a training set and a test set according to a predetermined ratio, and both the training set and the test set include all the image samples of workers wearing safety helmets and the image samples of workers not all wearing safety helmets.
[0038] In this embodiment, since the initially constructed initial recognition model needs to be trained with a large amount of data and verified before an accurate safety helmet recognition model can be obtained. Therefore, in this embodiment, the image samples of the unloading area of the sorting center need to be divided into a training set and a test set according to a predetermined ratio. In order to ensure that there is an adequate sample dataset for training the initial recognition model, the number of the training set should be greater than the number of the test set. Preferably, the ratio of the number of the training set to the number of the test set is 8:2 - 6:4. Among them, both the training set and the test set include all the image samples of workers wearing safety helmets and the image samples of workers not all wearing safety helmets. The image samples of workers all wearing safety helmets refer to that all the sorting personnel in the image samples of the unloading area of the sorting center wear safety helmets, and the image samples of workers not all wearing safety helmets refer to that at least one person in the image samples of the unloading area of the sorting center does not wear a safety helmet.
[0039] S300. Build an initial recognition model based on the RetinaNet network.
[0040] Specifically, RetinaNet proposes a loss function called Focal Loss, which is specifically designed to address the common class imbalance problem in object detection. In object detection tasks, the background region usually occupies most of the image, while the target objects (such as the safety helmets we want to detect) only occupy a small number of pixels, resulting in a serious imbalance between the foreground and background classes. In this case, with the traditional cross-entropy loss function, the model tends to predict a large number of background regions and ignore the small number of target objects. Focal Loss dynamically adjusts the cross-entropy loss, reducing the weight of easily classified samples (such as the background) and increasing the weight of difficult-to-classify samples (such as small or occluded targets). Specifically, for samples with correct classification and high probability, their loss values are reduced; while for samples with incorrect classification or low probability, the loss values are relatively increased. In this way, the model will pay more attention to those targets that are difficult to distinguish during the training process, thereby improving the detection ability for small and occluded targets. In the scenario of safety helmet recognition in the loading and unloading area of the sorting center, the number of people not wearing safety helmets may be relatively small compared to the background and people wearing safety helmets in the entire image, forming a situation of class imbalance. Therefore, in this embodiment, by using the Focal Loss of the RetinaNet network, the model can better focus on the minority class of people not wearing safety helmets, reduce the missed detection caused by class imbalance, and improve the accuracy of safety helmet wearing detection.
[0041] S400. Train the initial recognition model using the training set, adjust the parameters of the initial recognition model, and after testing with the test set, obtain the safety helmet recognition model.
[0042] S500. Input the image of the loading and unloading area of the sorting center to be recognized obtained in real time into the safety helmet recognition model, output the detection result, and determine whether to issue an alarm according to the detection result.
[0043] In this embodiment, after training and passing the test of the initial recognition model using the training set and the test set, the safety helmet recognition model is obtained. Based on this model, the image of the loading and unloading area of the sorting center to be recognized can be detected in real time and quickly, so as to obtain whether the image sample of the loading and unloading area of the sorting center to be recognized is a sample of all people wearing safety helmets. If the detection result is that it is not a sample of all people wearing safety helmets, an alarm will be issued to remind the sorting personnel not wearing safety helmets to wear safety helmets in time. If the detection result is a sample of all people wearing safety helmets, no alarm will be issued. The present invention provides a complete solution from sample collection and annotation, model construction and training to actual deployment and application, realizing real-time and intelligent monitoring of the safety helmet wearing situation in the loading and unloading area of the sorting center, effectively solving the blind area problem of manual supervision, and improving the efficiency and accuracy of safety production management.
[0044] Please refer to Figure 2 , the second embodiment of the safety helmet recognition method in the loading and unloading area in the embodiments of the present invention includes:
[0045] S110. Regularly collect image samples through a camera installed in the loading and unloading area of the sorting center to obtain image samples of the loading and unloading area of the sorting center;
[0046] S120. Perform data augmentation processing on the image samples of the loading and unloading area of the sorting center to obtain augmented picture samples;
[0047] S130. Perform annotation and classification processing on the augmented picture samples to obtain an image sample dataset of the loading and unloading area of the sorting center.
[0048] Specifically, in this embodiment, image samples are regularly collected through a camera installed in the loading and unloading area of the sorting center. The collection time should cover different time periods (such as day, night, early morning, etc.) and different lighting conditions (sunny, cloudy, rainy, direct strong light, shadow area, etc.) to ensure the diversity of the image samples of the loading and unloading area of the sorting center. For example, for a sorting center with non-stop operation for 24 hours, image samples can be automatically collected every 1-2 hours and continuously collected for one month to obtain sufficient and rich sample data, including images under various weather conditions and different operation busy levels; since the number of initially obtained image samples of the loading and unloading area of the sorting center is limited and is not conducive to training the constructed network, it is necessary to perform data augmentation processing on the obtained image samples of the loading and unloading area of the sorting center through data augmentation to obtain more augmented picture samples; then perform annotation processing and classification processing with special marks on these augmented picture samples, and place the processed augmented picture samples in the corresponding directories to form an image sample dataset of the loading and unloading area of the sorting center. In this embodiment, a machine learning or deep learning model is trained through a large number of annotated and classified image samples of the loading and unloading area of the sorting center, which can significantly improve the recognition accuracy of whether the sorting personnel wear safety helmets; at the same time, data augmentation processing enables the model to come into contact with more diverse image samples during the training process, so as to maintain a high recognition ability when facing various complex situations in actual applications.
[0049] Please refer to Figure 3 , the third embodiment of the safety helmet recognition method in the loading and unloading area in the embodiments of the present invention includes:
[0050] S121. Perform mirroring, rotation, scaling, cropping, translation or Gaussian noise processing on the image samples of the loading and unloading area of the sorting center to obtain deformed picture samples;
[0051] S122. Perform random cropping processing and / or shear fusion processing on the image samples of the loading and unloading area of the sorting center and the deformed picture samples to obtain augmented picture samples.
[0052] Specifically, in this embodiment, first, the initial image samples of the unloading area of the sorting center are processed by common data augmentation means such as mirroring, rotation, scaling, cropping, translation, or Gaussian noise to generate new pictures. Although this processing method results in new pictures for the network model and can achieve the purpose of expanding the dataset, the deformed picture samples after such processing do not change the labels of the original pictures, and the training effect is not good. Therefore, in this embodiment, these deformed picture samples and the initial image samples of the unloading area of the sorting center are further subjected to random cropping processing and / or shear fusion processing. In this way, the labels of the expanded picture samples obtained may be different from the labels of the original pictures, which is beneficial to the training of the network model and results in a network model with better prediction performance. Specifically, the random cropping processing (RandCrop) method refers to randomly cropping the image samples to increase the diversity of the data. Random cropping can not only simulate different shooting angles and distance changes but also enable the model to learn the features of the target in different local regions. For example, image regions of different sizes (such as 0.8 times, 0.6 times, 0.4 times, etc. of the original image) and positions are randomly cropped from an original image, and the cropped images should maintain the ratio of the original image to avoid deformation. After performing the RandCrop operation on 1000 original sample images, 3000 - 5000 enhanced images can be obtained, enriching the training data. The shear fusion processing (Mixup) method is to linearly combine two different samples to generate new samples. The specific operation is to randomly select two images and their corresponding labels and linearly combine the images and labels according to a certain ratio (such as 0.5) to obtain new training samples. This helps the model to generalize better and reduce overfitting. For example, there are image A and image B, and their labels are [1, 0] (indicating wearing a safety helmet) and [0, 1] (indicating not wearing a safety helmet) respectively. After the Mixup operation, a new image C = 0.5 * A + 0.5 * B is generated, and its label is [0.5, 0.5]. The model needs to judge this mixed sample during the learning process, thereby enhancing its adaptability to different sample combinations. In this embodiment, the RandCrop and Mixup data augmentation methods are combined and used for the image samples of the unloading area of the sorting center and the deformed picture samples, which can more effectively increase the diversity of the data and enable the model to better adapt to the safety helmet detection task in different scenarios. This data augmentation strategy can improve the generalization ability of the model and reduce the risk of overfitting compared with a single data augmentation method or the case where no data augmentation is used, thereby obtaining more stable and accurate detection results in practical applications.
[0053] Please refer to Figure 4 , the fourth embodiment of the method for identifying safety helmets in the loading and unloading area of express parcels in the embodiment of the present invention includes:
[0054] S131. Feature-label the head regions of sorting personnel in the augmented picture samples using the Labelme tool to obtain picture samples labeled with feature information;
[0055] S132. Classify the picture samples according to the feature information to obtain picture samples of all personnel wearing safety helmets and picture samples of those not all wearing safety helmets. The picture samples of all personnel wearing safety helmets and those not all wearing safety helmets constitute the picture sample dataset of the loading and unloading area of the sorting center.
[0056] In this embodiment, Labelme is an open-source image annotation tool that supports various annotation methods such as polygons, rectangles, and circles and is suitable for complex image annotation tasks. In this embodiment, the labeling software Labelme is used to feature-label the head regions of sorting personnel in the augmented picture samples to obtain picture samples labeled with feature information, accurately distinguishing between the two states of wearing a safety helmet and not wearing a safety helmet. During the annotation process, carefully mark the head regions of sorting personnel and assign different labels to different states. For example, "hat" indicates wearing a safety helmet, and "no_hat" indicates not wearing a safety helmet. If among a batch of 1,000 sample pictures, approximately 600 are in the state of all wearing safety helmets and 400 are in the state of not all wearing safety helmets, the annotator needs to ensure that the annotation of each picture is accurate to ensure the effectiveness of subsequent model training. Through automated or semi-automated image annotation and classification techniques, this embodiment can significantly improve the image recognition speed and accuracy and reduce manual intervention and misjudgment.
[0057] Please refer to Figure 5 , the fifth embodiment of the safety helmet recognition in the loading and unloading area in the embodiment of the present invention includes:
[0058] S310. Build a RetinaNet network framework. The RetinaNet network framework includes a feature extraction layer, a feature pyramid network, and a classification and regression sub-network. Among them, the feature extraction layer is composed of multiple convolutional layers and pooling layers;
[0059] S320. Introduce the Self-Attention mechanism on each layer of the feature pyramid network, enhance the expression ability of multi-scale features by calculating the correlation between feature positions at each location, and thus build an initial recognition model.
[0060] The RetinaNet network adopted in this embodiment has unique advantages in object detection, especially suitable for detecting small objects (safety helmets) in complex environments such as the loading and unloading areas of distribution centers. By optimizing its parameters and introducing improvement measures such as the Self-Attention mechanism, the detection accuracy and performance of the model can be further improved, and it has better performance than traditional object detection models (such as Faster R-CNN, YOLO, etc.) in the safety helmet detection scenario. Specifically, a RetinaNet network is built, which mainly includes a feature extraction layer, a Feature Pyramid Network (FPN), and a classification and regression sub-network part. Among them, the basic network architecture of the feature extraction layer can be a backbone network for feature extraction such as ResNet, VGG, etc. Taking ResNet as an example, it extracts the features of the image through a series of convolutional layers and residual blocks. Assuming ResNet-50 is selected, it contains multiple convolutional layers and 4 residual blocks (layer1, layer2, layer3, layer4 respectively), and each residual block is composed of multiple convolutional layers. The input image first undergoes initial feature extraction through the initial convolutional layer, and then passes through each residual block in turn. As the network depth increases, the degree of abstraction of the features gradually improves, and the extracted features contain various semantic information of the image, such as edges, textures, shapes, etc. These features will serve as the basis for subsequent processing. During the feature extraction process, the network will automatically learn the feature representations of different parts of the image. For example, for the features of a safety helmet, they may include information such as its shape, color, texture, etc., and these features will be further processed and utilized in subsequent network layers.
[0061] Build a Feature Pyramid Network (FPN) based on the feature extraction layer. The purpose of FPN is to fuse feature information at different scales to adapt to the detection of objects of different sizes. Obtain feature maps with different resolutions from the last few stages of the feature extraction layer (such as layer2, layer3, layer4 in ResNet-50), denoted as C2, C3, C4 respectively. Their resolutions decrease in turn (for example, the resolution of C2 is 1 / 4 of the input image, C3 is 1 / 8, and C4 is 1 / 16), but the semantic information gradually becomes richer. First, perform a 1x1 convolution operation on C4 to reduce its number of channels (for example, reduce the 2048 channels to 256 channels) to obtain P4. Then, upsample P4 (such as doubling its resolution using the nearest neighbor interpolation method) to make its resolution the same as that of C3. Next, add the upsampled P4 and C3 element by element to obtain the fused feature map F3, and then perform a 1x1 convolution operation on F3 to further adjust the features to obtain P3. Similarly, upsample P3 and fuse it with C2 to obtain P2. Finally, in order to obtain richer multi-scale features, P5 can also be obtained by downsampling on the basis of P2 (such as reducing the resolution by half using the max pooling operation). In this way, FPN constructs a pyramid structure containing features at different scales, which can effectively detect safety helmet targets of different sizes. The FPN in RetinaNet can effectively fuse multi-scale feature information. In a complex scene such as the loading and unloading area of a sorting center, the images captured by the camera may contain objects (workers and safety helmets) at different distances and sizes. FPN can fuse the features extracted from different stages of the backbone network to generate feature maps with different resolutions, enabling the model to simultaneously focus on objects of different scales. For example, for a small safety helmet in the distance, the high-level feature map has stronger semantic information and can better identify its category; while the low-level feature map has a higher resolution, which helps to accurately locate the position of the safety helmet. In this way, RetinaNet has good performance in small object detection and can more accurately detect small safety helmets in the image, reducing the situation of missed detection. For example, in the surveillance video, even if the worker is far from the camera and the safety helmet occupies fewer pixels in the image, RetinaNet can still detect it well.
[0062] On each layer (P2 - P5) of the FPN, a classification and regression sub - network is connected respectively. The classification sub - network is used to predict the class of the target at each feature point (i.e., whether it is a safety helmet), and the regression sub - network is used to predict the location information of the target (such as the coordinates of the bounding box). The classification sub - network usually consists of multiple convolutional layers and fully - connected layers. Taking P2 as an example, the output feature map undergoes feature transformation through a series of convolutional layers, then is converted into a fixed - length vector by a global average pooling layer, and finally a fully - connected layer is connected to output the probabilities of each feature point belonging to different classes (such as the probabilities of wearing a safety helmet and not wearing a safety helmet). The structure of the regression sub - network is similar, but it outputs the offset of the target location, which is used to determine the accurate location of the safety helmet in the image. During the training process, by calculating the loss between the prediction result and the true label (such as cross - entropy loss for classification and Smooth L1 loss for regression), and using the backpropagation algorithm to adjust the parameters of the network, the model gradually learns accurate classification and regression capabilities
[0063] Introduce the Self-Attention module at each layer of the Feature Pyramid Network (FPN) (such as P2 - P5). For the feature map of each layer, assuming its shape is (H, W, C) (where H represents height, W represents width, and C represents the number of channels), first perform a dimensionality reduction operation on it through a 1x1 convolutional layer to obtain a feature map with the shape of (H, W, C') (C' < C), denoted as Q (Query), K (Key), and V (Value). These three matrices represent the query, key, and value respectively, which are important concepts in the Self-Attention mechanism. Then calculate the attention weight matrix. Perform a matrix multiplication operation on the transpose of Q and K (Q * K^T) to obtain a matrix with the shape of (H, W, HW). Then scale this matrix (for example, divide it by the square root of C'), and then perform normalization processing through the softmax function to obtain the final attention weight matrix, whose shape is still (H, W, HW). This attention weight matrix represents the correlation between each position in the feature map and all other positions. The larger the value, the stronger the correlation. For example, for the head region where a safety helmet is located, the attention weights of the surrounding pixels will be relatively high because they are semantically closely related to the detection of the safety helmet. Perform a matrix multiplication operation on the obtained attention weight matrix and V (attention weight matrix * V) to obtain a weighted feature map with the shape of (H, W, C'). This process actually re-weights the original features according to the attention weights, enabling the model to pay more attention to the regional features related to the safety helmet detection. For example, if the attention weight of a certain region is high, then in the weighted summation process, the feature values of this region will be amplified, thus having a greater impact on the model's decision-making in subsequent processing. Finally, perform a dimensionality increase operation on the weighted feature map through a 1x1 convolutional layer to restore the number of channels to the number of channels of the original feature map (i.e., C), obtaining the feature map processed by the Self-Attention mechanism for subsequent classification and regression operations. Throughout the process, the Self-Attention mechanism can dynamically adjust the model's attention degree to different regional features, thereby improving the accuracy of safety helmet detection. Especially in complex backgrounds and multi-object scenarios, it can better capture the key features of the safety helmet, reducing the cases of false detection and missed detection. At the same time, due to the relatively complex calculation process of Self-Attention, after its introduction, it is necessary to appropriately adjust the depth and width of the network according to the actual situation to balance the computational complexity and model performance, ensuring the improvement of the detection effect without significantly increasing the consumption of computing resources.
[0064] Please refer to Figure 6 , the sixth embodiment of the safety helmet recognition in the loading and unloading area in the embodiment of the present invention includes:
[0065] S410. Assign initial values to all parameters in the initial recognition model using a random initialization method;
[0066] S420. Randomly select picture samples from the training set and input them into the initial recognition model for training. Calculate the prediction result through forward propagation, and then calculate the loss value according to the Focal Loss function;
[0067] S430. Use the backpropagation algorithm to update the parameters of the network to gradually reduce the loss value;
[0068] S440. Repeat the above training steps until a predetermined number of training epochs is reached to obtain a trained initial recognition model;
[0069] S450. Evaluate the trained initial recognition model using the test set, and adjust the hyperparameters according to the evaluation results to obtain an optimized safety helmet recognition model.
[0070] Specifically, before model training, the random initialization method is used to assign initial values to all parameters in the initial recognition model. For example, set the initial number of iterations (such as 5000 times), the initial learning rate (such as 10^-5), and batch_size (such as 300), as well as set the dataset path parameters and class parameters, and use parallel GPUs for training; randomly select a batch of images and corresponding bounding box annotation data from the training set and input them into the model for training; after the processing of the feature extraction and prediction branches, obtain the predicted class probabilities and bounding box positions; calculate the classification loss and regression loss according to the prediction results and the true annotations. The classification loss uses focal loss, and the regression loss uses smooth L1 loss. Among them, focal loss is used as the classification loss function, and the difference between the class probability predicted by the initial recognition model and the true class is measured through the classification loss function. The calculation formula of Focal loss is as follows: L_cls = -α(1 - p_t)^γlog(p_t), where p_t is the class probability predicted by the initial recognition model, and α and γ are hyperparameters used to adjust the weights of different difficult and easy samples. When p_t is close to 1, (1 - p_t)^γ will become smaller, making the loss value of this sample smaller; conversely, when p_t is close to 0, (1 - p_t)^γ will become larger, making the loss value of this sample larger; smooth L1 loss is used as the regression loss function, and the difference between the bounding box position predicted by the initial recognition model and the true bounding box position is measured through the regression loss function. The calculation formula of Smooth L1 loss is as follows: L_reg = 0.5*(IoU - 1)^2*1 if |IoU - 1| < 1, otherwise L_reg = |IoU - 1| - 0.5, where IoU represents the IOU value between the predicted bounding box and the true bounding box. When the IoU value is close to 1, it means that the predicted box is very close to the true box, and the loss value is smaller at this time; when the IoU value deviates from 1, the loss value will gradually increase. The Focal Loss proposed in the RetinaNet network performs excellently in dealing with the class imbalance problem. In object detection tasks, for example, in an image containing safety helmets and a large amount of background, the safety helmets may only occupy a few pixels and belong to the foreground class, while the number of background pixels is much larger than that of the foreground. Focal Loss can dynamically adjust the weights of different class samples by improving the standard cross-entropy loss function. For easily classified samples (such as background samples), their loss weights are reduced; while for difficult-to-classify samples (such as safety helmet samples), the loss weights are increased. This enables the model to pay more attention to difficult-to-classify samples during training, thereby improving the detection accuracy of minority classes such as safety helmets.
[0071] For example, in an image dataset of a loading and unloading area, the number of personnel not wearing safety helmets is relatively small, the number of personnel wearing safety helmets is large, and the background area occupies most of the image. Using Focal Loss, the model can better focus on the personnel not wearing safety helmets and avoid the decline in the model's detection ability for minority classes caused by class imbalance.
[0072] Then, according to the gradient of the loss function, backpropagate to all parameters of the model, update the parameter values, and repeat the training until the preset number of training epochs is reached; use the test set to evaluate the performance of the trained model, calculate metrics such as accuracy, recall, and mAP to evaluate the performance of the model; finally, test the verification results and adjust hyperparameters such as the learning rate and batch size to optimize the performance of the model and obtain an optimized safety helmet recognition model.
[0073] The present invention inputs the image of the loading and unloading area of the distribution center to be recognized obtained in real time into the safety helmet recognition model, outputs the detection result, and determines whether to issue an alarm according to the detection result. The present invention provides a complete solution from sample collection and annotation, model construction and training to actual deployment and application, realizes real-time and intelligent monitoring of the safety helmet wearing situation in the loading and unloading area of the distribution center, effectively solves the blind area problem of manual supervision, and improves the efficiency and accuracy of safety production management.
[0074] The method for recognizing safety helmets in the loading and unloading area in the embodiments of the present invention has been described above. Next, the device for recognizing safety helmets in the loading and unloading area in the embodiments of the present invention will be described. Please refer to Figure 7 , an embodiment of the device for recognizing safety helmets in the loading and unloading area in the embodiments of the present invention includes:
[0075] The annotation and classification module 10 is used to perform annotation and classification processing on the collected image samples of the loading and unloading area of the distribution center to obtain an image sample dataset of the loading and unloading area of the distribution center;
[0076] The data division module 20 is used to divide the image samples of the loading and unloading area of the distribution center into a training set and a test set according to a predetermined ratio, and both the training set and the test set include all image samples of personnel wearing safety helmets and image samples of personnel not all wearing safety helmets;
[0077] The model construction module 30 is used to construct an initial recognition model based on the RetinaNet network;
[0078] The training module 40 is used to train the initial recognition model through the training set, adjust the parameters of the initial recognition model, and obtain a safety helmet recognition model after testing with the test set;
[0079] The recognition module 50 is configured to input the image of the unloading area of the sorting center to be recognized obtained in real time into the safety helmet recognition model, output a detection result, and determine whether to issue an alarm according to the detection result.
[0080] The present invention provides a complete solution from sample collection and annotation, model construction and training to actual deployment and application, realizing real-time and intelligent monitoring of the wearing situation of safety helmets in the loading and unloading area of the sorting center, effectively solving the blind area problem of manual supervision, and improving the efficiency and accuracy of safety production management.
[0081] Please refer to Figure 8 , another embodiment of the safety helmet recognition device in the loading and unloading area in the embodiment of the present invention includes:
[0082] The annotation and classification module 10 is configured to perform annotation and classification processing on the collected image samples of the unloading area of the sorting center to obtain an image sample dataset of the unloading area of the sorting center;
[0083] The data division module 20 is configured to divide the image samples of the unloading area of the sorting center into a training set and a test set according to a predetermined ratio, and both the training set and the test set include all the image samples of wearing safety helmets and the image samples of not wearing safety helmets completely;
[0084] The model construction module 30 is configured to construct an initial recognition model based on the RetinaNet network;
[0085] The training module 40 is configured to train the initial recognition model through the training set, adjust the parameters of the initial recognition model, and obtain a safety helmet recognition model after testing by the test set;
[0086] The recognition module 50 is configured to input the image of the unloading area of the sorting center to be recognized obtained in real time into the safety helmet recognition model, output a detection result, and determine whether to issue an alarm according to the detection result.
[0087] In this embodiment, the annotation and classification module 10 includes:
[0088] The acquisition unit 11 is configured to regularly collect image samples through a camera installed in the loading and unloading area of the sorting center to obtain image samples of the unloading area of the sorting center;
[0089] The amplification unit 12 is configured to perform data amplification processing on the image samples of the unloading area of the sorting center to obtain expanded picture samples;
[0090] The classification unit 13 is configured to perform annotation and classification processing on the expanded picture samples to obtain an image sample dataset of the unloading area of the sorting center.
[0091] In this embodiment, the amplification unit 12 includes:
[0092] A deformation subunit 121, configured to perform mirroring, rotation, scaling, cropping, translation, or Gaussian noise processing on the image sample of the unloading area of the sorting center to obtain a deformed picture sample;
[0093] A transformation subunit 122, configured to perform random cropping processing and / or shear fusion processing on the image sample of the unloading area of the sorting center and the deformed picture sample to obtain an augmented picture sample.
[0094] In this embodiment, the classification unit 13 includes:
[0095] A labeling subunit 131, configured to perform feature labeling on the head area of the sorting personnel in the augmented picture sample through the Labelme tool to obtain an image sample labeled with feature information;
[0096] A classification subunit 132, configured to classify the image sample according to the feature information to obtain an image sample of all personnel wearing safety helmets and an image sample of those not all wearing safety helmets, and the image sample of all personnel wearing safety helmets and the image sample of those not all wearing safety helmets constitute the image sample dataset of the unloading area of the sorting center.
[0097] In this embodiment, the model construction module 30 includes:
[0098] A building unit 31, configured to build a RetinaNet network framework, and the RetinaNet network framework includes a feature extraction layer, a feature pyramid network, and a classification and regression sub-network, wherein the feature extraction layer is composed of multiple convolutional layers and pooling layers;
[0099] An introduction unit 32, configured to introduce the Self-Attention mechanism on each layer of the feature pyramid network, and enhance the expression ability of multi-scale features by calculating the correlation between feature positions at each position, so as to build an initial recognition model.
[0100] In this embodiment, the training module 40 includes:
[0101] An initialization unit 41, configured to assign initial values to all parameters in the initial recognition model by using a random initialization method;
[0102] A calculation unit 42, configured to randomly select a picture sample from the training set and input it into the initial recognition model for training, calculate the prediction result through forward propagation, and then calculate the loss value according to the Focal Loss function;
[0103] An update unit 43, configured to update the parameters of the network by using the backpropagation algorithm to gradually reduce the loss value;
[0104] The repeating unit 44 is used to repeat the above training steps until a predetermined number of training rounds is reached, and a trained initial recognition model is obtained;
[0105] The evaluation unit 45 is used to evaluate the trained initial recognition model using the test set, and adjust the hyperparameters according to the evaluation results to obtain an optimized safety helmet recognition model.
[0106] In this embodiment, the loss calculation unit 43 includes:
[0107] The classification loss calculation sub-unit 431 is used to adopt the focal loss as the classification loss function, and measure the difference between the class probability predicted by the initial recognition model and the true class through the classification loss function. The calculation formula of the Focal loss is as follows: L_cls = -α(1 - p_t)^γlog(p_t), where p_t is the class probability predicted by the initial recognition model, and α and γ are hyperparameters used to adjust the weights of different difficult and easy samples. When p_t is close to 1, (1 - p_t)^γ will become smaller, making the loss value of this sample smaller; conversely, when p_t is close to 0, (1 - p_t)^γ will become larger, making the loss value of this sample larger;
[0108] The regression loss calculation sub-unit 432 is used to adopt the smooth L1 loss as the regression loss function, and measure the difference between the predicted bounding box position of the initial recognition model and the true bounding box position through the regression loss function. The calculation formula of the SmoothL1 loss is as follows: L_reg = 0.5*(IoU - 1)^2*1 if |IoU - 1| < 1, otherwise L_reg = |IoU - 1| - 0.5, where IoU represents the IOU value of the predicted bounding box and the true bounding box. When the IoU value is close to 1, it means that the predicted box is very close to the true box, and the loss value is smaller at this time; when the IoU value deviates from 1, the loss value will gradually increase.
[0109] Above Figure 7 and Figure 8 The safety helmet recognition device in the loading and unloading area in the embodiment of the present invention is described in detail from the perspective of modular functional entities. Next, the safety helmet recognition device in the loading and unloading area in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0110] Figure 9FIG. 0 is a schematic structural diagram of a safety helmet recognition device in a loading and unloading area provided by an embodiment of the present invention. The safety helmet recognition device 100 in the loading and unloading area may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 11 (for example, one or more processors) and a memory 12, and one or more storage media 13 for storing application programs 133 or data 132 (for example, one or more mass storage devices). Among them, the memory 12 and the storage media 13 may be transient storage or persistent storage. The program stored in the storage media 13 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the safety helmet recognition device 100 in the loading and unloading area. Further, the processor 11 may be configured to communicate with the storage media 13 and execute a series of instruction operations in the storage media 13 on the safety helmet recognition device 100.
[0111] The safety helmet recognition device 100 in the loading and unloading area may further include one or more power supplies 14, one or more wired or wireless network interfaces 15, one or more input / output interfaces 16, and / or one or more operating systems 131, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that Figure 9 the shown device structure does not limit the safety helmet recognition device 100 in the loading and unloading area, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0112] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer, the computer is caused to execute the steps of the safety helmet recognition method in the loading and unloading area.
[0113] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system or device and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0114] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0115] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.
Claims
1. A method for identifying safety helmets in loading and unloading areas, characterized in that: The method for identifying safety helmets in a loading and unloading area comprises: The collected image samples of the unloading area of the distribution center are labeled and classified to obtain a sample data set of the image of the unloading area of the distribution center; Dividing the image samples of the unloading area of the distribution center into a training set and a test set according to a predetermined ratio, wherein the training set and the test set both include image samples of all wearing helmets and image samples of not all wearing helmets; Build an initial recognition model based on the RetinaNet network; The initial recognition model is trained by the training set, the parameters of the initial recognition model are adjusted, and after being tested by the test set, a helmet recognition model is obtained; The real-time acquired image of the unloading area of the distribution center to be identified is input into the helmet recognition model, the detection result is output, and it is determined whether to issue an alarm based on the detection result.
2. The method for identifying safety helmets in loading and unloading areas according to claim 1, characterized in that: The collected image samples of the unloading area of the distribution center are labeled and classified to obtain a sample data set of the image of the unloading area of the distribution center, including the following steps: The camera installed in the loading and unloading area of the distribution center is used to regularly collect image samples to obtain image samples of the unloading and loading area of the distribution center; Performing data augmentation processing on the image samples of the unloading area of the distribution center to obtain expanded image samples; The expanded image samples are labeled and classified to obtain an image sample dataset of the unloading area of the distribution center.
3. The method for identifying safety helmets in loading and unloading areas according to claim 2, characterized in that: Performing data amplification processing on the image samples of the unloading area of the distribution center to obtain expanded image samples includes the following steps: Mirroring, rotating, scaling, cropping, translating or Gaussian noise processing are performed on the image samples of the unloading area of the distribution center to obtain deformed image samples; The image samples of the unloading area of the distribution center and the deformed image samples are subjected to random cropping and / or shearing and fusion processing to obtain expanded image samples.
4. The method for identifying safety helmets in loading and unloading areas according to claim 2, characterized in that: The expanded image samples are labeled and classified to obtain an image sample dataset of the unloading area of the distribution center, including the following steps: Use the Labelme tool to perform feature annotation on the head area of the assigned personnel in the expanded image sample to obtain an image sample annotated with feature information; The image samples are classified according to the feature information to obtain image samples of all helmets being worn and image samples of not all helmets being worn, and the image samples of all helmets being worn and the image samples of not all helmets being worn constitute the image sample data set of the unloading area of the distribution center.
5. The method for identifying safety helmets in loading and unloading areas according to claim 1, characterized in that: Building an initial recognition model based on the RetinaNet network includes the following steps: Building a RetinaNet network framework, the RetinaNet network framework includes a feature extraction layer, a feature pyramid network and a classification regression sub-network, wherein the feature extraction layer is composed of multiple convolutional layers and pooling layers; A Self-Attention mechanism is introduced on each layer of the feature pyramid network to enhance the expression capability of multi-scale features by calculating the correlation between features at each position, thereby constructing an initial recognition model.
6. The method for identifying safety helmets in loading and unloading areas according to claim 1, characterized in that: The initial recognition model is trained by the training set, the parameters of the initial recognition model are adjusted, and after being tested by the test set, a helmet recognition model is obtained, including the steps of: Using a random initialization method to assign initial values to all parameters in the initial identification model; Randomly select image samples from the training set and input them into the initial recognition model for training, calculate the prediction results through forward propagation, and then calculate the loss value according to the Focal Loss function; Use the back propagation algorithm to update the parameters of the network so that the loss value gradually decreases; Repeat the above training steps until the predetermined training rounds are reached to obtain a trained initial recognition model; The trained initial recognition model is evaluated using the test set, and the hyperparameters are adjusted according to the evaluation results to obtain an optimized helmet recognition model.
7. The method for identifying safety helmets in loading and unloading areas according to claim 6, characterized in that: Calculate the loss value according to the Focal Loss function, including the following steps: Focal loss is used as the classification loss function, and the classification loss function is used to measure the difference between the category probability predicted by the initial recognition model and the actual category. The calculation formula of focal loss is as follows: L_cls = -α(1-p_t)^γlog(p_t), where p_t is the category probability predicted by the initial recognition model, α and γ are hyperparameters used to adjust the weights of samples of different difficulty levels. When p_t is close to 1, (1-p_t)^γ will become smaller, making the loss value of the sample smaller; conversely, when p_t is close to 0, (1-p_t)^γ will become larger, making the loss value of the sample larger; Smooth L1 loss is used as the regression loss function. The regression loss function is used to measure the difference between the bounding box position predicted by the initial recognition model and the true bounding box position. The calculation formula of Smooth L1 loss is as follows: L_reg=0.5*(IoU-1)^2*1if|IoU-1|<1,otherwise L_reg=|IoU-1|-0.5, where IoU represents the IOU value of the predicted bounding box and the true bounding box. When the IoU value is close to 1, it means that the predicted box is very close to the true box, and the loss value is small at this time; when the IoU value deviates from 1, the loss value will gradually increase.
8. A safety helmet identification device for a loading and unloading area, characterized in that: include: The labeling and classification module is used to label and classify the collected image samples of the unloading area of the distribution center to obtain the image sample data set of the unloading area of the distribution center; A data division module, used for dividing the image samples of the unloading area of the distribution center into a training set and a test set according to a predetermined ratio, wherein the training set and the test set both include image samples of all wearing helmets and image samples of not all wearing helmets; Model building module, used to build the initial recognition model based on the RetinaNet network; A training module, used to train the initial recognition model through the training set, adjust the parameters of the initial recognition model, and obtain a helmet recognition model after testing with the test set; The recognition module is used to input the real-time acquired image of the unloading area of the distribution center to be identified into the helmet recognition model, output the detection result, and determine whether to issue an alarm based on the detection result.
9. A helmet identification device for loading and unloading areas, characterized in that: comprising a memory and at least one processor, wherein the memory has computer-readable instructions stored therein; The at least one processor calls the computer-readable instructions in the memory to execute the various steps of the method for identifying a safety helmet in a loading and unloading area according to any one of claims 1-7.
10. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the steps of the method for identifying a safety helmet in a loading and unloading area according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Service-oriented robot target identification method based on visual analysis
CN120997792A
Vision analysis-based service robot target recognition method
CN120997792B