Power plant worker safety helmet and safety clothing intelligent detection method and system based on real-time visual Transform

Through the intelligent detection method of power plant workers' safety helmets and safety clothing based on real-time visual Transformer, the real-time monitoring problem of workers' safety helmets and safety clothing wearing status in power plants is solved, which realizes efficient and accurate detection and management, reduces the risk of accidents, and adapts to the intelligent needs of power plants.

CN120673328APending Publication Date: 2025-09-19CHINA YANGTZE POWER
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510666358.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies cannot achieve real-time and comprehensive monitoring of workers' helmet and safety clothing wearing status in power plants, resulting in increased accident risks and high intensity of manual inspections. Traditional machine learning methods also have low detection accuracy, high computational complexity, and long training time in complex environments.

Method used

A detection method based on real-time visual Transformer is adopted. Through data collection, labeling, model training and deployment, high-definition cameras are used for real-time detection. The PyQt5 framework is combined to realize data visualization and alarm functions, and support regular model updates and distributed storage.

Benefits of technology

It realizes real-time, high-precision monitoring of workers' helmet and safety clothing wearing status, reduces accident risks, reduces the workload of inspectors, improves detection efficiency and system interactivity, and adapts to the intelligent requirements of power plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673328A_ABST
    Figure CN120673328A_ABST
Patent Text Reader

Abstract

The invention discloses a power plant worker safety helmet and safety clothing intelligent detection method and system based on real-time visual Transform, and aims to solve the problems that the supervision of the wearing of worker safety helmets and safety clothing in a transmission power plant mainly depends on manual inspection, the real-time comprehensive monitoring cannot be realized, the efficiency is low and the like. A high-definition camera is deployed in a key area of a power plant to collect working image data of workers in real time. A real-time visual Transform model trained by a large amount of labeled data is utilized to intelligently analyze and process the image, and whether a worker correctly wears a safety helmet and a safety suit or not is accurately judged; the model has the advantages of high speed, high noise resistance, high detection precision and the like, and can adapt to a complex and changeable power plant environment; according to the invention, real-time and automatic monitoring of the wearing states of the safety helmets and the safety garments of the power plant workers is realized, the safety management level is effectively improved, the accident risk is reduced, and the system has a significant practical value and a wide market application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing and artificial intelligence technology, and in particular to a method and system for intelligently detecting safety helmets and safety clothing for power plant workers based on a real-time visual Transformer. Background Art

[0002] In the power industry, the safety of power plant workers is always a top priority. Properly wearing hard hats and safety clothing is a fundamental measure to ensure worker safety. However, in practice, due to the complexity of power plant environments and the diversity of worker behavior, how to comprehensively monitor workers' hard hat and safety clothing wearing status in real time has become a pressing technical challenge.

[0003] Currently, many power plants still rely on traditional manual inspections to check workers' helmets and safety clothing. This method, which typically relies on on-site observation and record-keeping by inspectors, is not only inefficient but also struggles to achieve real-time, comprehensive monitoring. Furthermore, manual inspections are labor-intensive and susceptible to subjective factors, leading to inaccurate and inconsistent monitoring results.

[0004] To overcome the limitations of manual inspections, some power plants have begun experimenting with traditional machine learning-based detection methods. These methods typically rely on manually designed feature templates and use image processing techniques to identify whether workers are wearing helmets or safety clothing. However, due to the complexity of power plant environments and the diversity of workers' clothing, traditional machine learning methods often face significant challenges in feature extraction and classification, resulting in low detection accuracy and difficulty meeting practical needs.

[0005] For example, CN117636472A discloses a method for intelligently detecting the wearing behavior of helmets and safety clothing. This method is based on visual recognition technology. By collecting a dataset of images of helmets and safety clothing being properly worn, the method uses the YOLOv5 algorithm for training to obtain a wearing analysis model, which is then deployed to a server. By establishing an image detection area and a human body recognition detection camera device, the method analyzes the video stream data within the detection area, combines the long-short memory algorithm to determine the wearing status of workers, and issues an alarm if the wearer is not wearing the helmet safely. However, this method has the following defects and shortcomings: Limited real-time performance: Although this method can parse video stream data in real time and determine the wearable status, in complex environments, such as target deformation, sudden movement, background clutter, occlusion, or video frame loss, it may lead to false detection or missed detection, affecting the accuracy of real-time monitoring.

[0006] Model generalization capability: Although the YOLOv5 model can adapt to different wearable detection scenarios to a certain extent, its generalization capability may be limited in extremely complex industrial environments, resulting in reduced detection accuracy.

[0007] Data processing efficiency: For large-scale video stream data, the parsing and processing process may require higher computing resources, affecting the overall operating efficiency of the system.

[0008] For example, CN118053176A discloses a method for detecting wearable devices based on multi-branch deep convolution and dual attention. This method constructs a sample library by taking pictures of substation work and network collection pictures, improves the YOLOv7 network using a multi-branch deep convolution network (MBDC), inserts a dual attention mechanism (MLF, Multi-Layer Feedforward, multi-layer feedforward network), and combines the focal loss and SIOU loss function to optimize the model to improve the detection ability of small targets, multiple targets, and occluded targets. However, this method has the following defects and shortcomings: Computational complexity: The introduction of multi-branch deep convolution and dual attention mechanism improves detection accuracy, but also significantly increases the computational complexity of the model, which may make it difficult to deploy on resource-constrained devices.

[0009] Training time: Due to the complexity of the model structure, the training process may take longer to converge, increasing the cost of development and deployment.

[0010] Small object detection: Although this method is optimized for small object detection, the detection performance may still be affected in extreme cases, such as when the small object is severely occluded or in a complex background.

[0011] Existing technologies have made some progress in intelligent detection of helmet and safety clothing wearing behavior, but they still face problems such as limited real-time performance, insufficient model generalization, low data processing efficiency, high computational complexity, and long training times. Therefore, developing an intelligent detection method for helmet and safety clothing wearing behavior that can ensure real-time performance while improving detection accuracy and generalization, while reducing computational complexity and training time, has important practical significance and application value. This paper proposes a method and system for intelligent detection of helmets and safety clothing for power plant workers based on a real-time visual Transformer. This method aims to achieve real-time, high-precision monitoring of helmet and safety clothing wearing by power plant workers through advanced artificial intelligence technology, thereby effectively improving the safety management level of power plants and reducing the incidence of accidents. Summary of the Invention

[0012] The technical problem to be solved by the present invention is to provide a method and system for intelligent detection of safety helmets and safety clothing for power plant workers based on real-time visual Transformer, so as to solve the specific technical problems existing in the field of supervision of the wearing of safety helmets and safety clothing by power plant workers in the power industry; specifically, the traditional manual inspection method is unable to monitor the wearing status of workers' safety helmets and safety clothing in real time and comprehensively, resulting in increased accident risks, and the manual inspection work is high in intensity and interferes with the normal work of workers.

[0013] In order to achieve the above objectives, the present invention adopts the following technical solutions: Pre-deployment steps: Data Collection: Select appropriate image sources, including on-site photos, public image libraries, social media, and surveillance videos, to ensure sample diversity and representativeness. Collect a large number of images of workers wearing and not wearing helmets in different scenarios, and remove redundant samples and images that do not meet classification criteria to ensure dataset quality.

[0014] Data annotation: Use professional image annotation tools to annotate images, labeling them as "safety helmet," "safety suit," and "worker," and draw bounding boxes. Perform data review to ensure the accuracy and consistency of the annotations.

[0015] Model training: On a server, use a high-performance processor and graphics card to train the model on the labeled data. Use bilinear interpolation to unify the image size and perform data augmentation. Select AdamW (Adaptive Moment Estimation with Weight Decay) as the optimizer, set the training parameters, and save the model weights after model training is complete.

[0016] System deployment: Installation and monitoring: Select cameras with high-definition, night vision, waterproof and other functions, and install them in key areas of the power plant. Ensure that the cameras can cover all areas that need to be monitored, and connect the cameras to the wireless network to transmit image data to the server.

[0017] Real-time detection: After acquiring image data, the image is resized to 256×256 using a bilinear interpolation algorithm. The trained model weights are then loaded into the real-time visual Transformer model for image detection. The model follows the encoder-decoder architecture of the Transformer. Image features are extracted using a convolutional neural network. After input embedding and positional encoding, the Transformer encoder and decoder perform feature processing and classification prediction.

[0018] Data visualization: The test results are displayed on the monitor, and the PyQt5 framework is used to implement the user login and registration function, obtain the monitoring name for the user to choose, realize real-time monitoring and playback, and display each test record on the interface.

[0019] In a preferred embodiment, the method further includes an alarm step. When it is detected that a worker is not wearing a safety helmet or safety clothing correctly, the system automatically sends an alarm signal. The alarm signal includes a sound alarm, a light alarm, or sending an alarm message to the manager's mobile phone.

[0020] In a preferred embodiment, the method further includes a model updating step, which involves regularly collecting new image data and updating and training the real-time visual Transformer model to adapt to feature changes that may occur in different time periods and different groups of workers, thereby ensuring the accuracy and reliability of detection. The model update adopts incremental learning or full learning, and records update logs.

[0021] In a preferred embodiment, the method further includes a data storage step, in which the collected image data, annotation data, and detection result data are stored in a database. The database adopts a distributed storage architecture to improve the security and scalability of data storage, and at the same time backs up the data to prevent data loss. The data storage module supports data encryption and access permission control.

[0022] The intelligent detection system for power plant workers' helmets and safety clothing based on real-time visual Transformers implements the aforementioned intelligent detection method for power plant workers' helmets and safety clothing based on real-time visual Transformers, including: Image acquisition module: used to collect image data of workers wearing or not wearing safety helmets and safety clothing in different scenarios; Data processing module: used to label and preprocess the collected image data; Model training module: used to train the real-time visual Transformer model on the server and save the trained model weights; System deployment module: used to install high-definition cameras in key areas of the power plant and connect the cameras to the wireless network; Real-time detection module: used to obtain image data, resize the image, load the trained model weights, and detect the image using the real-time visual Transformer model; Data display module: used to display the test results on the display, including real-time monitoring playback and test record display; Alarm module: used to send out an alarm signal when it detects that a worker is not wearing a safety helmet or safety clothing correctly; Model update module: used to regularly collect new image data and update the training of the real-time visual Transformer model; Data storage module: used to store the collected image data, annotation data, and detection result data in the database.

[0023] In a preferred solution, the image acquisition module is further associated with an environmental parameter acquisition unit, which acquires light intensity, temperature, and humidity parameters of the environment in which the camera is located, and stores them in association with the acquired image data.

[0024] In a preferred solution, the model training module is provided with a training monitoring unit, which monitors the loss function value and accuracy index during the model training process in real time, and automatically adjusts the training parameters or suspends the training and generates an abnormality report when the index is abnormal.

[0025] In a preferred solution, the data display module is provided with a statistical analysis unit, which performs statistical analysis on the detection records and generates violation rate statistical reports and violation type distribution charts for different time periods and different areas.

[0026] The method and system for intelligently detecting power plant workers' helmets and safety clothing based on real-time visual Transformer provided by the present invention have the following beneficial effects: 1. The present invention effectively solves the specific technical problems existing in the field of supervision of the wearing of safety helmets and safety clothing by power plant workers in the power industry, especially the problems that traditional manual inspection methods are unable to monitor the wearing status of workers' safety helmets and safety clothing in real time and comprehensively, resulting in increased accident risks, as well as the problems that manual inspection work is high in intensity and interferes with workers' normal work.

[0027] 2. The detection method and system of the present invention realize unmanned remote detection of safety helmets and safety clothing of power plant workers, significantly improving detection efficiency, reducing the workload of inspectors, and preventing inspections from interfering with the normal work of workers.

[0028] 3. Through the application of the real-time visual Transformer model, the present invention has the advantages of fast speed, easy deployment, strong noise resistance, and high detection accuracy. It adapts to the requirements of intelligent power plants and has excellent practicality and market prospects.

[0029] 4. This invention applies visual Transformer technology to the real-time detection of power plant workers' helmets and safety clothing for the first time, breaking through the limitations of traditional manual inspections and traditional machine learning-based detection methods.

[0030] 5. This invention improves the quality of the data set through diversified data sources and strict data review, providing a reliable foundation for model training.

[0031] 6. The model training optimization of the present invention adopts data enhancement technology and high-performance hardware configuration to improve the generalization ability and training efficiency of the model.

[0032] 7. The present invention rationally arranges cameras according to the actual needs of the power plant and realizes data visualization through the PyQt5 framework, thereby improving the interactivity and practicality of the system.

[0033] 8. The present invention uses real-time visual Transformer technology to achieve real-time monitoring of workers' helmet and safety clothing wearing status, solving the problem that traditional manual inspections cannot provide real-time and comprehensive monitoring.

[0034] 9. The present invention realizes unmanned remote detection, overcomes the problem that manual inspection is labor-intensive and easily interferes with workers' normal work, greatly reduces the workload of inspectors, and avoids the interference of inspection on workers' work.

[0035] 10. The present invention uses deep learning technology to solve the problem that traditional machine learning detection technology relies on manually designed feature templates and has low detection accuracy. It enables the model to adaptively learn appropriate features from manually labeled data, thereby improving detection accuracy.

[0036] 11. The system of the present invention can monitor workers' helmet and safety clothing wearing conditions in real time and comprehensively, provide high-precision detection results, and significantly improve detection efficiency.

[0037] 12. The present invention effectively reduces the risk of accidents caused by workers not wearing safety helmets or improperly wearing safety clothing through real-time monitoring and accurate detection.

[0038] 13. The implementation of the present invention significantly improves the safety management level and supervision information level of the power plant, and provides a strong guarantee for the safe production of the power plant.

[0039] 14. The present invention improves the interactivity and practicality of the system by rationally arranging cameras and realizing data visualization, making the system more adaptable to the actual needs of the power plant. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 Schematic diagram of the structure of the system of the present invention; Figure 2 This is the real-time visual Transformer detection model of the present invention. DETAILED DESCRIPTION

[0041] The technical solutions of the present invention are further described below with reference to the accompanying drawings and embodiments: Example 1 like Figures 1 to 2 As shown, this embodiment provides a method for intelligently detecting safety helmets and safety clothing for power plant workers based on a real-time visual Transformer, including the following steps: Step 1: Data Collection Step 1.1. Diversified selection of image sources On-site photography: Professional photographers were assigned to film various areas of the power plant. Shooting was conducted during all work hours, such as the morning, afternoon, and evening shifts, to capture images under varying lighting conditions. On sunny days, images were captured in both directly sunlit and shaded areas. On overcast days, the lighting was uniform. At night, the power plant's lighting equipment was utilized to simulate scenes of varying brightness levels. Workers were photographed from the front, side, back, and at various angles, ensuring that images of them in various positions, including those wearing helmets and safety clothing, were captured. For example, around the generator sets, workers might be operating in various postures, such as standing, bending over, and looking up, requiring detailed photography.

[0042] Public Image Library: Images related to power plant workers, hard hats, and safety clothing were screened from a professional image database. Based on actual power plant work scenes, images with similar industrial backgrounds and worker clothing styles were selected to enrich the dataset's diversity.

[0043] Social media: Search for pictures and videos related to power plant work on social media platforms and select images that meet the requirements. Be careful to select images with unique angles and representative scenes to increase the novelty of the dataset.

[0044] Surveillance Video: We collect existing surveillance videos from the power plant and use video processing software (such as FFmpeg) to extract them frame by frame, generating a large number of continuous images. These images can reflect the dynamic behavior of workers at work and help improve the generalization ability of the model.

[0045] Step 1.2, acquisition equipment and parameter settings Use a professional-grade digital camera, such as the Canon 5D Mark IV, which features high megapixels (30.4 effective megapixels), high sensitivity (ISO range of 100 to 32,000, expandable to 50 to 102,400), and excellent color reproduction. Choose a wide-angle lens, such as the 16-35mm f / 4L IS USM, to capture a wider scene within a limited space. When shooting, set the camera to manual mode and adjust the aperture, shutter speed, and ISO sensitivity to suit the lighting conditions. For example, on a bright sunny day, set the aperture to f / 8 to f / 11, the shutter speed to 1 / 250s to 1 / 500s, and the ISO sensitivity to 100. At night, in low light, set the aperture to f / 2.8 to f / 4, the shutter speed to 1 / 30s to 1 / 60s (use a tripod to stabilize the camera if necessary), and increase the ISO sensitivity to 800 to 1600.

[0046] Step 2: Data Labeling Step 2.1, annotation tools and processes The open-source image annotation tool LabelImg was used, which features a simple, easy-to-use interface and rich functionality. The annotator first opened LabelImg and loaded the image to be annotated. Using the mouse, they drew precise bounding boxes on the image, selecting the "safety helmet," "safety suit," and "worker" respectively. The bounding boxes were drawn to fit the edges of the target object as closely as possible, with an error of no more than 5 pixels. Once the annotation was complete, the annotation file was saved in the PASCAL VOC format, containing information such as the image path, bounding box coordinates, and category labels.

[0047] Step 2.2: Labeling quality control Establish a strict data review mechanism. Have multiple experienced labelers cross-check the same batch of labeled data. During the review process, focus on the accuracy of the bounding boxes, the correctness of the category labels, and the consistency of the annotations. If any labeling errors are found, such as bounding boxes that are too large or too small, incorrect category labels, etc., feedback will be promptly provided to the labelers for correction. At the same time, regular training will be provided to labelers, including training on labeling standards and feature recognition of target objects, to improve labeling quality. Set the labeling error rate threshold to 5%. If a labeler's error rate exceeds this threshold, they will be required to undergo labeling training again until the error rate meets the standard.

[0048] Step 3: Model training Step 3.1. Data preprocessing The images were resized to a uniform size of 512×512 pixels using a bilinear interpolation algorithm. The image's aspect ratio was maintained during the resizing process, and any portions exceeding the target size were padded with the image's average color to minimize the impact on image features. Data augmentation was performed. During random cropping, the size of the cropped region was randomly selected between 80% and 100% of the original image size, maintaining a 512×512 pixel size. Random rotations were performed within a range of -15° to 15°, with the image edges padded after rotation. Random flips included both horizontal and vertical flips, with a 50% probability of flipping.

[0049] Step 3.2, training parameter setting AdamW was chosen as the optimizer. It combines the Adam optimizer's adaptive learning rate feature with weight decay to effectively prevent model overfitting. The initial learning rate was set to 0.0001 and dynamically adjusted based on validation set performance during training. If the model's accuracy on the validation set did not improve for five consecutive epochs, the learning rate was reduced to 0.1 times the original value. The batch size was set to 16 to ensure that each training run fully utilized the parallel computing power of the GPU (Graphics Processing Unit). The number of epochs was set to 100 to allow the model sufficient time to learn the features in the dataset.

[0050] Step 3.3, cross validation and model evaluation Using the K-fold cross-validation method, the dataset is divided into training, validation, and test sets in a ratio of 7:2:1. For example, the dataset is randomly divided into 10 parts, 9 of which are selected each time as the training and validation sets (7 for training and 2 for validation), with the remaining part as the test set. This is repeated 10 times, and the average value is used as the model evaluation metric. During training, the model's accuracy and loss function value on the validation set are monitored in real time. Accuracy is evaluated using Exact Match Accuracy and Partial Match Accuracy. Exact Match Accuracy requires that the model's predicted bounding boxes and categories are completely consistent with the ground-truth annotations, while Partial Match Accuracy allows for a certain degree of deviation. The cross-entropy loss function is used as the loss function, and training is terminated when the model's accuracy on the validation set reaches 95%.

[0051] Step 4. System deployment Step 4.1, Camera selection and installation Install high-definition cameras in key areas of the power plant, such as around generators, transformer areas, and aerial work platforms. Choose the Hikvision DS-2CD3325F-I camera, which features high definition (resolution no less than 1080P), night vision (infrared night vision range up to 30 meters), and waterproofing (IP66 protection rating). Adjust the mounting angle based on specific monitoring needs. For example, around generators, the camera should be installed to provide clear coverage of the entire operating area, ideally with a front and side view of workers. The mounting height should generally be between 3 and 5 meters to avoid obstruction by equipment or personnel.

[0052] Step 4.2, Network connection and data transmission The cameras are connected to the server via a wireless network using 5G communication technology. A 5G (fifth-generation mobile communication technology) base station was built within the power plant to ensure signal coverage of all monitoring areas. The cameras were configured for network connectivity, with fixed IP (Internet Protocol) addresses and port numbers set to facilitate data reception by the server. Furthermore, data encryption technology was used to encrypt transmitted image data to prevent data leakage.

[0053] Step 5: Real-time detection Step 5.1: Image preprocessing After acquiring image data transmitted by the camera, we first perform denoising using a median filter with a filter window size of 3×3 pixels. For images containing salt-and-pepper noise, median filtering effectively removes noise points while preserving edge information. Contrast enhancement is then performed using a histogram equalization algorithm, with parameters adjusted based on the scene. In bright daylight, the contrast enhancement intensity of the histogram equalization algorithm is set to 0.8 to avoid overexposure; in low-light nighttime scenes, the intensity is set to 1.2 to increase image brightness.

[0054] Step 5.2, Model loading and testing Load the trained model weights. The model file is stored in PyTorch's .pth format. Use the real-time visual Transformer model to detect images. The model first extracts image features using a convolutional neural network (CNN). The convolution kernel size is set to 3×3, the stride is 1, and the padding is 1 to maintain the size of the feature map. Feature processing and classification prediction are then performed through the transformer encoder and decoder. The encoder layer is set to 6, and the decoder layer is set to 6. The self-attention mechanism captures the correlation information between different areas in the image. Finally, it determines whether the worker is wearing a safety helmet and safety clothing correctly. The detection results are output in the form of bounding boxes and category labels. Figure 2 In the figure, backbone is the backbone of the network, set of image features is the image feature set, positional encoding is the position encoding, object queries are object queries, prediction heads are prediction heads, FFN (Feed-Forward Neural Network) refers to the feedforward neural network, class box refers to the target, and noobject refers to the background.

[0055] Step 6: Data Visualization Step 6.1. User Interface Design The user login and registration functions are implemented using the PyQt5 framework, and the user interface design is simple and clear. The login screen includes a username and password input box and a login button. The registration screen includes input boxes for username, password, and confirm password. The user selects a monitoring name, which is displayed in a list on the interface. The user can select different monitoring areas by clicking the mouse.

[0056] Step 6.2, Real-time monitoring and detection record display The real-time monitoring screen is displayed in the form of a video stream with a frame rate set to 25 frames per second to ensure smooth viewing. At the same time, the inspection records are displayed on the interface in a list format, which includes fields such as inspection time, monitoring area, worker number, helmet wearing status, and safety clothing wearing status. Inspection record query and export functions are provided. The query function supports filtering by time (accurate to the minute), area (specific to the workshop, equipment number), inspection result (correct / not wearing helmet, correct / not wearing safety clothing), and other conditions. The export function supports CSV (Comma-Separated Values) and Excel formats. Users can select the export format as needed to facilitate data analysis and processing.

[0057] Example 2 In another preferred embodiment, based on the above embodiment 1, this embodiment provides an intelligent detection method for power plant workers' safety helmets and safety clothing based on real-time visual Transformer. The specific steps are as follows: 1. Data Collection 1.1. Diversified selection of image sources On-site photography: Professional photographers were assigned to film various areas of the power plant. Shooting was conducted during all work hours, such as the morning, afternoon, and evening shifts, to capture images under varying lighting conditions. On sunny days, images were captured in both directly sunlit and shaded areas. On overcast days, the lighting was uniform. At night, the power plant's lighting equipment was utilized to simulate scenes of varying brightness levels. Workers were photographed from the front, side, back, and at various angles, ensuring that images of them in various positions, including those wearing helmets and safety clothing, were captured. For example, around the generator sets, workers might be operating in various postures, such as standing, bending over, and looking up, requiring detailed photography.

[0058] Public Image Library: Images related to power plant workers, hard hats, and safety clothing were screened from a professional image database. Based on actual power plant work scenes, images with similar industrial backgrounds and worker clothing styles were selected to enrich the dataset's diversity.

[0059] Social media: Search for pictures and videos related to power plant work on social media platforms and select images that meet the requirements. Be careful to select images with unique angles and representative scenes to increase the novelty of the dataset.

[0060] Surveillance Video: We collect existing surveillance videos from the power plant and use video processing software (such as FFmpeg) to extract them frame by frame, generating a large number of continuous images. These images can reflect the dynamic behavior of workers at work and help improve the generalization ability of the model.

[0061] 1.2. Acquisition equipment and parameter settings Use a professional-grade digital camera, such as the Canon 5D Mark IV, which features high megapixels (30.4 effective megapixels), high sensitivity (ISO range of 100 to 32,000, expandable to 50 to 102,400), and excellent color reproduction. Choose a wide-angle lens, such as the 16-35mm f / 4L IS USM, to capture a wider scene within a limited space. When shooting, set the camera to manual mode and adjust the aperture, shutter speed, and ISO sensitivity to suit the lighting conditions. For example, on a bright sunny day, set the aperture to f / 8 to f / 11, the shutter speed to 1 / 250s to 1 / 500s, and the ISO sensitivity to 100. At night, in low light, set the aperture to f / 2.8 to f / 4, the shutter speed to 1 / 30s to 1 / 60s (use a tripod to stabilize the camera if necessary), and increase the ISO sensitivity to 800 to 1600.

[0062] 2. Data Labeling 2.1. Annotation Tools and Processes The open-source image annotation tool LabelImg was used, which features a simple, easy-to-use interface and rich functionality. The annotator first opened LabelImg and loaded the image to be annotated. Using the mouse, they drew precise bounding boxes on the image, selecting the "safety helmet," "safety suit," and "worker" respectively. The bounding boxes were drawn to fit the edges of the target object as closely as possible, with an error of no more than 5 pixels. Once the annotation was complete, the annotation file was saved in the PASCAL VOC format, containing information such as the image path, bounding box coordinates, and category labels.

[0063] 2.2. Labeling quality control Establish a strict data review mechanism. Have multiple experienced labelers cross-check the same batch of labeled data. During the review process, focus on the accuracy of the bounding boxes, the correctness of the category labels, and the consistency of the annotations. If any labeling errors are found, such as bounding boxes that are too large or too small, incorrect category labels, etc., feedback will be promptly provided to the labelers for correction. At the same time, regular training will be provided to labelers, including training on labeling standards and feature recognition of target objects, to improve labeling quality. Set the labeling error rate threshold to 5%. If a labeler's error rate exceeds this threshold, they will be required to undergo labeling training again until the error rate meets the standard.

[0064] 3. Model training 3.1 Data Preprocessing The images were resized to a uniform size of 512×512 pixels using a bilinear interpolation algorithm. The image's aspect ratio was maintained during the resizing process, and any portions exceeding the target size were padded with the image's average color to minimize the impact on image features. Data augmentation was performed. During random cropping, the size of the cropped region was randomly selected between 80% and 100% of the original image size, maintaining a 512×512 pixel size. Random rotations were performed within a range of -15° to 15°, with the image edges padded after rotation. Random flips included both horizontal and vertical flips, with a 50% probability of flipping.

[0065] 3.2 Training parameter settings AdamW was selected as the optimizer. It combines the Adam optimizer's adaptive learning rate feature with weight decay to effectively prevent model overfitting. The initial learning rate was set to 0.0001 and dynamically adjusted based on validation set performance during training. If the model's accuracy on the validation set did not improve for five consecutive epochs, the learning rate was reduced to 0.1 times the original value. The batch size was set to 16 to ensure that each training session fully utilized the parallel computing power of the GPU. The number of epochs was set to 100 to allow the model sufficient time to learn the features of the dataset.

[0066] 3.3 Cross-validation and model evaluation Using the K-fold cross-validation method, the dataset is divided into training, validation, and test sets in a ratio of 7:2:1. For example, the dataset is randomly divided into 10 parts, 9 of which are selected each time as the training and validation sets (7 for training and 2 for validation), with the remaining part as the test set. This is repeated 10 times, and the average is taken as the model evaluation metric. During training, the model's accuracy and loss function value on the validation set are monitored in real time. Accuracy is evaluated using Exact Match Accuracy and Partial Match Accuracy. Exact Match Accuracy requires that the model's predicted bounding boxes and categories are completely consistent with the ground-truth annotations, while Partial Match Accuracy allows for a certain degree of deviation. The cross-entropy loss function is used as the loss function, and training is terminated when the model's accuracy on the validation set reaches 95%.

[0067] 4. System deployment 4.1 Camera selection and installation Install high-definition cameras in key areas of the power plant, such as around generators, transformer areas, and aerial work platforms. Choose the Hikvision DS-2CD3325F-I model, which features high definition (resolution no less than 1080P), night vision (infrared night vision range up to 30 meters), and waterproofing (IP66 protection rating). The mounting angle should be adjusted based on specific monitoring needs. For example, around generators, cameras should be installed to provide clear coverage of the entire operating area, ideally with a front and side view of workers. The mounting height should generally be between 3 and 5 meters to avoid obstruction by equipment or personnel.

[0068] 4.2 Network connection and data transmission The cameras are connected to the server via a wireless network using 5G communication technology. A 5G base station is built within the power plant to ensure signal coverage of all monitoring areas. The cameras are configured for network connectivity, with fixed IP addresses and port numbers set to facilitate data reception by the server. Furthermore, data encryption technology is used to encrypt transmitted image data to prevent data leakage.

[0069] 5. Real-time detection 5.1 Image Preprocessing After acquiring image data transmitted by the camera, we first perform denoising using a median filter with a filter window size of 3×3 pixels. For images containing salt-and-pepper noise, median filtering effectively removes noise points while preserving edge information. Contrast enhancement is then performed using a histogram equalization algorithm, with parameters adjusted based on the scene. In bright daylight, the contrast enhancement intensity of the histogram equalization algorithm is set to 0.8 to avoid overexposure; in low-light nighttime scenes, the intensity is set to 1.2 to increase image brightness.

[0070] 5.2 Model loading and testing Load the trained model weights. The model file is stored in PyTorch's .pth format. Use the real-time visual Transformer model to detect images. The model first extracts image features using a convolutional neural network. The convolution kernel size is set to 3×3, with a stride of 1 and padding of 1 to maintain the size of the feature map. Feature processing and classification prediction are then performed using a transformer encoder and decoder. The encoder and decoder have 6 layers, and the self-attention mechanism captures correlations between different image regions. Finally, the model determines whether the worker is wearing the helmet and safety clothing correctly. The detection results are output as bounding boxes and category labels.

[0071] 6. Alarm steps 6.1. Alarm signal type If a worker is not wearing a safety helmet or safety clothing correctly, the system automatically triggers an alarm. The audible alarm uses a high-decibel buzzer set to above 90 decibels to ensure clear hearing in the noisy environment of the power plant. The visual alarm uses a red LED (Light Emitting Diode) warning light, prominently located in the monitoring room and key areas. When an alarm is triggered, the light flashes three times per second. Alarm messages are sent to managers' mobile phones via an SMS gateway or internal company instant messaging software. Alibaba Cloud's SMS service is used as the SMS gateway, and DingTalk is used as the company instant messaging software. When an alarm is triggered, the system automatically sends a text or instant message to a pre-defined manager's mobile phone containing the alarm time, location, worker ID, and specific violation details.

[0072] 7. Data Storage 7.1 Database Architecture A distributed storage architecture is used, such as the Hadoop Distributed File System (HDFS) and the Apache Cassandra database. HDFS stores large amounts of image and annotation data, offering high fault tolerance and scalability, meeting the power plant's massive data storage needs. Apache Cassandra, used to store inspection result data, offers high availability, linear scalability, and rapid response to query requests.

[0073] 7.2 Data Backup and Security Database data is backed up regularly, with a daily full backup and hourly incremental backups. Full backups are performed using a tape library, while incremental backups are performed using a disk array. Data is encrypted using the AES-256 encryption algorithm to ensure security during storage and transmission. Access control is implemented, assigning different access rights based on personnel responsibilities. For example, regular employees can only query test results, while management personnel can modify data and perform system configuration operations.

[0074] 8. Model Update 8.1 Data Collection and Update Methods New image data is collected regularly, with the collection cycle being weekly. Data is collected from newly added monitoring areas, newly hired workers, and seasonal changes in worker attire. Model updates are performed using either incremental learning or full-scale learning. Incremental learning fine-tunes the model using only newly collected data, with a batch size of 8, a learning rate of 0.00005, and an epoch of 20. Full-scale learning retrains the model using all data, with an initial learning rate of 0.0001, a batch size of 16, and an epoch of 100.

[0075] 8.2. Update log records Records information such as the time of model update, update method (incremental learning or full learning), amount of data used, and changes in model accuracy before and after the update. The update log is stored as a text file on the server and backed up to cloud storage to prevent data loss.

[0076] 9. System deployment (alarm, data storage, etc. integration) 9.1. Alarm and data storage integration The system architecture integrates the alarm module with the data storage module. When an alarm is triggered, not only is an alarm signal generated, but alarm information (including time, location, worker ID, violation type, etc.) is also stored in the database. Historical alarm data is also retrieved from the database for analysis and violation statistics.

[0077] 9.2 Real-time detection (including alarm) 9.3. Detection and alarm linkage During the real-time monitoring process, if a worker is detected not wearing a helmet or safety clothing correctly, an alarm is triggered immediately. Simultaneously, the detection results and alarm information are displayed in real time on a user interface implemented using the PyQt5 framework. The user interface displays alarm information in a list format, including the alarm time, monitoring area, worker number, and violation status.

[0078] 10. Data Visualization 10.1 User Interface Design The user login and registration functions are implemented using the PyQt5 framework, and the user interface design is simple and clear. The login screen includes a username and password input box and a login button. The registration screen includes input boxes for username, password, and confirm password. The user selects a monitoring name, which is displayed in a list on the interface. The user can select different monitoring areas by clicking the mouse.

[0079] 10.2 Real-time monitoring and detection record display The real-time monitoring screen is displayed in the form of a video stream with a frame rate of 25 frames per second to ensure smooth viewing. At the same time, the inspection records are displayed on the interface in a list format, which includes fields such as inspection time, monitoring area, worker number, helmet wearing status, and safety clothing wearing status. Inspection record query and export functions are provided. The query function supports filtering by time (accurate to the minute), area (specific to the workshop, equipment number), inspection result (correct wearing / not wearing helmet, correct wearing / not wearing safety clothing), and other conditions. The export function supports CSV and Excel formats. Users can select the export format as needed to facilitate data analysis and processing. At the same time, alarm information is also integrated into the visual interface, displaying alarm records in eye-catching colors (such as red) and icons (such as alarm bells) to facilitate timely processing by management personnel.

[0080] Example 3 In another preferred embodiment, based on the above-mentioned embodiments 1 and 2, this embodiment provides a real-time visual Transformer-based intelligent detection system for power plant workers' safety helmets and safety clothing, which is used to implement the detection methods of embodiments 1 and 2 and includes the following modules: 1. Image acquisition module The image acquisition module consists of multiple high-definition cameras located in key areas of the power plant. The cameras utilize Sony IMX series sensors, offering high resolution and low-light performance. The cameras are connected via Ethernet cables to a network switch, which transmits image data to a server. Each camera is equipped with an independent power supply and protective housing to ensure stable operation in harsh industrial environments.

[0081] 2. Data processing module The data processing module is deployed on the server and implemented using Python and OpenCV (Open Source Computer Vision Library). First, the captured image data is decoded and converted, converting all image formats to RGB. Next, data cleaning is performed to remove blurry, dark, or bright image data that doesn't meet inspection requirements. Labeled data is stored in JSON format for easy access and processing.

[0082] 3. Model training module The model training module is implemented using the deep learning framework PyTorch. A GPU computing environment is set up on the server, using NVIDIA Tesla V100 graphics cards for model training. During training, distributed training technology is used to distribute training tasks across multiple graphics cards for parallel computing, improving training efficiency. Furthermore, a visual monitoring system is used to monitor the training process, displaying real-time changes in metrics such as loss function value and accuracy.

[0083] 4. System deployment module The system deployment module is responsible for installing the cameras in the designated locations and performing network configuration. During installation, professional installation tools are used to ensure the cameras are securely installed. Network configuration utilizes DHCP (Dynamic Host Configuration Protocol), automatically assigning IP addresses to cameras for easy management and maintenance. Firewall rules are also set to only allow access to the server from specific ports and IP addresses, ensuring system security.

[0084] 5. Real-time detection module The real-time detection module receives image data transmitted by the camera, first preprocesses the image, and then loads the trained model weights for detection. Multi-threading technology is used during the detection process to improve real-time detection. If a worker is detected not wearing a hard hat or safety clothing correctly, the violation information (including image, time, location, etc.) is stored in the database and an alarm mechanism is triggered.

[0085] 6. Data display module The data display module utilizes a World Wide Web (WWW) interface, developed using HTML (Hypertext Markup Language), CSS (Cascading Style Sheets), and JavaScript for the front-end, and the Flask framework for the back-end. Users access the WWW interface through a browser to log in and select monitoring options. Streaming media technology is used to display real-time monitoring footage to ensure smooth viewing. Inspection records are displayed in tables and charts, allowing users to query, sort, and export.

[0086] 7. Alarm module The alarm module interacts with the real-time detection module and the database. When a violation is detected, the alarm module first retrieves the violation information from the database and then issues an alarm signal based on pre-set alarm rules (such as violation type and number of violations). Alarm signals include audible alarms (via audio equipment connected to the server), light alarms (alarm indicators set up in the monitoring center), and alarm messages sent to administrators' mobile phones (via SMS gateways or instant messaging tools).

[0087] 8. Model update module The model update module regularly retrieves new image data from the database to update and train the real-time visual Transformer model. This update and training process utilizes an incremental learning approach, fine-tuning the existing model using new data. Furthermore, an update log is maintained, including information such as the update time, updated data volume, and model performance changes, to facilitate subsequent tracing and analysis.

[0088] 9. Data storage module The data storage module utilizes distributed database systems, such as the Hadoop Distributed File System (HDFS) and HBase (Hadoop database). Collected image data is stored in HDFS, while labeled data and detection results are stored in HBase. The database utilizes master-slave replication and backup strategies to ensure data security and reliability. Data is also encrypted to prevent leakage.

[0089] The real-time visual Transformer model employed in this paper boasts stronger feature extraction and global information perception capabilities than traditional convolutional neural network models. Traditional convolutional neural networks primarily rely on local receptive fields for feature extraction, making them prone to missed and false detections when inspecting helmets and safety clothing in complex scenarios. However, the real-time visual Transformer model, through its self-attention mechanism, can capture correlations between different regions in an image, improving detection accuracy.

[0090] In terms of data collection, this invention not only collects common image data but also conducts specialized data collection for the specific scenarios of power plants, increasing sample diversity. Furthermore, during the data annotation process, more detailed annotations are provided for helmets and safety clothing in specific scenarios, enabling the model to better learn these special features and improving its generalization capabilities in these specific scenarios.

[0091] In the preferred solution, the structure of the real-time visual Transformer model in Step 3 follows the encoder-decoder structure of the transformer, including a convolutional neural network, input embedding, positional encoding, a transformer encoder, and a transformer decoder, wherein the convolutional neural network is used to extract image features, and the transformer encoder and decoder are used for feature processing and classification prediction; the above settings enable the model to efficiently extract key information from the input image, and perform deep feature fusion and sequence prediction through the encoder-decoder architecture, effectively improving the processing speed and accuracy of real-time visual tasks.

[0092] In the preferred solution, in Step 1, data collection uses images from sources including on-site photographs, public image libraries, social media, and surveillance videos. Worker images under varying lighting conditions and angles are collected to increase sample diversity. This approach aims to improve the generalization capabilities of the image recognition algorithm, ensuring accurate identification of worker status in a variety of environments. Simultaneously, image preprocessing, such as denoising, contrast enhancement, and normalization, is performed to improve image quality and reduce recognition errors. Furthermore, images of workers in various occupations and attire are collected to further enrich the sample library and provide a solid foundation for subsequent image recognition algorithm training.

[0093] In the preferred solution, in Step 2, data annotation is performed using professional image annotation tools. Data is audited to ensure accuracy and consistency. If the annotation error rate exceeds a preset threshold, re-annotation is performed. This setting aims to improve the quality of the dataset and lay a solid foundation for subsequent model training. This step also includes regular training for annotators to continuously optimize the annotation process and ensure both efficiency and quality.

[0094] In the preferred solution, during the Step 3 model training step, the image size is unified through bilinear interpolation, data augmentation is performed, AdamW is selected as the optimizer, training parameters are set, and cross-validation is used to evaluate model performance. Training is stopped when the model's accuracy on the validation set reaches the preset target value. These settings ensure that the model has good generalization ability and stability. In addition, an early stopping mechanism is introduced to prevent overfitting, and a learning rate scheduler is used to dynamically adjust the learning rate to further optimize the training process and improve training efficiency.

[0095] In the preferred solution, in Step 4 system deployment, the camera has high-definition, night vision, and waterproof functions, and is installed in a position that can cover all areas that need to be monitored, and the camera installation angle is adjustable to adapt to the monitoring needs of different scenarios; the above settings ensure that there are no blind spots in monitoring. At the same time, the system is equipped with intelligent analysis software that can automatically identify abnormal behavior and issue real-time alarms, greatly improving monitoring efficiency and safety; in addition, all devices are connected to the cloud management platform to achieve remote monitoring and data storage.

[0096] In the preferred solution, in Step 5, real-time detection, the image is preprocessed before being detected using the real-time visual Transformer model, including denoising and contrast enhancement operations, to improve detection accuracy. The above settings also incorporate an adaptive lighting adjustment mechanism to ensure that the image can present the best visual effect under different lighting conditions. In addition, the edge detection algorithm is used to further optimize image details, providing a more accurate analysis basis for the Transformer model.

[0097] In the preferred solution, the specific steps of Step 6 data visualization are: using the PyQt5 framework to implement the user login and registration function, obtaining the monitoring name for the user to select, and displaying the real-time monitoring and detection records, while providing the query and export functions of the detection records; the above settings enable the user to intuitively view the system status and historical data; in addition, the Matplotlib library is used to draw charts of the monitoring data, including trend charts, bar charts, etc., so that users can quickly identify data changes and assist in decision-making analysis.

[0098] In the preferred embodiment, the method further includes an alarm step. When it is detected that a worker is not wearing a safety helmet or safety clothing correctly, the system automatically sends an alarm signal, which includes a sound alarm, a light alarm, or an alarm message sent to the manager's mobile phone. The above settings ensure that the safety regulations of the construction site are strictly implemented and effectively prevent the occurrence of safety accidents. At the same time, the immediate transmission of the alarm information enables the management personnel to respond quickly and take corresponding measures to protect the lives and property of the workers.

[0099] In a preferred embodiment, the method also includes a model update step, regularly collecting new image data and updating and training the real-time visual Transformer model to adapt to feature changes that may occur in different time periods and worker groups, ensuring detection accuracy and reliability. Model updates utilize incremental learning or full learning, and update logs are recorded. This not only enhances the system's adaptability but also facilitates subsequent tracking and analysis of model performance changes. Furthermore, by setting thresholds to monitor the difference in performance before and after model updates, the effectiveness of updates is automatically evaluated, and optimization strategies are adjusted promptly to ensure steady improvements in production efficiency and quality.

[0100] In the preferred embodiment, the method also includes a data storage step, storing the collected image data, annotation data, and detection result data in a database. The database adopts a distributed storage architecture to improve the security and scalability of data storage, and at the same time backs up the data to prevent data loss, and the data storage module supports data encryption and access permission control; the above settings ensure the security and integrity of the data during the storage process. Furthermore, the stored data is deeply analyzed through the data analysis and mining module to extract valuable information, providing strong support for subsequent decision-making and optimization.

[0101] In the preferred solution, the image acquisition module is also associated with an environmental parameter acquisition unit, which collects the light intensity, temperature, and humidity parameters of the camera's environment, and stores them in association with the collected image data; the above settings can provide key environmental background information in subsequent image analysis, improve the accuracy and reliability of image recognition, and at the same time, provide data support for the intelligent maintenance of the monitoring system, making it easier to promptly discover and deal with image quality problems caused by environmental factors.

[0102] In the preferred solution, the model training module is provided with a training monitoring unit, which monitors the loss function value and accuracy index during the model training process in real time, and automatically adjusts the training parameters or suspends training and generates an abnormality report when the index is abnormal; the above settings ensure the efficiency and stability of model training; in addition, the training monitoring unit can also record training logs to provide data support for model optimization and fault diagnosis, further improving the convenience and reliability of model development.

[0103] In the preferred solution, the data display module is equipped with a statistical analysis unit, which performs statistical analysis on the detection records to generate statistical reports on violation rates and violation type distribution charts for different time periods and different areas; the above settings can intuitively reflect the temporal and spatial distribution patterns of violations and the main types of violations, providing a scientific basis for management personnel so that they can formulate targeted management measures and effectively curb the occurrence of violations.

[0104] In summary, the present invention provides an intelligent detection method and system for power plant workers' safety helmets and safety clothing based on real-time visual Transformer. Aiming at the problems in the field of power plant workers' safety helmet and safety clothing wearing supervision in the power industry, which are that traditional manual inspections cannot be fully monitored in real time, are labor-intensive and easily interfere with the normal work of workers, the visual Transformer technology is innovatively introduced. This technology realizes the automatic recognition of the wearing status of workers' safety helmets and safety clothing through deep learning target detection methods, significantly improves the monitoring efficiency and accuracy, and effectively reduces the risk of accidents. In terms of data collection, the present invention obtains samples from multiple channels such as on-site photos, public image libraries, social media and surveillance videos, ensuring the diversity and representativeness of the data. At the same time, professional image annotation tools are used for annotation, and strict data review is carried out to ensure the accuracy and consistency of the annotation, thereby improving the quality of the data set. During the model training process, the present invention uses data enhancement technologies such as rotation and scaling, and makes use of the AMD EPYC 7542 32-core high-performance processor and the NVIDIA GeForce RTX3090 (24G) high-performance graphics card to ensure the efficiency and effect of training; in addition, the present invention also rationally arranges cameras with high-definition, night vision, waterproof and other functions according to the actual needs of the power plant, realizing all-round monitoring; and realizes data visualization through the PyQt5 framework, providing practical functions such as user login and registration, monitoring name selection, real-time monitoring playback and detection record display, thereby enhancing the interactivity and practicality of the system; in response to the problem that traditional machine learning detection technology relies on manually designed feature templates and has low detection accuracy, the present invention uses deep learning technology to enable the model to adaptively learn appropriate features from manually labeled data, thereby improving detection accuracy; the real-time visual Transformer model designed by the present invention has a unique structure, which extracts image features through convolutional neural networks, and then converts them into image sequences through input embedding and positional encoding operations, and finally through transformer The encoder and decoder perform feature encoding and prediction, achieving efficient and accurate detection; the system can adapt to the complex working environment of the power plant, including different lighting conditions, shooting angles and the diversity of workers' clothing, and demonstrates strong robustness and practicality; ultimately, the present invention realizes unmanned remote detection of power plant workers' safety helmets and safety clothing, greatly improving detection efficiency, reducing the workload of inspectors, and avoiding the interference of inspections on workers' work, significantly improving the safety management level and supervision informationization level of the power plant, and providing strong guarantees for the safe production of the power plant.

Claims

1. An intelligent detection method for power plant workers' helmets and safety clothing based on real-time visual Transformer, characterized by: The following steps are involved: Step 1: Data collection: Select appropriate image sources and collect image data of workers wearing or not wearing safety helmets and safety clothing in different scenarios; Step 2: Data labeling: label the collected image data as "safety helmet", "safety suit" and "worker", and draw bounding boxes; Step 3: Model training: On the server, use the labeled data to train the real-time visual Transformer model and save the trained model weights. Step 4: System deployment: Install high-definition cameras in key areas of the power plant and connect them to a wireless network to transmit image data to the server. Step 5: Real-time detection: After acquiring the image data, resize the image, load the trained model weights, and use the real-time visual Transformer model to detect the image to determine whether the worker is wearing the helmet and safety clothing correctly. Step 6: Data visualization, displaying the test results on the monitor, including real-time monitoring playback and test record display.

2. The intelligent detection method for power plant workers' safety helmets and safety clothing based on real-time visual Transformer according to claim 1 is characterized by: The structure of the real-time visual Transformer model in Step 3 follows the encoder-decoder structure of the transformer, including a convolutional neural network, input embedding, positional encoding, a transformer encoder, and a transformer decoder. The convolutional neural network is used to extract image features, and the transformer encoder and decoder are used for feature processing and classification prediction.

3. The intelligent detection method for power plant workers' safety helmets and safety clothing based on real-time visual Transformer according to claim 2 is characterized by: In the Step 1 data collection step, image sources include on-site photos, public image libraries, social media, and surveillance videos. During the collection process, images of workers under different lighting conditions and angles are collected to increase sample diversity.

4. The intelligent detection method for power plant workers' safety helmets and safety clothing based on real-time visual Transformer according to claim 3 is characterized by: In Step 2, data annotation is performed using professional image annotation tools, and data review is performed to ensure the accuracy and consistency of the annotations. If the annotation error rate exceeds a preset threshold, re-annotation is performed.

5. The intelligent detection method for power plant workers' safety helmets and safety clothing based on real-time visual Transformer according to claim 4 is characterized by: In the Step 3 model training step, the image size is unified by bilinear interpolation, data enhancement is performed, AdamW is selected as the optimizer, and training parameters are set. At the same time, the cross-validation method is used to evaluate the model performance. Training is stopped when the accuracy of the model on the validation set reaches the preset target value.

6. The intelligent detection method for power plant workers' safety helmets and safety clothing based on real-time visual Transformer according to claim 5 is characterized by: In Step 4, the system deployment step, the camera has high-definition, night vision, and waterproof functions, and is installed in a position that can cover all areas that need to be monitored. The camera installation angle is adjustable to meet the monitoring needs of different scenarios.

7. The intelligent detection method for power plant workers' safety helmets and safety clothing based on real-time visual Transformer according to claim 6 is characterized by: In Step 5, real-time detection, before using the real-time visual Transformer model to detect the image, the image is preprocessed, including denoising and contrast enhancement operations, to improve detection accuracy.

8. The intelligent detection method for power plant workers' safety helmets and safety clothing based on real-time visual Transformer according to claim 7 is characterized by: The specific steps of Step 6 data visualization are: using the PyQt5 framework to implement the user login and registration function, obtaining the monitoring name for the user to choose, and displaying the real-time monitoring and detection records, while providing the query and export functions of the detection records.

9. The intelligent detection method for power plant workers' safety helmets and safety clothing based on real-time visual Transformer according to claim 8 is characterized by: The method also includes an alarm step. When it is detected that a worker is not wearing a safety helmet or safety clothing correctly, the system automatically sends an alarm signal, which includes a sound alarm, a light alarm, or sending an alarm message to the manager's mobile phone.

10. The intelligent detection method for power plant workers' safety helmets and safety clothing based on real-time visual Transformer according to claim 9 is characterized by: The method also includes a model updating step, which regularly collects new image data and updates and trains the real-time visual Transformer model to adapt to feature changes that may occur in different time periods and different worker groups, ensuring the accuracy and reliability of detection. The model update adopts incremental learning or full learning, and records update logs. The method also includes a data storage step, which stores the collected image data, labeled data, and detection result data in a database. The database adopts a distributed storage architecture to improve the security and scalability of data storage, and backs up the data to prevent data loss. The data storage module supports data encryption and access permission control.

11. The intelligent detection system for power plant workers' helmets and safety clothing based on real-time visual Transformer is characterized by: The method for intelligently detecting safety helmets and safety clothing for power plant workers based on real-time visual Transformer according to claim 10 is implemented, comprising: Image acquisition module: used to collect image data of workers wearing or not wearing safety helmets and safety clothing in different scenarios; Data processing module: used to label and preprocess the collected image data; Model training module: used to train the real-time visual Transformer model on the server and save the trained model weights; System deployment module: used to install high-definition cameras in key areas of the power plant and connect the cameras to the wireless network; Real-time detection module: used to obtain image data, resize the image, load the trained model weights, and detect the image using the real-time visual Transformer model; Data display module: used to display the test results on the display, including real-time monitoring playback and test record display; Alarm module: used to send out an alarm signal when it detects that a worker is not wearing a safety helmet or safety clothing correctly; Model update module: used to regularly collect new image data and update the training of the real-time visual Transformer model; Data storage module: used to store the collected image data, annotation data, and detection result data in the database.

12. The intelligent detection system for power plant workers' safety helmets and safety clothing based on real-time visual Transformer according to claim 11 is characterized by: The image acquisition module is also associated with an environmental parameter acquisition unit, which collects the light intensity, temperature, and humidity parameters of the camera environment and stores them in association with the collected image data; the model training module is provided with a training monitoring unit, which monitors the loss function value and accuracy index during the model training process in real time, and automatically adjusts the training parameters or suspends training and generates an abnormality report when the index is abnormal; the data display module has a statistical analysis unit, which performs statistical analysis on the detection records and generates violation rate statistical reports and violation type distribution charts for different time periods and different areas.

Citation Information

Patent Citations

  • Safety helmet and safety clothing wearing behavior intelligent detection method

    CN117636472A