Improved LightGBM imbalance classification method and system for factory safety early warning
Through the improved LightGBM unbalanced classification method, combined with deep separable convolutional CNN and FuzzySMOTE technology, the problems of false alarms and missed alarms in factory safety warnings are solved, and efficient and accurate safety warnings are achieved.
Patent Information
- Application Number
- CN202510078900.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing technology has problems such as false alarms, missed alarms, high computing resources, and poor detection results for complex violations in factory safety warnings, making it difficult to achieve efficient and accurate safety warnings.
The improved LightGBM unbalanced classification method is adopted, combined with deep separable convolutional CNN for feature extraction, FuzzySMOTE is used to resample a few class samples, and EFB and GOSS algorithm are introduced for feature bundling and sampling, and the loss function and parameter tuning are optimized.
It significantly improves the accuracy of identification of complex violations, reduces noise interference, enhances the identification ability of a few types of samples, improves the performance and generalization capabilities of the classification model, and achieves efficient and accurate safety warnings.
Smart Images

Figure CN120217172A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of imbalanced classification, and particularly relates to an improved LightGBM imbalanced classification method and system for plant area safety warning. Background Technique
[0002] Plant area safety has now become the most important part of enterprises. However, there are still many safety hazards in most modern production environments. Due to the fluke mentality and weak safety awareness of on-site workers, accidents will directly or indirectly occur, causing irreparable consequences. Therefore, while maximizing the safety awareness of workers, it is necessary to reduce safety hazards as much as possible. How to supervise and manage the on-site safety and timely give early warnings for potential accidents is an urgent problem that enterprises need to solve and are most concerned about.
[0003] Nowadays, many production plants still use a large number of manual labor for safety supervision, which not only consumes a large amount of manpower and financial resources of enterprises, but also, after long hours of work, the attention of safety supervisors will inevitably become inattentive, and their eyes and spirits will be fatigued, which may cause serious accidents. To avoid the occurrence of the above situation, on the premise of meeting the normal operation of the site, it is necessary to make the safety supervision achieve the purposes such as accurate and rapid early warning of safety hazards, low cost, and high supervision efficiency.
[0004] With the complexity of the plant area environment, the difficulty of safety supervision increases. Usually, algorithms need to be used to simulate and calculate problems to obtain the optimal results. At present, the solution methods for plant area safety warning can be divided into traditional object detection algorithms and detection algorithms based on deep learning. However, as the problem scale increases, traditional object detection algorithms seem powerless. Therefore, many domestic and foreign research scholars now focus on using deep learning or ensemble learning to conduct research on plant area safety warning. The main current solution methods for plant area safety warning problems and their defects are as follows:
[0005] Through the above analysis, the problems and defects existing in the prior art are:
[0006] (1) Detection algorithms based on convolutional neural network (CNN) and long short-term memory network (LSTM): long execution time, strong data dependence, and problems of false alarms and missed alarms. Therefore, this method consumes a large amount of time in plant area safety warning and has low efficiency;
[0007] (2) Detection algorithms based on autoencoder and generative adversarial network (GAN): sensitive to environmental factors, high demand for computing resources, unable to learn sufficient features for complex or rare violation behaviors, resulting in poor detection effects of abnormal behaviors. Therefore, the early warning quality solved by this method cannot be guaranteed and the efficiency is low;
[0008] (3) Detection algorithms based on reinforcement learning: Difficult to design the reward function, long training process, low sample efficiency, and excessive computational overhead.
[0009] (4) Detection algorithms based on LightGBM: An ensemble learning model, with complex parameter tuning and vulnerable to noise interference. Therefore, it is difficult for the original LightGBM model to find the optimal solution for factory area safety warning. Summary of the Invention
[0010] In view of the problems existing in the prior art, the present invention provides an improved LightGBM imbalance classification method for factory area safety warning.
[0011] The present invention is implemented as follows. An improved LightGBM imbalance classification method for factory area safety warning includes:
[0012] Step 1, use depthwise separable convolution CNN for feature extraction;
[0013] Step 2, use FuzzySMOTE to resample the minority class samples, generate n minority class samples, and initialize the sample weights;
[0014] Step 3, put the dataset with weights into the improved LightGBM model for training. The feature vectors after feature extraction are output as a set of bundled features through the EFB exclusive feature bundling of LightGBM. The specific algorithm is shown in Algorithm 1.
[0015] The specific algorithm process of the EFB algorithm is shown in Algorithm 1:
[0016]
[0017] Step 4, the set of bundled features after output is obtained as a sampled data subset through the GOSS sampling strategy of LightGBM, as shown in Algorithm 2 specifically.
[0018] The specific algorithm process of the GOSS algorithm is shown in Algorithm 2:
[0019]
[0020]
[0021] Step 5, calculate the split gain for the sampled data subset. The specific formula is as follows:
[0022]
[0023] where D represents the sample set of the current node; D L , D R represent the sample sets of the left and right child nodes respectively; nD , n L , n R respectively represent the sizes of the corresponding sample sets; gi is the gradient of the i-th sample.
[0024] Step 6, when the number of iterations T is satisfied, integrate the classifiers in the classifier pool according to the learning rate and the gradient of its loss function to obtain the final classifier;
[0025] Input: the number of generated minority class samples n, the number of iterations T, the initial population size of butterflies n_butterflies, the optimal butterfly ratio p_ratio, the retention ratio a of large-gradient data, the sampling ratio b of small-gradient data, the maximum conflict threshold K;
[0026] Output: the final classifier.
[0027] Furthermore, the feature extraction:
[0028] Introduce separable convolution to optimize the traditional convolution operation. Depthwise separable convolution decomposes the standard convolution into two steps: depthwise convolution and pointwise convolution.
[0029] Furthermore, the process of the improved FuzzySMOTE:
[0030] Calculate the fuzzy membership degree: For the minority class sample x i and its potential nearest neighbor sample x j ; Define the fuzzy membership degree function as:
[0031]
[0032] where x i is the minority class sample, is the j-th minority class nearest neighbor of x i , is the j-th majority class nearest neighbor of x i , k is the number of nearest neighbors, and d(·,·) is the Euclidean distance between samples;
[0033] Membership degree normalization: Process it to the interval [0,1], and the formula is as follows:
[0034]
[0035] The probability of selecting the reference sample: Divide the normalized membership degree of each sample by the sum of the membership degrees of all samples to obtain a probability distribution; that is, calculate the probability of selecting each minority class sample, and the formula is as follows:
[0036]
[0037] where, is the sum of all normalized membership degrees;
[0038] Generate a new sample: The sample generation formula is as follows:
[0039] x new = x i + λ(x j - x i )
[0040] where x i is the reference sample selected according to the probability P(x i ), x j is randomly selected from among the k nearest neighbors of x i , and λ is a random number in the range [0, 1].
[0041] Furthermore, the classification:
[0042] Based on the original LightGBM algorithm, introduce FocalLoss focusing loss to its original loss function. The formula is as follows:
[0043] FL(p t ) = -α t (1 - p t ) γ log(p t )
[0044] where p t represents the predicted probability of the model for the true class, and is defined as:
[0045]
[0046] where p is the probability of predicting the positive class, and y is the true label; α t is the balance factor, used to balance the importance of positive and negative samples; γ is the adjustment factor, controlling the degree of attention to easy and difficult samples; γ = 0 degenerates to the ordinary cross-entropy loss, γ > 0 increases the attention to difficult-to-classify samples and suppresses the contribution of easy-to-classify samples to the loss; and for key parameters such as the learning rate, introduce the Butterfly Optimization Algorithm (BOA) for optimization.
[0047] Another object of the present invention is to provide an improved LightGBM unbalanced classification system for factory area safety warning, including:
[0048] A feature extraction module, used to perform feature extraction using depthwise separable convolution CNN;
[0049] A sampling module, used to resample the minority class samples using FuzzySMOTE, generate n minority class samples, and initialize the sample weights;
[0050] The training module uses the EFB and GOSS algorithms to bundle and output a set of feature bundles for the feature vectors after feature extraction, and collects large-gradient samples and small-gradient samples accordingly to form a sampled data subset, and constructs a decision tree by calculating the split gain.
[0051] The iteration module, when the iteration number T is not satisfied, still adopts the above strategy to construct gradient samples and generate a decision tree.
[0052] The integration module is used to integrate all classifiers in the classifier pool according to the learning rate and the gradient of its loss function after the iteration number T is satisfied to obtain the final classifier;
[0053] Another object of the present invention is to provide a computer device, the computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the improved LightGBM unbalanced classification method for factory area safety warning.
[0054] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor executes the steps of the improved LightGBM unbalanced classification method for factory area safety warning.
[0055] Another object of the present invention is to provide an information data processing terminal, and the information data processing terminal is used to implement the improved LightGBM unbalanced classification system for factory area safety warning.
[0056] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0057] First, the present invention is a method improved based on LightGBM. In the original LightGBM algorithm, the positive and negative samples of the training data are usually balanced. When the data of illegal behaviors is much less than the data of normal behaviors, LightGBM may have a high bias towards normal behaviors, resulting in the problem of missed reports. When dealing with low-quality data, LightGBM is easily interfered by noise data, resulting in the model being unable to correctly identify illegal behaviors.
[0058] Inspired by deep learning, the present invention proposes an improved LightGBM algorithm for the problem of factory area safety warning. Deep learning has powerful feature extraction and pattern recognition capabilities when dealing with large amounts of data, and can automatically learn complex non-linear relationships from the data. The present invention introduces the multi-layer feature extraction ability of deep learning, deeply processes the original data through a neural network, extracts key features and inputs them into LightGBM, enabling LightGBM to learn in a richer feature space, thereby improving the recognition accuracy of complex violation behavior patterns.
[0059] The purpose of the present invention is to introduce an imbalanced classification model, improve its sampling method, feature extraction and classification algorithm, so as to effectively solve the problems in factory area safety warning. The present invention proposes a method for dynamically monitoring whether employees have violation behaviors for the safety behaviors in the factory area (such as whether to wear a safety helmet, whether to wear work clothes, etc.). By introducing the fuzzy logic idea to improve the original SMOTE method, introducing an improved CNN for feature extraction, improving the LightGBM loss function, and introducing the butterfly optimization algorithm to determine the setting of key parameters, this imbalanced classification model is determined. Since the SMOTE algorithm is used to process high-dimensional imbalanced data, the distribution of minority class samples is often sparser. Therefore, before resampling with the improved fuzzy SMOTE (FuzzySMOTE), the present invention extracts the features of the samples through a depthwise separable convolutional CNN to reduce the sample feature dimension. Then, the FuzzySMOTE algorithm and the improved LightGBM algorithm are combined, and the FuzzySMOTE algorithm is used in the process of weak classifier iteration to resample the minority class sample features extracted by the improved CNN, reducing the imbalance ratio of the samples. Experiments prove that the method proposed by the present invention well solves the deficiencies existing in the original model in the factory area safety problem and improves the performance and generalization ability of the classification model.
[0060] Inspired by a variety of advanced technologies, the present invention proposes an improved LightGBM balanced classification model aiming to solve the problem of factory area safety warning. Through innovative improvements in the data sampling stage, feature extraction stage and classification algorithm stage, the present invention not only improves the performance of the classification model, but also significantly enhances the generalization ability and adaptability of the model. Compared with the traditional factory area safety warning model, the improved scheme based on the present invention has the following significant advantages:
[0061] 1. Improve the anti-interference ability to noise points and enhance the minority class recognition ability
[0062] The improved process of the present invention has stronger anti-interference ability, especially outstanding in dealing with boundary noise points in complex industrial environments. Through optimization in the sampling and feature extraction stages, the improved model can effectively enhance the recognition ability of minority class samples, avoiding the problem of early warning failure caused by the neglect or inaccurate classification of minority class samples in traditional models. Compared with some complex deep learning end-to-end models, the process of the present invention adopts a design that separates feature extraction and classification. By using depthwise separable convolutions to extract high-quality features and combining with a lightweight FuzzyFSSLightGBM-like model, the model greatly reduces the computational complexity while ensuring classification accuracy, shortens the training and inference time, and significantly improves efficiency.
[0063] 2. Comprehensive optimization steps, focusing on solving the imbalanced classification problem
[0064] Regarding the problem of sample imbalance, the present invention introduces the FuzzySMOTE algorithm, which effectively balances the proportion of samples in each category, improves the generation quality of minority class samples, and thus enhances the recognition ability of the classification model for minority class samples. At the same time, depthwise separable convolutions perform in-depth processing on image data in the feature extraction stage, extracting more discriminative high-quality features, which provides richer information for subsequent classification. In addition, the LightGBM algorithm is introduced in the classification stage, and the Focal Loss focusing loss is adopted in its loss function, further optimizing the model's processing ability for difficult-to-classify samples, especially showing stronger robustness and classification accuracy in the classification of minority class samples. To further improve the model performance, the present invention also uses the Butterfly Optimization Algorithm (BOA) to finely tune the key parameters of the model, ensuring that each step can work efficiently in coordination, focusing on the key pain points of imbalanced classification, and thus overall improving the classification performance.
[0065] 3. Balance of high performance and low cost, meeting the requirements of industrial environments
[0066] The overall method of the present invention not only has high performance but also has the advantages of low cost, flexibility, and easy deployment, making it very suitable for imbalanced image classification tasks in industrial scenarios. Especially in high-demand application scenarios such as factory area safety monitoring, the model can not only ensure high classification accuracy but also balance resource consumption, enabling the method to perform excellently in practical environments with limited computing resources. Compared with traditional methods, the present invention can complete the efficient classification of complex data with lower computational overhead, ensuring the stability and accuracy of the factory area safety warning system, especially maintaining high classification accuracy and reliability in the face of imbalanced data and more noise points.
[0067] In summary, through improvements in multiple key aspects, the present invention has formed an efficient, accurate, and highly adaptable classification model, significantly enhancing the performance of the factory area safety warning system. By introducing advanced technical means and optimizing the model structure, the present invention can not only solve various problems encountered in traditional methods but also provide reliable technical support in a changing industrial environment, offering more accurate warning services for factory area safety.
[0068] Second, as auxiliary evidence of the inventiveness of the claims of the present invention, it is also reflected in the following important aspects:
[0069] (1) The expected benefits and commercial value after the transformation of the technical solution of the present invention are:
[0070] Expected benefits: In traditional factory areas, when employees violate regulations, it is often impossible to conduct efficient identification, detection, and system alarms. If serious violations occur and there are missed detections or false detections, it often leads to safety accidents and even huge economic losses. This solution uses the FuzzySMOTE algorithm to handle the problem of imbalanced data, especially strengthening the ability to identify various types of violations. Through high-quality features extracted by the depthwise separable convolutional network and combined with the improved LightGBM classification algorithm, it can accurately identify violations in surveillance images, greatly reducing the cases of missed detections and false detections. At the same time, the system constructed by the model can analyze violation data from multiple dimensions, including the personal dimension, department dimension, time dimension, etc., to generate personalized violation reports. This all-round data analysis ability enables employees to clearly understand their own violation situations, greatly improving their safety awareness. Managers can also carry out targeted safety education and training based on statistical data to improve the training effect. Therefore, through this solution, it is possible to effectively regulate employees' behaviors, greatly reduce the high incidence of safety accidents, avoid unnecessary economic losses, and provide good protection for factory area safety.
[0071] Commercial value: The system constructed by this method significantly reduces the workload of manual inspections by intelligently identifying violations in surveillance images. The daily / weekly / monthly violation analysis reports automatically generated by the system provide clear and intuitive data support for management. Through the multi-dimensional analysis function of the system, managers can quickly identify high-incidence violation areas and time periods and formulate targeted improvement measures. Based on the statistics of personal violation data, enterprises can accurately identify employees who need key training, avoiding the waste of resources caused by large-scale indiscriminate training. The violation behavior data continuously collected and analyzed by the system will form important digital assets of the enterprise. These data not only support current safety management but can also be used to establish prediction models to provide data support for future management decisions. Through the digital and intelligent management method of the system, it promotes the overall upgrade of the enterprise's safety management system, improves the enterprise's modern management level, and enhances the enterprise's competitiveness.
[0072] (2) The technical solution of the present invention fills the technical gaps at home and abroad in the industry:
[0073] This solution realizes the full-process optimization from data processing, feature extraction to classification decision-making, and establishes a complete technical solution. This systematic optimization strategy fills the technical gaps in the field of industrial image classification. By introducing fuzzy logic into the processing of unbalanced data in industrial scenarios, the limitations of the traditional SMOTE algorithm in processing complex industrial images are solved. This algorithm adaptively adjusts the sample generation strategy through fuzzy rules, significantly improving the quality of synthetic samples and providing a new technical idea for the industry. Based on AlexNet, a depthwise separable convolution structure is innovatively introduced, which not only maintains the original feature extraction ability but also reduces the model's computational complexity by more than 60%. This improvement enables the system to achieve real-time processing on ordinary hardware, lowering the threshold for technology application. In industrial image classification, Focal Loss is introduced into the LightGBM framework, creatively solving the problem of sample imbalance. This combination gives full play to the efficiency advantage of LightGBM and the focusing characteristics of Focal Loss, realizing the accurate identification of violations. The butterfly optimization algorithm is applied to system parameter optimization, achieving automatic parameter tuning. This mechanism greatly reduces the system debugging difficulty and improves the system's practicality. The overall solution realizes the intelligent identification and statistical analysis of violations, and this innovation provides a new technical means for industrial safety management.
[0074] (3) The technical solution of the present invention solves the technical problems that people have been eager to solve but have never succeeded in:
[0075] a. In industrial scenarios, the proportion of violation samples is often extremely small, and traditional methods are difficult to effectively handle. This solution innovatively solves this problem through the FuzzySMOTE algorithm, enabling the model to accurately identify various violations.
[0076] b. The monitoring images at industrial sites are often affected by factors such as lighting and occlusion, and traditional methods are difficult to ensure the recognition effect. This solution realizes stable recognition in complex environments through data augmentation, improved feature extraction networks, and multi-layer optimization strategies.
[0077] c. Previous methods either have insufficient accuracy or cannot meet the real-time processing requirements. This solution successfully realizes the unity of high accuracy and real-time processing through depthwise separable convolution and model optimization. Traditional models often overfit to specific scenarios and have poor generalization ability. This solution significantly improves the system's environmental adaptability through innovative data augmentation and model design.
[0078] d. Deep learning models are often criticized as "black boxes" and are difficult to interpret. This solution improves the interpretability of model decisions through visualization techniques and decision analysis tools.
[0079] (4) The technical solution of the present invention overcomes technical biases:
[0080] a. Sample bias: Through the FuzzySMOTE algorithm, the system overcomes the preference for mainstream samples and can treat various types of violations fairly. Especially for low-frequency but important violations, it maintains a high degree of sensitivity.
[0081] b. Feature bias: The improved deep learning model avoids over-reliance on specific features and ensures the comprehensiveness and fairness of recognition through multi-dimensional feature fusion.
[0082] c. Scenario bias: The system overcomes the preference for specific application scenarios through rich data augmentation and model optimization and can adapt to various industrial environments.
[0083] d. Decision bias: The analysis and decision-making mechanism based on objective data avoids biases brought by human subjective judgments and ensures the fairness of management decisions.
[0084] e. Evaluation index bias: The system adopts multi-dimensional evaluation indexes, avoids one-sidedness caused by a single index, and comprehensively and objectively reflects the system performance.
[0085] f. Effect evaluation bias: Through long-term data accumulation and multi-angle analysis, the objectivity and reliability of the system evaluation results are ensured.
[0086] Third, in the field of factory area safety warning, unbalanced data is a common problem. For example, the amount of data on dangerous events or violations is much less than that of normal data, resulting in traditional classification models being overly biased towards the majority class (normal events) and ignoring the detection of the minority class (abnormal events). When dealing with minority class samples, the existing methods usually have unsatisfactory effects and it is difficult to achieve a balance between high precision and high recall rate.
[0087] Traditional safety warning systems rely on simple statistics or manual feature engineering in feature extraction and are difficult to capture complex data patterns, especially when dealing with multi-source heterogeneous data (such as sensor data, image data), which limits the performance of classification models.
[0088] Existing classification models have poor adaptability in dynamic environments, especially lacking optimization in sample weight adjustment and weak classifier integration, resulting in classifiers being prone to overfitting or performance degradation in actual applications.
[0089] By combining the FuzzySMOTE resampling technique and the improved LightGBM model, the present invention optimizes the distribution of minority class samples, enabling the model to effectively identify minority class events, while reducing the impact of majority class samples on the classification results, and significantly improving the classification accuracy and recall rate.
[0090] The use of depthwise separable convolutional CNN enables efficient feature extraction of complex data, capable of capturing deeper pattern information, applicable to multi-source data such as plant sensor data or image data, and enhancing the overall performance of the classification model.
[0091] By dynamically adjusting the sample weights and optimizing the integration strategy of weak classifiers, the model is gradually optimized in multiple rounds of iteration to ensure the stability and generalization ability of the classifier. At the same time, the addition of sample weight threshold limits and a secondary resampling mechanism enhances the classifier's ability to detect abnormal samples.
[0092] In the classifier integration stage of the present invention, multiple weak classifiers are fused into an efficient final classifier using optimized parameters such as the learning rate, capable of adapting to dynamic environmental changes and ensuring reliability in long-term use.
[0093] Through the above technological advancements, the present invention provides an efficient and reliable solution for the field of plant safety warning, overcomes many bottlenecks in the prior art, and significantly improves the intelligent level of industrial safety management.
[0094] Third, in plant safety warning, the classification data usually has serious imbalance, with a low proportion of minority class samples, resulting in insufficient recognition ability of traditional classification methods for minority class samples. At the same time, the feature data has a high dimension and redundancy, which is prone to introducing noise and affecting the accuracy of the classification model. In addition, during the training process of large-scale data sets, the traditional gradient boosting algorithm (GBDT) has low computational efficiency and is difficult to meet the real-time requirements. These problems significantly limit the application effect of the prior art in plant safety warning.
[0095] The present invention proposes an optimization method for imbalanced classification problems by combining depthwise separable convolutional neural network (CNN), fuzzy minority oversampling technique (FuzzySMOTE), and improved LightGBM algorithm. This solution solves the problems of low feature extraction efficiency, decreased classification performance caused by sample imbalance, and insufficient computational efficiency of large-scale data, and significantly improves the recognition ability and operation efficiency of the plant safety warning model.
[0096] Compared with traditional methods, the technical solution of the present invention has achieved technological progress in many aspects. The depthwise separable convolutional CNN effectively reduces the model complexity and significantly improves the efficiency of feature extraction. The Fuzzy Minority Oversampling Technique (FuzzySMOTE) balances the distribution of minority class samples and improves the model's recognition ability for minority class samples. The improved LightGBM model, through Exclusive Feature Bundling (EFB) and Gradient-based One-Side Sampling (GOSS), significantly improves the training speed and computational efficiency while ensuring the model accuracy.
[0097] The technical solution of the present invention is applicable to the factory area safety warning system that needs to process imbalanced data and has broad application prospects in many fields such as industry, energy, and chemical industry. By improving the recognition ability of minority class samples, optimizing the computational efficiency, and enhancing the adaptability of the classification model, the present invention significantly improves the accuracy and real-time performance of the factory area safety warning, thereby providing technical support for industrial safety guarantee, effectively reducing the risk of safety accidents, and having important economic and social value. Brief Description of the Drawings
[0098] Figure 1 is the flowchart of the improved LightGBM imbalanced classification method for factory area safety warning provided by the embodiment of the present invention.
[0099] Figure 2 is the structural block diagram of the improved LightGBM imbalanced classification system for factory area safety warning provided by the embodiment of the present invention.
[0100] Figure 3 is the schematic diagram of feature extraction provided by the embodiment of the present invention.
[0101] Figure 4 is the flowchart of the imbalanced classification model based on the improved LightGBM provided by the embodiment of the present invention.
[0102] Figure 5 is the ablation experiment index performance graph under the SafetyHelmetWearingDataset dataset provided by the embodiment of the present invention.
[0103] Figure 6 is the comparative experiment index performance graph under the SafetyHelmetWearingDataset dataset provided by the embodiment of the present invention.
[0104] Figure 7 is the ablation experiment index performance graph under the Safety Helmet andReflective Jacket dataset provided by the embodiment of the present invention.
[0105] Figure 8It is the performance graph of the comparison experiment indicators under the Safety Helmet and Reflective Jacket dataset provided by the embodiments of the present invention.
[0106] Figure 9 It is the ablation experiment indicator performance graph under the ImVisible: Pedestrian Traffic Light Dataset provided by the embodiments of the present invention.
[0107] Figure 10 It is the performance graph of the comparison experiment indicators under the ImVisible: Pedestrian Traffic Light Dataset provided by the embodiments of the present invention.
[0108] Figure 11 It is the system detection effect diagram provided by the embodiments of the present invention.
[0109] Figure 12 It is the visualization analysis effect diagram provided by the embodiments of the present invention.
[0110] Figure 13 It is the camera capture effect diagram provided by the embodiments of the present invention. Detailed implementation manners
[0111] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0112] As Figure 1 shown, an improved LightGBM unbalanced classification method for factory area safety warning provided by the embodiments of the present invention includes the following steps:
[0113] S101, using a depthwise separable convolutional CNN for feature extraction;
[0114] S102, using FuzzySMOTE to resample the minority class samples, generating n minority class samples, and initializing the sample weights;
[0115] S103, putting the dataset with weights into the improved LightGBM model for training, and outputting the bundled feature bundle set through the EFB exclusive feature bundling of LightGBM for the feature vector after feature extraction;
[0116] Step 1, using a depthwise separable convolutional CNN for feature extraction;
[0117] Step 2: Use FuzzySMOTE to resample the minority class samples, generate n minority class samples, and initialize the sample weights;
[0118] Step 3: Put the dataset with weights into the improved LightGBM model for training. The feature vectors after feature extraction are output as a set of bundled features through the EFB (Exclusive Feature Bundling) of LightGBM. The specific algorithm is shown in Algorithm 1.
[0119] The specific algorithm process of the EFB algorithm is shown in Algorithm 1:
[0120]
[0121] Step 4: The set of output bundled features is used to obtain a sampled data subset through the GOSS (Gradient-based One-Side Sampling) strategy of LightGBM, as shown in Algorithm 2 specifically.
[0122] The specific algorithm process of the GOSS algorithm is shown in Algorithm 2:
[0123]
[0124] Step 5: Calculate the split gain for the sampled data subset. The specific formula is as follows:
[0125]
[0126] where D represents the sample set of the current node; D L , D R respectively represent the sample sets of the left and right child nodes; n D , n L , n R respectively represent the sizes of the corresponding sample sets; gi is the gradient of the i-th sample.
[0127] Step 6: After meeting the iteration times T, integrate all the classifiers in the classifier pool according to the learning rate and the gradient of their loss functions to obtain the final classifier;
[0128] Input: The number of generated minority class samples n, the iteration times T, the initial population size of butterflies n_butterflies, the optimal butterfly ratio p_ratio, the retention ratio a of large-gradient data, the sampling ratio b of small-gradient data, the maximum conflict threshold K;
[0129] Output: The final classifier.
[0130] S104: The set of output bundled features is used to obtain a sampled data subset through the GOSS strategy of LightGBM and calculate the split gain;
[0131] S105. After meeting the iteration number T, all classifiers in the classifier pool are integrated according to the learning rate and the gradient of their loss functions to obtain the final classifier;
[0132] Input: The number of generated minority class samples n, the iteration number T, the initial population size of butterflies n_butterflies, the optimal butterfly ratio p_ratio, the retention ratio a of large gradient data, the sampling ratio b of small gradient data, the maximum conflict threshold K;
[0133] Output: The final classifier.
[0134] Feature extraction provided by the embodiments of the present invention:
[0135] Introduce separable convolution to optimize the traditional convolution operation. Depthwise separable convolution decomposes the standard convolution into two steps: depthwise convolution and pointwise convolution.
[0136] Calculating its classification error rate and weights provided by the embodiments of the present invention:
[0137] Calculating the fuzzy membership degree: For the minority class sample x i and its potential nearest neighbor sample x j ; Define the fuzzy membership function as:
[0138]
[0139] where x i is the minority class sample, is the j-th minority class nearest neighbor of x i , is the j-th majority class nearest neighbor of x i , k is the number of nearest neighbors, and d(·,·) is the Euclidean distance between samples;
[0140] Membership degree normalization: Process it to the interval [0,1]. The formula is as follows:
[0141]
[0142] Probability of selecting the reference sample: Divide the normalized membership degree of each sample by the sum of the membership degrees of all samples to obtain a probability distribution; that is, calculate the probability of selecting each minority class sample. The formula is as follows:
[0143]
[0144] where, is the sum of all normalized membership degrees;
[0145] Generating new samples: The sample generation formula is as follows:
[0146] x new = xi + λ(x j - x i )
[0147] where x i is a reference sample selected according to the probability P(x i ), and x j is randomly selected from among the k nearest neighbors of x i , and λ is a random number in the range [0, 1].
[0148] The classification provided by the embodiments of the present invention is as follows:
[0149] Based on the original LightGBM algorithm, a FocalLoss focusing loss is introduced into its original loss function, and the formula is as follows:
[0150] FL(p t ) = -α t (1 - p t ) γ log(p t )
[0151] where p t represents the predicted probability of the model for the true class, and is defined as:
[0152]
[0153] where p is the probability of predicting the positive class, and y is the true label; α t is a balance factor used to balance the importance of positive and negative samples; γ is a regulation factor that controls the degree of attention to easy and difficult samples; when γ = 0, it degenerates into the ordinary cross-entropy loss, and when γ > 0, it increases the attention to difficult-to-classify samples and suppresses the contribution of easy-to-classify samples to the loss; and for key parameters such as the learning rate, the Butterfly Optimization Algorithm (BOA) is introduced for optimization.
[0154] 1. Chemical Plant Area Hazardous Substance Leakage Safety Early Warning System
[0155] In a chemical plant area, the leakage of hazardous substances may lead to serious accidents. Based on the improved LightGBM imbalanced classification method of the present invention:
[0156] Data feature extraction: Use depthwise separable convolution CNN to extract features from the data of plant area sensors (such as gas detectors, temperature and humidity sensors), including changes in leakage gas concentration, environmental parameters, etc.
[0157] Data balance processing: Use FuzzySMOTE to resample the data of the minority class (such as leakage events) to generate more minority class samples and avoid the bias of the model towards the data of the majority class (normal operation).
[0158] Classifier training: Through an improved LightGBM model, decision trees are iteratively constructed, and an efficient classifier is trained to identify the characteristic patterns of leakage events.
[0159] Application: Deploy the final classifier into the plant monitoring system to achieve real-time monitoring and early warning of hazardous material leakage.
[0160] 2. Power plant employee behavior safety monitoring system
[0161] In a power plant, employees' violation behaviors (such as not wearing protective equipment, misoperation) are important inducements for safety accidents. Based on the method of the present invention:
[0162] Data feature extraction: Use depthwise separable convolutional CNN to extract behavior features from the video data of employees' behaviors captured by the video monitoring system.
[0163] Data balancing processing: Perform FuzzySMOTE resampling on the sample data of the minority class (violation behaviors) to enhance the model's ability to identify violation behaviors.
[0164] Classifier training: Through an improved LightGBM model, train on the unbalanced behavior data set, gradually obtain weak classifiers and construct a classifier pool.
[0165] Application: Integrate the final classifier into the plant behavior monitoring system, classify and detect employees' behaviors in real time, automatically mark and record violation behaviors, and issue safety early warnings in a timely manner.
[0166] As Figure 2 shown, an improved LightGBM unbalanced classification system for plant safety early warning provided by an embodiment of the present invention includes:
[0167] A feature extraction module, used to perform feature extraction using depthwise separable convolutional CNN;
[0168] A sampling module, used to resample the minority class samples using FuzzySMOTE, generate n minority class samples, and initialize the sample weights;
[0169] A training module, using the EFB and GOSS algorithms to bundle and output a set of feature bundles for the feature vectors after feature extraction, and collect the large-gradient samples and small-gradient samples accordingly to form a sampled data subset, and construct a decision tree by calculating the split gain; perform the next round of loop;
[0170] An iteration module, when the iteration times T are not satisfied, still adopt the above strategy to construct gradient samples and generate decision trees.
[0171] An integration module, which is used to integrate the weak classifiers in the classifier pool according to the optimized parameters such as the learning rate after the iteration number T is satisfied, so as to obtain the final classifier;
[0172] In the feature extraction module, the system receives various sensor signals or monitoring data from the factory area (such as temperature and humidity sensors, gas detectors, video monitoring, etc.). The depthwise separable convolutional CNN is used to process the input data to extract key features. These features can capture deep patterns in the original data, such as the features of abnormal environmental changes or unsafe behaviors, and provide high-quality inputs for subsequent classification tasks.
[0173] The sampling module classifies the processed feature data, especially for minority class samples (such as abnormal event data). The FuzzySMOTE method is used to resample the minority class data to generate additional samples to balance the data distribution, and at the same time initialize the weights for each sample to ensure that the influence weights of the samples on the model can be dynamically adjusted during the subsequent training process. This step significantly enhances the system's ability to process imbalanced data.
[0174] In the training module, the resampled dataset is input into the improved LightGBM model for training. The EFB and GOSS algorithms, which are two prominent advantages of LightGBM, are respectively used to integrate the feature subsets, and the size of the split gain is calculated to judge the generation of decision trees.
[0175] When the iteration number T is not satisfied, the above strategy is still adopted for gradient sample construction and decision tree generation.
[0176] When the iteration number reaches the set threshold T, the integration module integrates all the weak classifiers in the classifier pool according to the learning rate and the gradient of the loss function to form the final classifier. The integration process optimizes the weight distribution among the classifiers to ensure that the final classifier has stronger robustness and generalization ability. The final classifier outputs the classification result, which is used as the decision basis for the factory area safety warning, can accurately identify and prompt potential safety hazards, and provides technical support for real-time monitoring and management.
[0177] The purpose of the present invention is to introduce an unbalanced classification model, improve its sampling method, feature extraction and classification algorithm, so as to effectively solve the problems of factory safety warning. The present invention proposes a method for dynamically monitoring whether employees have violations in terms of safety behaviors in the factory (such as whether to wear safety helmets, whether to wear work clothes, etc.). This unbalanced classification model is determined by introducing fuzzy logic ideas to improve the original SMOTE method, introducing improved CNN for feature extraction, improving LightGBM loss function, and introducing butterfly optimization algorithm to determine the setting of key parameters. Because the distribution of minority class samples is often sparser when the SMOTE algorithm processes high-dimensional unbalanced data. Therefore, the present invention extracts features of samples by deep separable convolutional CNN before resampling with the improved fuzzy SMOTE (FuzzySMOTE) to reduce the sample feature dimension. Then the FuzzySMOTE algorithm is combined with the improved LightGBM algorithm, and the FuzzySMOTE algorithm is used in the process of weak classifier iteration to resample the minority class sample features extracted by the improved CNN to reduce the imbalance ratio of samples. Experiments have shown that the method proposed in the present invention has effectively solved the shortcomings of the original model in factory safety issues and improved the performance and generalization ability of the classification model.
[0178] Basic steps of the original model:
[0179] Step 1. Dataset division: For the factory dataset, normal behaviors should account for a larger number, called majority class samples; while illegal behaviors should account for a smaller number, called minority class samples;
[0180] Step 2. Feature extraction: For high-dimensional images, traditional feature extraction methods are used to reduce the dimensionality, remove unnecessary complex features, and assign the same weight to all features;
[0181] Step 3. Oversampling: For the minority class samples after feature extraction, the original SMOTE method is used to expand their sample quantity because they have higher influence and higher misclassification cost.
[0182] Step 4. Traditional LightGBM algorithm classification: Apply the traditional LightGBM algorithm to classify the sample data set after oversampling.
[0183] Step 5. Determine the number of iterations: If the number of iterations is insufficient, continue iterative training to obtain a new weak classifier; if the number of iterations is reached, synthesize the classification weights and performance of all weak classifiers to obtain the final classifier;
[0184] When the original model is actually used to solve the plant safety problem, it may cause many problems. For example, the monitoring images may be blurred and have low resolution due to environmental noise; traditional feature extraction methods may be difficult to fully capture the complex image features in the plant area. Therefore, an improved model needs to be adopted to solve this problem in more detail.
[0185] Therefore, inspired by many factors, the present invention improves the original model, introduces the concept of fuzzy logic in the sample oversampling stage, and performs image enhancement, which not only improves the clarity of the image but also makes the newly generated samples closer to the feature distribution of the original samples, thereby improving the performance and generalization ability of the classification model. In the feature extraction stage, depthwise separable convolution can extract richer high-level semantic features, and through this layer-by-layer extraction, the features of the image can be expressed more effectively. At the classification algorithm level, the original LightGBM algorithm is greatly affected by key parameters such as the number of weak learners, the learning rate, and the maximum depth of each tree during training. Therefore, the butterfly optimization algorithm is introduced to optimize its key parameters; and the Focal Loss is introduced for its original loss function, which mainly adjusts the sample class weights, the weights of easy-to-classify samples, and the weights of difficult-to-classify samples by introducing parameters in the loss function to improve the classification accuracy of the model.
[0186] The steps of the improved LightGBM classification model (FuzzyFSSLightGBM) proposed by the present invention are as follows:
[0187] Input: The number of generated minority class samples n, the number of iterations T, the initial population size of butterflies n_butterflies, the optimal butterfly ratio p_ratio, the retention ratio a of large-gradient data, the sampling ratio b of small-gradient data, the maximum conflict threshold K;
[0188] Output: The final classifier
[0189] 1 Use depthwise separable convolution CNN for feature extraction;
[0190] 2 Use FuzzySMOTE to resample the minority class samples to generate n minority class samples and initialize the sample weights;
[0191] 3 Put the dataset with weights into the improved LightGBM model for training, and output the bundled feature beam set after bundling the feature vectors extracted by LightGBM's EFB exclusive feature bundling;
[0192] 4 The output feature beam set obtains the sampled data subset through LightGBM's GOSS sampling strategy;
[0193] 5 Calculate the split gain for the sampled data subset;
[0194] 6 If the iteration number T is not satisfied, continue the above steps; when the iteration number T is satisfied, all classifiers in the classifier pool are integrated according to the learning rate and the gradient of its loss function to obtain the final classifier.
[0195] Next, the technical solution will be described in detail, specifically elaborating on the implementation method of the proposed improvement plan and each key step. By gradually analyzing and elaborating on the operation process of each link, it aims to comprehensively demonstrate the working principle, innovation points and implementation methods of the plan. In particular, when explaining the improvement plan, it will deeply explore the technical breakthroughs and optimizations in different links to ensure that the design and implementation of each part can cooperate efficiently, so as to achieve the expected performance improvement. At the same time, by listing the implementation steps in detail, it is ensured that the plan can be effectively implemented in practical applications and has operability and popularization. This process not only covers the theoretical level of elaboration, but also through specific technical details and step-by-step instructions, demonstrates how the improvement plan solves problems, improves performance and promotes the development of related technical fields in practical applications.
[0196] Step 1, initialize parameter settings
[0197] In the present invention, it is necessary to set the number of generated minority class samples n, the iteration number T, the initial population size of butterflies n_butterflies, the optimal butterfly ratio p_ratio, the retention ratio a of large gradient data, the sampling ratio b of small gradient data, and the maximum conflict threshold K; the initialization parameter settings and the number of samples are as follows:
[0198] Table 1 Initialization parameter settings and number of samples table
[0199]
[0200] Step 2, depthwise separable convolution feature extraction
[0201] Introduce depthwise separable convolution to optimize the traditional convolution operation. Depthwise separable convolution decomposes the standard convolution into two steps: depthwise convolution and pointwise convolution, significantly reducing the parameter scale of the model. Specifically, traditional convolution needs to calculate simultaneously in the spatial dimension and channel dimension of the convolution kernel, while depthwise separable convolution separates these two calculation steps, making the calculation complexity of each step significantly reduced, thus significantly reducing the computational overhead. This optimization method can significantly improve the efficiency of feature extraction while maintaining the efficient feature extraction ability of the model.
[0202] More importantly, although the parameter scale and computational amount are reduced, depthwise separable convolution can still effectively capture the feature information of the input data, ensuring that the performance of the model in the classification task is not affected.
[0203] Through such improvements, the computational performance of the model has been further enhanced, while also enhancing its adaptability in different computing environments. For example, when running on resource-constrained embedded devices or mobile devices, it can also achieve efficient and stable performance output.
[0204] Figure 3 Schematic diagram of feature extraction
[0205] Step 3, membership calculation formula and sample generation formula
[0206] Calculate the fuzzy membership: For the minority class sample x i and its potential nearest neighbor sample x j . Define the fuzzy membership function as:
[0207]
[0208] where x i is the minority class sample, is the j-th minority class nearest neighbor of x i , is the j-th majority class nearest neighbor of x i , k is the number of nearest neighbors, and d(·,·) is the Euclidean distance between samples.
[0209] Membership normalization: Process it to the [0,1] interval, and the formula is as follows:
[0210]
[0211] Probability of selecting a reference sample: Divide the normalized membership of each sample by the sum of the memberships of all samples to obtain a probability distribution. That is, calculate the probability of selecting each minority class sample, and the formula is as follows:
[0212]
[0213] where, is the sum of all normalized memberships.
[0214] Generate new samples: The sample generation formula is as follows:
[0215] x new = x i + λ(x j - x i )
[0216] where x i is the reference sample selected according to the probability P(x i ), x j is randomly selected from one of the k nearest neighbors of x i , and λ is a random number between [0,1].
[0217] Step 4, FuzzyFSSLightGBM classification
[0218] Based on the original LightGBM algorithm, the Focal Loss is introduced into its original loss function, and the formula is as follows:
[0219] FL(p t ) = -α t (1 - p t ) γ log(p t )
[0220] where p t represents the predicted probability of the model for the true class, and is defined as:
[0221]
[0222] where p t is the probability of predicting the positive class, and y is the true label. α t is the balance factor, which is used to balance the importance of positive and negative samples. γ is the adjustment factor, which controls the attention degree of easy and difficult samples. When γ = 0, it degenerates into the ordinary cross-entropy loss. When γ > 0, it improves the attention to difficult-to-classify samples and suppresses the contribution of easy-to-classify samples to the loss. For key parameters such as the learning rate, the Butterfly Optimization Algorithm (BOA) is introduced for optimization.
[0223] The process of the unbalanced classification model based on the improved LightGBM is as follows:
[0224] Figure 4 Flowchart of the unbalanced classification model based on the improved LightGBM
[0225] Finally, the present invention verified the significant performance advantages of the improved LightGBM algorithm in the problem of factory area safety warning through experiments. Specifically, through comparative analysis based on three representative datasets designed in the experiments, the improved LightGBM algorithm showed obvious superiority in multiple key performance indicators such as classification accuracy, unbalanced sample processing ability, and model stability. The experimental results show that whether it is dealing with a dataset with extremely unbalanced sample category distribution or facing a scenario with a high feature dimension and complex sample features, the improved model can effectively improve the classification accuracy and robustness. Especially in terms of the recognition rate of minority class samples, compared with other common unbalanced classification models, the method proposed in the present invention shows strong adaptability and generalization performance. In addition, since LightGBM itself is a relatively lightweight model, the improved LightGBM algorithm is feasible and efficient in practical industrial scenarios. Therefore, the present invention not only outperforms traditional models in terms of comprehensive performance, but also provides an efficient and reliable solution for the factory area safety warning task. The experimental indicators of the three datasets are shown in the following table and figure.
[0226] Table 2 Index performance under the SafetyHelmetWearingDataset dataset
[0227]
[0228]
[0229] Table 3 Index performance under the Safety Helmet and Reflective Jacket dataset
[0230]
[0231] Table 4 Index performance under the ImVisible:Pedestrian Traffic Light Dataset dataset
[0232]
[0233] The present invention provides a factory area safety warning system that improves LightGBM, specifically for the safety management needs of high-risk industries such as industrial factories, chemical plants, and power plants. Traditional safety warning systems often rely on manual monitoring and simple rule systems, suffering from problems such as poor data processing ability, low recognition accuracy, and slow response speed, and are difficult to cope with the challenge of data imbalance.
[0234] Through the effective processing of imbalanced data, the present invention adopts an improved LightGBM classification algorithm, significantly enhancing the recognition ability of the safety warning system for minority-class accidents (such as not wearing a safety helmet, not wearing a warning suit). Since in the actual industrial environment, safety accidents are often few and unevenly distributed, therefore, how to effectively identify these potential safety risks from a large amount of monitoring data has become the core problem to be solved by the present invention. When dealing with imbalanced data, traditional classification algorithms tend to ignore the minority class, thus reducing the accuracy and effectiveness of safety warnings. By combining the FuzzySMOTE algorithm and the improved LightGBM algorithm, and using the FuzzySMOTE algorithm during the iteration of weak classifiers to resample the feature of minority-class samples extracted by the improved CNN, the present invention can effectively identify minority-class samples and improve the accuracy of the overall system.
[0235] In addition, the improved safety warning system can analyze in real time the information extracted from the video surveillance data stream, combine historical accident data, and predict potential risks through a machine learning model to provide accurate warnings. For industries such as industrial factory areas, the present invention has obvious advantages. Through the technical solution of the present invention, the system can identify these low-frequency, high-risk events in advance, ensuring that managers can obtain alerts in a timely manner and take corresponding measures in a timely manner, thereby avoiding the occurrence of major safety accidents.
[0236] More importantly, the factory area safety warning system of the present invention can not only effectively reduce the occurrence of accidents, but also greatly improve the intelligent level of factory area management. Through integrated real-time data analysis and intelligent decision support, the system can automatically adjust the warning strategy and dynamically optimize the classification model to adapt to the operation conditions of different factory areas. This intelligent upgrade makes safety management more efficient and reliable, and saves a large amount of accident prevention costs and emergency response costs for enterprises.
[0237] This solution designed a total of 7 groups of comparative experiments and ablation experiments to verify the effectiveness of this method. The ablation experiment was completed in a total of 4 parts, namely the original model, the model improved at the data sampling level, the model after data sampling and depthwise separable convolution (DSP), and the model after data sampling, depthwise separable convolution, and classification algorithm improvement; while the comparative experiment was completed for different classification algorithms and different optimization algorithms. First, while using the butterfly optimization algorithm to optimize key parameters and using the improved LightGBM model, a comparison was made with the relatively new classification algorithms such as DSPCNNNGBoost, DSPCNNCatBoost, DSPCNNSMOTEBoost, and DSPCNNRUSBoost; on the basis of selecting the improved LightGBM model, a comparison was made between the particle swarm optimization algorithm (PSO), grey wolf optimization algorithm (GWO), firefly algorithm (FA), and butterfly optimization algorithm (BOA). The overall technical solution has theoretical support, which can prove that the effect of this solution has been improved and is feasible. The ablation experiment of SafetyHelmetWearingDataset is as Figure 5 shown. The AUC value of the model improved using the FuzzySMOTE algorithm increased by 0.045, the AUC value of the model after combining the FuzzySMOTE algorithm and DSP increased by 0.07, and the AUC value of the LightGBM model improved by fusing the FuzzySMOTE algorithm and DSP increased by 0.084, indicating that this method has good classification performance. The comparative experiment of SafetyHelmetWearingDataset is as Figure 6 shown. Compared with DSPCNNNGBoost, DSPCNNCatBoost, DSPCNNSMOTEBoost, and DSPCNNRUSBoost, the AUC values of this method increased by 0.017, 0.007, 0.384, and 0.049 respectively, indicating that the classification effect of this method is the best. The ablation experiment of Safety Helmet andReflective Jacket is as Figure 7 shown. The AUC value of the model improved using the FuzzySMOTE algorithm is 0.849, the AUC value of the model after combining the FuzzySMOTE algorithm and DSP is 0.854, and the AUC value of the LightGBM model improved by fusing the FuzzySMOTE algorithm and DSP is 0.895. The overall improvement effect shows an upward trend, indicating that this method has good classification performance. The comparative experiment of Safety Helmet andReflective Jacket is as Figure 8As shown, compared with DSPCNNNGBoost, DSPCNNCatBoost, DSPCNNSMOTEBoost, and DSPCNNRUSBoost, the AUC values of this method are increased by 0.045, 0.01, 0.178, and 0.134 respectively, indicating that this method has significant advantages and better classification effects compared with existing methods. The ablation experiment of ImVisible:Pedestrian Traffic Light Dataset is as Figure 9 shown. The AUC value of the model improved by the FuzzySMOTE algorithm is 0.825, the AUC value of the model combined with the FuzzySMOTE algorithm and DSP is 0.873, and the AUC value of the LightGBM model improved by the fusion of the FuzzySMOTE algorithm and DSP is 0.911, indicating that this method has better classification effects. The comparative experiment of ImVisible:Pedestrian Traffic Light Dataset is as Figure 10 shown. Compared with DSPCNNNGBoost, DSPCNNCatBoost, DSPCNNSMOTEBoost, and DSPCNNRUSBoost, the AUC values of this method are increased by 0.021, 0.003, 0.033, and 0.04 respectively, indicating that this method has the best classification effect and good classification performance. For the detection, visual analysis, and camera capture of the system designed for this solution, as Figure 11 , 12 , shown in Figure 13.
[0238] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable logic devices such as field programmable gate arrays, or can be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software, such as firmware.
[0239] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be covered within the protection scope of the present invention.
Claims
1. An improved LightGBM imbalanced classification method for factory safety warning, characterized by: The following steps are involved: (1) Use deep separable convolutional neural network for feature extraction; (2) Use the fuzzy minority oversampling technique (FuzzySMOTE) to resample the minority class samples, generate a preset number of minority class samples, and initialize the sample weights; (3) performing feature bundling processing on the resampled data set, and using the exclusive feature bundling (EFB) algorithm to generate a bundled feature set, wherein the bundling process includes calculating feature conflict degree, feature sorting, feature allocation, and feature weight encoding; (4) Input the bundled feature set into the LightGBM model, use the gradient-aware subsampling strategy (GOSS) to generate large gradient and small gradient sample subsets, and construct a sampling data set based on a weighted method; (5) Calculate the splitting gain of the splitting point based on the node sample gradient and the splitting gain formula; (6) Perform multiple iterative training, integrate all classifiers according to the learning rate and loss function gradient, and output the final classifier.
2. The method according to claim 1, characterized in that The feature extraction step of the depthwise separable convolutional neural network includes: decomposing the standard convolution operation into a depthwise convolution operation and a pointwise convolution operation, wherein the depthwise convolution is performed only in the spatial dimension, and the pointwise convolution uses a 1×1 convolution kernel to complete the integration between channels.
3. The method according to claim 1, characterized in that In the mutually exclusive feature bundling algorithm, the conflict degree between features is calculated, and the features are assigned to existing bundling sets based on the criterion that the conflict degree does not exceed a preset threshold. If there is no set that meets the condition, a new bundling set is created.
4. The method according to claim 1, characterized in that: The gradient-aware subsampling strategy includes: sorting the samples in descending order according to their absolute values of gradient, retaining a preset proportion of large gradient samples, and randomly sampling a preset proportion of small gradient samples from the remaining samples, and constructing a sampling data set after assigning corresponding weights to different gradient samples.
5. The improved LightGBM imbalance classification method for factory safety warning as claimed in claim 1 is characterized in that: The classification error rate and weight are calculated as follows: Calculate fuzzy membership: For the minority class sample x i and its potential neighbor samples x j ; Define the fuzzy membership function as: where x i is a minority class sample, is x i The jth minority neighbor of is x i The jth majority class neighbor of , k is the number of neighbors, and d(·,·) is the Euclidean distance between samples; Membership normalization: process it to the [0,1] interval, the formula is as follows: Probability of selecting the benchmark sample: Divide the normalized membership of each sample by the sum of the memberships of all samples to obtain a probability distribution; that is, calculate the probability of selecting each minority class sample, the formula is as follows: in, is the sum of all normalized memberships; Generate new samples: The sample generation formula is as follows: x new =x i +λ(x j -x i ) Among them, x i According to P(x i ) The benchmark sample selected by probability, x j is x i is a random number selected from the k nearest neighbors of , and λ is a random number between [0,1].
6. The improved LightGBM imbalance classification method for factory safety warning as claimed in claim 1 is characterized in that: The categories described: Based on the original LightGBM algorithm, the Focal Loss is introduced into its original loss function. The formula is as follows: FL(p t )=-a t (1-p t ) γ log(p t ) Among them, p t Represents the model's predicted probability for the true category, defined as: Where p is the probability of predicting the positive class, y is the true label; α t is a balancing factor used to balance the importance of positive and negative samples; γ is an adjustment factor that controls the degree of attention paid to difficult and easy samples; γ = 0 degenerates into ordinary cross entropy loss, and γ > 0 increases the attention paid to difficult-to-classify samples and suppresses the contribution of easy-to-classify samples to the loss; for key parameters such as learning rate, the butterfly optimization algorithm (BOA) is introduced for optimization.
7. An improved LightGBM imbalance classification system for factory area safety warning implementing the improved LightGBM imbalance classification method for factory area safety warning as claimed in any one of claims 1 to 6, characterized in that: The improved LightGBM imbalance classification system for plant safety warning includes: Feature extraction module, used for feature extraction using depthwise separable convolutional CNN; The sampling module is used to resample the minority class samples using FuzzySMOTE, generate n minority class samples, and initialize the sample weights; The training module uses EFB and GOSS algorithms to bundle the feature vectors after feature extraction to output feature bundle sets, and collects large gradient samples and small gradient samples accordingly to form sampled data subsets, and constructs a decision tree by calculating the split gain; and then proceeds to the next cycle; In the iteration module, when the number of iterations T is not met, the above strategy is still adopted to construct gradient samples and generate decision trees. The integration module is used to integrate all classifiers in the classifier pool according to the learning rate and the gradient of its loss function to obtain the final classifier when the number of iterations T is met.
8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the improved LightGBM imbalance classification method for factory safety warning as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to execute the steps of the improved LightGBM imbalance classification method for factory safety warning as described in any one of claims 1 to 6.
10. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the improved LightGBM imbalance classification system for factory safety warning as described in claim 7.
Citation Information
Patent Citations
Bearing fault diagnosis method, system, device and terminal
CN113834656A
Intrusion detection method based on convolutional neural network and lightweight gradient elevator
CN113901448A
5G network fault prediction method for smart fishery
CN114676642A
Real-time early warning method and system for class imbalance well leakage based on deep learning
CN117786318A
Emergency fire-fighting fire risk assessment and early warning method driven by AI large model
CN118095865A